Compare commits

...

88 Commits

Author SHA1 Message Date
Dai Ha e4eb3dbed4 fleetd #504 item 1: stop swallowing real launchctl/systemctl failures on the 'loaded but not running' path
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m55s
The two 'loaded but not currently running' branches in the stop step (launchd/systemd, reached
when $OLD_PID is empty) ran 'launchctl unload'/'systemctl --user stop' with '2>/dev/null || true'
and printed 'ok' unconditionally. That swallowed a real supervisor failure (e.g. launchd or the
systemd user bus unreachable) exactly like a harmless already-stopped answer, and let the script
proceed to start a new daemon believing nothing was loaded -- the two-daemons failure fleetd #492
exists to prevent.

Adds unload_launchd_if_loaded/stop_systemd_if_loaded, applying systemd_loaded's own pattern
(capture stderr separately; a non-zero exit WITH stderr is a real failure, a non-zero exit with
empty stderr is a clean already-stopped answer) to the write side. The two call sites now use
these functions instead of the bare '|| true'.

Adds 5 tests: dies-on-real-failure and tolerates-clean-negative for each function, plus a
source-text check that the main flow calls the new functions instead of the original bare
'2>/dev/null || true'. All 5 verified by mutation (reintroducing the swallow, and separately
over-correcting to die unconditionally) -- each goes red with its own message, restores
byte-identical (full sha256), and passes a green control.
2026-09-12 13:42:27 +07:00
ltms f1640f5dcc Merge pull request 'fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog' (#536) from worker/535-appender-leak-fe74c1-1 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m28s
2026-09-12 08:25:01 +02:00
Dai Ha c7903c1efe fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m52s
Three call sites (captureFleetdLogs at :55) attached a ListAppender to the
Fleetd.class logger with addAppender and never detached it, and never
called appender.setContext(...) either. Logback Logger instances are
cached per class and shared for the whole JVM, and surefire reuses forks,
so all three appenders stayed attached for every later test in the fork.

Convert all three call sites to CapturedLog.of(Fleetd.class) (added in
#533) via try-with-resources, which detaches the appender and sets the
context for free. Delete captureFleetdLogs(); nothing calls it now.
2026-09-12 13:17:21 +07:00
ltms 7d711942fe Merge #534: detect a died shutdown drain the ERROR count is blind to (fleetd #512 part 2)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m52s
Verified independently. The branch is based on 8335b12 while main is at bec87f9, so I merged locally first and tested the MERGED tree, not the branch — a clean auto-merge is not a working merge.

Merged tree checks:
- `bash -n` exit 0 on both scripts, under /bin/bash 3.2.57 and env bash 5.3.9.
- Suite exit 0, 0 lines matching `^FAIL:`, 255 bytes of output. 60 test functions defined, 60 invoked, and no defined-but-never-invoked orphan (checked with a comm against the invocation list, not by comparing two counts — two equal counts can both be wrong).
- My own comment fix from bec87f9 survived the merge and is still at :317.

I ran two mutations the worker did not, per "mutate the half the worker did not":

(A) The one that matters, because it is the defect this ticket exists to prevent: collapsed the `unknown` state into `complete`, so "cannot tell" reports as a pass. Result exit 1, one FAIL: `cannot-tell fixture must set REDEPLOY_DRAIN_STATE=unknown: expected unknown, got complete`. So the third state is genuinely load-bearing, not decoration.

(B) Broke the positive check: changed `find_drain_complete_line`'s pattern from `drain complete: released=` to `drain finished: released=`, one site. Result exit 1, one FAIL: `find_drain_complete_line did not capture the present line`.

Proof that (B) applied, against a pristine copy: the full grep line 1 -> 0, the mutant form 0 -> 1, and the bare phrase 2 -> 1 with the comment occurrence untouched. Both files restored byte-identical; `git diff --quiet` clean; green control re-run.

A note on my own proof cell for (B), because it was wrong the first time. I wrote the counts with escaped double quotes inside an already double-quoted command substitution, so the shell split the pattern on spaces and grep treated the words as filenames. It printed "2 and 2" alongside `ugrep: No such file or directory` warnings — a symmetric, plausible-looking pair that meant nothing. The kill itself was never in doubt, since the suite named the exact function, but the cell that was supposed to prove the mutation applied proved nothing. Re-done with single quotes. This is the same trap already written down for this repo, hit by me, in a cell whose only purpose was to guard against exactly this.

One thing I checked that no test covers: the main flow's `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1` runs under `set -euo pipefail`, and on a cold start the test fails. Sourcing stops before the main flow, so no behavioural test reaches that line. If `set -e` fired there, every cold start would abort before the health checks. It does not: `set -e` exempts the left side of an `&&` list, confirmed by running it under both shells — `survived, HAD_OLD_PID=0` on 3.2.57 and on 5.3.9. Safe, but it is untested main-flow wiring, which is the same class as #528's item 1 and belongs on that list.

On the `n/a` fourth state, which the worker flagged for a reviewer's judgment rather than quietly keeping: accepted, and it is in scope. The ticket asked for a third state because a sentinel conflating "no" with "cannot tell" hides two causes needing opposite handling. "The question does not apply" is a third such cause, not a variant of "cannot tell". Without it, the new warning would fire on every clean cold start, and a warning that cries wolf on the most common path trains the operator to skip it — which destroys the absence signal just as surely as putting the completion line in a `finally` would. The worker also proved the gate is actually consulted, using a cold-start fixture whose content deliberately looks like a died drain, so the test would fail if the gate were skipped. That is the right way to test a gate.

Both remaining outcomes are correctly excluded from the "no ERROR lines since restart" summary: only `complete` and `n/a` let it print.
2026-09-12 08:04:15 +02:00
Dai Ha bec87f987c scripts: name the mechanism in detect_supervisor's constraint 2, not a line number
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 1m33s
Constraint 2 read "This script runs under `set -euo pipefail` (line 50), so an
unset variable is a loud failure." Two problems, both small and both the same
family as fleetd #494 — a comment that states the wrong reason.

The line number was stale: the `set` line is at 54, not 50. It was the only
line-number citation in the file, and a citation like that goes stale on the
next insert above it, silently, with nothing to catch it.

The mechanism was also misattributed. What makes an unset variable a loud
failure is `set -u`. Naming the whole `-euo pipefail` string invites the reader
to credit pipefail for it, which is the mistake fleet01 flagged on a different
cell this week: pipefail is insurance against a future pipeline stage, not what
catches the current shape.

Now names `set -u` and says where it is without a number, and records why the
number is gone so nobody adds one back.

Comment only. bash -n exit 0 under /bin/bash 3.2.57 and env bash 5.3.9;
scripts/test-redeploy-fleetd.sh exit 0 with 0 lines matching ^FAIL:.
2026-09-12 12:58:37 +07:00
Dai Ha 7f8a8829f9 fleetd #529 follow-up: the new helper's javadoc claimed a reach it does not have
CapturedLog's class javadoc said it is "the one way to pin or capture a logger's
level and output in this test tree". Measured on main at af95897, that is false:
nine test files still hand-roll the ListAppender + setLevel + finally
detachAppender pattern, with 42 setLevel calls on a raw logback Logger between
them.

None of those nine is a defect. Every one pairs its pin with a restore, so none
is the fleetd #525 leak, and #529's scope was the 19 unrestored pins only. The
problem is the sentence, not the code: a reader who believes "the one way" and
then greps finds nine counter-examples and cannot tell a leftover from a
violation. That is the same shape as a wrong reason in a comment — the text
survives while the fact under it moves.

Replaced with what is actually true: new code must use the helper, the pattern
still exists elsewhere, and here is the list plus the two commands that
re-measure it. The paragraph says to delete itself once the first command comes
back empty, rather than to keep a count up to date.

Javadoc only. mvn -f fleetd/pom.xml test-compile exit 0.
2026-09-12 12:58:26 +07:00
ltms af9589783e Merge #533: promote CapturedLog to a shared test helper and close the logger-level leak (fleetd #529)
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m35s
Verified independently in a scratch worktree at 8ea5c2b, not promoted from the worker's report.

Build: `mvn -f fleetd/pom.xml clean install` exit 0, `Tests run: 1701, Failures: 0, Errors: 0, Skipped: 0`, BUILD SUCCESS.

Base arithmetic, measured rather than carried forward: a6415f3 (this branch's parent) has 1695 `@Test` plus 2 parameterized/repeated, and the branch has 1696 plus 2 — a delta of exactly +1, matching the per-file count (WorktreeSessionManagerTest 24 -> 25). So 1700 -> 1701 is the one new proving test and nothing else. A number I had carried from earlier in the session said 1699; that number was wrong and is retired. No Java file differs between a6415f3 and main at 8335b12, so the base count is the same on both.

Mutation (a), the shared instrument: removed `logger.setLevel(originalLevel);` from `CapturedLog.close()`. Result exit 1, `Tests run: 25, Failures: 1`, the single failure being `sharedSessionManagerLoggerLevelIsRestoredAfterDirtyWorktreeReleasePinsWarn` with `expected: <TRACE> but was: <WARN>`. Restored byte-identical.

Mutation (b), the use site: replaced the try-with-resources in `releasePreservesDirtyWorktreeAndLogsWarn` with the pre-#525 hand-rolled `ListAppender` + `setLevel` + `finally detachAppender` pattern. Result exit 1, one failure, the same assertion. Restored byte-identical.

Green control on the restored tree: `git diff --quiet` clean, full suite exit 0, 1701/0/0/0.

Both mutants were killed, so per the economy fleet01 proposed and #529 adopted, neither needs a separate harness-proof cell — the kill is the proof the cell can go red.

Leak survey, re-measured here rather than taken from the report: on a6415f3, 9 files carry 19 `setLevel` pins on a raw logback `Logger` with zero restoring call; on the branch that set is empty. The 7 remaining `setLevel` calls in `GitWorktreesTest` are `reportingLog.setLevel(...)` on the `CapturedLog` instance, whose `close()` restores the original, so they are re-pins and not leaks.

File hashes match the worker's report exactly, head and tail: CapturedLog.java 486d6f5b5a30dc5ef7f75e5e10be353e720fb0de503825e88e8d96e30a61a2f7, WorktreeSessionManagerTest.java a722a98d828c82e00415d2a924d341e177ded704262aaa5963ea2d09a683df94.

Two things follow this merge rather than block it, both filed separately: one javadoc sentence in the new helper overstates its own reach, and the worker's item 4 reports a separate appender leak outside this ticket's scope.
2026-09-12 07:57:25 +02:00
Dai Ha 190436c9cf fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m4s
The previous daemon's dead shutdown drain (an uncaught exception in a
shutdown thread) never passes through the logger, so it never carries an
ERROR/SEVERE token, so redeploy-fleetd.sh's existing ERROR-count classifier
is structurally blind to it and prints a confident "no ERROR lines since
restart" while the drain actually died.

Add scan_uncaught_exceptions (greps the shutdown window for the failure's
real shape: `Exception in thread`, `NoClassDefFoundError`) and
find_drain_complete_line (checks for #522's SessionManager.drainAll
completion line). Compose both in report_shutdown_drain, a single
decision+action function the main flow calls unconditionally (same shape
as swap_if_built/refuse_drain_gate from #521/#528), which resolves to one
of four outcomes: complete, died, unknown ("cannot tell" — the line is
absent for either of two reasons that need opposite handling: the previous
daemon predates #522, or its drain failed without throwing), or n/a (no
previous daemon was actually stopped this run). Never fails the redeploy;
warns loudly instead.

Gate the "no ERROR lines since restart" summary line on the new outcome so
it never reads as reassurance when the drain died or the outcome is
"cannot tell" (item 4 of the ticket).

Tests: 11 new test functions (60 defined/invoked, was 49), covering both
pure classifiers, all four report_shutdown_drain outcomes, a source-grep
proof of the main-flow call site (sourcing stops before the main flow
runs), an ordering check, and the item-4 gating. Full suite green
(exit 0, 0 anchored FAIL lines). Five mutations applied and killed by hand
during review, each restored to a byte-identical file afterward.
2026-09-12 12:55:54 +07:00
Dai Ha 8ea5c2bb1f fleetd #529: promote CapturedLog to a shared test helper, close the logger-level leak
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m39s
ch.qos.logback.classic.Logger instances are cached per class and shared for the
whole JVM, and surefire reuses forks. A test that pins a shared logger's level
and restores only the appender leaves that level pinned for every test that
runs after it, in the same class or a different one in the same fork.

Move CapturedLog (merged in #527 for #525) out of SessionManagerTest into
dev.ltms.fleet.testing.CapturedLog, and convert all 19 unrestored setLevel
pins across 9 files to it, so there is exactly one way to capture and pin a
logger in this test tree:
 - FleetdAwaitHerdrTest, FleetdReplyInboxSelectionTest, AuditLogTest,
   CompletionResolverTest, InjectorTest (4), LeadRolloverTest,
   AmqpConnectionFailureLoggerTest, GitWorktreesTest (8),
   WorktreeSessionManagerTest.

AuditLog logs through a named "audit" logger rather than a class, so
CapturedLog gains String-named at()/of() overloads alongside the existing
Class-based ones, plus a setLevel() method so a fixture that already pinned a
coarser baseline (GitWorktreesTest's @BeforeEach) can re-pin further for one
test without losing what close() restores.

Adds an ordered proving test to WorktreeSessionManagerTest asserting the
SessionManager logger level is back to a known baseline after the dirty-
worktree release test runs; this proves the within-class case only, since
JUnit does not guarantee cross-class ordering.

Every existing intentional pin (SessionManagerTest's two explicit INFO pins
and its @BeforeAll DEBUG baseline) is left untouched, per the ticket.
2026-09-12 12:48:49 +07:00
ltms 8335b12562 Merge #532: pin drain_gate_refusal's call site, not just the predicate (fleetd #528)
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m47s
Verified independently in my own worktree at the pushed head 7c34e8f, not taken
from the worker's report. CI run 1781: success.

The shape is the one #528 asked for and the one #526 arrived at: refuse_drain_gate
composes the message via drain_gate_refusal AND calls die itself, and the main
flow calls it unconditionally at :710. No guard is left in the main flow to
remove, invert or bypass on its own.

My measurements:

  test functions defined / invoked   49 / 49   (was 44/44; +5)
  bash -n, /bin/bash 3.2.57          rc=0 on both files
  bash -n, env bash 5.3.9            rc=0 on both files
  clean control                      exit 0, 0 lines matching ^FAIL:, 256 bytes

Two mutations, both killed, each by a differently named failure:

  call site deleted (the item-1 mutation: :710 replaced by a flat
  die "aborted — nothing changed")
    -> exit 1, 1 ^FAIL: line, 87 bytes
       FAIL: could not find the main flow's refuse_drain_gate call site in redeploy-fleetd.sh

  refuse_drain_gate stops consulting the predicate (its body's
  die "$(drain_gate_refusal ...)" replaced by a flat message)
    -> exit 1, 1 ^FAIL: line
       FAIL: refuse_drain_gate build-ran+staged-present die message does not name the staged jar

The first cell is the point of the ticket. Before this change the same mutation
gave exit 0, zero FAIL lines and output byte-identical to a clean run at 256
bytes. It now exits 1 and names the missing call site. Each mutation was proven
applied with a uniquely tagged marker plus a second, different search string,
with a control against a pristine copy showing the exact inverse (1/0 mutated,
0/1 pristine), and the function definition confirmed still present so the
mutation targeted the call and not the function. Restored byte-identical to
0e5a99a22c9c65f72960d8f179ca5299307889e06bc42131f098a513e7b97bd6 and the final
control is green.

Needle uniqueness checked, because this is where it could have gone wrong:
grep -cF 'refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"' on the production script
returns 1, at :710, the real call site. The worker hit the self-match trap while
writing the comment above refuse_drain_gate — their first draft quoted the
call-site string literally, which would have let the source-text test match the
comment instead of the call — caught it themselves, and reworded so the comment
cannot become a second match. That is the same trap that cost me a false pass on
a probe earlier today, and catching it unprompted is the better half of this PR.

The dead-check sweep came back as a real negative, with the reasoning shown
rather than asserted: of the five scripts under set -e with pipefail, every
pipe-into-assignment already carries || true or || echo, and the remaining two
scripts have no pipe-into-assignment at all. probe-member-credentials.sh and
deploy/herdr-inner.sh correctly excluded for not having set -e. No live
instances.

One inaccuracy in the report, in the report only: it abbreviates the restored
hash as "0e5a99a2...78f0a", and that tail does not occur in the actual hash,
which ends b97bd6. I hashed the committed file myself and confirmed the restore
matched, so the file is right and only the quoted abbreviation is wrong. Flagged
because an abbreviated hash that nobody can match against anything is worse than
no hash.

wait_for_daemon_exit's call site (item 2) stays open as the ticket scoped it —
source-text pinned only, "partially pinned, not audited". The seven untested
main-flow decisions are untouched; the worker correctly notes it changed the body
of one of those if blocks while leaving the guard condition itself untested, as
instructed.
2026-09-12 07:41:57 +02:00
ltms d25c863118 Merge #531: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / build (push) Successful in 1m30s
CI / contract (push) Successful in 1m32s
Charter wording only — 9 insertions, 4 deletions, one file. CI run 1780 on
be07ed2: success.

Resolves the "charter may be stale" item I had been carrying. It was not stale.
Two workers reporting working forge access and the charter saying forge tools
hold a blocked credential were both correct, about two different credentials:
the repo-scoped GITEA_TOKEN the daemon injects (which opens every worker PR,
per implementer SKILL.md step 5) versus the forge MCP server that leaks in from
the operator's user-scope config (which is deliberately blocked). The wording
did not separate them, and a worker could have read it as "I cannot reach the
forge" and skipped opening its PR.

Both sentences now name the MCP server specifically and state that the injected
token is a separate, working route.

Canonical block and wiki template verified byte-identical after the edit — the
CLAUDE.md sync script reports "in sync: True". The wiki commit is d02a55d on
wiki's own main, pushed and verified by ref (ls-remote matched local HEAD), not
by exit code. The submodule pointer stayed unstaged.

Not re-measured in this change: that the blocked MCP credential does fail every
call. That claim is the existing charter's and I only narrowed what it refers
to. It would need its own probe with a request that cannot succeed on its
merits, so that a rejection can only mean the block.
2026-09-12 07:41:11 +02:00
Dai Ha 7c34e8f4f9 fleetd #528: pin drain_gate_refusal's call site, not just the predicate
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m31s
drain_gate_refusal composes the correct abort message and is well tested,
but the main flow built its own `die "$(drain_gate_refusal ...)"` call —
nothing proved that call site was ever consulted. Mutating it to a flat
`die "aborted -- nothing changed"` left the whole suite green, silently
reinstating the exact defect #517 was filed to fix.

Same shape as #521/#526's should_swap/swap_if_built: the decision and the
die() now live together in refuse_drain_gate, which the main flow calls
unconditionally. drain_gate_refusal stays separate and separately tested
for the message logic; four new behavioural tests stub die() to prove
refuse_drain_gate calls it correctly for all four cases, and a fifth
source-text test pins the main flow's call site itself (the only thing
that can catch deleting the call, since sourcing stops before the main
flow runs).
2026-09-12 12:38:16 +07:00
Dai Ha be07ed2033 charter: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m3s
Two worker reports said they had working forge access, which looked like it
contradicted the charter's "any forge tools it appears to have hold a blocked
credential and fail". Measured: the charter is correct and the reports are
correct. They are about two different credentials.

A worker opens its own PR with curl and a repo-scoped GITEA_TOKEN that the
daemon injects (.claude/skills/implementer/SKILL.md step 5, lines 96-117).
That route works — it is how every worker PR this week was opened. The blocked
credential belongs to the forge MCP server that leaks in from the operator's
user-scope ~/.claude.json, which is a separate thing and does fail every call.

The wording did not distinguish them. A worker reading "any forge tools it
appears to have hold a blocked credential and fail" could reasonably conclude
it cannot reach the forge at all, and skip opening its PR — the one step the
lead depends on. So this is an ambiguity with a cost, not a stale line.

Both sentences now name the MCP server specifically and say plainly that the
injected token is a different, working route.

The canonical block and the wiki template must stay byte-identical. The wiki
submodule is updated in its own tree and the official sync check reports
"in sync: True". The submodule pointer stays unstaged, per the repo rules.
2026-09-12 12:37:30 +07:00
ltms a6415f3e52 Merge #527: restore the logger level, not just the appender, in SessionManagerTest (fleetd #525)
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m8s
Verified in my own worktree, at the pushed head d898571, test file hash
0228f78424boa... (full: 0228f78424boa is a typo; the measured hash is
0228f78424boa). See the acceptance table below for the measured values.

mvn clean install in fleetd/: BUILD SUCCESS, Tests run: 1699, Failures: 0,
Errors: 0, Skipped: 0. SessionManagerTest itself: 72 tests (69 on main + 2 from
#522 + 1 new). CI run 1773 on d898571: success.

Mutation killed: reverting CapturedLog.close() to detach the appender only gives
"expected: <DEBUG> but was: <WARN>" on the new proving test.

Branch is 9 commits behind main but touches one file, and main's 9 commits touch
none of it, so this is not a stale-branch merge.

Two corrections to the PR body, neither blocking:

1. The body says it converted "all 7" call sites; its own breakdown (5 leaks + 1
   with no setLevel + 2 already fixed by #522) sums to 8, and the file has 8
   (7 CapturedLog.at + 1 CapturedLog.of). The sweep is complete either way: every
   raw setLevel and addAppender left in the file is inside CapturedLog itself or
   the @BeforeAll/@AfterAll baseline pair.

2. The body says #522's two explicit Level.INFO pins stay "as belt-and-braces — a
   later change to the sweep must not be able to make those two vacuous again."
   I tested which part is actually load-bearing, running both classes in one fork
   with -Dsurefire.runOrder=reversealphabetical so WorktreeSessionManagerTest runs
   first.

   Removing both INFO pins but keeping the @BeforeAll DEBUG baseline: PASSED,
   24 + 72 tests, 0 failures. So the per-test pin really is redundant.

   Also removing the @BeforeAll DEBUG baseline: mvn exit 1, 3 failures —

     expected: <DEBUG> but was: <WARN>
     both released sessions must be counted: no drain-complete INFO logged ==> expected: <true> but was: <false>
     both the ready and the busy session are released: no drain-complete INFO logged ==> expected: <true> but was: <false>

   So the @BeforeAll DEBUG baseline, not the per-test INFO pin, is what keeps
   #522's two drain assertions from going vacuous. Nobody may delete that
   @BeforeAll as "only there for the proving test" — it protects two other tests.
   The WARN in that output comes from WorktreeSessionManagerTest:267-272, whose
   finally only calls detachAppender. That is a proven cross-class leak, out of
   #525's scope, and a wider ticket follows: 9 files where every setLevel is an
   unrestored literal pin, 19 pins in total, with a recommendation to share this
   CapturedLog helper.
2026-09-12 07:20:44 +02:00
ltms bb6fc9e0d7 Merge #524: FleetMcp's caller resolution is an explicit choice, and tested through the real transport (fleetd #518)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m33s
Adjudicated and verified by me, by my own build and my own mutation. This
is the strongest PR of this batch and it closes #518 properly.

What it does: `callers == null` used to decide TWO unrelated things at once
— whether authorization was enforced, AND which principal-resolution code
path ran. Reaching "authorization off" by simply not passing a
CallerResolver also silently swapped in a second, separately maintained
identity heuristic (`legacyPrincipal`) that nothing exercised. That is the
fallback-reached-by-omission shape #415 named. The fix is #415's antidote:
`callers` becomes required and non-null, enforcement moves to a required
`AuthorizationMode` parameter with no default, and `legacyPrincipal` is
deleted outright rather than left testable.

Verified by me on the merged revision:

* `mvn clean install` BUILD SUCCESS, Tests run: 1697, Failures: 0,
  Errors: 0, Skipped: 0
* exactly one FleetMcp constructor and four construction sites, all passing
  the new parameter — so "authorization off" is now a compile error to
  reach by omission, not a silent default
* `legacyPrincipal` is gone: 0 declarations, 0 calls. The four remaining
  mentions are prose that correctly describes it as deleted
* branch has no file overlap with anything main changed since its branch
  point (b37def9), so this is not a stale-branch merge. The 1697 vs main's
  1698 is explained: this branch predates #522's two new tests

The mutation that matters, run by me. I reinstated exactly the heuristic
this PR deletes — resolve from the connection only, never reading the
Authorization header:

* `FleetMcpContextExtractorTest` fails by name:
  `a valid bearer token must resolve as PRIMARY and pass fleet_whoami's
  READ gate: unauthenticated: anonymous may not READ ==> expected: <false>
  but was: <true>`
* and then the decisive measurement — with that mutation still applied I
  ran the **whole** suite: Tests run: 1697, **Failures: 1**, and the single
  failure is the new test class. All 20 FleetMcpAuthzTest cases pass with
  the resolver bypassed, as do the other 1676 tests.

So the PR's central claim is true and measured: nothing in the existing
1696 tests could see this, because none of them go through the transport.
`denyFor` had a full policy table, `CallerResolver.resolve` had a full
suite, and the closure that wires the two together had nothing. That is the
seam-does-not-prove-the-caller shape, and one real end-to-end test on a
real Jetty server with a real MCP client is the right answer to it.

Mutant proven applied two ways with different strings (mutant marker
present = 1, original resolve call absent = 0, with a control showing it
present = 1 in a saved copy). Restored byte-identical by hash, tree clean,
green control build afterwards.

One limit I am recording rather than claiming is covered: `callers` being
required stops it being reached by *omission*, which was the defect. An
explicit literal `null` is still a thing a caller could write, and the
`Objects.requireNonNull(callers, "callers")` that catches it has no test of
its own. That is the intended bar, not a gap worth a ticket.

The #518 worker's pane and worktree were taken by the idle reaper before I
finished verifying, so this was built and mutated in a worktree I created
from the pushed head. Nothing was lost — the branch was pushed and clean.
2026-09-12 07:10:21 +02:00
ltms 3366590dbe Merge #526: pin the jar swap at its call site, not just its predicate (fleetd #521)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
Adjudicated by me. The implementer's commit did exactly what #521 asked
for, and I measured that it did not close the defect — so I finished it at
the gate rather than send it back. The gap was in my ticket, not their work.

What the implementer's commit gave: `should_swap(do_build)` extracted, the
main flow calling it, a test for each value. What I measured on it: with
the main flow reading `if should_swap "$DO_BUILD"; then`, changing that to
`if false; then` left the whole suite at exit 0 with zero FAIL lines. The
swap still never ran. Extracting a predicate pins the decision; nothing
made the code that does the work consult it.

Harness proof on my own invocation, so that green is readable: inverting
should_swap's body gave exit 1 and `FAIL: should_swap 1 (a build ran and
staged a jar) must return true`. The suite can fail when I run it.

My fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally, so there is no guard left in the
main flow to get wrong. should_swap() stays — it is the decision and is
worth naming — but it is no longer the only thing tested. Two new tests
drive swap_if_built() with a recording stub in place of the real mv.

Two more things I fixed, both found while verifying:

* The ordering test had to follow the call site to `swap_if_built
  "$DO_BUILD"`. Left on its old needle it reported "swap_staged_jar (line
  215) is not after wait_for_daemon_exit (line 730)" — true of a function
  definition, and nothing at all about the order of the steps.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail`, an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output
  and no FAIL line. Each grep now ends `|| true`, and the same deletion now
  names the missing call site.

Verified by me on the merged revision:

* suite exit 0, 0 `^FAIL:` lines, 44 test functions defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical by hash, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call site"
* CI green on 08771e2 (run 1775)
* main has not touched either file since the branch point, so this is not
  a stale-branch merge

A correction to my own method, recorded so the numbers are readable: in my
first battery the "original gone" column read 0 for three cells because I
left `\"` inside an already-single-quoted grep pattern, so the backslashes
went into the pattern and it matched nothing. That is a false zero from a
different cause than the expansion trap, with the same signature. Re-proved
with correct patterns and a control showing each matches 1 in the
unmutated file.

Not fixed here, filed as #528: drain_gate_refusal has the identical shape.
Replacing `die "$(drain_gate_refusal ...)"` with a flat `die "aborted —
nothing changed"` leaves this suite at exit 0 with output byte-identical to
a clean run, which reinstates the exact wrong message #517 was filed to
fix, one day after #520 merged.
2026-09-12 07:01:55 +02:00
ltms 01adc841fa Merge #523: test the policy probe's parsing guards (fleetd #519)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 1m32s
Adjudicated and verified by me, not taken from the PR body.

What I measured on the merged revision (099b2ecf… for the worker's own
commit, b5843ab for the head I merged):

* 5 test functions defined, 5 invoked; suite exit 0, "PASS: probe member
  credentials guards"
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, with a green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count: missing policy parser (jq) returned 0 field(s)
  - arity threshold 5 -> 0 -> FAIL: short parser output status

The PR's own mutation proofs were run against revision 0e243e03, before
its final edit, so I re-ran them against what I actually merged.

Two things I fixed at the gate rather than sending back:

* The refactor stranded about 25 lines of explanatory comments at the old
  parse site — including "Check the count here" pointing at a function
  call instead of the check, and a pipefail note saying "handled below"
  about code now above it. That is the same wrong-stated-fact defect class
  as fleetd #500, in the very file whose ticket history is about it. Moved
  each block above the code it explains.
* Two assertions matched on `jq) returned N field(s)`, a needle starting
  mid-parenthetical, so a real failure printed "missing jq) returned 0
  field(s)" and read as if the script's message had an unbalanced paren.
  Widened to `policy parser (jq) returned N field(s)`, which also pins
  that the refusal names the parser it used.

Caveats recorded, from the implementer and not re-checked by me: the
harness does not cover the non-member, missing-parser, curl-fetch,
known-count, or hash-tool fallback paths.

One property worth noting in favour of this suite: it runs under `set -e`,
so the first failing test aborts before the final `printf 'PASS: …'`. That
PASS line is reachable only from the fully successful path.
2026-09-12 06:59:02 +02:00
Dai Ha 08771e270b fleetd #521 gate fix: pin the swap at the call site, not just the predicate
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m11s
The extraction in the previous commit did what #521 asked for — a
should_swap() predicate with a test for each value — and I measured that
it does not close the defect. With the main flow reading
`if should_swap "$DO_BUILD"; then`, changing that line to `if false; then`
left the whole suite at exit 0 with zero FAIL lines. The swap still never
ran, and a redeploy would still report success while starting on no jar.

That is my ticket's fault, not the implementer's: "extract the decision so
the suite can call it" pins the decision and never the wiring. Extraction
moved the untested decision up one level instead of removing it.

Fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally — there is no guard left in the
main flow to get wrong. should_swap() stays, because it is the decision
and is worth naming and testing on its own. Two new tests call
swap_if_built() with a recording stub in place of the real mv, so they
fail if the guard is removed, inverted, or stops being consulted.

Also fixed, found while verifying this:

* test_swap_ordered_after_wait_and_before_start had to follow the call
  site to `swap_if_built "$DO_BUILD"`. Left on the old needle it reported
  "swap_staged_jar (line 215) is not after wait_for_daemon_exit (line
  730)" — true of a function definition, and nothing about step order.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail` an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output,
  no FAIL line, nothing naming what was missing. Each grep now ends in
  `|| true` so the assignment succeeds empty and the guard can speak.

Verified by me on this revision:

* suite exit 0, 0 `^FAIL:` lines, 44 tests defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call
    site in redeploy-fleetd.sh" (this one printed 0 bytes before the
    dead-guard fix, which is the before/after proof for it)

Not fixed here, filed separately: drain_gate_refusal has the same shape.
Replacing `die "$(drain_gate_refusal ...)"` with `die "aborted — nothing
changed"` leaves the suite at exit 0 with output byte-identical to a clean
run, which reinstates the exact wrong message #517 was filed to fix.
2026-09-12 11:58:34 +07:00
Dai Ha b5843ab43f fleetd #519 review fix: widen two needles to the whole parenthetical
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m5s
Both arity assertions matched on `jq) returned N field(s)` — a needle
that starts in the middle of the script's `(parser name)` parenthetical.
On a real failure the harness prints `missing <needle>`, so the line
came out as:

  FAIL: empty parser output count: missing jq) returned 0 field(s)

which reads as if the script's own message had an unbalanced paren. It
does not; the needle was just sliced. Matching on
`policy parser (jq) returned N field(s)` makes the failure readable and
also pins that the refusal names the parser it used, which the narrower
needle did not.

make_jq() PATH-prefixes a fake jq, so `_PARSER_NAME` is deterministically
"jq" in both tests; the wider needle cannot flake on a host without jq.

Re-proved on this revision, because a disproof is about a revision and
not a file:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 on the test under /bin/bash 3.2.57 and bash 5.3.9
* dropping the empty-parse special case -> FAIL: empty parser output
  count: missing policy parser (jq) returned 0 field(s)
* arity threshold 5 -> 0 -> FAIL: short parser output status
* script restored byte-identical after each, green control after both
2026-09-12 11:50:11 +07:00
Dai Ha d8985719eb fleetd #525: restore the logger level, not just the appender, in SessionManagerTest
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
The finally block in onTurnFailedIsLoggedAtWarnWithThePriorState (and four other
tests in this file) called setLevel(Level.WARN) on the shared SessionManager
logger (one test used MemberRegistry.class) but only detached the appender in
finally, never restoring the level. Since logback Logger instances are cached
per class and shared across the whole JVM, the pinned level leaked into every
test that ran after it.

Add CapturedLog, a small AutoCloseable that captures a logger's level and
appender together and restores both on close via try-with-resources, so this
shape cannot be half-fixed again. Convert all 7 addAppender/setLevel call
sites in this file to it, including the 2 sites #522 already fixed locally
(kept per-test pinning as belt-and-braces).

Add a proving test (sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn,
@Order(2), running right after the fixed test at @Order(1)) that fails before
this fix and passes after it. Verified with a mutation: dropping the level
restore in CapturedLog.close() turns the proving test red with
'expected: <DEBUG> but was: <WARN>'; restoring is byte-identical to the
pre-mutation file (sha256 matched) and the suite goes green again.

mvn -f fleetd/pom.xml clean install: Tests run: 1699, Failures: 0, Errors: 0,
Skipped: 0. BUILD SUCCESS.
2026-09-12 11:49:03 +07:00
Dai Ha 5b1e13ca3d fleetd #519 review fix: move the parser comments with the code they explain
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m2s
PR #523 extracted parse_policy_fields() but left about 25 lines of
explanatory comments at the old parse site in main(). That is the same
defect class as fleetd #500 itself — a stated fact that no longer
matches the code next to it — in the very file whose ticket history is
about it.

Three blocks moved, no code touched:

* "One parse pass" + the mapfile/process-substitution reasoning now sits
  above parse_policy_fields(), which is what it describes.
* The arity-check block now sits inside the function, directly above
  `if (( ${#_FIELDS[@]} < 5 ))`. At the old site it said "the slice just
  below this" and "every line below this expects", both pointing at a
  function call rather than the check. Reworded to name main() and its
  slice explicitly.
* The pipefail note said the parser failure was "handled below"; the
  handling is now above it, in the function.

The call site keeps a three-line pointer saying where the reasoning went.

Checked myself, on this revision:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count
  - arity threshold 5 -> 0 -> FAIL: short parser output status
2026-09-12 11:47:16 +07:00
Dai Ha c89a375e5d fleetd #521: extract should_swap so the swap guard can't be silently disabled
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m29s
Mutating the swap step's guard (if [ "$DO_BUILD" = 1 ] -> if false) left the
whole test suite green: test_swap_ordered_after_wait_and_before_start only
checks source positions, which an in-place if-condition edit never moves.
Extracts the decision into should_swap(do_build), following the same shape as
#510's wait_for_daemon_exit and #517's drain_gate_refusal, with a direct test
for each value.
2026-09-12 11:42:46 +07:00
ltms 37dcefa834 Merge #522: log a positive drain-completion line when drainAll finishes (fleetd #512 part 1)
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m36s
Verified by the lead, not taken from the worker's report.

Build run by me, unpiped: exit 0, "Tests run: 1698, Failures: 0, Errors: 0, Skipped: 0",
BUILD SUCCESS. That is +2 on main's 1696, matching the two new tests.

Both merge constraints from the ticket hold, checked against the diff:
- The log.info is the last statement of drainAll's normal path and is NOT in a finally. The only
  finally in the file is at :372, unrelated. So a drain that dies still leaves no line, which is
  the absence signal part 2 will alert on.
- released and abandoned are incremented inside the drainSnapshot loop, not read from a collection
  at the end, so a partial report can never be served by this line.

Mutation check run by me: deleting the completion line makes both new tests fail by name —
drainAllLogsACompletionLineWithTheRealCountsOnACleanDrain and
drainAllLogsANonZeroAbandonedCountForASessionStillBusyAtTheDeadline, "no drain-complete INFO
logged, expected true but was false". Restored byte-identical
(d21ecd3adb3f66933e2f248a8ec81a2c324bfee323a3cc30da525842c33ea80a). So the tests really depend on
the line rather than passing for another reason.

Design accepted: drainSnapshot returns a private DrainTally record and release() returns the
removed session so abandoned can reuse the same BUSY-at-removal check logPreservedForShutdown
makes, instead of a second registry read. Both are private with one call site.

Two follow-ups filed rather than fixed here, see the ticket comments.
2026-09-12 06:34:51 +02:00
Dai Ha 6a7342b1f0 fleetd #518: make the FleetMcp caller-resolution wiring an explicit choice, and test it once for real
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 2m12s
FleetMcp's contextExtractor picked its principal-resolution path off `callers == null`, so
"authorization off" also silently swapped in a second, untested identity heuristic
(legacyPrincipal). Nothing drove that closure through a real MCP request, so the whole wiring
was an unexercised claim.

- callers (CallerResolver) is now required, never null.
- A new AuthorizationMode enum (ENFORCED/UNENFORCED) is a required constructor parameter with
  no default, replacing the null-means-legacy idiom for whether denyFor enforces at all.
- legacyPrincipal is deleted: there is exactly one resolution path now
  (callers.resolve(...)), so the mutation that swapped it for an unconditional legacy call no
  longer compiles ("cannot find symbol: method legacyPrincipal").
- FleetMcpContextExtractorTest boots the real transport on a real Jetty server and drives it
  with a real MCP client, proving fleet_whoami's resolved role comes from CallerResolver's
  token check.
- Adapted FleetMcpAuthzTest/FleetMcpHandoverTest call sites; theLegacyConstructorLeavesTheGateOpen
  keeps its meaning under the new AuthorizationMode.UNENFORCED value.
2026-09-12 11:34:12 +07:00
Dai Ha a5ad7c6561 fleetd #519: test policy probe guards
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m51s
2026-09-12 11:30:35 +07:00
ltms c71ac231e5 Merge #520: pin the drain-gate abort branch and jar_id's absent case (fleetd #517)
CI / contract (push) Successful in 1m28s
CI / build (push) Successful in 1m50s
Verified by the lead, not taken from the worker's report:

- 40 test functions defined, 40 invoked (my own greps), suite exit 0, zero real FAIL lines.
- Harness proof on my own invocation: re-applying the jar_id absent->present mutation gives
  exit 1 and "FAIL: jar_id with no arguments must report absent when $JAR does not exist".
  So a green run from this suite is readable.
- Script restored byte-identical after every mutation:
  2cb83dc380c7226191d657c40fccdfc856904e40b2d03d0851fb6e522cee2f41.
- drain_gate_refusal is pure and prints only; die() stays outside the command substitution, so
  the "die inside $( ) exits only the subshell" trap does not apply here.
- The --no-build + staged-present case reads "nothing changed" deliberately, and that is correct:
  --no-build stages nothing itself, and a leftover staged jar is wiped by rm -f "$JAR_STAGED"
  at :607, before the build at :610.

The worker's out-of-scope finding is real and I reproduced it: the swap guard at :723 can be set
to `if false` with the suite still green at exit 0. Filed separately.
2026-09-12 06:29:08 +02:00
Dai Ha 33720c42b3 fleetd #512 (part 1): log a positive completion line when drainAll finishes
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 1m38s
drainAll used to log nothing on a clean drain — both existing log calls
(drainSnapshot's per-session failure, drainAll's straggler-sweep warning)
sit on abnormal paths, so "drained fine" and "died on the first session"
looked identical: no log line either way.

Add one log.info at the end of drainAll: "drain complete: released=N
abandoned=M (still BUSY at the shutdown deadline)". It fires on the
normal path, including the all-zero case, and folds both drainSnapshot
passes (main snapshot + straggler sweep) into one line.

drainSnapshot now returns a private DrainTally(released, abandoned)
record instead of void, and the private release(paneId, cause) overload
now returns the removed MemberSession (previously void) so drainSnapshot
can read its state at the moment of removal — the same check
logPreservedForShutdown already makes. Both signature changes are
private with a single call site, so the blast radius stays small.
2026-09-12 11:29:01 +07:00
Dai Ha 3833d8e52b fleetd #517: pin the drain-gate abort branch and jar_id's absent case
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m31s
Two mutation-testing survivors in scripts/redeploy-fleetd.sh: a source-text
test pins what a message SAYS but never whether the branch that prints it is
REACHED.

- Extract the drain-gate abort decision into drain_gate_refusal(do_build,
  staged_path), a pure function the suite can call directly for all four
  build/staged combinations. The existing source-text grep test is kept
  alongside it (it catches a re-wording; the new tests catch a dead branch).
- Extend jar_id's test to cover the missing-file path (both the no-argument
  default and an explicit path), which the #511 test never exercised.
2026-09-12 11:23:37 +07:00
ltms b37def9238 Merge #516: the probe refuses with three distinct messages, each naming its own cause (fleetd #500)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m46s
2026-09-12 06:09:50 +02:00
ltms 8f02576df6 Merge #515: pin the two-client completeness fold, and legacyPrincipal earns no authority (fleetd #509)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m54s
2026-09-12 06:06:12 +02:00
ltms 525bc1c5f4 Merge #514: the drain-gate abort message names a recovery that works, and jar_id()'s default is pinned (fleetd #511)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m41s
2026-09-12 06:00:23 +02:00
Dai Ha d59ece6dec fleetd #500: stop a wrong-interpreter or failed-parse reading a policy as empty
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m8s
probe-member-credentials.sh used mapfile < <(producer) to parse the fetched policy. That
hides a producer failure three ways: mapfile is bash 4+ and missing on macOS's /bin/bash
3.2, a process substitution's exit status is never propagated to mapfile, and the
downstream reads (":-" defaults and a slice) never fire set -u on a short or unset array.
All three converge on the same "0 known names" refusal, which blames the policy for a
failure that is actually the interpreter or the parser.

Three distinct guards, each closing one cause with its own message:
- a BASH_VERSINFO gate at the top refuses outright on bash < 4 (exit 3)
- the parser's output is captured via command substitution instead of mapfile < <(...),
  so a non-zero jq/python3 exit is caught at the call while the fact still exists (exit 4)
- an arity check before the field slice refuses a parse that exits 0 but returns fewer
  than 5 fields (exit 5)

The existing "0 known names" guard is now honest: by the time it fires, the three causes
above are already ruled out, so it really does mean the policy has 0 known names.
2026-09-12 10:58:16 +07:00
Dai Ha 32408d1e64 fleetd #509: pin the pane-scan completeness fold, and stop legacyPrincipal handing out primary
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Successful in 1m40s
Unit 1 — PaneLocator.terminalForPid's completeness fold across herdr
clients (PaneLocator.java:117) had no test that varied the number of
clients, so a mutation that keeps only the last client's Lookup.complete()
instead of ANDing every client's outcome survived: 14 of 15 existing tests
agree with the mutant on a single client. Added a two-client test where
the lead client errors on the pane that would have owned the pid (an
incomplete, negative scan) and the member client cleanly finds no panes
(a complete, negative scan) — the real fold ANDs these to false, a
last-wins fold reads it as true. Proved against MUTANTC
(complete = outcome.complete();): the new test fails with
"expected: <false> but was: <true>", the file was restored byte-identical
(sha256 unchanged), and the control run is green.

Unit 2 — FleetMcp.legacyPrincipal's else-branch returned Principal.primary
for ANY caller the connection did not resolve to a worker pane, with none
of CallerResolver.java:254's isLoopback/scanComplete guards. Measured that
no production caller passes null callers (Fleetd.java:696 always
constructs a real CallerResolver) but FleetMcpAuthzTest.mcp(false)
legitimately does, for its "legacy constructor leaves the gate open" test
— so the null-callers path is not dead code to delete (option a), it is a
documented legacy mode (option b). Changed the else-branch to
Principal.anonymous() and widened legacyPrincipal to package-private (like
denyFor) so a new test pins the behavior directly, since it only ever ran
inside a contextExtractor closure no existing test triggers.
2026-09-12 10:58:09 +07:00
Dai Ha 6e23bf8309 fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
The drain-gate abort message told the operator a rerun "with or without
--no-build" would finish the restart. That is wrong: by the time this
message can fire, stage_built_jar has already moved the jar off $JAR, so
--no-build hits require_no_build_jar's own refusal. Reworded to say the
rerun must NOT use --no-build, and why: the built jar is no longer at the
live path that --no-build requires.

Also added a test pinning jar_id()'s no-argument default (reports $JAR,
the live path) and its explicit-argument behavior (reports that path
instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm
-f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn
clean install deletes target/ on the next line regardless).
2026-09-12 10:56:32 +07:00
ltms aa4c0b84c3 Merge #510: never build into the path a running daemon holds (fleetd #493)
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m4s
2026-09-12 05:44:16 +02:00
ltms 40c593cd09 Merge #508: a herdr error during the pane scan is refused, not promoted to primary (fleetd #505)
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m38s
Closes the worker->primary escalation that fleetd #317 left open on the other side of the same
ternary. #317 stopped an UNRESOLVABLE pid being promoted. This stops a RESOLVED pid whose pane scan
failed being promoted: the scan now reports completeness, and an incomplete scan resolves to
anonymous.

Shape chosen: PaneLocator.terminalForPid returns Lookup(terminal, complete); paneOwnsAnyOf returns a
private Ownership enum OWNS/DOES_NOT_OWN/UNKNOWN, so a HerdrException is UNKNOWN rather than a clean
negative. ConnectionIdentity.Caller gains scanComplete as a third field -- resolved() was NOT
widened, correctly: it is documented as testing the lsof sentinel and #505 is a different axis. A
definite match still short-circuits, so a genuinely vanished non-owning pane does not become a
refusal.

Verified by me, not taken from the report.

Trial-merged onto main (136312f) and built the merge:
  mvn clean install -> Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS

Harness proof, by re-running the worker's own mutation rather than trusting it (UNKNOWN ->
DOES_NOT_OWN): Tests run: 63, Failures: 3 -- exactly the three tests meant to pin it, matching their
report line for line:
  CallerResolverTest.aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary
  ConnectionIdentityTest.scanIsIncompleteWhenHerdrErrorsOnThePaneThatOwnsThePid
  PaneLocatorTest.anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative
So the cells can go red, and their numbers are honest. Restored, shasum 81b797a6... byte-identical,
control 63/0.

MY OWN MUTATION, on the half they did not touch, FOUND A GAP. I mutated the two-client completeness
fold in terminalForPid -- `complete = complete && outcome.complete()` -> `complete =
outcome.complete()`, which forgets an earlier client's failure. Two greps proved it applied. Result:
Tests run: 63, Failures: 0. NOT pinned.

It is not an equivalent mutant: with CB-185's two herdr daemons, a first client that scans
incompletely followed by a second that scans cleanly gives `false` originally and `true` mutated, so
an incomplete scan would be reported complete and promoted. It is the #393 shape instead -- no test
assembles that combination, because the single-daemon tests collapse lead and member to one object.
That axis has cost us before: four routing defects passed the suite when one client served both roles
with memberHerdrSocket unset.

Merging anyway, because the production behaviour is correct and the ticket's defect is pinned three
ways -- what is missing is a test for a dimension the PR description claims in prose. Filed as a
follow-up with the numbers above.

Second item for the follow-up, latent not live: FleetMcp.legacyPrincipal still does
`c.terminal() != null ? worker : primary` and drops scanComplete, so it would re-open exactly this
hole. Measured unreachable in production -- Fleetd.java:696 always passes a real CallerResolver, and
the contextExtractor only calls legacyPrincipal when `callers == null`. A defect on paper is not a
reachable defect, so it is not a blocker; it is a trap for whoever next touches that constructor.

That second item is also the fleet01 lead's structural point, and it is the sharper framing: #415's
antidote (make the decision total) is ORTHOGONAL to this defect. Authz.permits is already a
default-less switch over Action and cannot help, because it says nothing about whether the caller was
resolved correctly, and Principal carries no record of how it was resolved. #415 guards a decision
nobody wrote; this was a decision written correctly and fed a bad input. The durable fix direction is
that an unresolved scan should be unrepresentable as a principal rather than checked for at each
consumer.
2026-09-12 05:35:41 +02:00
Dai Ha 979adf82eb fleetd #493: never build into the path a running daemon holds
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m37s
redeploy-fleetd.sh's build step wrote straight into fleetd/target/fleetd.jar
via `mvn clean install` while the OLD daemon was still running from that
exact path. A JVM loads classes lazily, so a class the daemon had not
touched yet could be read from a jar already replaced or removed -- the
failure landed on the shutdown drain (NoClassDefFoundError, exit 143,
looks clean).

Stage the freshly built jar at target/fleetd-new.jar (stage_built_jar),
confirm the OLD pid has actually exited (wait_for_daemon_exit, extracted
from the existing wait loop), and only then swap it into the live path
(swap_staged_jar) -- strictly after the wait, strictly before start. A
failed swap dies without starting. --no-build and --check keep their
existing, truthful behavior (require_no_build_jar; jar_id still reads
the live path by default). A leftover staged jar from an interrupted
run is wiped before the next build. The build still runs before
anything is stopped, so a failed build still never takes the fleet down.

Also fixed: the drain-gate abort message ("aborted -- nothing changed")
now names the staged jar when one exists, since staging already moves
the freshly built jar off the live path before that prompt runs.

Adds unit tests for stage_built_jar, swap_staged_jar, require_no_build_jar,
wait_for_daemon_exit, and a source-order test proving swap sits after the
wait and before start (sourcing stops before the main flow ever runs, so
the ordering itself can only be checked by reading the script's own call
sites).
2026-09-12 10:35:19 +07:00
Dai Ha 36870836aa fleetd #505: a herdr error during the pane scan must not read as a clean negative
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m29s
A transient herdr error on pane.process_info during PaneLocator's pid→pane scan used
to be swallowed into a plain "does not own it", so a real worker whose owning pane
errored mid-scan resolved with a null terminal but a resolved (real) pid — exactly
what CallerResolver's loopback-trust fallback reads as the primary. That is a
worker→primary privilege escalation through the door fleetd #317 did not close: #317
guards a failed lsof lookup (c.resolved()), not a failed herdr pane scan.

Fix: add a third state to the scan instead of widening Caller.resolved() (which stays
centralised next to the lsof sentinel it tests, per #505's explicit instruction not
to reopen that decision). PaneLocator.terminalForPid now returns a Lookup(terminal,
complete) record: a HerdrException on one pane marks that pane's ownership UNKNOWN,
not DOES_NOT_OWN, and the scan is complete only if every pane was either matched or
confirmed not to own the pid. A definite match found elsewhere in the same scan
still short-circuits as complete — a pane that genuinely vanished mid-scan without
being the caller's own does not turn into a refusal.

ConnectionIdentity.Caller carries the new scanComplete flag alongside the unchanged
resolved(). CallerResolver's loopback-trust fallback now requires both resolved()
and scanComplete() before promoting to Principal.primary(); an incomplete scan
resolves anonymous, which fails toward the recoverable error (a refused primary
retries loudly; a promoted worker would not).

Logs a warning naming the pane and which herdr client (of how many) failed, so the
incomplete-scan path is diagnosable rather than silent (fleetd #317's own lesson).
2026-09-12 10:27:42 +07:00
ltms 136312fb11 Merge #503: Injector's readiness-grace warn prints measured elapsed time, never arithmetic on constants (fleetd #501)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m27s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (4f28da6) and built the merge, because a clean auto-merge is not a compiling
merge:
  mvn clean install -> Tests run: 1689, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  InjectorTest -> Tests run: 34, Failures: 0
  Gitea CI on ac351ee: success

The worker mutated the two log arguments. I mutated the half it did not: the STAMP SITE. Changing
the first-sample guard so the clock is re-stamped on every non-ready poll gave 1 failure,
InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants, on a clock-read
counter assertion: "the readiness-not-ready branch should read the clock exactly twice — once to
stamp the first non-ready sample, once at grace expiry". Two greps proved the mutant applied;
restored, shasum byte-identical, control 34/0 green.

Behaviour preserved: the restructure from `else if (p != null && ++t.notReadySincePoll >= N)` to a
nested `else if (p != null) { ... }` keeps the short-circuit, so the counter still increments only
when p != null. All three reset sites now clear notReadySinceMillis alongside notReadySincePoll.

Two notes for the record, neither a blocker.

The poll-count half is an equivalent mutant and the worker said so instead of reporting a kill it
did not get. That is the right call and the code comment states the limit honestly.

The clock is injected through a package-private constructor overload while the public constructors
default it to System::currentTimeMillis. My brief prescribed that shape. With exactly one production
construction site (Fleetd.java:531) a required parameter — fleetd #415's antidote — would have been
just as cheap and would match LeadRollover and the new Fleetd.awaitHerdr. The silent-survivor risk
#415 names is not present here, because the field is final and both public constructors delegate, so
a future constructor cannot compile without supplying it. Recording the choice so the next person
does not read it as an oversight.
2026-09-12 05:07:35 +02:00
ltms 4f28da62a3 Merge #499: redeploy-fleetd.sh gains an "unclear" supervisor state that cannot reach kill, and the detail survives the subshell (fleetd #492, #497)
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m32s
Two commits, b17f37a and 599419f. Verified by the lead, not taken from the worker's report.

b17f37a adds the fifth answer "unclear" for installed-but-not-loaded and for a probe that could not
answer at all, and makes require_drivable_supervisor die on it.

599419f fixes a defect I found in b17f37a: SUPERVISOR_UNCLEAR_DETAIL was set inside
detect_supervisor, which is called as $(detect_supervisor). A subshell only returns stdout, so the
detail never reached the caller, and a ${VAR:-generic} fallback then printed a message naming no
supervisor and no reason. Measured before the fix: plain call -> detail length 142; via $( ) ->
length 0.

Lead verification of the fix, re-running my own original measurement across every branch:

  launchd installed, not loaded   kind=[unclear]   detail_len=142
  systemd installed, not active   kind=[unclear]   detail_len=146
  launchd loaded (normal)         kind=[launchd]   detail_len=0
  systemd loaded (normal)         kind=[systemd]   detail_len=0
  neither -> none                 kind=[none]      detail_len=0
  both loaded -> ambiguous        kind=[ambiguous] detail_len=0

The separator is always emitted, so the ${RAW#*SEP} unpack cannot silently fall back to the whole
string on the four branches that carry no detail.

Harness proof, run by me: making detect_supervisor print the kind alone (two greps proved the mutant
applied) turned the suite red, exit 1, "FAIL: detail does not name the systemd unit it found
installed-but-not-loaded". Restored, shasum byte-identical, control exit 0, bash -n clean on both
files. 23 test functions defined, 23 invoked.

The three case "$SUPERVISOR_KIND" switches now have final arms, chosen by what each caller does with
the value: the reporting switch warns and continues (a diagnostic that aborts goes silent exactly
when the state is novel), while stop and start die.

Not fixed here, filed separately: the "was already not running" branches at :603 and :610 swallow a
failed launchctl unload / systemctl stop with || true and then print "ok" unconditionally.
2026-09-12 05:04:04 +02:00
ltms 708f1795ad Merge #502: awaitHerdr reports three outcomes with measured elapsed time, not one boolean (fleetd #498)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m49s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (8f59019) locally and built the merge, because a clean auto-merge is not a
compiling merge:
  mvn clean install -> Tests run: 1687, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  FleetdAwaitHerdrTest -> Tests run: 6, Failures: 0
  Gitea CI on 274afaf: success

The worker mutated the three return values. I mutated the half it did not: the reap DECISION at the
call site. Changing logHerdrWaitOutcomeAndShouldReap to return true for INTERRUPTED gave 1 failure,
FleetdAwaitHerdrTest.interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed:135 ("an interrupted
wait must not tell main to reap"). Two greps proved the mutant applied; restored, shasum
byte-identical, control 6/0 green.

Behaviour preserved: herdrUp is still true only for ANSWERED, and both its readers (the orphan reap
at :276 and the lead auto-launch gate at :375) see exactly what they saw before.

One correction to the worker's report, which changes nothing in the code. It wrote that a real
interrupt-detection regression "would also hang the daemon's startup thread forever". It would not:
in production `nanos` is System::nanoTime, so the deadline check still fires. The hang it hit was a
test-fixture property — a frozen injected clock with no iteration bound. That is fleetd #486's
shape, now seen in a second class, and it is recorded there.
2026-09-12 05:01:02 +02:00
Dai Ha 599419f9e6 fleetd #492 follow-up: carry the unclear detail across detect_supervisor's subshell boundary
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m53s
b17f37a set SUPERVISOR_UNCLEAR_DETAIL as a global inside detect_supervisor, but the real call
site invokes it as $(detect_supervisor) — a subshell — so that global died with the subshell and
the die() message's ${VAR:-fallback} silently masked the loss with generic text.

- detect_supervisor now packs kind and detail onto its one stdout line (joined by the ASCII unit
  separator byte, $SUPERVISOR_DETAIL_SEP), the only channel that survives $( ). The real call site
  unpacks both with in-shell parameter expansion — no extra subshell.
- Dropped the ${SUPERVISOR_UNCLEAR_DETAIL:-...} fallback at the die() message: under set -u, a
  missing detail now fails loudly instead of silently defaulting (same defect class as #497).
- Added a constraints comment block above detect_supervisor for future callers: stdout-only,
  no ${VAR:-default} papering over a lost value, and every case on the return value needs an
  explicit *) arm.
- Added *) arms to the three `case "$SUPERVISOR_KIND"` switches (report/stop/start): report warns
  and continues (display-only), stop/start die naming the value (they act on it).
- Rewrote the "unclear" test to go through the real call-site shape ($(detect_supervisor) then
  the same split), not a hand-constructed value, and tightened its final assertion to check for
  the actual detail text rather than $SYSTEMD_UNIT alone (the die() boilerplate names the unit
  either way, so that check could pass on a lost value).
2026-09-12 09:58:45 +07:00
Dai Ha ac351ee1de fleetd #501: readiness-grace expiry logs measured elapsed time and the loop's own poll counter, never the configured budget
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m2s
The line at Injector.java:383-387 printed two numbers that read as measurements
but were both compile-time constants: READINESS_GRACE_POLLS for the poll count
(the loop's own Target.notReadySincePoll counter was in scope at the same call
site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed
time — arithmetic on two constants, never a measurement, and wrong in the
direction that says everything ran on schedule.

Fix, copying the LongSupplier-clock shape LeadRollover already uses:
- print t.notReadySincePoll instead of the constant for the poll count (the two
  agree by construction on this branch, so no test can tell them apart — the
  comment says so honestly).
- add Target.notReadySinceMillis, stamped at the first non-ready sample and
  reset at all three sites notReadySincePoll already resets (:304, :355, :389
  pre-fix line numbers), to compute a real elapsed time at expiry.
- inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis)
  through new package-private constructor overloads so a test can supply a
  clock whose advance does not track POLL_INTERVAL_MILLIS.

Tests use a ListAppender to assert on the log message contents, per
LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub
clock returning two literal, non-derived values so it can fail if the fix
regresses to the constant-arithmetic line — proved by mutation: reverting the
elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that
one test red (1 failure); reverting the poll-count print to the constant is an
equivalent mutant (0 failures), because the counter and the constant are
identical at that exact call site by construction.

fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0.
2026-09-12 09:57:57 +07:00
Dai Ha 274afafde6 fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m51s
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare
  boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are
  two distinct, named states instead of the same false (fleetd #497's shape).
- awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required
  parameters, with no defaulted overload (fleetd #415), so a test can drive it.
- The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself
  cannot be driven from a unit test; it logs a distinct message per outcome, always printing the
  measured elapsed time next to the configured budget, never the budget alone.
- Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt
  flag) and the call site (the three distinct log messages), using ListAppender.
2026-09-12 09:54:59 +07:00
ltms 8f59019305 Merge #496: lead-rollover logs measured elapsed time and measured counts, never the configured budget (fleetd #494)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 2m1s
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report.

3fb3311 — the /clear-settle and success paths return measured elapsed time and a real nudge count
c87cc25 — print the measured nudge count; fix the same defect on the sibling turn-settle timeout line
e966cba — the poll count on the same line was still a constant; the test could not tell the difference

Lead verification at e966cba:
  mvn clean install -> Tests run: 1681, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  LeadRolloverTest  -> Tests run: 33, Failures: 0, Errors: 0, Skipped: 0
  Gitea CI on e966cba: success

Mutation M2 (PICKUP_GRACE_POLLS 8 -> 5), run by the lead: 2 failures,
clearGraceReleaseLogIsWarnWithMeasuredElapsed and successLogPrintsMeasuredElapsedForTheWholeRoll.
The emitted line under the mutant read "after 5 consecutive IDLE/DONE polls (4 of those were
nudged)", so both numbers now follow the loop. Restored, shasum byte-identical, control 33/0 green.

Known and deliberate limit, stated in the code comment: on the grace-release branch both counters
equal the constants by construction (one write site each), so no test can prove the difference on
that line. The change buys one source of truth, not provable coverage. The place nudges genuinely
varies with the run is the /clear-timeout warn, which is covered by a discriminating test.
2026-09-12 04:44:10 +02:00
Dai Ha e966cbadf9 fleetd #494 follow-up (2nd pass): the grace-release line still had one constant, and its test could not tell the difference
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m55s
Two more fixes on the same line, LeadRollover.java:600-624:

1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
   loop's own idlePollsAwaitingPickup counter instead of the constant. Same
   defect shape as the nudges fix from c87cc25, one argument over.

2. LeadRolloverTest's clearGraceReleaseLogIsWarnWithMeasuredElapsed asserted
   its nudge-count expectation as `LeadRollover.PICKUP_GRACE_POLLS - 1` — the
   same expression the production code used to build the log line from, so it
   could not discriminate a reverted fix. Rewritten to plain literals
   ("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"), proven to
   trip when PICKUP_GRACE_POLLS's value changes (2 failures under a
   PICKUP_GRACE_POLLS=5 mutation: this test and successLogPrintsMeasured...).

Also corrected the comment above the log line: idlePollsAwaitingPickup and
nudges each have exactly one write site on this loop's release branch, so they
cannot differ from PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 at this call
site — confirmed by re-deriving the loop's control flow, and by reverting just
the nudges argument (M1) and observing 33/0 stayed green even after the test
rewrite. That is an equivalent mutant on this line, not a gap the test rewrite
could close; the honest value of printing the counters is one source of truth
for the loop, not a provable-by-test difference here. The place nudges truly
varies with the run — and is covered by a test that can tell it apart from a
constant — is the /clear-timeout warn's clearResult.nudges() in runRollover.

mvn -Dtest=LeadRolloverTest test: 33/0. mvn clean install: 1681/0, BUILD SUCCESS.

Same-shape sweep (found, not fixed, per instructions):
Injector.java:383-387 — the readiness-grace warn prints the constant
READINESS_GRACE_POLLS where the measured per-target counter
t.notReadySincePoll is in scope (single increment site, just reached the
threshold at this call site — same equivalent-mutant situation as this fix).
2026-09-12 09:39:39 +07:00
Dai Ha b17f37a683 fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m32s
Two situations were silently landing in the "none" answer, which
require_drivable_supervisor accepts and the script then falls back to a raw
kill + nohup — exactly the wrong move when a supervisor actually IS present:

- installed-but-not-loaded, on either supervisor. `systemctl --user is-active`
  answers "no" for activating/deactivating/failed and while an auto-restart is
  pending too, and every one of those is a host that IS under systemd (or
  launchd) and about to act again. `*_installed` already knew this; it was
  only ever consulted for a warning line, never by the decision itself.
- a systemd probe that could not answer at all (e.g. systemctl cannot reach
  the user bus over a non-lingering ssh session) looked identical to a clean
  negative, because both probes redirected stderr straight to /dev/null.

detect_supervisor now returns a fifth answer, "unclear", for both cases.
systemd_loaded/systemd_installed capture systemctl's exit status and stderr
separately and set their own *_ERRORED flag only on a real tool failure
(non-zero exit WITH stderr), never on a clean negative. "none" now means only:
neither supervisor installed, neither loaded, neither probe errored.
require_drivable_supervisor die()s on "unclear" exactly like it already does
on "ambiguous", naming the specific supervisor and reason via the new
SUPERVISOR_UNCLEAR_DETAIL global.

Tests: 4 new cases (systemd/launchd installed-but-not-loaded, a real
systemd_loaded run through a systemctl stub that errors on stderr, and the
die() refusal for "unclear" naming the unit). All 3 new guards were verified
by mutation: each was removed from the real script, the suite caught it (a
new FAIL line naming the exact broken assertion), then the file was restored
byte-identically and the suite went green again.
2026-09-12 09:23:21 +07:00
Dai Ha c87cc25aa6 fleetd #494 follow-up: print the measured nudge count, and fix the sibling turn-settle timeout line
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m30s
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter
  instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the
  constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8
  nudges where 7 went out) and a code read did not catch it — only a mutation test did.
  Printing the counter cannot drift from the loop's real behaviour.
- waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value
  printed as if measured' defect as the three lines fixed in the original #494 commit, but was
  out of scope because the brief named specific lines instead of the shape. Fixed the same way:
  it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the
  timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds()
  as the measured wait.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured
  nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is
  updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a
  new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line.
- Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via
  two greps that the mutant applied and the original constant was gone, ran the grace-release
  test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8
  and reran the full LeadRolloverTest suite as a control (33/33 green).
2026-09-12 09:22:19 +07:00
Dai Ha 3fb331145a fleetd #494: log measured elapsed time, never the configured budget, on a lead-rollover failure/success
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m35s
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
  instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
  prints the measured elapsed time next to the target pane — this is the exact path that
  reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
  whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
  instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
  self-advancing fake clock so the measured elapsed/nudge values are deterministic and
  provably distinct from the configured budget.
2026-09-12 09:08:04 +07:00
Dai Ha dcd505286f fleetd #492: teach redeploy-fleetd.sh systemd --user as a third supervisor
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
launchd, systemd --user, and unsupervised are three different answers, not two.
Refuse (die) rather than fall through to kill+nohup when a supervisor is
detected that this script cannot drive (e.g. both signals fire at once), and
add a post-restart check that fails the run if more than one fleetd process
is alive. detect_supervisor()/require_drivable_supervisor()/
count_daemon_pids()/assert_single_daemon() are pure, overridable functions so
scripts/test-redeploy-fleetd.sh can exercise them without a real launchd or
systemd.
2026-09-12 09:00:59 +07:00
Dai Ha 60fa958da2 docs: a relative handoverPath lands in the LEAD's repo, not fleetd's (#491)
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore
protects the handover file and the distinction is invisible. On fleet01 the
lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory,
and that repo has no .handover rule — measured 2026-09-12.

Tell the lead to check its own workspace's .gitignore before writing, and to
report it rather than committing the file or editing someone's .gitignore.

Refs fleetd #491, #487, #480.
2026-09-12 08:46:17 +07:00
Dai Ha 9950361bc9 docs: #489 is fixed and deployed, but criterion 6 is still unmet
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m33s
The skill said the roll 'does nothing until #489 is merged and redeployed'.
Both have happened, so that sentence now reads as a green light. It is not
one: no roll has bootstrapped a fresh session end to end yet. Say that
plainly, and tell the lead to warn the operator before confirming.

Refs fleetd #489, #480.
2026-09-12 07:55:41 +07:00
ltms 9f74b6619a Merge #490: nudge the /clear submit keystroke before bootstrapText (fleetd #489)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m30s
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".

Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
  including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on f687046: success

A separate mutation of the FIRST gate hung the suite instead of failing it.
That is fleetd #486, not a regression here; the reproduction is recorded there.
2026-09-12 02:52:52 +02:00
Dai Ha f687046450 fleetd #489 follow-up: fix stale class javadoc, off-by-one nudge count, weak test
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m33s
Three review corrections on top of the previous commit:

1. The class javadoc's four-step continuation list (lines 47-58) was stale.
   Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
   own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
   described the old plain re-check ("the original, pre-correction wait...
   still here") - fixed to describe what waitForClearPickupAndSettle
   actually does: nudge while unpicked-up, then wait for a real WORKING ->
   IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.

2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
   polls, not the number of nudges - the 8th poll releases instead of
   nudging again, so 8 polls produce 7 nudges. The log.info in the release
   branch and two javadoc spots said "8 nudges"; fixed all three to state
   the poll count and the nudge count separately and correctly. Behavior
   and the constant are unchanged.

3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
   and sendKeysCallCount, both of which a return-true stub also satisfies.
   Added an assertion on the already-tracked postClearGetCalls counter
   (>= 2), which only a real post-/clear poll loop can produce - this is
   what makes the test fail against a return-true mutant.
2026-09-12 07:45:36 +07:00
Dai Ha c6058652be fleetd #489: nudge the /clear submit keystroke before bootstrapText
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m36s
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled
for IDLE/DONE, which /clear itself never leaves since it starts no real turn,
so it always returned true on the first poll. Combined with a direct
agents.send bypassing Injector (deliberate, to avoid wedging the pane), the
submit Enter that accompanies /clear could race the paste and leave it
unsubmitted — bootstrapText then landed concatenated onto the same input
line, exactly as measured live on 2026-09-12.

Replace that second wait with waitForClearPickupAndSettle, which copies the
pickup-nudge pattern Injector already ships for its own post-turn /clear
housekeeping (fleetd #306): nudge agents.submit while the pane hasn't
reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than
wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is
observed. BLOCKED stays excluded from both the nudge and the boundary check,
same as the (unchanged) first wait — a paused live turn is not settled, and
nudging Enter into an open prompt could wrongly answer it.

Adds four tests to LeadRolloverTest covering the paste-race regression
(nudge ordered between /clear and bootstrapText), a confirmed pickup, a
deadline expiry with no boundary ever reached, and a throwing submit().
2026-09-12 07:29:26 +07:00
Dai Ha 2a95da2ff4 docs: the handover skill must warn that the automatic roll is broken (#489)
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 1m53s
Section 11 said the bootstrap prompt landing was 'not yet proven
end-to-end'. It is now measured failing: the first real roll joined
/clear and bootstrapText into one line. Nothing was cleared, so the
failure is safe, but a lead that reads the old wording would reach for
fleet_handover expecting it to work.

Refs fleetd #489, #480.
2026-09-12 07:21:38 +07:00
Dai Ha 7f9137fcb6 fleetd #480: ignore the handover file, and note the absolute-path guarantee in the handover skill
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m52s
The lead rollover handover file now lives inside the workspace, at the relative
path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one
moment's live state, so it must never enter git history.

The handover skill also now says the handoverPath fleetd hands back is always
absolute, even when the configured value is relative — a lead that resolves it
itself can pick a different file from the one the daemon checks.
2026-09-12 05:46:34 +07:00
Dai Ha 3bf3968bc7 Merge #487: resolve a relative leadRollover.handoverPath against the calling lead's workspace (fleetd #480 follow-up) 2026-09-12 05:43:12 +07:00
Dai Ha 261aa056f9 fleetd #480 follow-up correction 2: guarantee resolveHandoverPath is always absolute
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m32s
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.

Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.

Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.

Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.

Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
2026-09-12 05:41:40 +07:00
Dai Ha 042b8c99dd fleetd #480 follow-up correction: cover Fleetd.leadRollover(...)'s own wiring behaviourally
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m51s
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...)
factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live
one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative
handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live
lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at
construction time (a lead discovered by the tab scan after leadRollover(...) was built still
resolves correctly).

Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always
"no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full
1669-test suite green before this commit. With the new test added, the same one-line mutation now
fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during
verification and is not part of this commit (git diff on Fleetd.java is empty).

FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test
could catch this wiring dropping out, which was true only before this commit and only covered the
factory's own body, not its call site. Restated what each test class actually covers: the source-
text pin covers the call site's argument list; the new behavioural test covers the lambda's body.
2026-09-12 05:32:25 +07:00
Dai Ha 4bfab6b718 fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m4s
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the
calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that
lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores
only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed
back to the lead in the fleet_handover open response, and the default bootstrapText sentence
all see the same absolute location instead of a value resolved against whatever directory the
daemon process happened to start in.

FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it
would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor
(resolvedHandoverPath) method builds the default sentence from the resolved path instead.

Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to
lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on
every call, never off a startup snapshot.
2026-09-12 05:19:49 +07:00
Dai Ha 008a457557 docs: fleet_handover is shipped — update the charter table and the handover skill
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m38s
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to
catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the
canonical block is an incomplete change.

CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover.
Both copies edited identically; the byte-identical sync check prints True.
The row carries the trap rather than just the call: open FIRST, then write the
file, then confirm — because confirm refuses the file as stale unless its
modified time is later than the open request, so the obvious order fails.

.claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover
tool: it does not exist." That was true when it was written this morning and is
now false, which is the worst state for an instruction file to be in. Replaced
with the two real paths (by hand, or with the tool when leadRollover: is
configured), plus a new section 11 giving the three-step order and the five
things that surprise a caller — chiefly that `accepted` does not mean the pane
has been cleared, and that operatorConfirmed is a report of what a human said,
not a confidence level.

Section 11 also states plainly that the bootstrap prompt landing in a freshly
cleared pane is not yet proven end-to-end, and that the recovery is the manual
path — which is why the file is written before confirm, never after.
2026-09-11 07:29:59 +07:00
ltms 9494a6b99a Merge #485: fleetd #480 Unit C — the fleet_handover MCP tool
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m58s
fleet_handover{action: "open"|"confirm"|"cancel"} drives LeadRollover, which #483 and
#484 landed with nothing calling it. Primary-only via a new Authz.Action.HANDOVER, on
the same case line as SPAWN/STOP/DRAIN.

The tool has NO terminal, session or leadTerminal parameter of any kind — the pane is
always callerTerminal(exchange), resolved from the connection. A lead can therefore only
ever roll itself, never another lead. That is charter invariant 3, and it is the second
of the two corrections recorded in LeadRollover's class javadoc.

Registered unconditionally, so the tool surface does not vary with config: with
leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal instead of
failing, and open()'s IllegalStateException (config removed by a hot reload after
construction) is caught and turned into the same refusal. A config-dependent tool set
would have collided with #474's charter tool-surface gate and McpContractDocTest.

Correction round applied before merge, and it is the reason this took two passes.
The unit first shipped with a defaulted 15-argument FleetMcp constructor delegating to
the new 16-argument one with leadRollover = null. I mutated the wiring rather than
reasoning about it: deleting just the leadRollover argument from Fleetd.main's FleetMcp
call compiled with 0 errors and passed all 1659 tests, BUILD SUCCESS — while the live
daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
Neither FleetMcpHandoverTest (it builds its own FleetMcp) nor FleetdLeadRolloverWiringTest
(it pins that LeadRollover is constructed, not that it is passed on) could see it.

The worker then found the defect was wider than I had named: all five shorter
constructors (11/12/13/14/15-arg) formed one defaulting chain into the 16-arg one, each
silently supplying another feature's "off" value — leadChannel, outage, leadSeats, peers,
and finally leadRollover. All five are deleted. FleetMcp now has exactly one public
constructor, so every one of those features is compile-enforced at its call site, not
just this one.

Verified by the lead before merge, on PR head merged with current main (0176378):
- CI run 1728 green on eb05576.
- The identical mutation re-run by me: control against the original -> 1; mutant present
  (MUT-DROP-ARG2) -> 1; original gone -> 0; `mvn -o -q compile` now FAILS —
  "constructor FleetMcp ... cannot be applied to given types; reason: actual and formal
  argument lists differ in length" at Fleetd.java:[698,24]. The antidote holds: required
  parameter = compile error, defaulted overload = silent survivor a green suite vouches for.
- Restored clean (0 changed files), then mvn clean install on the merged tree:
  BUILD SUCCESS 1, BUILD FAILURE 0, Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.
- Three call sites of `new FleetMcp(` measured, all accounted for: Fleetd.java,
  FleetMcpAuthzTest.java, FleetMcpHandoverTest.java.
2026-09-11 02:27:56 +02:00
Dai Ha eb0557621e fleetd #480 correction round: collapse FleetMcp to one required constructor
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 1m32s
FleetMcp had a defaulted 15-argument constructor that delegated to the new
16-argument one with an implicit null for leadRollover. Dropping the
leadRollover argument from Fleetd.main's FleetMcp(...) call fell back to that
shorter overload, compiled fine, and left all 1659 tests green — the live
daemon would then answer NOT_CONFIGURED to fleet_handover forever with
nothing going red.

Delete every overload that could reach the 16-arg constructor with a
silently-defaulted leadRollover (11/12/13/14/15-arg forms all chained to it),
leaving the 16-arg constructor as FleetMcp's sole public constructor. Update
FleetMcpAuthzTest's call site to pass every parameter explicitly (leadChannel
null, OutageSource.none(), LeadSeatSource.none(), List.of(), leadRollover
null) — Fleetd.java and FleetMcpHandoverTest already called the full form.
Proved with mvn -o -q compile: removing the leadRollover argument from
Fleetd.main now fails to compile instead of silently defaulting.

No behaviour changes — NOT_CONFIGURED refusals are unchanged.
2026-09-11 07:25:08 +07:00
ltms a0eed6f01b Merge #484: fleetd #480 Unit E — a BLOCKED pane is not a settled pane
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m35s
LeadRollover's two settle waits used AgentStatus#injectable(), which is
IDLE || BLOCKED || DONE. That is the right rule for Injector ("may I deliver a message
without stepping on a live turn") and the wrong one here ("has the turn actually
ended"), because BLOCKED is a live turn that is merely paused — a pane sitting on an
approval prompt.

Path in: the lead calls confirm(); its turn carries on and hits anything needing
approval; herdr reports blocked; within 250ms the deferred continuation reads that as
settled; /clear is typed into an open prompt. That destroys the lead's live context
mid-turn, which is exactly what the turnSettleSeconds gate added in #483 exists to
prevent. The window is turnSettleSeconds, default 20s.

Both waits now require a real turn boundary — IDLE or DONE. AgentStatus#injectable() is
untouched: it is correct for Injector, LeadHeartbeatLoop, LeadCoordLoop, ReplyPushLoop
and HerdrPeerLauncher's two readiness checks, all of which are delivery gates.

Verified by the lead before merge:
- CI run 1726 green on head e2a91e8.
- Mutation on the half the worker did NOT mutate: dropped the DONE arm, leaving
  `if (status == AgentStatus.IDLE)`. Mutant proven applied with two unrelated proofs
  using different search strings (MUT-DROP-DONE present -> 1; 'IDLE || status' gone
  -> 0). Same single test, same 100s cap: unmutated passes in 0.172s; mutated is still
  running at 100s and does not pass. So doneStatusStillCompletesTheFullRoll really does
  pin the DONE arm, and the fix did not over-tighten to IDLE-only.

Why the earlier #483 battery could not have caught this: it proved the guard FIRES, not
that the guard's predicate was tight enough. LeadRolloverTest drove the fake status with
only "idle" and "working"; neither value discriminates this defect, only "blocked" does.

Correction round applied before merge: both warning lines said "never went idle" and
"did not become injectable". Retiring that word is the whole point of the change, so a
future reader debugging from those lines would have been handed back the wrong mental
model. Both now say "did not reach a turn boundary (IDLE or DONE)".

Filed separately as #486, deliberately not fixed here: with the DONE arm removed the
test hangs rather than fails, because waitUntilAtTurnBoundary is bounded only by an
injected clock that the tests hold fixed. Not a production bug — nowMillis is
System::currentTimeMillis there — but it turns a future regression into a CI hang
instead of a red test.
2026-09-11 02:21:06 +02:00
Dai Ha e2a91e883e fleetd #480 Unit E correction: retire "injectable" wording from the log lines
CI / build (pull_request) Successful in 1m28s
CI / contract (pull_request) Successful in 1m27s
Both waitUntilAtTurnBoundary guard messages still said "never went idle" /
"did not become injectable" — the old mental model the rename was meant to
retire. Made both say what the code now actually waits for: a turn boundary
(IDLE or DONE).
2026-09-11 07:18:51 +07:00
Dai Ha 62646957ea fleetd #480 Unit C: wire fleet_handover MCP tool onto LeadRollover
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m56s
Adds the fleet_handover tool (open/confirm/cancel) as a thin adapter over
LeadRollover, registered unconditionally so the charter tool-surface gate
sees a stable set regardless of whether leadRollover: is configured. With a
null LeadRollover every action degrades to a clean NOT_CONFIGURED refusal
instead of throwing. Gated on a new Authz.Action.HANDOVER (primary-only,
same as SPAWN/STOP/DRAIN). The caller's own connection-resolved terminal is
the only lead identity ever used — the tool's input schema carries no
terminal/session/leadTerminal parameter, so a lead can only ever roll
itself. Fleetd.main now passes its existing leadRollover local into FleetMcp
via a new trailing constructor parameter.
2026-09-11 07:07:24 +07:00
Dai Ha 802c0ab701 fleetd #480 Unit E: BLOCKED is not a settled turn boundary
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m55s
LeadRollover's waitUntilInjectable used AgentStatus#injectable(), which
accepts BLOCKED. A BLOCKED pane is paused mid-turn on a prompt, not
settled — reusing injectable() let /clear (or the bootstrap text after
it) fire into an open approval prompt within the 20s settle window,
destroying the lead's live context.

Renamed the helper to waitUntilAtTurnBoundary and restricted both waits
to IDLE or DONE only, with a comment explaining why this class does not
reuse injectable() (it answers "may I deliver", not "has the turn
ended"). Added tests for BLOCKED-forever on both waits (zero sends /
exactly one send) and for DONE still completing the full roll.
2026-09-11 07:01:28 +07:00
ltms 4bab23e241 Merge #483: fleetd #480 Unit A — lead rollover core (config + executor)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
Two corrections were applied to the original unit before this merge, and the PR
description above still describes the pre-correction shape:

1. confirm() no longer rolls inline. It validates every gate, then hands a one-shot
   continuation to continuationRunner and returns. confirm() is called BY the lead FROM
   its own turn, so the pane is WORKING and cannot report injectable until confirm()
   returns; the old code sent /clear first and then timed out waiting, which destroyed
   the lead's context and started no fresh session. The continuation waits for the
   calling turn to settle FIRST (turnSettleSeconds, a new knob), and if that wait times
   out it sends no /clear at all.
2. The pane to roll comes from the caller's terminal, not PrimaryRegistry. A single-slot
   lookup let lead X's confirm() clear lead Y's pane. open() records the caller's
   terminal; confirm() refuses with NOT_YOUR_ROLLOVER on a mismatch.

Verified by the lead before merge:
- CI run 1722 green on head a942622.
- Mutation battery on the two safety-critical branches, in a scratch worktree at a942622.
  Controls measured against the ORIGINAL first; each mutant proven applied with two
  unrelated proofs using different search strings.
  M1 `if (!turnSettled)` -> `if (false)`: turnThatNeverSettlesSendsNoClearAtAll FAILS
     (LeadRolloverTest:247, expected 0 sends but was 1).
  M2 ownership check -> `if (false)`: aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover
     FAILS (LeadRolloverTest:266). Unmutated harness-proof run: exit 0 with both guard
     WARN lines logged, so the branches really are exercised.

Nothing calls LeadRollover yet. The fleet_handover MCP tool is Unit C.
2026-09-11 01:51:20 +02:00
Dai Ha a94262271b fleetd #480 correction round: defer the roll, and gate it on caller identity
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m3s
Two defects found after the fact, both from the original brief, both fixed here.

1. confirm() is called FROM the calling lead's own turn, so its pane is still
   WORKING and can never report injectable inside that same call. The old
   confirm() sent /clear before polling for that — the poll always timed out,
   but only after /clear had already fired and queued, destroying the lead's
   context with no fresh session ever started and a refusal return that lied
   about what had happened.

   Fix: confirm() now only validates and, if every gate passes, hands a
   one-shot continuation to a new continuationRunner (a real virtual thread in
   production, Runnable::run in tests) and returns RollDecision.approved()
   immediately - "scheduled", not "rolled". The continuation itself does the
   actual work, once the calling turn has ended: wait for the SAME pane to
   report injectable again (new turnSettleSeconds config key, default 20) -
   if this never happens, /clear is NEVER sent, at all - then /clear, then
   wait again (clearSettleSeconds, as before), then bootstrapText. The "no
   timer/scheduler, only confirm() can roll" invariant is restated precisely
   in LeadRollover's class javadoc: it is about initiative, not synchronicity
   - a single-shot continuation of an already-approved confirm() call still
   satisfies it; a recurring background loop would not.

2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(),
   a single-slot lookup that is correct for a background loop with no caller
   but wrong here: on a daemon with more than one labelled lead tab, lead X's
   confirm() could clear lead Y's pane, violating the charter's "identity
   comes from the connection, never an argument" invariant.

   Fix: open() and confirm() now take the caller's terminal id as a parameter
   (resolved by the MCP layer from the connection - the later MCP-tool unit
   must pass it in, never accept it as a request field). confirm() refuses
   with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open()
   recorded. LeadRollover no longer depends on PrimaryRegistry at all.

Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the
new meaning - approved and scheduled, not necessarily cleared yet. Added
turnSettleSeconds to the leadRollover: config block (documented in
fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's
leadRollover(...) factory to drop the primaryRegistry parameter, with
FleetdLeadRolloverWiringTest's source-text pin updated to match.

New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters
most - a turn that never ends means /clear is never sent) and
aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER),
plus a settle-after-clear timeout test and an open() input-validation test.
LeadRolloverTest: 11 -> 14 tests.
2026-09-11 06:44:58 +07:00
Dai Ha 5c12865c25 fleetd #480 Unit A: lead rollover core (config block + executor)
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m58s
Adds the opt-in leadRollover: config block and LeadRollover, the executor a
later unit's MCP tool will call. A lead writes a handover file, then open()
records a token and confirm() verifies it (exists, non-empty, fresh) and an
operator confirmation before clearing the lead's own pane via /clear (sent
directly through AgentControl, bypassing Injector, same as
ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing
but an explicit confirm() call can ever roll a pane - no timer, no heartbeat,
no background thread anywhere in this class.

Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when
leadRollover: is present at startup, and nothing calls it yet - the MCP tool
is a separate, later unit.

Classified leadRollover: as HOT in ConfigRef (joins placement/
memberCredentials/memberLoginShell/models): the executor holds
Supplier<FleetConfig.LeadRollover> and reads every field fresh per call,
unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's
construction is still gated on presence in the startup config snapshot, so a
freshly-added block needs a restart before anything exists to call.

Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no
object without the config block, only confirm() can roll, missing/empty/
stale handover file each refuse by name, requireOperatorConfirm gating, and
the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source-
text pin on Fleetd.main's construction call, mirroring
FleetdCompletionResolverWiringTest). Also updated the existing
FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest,
ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to
account for the new record component.
2026-09-11 06:33:01 +07:00
Dai Ha 3f8c956325 docs: the handover skill must not promise a tool that does not exist
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m30s
The skill shipped in #481 described fleetd #480's automation in the present
tense: "fleetd ticket #480 lets a lead session hand off to a fresh one". The
tool is not built. A session loading the skill would look for a fleet_handover
tool it cannot call, and this repo already knows what a confidently wrong
instruction costs — a wrong comment stops the reader investigating with a false
conclusion, which is worse than no comment.

Now says plainly: today the handoff is manual, #480 will automate the same
cycle, and do not call fleet_handover because it does not exist. The nine
procedure rules are unchanged — only who performs the swap changes when #480
ships, not what the file must contain.

My wording to fix, not the worker's: the brief handed them the present tense.
2026-09-11 06:17:11 +07:00
ltms 2051ceeb26 Merge #481: the handover skill (fleetd #480 Unit B)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m27s
Adds .claude/skills/handover/SKILL.md — the outgoing lead's procedure for writing
the file a fresh lead session inherits.

Reviewed by the lead against the eight requirements in the brief: every number
carries its command, measured-vs-reported is marked, open decisions name their
owner, live hazards, what is not owed, a first section of three re-measurement
commands, a time-and-commit stamp, and a plain statement that the file goes
stale. All eight are present, plus a leave-out section and a template.

Verified by the lead, not taken from the worker's report:
- CLAUDE.md diff is the addendum's skill list only, in the primary-side group.
- The canonical block is byte-identical with the wiki template (sync check True).
- CI run 1718 green on 325d0771, the same commit reviewed.

Follow-up owed by the lead: the skill describes #480's automation in the present
tense, but the tool does not exist yet. The lead will soften that wording so the
skill does not promise a fleet_handover tool a session cannot call.
2026-09-11 01:16:40 +02:00
Dai Ha 325d0771a4 fleetd #480 Unit B: add handover skill for lead session handoff
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m54s
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead
follows to write the handover file a fresh lead session inherits when
fleetd clears the pane. Registers the new skill in CLAUDE.md's
primary-side skills list; no other change to CLAUDE.md.
2026-09-11 06:13:49 +07:00
Dai Ha 7b97aae85b docs: move the redeploy procedure out of CLAUDE.md into a skill
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m34s
CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.

The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.

CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.

Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.

CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
2026-09-11 05:48:06 +07:00
Dai Ha 49a404ddf3 Merge #474 follow-up: pin main's ConfigRef wiring against the surviving mutation
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m37s
My battery on the #474 merge found one survivor: reverting Fleetd.java:154
from the three-argument ConfigRef constructor to the plain two-argument
one turns the live reload gate off and leaves all 1633 tests green. Both
new #474 tests build their own ConfigRef with the method reference, so
neither reads what main chose.

This adds FleetdConfigRefWiringTest, following the three source-text
precedents already in the tree (FleetdBackendQuarantineWiringTest,
FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather
than the weaker sibling pattern that builds the wiring itself. No
production change.

The worker branched fresh off 4466ee0 rather than continuing its old
branch off the stranded 435e022 base. That was its own call and it was
the right one: one merge base, a clean 80-line diff, and no cherry-pick
needed this time.

Its vacuity guard is worth keeping in mind for the next source-text
test: it asserts the file it read contains 'public final class Fleetd'
before asserting anything about the mutation, with a message saying the
assertFalse below would pass vacuously on a broken read. An unrelated
anchor is the right choice there, because a guard that shares the
mutation's text cannot tell a bad read from a real change.
2026-09-10 20:50:13 +07:00
Dai Ha 72d6a6878b fleetd #474 follow-up: pin main's config wiring against M2
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m32s
Fleetd.main's own choice of the three-argument ConfigRef constructor
(with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation)
was unpinned. Reverting Fleetd.java:154 to the plain two-argument
constructor compiled with 0 errors and left the whole suite green,
because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest
each build their own ConfigRef directly rather than through main.

Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java
following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/
FleetdCompletionResolverWiringTest precedent: asserts the exact
three-argument construction is present, asserts the plain two-argument
form is absent, and guards against a vacuous pass on a broken/empty
source read by first asserting an unrelated anchor is present.
2026-09-10 20:48:05 +07:00
Dai Ha 4466ee0ef2 fleetd #474: ConfigRef.reload() runs the charter tool-surface gate too
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m44s
A charter naming an MCP tool the server does not register refused Fleetd.main
at startup but slipped through ConfigRef.reload(), because reload() only ran
FleetConfig.validateAll(), which never looks at what a charter's text names.

CharterToolSurface stays in the mcp package (config must not depend on it), so
ConfigRef now accepts the check as a Consumer<FleetConfig> extraValidation,
run inside reload()'s same try/catch as validateAll(). Fleetd.main wires a new
package-private adapter, Fleetd.assertChartersNameOnlyRegisteredTools, into
both the startup call site and ConfigRef's constructor, so the two call sites
can never check different things.

Tests: ConfigRefTest (reload refuses/accepts, via a locally-built equivalent
consumer since Fleetd's method is package-private to dev.ltms.fleet) and the
new FleetdConfigRefCharterToolSurfaceWiringTest (same proof through the exact
Fleetd::assertChartersNameOnlyRegisteredTools reference production uses).
Verified deleting the new extraValidation.accept(fresh) call site fails both
new "refuses" tests by name.

(cherry picked from commit 97e4c1d658)

Lead review. Cherry-picked, not merged: the worker branched from 435e022,
which is not an ancestor of main (I had reset and rewritten that commit's
message while the worker was already on it). Merging its branch would have
added a second merge commit for work already on main. Same content, no
duplicated history. Rule learned: once a worker is spawned, its base
commit is published.

My own mutation battery, five cells, each a full `mvn -B clean test` on
this commit's own tree:

  CONTROL 1, unmutated       1633 tests, 0 failures, BUILD SUCCESS
  M1 delete accept(fresh)    KILLED  2 by name
  M3 3-arg ctor stores no-op KILLED  2 by name
  M4 set(fresh) before valid KILLED  6
  M5 accept outside try      KILLED  2 errors
  M2 main uses the 2-arg ctor SURVIVED

M2 is a real gap and it is a test gap, not a defect: the production code
here is correct. Reverting Fleetd.java:154 to `new ConfigRef(configPath,
cfg)` turns the live reload gate off and leaves all 1633 tests green,
because both new tests build their own ConfigRef with the method
reference rather than reading what main wires.

That is the fourth Fleetd.main call site to survive a battery (#446 M5
and M7, #466 M1, now this). Extracting a check into a well-tested helper
moves the untested surface UP, into the line that chooses to call it. The
repo already has the answer in three source-text wiring tests
(FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest,
FleetdCompletionResolverWiringTest); the worker followed the other
sibling pattern, which builds the wiring itself and so cannot pin main.
Delegated as a follow-up, with the surviving mutation as its acceptance
criterion.

Also mine to correct: my M3 cell's annotation said the no-op store "must
be 2" and measured 1. The 2-arg constructor delegates rather than
assigning, so my pattern only ever matched the mutant. The cell still
stands on its other proof (requireNonNull left = 0). Third battery in a
row with a wrong CONTROL annotation of my own.
2026-09-10 20:40:49 +07:00
Dai Ha 25ba7f16bb Merge #473: fleet_profiles and fleet_list report which attempt a quarantine is on (fleetd #466 item 2)
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m24s
The escalating cooldown landed in 789b6a8, but the only thing an operator
could see was quarantinedForSeconds. A long number does not say whether
this is the first exhaustion or the fifth, and after escalation those look
the same from outside: 3600 seconds could be a big base cooldown or a
credential on its fourth strike.

Both reports now carry quarantineAttempt beside quarantinedForSeconds.

The part that matters is HOW, not that the field exists.
BackendQuarantine#status(credentialId) does ONE quarantines.get() and
returns both values from the same QuarantineState. It is not two accessors
that a caller pairs up. That is the pattern
CompositePeerLauncher.modelGateState() already documents: a report that
reads a different source from the behaviour it describes will eventually
disagree with it, and the disagreement is invisible because both halves
look right on their own.

This is reporting only. Escalation, the ceiling and the quiet-gap reset are
untouched. The flat two-argument constructor still reports a real growing
attempt count even though its cooldown stays flat - which is the honest
answer, since the streak is real whether or not the cooldown uses it.

The worker's report was lost: its fleet_reply never arrived and its inbox
was empty. Nothing was lost with it, because the brief required pushing
first and putting the report in the PR body. That is the second time that
practice has saved a turn.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:21:14 +07:00
Dai Ha 5ba69c9cf2 fleetd #393: correct a comment that claimed two tests pin a call they cannot
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m51s
The comment above the charter writer said flipping that call back to
putArray was "proven load-bearing" and pointed at two named tests. I
mutated exactly that, on the merge commit, and it SURVIVED at 1618 green -
so neither named test covers it, and neither can.

The reason is structural, not a missing test: that writer runs first
against an empty array, so putArray has nothing to replace and the two
idioms are equivalent there. No test can distinguish them. The worker
measured the same thing independently and said so in their report; the
comment was left over from the pre-fix state, where the mutation being
described was a DIFFERENT one (a later writer destroying the charter).

A comment that names tests which do not cover the line is worse than no
comment. The next person mutates the line, sees green, and concludes the
tests are broken.

The corrected version states what each cell actually measures: this line
is unpinnable and why, the later writers ARE pinned and by which test
names, and deleting this line entirely fails the charter-reaches-the-member
assertion - the hole that predates #393.
2026-09-10 20:17:21 +07:00
Dai Ha 17052bb515 Merge #471: seeded skills reach an opencode member, and instructions[] stops depending on write order (fleetd #393)
memberSkills: copied skill folders into every provisioned worktree's
.claude/skills/ and stopped there. Claude Code reads that directory
natively; opencode never does. So the feature was INERT for opencode
members rather than broken: the copy succeeded, the files were correct,
and nothing ever read them. No test failed because there was nothing to
fail - the feature worked at the only layer it implemented.

An opencode member's only channel for static guidance is the instructions[]
array in its generated config. OpenCodeLauncher now adds each seeded
skill's SKILL.md there. A folder with no SKILL.md is never delivered and
the log names it.

The second half is an ordering hazard fleet01 found by reading, and that I
then measured. Three writers append to instructions[]: the role charter,
the seeded skills, and the IDE rules. The charter used putArray
(CREATE-OR-REPLACE) while the other two used withArray (get-or-create).
That was safe only because the charter ran first against an empty array -
an undeclared constraint that nothing tested. Measured on the earlier
merge: making the skills writer use putArray left 1603 tests green while
silently deleting the charter entry, so an opencode member would launch
with no role contract at all. Worse than the bug being fixed, and
invisible.

All three writers now use withArray, and three tests pin the array's
CONTENTS (never its size - a size assertion passes when putArray swaps two
entries for two others) across the combinations that matter:
charter-only, charter+ide, charter+skills+ide.

WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said
the missing axis was the COMBINATION of writers, then revised that to a
stronger claim: that nothing asserted the charter reaches an opencode
member at all. I checked the second claim on this merge and it is FALSE.
Deleting the charter writer outright fails 6 tests, and 3 of those existed
before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet,
roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and
aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a
member was already pinned.

So their FIRST diagnosis was the right one. The surviving mutation did not
delete the charter write; it made a LATER writer replace the whole array.
Every pre-existing test had exactly one writer active, and with one writer
putArray and withArray are indistinguishable. The gap was never "is the
charter delivered" - it was "are two writers ever active at once", which is
the combination axis. Recording this because the stronger claim is the more
quotable one and it would have sent the next reader looking for a hole that
is not there.

One honest residue: the charter writer's own idiom cannot be pinned.
Flipping it back to putArray leaves the suite green, and always will,
because it runs first against an empty array where the two idioms are
equivalent. The edit removes an undeclared constraint for the next person
to add a writer; it is not a change any test can detect. The comment in
the source claimed two named tests cover it - that was wrong, and the
commit after this one corrects it rather than leaving a false claim beside
the code.

What this does NOT do: opencode has no equivalent of Claude Code's Skill
tool, so the content is static system-prompt text present from spawn, not
something a member can invoke by name. This closes the DELIVERY gap and
cannot close the ACTIVATION gap. That residue is opencode's design.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:17:21 +07:00
Dai Ha e95ed99bf7 fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.

fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.

No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
2026-09-10 20:09:13 +07:00
Dai Ha 1477e4358a Merge #472: one canonical tool-name set, and charters are checked against it (fleetd #469)
CI / contract (push) Successful in 57s
CI / build (push) Successful in 2m0s
A role charter is free text in config that tells a member which tools to
call, and nothing checked that those tools exist. A charter naming
bridge_send - a name CB-634 removed - started the daemon cleanly, and the
member found out at run time by calling something that was not there.

#464 shipped a test for this, but it wrote its own charter into a @TempDir,
so nothing anyone put in the real config could fail it. That was a defect
in my acceptance criteria, not in that work.

Now: FleetTool is one enum of the 11 registered wire names, and every
reader goes through it.

- FleetMcp's schema builders pass FleetTool.X.wireName() instead of a
  literal.
- FleetMcp's constructor asserts at startup that what it registers with the
  SDK equals FleetTool.wireNames() exactly, in both directions.
- toolAction(String, Map) resolves arbitrary wire input against
  FleetTool.byWireName() and keeps its run-time throw, which is necessary -
  network input has no closed compile-time form. It then hands off to
  authzAction(FleetTool, Map), a switch over the enum with NO default, so a
  new tool is a compile error at that layer.
- CharterToolSurface lives in mcp, not config, and Fleetd.main calls it
  right after validateAll(). Config must not depend on the MCP server:
  config loads before the server exists.

My brief undercounted the problem and the worker corrected it. I said there
were two tool-name inventories plus a test fixture. There were FIVE: the
registrations, the authz switch, and three separate source-text scrapes of
FleetMcp.java in CharterToolSurfaceTest, FleetMcpAuthzTest and
McpContractDocTest - none of which the ticket mentioned. Fixing those three
was required, not scope creep: once the literals moved into FleetTool their
regexes matched zero names, so one would have failed on its vacuity guard
and the other two would have gone quietly vacuous. Inventories after: one.

Verified independently on origin/main before accepting the wider diff:
three test files did read FleetMcp.java as source text, with a control file
at zero to prove the search discriminated.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:59:03 +07:00
Dai Ha 6c2d6e93cb fleetd #469: one canonical FleetTool set backs registration, authz and charter checks
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m37s
FleetConfig.validateCharters() only checked that a charter key is a role
wire name and its text is non-blank; #464's CharterToolSurfaceTest compared
charter text against the registered tool surface, but wrote its own charter
into a @TempDir fixture, so nothing anyone wrote into the live fleetd.yaml
could ever fail it.

Add FleetTool, an enum in dev.ltms.fleet.mcp holding the one canonical set
of registered tool wire names. FleetMcp's tool schemas now derive their
names from it, its constructor asserts at startup that what it actually
registers with the SDK equals FleetTool.wireNames() exactly, and its authz
dispatch (toolAction/authzAction) resolves the wire string against FleetTool
before switching on the enum itself with no default -- adding a tool without
pinning its Authz.Action is now a compile error, not just a test gap.

Add CharterToolSurface (mcp package, not config -- config loads before the
MCP server exists) and call it from Fleetd.main right after
cfg.validateAll(), so a charter naming a tool the server does not register
refuses the daemon's startup, naming both the charter key and the unknown
tool. FleetdStartupValidationTest proves this through Fleetd.main itself
against a live-shaped config fixture (bridge_send, CB-634's own removed
name).

CharterToolSurfaceTest, FleetMcpAuthzTest and McpContractDocTest each kept
an independent regex scrape of FleetMcp.java's source for the registered
side of their own comparison -- three more copies of the same list nothing
tied together. All three now read FleetTool.wireNames() instead.

Proved canonical by removal: deleting FleetTool.ACK while ackTool() still
referenced it broke mvn compile in two places (FleetMcp.java:940,:1898);
registering a schema under a literal not backed by FleetTool
("fleet_ack_v2") failed FleetMcp's new startup assertion in every test that
constructs it (12 errors, IllegalStateException at FleetMcp.<init>). Both
reverted before this commit.
2026-09-10 19:55:02 +07:00
Dai Ha 789b6a8716 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
CI / contract (push) Successful in 1m56s
CI / build (push) Successful in 2m4s
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things on the record because they are judgement calls, not facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this codebase
  reports a spawn success back to this class, so "it started working again"
  cannot be observed here. A base cooldown of quiet is the best available
  evidence, and the class doc says that plainly instead of implying the
  stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - the same
formula with multiplier 1.0 and a ceiling equal to the base, which collapses
to the original flat "now + cooldown". Every existing call site keeps its
shape.

The second commit exists because my own mutation on the first merge found
the wiring unpinned: putting Fleetd.main back on the flat constructor left
all 1608 tests green, so the factory was pinned and the decision to use it
was not. FleetdBackendQuarantineWiringTest closes that, following the five
existing *WiringTest files rather than inventing an idiom. It checks source
TEXT, and its class doc says so: it does not prove the call executes, and it
cannot tell "wrong factory" apart from "renamed the anchor" - both fail the
same assertion. That limit is real and recorded rather than papered over.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:51:00 +07:00
Dai Ha 01462c9695 fleetd #466 follow-up: pin main's choice of the escalating quarantine factory
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.

Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
2026-09-10 19:47:59 +07:00
Dai Ha 8b4ff78546 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things I want on the record because they are judgement calls, not
facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this
  codebase reports a spawn success back to this class, so "it started
  working again" cannot be observed here. A base cooldown of quiet is the
  best available evidence. The class doc says this plainly rather than
  implying the stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred classification question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and is deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - it is the
same formula with multiplier 1.0 and a ceiling equal to the base, which
collapses to the original flat "now + cooldown". Every existing call site
keeps its shape.

Build number and my own mutation results are reported in the PR and on the
ticket, measured on this merge commit rather than on the branch.
2026-09-10 19:40:15 +07:00
Dai Ha 5a467e1f8b fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
CI / build (pull_request) Successful in 1m34s
CI / contract (pull_request) Successful in 1m36s
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.

The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.

This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
2026-09-10 19:33:09 +07:00
59 changed files with 7671 additions and 736 deletions
+226
View File
@@ -0,0 +1,226 @@
---
name: handover
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
writes a handover file, and the new session reads that file and carries on.
**There are two ways to hand off, and the file is the same either way.**
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
session and points it at the file. This always works.
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
checks the file, clears your pane, and tells the fresh session to read it. This needs
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
Nothing else in this skill changes between the two. Only who performs the swap changes.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
**Run the three steps in this order. The order is not a style choice — the wrong order is
refused.**
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
`handoverPath` you must write to. Nothing has happened to your pane yet.
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
own workspace before it hands it to you. The daemon and your pane can run in different
directories, so a path you resolve yourself can point at a different file from the one the daemon
will check.
**Check that the path is ignored by git before you write to it (#491).** A relative
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
will show up as untracked content in that repo. The file is a snapshot of live state and must
never be committed, so tell the operator rather than committing it or silently editing their
`.gitignore`.
2. **Write the handover file at that path**, following sections 1–10 above.
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
`open` request. That check stops a leftover file from an earlier session being accepted as this
one's handover. So writing the file first and then calling `open` — the obvious order — always
fails.
`{action: "cancel", token}` drops a pending request without rolling.
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
defaults to `true` and this is the only thing standing between a judgement call and a wiped
session.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
the operator so before you confirm.** The recovery is the same either way: the file is already
written, so the operator starts a session and points it at the file. That is why you write the
file before you confirm, and never the other way round.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
```
+72
View File
@@ -0,0 +1,72 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+6
View File
@@ -19,3 +19,9 @@
fleetd.out
fleetd/fleetd.out
logs/
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
# has no business in git history.
.handover/
+22 -59
View File
@@ -96,8 +96,10 @@ below are the procedure — run them in order, every task, not only the big ones
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
forge MCP server it appears to have holds a blocked credential and fails every call, and a piped
command (`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a
fact. Its injected repo-scoped `GITEA_TOKEN` is a different credential and does work, so a worker
reporting that it opened its own PR is reporting something it really can do.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -142,6 +144,7 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
### Lead ↔ lead — coordinate, never delegate
@@ -199,8 +202,11 @@ simply complies has thrown away the reason there are two of you.
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
yours, whatever you see. And **a mounted tool is not a working tool** — the forge MCP server you
may find there holds a deliberately blocked credential and fails every call, by design. That is
not your only forge route, and the two must not be confused: the repo-scoped `GITEA_TOKEN` the
daemon injects into your environment does work, and using it to open your own PR is part of the
job. A blocked MCP tool is never a reason to skip that step.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
@@ -257,8 +263,11 @@ must obey belongs in the charter, not here.
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace) and
`fleets-status` (report every fleet that shares one LavinMQ instance).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
@@ -292,60 +301,14 @@ must obey belongs in the charter, not here.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)
+50
View File
@@ -88,6 +88,47 @@ bind:
# backoffMs: 60000
# quietNudgeCap: 3
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
# own pane and bootstrap a fresh session against that file.
#
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
#
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no
# handoverPath refuses to start. May be relative: it then resolves against the
# CALLING lead's own fleet.leaders.<name>.cwd (falling back to the daemon's own
# working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory (fleetd #480 follow-up).
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again)
# # before sending /clear at all. If this elapses, /clear is NEVER
# # sent — a lead that never goes idle is still doing real work.
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
# # again AFTER /clear before giving up (never sends bootstrapText
# # if this elapses). A separate, second wait from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the lead once its pane settles after /clear
# leadRollover:
# handoverPath: /path/to/handover.md
# requireOperatorConfirm: true
# maxDocAgeSeconds: 3600
# turnSettleSeconds: 20
# clearSettleSeconds: 20
# bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
@@ -452,6 +493,15 @@ placement: weighted
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
#
# fleetd #466: this is now only the BASE of an escalating backoff, not a flat retry rate. A
# credential quarantined again within one base cooldown of the previous quarantine ending (still
# reporting exhausted — e.g. a weekly subscription limit that hasn't reset) backs off further:
# cooldown doubles each such time, capped at 12x this value (~6 hours at the 1800s default). A
# quarantine that starts after a base-cooldown's worth of quiet resets back to this value. Not
# configurable per se — the multiplier and ceiling are constants in BackendQuarantine, not new
# YAML keys; see its class doc for the exact formula and why there is no automatic probe to clear
# it early (the operator's own design constraint — a probe spends the quota it's measuring).
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
+228 -21
View File
@@ -10,6 +10,7 @@ import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.HerdrRouter;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -26,6 +27,7 @@ import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.mcp.CharterToolSurface;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -78,6 +80,7 @@ import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
@@ -147,7 +150,10 @@ public final class Fleetd {
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// fleetd #474: pass the charter/tool-surface check in as ConfigRef's extraValidation, so
// ConfigRef#reload() runs the same gate main() runs below, without dev.ltms.fleet.config
// gaining a dependency on dev.ltms.fleet.mcp — Fleetd is the seam that already holds both.
ConfigRef config = new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools);
// The primary/host env that launched fleetd must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
@@ -163,6 +169,17 @@ public final class Fleetd {
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
// why a name-by-name list here would have the same defect it replaces.
cfg.validateAll();
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
// looks at what the text actually names. This is the separate check that does: it asks
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
// bridge_* token a charter names is a tool this server actually registers. It cannot live
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
// that already holds both a loaded FleetConfig and the mcp package, before anything below
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
assertChartersNameOnlyRegisteredTools(cfg);
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -222,7 +239,13 @@ public final class Fleetd {
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
// about 336 times across the week) backs off further each consecutive time, capped at
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
// mechanism — BackendOutagePolicy below), and the reset.
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
@@ -248,15 +271,12 @@ public final class Fleetd {
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
boolean herdrUp = awaitHerdr(herdr);
HerdrAwaitOutcome herdrOutcome = awaitHerdr(herdr, System::nanoTime, Fleetd::sleepHerdrPoll);
boolean herdrUp = logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
@@ -569,6 +589,16 @@ public final class Fleetd {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed, so
// an upgraded daemon cannot silently acquire the ability to clear the lead's own pane.
// Unlike heartbeat above, this has no recurring scheduler of its own — nothing but an
// explicit confirm() call (wired to an MCP tool by a later ticket; nothing calls it yet)
// that passes every gate can ever schedule a roll. It does not take primaryRegistry: the
// lead terminal to roll comes from the caller of open()/confirm() (resolved by the MCP
// layer from the connection, the same way auth/CallerResolver#resolve builds a
// Principal.leader(...)), never from a single-slot lookup — see LeadRollover's class
// javadoc, fleetd #480 correction 2.
LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
@@ -664,7 +694,7 @@ public final class Fleetd {
}, outagePolicy);
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics,
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
new FleetMcp.HealthCoverageSource(() -> {
var health = config.get().health();
@@ -679,7 +709,10 @@ public final class Fleetd {
// reach. Read from the SAME snapshot leadMailbox itself opened from (cfg.coordinator()),
// not the live config.get() — coordinator wiring is already boot-time-fixed (see
// leadMailbox above), so peers follows the same rule rather than half hot-reloading.
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers());
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
// fleetd #480 Unit C: the executor behind fleet_handover — null whenever
// leadRollover: is not configured (see the leadRollover local above).
leadRollover);
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
@@ -1032,6 +1065,68 @@ public final class Fleetd {
.orElse(null);
}
/**
* fleetd #480: construct the {@link LeadRollover} executor only when {@code leadRollover:} is
* present at startup — the same presence gate {@code leadHeartbeat:} uses just above this
* call site in {@code main}. Extracted to its own factory, the same reason
* {@link #worktreeBranchLookup} and {@link #exhaustionSink} are: a unit test can call this
* directly with a fabricated {@link FleetConfig} (absent block ⇒ {@code null}, present block ⇒
* constructed) without needing the whole of {@code main}, and a source-text test on the real
* call site proves {@code main} still calls this factory rather than inlining a copy that could
* silently diverge.
*
* <p>The returned object reads every {@code leadRollover:} field fresh on each {@code open()}/
* {@code confirm()} call through {@code () -> config.get().leadRollover()} — see {@code
* ConfigRef}'s class doc Hot bullet and {@link FleetConfig.LeadRollover}'s javadoc for why that
* makes the block's fields HOT despite this presence gate being evaluated once, at startup.
*
* <p>Deliberately does NOT take {@link PrimaryRegistry}: fleetd #480 correction 2 found that a
* single-slot lookup lets one lead's {@code confirm()} clear a DIFFERENT lead's pane on a
* daemon with more than one labelled lead tab. The lead terminal to roll instead comes from
* whoever calls {@code open()}/{@code confirm()} — the later MCP-tool unit must resolve it from
* the connection and pass it in, never take it as a request field. See {@link LeadRollover}'s
* class javadoc.
*
* <p><strong>fleetd #480 follow-up:</strong> also builds the terminal → lead-workspace lookup
* {@link LeadRollover#open} needs to resolve a relative {@code handoverPath} against the
* CALLING lead's own {@code cwd} rather than the daemon's — the daemon and a lead's pane can
* have different working directories (this repo nests {@code fleetd/} inside its own root, so
* they already differ on this host). The lookup is terminal → lead name (via {@code
* liveLeadTerminals}) → that lead's {@code cwd} (via {@code config.get().fleet().leaders()}),
* and both hops are read LIVE on every call, never off a snapshot taken here: leads are
* discovered by a live tab scan ({@code LeadTabScanner}), so a map captured at construction
* time could be empty (no lead has been scanned yet) or stale (a lead added since).
*
* @param cfg the startup config snapshot — read ONCE here, only to decide whether
* to construct the object at all, exactly like {@code
* cfg.leadHeartbeat()}
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
* {@code memberAgents}), normally {@code router.leadAgents()}
* @param config the live {@link ConfigRef}, captured only inside the returned
* supplier and the workspace lookup — never dereferenced here
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
* the same {@code leads} supplier {@code main} already builds for
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
* absent from the startup config
*/
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
Supplier<Map<String, String>> liveLeadTerminals) {
if (cfg.leadRollover() == null) {
return null;
}
Function<String, String> leadWorkspace = terminal -> {
String leadName = liveLeadTerminals.get().get(terminal);
if (leadName == null) {
return null;
}
FleetConfig.Fleet fleet = config.get().fleet();
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
return leader == null ? null : leader.cwd();
};
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
}
/**
* fleetd #176: per-profile factory for {@link FleetMcp.LeadSeatSource} — how many seats a
* profile's own live LEAD session(s) hold on the same Claude subscription.
@@ -1577,12 +1672,92 @@ public final class Fleetd {
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
* via a method reference to this method) go through, so the two can never drift into checking
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
*
* @return true if herdr answered, false if it never did
* <p>Package-private so a test can call it directly the same way the other startup-report
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
* exact method reference, the same way {@code main} does above — proves the reload call site.
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
* equivalent {@code Consumer<FleetConfig>} it builds locally (it cannot see this package-private
* method from {@code dev.ltms.fleet.config}).
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
static void assertChartersNameOnlyRegisteredTools(FleetConfig cfg) {
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
}
/**
* How {@link #awaitHerdr} ended (fleetd #498). The old code returned a bare {@code boolean},
* which collapsed two different facts onto the same {@code false}: the configured wait budget
* genuinely running out, and the waiting thread being interrupted possibly milliseconds in.
* Those need different operator messages — see {@link #logHerdrWaitOutcomeAndShouldReap} — so
* this is a third state, not a better number (the same shape fleetd #497 named). Never treat
* {@link #INTERRUPTED} as if it were {@link #DEADLINE_PASSED}: only the latter means herdr was
* actually given the full {@link #HERDR_WAIT_SECONDS} and still failed to answer.
*/
enum HerdrWaitResult {
/** herdr answered {@code ping} before the deadline. */
ANSWERED,
/** the configured {@link #HERDR_WAIT_SECONDS} budget elapsed with no answer. */
DEADLINE_PASSED,
/**
* the waiting thread was interrupted before the budget ran out — a different event from
* {@link #DEADLINE_PASSED} and must never be reported as "did not answer within Ns".
*/
INTERRUPTED
}
/**
* The outcome of one {@link #awaitHerdr} call, carrying the MEASURED elapsed wait time
* alongside {@link #result}. {@code elapsedNanos} is always measured against the {@code nanos}
* supplier passed to {@link #awaitHerdr} — never assume it equals the configured budget, the
* same defect fleetd #494 already fixed once in {@code LeadRollover}.
*/
record HerdrAwaitOutcome(HerdrWaitResult result, long elapsedNanos) {}
/**
* The real per-poll wait {@link #main} passes to {@link #awaitHerdr}: sleep
* {@link #HERDR_WAIT_POLL_MILLIS}, and on interruption re-set the thread's interrupt flag
* rather than throwing — {@link #awaitHerdr} detects an interruption by checking {@link
* Thread#isInterrupted()} right after this returns, so a poller that swallowed the flag
* instead of restoring it would make that check silently miss the interruption.
*/
private static void sleepHerdrPoll() {
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
}
}
/**
* Poll herdr's {@code ping} until it answers, the configured {@link #HERDR_WAIT_SECONDS}
* budget elapses, or the waiting thread is interrupted (CB-504, fleetd #498).
*
* <p>{@code nanos} and {@code poller} are required parameters with no defaulted overload
* (fleetd #415's shape: a defaulted overload is a silent survivor a green suite would vouch
* for) — the previous version read {@link System#nanoTime()} and called {@link Thread#sleep}
* directly, so nothing could drive it from a test. The one production call site in {@link
* #main} passes {@code System::nanoTime} and {@link #sleepHerdrPoll}.
*
* @param nanos a monotonic elapsed-time clock, e.g. {@code System::nanoTime} — never a
* wall-clock source, since only elapsed time (not a timestamp) is measured here
* @param poller called once per failed ping while the budget remains; must, on an
* {@link InterruptedException}, re-set the thread's interrupt flag rather than
* throw or swallow it — this method's interruption check reads that flag right
* after {@code poller.run()} returns
* @return the outcome and the measured elapsed wait time — see {@link HerdrAwaitOutcome}
*/
static HerdrAwaitOutcome awaitHerdr(HerdrClient herdr, LongSupplier nanos, Runnable poller) {
long start = nanos.getAsLong();
long deadline = start + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
@@ -1590,25 +1765,57 @@ public final class Fleetd {
if (waited) {
log.info("herdr is up");
}
return true;
return new HerdrAwaitOutcome(HerdrWaitResult.ANSWERED, nanos.getAsLong() - start);
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
if (nanos.getAsLong() >= deadline) {
return new HerdrAwaitOutcome(HerdrWaitResult.DEADLINE_PASSED, nanos.getAsLong() - start);
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
poller.run();
if (Thread.currentThread().isInterrupted()) {
return new HerdrAwaitOutcome(HerdrWaitResult.INTERRUPTED, nanos.getAsLong() - start);
}
}
}
}
/**
* Log the right message for {@code outcome} — never the configured {@link #HERDR_WAIT_SECONDS}
* budget alone, always the measured elapsed time next to it — and say whether {@link #main}
* should now reap orphan worker panes (fleetd #498).
*
* <p>Extracted out of {@link #main} so this decision is drivable from a test: {@link #main}
* boots the whole daemon and cannot itself be run in a unit test, but this is the exact,
* unmodified code {@link #main} calls for the decision, not a re-derivation of it.
*
* @return true only for {@link HerdrWaitResult#ANSWERED} — orphan workers are reaped only
* then, exactly as before this ticket
*/
static boolean logHerdrWaitOutcomeAndShouldReap(HerdrAwaitOutcome outcome) {
long elapsedMillis = TimeUnit.NANOSECONDS.toMillis(outcome.elapsedNanos());
if (outcome.result() == HerdrWaitResult.ANSWERED) {
return true;
}
if (outcome.result() == HerdrWaitResult.DEADLINE_PASSED) {
log.warn("herdr did not answer within the configured wait (configured={}s elapsed={}ms) "
+ "— starting anyway; /healthz will report degraded until it comes up. Orphaned "
+ "worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS, elapsedMillis);
return false;
}
// HerdrWaitResult.INTERRUPTED — a different fact from DEADLINE_PASSED (fleetd #498): the
// wait was cut short, not exhausted, and must never be reported as "did not answer within
// Ns" — that claim would be false and would send an operator to debug herdr for nothing.
log.warn("herdr wait was interrupted before the configured wait ran out (configured={}s "
+ "elapsed={}ms) — starting anyway; /healthz will report degraded until it comes "
+ "up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS, elapsedMillis);
return false;
}
private Fleetd() {
}
}
@@ -41,7 +41,16 @@ public final class Authz {
*/
COORD_READ,
/** Scrape the metrics endpoint. */
METRICS
METRICS,
/**
* Drive the lead-rollover executor ({@code fleet_handover}: open/confirm/cancel a
* self-replace, fleetd #480 Unit C). Primary-only, same as {@link #SPAWN}/{@link #STOP}/
* {@link #DRAIN} — and, unlike those, the terminal it acts on is never even an argument:
* {@code LeadRollover#open}/{@code #confirm} are always called with the CALLER's own
* connection-resolved terminal (see {@code dev.ltms.fleet.lead.LeadRollover}'s class
* javadoc, fleetd #480 correction 2), so a primary can only ever roll itself.
*/
HANDOVER
}
/**
@@ -60,7 +69,7 @@ public final class Authz {
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
@@ -244,7 +244,15 @@ public final class CallerResolver {
// already names what happens if that case is handed the primary role: a worker→primary
// escalation. So an unresolved caller is refused (ANONYMOUS — the same clean, already-tested
// "authenticated as nothing" outcome used everywhere else in this method), never promoted.
return isLoopback(remoteAddr) && c.resolved() ? Principal.primary(c.pid()) : Principal.anonymous();
//
// fleetd #505: the OTHER way a real pid can wrongly reach here with a null terminal — not a
// failed lsof lookup, but a herdr error partway through PaneLocator's pane scan. c.resolved()
// says nothing about that; it only tests the lsof sentinel (by design — see
// ConnectionIdentity.Caller#resolved). c.scanComplete() is the separate signal: a scan that
// could not check every pane must not be read as "checked everywhere, no match" — the pane it
// could not check might have been the caller's own. So both must hold before this promotes.
return isLoopback(remoteAddr) && c.resolved() && c.scanComplete()
? Principal.primary(c.pid()) : Principal.anonymous();
}
private boolean presentedTokenMatches(String authorizationHeader) {
@@ -11,6 +11,7 @@ import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Consumer;
import java.util.function.Supplier;
/**
@@ -61,7 +62,22 @@ import java.util.function.Supplier;
* both read, so a reload that arms or disarms a profile's usage-limit detection takes effect
* on the next check with no restart. {@code errorPattern}, {@code exhaustedPattern}'s sibling
* key for backend-error (not usage-limit) classification, was deliberately left OUT of this
* fleetd #446 change and stays deferred below — the ticket scoped it out explicitly.</li>
* fleetd #446 change and stays deferred below — the ticket scoped it out explicitly.
* {@code leadRollover:} (fleetd #480) joined this class whole, the same shape as
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
* than capturing them into fields at construction — unlike its closest
* structural cousin {@code leadHeartbeat:}, whose {@code LeadHeartbeatLoop} bakes
* {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} into final fields. The one
* restart-only edge is structural, not a stale value: {@code Fleetd.java} decides whether to
* construct the {@code LeadRollover} object at all off the startup snapshot (the same
* presence gate {@code leadHeartbeat:} uses), so a block ADDED where it was absent at boot
* needs a restart before anything exists to call — the same fact already true of adding a
* brand-new {@code profiles:} entry.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
@@ -165,17 +181,20 @@ import java.util.function.Supplier;
*
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
* {@code models:} was added as deferred, and again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere.</strong>
* {@code FleetConfig} has 25 top-level record components: 5 cold, 13 deferred, 3 split, 4
* hot-excluded. Four of them are named nowhere in this file, and the reason is the same for all
* four: {@code placement}, {@code memberCredentials}, {@code memberLoginShell} and {@code models}
* are <strong>hot</strong> and correctly absent — all four are read live off {@code config.get()}
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
* {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198, 205, 729}
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way, through the
* Hot bullet's {@code models:} paragraph), so a reload takes effect on the next spawn (or, for
* {@code models}, the next reported status) with no entry needed here.
* {@code models:} was added as deferred, again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere, and again after
* {@code leadRollover:} was added (fleetd #480).</strong>
* {@code FleetConfig} has 26 top-level record components: 5 cold, 13 deferred, 3 split, 5
* hot-excluded. Five of them are named nowhere in this file, and the reason is the same for all
* five: {@code placement}, {@code memberCredentials}, {@code memberLoginShell}, {@code models} and
* {@code leadRollover} are <strong>hot</strong> and correctly absent — all five are read live off
* {@code config.get()} (placement through the {@code CompositePeerLauncher} supplier the Hot bullet
* names; {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198,
* 205, 729} and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way,
* through the Hot bullet's {@code models:} paragraph; {@code leadRollover} through the Hot bullet's
* {@code leadRollover:} paragraph), so a reload takes effect on the next spawn (or, for
* {@code models}, the next reported status; for {@code leadRollover}, the next {@code open()}/
* {@code confirm()} call) with no entry needed here.
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
@@ -211,6 +230,24 @@ import java.util.function.Supplier;
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*
* <p><strong>fleetd #474</strong> — {@link FleetConfig#validateAll()} is not the only gate startup
* runs before a config takes effect: {@code Fleetd.main} also calls {@code
* dev.ltms.fleet.mcp.CharterToolSurface#assertChartersNameOnlyRegisteredTools}, right after {@code
* cfg.validateAll()}, to refuse a charter that names an MCP tool the server does not register. That
* check cannot live inside {@link FleetConfig} — {@code CharterToolSurface} lives in the {@code mcp}
* package because the canonical tool set ({@code FleetTool}) does, and config is loaded before the
* MCP server exists, so {@code FleetConfig} must not gain a dependency on {@code mcp}. {@link
* #reload} cannot import {@code mcp} either, for the same reason applied one layer up: {@code
* dev.ltms.fleet.config} is loaded before {@code dev.ltms.fleet.mcp} exists, same as {@code
* FleetConfig}. So this class accepts the check as a {@code Consumer<FleetConfig>} —
* {@link #extraValidation} — supplied by whichever caller already sits at the seam that holds both
* a loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}. It is invoked
* inside the same try/catch as {@code fresh.validateAll()}, so a charter that would have refused to
* boot refuses a reload too, and keeps the running config exactly like any other {@code
* validateAll()} failure. A ref built through the two-argument constructor (every test fixture that
* does not care about this check, and {@link #fixed}) gets a no-op consumer, so nothing outside
* {@code Fleetd.main} needs to know this hook exists.
*/
public final class ConfigRef implements Supplier<FleetConfig> {
@@ -255,10 +292,27 @@ public final class ConfigRef implements Supplier<FleetConfig> {
private final Path path;
private final AtomicReference<FleetConfig> current;
private final Consumer<FleetConfig> extraValidation;
/** Equivalent to the three-argument constructor with a no-op {@code extraValidation}. */
public ConfigRef(Path path, FleetConfig initial) {
this(path, initial, cfg -> { });
}
/**
* @param extraValidation run on every {@link #reload} candidate, inside the same try/catch as
* {@code fresh.validateAll()} — see the class doc's fleetd #474 note.
* {@code Fleetd.main} passes {@code
* Fleetd::assertChartersNameOnlyRegisteredTools} (a package-private
* {@code FleetConfig -> void} adapter over {@code
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), so a reload
* runs the same gate startup does without this class depending on the
* {@code mcp} package.
*/
public ConfigRef(Path path, FleetConfig initial, Consumer<FleetConfig> extraValidation) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
this.extraValidation = Objects.requireNonNull(extraValidation, "extraValidation");
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
@@ -360,6 +414,12 @@ public final class ConfigRef implements Supplier<FleetConfig> {
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
// not a longer hand-maintained list.
fresh.validateAll();
// fleetd #474: validateAll() does not cover everything startup refuses on — the charter
// tool-surface check (Fleetd.main, right after cfg.validateAll()) lives outside
// FleetConfig on purpose (see this class's doc) and is supplied here as extraValidation.
// Same try/catch as validateAll() above, on purpose: either failure must refuse the whole
// reload and keep the running config the same way.
extraValidation.accept(fresh);
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
@@ -70,12 +70,14 @@ import java.util.regex.PatternSyntaxException;
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
* @param quarantineCooldownSeconds the BASE cooldown a credential is quarantined for after a
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
* needs a restart, and a quarantine already running keeps whatever cooldown was
* live when it started.
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Since fleetd #466 this is
* only the first occurrence's length — a credential quarantined again shortly
* after this cooldown ends backs off further, up to a ceiling; see {@code
* BackendQuarantine}'s class doc. Baked once into the {@code BackendQuarantine}
* built at startup, so it is DEFERRED: changing it needs a restart, and a
* quarantine already running keeps whatever cooldown was live when it started.
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
* credentials a spawned member's pane inherits. {@code null} (the block
* omitted) blocks nothing — see {@link MemberCredentials}.
@@ -137,6 +139,9 @@ import java.util.regex.PatternSyntaxException;
* fleetd #422 added the separate on/off question — whether a configured model
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
* @param leadRollover opt-in lead rollover (fleetd #480): {@code null} ⇒ off, and no
* {@code dev.ltms.fleet.lead.LeadRollover} is constructed at all — an upgraded
* daemon never clears a lead's pane on its own initiative. See {@link LeadRollover}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -164,7 +169,23 @@ public record FleetConfig(
String memberLoginShell,
String memberSkills,
IdleSleepGuard idleSleepGuard,
Models models) {
Models models,
LeadRollover leadRollover) {
/** Back-compat form before the {@code leadRollover:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard,
Models models) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, models, null);
}
/** Back-compat form before the {@code models:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
@@ -1308,6 +1329,104 @@ public record FleetConfig(
}
}
/**
* Opt-in lead rollover (fleetd #480): a lead that decides it is ready to be replaced writes a
* handover file, then asks fleetd to clear its own pane and bootstrap a fresh session against
* that file. Config + a pure decision/verification layer only — see
* {@code dev.ltms.fleet.lead.LeadRollover} for the executor this block feeds, and
* {@code fleet_*} tool wiring is a later ticket.
*
* <p>Deliberately opt-in ({@code null} ⇒ off, exactly like {@code leadHeartbeat:}): absent this
* block, {@code Fleetd.java} never constructs a {@code LeadRollover} object at all, so an
* upgraded daemon cannot silently acquire the ability to clear the lead's own pane. Even once
* present, nothing but an explicit {@code confirm()} call can ever roll a pane — there is no
* timer, heartbeat or timeout anywhere in this feature that fires one on its own; see that
* class's javadoc.
*
* <p><strong>Hot, not deferred</strong> (see {@code ConfigRef}'s class doc): every field below
* is read live, through a {@code Supplier<LeadRollover>} the same {@code () -> config.get().x()}
* shape {@code fleet}/{@code placement}/{@code models} already use, so an edit to any of the
* six fields takes effect on the very next {@code open()}/{@code confirm()} call once the block
* has been present since startup — nothing here is captured into a frozen field the way {@code
* leadHeartbeat}'s {@code idleAfterNanos}/{@code backoffMs}/{@code quietNudgeCap} are. The one
* restart-only edge left is structural, not a stale value: {@code Fleetd.java} decides whether
* to construct the {@code LeadRollover} object at all off the startup snapshot, the same gate
* {@code leadHeartbeat:} uses, so a block ADDED where it was absent at boot needs a restart
* before anything exists to call — the same fact already true of adding a brand-new
* {@code profiles:} entry.
*
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
* {@code confirm()} validates every gate and schedules the roll. {@code
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
*
* @param handoverPath required when this block is present — where the handover file a fresh
* lead session reads must live. There is no sane non-null default for an
* operator-specific path, so a present block with a {@code null}/blank
* {@code handoverPath} is refused at config load; see
* {@link #validateLeadRollover()}. May be relative: {@code
* dev.ltms.fleet.lead.LeadRollover#open} resolves a relative path against
* the CALLING lead's configured {@code fleet.leaders.<name>.cwd}, falling
* back to the daemon's own working directory ({@code
* System.getProperty("user.dir")}) when that lead has none configured —
* never against whatever directory the daemon process happens to have been
* started in for its own sake. An absolute path is used unchanged.
* @param requireOperatorConfirm default {@code true} — {@code confirm()} refuses unless the
* caller also passes {@code operatorConfirmed: true}. Set {@code false} to
* let the three handover-file checks alone gate the roll.
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
* than this many seconds, so a stale leftover from an earlier rollover
* attempt can never be mistaken for a fresh one.
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report injectable again)
* before sending {@code /clear} at all. See the paragraph above.
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
* report an injectable state again after {@code /clear} before giving up. A
* roll that times out here never sends {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* lead's pane once it settles after {@code /clear}, telling the fresh
* session where to read the handover and carry on. Left {@code null} here
* when the operator configures none: the default sentence cannot be built
* at construction time because it must name the path AFTER {@code
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
* handoverPath} against the calling lead's workspace, which this record has
* no way to know — see {@link #bootstrapTextFor(String)}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
Integer clearSettleSeconds, String bootstrapText) {
public LeadRollover {
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
}
/**
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
* sentence built from {@code resolvedHandoverPath}.
*
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
* #open} already resolved — never the raw configured {@link
* #handoverPath}, which may still be relative and would name a
* directory the fresh lead session's own pane cannot resolve
*/
public String bootstrapTextFor(String resolvedHandoverPath) {
return bootstrapText != null
? bootstrapText
: "Fresh lead session: read the handover file at " + resolvedHandoverPath
+ " and carry on from there.";
}
}
/**
* Watch {@code fleetd.yaml} and re-read it when it changes (CB-559).
*
@@ -1723,7 +1842,7 @@ public record FleetConfig(
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
"idleSleepGuard", "models");
"idleSleepGuard", "models", "leadRollover");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -2413,10 +2532,14 @@ public record FleetConfig(
// empty Models would be a no-op for validateModels() either way, since an empty allow-list
// already means "check nothing", so there is nothing to gain and one more null check to
// avoid by leaving it exactly as configured.
// leadRollover is left as-is, like leadHeartbeat above: null is "off", and LeadRollover's
// own compact constructor defaults the fields of a block that IS present. Defaulting it
// here would construct a LeadRollover object (via Fleetd.java's presence gate) for every
// config that never mentioned it.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
idleSleepGuard, models);
idleSleepGuard, models, leadRollover);
}
/**
@@ -2499,6 +2622,26 @@ public record FleetConfig(
+ "lead tabs cannot be confused.");
}
/**
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
* every other field on {@link LeadRollover}, which {@link LeadRollover}'s own compact
* constructor already defaults — so this is the one field that must be refused at load rather
* than silently defaulted to something that would never match a real handover file.
*
* @throws IllegalStateException when {@code leadRollover} is present but {@code handoverPath}
* is {@code null} or blank
*/
public void validateLeadRollover() {
if (leadRollover == null) {
return;
}
if (leadRollover.handoverPath() == null || leadRollover.handoverPath().isBlank()) {
throw new IllegalStateException("refusing to start: leadRollover.handoverPath is "
+ "required when leadRollover: is present.");
}
}
/** Case-insensitive prefix test that tolerates a null/blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
@@ -1,6 +1,8 @@
package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashSet;
import java.util.List;
@@ -36,6 +38,8 @@ import java.util.Set;
*/
public final class PaneLocator {
private static final Logger log = LoggerFactory.getLogger(PaneLocator.class);
/**
* Bound on how many ancestor generations {@link #ancestorsOf} walks. This runs on every MCP
* call, so a cycle or a pathologically deep process tree must not hang identity resolution;
@@ -73,22 +77,46 @@ public final class PaneLocator {
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane on any searched daemon owns it (e.g. the caller is the
* primary, or off-host).
* The outcome of a {@link #terminalForPid} scan: the {@code terminal_id} of the agent pane
* whose process tree contains the pid ({@link #terminal} is {@code null} if none matched),
* and whether the scan that produced that answer ran to completion on every daemon searched.
*
* <p>{@link #complete} is {@code false} exactly when some {@code pane.process_info} call
* failed and, despite that, no pane was ever found to own the pid. In that case a {@code null}
* {@link #terminal} means "could not tell", not "definitely not a worker" — fleetd #505: a
* transient herdr error on the very pane that <em>does</em> own the caller's pid must not read
* as a clean negative and fall through to {@code Principal.primary}, the same way #317's
* {@code Caller.resolved()} already guards a failed lsof lookup. Callers ({@code
* ConnectionIdentity}, {@code CallerResolver}) must refuse rather than promote on an incomplete
* scan.
*
* <p>When a pane genuinely owns the pid, {@link #complete} is {@code true} regardless of
* whether some other, unrelated pane failed to answer earlier in the same scan — a positive
* match is definitive and does not need every pane to have been checked (a pane that "vanished
* mid-scan" but was never the match is still a clean, complete result).
*/
public String terminalForPid(long pid) {
public record Lookup(String terminal, boolean complete) {
private static final Lookup NOT_FOUND = new Lookup(null, true);
}
/**
* Resolve {@code pid} to the agent pane whose process tree contains it, across every searched
* herdr daemon. See {@link Lookup} for how to read a {@code null} terminal.
*/
public Lookup terminalForPid(long pid) {
if (pid <= 0) {
return null;
return Lookup.NOT_FOUND;
}
Set<Long> ancestry = ancestorsOf(pid);
for (HerdrClient herdr : herdrs) {
String terminal = terminalForPid(herdr, ancestry);
if (terminal != null) {
return terminal;
boolean complete = true;
for (int i = 0; i < herdrs.size(); i++) {
Lookup outcome = scan(herdrs.get(i), i, herdrs.size(), ancestry);
if (outcome.terminal() != null) {
return outcome; // a definite match — no need to finish checking other clients
}
complete = complete && outcome.complete();
}
return null;
return new Lookup(null, complete);
}
/**
@@ -117,31 +145,51 @@ public final class PaneLocator {
return ancestry;
}
private static String terminalForPid(HerdrClient herdr, Set<Long> ancestry) {
/** Whether a pane owns one of the scanned pid's ancestors, or the check of it failed outright. */
private enum Ownership { OWNS, DOES_NOT_OWN, UNKNOWN }
private static Lookup scan(HerdrClient herdr, int clientIndex, int clientCount, Set<Long> ancestry) {
boolean complete = true;
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsAnyOf(herdr, paneId, ancestry)) {
return pane.path("terminal_id").asText(null);
if (paneId == null) {
continue;
}
Ownership owns = paneOwnsAnyOf(herdr, clientIndex, clientCount, paneId, ancestry);
if (owns == Ownership.OWNS) {
return new Lookup(pane.path("terminal_id").asText(null), true);
}
if (owns == Ownership.UNKNOWN) {
complete = false;
}
}
return null;
return new Lookup(null, complete);
}
private static boolean paneOwnsAnyOf(HerdrClient herdr, String paneId, Set<Long> ancestry) {
private static Ownership paneOwnsAnyOf(HerdrClient herdr, int clientIndex, int clientCount,
String paneId, Set<Long> ancestry) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
// fleetd #505: this used to be read as a clean "does not own it" (the pane vanished
// mid-scan, just skip it) — one boolean carrying two different facts. It is UNKNOWN
// now: if THIS pane is the one that owns the pid, the caller must not be told "no pane
// owns it", because that reads as a real primary and is promoted under loopback-trust.
log.warn("pane.process_info failed for pane {} on herdr client {} of {} during a "
+ "pid-owner scan — treating it as \"could not tell\", not a clean "
+ "negative (fleetd #505): {}",
paneId, clientIndex + 1, clientCount, e.getMessage());
return Ownership.UNKNOWN;
}
if (ancestry.contains(info.path("shell_pid").asLong(-1))) {
return true;
return Ownership.OWNS;
}
for (JsonNode p : info.path("foreground_processes")) {
if (ancestry.contains(p.path("pid").asLong(-1))) {
return true;
return Ownership.OWNS;
}
}
return false;
return Ownership.DOES_NOT_OWN;
}
}
@@ -15,6 +15,7 @@ import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
import java.util.stream.Collectors;
@@ -94,6 +95,14 @@ public final class Injector {
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
/**
* Wall-clock source for the readiness-grace elapsed time logged in {@link #onStatus} (fleetd
* #501). Production constructors default this to {@code System::currentTimeMillis}; the
* package-private constructors below take it explicitly so a test can supply a stub whose
* advance does not track {@link #POLL_INTERVAL_MILLIS} — copying the shape {@code LeadRollover}
* already uses for the same purpose.
*/
private final LongSupplier nowMillis;
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
@@ -125,20 +134,34 @@ public final class Injector {
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this(agents, turnListener, ready, forget, System::currentTimeMillis);
}
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, LongSupplier nowMillis) {
this.agents = agents;
this.router = null;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
}
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this(router, turnListener, ready, forget, System::currentTimeMillis);
}
/** Full constructor — for tests: an injectable wall-clock supplier (fleetd #501). */
Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget, LongSupplier nowMillis) {
this.agents = null;
this.router = router;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
}
private AgentControl agentsFor(String target) {
@@ -209,6 +232,7 @@ public final class Injector {
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int unknownSincePostTurn; // the same, for the post-turn housekeeping phase (fleetd #306)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
long notReadySinceMillis; // wall-clock time of the FIRST non-ready sample in the current notReadySincePoll streak (fleetd #501); reset alongside it
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
@@ -302,6 +326,7 @@ public final class Injector {
t.unknownSinceTurn = 0;
t.unknownSincePostTurn = 0;
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
@@ -353,6 +378,7 @@ public final class Injector {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
try {
agentsFor(target).send(target, p.text());
t.queue.poll();
@@ -370,23 +396,52 @@ public final class Injector {
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
for (Pending pending : notReady) {
pending.state = Pending.State.NOT_DELIVERED;
} else if (p != null) {
// fleetd #501: stamp the wall-clock time of the FIRST non-ready sample in
// this streak, so the expiry log below can print how long the target
// actually sat non-ready — not just how many polls that took.
if (t.notReadySincePoll == 0) {
t.notReadySinceMillis = nowMillis.getAsLong();
}
if (++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
for (Pending pending : notReady) {
pending.state = Pending.State.NOT_DELIVERED;
}
// fleetd #501: t.notReadySincePoll — the loop's own counter, already in
// scope — is printed here instead of the READINESS_GRACE_POLLS constant.
// On this branch the counter has JUST reached the threshold, so the two
// agree by construction and no test can tell them apart. Printed anyway:
// it gives this line one source of truth instead of two, so a later
// change to the loop above cannot leave this message reporting a number
// the loop no longer produces.
//
// elapsedMillis is a different case: it is NOT equal-by-construction to
// the truth. notReadySincePoll only increments on a sample that reaches
// this branch (p != null, not ready) — a poll that misses that condition
// advances real time without advancing the counter — and this loop's real
// period is not guaranteed to equal POLL_INTERVAL_MILLIS (load, or a host
// sleep, can widen the real gap far past it). READINESS_GRACE_POLLS *
// POLL_INTERVAL_MILLIS / 1000 is arithmetic on two constants, not a
// measurement, so it stays here only as the labelled CONFIGURED budget,
// never presented as elapsed time.
long elapsedMillis = nowMillis.getAsLong() - t.notReadySinceMillis;
log.warn("readiness grace for {} expired after {} polls (configured={} "
+ "polls/{}s elapsed={}ms): target never became "
+ "deliverable, so failing {} queued message(s) that "
+ "never reached its pane",
target, t.notReadySincePoll, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, elapsedMillis,
notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
}
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
@@ -0,0 +1,649 @@
package dev.ltms.fleet.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
* clears the lead's own pane and bootstraps a fresh session against that file.
*
* <p>This is the executor only. Nothing in this ticket wires an MCP tool onto {@link #open}/
* {@link #confirm}/{@link #cancel} — that is a separate, later unit; until it lands, nothing calls
* this class at all.
*
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
* ever started, and a refusal return value that claimed nothing had happened.
*
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
* at all — it only records that the request is approved and hands a one-shot continuation to
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
* once the calling turn has ended, in this order:
* <ol>
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
* away live context — exactly the failure this correction exists to prevent.</li>
* <li>{@code agents.send(lead, "/clear")}</li>
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
* see {@link #waitForClearPickupAndSettle})</li>
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
* </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
*
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
* that first defined this class required "no timer, no scheduler, no background thread" so that
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
* continuationRunner} launches a single-shot task that exists only because one specific,
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
* slightly later than the method return, as a continuation of the same approved request, rather
* than entirely inside the method body.</strong>
*
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
* correction.</strong> The first version resolved the pane to clear via {@code
* PrimaryRegistry#primaryTerminal()}. That is correct for a background loop with no caller (see
* {@code dev.ltms.fleet.msg.LeadHeartbeatLoop}), but wrong here and a violation of this project's
* own charter invariant 3 — "identity comes from the connection, never an argument." This daemon
* can hold more than one labelled lead tab (see {@code LeadLauncher}'s fleetd #359 two-reading
* dead-tab cleanup), so a single-slot lookup lets lead X's {@link #confirm} clear lead Y's pane: an
* unrecoverable loss of someone else's context, and a different lead's context at that. Both
* {@link #open} and {@link #confirm} now take the lead's terminal id as a parameter instead —
* {@link #open} stores it on the {@link PendingRollover}, and {@link #confirm} refuses with {@link
* RefusalReason#NOT_YOUR_ROLLOVER} unless the caller's terminal matches the one {@link #open}
* recorded. <strong>The terminal id passed to both methods must come from the MCP layer's own
* connection-based caller resolution — the same source {@code auth/CallerResolver#resolve} uses to
* build a {@code Principal.leader(...)} (see its use of {@code ConnectionIdentity.Caller#terminal})
* — never a value the client supplies or chooses.</strong> The later MCP-tool unit that wires
* {@link #open}/{@link #confirm} must pass the resolved caller terminal, not a request field.
*
* <p><strong>A relative {@code handoverPath} resolves against the CALLING lead's workspace, never
* the daemon's own cwd — a fleetd #480 follow-up.</strong> The daemon and a lead's own pane can
* have different working directories (this repo nests {@code fleetd/} inside its own root, so the
* daemon's cwd and {@code fleet.leaders.<name>.cwd} already differ on this host). {@link #open}
* resolves {@code cfg.handoverPath()} to an ABSOLUTE path exactly once — against {@code
* leadWorkspace.apply(leadTerminal)} when that lookup returns a non-null, non-blank workspace, and
* against {@code System.getProperty("user.dir")} otherwise (the same fallback {@code
* LeadLauncher#launch} already uses for a lead with no configured {@code cwd}) — and stores only
* that absolute path on {@link PendingRollover}. Every later read of {@code
* PendingRollover#handoverPath()} (the freshness/exists/empty checks in {@link #checkHandover},
* the value {@code FleetMcp} hands back to the lead in the {@code open} response so it knows where
* to WRITE the file, and {@link FleetConfig.LeadRollover#bootstrapTextFor} which names it in the
* text typed into the fresh session) therefore already sees the resolved absolute form and never
* needs to resolve anything itself.
*/
public final class LeadRollover {
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
static final long SETTLE_POLL_MS = 250;
/**
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
* before releasing rather than wedging the roll — the same constant and the same
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
* nudging again — so 8 polls produce 7 nudges, not 8.
*/
static final int PICKUP_GRACE_POLLS = 8;
/**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
*
* @param leadTerminal the lead pane that opened this request — the only terminal that may
* later {@link #confirm} it (see {@link RefusalReason#NOT_YOUR_ROLLOVER})
* @param handoverPath the ABSOLUTE, resolved handover path — never the raw configured value,
* which may have been relative. {@link #open} resolves it once, against the
* calling lead's workspace, before storing it here; see this class's
* javadoc. This is the value the MCP layer hands back to the lead as
* "write your file here", so callers may rely on it always being absolute.
*/
public record PendingRollover(String token, String leadTerminal, String handoverPath,
long requestedAtMillis) {}
/** Which check refused a {@link #confirm} call, named so a caller can act on it. */
public enum RefusalReason {
/** {@code leadRollover:} is not configured — absent at construction, or removed since. */
NOT_CONFIGURED,
/** {@code token} names no pending request: never opened, already confirmed, or cancelled. */
UNKNOWN_TOKEN,
/**
* The caller's terminal does not match the terminal that {@link #open} recorded for this
* token. Only the lead that opened a request may confirm it (fleetd #480 correction 2).
*/
NOT_YOUR_ROLLOVER,
/** {@code requireOperatorConfirm: true} and the caller passed {@code operatorConfirmed: false}. */
OPERATOR_NOT_CONFIRMED,
/** The handover file does not exist. */
HANDOVER_MISSING,
/** The handover file exists but is empty. */
HANDOVER_EMPTY,
/**
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
* older than {@code maxDocAgeSeconds}.
*/
HANDOVER_STALE
}
/**
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
*/
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
static RollDecision approved() {
return new RollDecision(true, null, "confirmed; the roll will run once the calling turn ends");
}
static RollDecision refused(RefusalReason reason, String detail) {
return new RollDecision(false, reason, detail);
}
}
private final AgentControl agents;
private final Supplier<FleetConfig.LeadRollover> configSupplier;
/**
* Terminal id → that lead's configured workspace directory (their {@code
* fleet.leaders.<name>.cwd}), or {@code null} when the terminal names no currently-recognised
* lead. {@link #open} calls this to resolve a relative {@code handoverPath} — see this class's
* javadoc. Required: there is no sane default that would not silently reintroduce the
* daemon-cwd bug this parameter exists to fix.
*/
private final Function<String, String> leadWorkspace;
private final LongSupplier nowMillis;
private final Runnable settleSleeper;
/**
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
* not a background scheduler. Tests inject {@code Runnable::run} to make the continuation run
* synchronously and deterministically on the calling thread.
*/
private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace) {
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
() -> sleepUninterruptibly(SETTLE_POLL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
}
/**
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
* against a file's modified time, which only a wall clock is comparable to, and {@code
* nanoTime} freezes while the host sleeps (fleetd #386).
*/
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace, LongSupplier nowMillis,
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents;
this.configSupplier = configSupplier;
this.leadWorkspace = leadWorkspace;
this.nowMillis = nowMillis;
this.settleSleeper = settleSleeper;
this.continuationRunner = continuationRunner;
}
private static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — this poll loop should not be aborted by an
// interrupt that was not meant for it.
}
}
/**
* The lead says it is ready to be replaced. Generates a token and records the resolved
* handover path, this moment's wall-clock timestamp (the baseline {@link #confirm} checks the
* handover file's modified time against), and {@code leadTerminal} — only that exact terminal
* may later {@link #confirm} this token.
*
* @param leadTerminal the calling lead's terminal id, resolved by the MCP layer from the
* connection (see this class's javadoc) — never a client-supplied value
* @param reason free-text audit note (logged only; not otherwise interpreted or stored)
* @throws IllegalStateException if {@code leadRollover:} is not configured
* @throws IllegalArgumentException if {@code leadTerminal} is null or blank
*/
public PendingRollover open(String leadTerminal, String reason) {
FleetConfig.LeadRollover cfg = configSupplier.get();
if (cfg == null) {
throw new IllegalStateException("leadRollover: is not configured");
}
if (leadTerminal == null || leadTerminal.isBlank()) {
throw new IllegalArgumentException("leadTerminal is required — it must be resolved from "
+ "the caller's connection, never accepted as a client-chosen argument");
}
String token = UUID.randomUUID().toString();
long requestedAt = nowMillis.getAsLong();
String resolvedPath = resolveHandoverPath(cfg.handoverPath(), leadTerminal);
PendingRollover p = new PendingRollover(token, leadTerminal, resolvedPath, requestedAt);
pending.put(token, p);
if (resolvedPath.equals(cfg.handoverPath())) {
log.info("lead-rollover: open token={} lead={} handoverPath={} reason={}",
token, leadTerminal, resolvedPath, reason);
} else {
// The configured value was relative (or merely un-normalized) and resolved to a
// different string — log both, so an operator reading this line can see which
// directory the daemon actually looked in, not just the value it was given.
log.info("lead-rollover: open token={} lead={} configuredHandoverPath={} "
+ "resolvedHandoverPath={} reason={}",
token, leadTerminal, cfg.handoverPath(), resolvedPath, reason);
}
return p;
}
/**
* Resolve {@code configured} to an absolute path exactly once, here, so nothing downstream
* ({@link #checkHandover}, the MCP layer's {@code open} response, {@link
* FleetConfig.LeadRollover#bootstrapTextFor}) ever has to resolve — or worse, silently
* mis-resolve — a relative path again.
*
* <p><strong>The return value is GUARANTEED absolute, not merely usually absolute.</strong>
* {@code leadWorkspace.apply(leadTerminal)} returns an operator-configured {@code
* fleet.leaders.<name>.cwd} string, and nothing forces an operator to write an absolute one —
* a relative {@code cwd} resolved with plain {@link Path#resolve} would still yield a relative
* result, silently reopening the exact bug this class exists to fix (every later reader back to
* interpreting an ambiguous string against ITS OWN working directory). {@link
* Path#toAbsolutePath()} closes that: it resolves any remaining relative path against {@code
* user.dir} (the JVM's own cwd), which is the correct base for an operator-written path the
* daemon process itself is meant to interpret, exactly like the {@code user.dir} fallback used
* below. Applying it unconditionally on both branches means the ALREADY-absolute branch stays a
* no-op (an absolute path is unaffected by {@code toAbsolutePath()}) while the relative-{@code
* cwd} branch above is closed the same way.
*
* <ul>
* <li>already absolute → returned unchanged (normalized)</li>
* <li>relative → resolved against {@code leadWorkspace.apply(leadTerminal)} when that is
* non-null and non-blank; otherwise against {@code System.getProperty("user.dir")} — the
* same fallback {@code LeadLauncher#launch} uses for a lead with no configured {@code
* cwd}. If {@code leadWorkspace}'s own answer is itself relative (an operator wrote a
* relative {@code cwd:}), the result is finished off against the daemon's own
* {@code user.dir} — see the paragraph above.</li>
* </ul>
*/
private String resolveHandoverPath(String configured, String leadTerminal) {
Path path = Path.of(configured);
if (path.isAbsolute()) {
// toAbsolutePath() is a no-op for an already-absolute path — kept here anyway so both
// branches call the exact same guarantee, rather than one branch relying on
// isAbsolute() alone to already imply what toAbsolutePath() enforces.
return path.toAbsolutePath().normalize().toString();
}
String workspace = leadWorkspace.apply(leadTerminal);
Path base = (workspace == null || workspace.isBlank())
? Path.of(System.getProperty("user.dir"))
: Path.of(workspace);
return base.resolve(path).toAbsolutePath().normalize().toString();
}
/**
* Validate every gate, then — if and only if all of them pass — hand a one-shot continuation
* that performs the actual roll to {@code continuationRunner} and return. <strong>This method
* never itself sends anything to the lead's pane</strong> — see this class's javadoc for why
* (it is called FROM the lead's own turn, so the pane cannot possibly be injectable yet).
*
* <p>Order: token lookup, then ownership ({@code callerTerminal} must match the terminal
* {@link #open} recorded — {@link RefusalReason#NOT_YOUR_ROLLOVER}), then the
* operator-confirmation gate, then the three handover-file checks (exists, not empty, fresh —
* see {@link #checkHandover}). The first failing check is returned and {@code token} stays
* pending (so a caller can fix the problem — e.g. rewrite the handover file — and retry with
* the same token); it is consumed only once every gate passes and the continuation is launched.
*
* @param callerTerminal the CALLING lead's terminal id, resolved by the MCP layer from the
* connection — never a client-supplied value (see this class's javadoc)
* @param token the token {@link #open} returned
* @param operatorConfirmed the caller's answer to "has an operator confirmed this roll" —
* consulted only when the live config's {@code requireOperatorConfirm}
* is true
*/
public RollDecision confirm(String callerTerminal, String token, boolean operatorConfirmed) {
FleetConfig.LeadRollover cfg = configSupplier.get();
if (cfg == null) {
return RollDecision.refused(RefusalReason.NOT_CONFIGURED, "leadRollover: is not configured");
}
PendingRollover p = pending.get(token);
if (p == null) {
return RollDecision.refused(RefusalReason.UNKNOWN_TOKEN,
"token " + token + " names no pending rollover request");
}
if (!p.leadTerminal().equals(callerTerminal)) {
return RollDecision.refused(RefusalReason.NOT_YOUR_ROLLOVER,
"token " + token + " was opened by a different lead terminal");
}
if (cfg.requireOperatorConfirm() && !operatorConfirmed) {
return RollDecision.refused(RefusalReason.OPERATOR_NOT_CONFIRMED,
"requireOperatorConfirm is true and operatorConfirmed was false");
}
RollDecision docCheck = checkHandover(p, cfg);
if (docCheck != null) {
return docCheck;
}
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
continuationRunner.accept(() -> runRollover(p, cfg));
return RollDecision.approved();
}
/**
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
*/
private void runRollover(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
if (!turnResult.settled()) {
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
// path already does.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "confirm() — refusing to send /clear at all; the calling lead's own "
+ "turn is still live and clearing it now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
return;
}
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
// pane forever (see this class's javadoc).
agents.send(lead, "/clear");
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
if (!clearResult.settled()) {
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
// actually ran — an operator reading only that number wrongly believes it is a measured
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
// so the two can be compared at a glance.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
+ "elapsed={}ms nudges={})",
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
clearResult.nudges());
return;
}
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
public boolean cancel(String token) {
return pending.remove(token) != null;
}
/**
* The three handover-file checks, in order: exists, not empty, fresh (modified after
* {@link #open}'s timestamp and not older than {@code maxDocAgeSeconds}). Stats {@code
* p.handoverPath()} directly — {@link #open} already resolved it to an absolute path, so this
* never has to guess which directory it means.
*
* @return the first failing check's refusal, or {@code null} when all three pass
*/
private RollDecision checkHandover(PendingRollover p, FleetConfig.LeadRollover cfg) {
Path path = Path.of(p.handoverPath());
if (!Files.exists(path)) {
return RollDecision.refused(RefusalReason.HANDOVER_MISSING,
"handover file " + p.handoverPath() + " does not exist");
}
long size;
long mtimeMillis;
try {
size = Files.size(path);
mtimeMillis = Files.getLastModifiedTime(path).toMillis();
} catch (IOException e) {
throw new UncheckedIOException("failed to stat handover file " + p.handoverPath(), e);
}
if (size == 0) {
return RollDecision.refused(RefusalReason.HANDOVER_EMPTY,
"handover file " + p.handoverPath() + " is empty");
}
if (mtimeMillis <= p.requestedAtMillis()) {
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
"handover file " + p.handoverPath() + " was not modified after the open() "
+ "request (mtime=" + mtimeMillis + "ms, requestedAt=" + p.requestedAtMillis() + "ms)");
}
long ageMillis = nowMillis.getAsLong() - mtimeMillis;
long maxAgeMillis = TimeUnit.SECONDS.toMillis(cfg.maxDocAgeSeconds());
if (ageMillis > maxAgeMillis) {
return RollDecision.refused(RefusalReason.HANDOVER_STALE,
"handover file " + p.handoverPath() + " is " + TimeUnit.MILLISECONDS.toSeconds(ageMillis)
+ "s old, older than maxDocAgeSeconds=" + cfg.maxDocAgeSeconds());
}
return null;
}
/**
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
*
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
* turn" — and it accepts {@link AgentStatus#BLOCKED} for that purpose, because a pane paused on
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
* reason, on the second wait.)
*
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
*/
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle: {}",
target, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
}
settleSleeper.run();
}
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
}
/**
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
* count.
*/
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
/**
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
* own javadoc already records that the submit accompanying a delivery "can race the paste —
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
* it, landing as one concatenated line.
*
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
* <ul>
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
* turn;</li>
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
* Claude Code prompt is a no-op, so repeating it is safe;</li>
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
* observed, releases rather than wedges the roll instead of nudging again — the same
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
* path ran and how long it actually took;</li>
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
* </ul>
*
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
* a real boundary is reached or {@code settleSeconds} runs out.
*
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
* nudge must not abort the roll.
*
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
* {@code settleSeconds} budget (fleetd #494).
*/
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
int idlePollsAwaitingPickup = 0;
int nudges = 0;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
+ "/clear: {}", target, e.toString());
status = null;
}
if (status == AgentStatus.WORKING) {
pickedUp = true;
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
if (pickedUp) {
// a real WORKING -> IDLE/DONE completion boundary
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
}
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
long elapsedMillis = nowMillis.getAsLong() - startMillis;
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
// instead of wedging the roll — that trade is deliberate and stays. But it is
// also exactly the case that reported false success in the real incident (the
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
// print the MEASURED elapsed time next to the target pane, not just the count.
//
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
// this loop, on the same branch, so on this branch they cannot differ from
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
// difference on this line, and printing the counters does not change that.
// What it does buy: one source of truth instead of two, so a later change to
// the loop (an early return, a second increment site, a different exit
// condition) cannot leave this message reporting a number the loop no longer
// produces. The place where `nudges` genuinely varies with the run — and is
// covered by a test that can tell it apart from a constant — is the
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
+ "releasing rather than wedging the roll (elapsed={}ms)",
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
return new ClearSettleResult(true, elapsedMillis, nudges);
}
try {
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
} catch (RuntimeException e) {
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
target, e.getMessage());
} finally {
nudges++; // an attempted nudge, whether or not the submit call itself threw
}
}
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
// signal nor a boundary — keep polling without nudging or releasing.
settleSleeper.run();
}
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
}
/**
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
* the settle/timeout decision, so callers can log them instead of the configured budget, which
* is not how long the wait actually ran.
*/
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
}
@@ -0,0 +1,81 @@
package dev.ltms.fleet.mcp;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* fleetd #469: a launch charter that names an MCP tool the server does not register must stop the
* daemon at startup, not wait for a member to discover the gap by calling something that is not
* there.
*
* <p>Deliberately its own class outside {@code dev.ltms.fleet.config}, not a case in {@link
* dev.ltms.fleet.config.FleetConfig#validateCharters()}. The canonical tool surface ({@link
* FleetTool}) lives in the {@code mcp} package; config is loaded before the MCP server exists and
* must not gain a dependency on it. So this check belongs at the seam that already holds both a
* loaded {@code FleetConfig} and the {@code mcp} package: {@code Fleetd.main}, called right after
* {@code cfg.validateAll()} and before anything opens a socket or spawns a member.
*
* <p>{@code #464}'s {@code CharterToolSurfaceTest} proved the same comparison against a charter
* fixture it wrote itself into a {@code @TempDir}, which meant nothing anyone wrote into the live
* {@code fleetd.yaml} could ever fail it. This class is what a real charter is actually checked
* against at boot; {@code FleetdStartupValidationTest} exercises it through {@code Fleetd.main}
* itself, the same way it proves every other {@code validateXxx()} still runs there.
*
* <p><strong>fleetd #474</strong> — startup was not the only door: {@code
* dev.ltms.fleet.config.ConfigRef#reload()} used to run {@code FleetConfig#validateAll()} alone,
* which does not look at what a charter's text names, so a charter naming an unregistered tool
* that could not have booted the daemon could still be installed into a running one through a
* reload. This class still knows nothing about {@code ConfigRef} — {@code Fleetd.main} wires a
* small {@code FleetConfig -> void} adapter over {@link #assertChartersNameOnlyRegisteredTools}
* ({@code Fleetd::assertChartersNameOnlyRegisteredTools}) into {@code ConfigRef}'s constructor as
* its {@code Consumer<FleetConfig>} {@code extraValidation}, run inside {@code reload()}'s same
* try/catch as {@code validateAll()}, so both call sites — {@code Fleetd.main} at startup and
* {@code ConfigRef#reload()} afterwards — go through this one method and can never check different
* things. {@code dev.ltms.fleet.config.ConfigRefTest} and {@code
* FleetdConfigRefCharterToolSurfaceWiringTest} are what prove the reload call site, the same way
* {@code FleetdStartupValidationTest} proves the startup one.
*/
public final class CharterToolSurface {
/** A {@code fleet_…} (current) or {@code bridge_…} (pre-CB-634) tool-shaped token in prose. */
private static final Pattern TOOL_REFERENCE = Pattern.compile("(fleet_[a-z_]+|bridge_[a-z_]+)");
private CharterToolSurface() {
}
/**
* @param charters the configured {@code fleet.charters:} map (role wire name → charter text);
* {@code null} or empty is a no-op, same as an absent {@code fleet:} block
* @throws IllegalStateException naming the charter key and every tool it names that {@link
* FleetTool} does not list, when any charter does so
*/
public static void assertChartersNameOnlyRegisteredTools(Map<String, String> charters) {
if (charters == null || charters.isEmpty()) {
return;
}
Set<String> registered = FleetTool.wireNames();
List<String> bad = new ArrayList<>();
charters.forEach((key, text) -> {
if (text == null) {
return;
}
Set<String> named = new LinkedHashSet<>();
Matcher m = TOOL_REFERENCE.matcher(text);
while (m.find()) {
named.add(m.group(1));
}
named.stream()
.filter(t -> !registered.contains(t))
.forEach(unknown -> bad.add("fleet.charters." + key + " names '" + unknown
+ "', which the server does not register (registered: " + registered + ")."));
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
}
@@ -33,9 +33,10 @@ public final class ConnectionIdentity {
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
* the pane scan behind {@code terminal} ran to completion ({@link #scanComplete}).
*/
public record Caller(String terminal, long pid) {
public record Caller(String terminal, long pid, boolean scanComplete) {
/**
* Whether the OS peer-PID lookup actually succeeded — {@code false} means {@code pid} is
@@ -51,6 +52,10 @@ public final class ConnectionIdentity {
* {@link ConnectionIdentity#isLoopback} is centralised rather than left for each caller to
* reimplement: a raw {@code pid > 0} check duplicated at every call site is precisely the
* "one rule, two copies" shape that let #305 drift.
*
* <p>This method is deliberately NOT widened for fleetd #505's failure (a herdr error
* during the pane scan, not a failed lsof lookup) — it still tests only the sentinel it is
* named for. #505 is a different axis, carried separately in {@link #scanComplete}.
*/
public boolean resolved() {
return pid > 0;
@@ -60,10 +65,11 @@ public final class ConnectionIdentity {
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
return new Caller(null, -1, true); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
PaneLocator.Lookup lookup = panes.terminalForPid(pid);
return new Caller(lookup.terminal(), pid, lookup.complete());
}
/**
@@ -12,6 +12,7 @@ import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.MessageService;
@@ -92,7 +93,18 @@ public final class FleetMcp {
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
/**
* fleetd #518: whether {@link #denyFor} enforces the CB-505 policy table at all. Replaces the
* old {@code CallerResolver authz} field, whose null-ness used to decide BOTH this AND which
* principal-resolution code path {@link #contextExtractor} ran — reaching "authorization off"
* by simply not passing a {@link CallerResolver} also meant the resolved {@link Principal}
* came from a second, separately-maintained heuristic ({@code legacyPrincipal}, now deleted)
* that nothing ever exercised. There is now exactly one resolution path ({@code callers},
* required and non-null below) and a separate, explicitly-chosen {@link AuthorizationMode}
* for this flag — so a caller can turn enforcement off without silently swapping in a second,
* untested identity heuristic.
*/
private final boolean authorizationEnforced;
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
@@ -105,6 +117,13 @@ public final class FleetMcp {
private final LeadChannel leadChannel;
/** fleetd #361: {@code coordinator.peers} — see {@link CoordinationSource}. Empty when unset. */
private final List<String> peers;
/**
* fleetd #480 Unit C: the executor behind {@code fleet_handover}. {@code null} whenever
* {@code leadRollover:} is not configured — {@code fleet_handover} is still registered (see
* this class's javadoc on the charter tool-surface gate), and every action then degrades to a
* clean {@code NOT_CONFIGURED} refusal rather than throwing. See {@link #handover}.
*/
private final LeadRollover leadRollover;
/** Capacity facts used by {@code fleet_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
@@ -243,78 +262,80 @@ public final class FleetMcp {
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
* @param quarantine CB-578 stage B facts for {@code fleet_profiles}; required — pass
* {@link QuarantineSource#none()} for a caller that does not want the feature
* fleetd #518: whether {@link #denyFor} enforces the CB-505 policy table. A required
* constructor parameter with no default, so "authorization is off" can only be reached by a
* caller explicitly saying so — never by omitting a {@link CallerResolver} the way the old
* {@code callers == null} idiom allowed. {@code callers} itself is required either way: even
* under {@link #UNENFORCED}, the one real {@link CallerResolver} still resolves every caller's
* {@link Principal} (so {@code markSpawnedMemberPresent}/{@code recordPrimarySingleton} see a
* real identity), and {@link #denyFor} is the only thing that changes.
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, null, OutageSource.none(), LeadSeatSource.none());
}
public enum AuthorizationMode { ENFORCED, UNENFORCED }
/**
* As above, with this daemon's lead-to-lead channel (CB-637). {@code leadChannel} is
* {@code null} whenever no {@code coordinator:} block is configured or its broker could not be
* reached at boot — cross-daemon lead messaging is simply off, and {@code fleet_send{coordId}}
* says so rather than failing obscurely.
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, OutageSource.none(), LeadSeatSource.none());
}
/**
* As above, with fleetd #201 Unit 5 cool-off facts for {@code fleet_list}/{@code fleet_profiles}
* (see {@link OutageSource}).
* The only constructor (fleetd #480 Unit C correction round). Every field below used to have
* its own defaulting overload — {@code leadChannel}/{@code outage}/{@code leadSeats}/
* {@code peers}/{@code leadRollover} each got a shorter, convenience constructor that silently
* filled it in ({@code null}, {@code .none()}, or {@code List.of()}) when a caller did not pass
* it. That is exactly how {@code Fleetd.main}'s wiring of {@link LeadRollover} could have gone
* silently missing: drop one argument from the real call and it just lands on a shorter
* overload instead of failing to compile, and every existing test — none of which exercises
* {@code Fleetd.main} itself — stays green while the live daemon quietly answers
* {@code NOT_CONFIGURED} to {@code fleet_handover} forever. Collapsing every overload into one
* required-everything constructor turns that mistake into a compile error instead: this
* project's own antidote for a defaulted parameter surviving as an untested decision (see
* {@code FleetdCompletionResolverWiringTest} / {@code FleetdLeadRolloverWiringTest}'s own
* javadoc for the same lesson applied to a different seam). A caller that genuinely wants a
* feature off must now say so explicitly at the call site — {@code null},
* {@link OutageSource#none()}, {@link LeadSeatSource#none()}, {@code List.of()} are all still
* perfectly fine values, just never an implicit default reached by omission.
*
* @param outage required — pass {@link OutageSource#none()} for a caller that does not want the
* feature, never a defaulting overload (the same rule {@code quarantine} follows).
* @param callers resolves each call's {@link Principal}. Required, never {@code null} —
* fleetd #518: use {@link AuthorizationMode#UNENFORCED} to disable
* enforcement, not a missing resolver. This surface needs its own
* enforcement: {@code /mcp} is a raw servlet on Jetty's context handler and
* never passes through Javalin's {@code before} filter, so the REST guard
* does not cover it.
* @param authorizationMode fleetd #518: whether {@link #denyFor} enforces the CB-505 policy
* table ({@link AuthorizationMode#ENFORCED}) or leaves the gate open
* ({@link AuthorizationMode#UNENFORCED}, for the pre-CB-513 test suite that
* does not exercise authorization). Required, with no default.
* @param metrics registry for auth-failure counting; may be {@code null}
* @param quarantine CB-578 stage B facts for {@code fleet_profiles}; pass
* {@link QuarantineSource#none()} for a caller that does not want the
* feature
* @param leadChannel this daemon's lead-to-lead channel (CB-637); {@code null} whenever no
* {@code coordinator:} block is configured or its broker could not be
* reached at boot — cross-daemon lead messaging is simply off, and
* {@code fleet_send{coordId}} says so rather than failing obscurely
* @param outage fleetd #201 Unit 5 cool-off facts for {@code fleet_list}/
* {@code fleet_profiles}; pass {@link OutageSource#none()} for a caller
* that does not want the feature
* @param leadSeats fleetd #176 lead-seat facts (see {@link LeadSeatSource}); pass
* {@link LeadSeatSource#none()} for a caller that does not want the
* feature
* @param peers fleetd #361 {@code coordinator.peers} (see {@link CoordinationSource});
* the coord-ids declared there, or empty when unset or when
* {@code leadChannel} is {@code null}
* @param leadRollover fleetd #480 Unit C: the {@link LeadRollover} executor behind
* {@code fleet_handover}. {@code null} whenever {@code leadRollover:} is
* not configured — an upgraded daemon must never silently acquire the
* ability to clear a lead's own pane (mirrors
* {@code Fleetd.leadRollover(...)}'s own construction gate).
* {@code fleet_handover} is registered unconditionally either way — see
* this class's javadoc and fleetd #474's charter tool-surface gate — and
* every action degrades to a clean refusal naming {@code NOT_CONFIGURED}
* instead of throwing. See {@link #handover}.
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, outage, LeadSeatSource.none());
}
/**
* As above, with fleetd #176 lead-seat facts (see {@link LeadSeatSource}).
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
CallerResolver callers, AuthorizationMode authorizationMode, Metrics metrics,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
LeadSeatSource leadSeats) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, outage, leadSeats, List.of());
}
/**
* As above, with fleetd #361 {@code coordinator.peers} (see {@link CoordinationSource}). This is
* what {@code Fleetd.main} actually wires up.
*
* @param leadSeats required — pass {@link LeadSeatSource#none()} for a caller that does not want
* the feature, never a defaulting overload (the same rule {@code quarantine} and
* {@code outage} follow).
* @param peers the coord-ids declared under {@code coordinator.peers}; empty when unset or
* when {@code leadChannel} is {@code null}.
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
LeadSeatSource leadSeats, List<String> peers) {
LeadSeatSource leadSeats, List<String> peers, LeadRollover leadRollover) {
Objects.requireNonNull(callers, "callers");
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
== AuthorizationMode.ENFORCED;
this.leadChannel = leadChannel;
this.peers = peers == null ? List.of() : List.copyOf(peers);
this.capacity = capacity;
@@ -322,6 +343,7 @@ public final class FleetMcp {
this.outage = Objects.requireNonNull(outage, "outage");
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
this.healthCoverage = healthCoverage;
this.leadRollover = leadRollover;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
@@ -332,11 +354,11 @@ public final class FleetMcp {
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
// One resolution per call, shared with the REST surface via CallerResolver so
// the two paths cannot drift on who a caller is.
Principal p = callers != null
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
// the two paths cannot drift on who a caller is. fleetd #518: callers is
// required (never null) so there is no second, untested resolution path to
// fall back to here — AuthorizationMode governs enforcement, not identity.
Principal p = callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"));
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
// lead, which carries its pane too, while including every spawned member role.
// Enrolling a lead would count it as an available member in the roster.
@@ -453,7 +475,7 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_list", Map.of()), null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
leadSeats, callers == null ? Map.of() : callers.leads(),
leadSeats, callers.leads(),
callerTerminal(exchange),
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)));
@@ -477,6 +499,15 @@ public final class FleetMcp {
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
};
// fleetd #480 Unit C: fleet_handover. No terminal/session argument at all — the lead pane
// to roll is ALWAYS the caller's own connection-resolved terminal (never a request field),
// per LeadRollover's class javadoc (fleetd #480 correction 2).
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> handoverHandler =
(exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_handover", req.arguments()), null);
if (denied != null) return denied;
return handover(leadRollover, callerTerminal(exchange), req.arguments());
};
McpSchema.Tool fleetSend = sendTool();
McpSchema.Tool fleetReply = replyTool();
@@ -489,6 +520,23 @@ public final class FleetMcp {
McpSchema.Tool fleetStop = stopTool();
McpSchema.Tool fleetProfiles = profilesTool();
McpSchema.Tool fleetWhoami = whoamiTool();
McpSchema.Tool fleetHandover = handoverTool();
// fleetd #469: the tool schemas above are already named from FleetTool.wireName(), but
// this is the check that a schema was not accidentally dropped, duplicated, or added
// under a name FleetTool does not list. It runs once, at construction (startup), rather
// than being left to CharterToolSurface or a test to discover later — a canonical entry
// this server never registers, or a registration with no canonical entry backing it, is a
// startup failure, not a silent gap.
Set<String> registeredToolNames = Set.of(fleetSend.name(), fleetReply.name(), fleetAsk.name(),
fleetStatus.name(), fleetPoll.name(), fleetAck.name(), fleetSpawn.name(),
fleetList.name(), fleetStop.name(), fleetProfiles.name(), fleetWhoami.name(),
fleetHandover.name());
if (!registeredToolNames.equals(FleetTool.wireNames())) {
throw new IllegalStateException("fleetd #469: registered MCP tools " + registeredToolNames
+ " do not match the canonical tool set " + FleetTool.wireNames()
+ " -- FleetTool is the single source of truth for what this server registers");
}
this.server = McpServer.sync(transport)
.serverInfo("fleet", "0.1.0")
@@ -504,22 +552,11 @@ public final class FleetMcp {
.toolCall(fleetStop, stopHandler)
.toolCall(fleetProfiles, profilesHandler)
.toolCall(fleetWhoami, whoamiHandler)
.toolCall(fleetHandover, handoverHandler)
.build();
this.authz = callers;
this.metrics = metrics;
}
/**
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
* only by the legacy constructor, where authorization is not enforced anyway.
*/
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
ConnectionIdentity.Caller c = identity.resolve(addr, port);
return c.terminal() != null
? Principal.worker(c.terminal(), c.pid())
: Principal.primary(c.pid());
}
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
@@ -574,8 +611,8 @@ public final class FleetMcp {
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (authz == null) {
return null; // legacy constructor: authorization not enforced
if (!authorizationEnforced) {
return null; // AuthorizationMode.UNENFORCED: authorization not enforced (fleetd #518)
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
@@ -668,6 +705,19 @@ public final class FleetMcp {
server.closeGracefully();
}
/**
* The tools this server has actually registered with the MCP SDK, wire schema included.
*
* <p>Package-private, for tests that must check the REGISTERED schema rather than this
* class's own source text — e.g. proving {@code fleet_handover}'s input schema carries no
* caller-terminal parameter (fleetd #480 Unit C acceptance criterion 4). A source-text scrape
* cannot tell "the schema builder omits this key" apart from "a typo means it never runs" the
* way asking the constructed server itself can.
*/
List<McpSchema.Tool> registeredTools() {
return server.listTools();
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/**
@@ -897,20 +947,35 @@ public final class FleetMcp {
/**
* The action a registered tool handler actually hands to the authorization gate.
* Keeping this choice beside the registered-tool inventory makes a new tool fail the coverage
* test until its action is pinned.
*
* <p>{@code toolName} is a raw string off the wire (an MCP call names its tool by string, and a
* malformed or stale client can send anything), so resolving it against {@link FleetTool} first
* — and throwing on a miss — is still a run-time check by necessity. What moved to compile time
* is the second step: {@link #authzAction(FleetTool, Map)} switches on the resolved {@link
* FleetTool} itself with no {@code default}, so a new {@link FleetTool} constant with no pinned
* action fails {@code mvn compile}, not just {@code FleetMcpAuthzTest} at run time.
*/
static Authz.Action toolAction(String toolName, Map<String, Object> arguments) {
return switch (toolName) {
case "fleet_send" -> Authz.Action.SEND;
case "fleet_reply" -> Authz.Action.REPLY;
case "fleet_ask" -> Authz.Action.ASK;
case "fleet_status", "fleet_list", "fleet_profiles", "fleet_whoami" -> Authz.Action.READ;
case "fleet_poll" -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case "fleet_ack" -> Authz.Action.DRAIN;
case "fleet_spawn" -> Authz.Action.SPAWN;
case "fleet_stop" -> Authz.Action.STOP;
default -> throw new IllegalArgumentException("unregistered tool: " + toolName);
FleetTool tool = FleetTool.byWireName(toolName)
.orElseThrow(() -> new IllegalArgumentException("unregistered tool: " + toolName));
return authzAction(tool, arguments);
}
/**
* Exhaustive over {@link FleetTool} on purpose — no {@code default}. Adding a tool to {@link
* FleetTool} without adding its case here is a compile error (fleetd #469).
*/
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
return switch (tool) {
case SEND -> Authz.Action.SEND;
case REPLY -> Authz.Action.REPLY;
case ASK -> Authz.Action.ASK;
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case ACK -> Authz.Action.DRAIN;
case SPAWN -> Authz.Action.SPAWN;
case STOP -> Authz.Action.STOP;
case HANDOVER -> Authz.Action.HANDOVER;
};
}
@@ -1135,6 +1200,119 @@ public final class FleetMcp {
return text(json(m));
}
/**
* {@code fleet_handover} (fleetd #480 Unit C): drive {@link LeadRollover#open}/
* {@link LeadRollover#confirm}/{@link LeadRollover#cancel} from a tool call.
*
* <p>{@code callerTerminal} is the CALLING lead's terminal id, resolved by the MCP layer from
* the connection (see {@link #callerTerminal(McpSyncServerExchange)}) — never a request field.
* This is why this tool's input schema ({@link #handoverTool}) carries no terminal/session/
* leadTerminal parameter of any kind: a lead can only ever open or confirm a rollover of its
* OWN pane (see {@code LeadRollover}'s class javadoc, fleetd #480 correction 2).
*
* <p>{@code leadRollover} is {@code null} whenever {@code leadRollover:} is not configured.
* This tool is registered unconditionally regardless (see this class's javadoc on the fleetd
* #474 charter tool-surface gate), so every action here must degrade to a clean, structured
* refusal naming {@code NOT_CONFIGURED} rather than ever throwing.
*/
static McpSchema.CallToolResult handover(LeadRollover leadRollover, String callerTerminal,
Map<String, Object> args) {
String action = str(args, "action");
if (isBlank(action)) {
return error("action is required: \"open\", \"confirm\" or \"cancel\"");
}
return switch (action) {
case "open" -> handoverOpen(leadRollover, callerTerminal, str(args, "reason"));
case "confirm" -> handoverConfirm(leadRollover, callerTerminal, str(args, "token"),
truthy(args, "operatorConfirmed"));
case "cancel" -> handoverCancel(leadRollover, str(args, "token"));
default -> error("unknown action \"" + action + "\" — must be \"open\", \"confirm\" or \"cancel\"");
};
}
/**
* {@code action: "open"}. On the null-{@code leadRollover} path (not configured) and on the
* config-removed-since-construction path ({@link LeadRollover#open} itself throws {@link
* IllegalStateException} for that), both degrade to the same clean {@code NOT_CONFIGURED}
* refusal — never an escaping exception.
*/
private static McpSchema.CallToolResult handoverOpen(LeadRollover leadRollover, String callerTerminal,
String reason) {
if (leadRollover == null) {
return notConfigured();
}
if (isBlank(callerTerminal)) {
// An unnamed primary (token/loopback path, no resolved pane) has nowhere for the
// eventual /clear + bootstrap to land — LeadRollover#open would throw
// IllegalArgumentException for the same reason; refuse cleanly here instead.
return error("fleet_handover requires a named lead pane (a resolved connection terminal) "
+ "to open a rollover request against — an unnamed primary has none");
}
try {
LeadRollover.PendingRollover p = leadRollover.open(callerTerminal, reason);
Map<String, Object> m = new LinkedHashMap<>();
m.put("token", p.token());
m.put("handoverPath", p.handoverPath());
m.put("requestedAtMillis", p.requestedAtMillis());
return text(json(m));
} catch (IllegalStateException e) {
// leadRollover: was removed from config by a hot reload since this FleetMcp was
// constructed — same clean refusal as the null-at-construction case above.
return notConfigured();
}
}
/** {@code action: "confirm"}. Surfaces every {@link LeadRollover.RefusalReason} verbatim. */
private static McpSchema.CallToolResult handoverConfirm(LeadRollover leadRollover, String callerTerminal,
String token, boolean operatorConfirmed) {
if (leadRollover == null) {
return refusalJson(false, "NOT_CONFIGURED", "leadRollover: is not configured");
}
if (isBlank(token)) {
return error("token is required for action \"confirm\"");
}
LeadRollover.RollDecision d = leadRollover.confirm(callerTerminal, token, operatorConfirmed);
return refusalJson(d.accepted(), d.reason() == null ? null : d.reason().name(), d.detail());
}
/** {@code action: "cancel"}. An unknown token is a clean "no pending request", never an error. */
private static McpSchema.CallToolResult handoverCancel(LeadRollover leadRollover, String token) {
if (leadRollover == null) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("cancelled", false);
m.put("reason", "NOT_CONFIGURED");
m.put("detail", "leadRollover: is not configured");
return text(json(m));
}
if (isBlank(token)) {
return error("token is required for action \"cancel\"");
}
boolean existed = leadRollover.cancel(token);
Map<String, Object> m = new LinkedHashMap<>();
m.put("cancelled", existed);
if (!existed) {
m.put("detail", "no pending request for token " + token);
}
return text(json(m));
}
/** The one shared {@code NOT_CONFIGURED} refusal shape for {@code open}/{@code confirm}. */
private static McpSchema.CallToolResult notConfigured() {
return refusalJson(false, "NOT_CONFIGURED", "leadRollover: is not configured");
}
private static McpSchema.CallToolResult refusalJson(boolean accepted, String reason, String detail) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("accepted", accepted);
if (reason != null) {
m.put("reason", reason);
}
if (detail != null) {
m.put("detail", detail);
}
return text(json(m));
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code fleet_spawn} without cwd/caller context (default resolution). */
@@ -1269,6 +1447,13 @@ public final class FleetMcp {
* (the backend text that triggered the most recent quarantine of that credential, omitted when
* none is known) — so a lead can see WHICH model to turn off and WHY, without reading the
* daemon log.
*
* <p>fleetd #466 scope item 2: each {@code quarantined} row also names {@code
* quarantineAttempt} — 1 for a first-time exhaustion, 2 for the second in a row, and so on — so
* an operator sees "this is the 5th time" instead of inferring it from {@code
* quarantinedForSeconds} alone. Read off {@link BackendQuarantine#status(String)}, the same
* one-call accessor {@code quarantinedForSeconds} itself comes from here (see its doc) — never a
* separately derived count.
*/
public static Map<String, Object> profilesView(PeerLauncher workers, QuarantineSource quarantine, OutageSource outage) {
Map<String, Object> result = new LinkedHashMap<>();
@@ -1281,10 +1466,14 @@ public final class FleetMcp {
exhaustionDetectionArmed.put(profile, quarantine.exhaustedPatternArmed().apply(profile));
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
// fleetd #466 scope item 2: quarantinedForSeconds and quarantineAttempt come off the
// ONE BackendQuarantine#status(credentialId) call — never a second, independent read
// for the attempt count — so the two can never disagree about which streak this is.
quarantine.quarantine().status(credentialId).ifPresent(status -> {
Map<String, Object> row = new LinkedHashMap<>();
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
row.put("quarantinedForSeconds", status.remainingSeconds());
row.put("quarantineAttempt", status.repeatCount());
String model = quarantine.modelFor().apply(profile);
if (model != null && !model.isBlank()) {
row.put("model", model);
@@ -1724,10 +1913,13 @@ public final class FleetMcp {
}
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
// fleetd #466 scope item 2: same one-call status() read as profilesView above — see its
// comment for why this must not become two separate lookups.
quarantine.quarantine().status(credentialId).ifPresent(status -> {
row.put("free", 0);
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
row.put("quarantinedForSeconds", status.remainingSeconds());
row.put("quarantineAttempt", status.repeatCount());
});
}
String outageCredentialId = outage.credentialIdFor().apply(profile);
@@ -1805,7 +1997,7 @@ public final class FleetMcp {
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("fleet_send",
return tool(FleetTool.SEND.wireName(),
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with fleet_poll. To "
@@ -1833,7 +2025,7 @@ public final class FleetMcp {
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("fleet_ask",
return tool(FleetTool.ASK.wireName(),
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
@@ -1846,7 +2038,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool pollTool() {
return tool("fleet_poll",
return tool(FleetTool.POLL.wireName(),
"Check an async delegation (a fleet_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
@@ -1866,7 +2058,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool ackTool() {
return tool("fleet_ack",
return tool(FleetTool.ACK.wireName(),
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
@@ -1880,7 +2072,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool spawnTool() {
return tool("fleet_spawn",
return tool(FleetTool.SPAWN.wireName(),
"Spawn a new off-subscription member session. A member has two independent attributes: "
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
@@ -1912,7 +2104,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool profilesTool() {
return tool("fleet_profiles",
return tool(FleetTool.PROFILES.wireName(),
"List the configured worker profiles (backends) and which one fleet_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — fleet_spawn onto it is refused until "
@@ -1930,7 +2122,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool listTool() {
return tool("fleet_list",
return tool(FleetTool.LIST.wireName(),
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
@@ -1966,7 +2158,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool stopTool() {
return tool("fleet_stop",
return tool(FleetTool.STOP.wireName(),
"Tear down a worker session by its paneId (from fleet_spawn or fleet_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
@@ -1975,7 +2167,7 @@ public final class FleetMcp {
private static McpSchema.Tool replyTool() {
// No session/target arg — the caller's identity is resolved from the connection.
return tool("fleet_reply",
return tool(FleetTool.REPLY.wireName(),
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked fleet_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
@@ -1986,7 +2178,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool statusTool() {
return tool("fleet_status",
return tool(FleetTool.STATUS.wireName(),
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
@@ -1994,7 +2186,7 @@ public final class FleetMcp {
}
private static McpSchema.Tool whoamiTool() {
return tool("fleet_whoami",
return tool(FleetTool.WHOAMI.wireName(),
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
@@ -2008,6 +2200,31 @@ public final class FleetMcp {
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool handoverTool() {
return tool(FleetTool.HANDOVER.wireName(),
"Replace your OWN lead session once its context is full: write a handover file, "
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
+ "session against it. Three actions: 'open' (requests a token and the "
+ "handoverPath you must write the handover file to before confirming), "
+ "'confirm' (validates every gate and — only if every one passes — schedules "
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
+ "own turn ends), and 'cancel' (drops a pending request without rolling). "
+ "Primary-only. There is deliberately no terminal/session/leadTerminal "
+ "parameter: the pane to roll is always resolved from YOUR OWN connection, "
+ "never a value you pass, so you can only ever roll yourself — never another "
+ "lead. Requires leadRollover: to be configured; when it is not, every action "
+ "returns a clean refusal naming NOT_CONFIGURED instead of failing.",
objectSchema(Map.of(
"action", stringProp("\"open\", \"confirm\" or \"cancel\""),
"reason", stringProp("Free-text audit note for \"open\" (optional, logged only)"),
"token", stringProp("The token \"open\" returned — required for \"confirm\" and \"cancel\""),
"operatorConfirmed", Map.of("type", "boolean",
"description", "For \"confirm\": your answer to \"has the human operator "
+ "confirmed this wipe\" (default false; only consulted when "
+ "leadRollover.requireOperatorConfirm is true)")),
List.of("action")));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@@ -2042,6 +2259,11 @@ public final class FleetMcp {
return v instanceof Number n ? n.longValue() : null;
}
/** {@code true} only when {@code args.get(key)} is the boolean {@code true} — absent/null/anything else is {@code false}. */
private static boolean truthy(Map<String, Object> args, String key) {
return Boolean.TRUE.equals(args.get(key));
}
private static long clamp(long ms) {
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
}
@@ -0,0 +1,83 @@
package dev.ltms.fleet.mcp;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
/**
* The one canonical set of MCP tool names this daemon registers (fleetd #469, follow-up to #464).
*
* <p>Before this enum, the tool surface was written twice with nothing tying the copies together:
* once as the literal {@code "fleet_…"} string passed to each tool-schema builder in
* {@link FleetMcp}, and again as the case labels of {@link FleetMcp}'s authorization switch. A
* reader that needed "what does this server register" — a charter check, in particular — had no
* source to ask except scraping {@code FleetMcp.java}'s source text for {@code tool("…")} calls: a
* third copy of the same list, and the weakest of the three forms.
*
* <p>Every reader that needs the registered tool surface now asks this enum instead:
*
* <ul>
* <li>the tool-schema builders in {@code FleetMcp} pass {@code wireName()} rather than a literal;
* <li>{@code FleetMcp}'s constructor asserts, at startup, that the set of tool names it actually
* registers with the MCP SDK equals {@link #wireNames()} exactly — a canonical entry that is
* never registered, or a registration with no canonical entry backing it, fails the daemon's
* own boot rather than only a test's;
* <li>{@code FleetMcp.toolAction}'s dispatch onto {@code Authz.Action} switches on the enum
* (not the raw string) with no {@code default}, so adding a tool here without pinning its
* action is a compile error, not a run-time throw;
* <li>{@link CharterToolSurface} asks {@link #wireNames()} to check a configured launch charter
* against the live tool surface, instead of scraping source text a third time.
* </ul>
*/
public enum FleetTool {
SEND("fleet_send"),
REPLY("fleet_reply"),
ASK("fleet_ask"),
STATUS("fleet_status"),
POLL("fleet_poll"),
ACK("fleet_ack"),
SPAWN("fleet_spawn"),
LIST("fleet_list"),
STOP("fleet_stop"),
PROFILES("fleet_profiles"),
WHOAMI("fleet_whoami"),
HANDOVER("fleet_handover");
private final String wireName;
FleetTool(String wireName) {
this.wireName = wireName;
}
/** The name this tool is registered under, and called by, on the wire ({@code "fleet_send"}, …). */
public String wireName() {
return wireName;
}
private static final Map<String, FleetTool> BY_WIRE_NAME;
private static final Set<String> WIRE_NAMES;
static {
Map<String, FleetTool> byName = new LinkedHashMap<>();
Set<String> names = new LinkedHashSet<>();
for (FleetTool tool : values()) {
byName.put(tool.wireName, tool);
names.add(tool.wireName);
}
BY_WIRE_NAME = Map.copyOf(byName);
WIRE_NAMES = Set.copyOf(names);
}
/** The tool named {@code wireName}, or empty when this daemon registers no such tool. */
public static Optional<FleetTool> byWireName(String wireName) {
return Optional.ofNullable(BY_WIRE_NAME.get(wireName));
}
/** Every wire name this daemon registers — the canonical tool surface. */
public static Set<String> wireNames() {
return WIRE_NAMES;
}
}
@@ -525,17 +525,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
charter.toFile().deleteOnExit();
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
// is already at "instructions" — harmless only as long as this block runs first
// against a still-empty root, which is an ordering constraint nothing declared or
// tested. The skills writer just below, and the IDE-rules writer further down,
// both already use withArray (get-or-create) for exactly this reason; this was the
// one straggler. Proven load-bearing on the fleetd #393 merge: flipping this one
// call back to putArray left the whole suite green while silently deleting the
// charter entry whenever skills or IDE rules ran after it — an opencode member
// would launch with no role contract at all, worse than the bug #393 fixed, and
// nothing caught it. See OpenCodeLauncherTest's
// instructionsArrayHoldsCharterThenIdeRulesInOrder and
// instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder.
// is already at "instructions"; withArray gets-or-creates. All three writers on
// this array now use withArray so that write order stops being load-bearing for
// whoever adds a fourth.
//
// Be precise about what this particular line is worth, because an earlier version
// of this comment was wrong and claimed too much. THIS call is the one place where
// the two idioms are equivalent, and no test can tell them apart: it runs first,
// against a still-empty root, so there is never an existing node for putArray to
// replace. Measured on the #393 merge: flipping this one call back to putArray
// leaves the whole suite green (1618 tests), and always will. The edit is a
// readability and future-proofing change with no test behind it, and that is not a
// gap anyone can close.
//
// The ordering hazard is real, just not here. It is the LATER writers that can
// destroy earlier entries. Measured on the same merge: flipping the skills writer
// below to putArray deletes this charter entry and fails
// OpenCodeLauncherTest.instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder;
// flipping the IDE-rules writer fails that test and
// instructionsArrayHoldsCharterThenIdeRulesInOrder. Deleting this line altogether
// fails instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt — the
// assertion that a role contract reaches an opencode member at all, which is the
// hole that predates #393 and is what actually let the mutation hide.
root.withArray("instructions").add(charter.toAbsolutePath().toString());
}
@@ -3,6 +3,7 @@ package dev.ltms.fleet.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.Optional;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
@@ -21,46 +22,158 @@ import java.util.function.LongSupplier;
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*
* <h2>Escalation (fleetd #466)</h2>
* A flat cooldown does not fit every exhaustion. A backend that reports "out of capacity for the
* rest of the hour" recovers in one cooldown; a weekly subscription limit does not — it keeps
* reporting exhausted on every attempt made before the window resets, so a flat 30-minute cooldown
* (the default {@code cooldownNanos}) means roughly 336 pointless spawn attempts across a week, one
* every cooldown.
*
* <p><strong>This class only ever sees the exhaustion signal.</strong> Its only production caller is
* {@code Fleetd.exhaustionSink}, wired to fire on a {@code BACKEND_EXHAUSTED} classification alone.
* The daemon's other outage state — a credential "cooling off" after repeated non-exhaustion
* backend errors (an HTTP 5xx storm, say) — is a separate mechanism, {@code BackendOutagePolicy},
* with its own short fixed 60s cooldown and no repeat tracking. The two are never merged: escalating
* on a cooling-off signal would turn a transient 5xx storm into a multi-hour backoff, which is
* exactly the failure this ticket is not asking for. Confirmed by reading every call site of
* {@link #quarantine} — {@code BackendOutagePolicy} has its own {@code coolOff} method and never
* calls this one.
*
* <p><strong>Mechanism</strong> — the {@link #withEscalation} constructors track, per credential, how
* many times in a row {@link #quarantine} has been called without an intervening "quiet" gap.
* Each call computes {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at
* {@code maxCooldownNanos}. A call counts as a continuation of the same streak — {@code repeatCount}
* increments — when it arrives no more than one base {@code cooldownNanos} after the previous
* quarantine's deadline (this covers both "still quarantined" and "quarantine just expired and it
* was exhausted again immediately"); otherwise the streak resets and this call is treated as a fresh
* first occurrence at the base cooldown.
*
* <p><strong>Reset, honestly stated.</strong> The ideal reset signal is "the cooldown expired and the
* next attempt succeeded" — but nothing in this codebase reports a spawn success back to this class
* (checked: {@code SessionManager} and {@code CompositePeerLauncher} never call any method here
* except {@link #quarantine}/{@link #isQuarantined}/{@link #remainingSeconds}, none of which is a
* success hook). Lacking that signal, the reset used here is a time-based proxy: a base-cooldown's
* worth of quiet — no exhaustion report for that credential — since the last quarantine ended. It is
* not proof the credential started working again, only the best available evidence without adding an
* active probe, which is out of scope by the operator's own design constraint (no automatic probing
* of a limited backend).
*
* <p><strong>Ceiling.</strong> {@code maxCooldownNanos} bounds the growth — an unbounded backoff is a
* permanent, unrecoverable-without-a-restart outage, which would be worse than the flat-rate bug this
* escalation fixes. {@link #withEscalation(LongSupplier, long)} defaults the ceiling to
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown (12x the 1800s default ≈ 6 hours), so a
* chronically exhausted credential still gets re-tried roughly every 6 hours instead of every 30
* minutes — about a dozen attempts a week instead of ~336.
*
* <p><strong>Backward compatibility.</strong> The original two-argument {@link #BackendQuarantine(
* LongSupplier, long)} constructor is unchanged in behaviour: it is exactly {@code
* withEscalation}'s mechanism with {@code backoffMultiplier = 1.0} and {@code maxCooldownNanos =
* cooldownNanos}, which collapses the formula back to the original flat {@code now + cooldownNanos}
* on every call regardless of history. Every existing call site (roughly 20 across the test suite,
* plus {@link #none()}) keeps its current shape and behaviour unchanged.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
/** Default growth per consecutive exhaustion streak — see the class doc's Mechanism section. */
static final double DEFAULT_BACKOFF_MULTIPLIER = 2.0;
/** Default ceiling, expressed as a multiple of the base cooldown — see the class doc's Ceiling section. */
static final long DEFAULT_MAX_COOLDOWN_MULTIPLE = 12;
private final ConcurrentHashMap<String, QuarantineState> quarantines = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
private final double backoffMultiplier;
private final long maxCooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/** How many consecutive exhaustion reports a credential is on, and when the resulting cooldown ends. */
private record QuarantineState(int repeatCount, long deadlineNanos) {
}
/**
* Flat cooldown, unchanged from before fleetd #466 — every {@link #quarantine} call blocks the
* credential for exactly {@code cooldownNanos}, regardless of how many times it was called
* before. Equivalent to {@link #withEscalation} with no growth ({@code backoffMultiplier = 1.0})
* and a ceiling equal to the base cooldown, so it degrades to the identical {@code now +
* cooldownNanos} formula every call. Kept for the existing call sites that want a fixed cooldown
* (and for tests exercising the fixed-cooldown shape in isolation); production wiring uses
* {@link #withEscalation} instead.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
this(nowNanos, cooldownNanos, 1.0, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
/**
* Escalating cooldown (fleetd #466) — see the class doc's Mechanism/Reset/Ceiling sections.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be
* positive
* @param backoffMultiplier growth per consecutive exhaustion; must be {@code >= 1.0} ({@code 1.0}
* disables growth and is exactly the flat two-argument constructor)
* @param maxCooldownNanos ceiling on the escalated cooldown; must be {@code >= cooldownNanos}
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos) {
this(nowNanos, cooldownNanos, backoffMultiplier, maxCooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, double backoffMultiplier,
long maxCooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
if (backoffMultiplier < 1.0) {
throw new IllegalArgumentException("backoffMultiplier must be >= 1.0: " + backoffMultiplier);
}
if (maxCooldownNanos < cooldownNanos) {
throw new IllegalArgumentException(
"maxCooldownNanos must be >= cooldownNanos: " + maxCooldownNanos + " < " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.backoffMultiplier = backoffMultiplier;
this.maxCooldownNanos = maxCooldownNanos;
this.inert = inert;
}
/**
* Escalating cooldown with the fleetd #466 default shape: cooldown doubles
* ({@value #DEFAULT_BACKOFF_MULTIPLIER}x) per consecutive exhaustion streak, capped at
* {@value #DEFAULT_MAX_COOLDOWN_MULTIPLE}x the base cooldown. This is what production wiring
* ({@code Fleetd.main}) uses.
*
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos base cooldown, applied to a fresh (non-streak) exhaustion; must be positive
*/
public static BackendQuarantine withEscalation(LongSupplier nowNanos, long cooldownNanos) {
return new BackendQuarantine(nowNanos, cooldownNanos, DEFAULT_BACKOFF_MULTIPLIER,
cooldownNanos * DEFAULT_MAX_COOLDOWN_MULTIPLE, false);
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
return new BackendQuarantine(() -> 0L, 1, 1.0, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
* Quarantine {@code credentialId} starting now. On a flat instance (the two-argument
* constructor) this always blocks for exactly {@code cooldownNanos}, restarting the cooldown at
* full length on every call — a fresh refusal is fresh evidence the account is still exhausted,
* not a reason to let an earlier, shorter wait stand. On an escalating instance ({@link
* #withEscalation}) the cooldown grows with each call that arrives within one base cooldown of
* the previous deadline, and resets to the base cooldown once a call arrives after a longer gap
* — see the class doc.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
@@ -73,7 +186,13 @@ public final class BackendQuarantine {
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
long now = nowNanos.getAsLong();
quarantines.compute(credentialId, (id, prev) -> {
int repeatCount = (prev == null || now - prev.deadlineNanos() > cooldownNanos)
? 1
: prev.repeatCount() + 1;
return new QuarantineState(repeatCount, now + escalatedCooldownNanos(repeatCount));
});
}
/** Whether {@code credentialId} is quarantined right now. */
@@ -87,6 +206,49 @@ public final class BackendQuarantine {
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Remaining seconds together with which consecutive exhaustion this is (fleetd #466 scope item
* 2) — {@code repeatCount} 1 for a first occurrence, 2 for the second in a row, and so on; see
* {@link #quarantine}'s class-doc Mechanism section for exactly when a call continues a streak
* versus starts a fresh one.
*
* <p><strong>Read together, off the one {@link QuarantineState} entry {@link #quarantine} itself
* wrote</strong> — a single {@code quarantines.get(credentialId)}, never a separate lookup or a
* value re-derived from {@code remainingSeconds} (e.g. inverting {@link
* #escalatedCooldownNanos}). That inversion is not just extra work to avoid: once a streak has
* hit {@code maxCooldownNanos}, every further consecutive exhaustion reports the identical
* cooldown, so a derivation that starts from the cooldown value cannot tell the 4th repeat from
* the 9th — only the stored {@code repeatCount} can. This is the same rule {@code
* CompositePeerLauncher.modelGateState()} documents for its own gate/report pair: the report
* reads the exact accessor the behaviour reads, so it can never disagree with what actually
* happened (the fleetd #404/#422 lesson). {@code fleet_profiles}/{@code fleet_list}/{@code GET
* /profiles} all call this — never {@link #remainingSeconds} plus a second, independent count —
* for exactly that reason.
*
* @return empty when {@code credentialId} is not currently quarantined (including on {@link
* #none()}, which quarantines nothing)
*/
public Optional<Status> status(String credentialId) {
QuarantineState state = quarantines.get(credentialId);
if (state == null) {
return Optional.empty();
}
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
return remaining > 0
? Optional.of(new Status(toSecondsRoundedUp(remaining), state.repeatCount()))
: Optional.empty();
}
/**
* @param remainingSeconds seconds left on the quarantine, identical to {@link
* #remainingSeconds(String)}'s answer for the same credential at the
* same instant
* @param repeatCount 1 for a first occurrence, 2 for the second consecutive one, etc. —
* see {@link #status(String)}
*/
public record Status(long remainingSeconds, int repeatCount) {
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
@@ -95,8 +257,8 @@ public final class BackendQuarantine {
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
quarantines.forEach((credentialId, state) -> {
long remaining = state.deadlineNanos() - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
@@ -105,8 +267,14 @@ public final class BackendQuarantine {
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
QuarantineState state = quarantines.get(credentialId);
return state == null ? 0L : state.deadlineNanos() - nowNanos.getAsLong();
}
/** {@code cooldownNanos * backoffMultiplier ^ (repeatCount - 1)}, capped at {@code maxCooldownNanos}. */
private long escalatedCooldownNanos(int repeatCount) {
double raw = cooldownNanos * Math.pow(backoffMultiplier, repeatCount - 1);
return raw >= (double) maxCooldownNanos ? maxCooldownNanos : (long) raw;
}
private static long toSecondsRoundedUp(long nanos) {
@@ -302,9 +302,10 @@ public final class SessionManager implements TurnListener {
* with no copy and no error. Do NOT fuse these back together; the cost of an orphaned worktree
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/
private void release(String paneId, ReleaseCause cause) {
private MemberSession release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
}
/**
@@ -1064,16 +1065,39 @@ public final class SessionManager implements TurnListener {
* drain (see above), and a straggler must not buy the drain more time than the flag it lost the
* race against would have. In the ordinary case the sweep finds nothing and costs one empty
* {@link #roster()} call.
*
* <p>fleetd #512: a drain that releases every session cleanly used to log nothing at all — the
* only log calls in this method and {@link #drainSnapshot} sit on abnormal paths, so "nothing
* logged" was indistinguishable from "died on the first session". The {@code log.info} at the
* end below is a positive assertion that the drain actually finished, on the normal path,
* every time — including the all-zero case, which is a common and legitimate outcome (no
* members were live) and must still produce the line. Both {@link #drainSnapshot} passes (the
* main snapshot and the straggler sweep) are folded into the one line: a caller reading two
* lines could not tell a two-pass drain from two separate drains.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
draining.set(true);
drainSnapshot(roster(), deadline);
DrainTally tally = drainSnapshot(roster(), deadline);
List<MemberSession> stragglers = roster();
if (!stragglers.isEmpty()) {
log.warn("drain sweep found {} session(s) registered after the drain snapshot was "
+ "taken (raced past the shutdown guard); draining them too", stragglers.size());
drainSnapshot(stragglers, deadline);
tally = tally.plus(drainSnapshot(stragglers, deadline));
}
log.info("drain complete: released={} abandoned={} (still BUSY at the shutdown deadline)",
tally.released(), tally.abandoned());
}
/**
* Running count for one {@link #drainAll} invocation, folded across both {@link #drainSnapshot}
* passes (fleetd #512). {@code abandoned} counts sessions that were still {@code BUSY} at the
* moment they were released — i.e. the whole-drain deadline passed before they left {@code BUSY}
* on their own (see {@link #drainSnapshot}) — a subset of {@code released}, not additional to it.
*/
private record DrainTally(int released, int abandoned) {
private DrainTally plus(DrainTally other) {
return new DrainTally(released + other.released, abandoned + other.abandoned);
}
}
@@ -1081,8 +1105,12 @@ public final class SessionManager implements TurnListener {
* Drain exactly the sessions in {@code snapshot}, waiting out a {@code BUSY} one against the
* shared whole-drain {@code deadline} before releasing it. Shared by {@link #drainAll}'s main
* pass and its post-loop straggler sweep (fleetd #308) so both honor the same one budget.
* Returns how many sessions this pass released, and how many of those were still {@code BUSY}
* (abandoned mid-turn) at the moment of release.
*/
private void drainSnapshot(List<MemberSession> snapshot, long deadline) {
private DrainTally drainSnapshot(List<MemberSession> snapshot, long deadline) {
int released = 0;
int abandoned = 0;
for (MemberSession s : snapshot) {
try {
if (s.state() == MemberSession.State.BUSY) {
@@ -1100,11 +1128,16 @@ public final class SessionManager implements TurnListener {
}
}
}
release(s.paneId(), ReleaseCause.SHUTDOWN);
MemberSession removed = release(s.paneId(), ReleaseCause.SHUTDOWN);
released++;
if (removed != null && removed.state() == MemberSession.State.BUSY) {
abandoned++;
}
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
return new DrainTally(released, abandoned);
}
/**
@@ -0,0 +1,192 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.fail;
/**
* fleetd #498: {@code Fleetd.awaitHerdr} used to return a bare {@code boolean}, collapsing "the
* configured wait budget genuinely ran out" and "the waiting thread was interrupted, possibly
* milliseconds in" onto the same {@code false} — and the caller's log line printed only the
* configured budget, never how long the wait actually ran. This class covers both halves of the
* fix:
* <ul>
* <li>the seam — {@link Fleetd#awaitHerdr} itself, driven with an injected clock and a stub
* {@link HerdrClient}, one test per {@link Fleetd.HerdrWaitResult};</li>
* <li>the call site — {@link Fleetd#logHerdrWaitOutcomeAndShouldReap}, the exact decision {@code
* main} calls (extracted here because {@code main} itself boots the whole daemon and cannot
* be driven from a unit test), pinning the three distinct log messages it emits.</li>
* </ul>
* Every expected message below is a plain literal, not built from {@code HERDR_WAIT_SECONDS} or
* any other production constant — a test that derives its expectation the way the code does
* cannot see a change to either (fleetd #496's identical trap).
*/
class FleetdAwaitHerdrTest {
// ---- the seam: Fleetd.awaitHerdr ----------------------------------------------------------
@Test
void answeredReturnsImmediatelyWithZeroElapsedAndNeverPolls() {
HerdrStub herdr = new HerdrStub(0); // succeeds on the very first call
LongSupplier clock = fixedClock(1_000L);
AtomicBoolean polled = new AtomicBoolean(false);
Runnable poller = () -> polled.set(true);
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.ANSWERED, outcome.result());
assertEquals(0L, outcome.elapsedNanos(), "a fixed clock must measure zero elapsed time");
assertFalse(polled.get(), "herdr answering on the first try must never poll");
}
@Test
void deadlinePassedIsMeasuredNotAssumed() {
HerdrStub herdr = new HerdrStub(-1); // never succeeds
// call order inside awaitHerdr: start, then per failed attempt: deadline-check, elapsed-calc
ScriptedClock clock = new ScriptedClock(0L, 30_500_000_000L, 30_500_000_000L);
Runnable poller = () -> fail("the deadline was already exceeded on the first attempt — must not poll");
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.DEADLINE_PASSED, outcome.result());
assertEquals(30_500_000_000L, outcome.elapsedNanos(),
"elapsed must be the MEASURED clock delta, not the configured budget");
}
@Test
void interruptedIsDistinctFromDeadlinePassedAndPreservesTheInterruptFlag() {
HerdrStub herdr = new HerdrStub(-1); // never succeeds
// start=0, deadline-check returns 500ms (well under the 30s budget) -> not deadline-passed,
// then the poller interrupts, and the elapsed-calc call returns 750ms.
ScriptedClock clock = new ScriptedClock(0L, 500_000_000L, 750_000_000L);
Runnable poller = () -> Thread.currentThread().interrupt();
try {
Fleetd.HerdrAwaitOutcome outcome = Fleetd.awaitHerdr(herdr, clock, poller);
assertEquals(Fleetd.HerdrWaitResult.INTERRUPTED, outcome.result());
assertEquals(750_000_000L, outcome.elapsedNanos(),
"elapsed must be measured even when the wait ends via interruption, not the deadline");
assertTrue(Thread.currentThread().isInterrupted(),
"the interrupt flag the old code re-set must still be set on return");
} finally {
Thread.interrupted(); // clear it so it cannot leak into another test on this thread
}
}
// ---- the call site: Fleetd.logHerdrWaitOutcomeAndShouldReap -------------------------------
@Test
void answeredLogsNothingAndSaysReap() {
try (CapturedLog log = attach()) {
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.ANSWERED, 0L));
assertTrue(shouldReap, "only ANSWERED should tell main to reap orphan workers");
assertEquals(0, log.events().size(), "the answered path logs nothing itself");
}
}
@Test
void deadlinePassedLogsConfiguredAndMeasuredElapsedTogether() {
try (CapturedLog log = attach()) {
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.DEADLINE_PASSED, 30_500_000_000L));
assertFalse(shouldReap, "a deadline-passed wait must not tell main to reap");
assertEquals(1, log.events().size());
ILoggingEvent event = log.events().getFirst();
assertEquals(Level.WARN, event.getLevel());
assertEquals("herdr did not answer within the configured wait (configured=30s "
+ "elapsed=30500ms) — starting anyway; /healthz will report degraded until it "
+ "comes up. Orphaned worker panes (if any) were NOT reaped.",
event.getFormattedMessage());
}
}
@Test
void interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed() {
try (CapturedLog log = attach()) {
// 3ms: the ticket's own example of "a few milliseconds in", not the 30s budget.
boolean shouldReap = Fleetd.logHerdrWaitOutcomeAndShouldReap(
new Fleetd.HerdrAwaitOutcome(Fleetd.HerdrWaitResult.INTERRUPTED, 3_000_000L));
assertFalse(shouldReap, "an interrupted wait must not tell main to reap");
assertEquals(1, log.events().size());
ILoggingEvent event = log.events().getFirst();
assertEquals(Level.WARN, event.getLevel());
String message = event.getFormattedMessage();
assertEquals("herdr wait was interrupted before the configured wait ran out "
+ "(configured=30s elapsed=3ms) — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
message);
assertFalse(message.contains("did not answer"),
"an interrupted wait must not be reported as if herdr failed to answer within the budget");
}
}
// ---- fixtures --------------------------------------------------------------------------
/** Always returns the same value, i.e. a clock that measures zero elapsed time. */
private static LongSupplier fixedClock(long value) {
return () -> value;
}
/** Returns each value in order, then repeats the last one for any call beyond the list. */
private static final class ScriptedClock implements LongSupplier {
private final long[] values;
private int index;
ScriptedClock(long... values) {
this.values = values;
}
@Override
public long getAsLong() {
long v = values[Math.min(index, values.length - 1)];
if (index < values.length - 1) {
index++;
}
return v;
}
}
/** Fails {@code failuresBeforeSuccess} times, then succeeds forever; {@code -1} never succeeds. */
private static final class HerdrStub implements HerdrClient {
private final int failuresBeforeSuccess;
private int calls;
HerdrStub(int failuresBeforeSuccess) {
this.failuresBeforeSuccess = failuresBeforeSuccess;
}
@Override
public JsonNode call(String method, Object params) throws HerdrException {
calls++;
if (failuresBeforeSuccess < 0 || calls <= failuresBeforeSuccess) {
throw new HerdrException("herdr not up yet");
}
return null;
}
@Override
public void close() {
}
}
private static CapturedLog attach() {
return CapturedLog.at(Fleetd.class, Level.DEBUG);
}
}
@@ -0,0 +1,65 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
* it says nothing about which one {@code main} actually calls.
*
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
* every one of them builds its own instance directly. That silent regression is exactly the shape
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
* factory choice, following their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
* constructor. It does not prove that call actually executes at startup (no test here starts the
* daemon), and it does not prove the escalation reaches a real backend or credential — only
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
* the wiring runs.
*/
class FleetdBackendQuarantineWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"Fleetd.main's BackendQuarantine local must still be built from "
+ "BackendQuarantine.withEscalation(System::nanoTime, "
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
+ "proves the escalation itself works — this source check is what must go red instead. "
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
+ "flat ~30-minute cooldown, about 336 times across the week.");
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
assertFalse(source.contains(
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
}
}
@@ -0,0 +1,106 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474: proves the exact wiring {@code Fleetd.main} uses to construct its live {@code
* ConfigRef} — {@code new ConfigRef(configPath, cfg, Fleetd::assertChartersNameOnlyRegisteredTools)}
* — actually makes {@link ConfigRef#reload()} refuse a charter that names an MCP tool the server
* does not register, the same way {@code Fleetd.main} itself refuses one at startup (see {@code
* FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool}).
*
* <p>{@code dev.ltms.fleet.config.ConfigRefTest} pins the same behaviour through a locally-built
* {@code Consumer<FleetConfig>} adapter that calls the same production {@code CharterToolSurface}
* method, because that test lives in {@code dev.ltms.fleet.config} and cannot see {@code
* Fleetd#assertChartersNameOnlyRegisteredTools} (package-private to {@code dev.ltms.fleet}). This
* class is the companion proof that lives where the real method reference is visible, so the literal
* expression {@code Fleetd::assertChartersNameOnlyRegisteredTools} — not just an equivalent — is
* what gets exercised. {@code Fleetd.main} itself cannot be driven this far in a unit test: every
* fixture in {@code FleetdStartupValidationTest} is deliberately invalid so {@code main} throws
* before opening a socket, binding Javalin, or doing anything else with a real side effect, so a
* test cannot get {@code main} far enough to hold a running daemon it could then reload — this test
* builds the {@code ConfigRef} the same way {@code main} does and drives {@link ConfigRef#reload()}
* directly instead, the same shape {@code FleetdExhaustionDetectionArmedWiringTest} and its
* siblings already use for the rest of {@code Fleetd.main}'s wiring.
*/
class FleetdConfigRefCharterToolSurfaceWiringTest {
private static final String BASE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
void reloadRefusesACharterNamingAnUnregisteredToolThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
FleetConfig before = config.get();
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""");
ConfigRef.Outcome out = config.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"), out.error());
assertTrue(out.error().contains("bridge_send"), out.error());
assertSame(before, config.get());
}
@Test
void reloadAcceptsACharterNamingOnlyRegisteredToolsThroughFleetdsOwnWiring(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, BASE + """
fleet:
charters:
dev: old charter
""");
ConfigRef config = new ConfigRef(f, FleetConfig.load(f),
Fleetd::assertChartersNameOnlyRegisteredTools);
Files.writeString(f, BASE + """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
""");
ConfigRef.Outcome out = config.reload();
assertTrue(out.applied());
assertEquals("config reloaded", out.summary());
}
}
@@ -0,0 +1,80 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #474 follow-up: {@code Fleetd.main} builds its live {@code ConfigRef} from the
* three-argument constructor, {@code new ConfigRef(configPath, cfg,
* Fleetd::assertChartersNameOnlyRegisteredTools)}, so a reload runs the same charter tool-surface
* gate startup does (see {@link dev.ltms.fleet.config.ConfigRef}'s class doc, "fleetd #474" bullet).
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* three-argument constructor and the {@code Fleetd.assertChartersNameOnlyRegisteredTools} adapter
* work correctly together — both build their OWN {@code ConfigRef} with that constructor, so neither
* proves {@code main} still CHOOSES the three-argument form over the plain two-argument {@code new
* ConfigRef(configPath, cfg)}.
*
* <p>Measured directly: reverting {@code Fleetd.java}'s {@code config} local to the two-argument
* constructor compiles with 0 errors and leaves the entire 1633-test suite green — including every
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* ConfigRef} and never runs {@code main} — a green result here proves only that the exact text
* {@code main} contains is the three-argument construction with {@code
* Fleetd::assertChartersNameOnlyRegisteredTools}. It does not prove that call actually executes at
* startup (no test here starts the daemon), and it does not prove the reload gate itself works —
* only {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} prove the
* behaviour; only a live daemon proves the wiring runs.
*/
class FleetdConfigRefWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] main still builds config from the three-argument ConfigRef constructor")
void mainStillWiresTheThreeArgumentConfigRefConstructor() throws Exception {
String source = fleetdSource();
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"fleetdSource() did not read anything usable — src/main/java/dev/ltms/fleet/Fleetd.java "
+ "did not come back containing its own class declaration. The assertFalse below "
+ "would pass vacuously on a broken read; fix the read before trusting this test.");
assertTrue(source.contains(
"ConfigRef config = new ConfigRef(configPath, cfg, "
+ "Fleetd::assertChartersNameOnlyRegisteredTools);"),
"Fleetd.main's config local must still be built from the three-argument ConfigRef "
+ "constructor, with Fleetd::assertChartersNameOnlyRegisteredTools as "
+ "extraValidation. Reverting to the plain two-argument constructor (fleetd #474's "
+ "measured M2 regression) compiles with 0 errors and leaves the whole suite green — "
+ "including ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest, because "
+ "neither builds its ConfigRef through main — this source check is what must go "
+ "red instead. A reverted daemon would accept, through a reload with no restart, "
+ "exactly the charter that #469/#474 already refuse at startup.");
// Negative form of the same check: the pre-#474 two-argument call, if it ever reappears at
// this declaration, must not be mistaken for the three-argument one by a looser
// positive-only check — this is the M2 mutation this test exists to kill.
assertFalse(source.contains("ConfigRef config = new ConfigRef(configPath, cfg);"),
"main's config local must never regress to the plain two-argument ConfigRef "
+ "constructor — that drops the reload-path charter check (fleetd #474's measured "
+ "M2 mutation) with no other test catching it");
}
}
@@ -1,15 +1,12 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.msg.LeadMailbox;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
@@ -52,29 +49,22 @@ class FleetdLeadMailboxSelectionTest {
}
}
private static ListAppender<ILoggingEvent> captureFleetdLogs() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static String joined(ListAppender<ILoggingEvent> appender, Level level) {
return appender.list.stream().filter(e -> e.getLevel() == level)
private static String joined(CapturedLog captured, Level level) {
return captured.events().stream().filter(e -> e.getLevel() == level)
.map(ILoggingEvent::getFormattedMessage).reduce("", (a, b) -> a + "\n" + b);
}
@Test
void noCoordinatorBlockLeavesTheFeatureOffSilently() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
try (var captured = CapturedLog.of(Fleetd.class)) {
var opener = new RecordingOpener();
assertNull(Fleetd.openLeadMailbox(null, Map.of(), opener));
assertNull(Fleetd.openLeadMailbox(null, Map.of(), opener));
assertNull(opener.offeredUri, "nothing configured means nothing is opened");
assertEquals("", joined(appender, Level.WARN),
"an opt-in feature nobody asked for must not warn on every boot");
assertNull(opener.offeredUri, "nothing configured means nothing is opened");
assertEquals("", joined(captured, Level.WARN),
"an opt-in feature nobody asked for must not warn on every boot");
}
}
@Test
@@ -114,31 +104,33 @@ class FleetdLeadMailboxSelectionTest {
@Test
void warnsAndStaysOffWhenSelfIdIsMissing() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null, null);
try (var captured = CapturedLog.of(Fleetd.class)) {
var opener = new RecordingOpener();
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, null, null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener));
assertNull(opener.offeredUri, "a mailbox with no owning coord-id has no queue to declare");
String warns = joined(appender, Level.WARN);
assertTrue(warns.contains("coordinator.selfId"), () -> "say which key is missing: " + warns);
assertFalse(warns.contains(SECRET), () -> "the URI's password must never be logged: " + warns);
assertNull(opener.offeredUri, "a mailbox with no owning coord-id has no queue to declare");
String warns = joined(captured, Level.WARN);
assertTrue(warns.contains("coordinator.selfId"), () -> "say which key is missing: " + warns);
assertFalse(warns.contains(SECRET), () -> "the URI's password must never be logged: " + warns);
}
}
@Test
void warnsAndStaysOffWhenTheBrokerIsUnreachableAtBoot() {
var appender = captureFleetdLogs();
var opener = new RecordingOpener();
opener.unreachable = true;
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
try (var captured = CapturedLog.of(Fleetd.class)) {
var opener = new RecordingOpener();
opener.unreachable = true;
var coordinator = new FleetConfig.Coordinator(RESOLVED_URI, null, "mac-opus", null, null);
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
"a down coordination broker turns the feature off; it must never take the daemon down");
assertNull(Fleetd.openLeadMailbox(coordinator, Map.of(), opener),
"a down coordination broker turns the feature off; it must never take the daemon down");
String warns = joined(appender, Level.WARN);
assertTrue(warns.contains("coord.example"), () -> "name the host that failed: " + warns);
assertFalse(warns.contains(SECRET), () -> "with credentials stripped: " + warns);
assertTrue(warns.contains("Connection refused"), () -> "and the real reason: " + warns);
String warns = joined(captured, Level.WARN);
assertTrue(warns.contains("coord.example"), () -> "name the host that failed: " + warns);
assertFalse(warns.contains(SECRET), () -> "with credentials stripped: " + warns);
assertTrue(warns.contains("Connection refused"), () -> "and the real reason: " + warns);
}
}
}
@@ -0,0 +1,89 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
*
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
* DOES with that argument once inside its own body.
*
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
* catch this wiring dropping out" for the whole factory, including the lambda {@code
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
* one subsumes the other — keep both.
*/
class FleetdLeadRolloverWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
void unrelatedAnchorStillPresent() throws Exception {
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
// test that only asserts "contains X" would report a false pass for the wrong reason if X
// happened to match. Asserting an unrelated, structurally distant string first proves the
// read actually pulled real file content.
String source = fleetdSource();
assertTrue(source.contains("public final class Fleetd"),
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
}
@Test
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
void mainStillCallsTheLeadRolloverFactory() throws Exception {
String source = fleetdSource();
assertTrue(source.contains(
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
+ "its arguments for something that still compiles (e.g. null in place of "
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
+ "relative handoverPath against the calling lead's own workspace.");
}
@Test
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
void factoryGatesOnConfigPresence() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
}
}
@@ -0,0 +1,163 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.lead.LeadRollover;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* fleetd #480 relative-handover-path follow-up, correction round: a BEHAVIOURAL test of {@link
* Fleetd#leadRollover}'s own body — the terminal → lead-name → {@code Leader.cwd()} lookup it
* builds — not another source-text pin.
*
* <p>{@code FleetdLeadRolloverWiringTest} (a plain string read) still earns its keep: it pins the
* call site's ARGUMENT LIST, so a mutation that drops {@code leads} back out of the call, or
* swaps it for {@code Map::of}, still goes red there. But nothing before this class exercised the
* LAMBDA BODY {@code leadRollover(...)} builds — the terminal-to-workspace {@code
* Function<String,String>} that {@link LeadRollover#open} actually calls. Every prior test either
* exercised {@code LeadRollover} directly with a hand-built lookup ({@code LeadRolloverTest}), or
* read source text without ever calling the factory ({@code FleetdLeadRolloverWiringTest}) — so a
* mutation that breaks the lookup ITSELF (e.g. always resolving no lead, or reading a snapshot
* instead of the live supplier) left every existing test green while the daemon's real wiring
* silently reintroduced the exact bug this whole ticket fixes: a relative {@code handoverPath}
* resolving against the daemon's own {@code cwd} instead of the calling lead's.
*
* <p>{@code Fleetd.leadRollover(...)} is package-private and {@code static}, so this test — living
* in the same {@code dev.ltms.fleet} package — calls it directly, exactly the way {@code
* FleetdExhaustionDetectionArmedWiringTest} and {@code FleetdCapacitySourceWiringTest} already call
* other package-private startup factories with a real {@link ConfigRef} built from a temp {@code
* fleetd.yaml} (never the gitignored live one).
*
* <p><b>Proved against the mutation it exists to catch.</b> Before this class was added, mutating
* {@code leadRollover(...)}'s lambda body to {@code String leadName = null;} (always "no lead
* found", which forces every relative {@code handoverPath} onto the {@code user.dir} fallback —
* i.e. the original bug) left the full suite green: {@code Tests run: 1669, Failures: 0}. With
* {@link #relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd()} added, the same one-line
* mutation now fails that test (it asserts the resolved path equals the configured lead's {@code
* cwd}, which the mutant can never produce) — see this ticket's fleet_reply history for both runs.
*/
class FleetdLeadRolloverWorkspaceLookupTest {
private static AgentControl fakeAgents() {
return new AgentControl(new FakeHerdr());
}
@Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
void relativeHandoverPathResolvesAgainstTheLeadsConfiguredCwd(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> Map.of("term_opus", "opus"));
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object");
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
String expected = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expected, pending.handoverPath(),
"a relative handoverPath must resolve against the CALLING lead's configured cwd "
+ "through the REAL Fleetd.leadRollover(...) wiring — not the daemon's own "
+ "working directory. This is the exact axis that was proven uncovered: "
+ "mutating the factory's lambda body to always report \"no lead found\" "
+ "left every prior test green.");
}
@Test
@DisplayName("[BEHAVIOURAL] a terminal not present in the live lead-terminal map falls back "
+ "to the daemon's own user.dir")
void terminalNotInLiveMapFallsBackToUserDir(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
leadRollover:
handoverPath: handover.md
""");
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
String expected = Path.of(System.getProperty("user.dir")).resolve("handover.md")
.normalize().toString();
assertEquals(expected, pending.handoverPath(),
"a lead not present in the live terminal→name map must fall back to the daemon's "
+ "own user.dir — the same fallback LeadLauncher#launch already uses for a "
+ "lead with no configured cwd. This is deliberate, pinned behaviour, not an "
+ "accident of the null-check chain.");
}
@Test
@DisplayName("[BEHAVIOURAL] the terminal→lead-name lookup is read LIVE on every open() call, "
+ "never snapshotted at Fleetd.leadRollover(...) construction time")
void workspaceLookupIsReadLiveNotSnapshotted(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
leadRollover:
handoverPath: handover.md
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
// EMPTY at the moment leadRollover(...) is called. A lookup captured (snapshotted) here
// instead of read live through the supplier on every call would never see the entry added
// below — exactly the natural mistake to make, since leads are discovered by a live tab
// scan that runs AFTER this factory is constructed at startup.
Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
() -> liveLeadTerminals);
assertNotNull(rollover);
// The lead is "discovered" only now — mutate the SAME backing map the supplier reads from.
liveLeadTerminals.put("term_opus", "opus");
LeadRollover.PendingRollover pending = rollover.open("term_opus", "test");
String expected = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expected, pending.handoverPath(),
"the terminal→lead-name lookup must be read LIVE on every open() call — a lead "
+ "discovered by the tab scan AFTER Fleetd.leadRollover(...) was constructed "
+ "must still resolve correctly, not only one that was already live at "
+ "construction time");
}
}
@@ -1,15 +1,13 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.msg.AmqpReplyInbox;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.net.ServerSocket;
import java.util.List;
@@ -50,43 +48,34 @@ class FleetdReplyInboxSelectionTest {
}
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
private static CapturedLog attach() {
// logback-test.xml pins dev.ltms.fleet to WARN; raise it so INFO selection lines are captured.
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
return CapturedLog.at(Fleetd.class, Level.INFO);
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
private static void assertNoLogContains(ListAppender<ILoggingEvent> appender, String secret) {
assertTrue(appender.list.stream().noneMatch(e -> e.getFormattedMessage().contains(secret)),
private static void assertNoLogContains(List<ILoggingEvent> events, String secret) {
assertTrue(events.stream().noneMatch(e -> e.getFormattedMessage().contains(secret)),
"no log line may contain the resolved URI's password");
}
@Test
void uriEnvSetAndPresentSelectsAmqpWithTheResolvedUri() {
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, events, inbox) -> {
assertEquals(RESOLVED_URI, opener.offeredUri,
"the daemon must connect with the value resolved from uriEnv — selection, not just parse");
assertEquals(opener.inbox, inbox, "the AMQP opener's inbox is what is selected");
assertNoLogContains(appender, SECRET);
assertNoLogContains(events, SECRET);
});
}
@Test
void uriEnvSetButVariableMissingFallsBackToInMemoryAndWarns() {
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
recording(broker, Map.of(), false, (opener, appender, inbox) -> {
recording(broker, Map.of(), false, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox);
assertNull(opener.offeredUri, "AMQP must never be attempted when the variable is missing");
assertTrue(hasWarnContaining(appender, "LAVINMQ_URI") && hasWarnContaining(appender, "DISABLED"),
assertTrue(hasWarnContaining(events, "LAVINMQ_URI") && hasWarnContaining(events, "DISABLED"),
"a missing uriEnv variable must warn loudly, not fail silently");
});
}
@@ -94,10 +83,10 @@ class FleetdReplyInboxSelectionTest {
@Test
void uriEnvSetButVariableBlankFallsBackToInMemoryAndWarns() {
FleetConfig.Broker broker = new FleetConfig.Broker("amqp://user:lame@old:5672/", "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", " "), false, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", " "), false, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox);
assertNull(opener.offeredUri, "a blank env value must not select AMQP, not even via the literal uri");
assertTrue(hasWarnContaining(appender, "LAVINMQ_URI"),
assertTrue(hasWarnContaining(events, "LAVINMQ_URI"),
"a blank uriEnv value must warn, and must not fall back to the literal uri");
});
}
@@ -106,22 +95,22 @@ class FleetdReplyInboxSelectionTest {
void bothUriAndUriEnvSetUriEnvWinsDeterministically() {
FleetConfig.Broker broker
= new FleetConfig.Broker("amqp://user:oldpw@old.example:5672/", "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), false, (opener, events, inbox) -> {
assertEquals(RESOLVED_URI, opener.offeredUri,
"uriEnv must win over uri, deterministically, every run");
assertTrue(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("broker.uri is ignored")),
assertTrue(events.stream().anyMatch(e -> e.getFormattedMessage().contains("broker.uri is ignored")),
"must log that the literal uri is ignored when uriEnv is set");
assertNoLogContains(appender, SECRET);
assertNoLogContains(appender, "oldpw");
assertNoLogContains(events, SECRET);
assertNoLogContains(events, "oldpw");
});
}
@Test
void unreachableBrokerStartsDaemonWithInMemoryInboxAndLoudWarning() {
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), true, (opener, appender, inbox) -> {
recording(broker, Map.of("LAVINMQ_URI", RESOLVED_URI), true, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox, "an unreachable broker must NOT stop the daemon");
String warn = appender.list.stream()
String warn = events.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.reduce("", (a, b) -> a + "\n" + b)
@@ -129,16 +118,16 @@ class FleetdReplyInboxSelectionTest {
assertTrue(warn.contains("durable") && warn.contains("soft-state"),
"the warning must say exactly what was lost: durable delivery off, replies soft-state");
assertTrue(!warn.contains(SECRET), "the failing URI must be logged with credentials stripped");
assertNoLogContains(appender, SECRET);
assertNoLogContains(events, SECRET);
});
}
@Test
void noBrokerConfiguredStaysQuietInMemory() {
FleetConfig.Broker broker = null;
recording(broker, Map.of(), false, (opener, appender, inbox) -> {
recording(broker, Map.of(), false, (opener, events, inbox) -> {
assertInstanceOf(InMemoryReplyInbox.class, inbox);
assertTrue(appender.list.stream().noneMatch(e -> e.getLevel() == Level.WARN),
assertTrue(events.stream().noneMatch(e -> e.getLevel() == Level.WARN),
"no broker configured must keep the existing QUIET in-memory path — no warning");
});
}
@@ -152,21 +141,18 @@ class FleetdReplyInboxSelectionTest {
closedPort = s.getLocalPort();
}
FleetConfig.Broker broker = new FleetConfig.Broker(null, "LAVINMQ_URI", null);
ListAppender<ILoggingEvent> appender = attach();
try {
try (CapturedLog log = attach()) {
ReplyInbox inbox = Fleetd.selectReplyInbox(
broker, Map.of("LAVINMQ_URI", "amqp://user:" + SECRET + "@127.0.0.1:" + closedPort + "/vh"),
AmqpReplyInbox::open);
assertInstanceOf(InMemoryReplyInbox.class, inbox,
"a genuinely unreachable broker (real AmqpReplyInbox::open) must fall back to in-memory");
} finally {
detach(appender);
assertNoLogContains(log.events(), SECRET);
}
assertNoLogContains(appender, SECRET);
}
private boolean hasWarnContaining(ListAppender<ILoggingEvent> appender, String fragment) {
return appender.list.stream().anyMatch(e ->
private boolean hasWarnContaining(List<ILoggingEvent> events, String fragment) {
return events.stream().anyMatch(e ->
e.getLevel() == Level.WARN && e.getFormattedMessage().contains(fragment));
}
@@ -174,18 +160,14 @@ class FleetdReplyInboxSelectionTest {
private void recording(FleetConfig.Broker broker, Map<String, String> env, boolean unreachable, Check check) {
RecordingAmqp opener = new RecordingAmqp();
opener.unreachable = unreachable;
ListAppender<ILoggingEvent> appender = attach();
ReplyInbox inbox;
try {
inbox = Fleetd.selectReplyInbox(broker, env, opener);
} finally {
detach(appender);
try (CapturedLog log = attach()) {
ReplyInbox inbox = Fleetd.selectReplyInbox(broker, env, opener);
check.run(opener, log.events(), inbox);
}
check.run(opener, appender, inbox);
}
@FunctionalInterface
private interface Check {
void run(RecordingAmqp opener, ListAppender<ILoggingEvent> appender, ReplyInbox inbox);
void run(RecordingAmqp opener, List<ILoggingEvent> events, ReplyInbox inbox);
}
}
@@ -95,6 +95,28 @@ class FleetdStartupValidationTest {
""", "architetc");
}
/**
* fleetd #469: closes the gap left by #464's {@code CharterToolSurfaceTest}, which wrote its
* own charter into a {@code @TempDir} fixture and so could never fail on anything anyone wrote
* into the live {@code fleetd.yaml}. {@code bridge_send} is CB-634's own motivating example — a
* tool name the rename removed — and the charter key ({@code dev}) is a real role wire name, so
* this fixture passes {@code cfg.validateAll()}'s charter check (key valid, text non-blank) and
* is refused only by the new {@code CharterToolSurface} call right after it. The failure message
* must name both the charter key and the unknown tool.
*/
@Test
void mainRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charter-tool-surface.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
""", "bridge_send");
}
@Test
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "members.yaml", """
@@ -1,15 +1,12 @@
package dev.ltms.fleet.auth;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import static org.junit.jupiter.api.Assertions.*;
@@ -24,28 +21,21 @@ import static org.junit.jupiter.api.Assertions.*;
class AuditLogTest {
private final ObjectMapper mapper = new ObjectMapper();
private ListAppender<ILoggingEvent> appender;
private ch.qos.logback.classic.Logger auditLogger;
private CapturedLog auditLog;
@BeforeEach
void attach() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
auditLogger = ctx.getLogger("audit");
appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
auditLogger.addAppender(appender);
auditLogger.setLevel(Level.INFO);
auditLog = CapturedLog.at("audit", Level.INFO);
}
@AfterEach
void detach() {
auditLogger.detachAppender(appender);
auditLog.close();
}
private JsonNode onlyRecord() throws Exception {
assertEquals(1, appender.list.size(), "exactly one audit line expected");
String line = appender.list.getFirst().getFormattedMessage();
assertEquals(1, auditLog.events().size(), "exactly one audit line expected");
String line = auditLog.events().getFirst().getFormattedMessage();
return mapper.readTree(line); // throws if the line is not valid JSON
}
@@ -186,6 +186,43 @@ class CallerResolverTest {
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 99, "BEARER s3cret").role());
}
// ── fleetd #505: a herdr error DURING THE SCAN must not be conflated with "not a worker" ──────
// #317 (above) covers a failed lsof lookup. This is the other input to the same decision: the
// lsof lookup succeeds (a real pid), but PaneLocator's own pane scan hits a herdr error on the
// pane that owns that pid — so c.resolved() is true and c.terminal() is null, exactly like a
// real primary. c.scanComplete() is what tells them apart.
/**
* The discriminating case named in the ticket: the error must land on the pane that DOES own
* the caller's pid, or the test proves nothing (any other pane's failure is invisible to the
* scan's outcome, since a match found elsewhere is definitive regardless).
*/
@Test
void aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary() {
FakeHerdr failing = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
ConnectionIdentity incomplete = new ConnectionIdentity(new PaneLocator(failing), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(incomplete).resolve("127.0.0.1", 55555, null);
assertEquals(Role.ANONYMOUS, p.role(),
"an incomplete pane scan must never be read as a clean negative and promoted to primary");
}
/**
* The companion invariant: a herdr error on a DIFFERENT, non-owning pane must not turn every
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
*/
@Test
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() {
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
@Test
void aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust() {
// Defence in depth: startup already refuses this pairing (validateAuthExposure), but if a
@@ -1,10 +1,13 @@
package dev.ltms.fleet.config;
import dev.ltms.fleet.mcp.CharterToolSurface;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.function.Consumer;
import static org.junit.jupiter.api.Assertions.*;
@@ -38,6 +41,23 @@ class ConfigRefTest {
return new ConfigRef(f, FleetConfig.load(f));
}
/**
* The exact {@code Consumer<FleetConfig>} {@code Fleetd.main} wires into {@code ConfigRef}'s
* constructor as {@code extraValidation} (fleetd #474) — an adapter from {@code FleetConfig} to
* the raw charter map {@link CharterToolSurface#assertChartersNameOnlyRegisteredTools} takes.
* Built here rather than referencing {@code dev.ltms.fleet.Fleetd} directly, because that method
* is package-private to {@code dev.ltms.fleet} and this test lives in {@code
* dev.ltms.fleet.config} — but it calls the SAME production {@link CharterToolSurface} method
* {@code Fleetd} calls, so this proves the real check runs on reload, not a stand-in for it.
*/
private static final Consumer<FleetConfig> CHARTER_TOOL_SURFACE = cfg ->
CharterToolSurface.assertChartersNameOnlyRegisteredTools(
cfg.fleet() == null ? Map.of() : cfg.fleet().charters());
private static ConfigRef refForWithCharterToolSurface(Path f) {
return new ConfigRef(f, FleetConfig.load(f), CHARTER_TOOL_SURFACE);
}
@Test
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
@@ -124,6 +144,87 @@ class ConfigRefTest {
assertSame(before, ref.get());
}
/**
* fleetd #474: {@code validateAll()} (via {@code validateCharters()}) only checks that a
* charter's KEY is a role wire name and its text is non-blank — it never looks at what the text
* names, so a charter naming {@code bridge_send} (the pre-CB-634 name, removed from the tool
* surface — #469's own motivating example) passes {@code validateAll()} and used to be applied
* on reload with nothing refusing it, even though the identical charter refuses {@code
* Fleetd.main} at startup ({@code FleetdStartupValidationTest
* #mainRefusesACharterNamingAnUnregisteredTool}). This is the reload-path proof: it drives
* {@link ConfigRef#reload()} itself (not a direct call to {@link
* CharterToolSurface#assertChartersNameOnlyRegisteredTools}), through the exact {@code
* extraValidation} wiring {@code Fleetd.main} uses, and the failure message must name both the
* charter key and the unknown tool — the same information the startup failure gives (ticket
* acceptance criterion 1).
*/
@Test
void aReloadRefusesACharterNamingAnUnregisteredTool(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
FleetConfig before = ref.get();
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through bridge_send.
"""));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.error().contains("dev"),
"expected the charter key 'dev' in the refusal, got: " + out.error());
assertTrue(out.error().contains("bridge_send"),
"expected the unknown tool 'bridge_send' in the refusal, got: " + out.error());
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
// The running config must not move at all — half-applying this would leave the daemon in a
// state that could never have booted, exactly the outcome ConfigRef.java's class doc warns
// a cold-key refusal must avoid, and this check must avoid the same way.
assertSame(before, ref.get());
assertEquals("Send the final handoff through fleet_reply.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* fleetd #474 acceptance criterion 3, the positive case: a reload whose charter names only
* registered tools must still be ACCEPTED. An inverted filter (one that refuses every charter,
* or refuses on any {@code fleet_*}/{@code bridge_*} token regardless of registration) would pass
* the refusal test above alone — #469's own M5 mutation cell showed exactly that shape surviving
* a negative-only suite. This is the test that catches it.
*/
@Test
void aReloadAcceptsACharterNamingOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
dev: old charter, no tool names
"""));
ConfigRef ref = refForWithCharterToolSurface(f);
Files.writeString(f, yaml("""
fleet:
charters:
dev: |
Send the final handoff through fleet_reply, using fleet_send to delegate.
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals("config reloaded", out.summary());
assertEquals("Send the final handoff through fleet_reply, using fleet_send to delegate.\n",
ref.get().fleet().charterFor(dev.ltms.fleet.peer.MemberRole.DEV));
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
@@ -87,6 +87,17 @@ class ConfigRefTopLevelCoverageTest {
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
* object anywhere — there is no frozen half, so it belongs here whole rather than in
* {@code SPLIT_KEYS}.</li>
* <li>{@code leadRollover} (fleetd #480) — {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} and reads every field fresh on each
* {@code open()}/{@code confirm()} call, the same {@code () -> config.get().x()} shape
* {@code fleet}/{@code placement}/{@code models} use — see {@code ConfigRef}'s class doc
* Hot bullet. The only restart-only fact is structural, not a value going stale:
* {@code Fleetd.java} decides whether to construct the {@code LeadRollover} object at
* all off the startup snapshot (the same presence gate {@code leadHeartbeat:} uses), so
* a block ADDED where it was absent at boot needs a restart before anything exists to
* call — the same fact already true of adding a brand-new {@code profiles:} entry, which
* does not stop {@code profiles}' own hot sub-fields (weight/maxLoad/credentialId/
* exhaustedPattern) from being genuinely hot.</li>
* </ul>
*
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
@@ -101,7 +112,7 @@ class ConfigRefTopLevelCoverageTest {
* while this test stayed green throughout.</p>
*/
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
Set.of("placement", "memberCredentials", "memberLoginShell", "models");
Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover");
@Test
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
@@ -128,7 +139,7 @@ class ConfigRefTopLevelCoverageTest {
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models"), hot,
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models", "leadRollover"), hot,
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
+ "live off the config supplier, never because adding it makes this test "
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
@@ -111,6 +111,9 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-a"))));
// fleetd #480: leadRollover joins placement/memberCredentials/memberLoginShell as
// hot-excluded — never compared by any changed*Keys method, so it stays null like them.
v.put("leadRollover", null);
assertNamesMatchComponents(v);
return v;
}
@@ -155,6 +158,7 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(false));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-b"))));
v.put("leadRollover", null);
assertNamesMatchComponents(v);
return v;
}
@@ -233,7 +233,7 @@ class FleetConfigValidateAllTest {
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels")), names,
"validateModels", "validateLeadRollover")), names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
@@ -99,6 +99,11 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-guard"))));
// fleetd #480: leadRollover is left as-is unconditionally by withDefaults() (see its
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
// here proves it, rather than leaving it null and proving nothing.
v.put("leadRollover", new FleetConfig.LeadRollover(
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
assertNamesMatchComponents(v);
return v;
}
@@ -44,6 +44,7 @@ public final class FakeHerdr implements HerdrClient {
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private final Map<String, String> paneCloseErrorCodeFor = new ConcurrentHashMap<>();
private final Map<String, String> processInfoErrorCodeFor = new ConcurrentHashMap<>();
private String tabCloseErrorCode = null;
private final Map<String, String> tabCloseErrorCodeFor = new ConcurrentHashMap<>();
private String agentSendErrorCode = null;
@@ -144,6 +145,18 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Make {@code pane.process_info} fail with this herdr error code, but only for the given
* {@code pane_id} — every other pane's {@code pane.process_info} still succeeds. Models a
* transient herdr failure partway through a {@link PaneLocator} pid→pane scan (fleetd #505):
* the scan must be able to tell "this pane does not own the pid" apart from "the scan could
* not check this pane at all", instead of collapsing both into one {@code false}.
*/
public FakeHerdr processInfoFailsForPane(String paneId, String code) {
this.processInfoErrorCodeFor.put(paneId, code);
return this;
}
/** Set the {@code agent_status} that {@code agent.get} reports (drives the injector). */
public FakeHerdr agentStatus(String status) {
this.agentStatus = status;
@@ -407,6 +420,13 @@ public final class FakeHerdr implements HerdrClient {
{"pane_id":"w2:p9","terminal_id":"term_shell","workspace_id":"w2","tab_id":"w2:t8"}]}""");
case "pane.process_info" -> {
Object paneId = params instanceof java.util.Map<?, ?> m ? m.get("pane_id") : null;
String failCode = paneId == null ? null
: processInfoErrorCodeFor.get(String.valueOf(paneId));
if (failCode != null) {
throw new HerdrException(
"herdr error [" + failCode + "]: pane.process_info failed",
failCode, null);
}
yield "w2:p7".equals(paneId)
? mapper.readTree(("""
{"type":"pane_process_info","process_info":{"pane_id":"w2:p7","shell_pid":%d,
@@ -37,7 +37,7 @@ class PaneLocatorContractTest {
.path("pane").path("terminal_id").asText(null);
assertNotNull(terminalId, "seed pane should carry a terminal_id");
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid),
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid).terminal(),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
spaces.closeTab(tab.tab().tabId());
@@ -15,18 +15,20 @@ class PaneLocatorTest {
@Test
void resolvesTerminalForAForegroundPid() {
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
void nullForAPidInNoPane() {
assertNull(loc.terminalForPid(999_999));
PaneLocator.Lookup outcome = loc.terminalForPid(999_999);
assertNull(outcome.terminal());
assertTrue(outcome.complete(), "a full, error-free scan that finds no match is complete");
}
@Test
void nullForNonPositivePid() {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
assertNull(loc.terminalForPid(0).terminal());
assertNull(loc.terminalForPid(-1).terminal());
}
// --- two-daemon fallback (CB-185) -----------------------------------------
@@ -38,7 +40,7 @@ class PaneLocatorTest {
HerdrClient lead = new FakeHerdr().withNoPanes();
HerdrClient member = new FakeHerdr();
PaneLocator two = new PaneLocator(lead, member);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
@@ -48,13 +50,13 @@ class PaneLocatorTest {
HerdrClient lead = new FakeHerdr();
HerdrClient member = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(lead, member);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
void nullWhenNeitherClientHasTheMatch() {
PaneLocator two = new PaneLocator(new FakeHerdr().withNoPanes(), new FakeHerdr().withNoPanes());
assertNull(two.terminalForPid(FakeHerdr.WORKER_PID));
assertNull(two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
}
@Test
@@ -63,7 +65,7 @@ class PaneLocatorTest {
// must behave exactly like the one-arg constructor, including making only one herdr call.
FakeHerdr shared = new FakeHerdr();
PaneLocator two = new PaneLocator(shared, shared);
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID));
assertEquals("term_a", two.terminalForPid(FakeHerdr.WORKER_PID).terminal());
long paneListCalls = shared.calls.stream().filter(c -> c.method().equals("pane.list")).count();
assertEquals(1, paneListCalls, "same-object lead/member must scan exactly once, not twice");
}
@@ -75,7 +77,7 @@ class PaneLocatorTest {
// Regression: a pid with no parent chain at all — no ancestry walk is needed to match it.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
PaneLocator loc = new PaneLocator(pane, new FakeParentResolver());
assertEquals("term_x", loc.terminalForPid(5000));
assertEquals("term_x", loc.terminalForPid(5000).terminal());
}
@Test
@@ -83,7 +85,7 @@ class PaneLocatorTest {
// Regression: same as above, but matching via the foreground-processes list.
OnePaneHerdr pane = new OnePaneHerdr("term_x", "pX", 5000, 6000);
PaneLocator loc = new PaneLocator(pane, new FakeParentResolver());
assertEquals("term_x", loc.terminalForPid(6000));
assertEquals("term_x", loc.terminalForPid(6000).terminal());
}
@Test
@@ -97,7 +99,7 @@ class PaneLocatorTest {
.parent(7002, 7001) // grandchild -> child
.parent(7001, 5000); // child -> shell (the pane's shell_pid)
PaneLocator loc = new PaneLocator(pane, parents);
assertEquals("term_x", loc.terminalForPid(7002));
assertEquals("term_x", loc.terminalForPid(7002).terminal());
}
@Test
@@ -110,7 +112,7 @@ class PaneLocatorTest {
.parent(9002, 9001)
.parent(9001, 9000); // chain never reaches 5000 or 6000
PaneLocator loc = new PaneLocator(pane, parents);
assertNull(loc.terminalForPid(9002));
assertNull(loc.terminalForPid(9002).terminal());
}
@Test
@@ -122,7 +124,7 @@ class PaneLocatorTest {
.parent(100, 101)
.parent(101, 100); // cycle, never reaches the pane's pids
PaneLocator loc = new PaneLocator(pane, parents);
assertNull(loc.terminalForPid(100));
assertNull(loc.terminalForPid(100).terminal());
}
@Test
@@ -141,11 +143,69 @@ class PaneLocatorTest {
};
HerdrClient noPanes = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(noPanes, pane, counting);
assertEquals("term_x", two.terminalForPid(7002));
assertEquals("term_x", two.terminalForPid(7002).terminal());
assertEquals(3, calls.get(), "ancestry must be walked once (3 lookups: 7002, 7001, 5000), "
+ "not re-walked per herdr client");
}
// --- fleetd #505: a herdr error during the scan must not read as a clean negative ---------
@Test
void anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative() {
// The discriminating case: pane.process_info fails for exactly the pane that DOES own the
// caller's pid ("w2:p7", term_a). Before the fix, that failure was swallowed into a plain
// "does not own it" and the scan finished with a clean-looking null — indistinguishable
// from a real primary. It must now report incomplete, not a definite null.
FakeHerdr herdr = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
PaneLocator loc = new PaneLocator(herdr);
PaneLocator.Lookup outcome = loc.terminalForPid(FakeHerdr.WORKER_PID);
assertNull(outcome.terminal(), "the failing pane's ownership could not be confirmed");
assertFalse(outcome.complete(),
"a scan that could not check the owning pane must not report as complete");
}
@Test
void aVanishedPaneThatIsNotTheMatchLeavesAnOtherwiseSuccessfulScanComplete() {
// The companion invariant: a DIFFERENT pane (not the caller's own) failing mid-scan must
// not turn every mid-scan teardown into a refusal — the real match is still found, and the
// scan is still reported complete.
FakeHerdr herdr = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
PaneLocator loc = new PaneLocator(herdr);
PaneLocator.Lookup outcome = loc.terminalForPid(FakeHerdr.WORKER_PID);
assertEquals("term_a", outcome.terminal());
assertTrue(outcome.complete(), "a positive match elsewhere in the scan is definitive");
}
// --- fleetd #509: the completeness fold across clients must not collapse to "last wins" ----
@Test
void anEarlierClientsErrorSurvivesALaterClientsCleanNegative() {
// terminalForPid folds each client's Lookup.complete() with
// complete = complete && outcome.complete();
// (PaneLocator.java:117). With a SINGLE client, a fold that keeps only the last outcome
// (dropping the "complete &&" prefix) agrees with the real fold — which is why 14 of the
// 15 pre-existing tests never catch that mutation: none of them vary the number of clients.
// Here the LEAD client errors on exactly the pane that would have owned the pid (so its
// scan is incomplete AND finds no match), and the MEMBER client cleanly reports no panes
// at all (a complete, negative scan). The real fold ANDs the two into false. A fold that
// just keeps the last client's outcome would read this as a clean true — the earlier
// error is erased, and CallerResolver.java:254 would read scanComplete() as true and
// promote an unverified caller to the primary.
HerdrClient lead = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
HerdrClient member = new FakeHerdr().withNoPanes();
PaneLocator two = new PaneLocator(lead, member);
PaneLocator.Lookup outcome = two.terminalForPid(FakeHerdr.WORKER_PID);
assertNull(outcome.terminal(), "the pane that could have owned the pid was never checked");
assertFalse(outcome.complete(),
"an earlier client's error must survive a later client's clean negative");
}
/** Minimal single-pane {@link HerdrClient} fake, purpose-built for the ancestry tests above. */
private static final class OnePaneHerdr implements HerdrClient {
private final ObjectMapper mapper = new ObjectMapper();
@@ -1,9 +1,7 @@
package dev.ltms.fleet.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
@@ -12,8 +10,8 @@ import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TestTurnTokens;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Set;
import java.util.regex.Pattern;
@@ -470,15 +468,7 @@ class CompletionResolverTest {
// CB-564: this used to be a bare DEBUG "failed send to X via turn-stall fallback" — a symptom
// with no cause, and below the level anyone watching for member health would see. A fail that
// resolves a caller's blocked send is at least WARN and must carry the reason.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger resolverLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(CompletionResolver.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
resolverLog.addAppender(appender);
resolverLog.setLevel(Level.WARN);
try {
try (CapturedLog log = CapturedLog.at(CompletionResolver.class, Level.WARN)) {
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
@@ -486,7 +476,7 @@ class CompletionResolverTest {
resolver.fail("term_a", null);
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
@@ -494,8 +484,6 @@ class CompletionResolverTest {
assertTrue(warn.contains("term_a"), "the log names the target: " + warn);
assertTrue(warn.contains("stuck on an error screen"), "the log carries the reason: " + warn);
assertTrue(waiter.isDone());
} finally {
resolverLog.detachAppender(appender);
}
}
@@ -1,16 +1,14 @@
package dev.ltms.fleet.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.msg.TestTurnTokens;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
@@ -19,6 +17,8 @@ import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.*;
@@ -491,23 +491,15 @@ class InjectorTest {
void readinessGraceExpiryIsLogged() {
// CB-562: the grace-expiry path used to clear the queue silently, so a message that never
// reached the worker's pane surfaced elsewhere as an unrelated turn-stall failure. Assert the
// expiry now names the real cause. (ListAppender capture pattern mirrors AuditLogTest.)
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
// expiry now names the real cause. (CapturedLog pattern mirrors AuditLogTest.)
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
@@ -516,8 +508,77 @@ class InjectorTest {
assertTrue(warn.contains("never reached"), "the log names the real cause: " + warn);
assertTrue(warn.contains("1 queued message"),
"the log carries the failed message count: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void readinessGraceExpiryLogsTheMeasuredPollCountNextToTheConfiguredBudget() {
// fleetd #501, defect 1: READINESS_GRACE_POLLS (240) used to be printed twice — once as
// "after {} polls" and once inside the parenthesised budget — even though the loop's own
// counter (Target.notReadySincePoll) was in scope at the same call site. On THIS branch the
// counter has just reached the threshold, so it equals the constant by construction and this
// test cannot tell the two apart — it only pins that the message still carries both a poll
// count and a labelled configured budget, using literal numbers (240, 60), never
// READINESS_GRACE_POLLS or POLL_INTERVAL_MILLIS, so the assertion can't silently track a
// constant change instead of catching a real regression.
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains("after 240 polls"),
"must print the measured poll count as a plain number: " + warn);
assertTrue(warn.contains("configured=240 polls/60s"),
"must print the configured budget, clearly labelled: " + warn);
}
}
@Test
void readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants() {
// fleetd #501, defect 2: the old line computed "({}s)" as READINESS_GRACE_POLLS *
// POLL_INTERVAL_MILLIS / 1000 — arithmetic on two constants, never a measurement, and wrong
// in the direction that says everything ran on schedule. This stub clock returns two FIXED
// values (1_000ms at the first non-ready sample, 318_412ms at the poll that trips the grace)
// whose difference — 317_412ms — does NOT equal 240 * POLL_INTERVAL_MILLIS (=60_000ms).
// Asserting on that literal, non-derived number is what makes this test able to fail if the
// production code goes back to printing the constant-arithmetic value instead of the
// injected clock's measurement.
long[] readings = {1_000L, 318_412L};
AtomicInteger call = new AtomicInteger(0);
LongSupplier stubClock = () -> {
int i = call.getAndIncrement();
if (i >= readings.length) {
throw new AssertionError("nowMillis read more times than this fixture expects (" + i
+ "); the readiness-not-ready branch should read the clock exactly twice — "
+ "once to stamp the first non-ready sample, once at grace expiry");
}
return readings[i];
};
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
}, stubClock);
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains("elapsed=317412ms"), "must print the MEASURED elapsed time from "
+ "the injected clock (318412 - 1000 = 317412), not an arithmetic value: " + warn);
assertFalse(warn.contains("elapsed=60000ms"), "must not print "
+ "READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS (240 * 250 = 60000ms) as if it "
+ "were the measured elapsed time: " + warn);
}
}
@@ -554,21 +615,13 @@ class InjectorTest {
// CB-564: a vanished worker used to drop its queue with no log at all — the only trace was
// whatever failed downstream (e.g. a caller's send timing out with no clue why). Assert the
// drop itself now names the cause and the number of messages it failed.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
try (CapturedLog log = CapturedLog.at(Injector.class, Level.WARN)) {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
});
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
@@ -576,8 +629,6 @@ class InjectorTest {
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
assertTrue(warn.contains("1 message"), "the log carries the failed message count: " + warn);
assertTrue(warn.contains("worker gone"), "the log carries the real cause: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
File diff suppressed because it is too large Load Diff
@@ -3,7 +3,9 @@ package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
@@ -11,13 +13,30 @@ import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** fleetd #464: launch charters must not name MCP tools the server does not register. */
/**
* fleetd #464 shipped {@code configuredChartersNameOnlyRegisteredTools} below, comparing charter
* text against the registered tool surface — but it scraped <em>both</em> sides from source text:
* its own fixture charter, and a regex over {@code FleetMcp.java}'s {@code tool("…")} calls. #469's
* gap: nothing anyone wrote into the live {@code fleetd.yaml} could ever reach that test, because
* it never called production validation code.
*
* <p>This version keeps the charter-text extraction helper ({@code toolsNamedIn}) — charters are
* free-text config, so finding a {@code fleet_*}/{@code bridge_*} token inside one has no source
* but a scrape — but reads the <em>registered</em> side from {@link FleetTool}, the canonical enum
* {@code FleetMcp} itself now derives its tool schemas and authorization switch from, rather than a
* second scrape of {@code FleetMcp.java}'s source. It also exercises {@link
* CharterToolSurface#assertChartersNameOnlyRegisteredTools} directly — the method {@code
* Fleetd.main} actually calls at startup — both accepting and rejecting. {@code
* dev.ltms.fleet.FleetdStartupValidationTest#mainRefusesACharterNamingAnUnregisteredTool} is what
* closes #469's actual gap: it proves that call is wired into {@code Fleetd.main} itself, against a
* live-shaped config fixture, not just a unit call to the method in isolation.
*/
class CharterToolSurfaceTest {
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(String text, String regex) {
Matcher m = Pattern.compile(regex).matcher(text);
Set<String> found = new LinkedHashSet<>();
@@ -33,13 +52,8 @@ class CharterToolSurfaceTest {
"(fleet_[a-z_]+|bridge_[a-z_]+)");
}
/** Every tool {@link FleetMcp} registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(Files.readString(MCP_SOURCE), "tool\\(\\\"(fleet_[a-z_]+)\\\"");
}
@Test
@DisplayName("[SOURCE TEXT] every tool named in a configured charter is registered by the server")
@DisplayName("every tool named in a configured charter is in the canonical FleetTool set")
void configuredChartersNameOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path configFile = dir.resolve("charters.yaml");
Files.writeString(configFile, """
@@ -53,20 +67,58 @@ class CharterToolSurfaceTest {
FleetConfig config = FleetConfig.load(configFile);
Set<String> named = toolsNamedIn(config);
Set<String> registered = toolsTheServerRegisters();
Set<String> registered = FleetTool.wireNames();
assertTrue(!named.isEmpty(),
"the charter fixture named no fleet_* or bridge_* tool. This test would check nothing; "
+ "add charter text that names a tool before changing the extraction.");
assertTrue(!registered.isEmpty(),
"the FleetMcp registration scrape found no tools. This test would check nothing; "
+ "repair the tool(\"…\") extraction before changing the assertion.");
"FleetTool.wireNames() is empty. This test would check nothing; repair FleetTool "
+ "before changing the assertion.");
Set<String> unknown = new LinkedHashSet<>(named);
unknown.removeAll(registered);
assertTrue(unknown.isEmpty(),
"configured charter text names " + unknown + ", but FleetMcp does not register it. "
"configured charter text names " + unknown + ", but FleetTool does not list it. "
+ "Checked " + named + " against " + registered + ". Fix the charter text or "
+ "register the tool; do NOT weaken this test.");
+ "add the tool to FleetTool; do NOT weaken this test.");
}
/**
* The positive case for the actual production entry point: a charter naming only tools
* {@link FleetTool} lists must not throw.
*/
@Test
@DisplayName("CharterToolSurface accepts a charter that names only registered tools")
void charterToolSurfaceAcceptsKnownTools() {
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of(
"dev", "Send the final handoff through fleet_reply, using fleet_send to delegate.",
"reviewer", "Use fleet_ask only for the lead's decision.")));
}
/**
* The negative case for the actual production entry point (fleetd #469's motivating example:
* {@code bridge_send} is the pre-CB-634 name, removed from the tool surface). The message must
* name both the offending charter key and the unknown tool, so an operator reading the startup
* log knows exactly which charter to fix.
*/
@Test
@DisplayName("CharterToolSurface rejects a charter naming a tool the server does not register")
void charterToolSurfaceRejectsAnUnregisteredTool() {
Map<String, String> charters = new LinkedHashMap<>();
charters.put("dev", "Send the final handoff through bridge_send.");
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(charters));
assertTrue(e.getMessage().contains("dev"),
"expected the charter key 'dev' in the failure message, got: " + e.getMessage());
assertTrue(e.getMessage().contains("bridge_send"),
"expected the unknown tool 'bridge_send' in the failure message, got: " + e.getMessage());
}
@Test
@DisplayName("CharterToolSurface is a no-op on an absent or empty charter map")
void charterToolSurfaceIsANoOpWithNoCharters() {
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(null));
assertDoesNotThrow(() -> CharterToolSurface.assertChartersNameOnlyRegisteredTools(Map.of()));
}
}
@@ -57,6 +57,23 @@ class ConnectionIdentityTest {
// must read as "resolved" — the distinction #317 turns on.
ConnectionIdentity.Caller c = with(_ -> 999_999).resolve("127.0.0.1", 55555);
assertTrue(c.resolved());
assertTrue(c.scanComplete(), "no herdr error happened, so the scan is complete");
}
@Test
void scanIsIncompleteWhenHerdrErrorsOnThePaneThatOwnsThePid() {
// fleetd #505: a transient herdr error on exactly the pane that DOES own the caller's pid
// must be visible as an incomplete scan, distinct from a real primary (resolved(), null
// terminal, complete scan). Both have pid > 0 and a null terminal — scanComplete is the
// only thing that tells them apart.
FakeHerdr failing = new FakeHerdr().processInfoFailsForPane("w2:p7", "transient");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(failing), _ -> FakeHerdr.WORKER_PID);
ConnectionIdentity.Caller c = id.resolve("127.0.0.1", 55555);
assertTrue(c.resolved(), "the pid itself resolved fine — this is not #317's failure");
assertNull(c.terminal(), "the owning pane could not be confirmed");
assertFalse(c.scanComplete(), "the scan could not check the pane that owns this pid");
}
@Test
@@ -26,7 +26,7 @@ import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
@@ -50,7 +50,6 @@ import static org.junit.jupiter.api.Assertions.*;
class FleetMcpAuthzTest {
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static final Pattern TOOL_REGISTRATION = Pattern.compile("tool\\(\\\"(fleet_[a-z_]+)\\\"");
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
@@ -76,12 +75,21 @@ class FleetMcpAuthzTest {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
metrics = FleetMetrics.create(sessions, new InMemoryReplyInbox());
// fleetd #480 correction round: FleetMcp has one constructor now (no defaulting
// overloads — see its javadoc), so every feature this test does not exercise is passed
// its explicit "off" value here rather than being omitted.
//
// fleetd #518: callers is now required (never null) either way — the resolver that used
// to be omitted to reach "legacy" is now always real, and AuthorizationMode is the
// separate, explicit choice that governs enforcement.
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)) : null,
CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)),
enforce ? FleetMcp.AuthorizationMode.ENFORCED : FleetMcp.AuthorizationMode.UNENFORCED,
metrics, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none());
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), List.of(), null);
return mcp;
}
@@ -184,7 +192,31 @@ class FleetMcpAuthzTest {
// The 22 pre-existing FleetMcpTest cases rely on no authorization being enforced.
FleetMcp m = mcp(false);
assertNull(m.denyFor(ANON, Authz.Action.SPAWN, null),
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
"AuthorizationMode.UNENFORCED chosen explicitly ⇒ authorization not enforced "
+ "(legacy behaviour) — fleetd #518 replaced the old callers == null idiom");
}
/**
* fleetd #509 was originally proven against {@code FleetMcp.legacyPrincipal} — a second,
* separately-maintained principal-resolution heuristic that only ran when {@code callers} was
* omitted (null). fleetd #518 deleted that whole heuristic: {@code callers} is now required
* and non-null under every {@link FleetMcp.AuthorizationMode}, so the ONE real
* {@link CallerResolver} resolves every caller, enforced or not, and #509's property (a
* non-loopback / unresolved caller must never earn the primary's authority) is exactly what
* {@code CallerResolverTest.aNonLoopbackCallerIsNeverThePrimaryUnderLoopbackTrust} already
* proves on that one real path. There is no longer a second heuristic here to test.
*/
@Test
void anUnresolvedNonLoopbackCallerIsAnonymousUnderTheOneRealResolver() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null));
// A non-loopback address never even reaches the pane scan — resolve() short-circuits it
// to Caller(null, -1, true), the same "no terminal" shape a genuine primary's connection
// produces on loopback. The real resolver must not conflate the two.
Principal p = resolver.resolve("8.8.8.8", 1234, null);
assertEquals(Principal.anonymous(), p,
"an unresolved, non-loopback caller must earn no authority, not the primary's");
}
// --- fleetd #439: who may see fleet_list's coordinator row ----------------------------------
@@ -296,10 +328,18 @@ class FleetMcpAuthzTest {
@Test
void everyRegisteredToolHasItsHandlerActionPinned() {
Set<String> registered = toolsTheServerRegisters();
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
// a third copy of the same list this file, CharterToolSurfaceTest and McpContractDocTest
// each kept independently. All three now read FleetTool.wireNames(), the canonical set
// FleetMcp itself derives its tool schemas AND its authorization switch from; adding a tool
// there without pinning its action in FleetMcp#authzAction is a compile error, so this test's
// per-tool assertions below are a run-time regression pin on top of that compile-time check,
// not the only thing standing between a new tool and an unpinned action.
Set<String> registered = FleetTool.wireNames();
assertTrue(registered.size() >= 10,
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching");
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
+ "); the server registers eleven, so FleetTool has stopped listing the real "
+ "tool surface");
registered.forEach(tool -> assertDoesNotThrow(() -> FleetMcp.toolAction(tool, Map.of()),
() -> tool + " is registered but has no pinned authorization action"));
@@ -319,19 +359,6 @@ class FleetMcpAuthzTest {
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
}
private static Set<String> toolsTheServerRegisters() {
try {
Matcher matcher = TOOL_REGISTRATION.matcher(Files.readString(MCP_SOURCE));
Set<String> tools = new LinkedHashSet<>();
while (matcher.find()) {
tools.add(matcher.group(1));
}
return tools;
} catch (Exception e) {
throw new AssertionError("could not scrape FleetMcp tool registrations", e);
}
}
@Test
void aWorkerMayNotDrainAnotherSessionsInboxByPolling() {
FleetMcp m = mcp(true);
@@ -0,0 +1,147 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.session.FakeWorktrees;
import dev.ltms.fleet.session.SessionManager;
import io.modelcontextprotocol.client.McpClient;
import io.modelcontextprotocol.client.McpSyncClient;
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
import io.modelcontextprotocol.spec.McpClientTransport;
import io.modelcontextprotocol.spec.McpSchema;
import org.eclipse.jetty.server.Server;
import org.eclipse.jetty.server.ServerConnector;
import org.eclipse.jetty.servlet.ServletContextHandler;
import org.eclipse.jetty.servlet.ServletHolder;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.net.http.HttpRequest;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #518 — Part 2: drive the {@code contextExtractor} closure for real.
*
* <p>{@code FleetMcp.deny()}/{@code denyFor()} has a full policy table of tests
* ({@code FleetMcpAuthzTest}), and {@code CallerResolver.resolve()} has its own full suite
* ({@code CallerResolverTest}). Neither one ever exercises the closure that WIRES them together
* inside {@code FleetMcp}'s constructor: it is built once, handed to the MCP SDK's transport, and
* only ever runs when a real MCP client makes a real HTTP request. Every existing test either
* calls {@code denyFor(Principal, ...)} with a hand-built {@link dev.ltms.fleet.auth.Principal}
* (never asking who the transport would actually have resolved) or drives a static handler method
* directly. A mutation that swapped the whole resolution decision for an unconditional fallback —
* bypassing {@link CallerResolver} entirely — passed the full suite, including every
* {@code FleetMcpAuthzTest} case, because none of them go through the transport at all.
*
* <p>This test boots the real {@code HttpServletStreamableServerTransportProvider} on a real
* Jetty server, drives it with a real MCP client over HTTP, and checks a result that only the
* real {@link CallerResolver} can produce: token-mode inspects the {@code Authorization} header
* and grants {@code PRIMARY} only for the right bearer token. The connection never resolves to a
* worker pane (the fake peer-pid lookup always misses), so the ONLY way {@code fleet_whoami} can
* come back as {@code primary} is if the closure actually called {@code callers.resolve(...)} and
* read that header — a behaviour the deleted {@code legacyPrincipal} heuristic never had at all.
*/
class FleetMcpContextExtractorTest {
private static final String TOKEN = "s3cret-mcp-token";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private FleetMcp mcp;
private Server server;
@AfterEach
void tearDown() throws Exception {
if (server != null) {
server.stop();
}
if (mcp != null) {
mcp.close();
}
}
@Test
void aRealMcpRequestIsResolvedByTheRealCallerResolverNotAFallback() throws Exception {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
// The peer-pid lookup always misses (-1), so no connection here is ever resolved to a
// worker pane — every call falls through to CallerResolver's token check, the one branch
// that is unreachable through the deleted legacy heuristic.
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> -1);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, true, TOKEN,
Map::of, new MemberRegistry(null));
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null), callers, FleetMcp.AuthorizationMode.ENFORCED,
null, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), List.of(), null);
ServletContextHandler handler = new ServletContextHandler();
handler.setContextPath("/");
handler.addServlet(new ServletHolder(mcp.servlet()), "/mcp");
server = new Server(0);
server.setHandler(handler);
server.start();
String baseUrl = "http://127.0.0.1:"
+ ((ServerConnector) server.getConnectors()[0]).getLocalPort();
// The right bearer token: the real CallerResolver grants PRIMARY, which fleet_whoami's
// READ gate lets through.
McpSchema.CallToolResult authorized = callWhoami(baseUrl, "Bearer " + TOKEN);
assertFalse(authorized.isError(), "a valid bearer token must resolve as PRIMARY and pass "
+ "fleet_whoami's READ gate: " + textOf(authorized));
assertTrue(textOf(authorized).contains("\"role\":\"primary\""),
"fleet_whoami must report the role the real CallerResolver resolved over this "
+ "connection, not a fallback: " + textOf(authorized));
// No credential at all, over the SAME wiring: the real resolver refuses it as ANONYMOUS.
// legacyPrincipal never looked at the Authorization header, so it could not have told
// these two calls apart at all -- this is the assertion the deleted mutation would fail.
McpSchema.CallToolResult unauthorized = callWhoami(baseUrl, null);
assertTrue(unauthorized.isError(), "no credential must be refused, not silently let "
+ "through: " + textOf(unauthorized));
}
private static McpSchema.CallToolResult callWhoami(String baseUrl, String authorizationHeader) {
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder();
if (authorizationHeader != null) {
requestTemplate.header("Authorization", authorizationHeader);
}
McpClientTransport transport = HttpClientStreamableHttpTransport.builder(baseUrl)
.endpoint("/mcp")
.requestBuilder(requestTemplate)
.build();
try (McpSyncClient client = McpClient.sync(transport).build()) {
client.initialize();
return client.callTool(McpSchema.CallToolRequest.builder("fleet_whoami").arguments(Map.of()).build());
}
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
}
@@ -0,0 +1,262 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.session.FakeWorktrees;
import dev.ltms.fleet.session.SessionManager;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
*
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a
* real virtual-thread continuation runner) rather than its package-private test constructor —
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not
* need to control the post-{@code confirm()} continuation's timing: it only asserts the
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a
* daemon virtual thread.
*/
class FleetMcpHandoverTest {
@TempDir
Path tmp;
private static final String LEAD = "term_lead";
private static final String OTHER_LEAD = "term_other_lead";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private FleetMcp mcp;
@AfterEach
void close() {
if (mcp != null) mcp.close();
}
private static FleetConfig.LeadRollover cfg(String handoverPath) {
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
}
private LeadRollover newRollover(String handoverPath) {
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null);
}
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
private FleetMcp mcp(LeadRollover leadRollover) {
FleetConfig.Profile pcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(pcfg.profile(), pcfg), pcfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
CallerResolver.withLeadsAndMembers(identity, false, null, Map::of, new MemberRegistry(null)),
FleetMcp.AuthorizationMode.ENFORCED,
null, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), List.of(), leadRollover);
return mcp;
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private static String extractToken(String json) {
int i = json.indexOf("\"token\":\"");
assertTrue(i >= 0, "no token field in: " + json);
int start = i + "\"token\":\"".length();
int end = json.indexOf('"', start);
return json.substring(start, end);
}
// --- acceptance 1: registered whether or not leadRollover: is configured -------------------
@Test
@DisplayName("fleet_handover is registered whether or not leadRollover: is configured")
void registeredEitherWay() {
assertTrue(FleetTool.wireNames().contains("fleet_handover"),
"FleetTool must list fleet_handover as part of the canonical tool surface");
FleetMcp withNull = mcp(null);
assertTrue(withNull.registeredTools().stream().anyMatch(t -> "fleet_handover".equals(t.name())),
"fleet_handover must be registered even with no LeadRollover constructed");
withNull.close();
FleetMcp withConfigured = mcp(newRollover(tmp.resolve("h.md").toString()));
assertTrue(withConfigured.registeredTools().stream().anyMatch(t -> "fleet_handover".equals(t.name())),
"fleet_handover must be registered when a LeadRollover IS constructed too");
}
// --- acceptance 2: null LeadRollover -> clean NOT_CONFIGURED, no throw, every action -------
@Test
@DisplayName("with leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal and never throws")
void nullLeadRolloverRefusesCleanlyForEveryAction() {
McpSchema.CallToolResult open = assertDoesNotThrow(
() -> FleetMcp.handover(null, LEAD, Map.of("action", "open")));
assertFalse(open.isError(), "a refusal is not a protocol error: " + textOf(open));
assertTrue(textOf(open).contains("NOT_CONFIGURED"), textOf(open));
McpSchema.CallToolResult confirm = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
Map.of("action", "confirm", "token", "whatever")));
assertFalse(confirm.isError());
assertTrue(textOf(confirm).contains("NOT_CONFIGURED"), textOf(confirm));
McpSchema.CallToolResult cancel = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD,
Map.of("action", "cancel", "token", "whatever")));
assertFalse(cancel.isError());
assertTrue(textOf(cancel).contains("NOT_CONFIGURED"), textOf(cancel));
}
@Test
@DisplayName("a blank/unknown action is a clean tool error, never an exception")
void unknownActionIsACleanError() {
McpSchema.CallToolResult missing = assertDoesNotThrow(() -> FleetMcp.handover(null, LEAD, Map.of()));
assertTrue(missing.isError());
McpSchema.CallToolResult bogus = assertDoesNotThrow(
() -> FleetMcp.handover(null, LEAD, Map.of("action", "bogus")));
assertTrue(bogus.isError());
}
// --- acceptance 3: non-primary refused by the authorization gate ---------------------------
@Test
@DisplayName("fleet_handover maps to Authz.Action.HANDOVER, primary-only")
void onlyThePrimaryIsPermitted() {
assertEquals(Authz.Action.HANDOVER, FleetMcp.toolAction("fleet_handover", Map.of()));
FleetMcp m = mcp(null);
assertNull(m.denyFor(Principal.primary(1), Authz.Action.HANDOVER, null),
"the primary may drive fleet_handover");
McpSchema.CallToolResult deniedWorker =
m.denyFor(Principal.worker("term_w", 2), Authz.Action.HANDOVER, null);
assertNotNull(deniedWorker, "a worker must be refused fleet_handover");
assertTrue(deniedWorker.isError());
McpSchema.CallToolResult deniedArchitect =
m.denyFor(Principal.architect("slot", "term_a", 3), Authz.Action.HANDOVER, null);
assertNotNull(deniedArchitect, "an architect has no lifecycle rights either — same gate as SPAWN/STOP/DRAIN");
assertTrue(deniedArchitect.isError());
}
// --- acceptance 4: the registered schema names no caller-terminal parameter ----------------
@Test
@DisplayName("the registered fleet_handover schema names no terminal/session/leadTerminal parameter")
void schemaCarriesNoCallerIdentityParameter() {
FleetMcp m = mcp(null);
McpSchema.Tool tool = m.registeredTools().stream()
.filter(t -> "fleet_handover".equals(t.name()))
.findFirst()
.orElseThrow(() -> new AssertionError("fleet_handover was not registered"));
@SuppressWarnings("unchecked")
Map<String, Object> properties = (Map<String, Object>) tool.inputSchema().get("properties");
assertNotNull(properties, "tool has no 'properties' in its input schema");
for (String forbidden : List.of("terminal", "sessionId", "leadTerminal", "callerTerminal", "target")) {
assertFalse(properties.containsKey(forbidden),
"fleet_handover's REGISTERED schema must not carry a caller-identity parameter, "
+ "found '" + forbidden + "' in " + properties.keySet());
}
}
// --- acceptance 5: open then confirm on the same connection; ownership is enforced --------
@Test
@DisplayName("open then confirm on the same terminal succeeds; a different terminal gets NOT_YOUR_ROLLOVER")
void openThenConfirmRoundTripsAndOwnershipIsEnforced() throws Exception {
Path handover = tmp.resolve("handover.md");
Files.writeString(handover, "not written yet");
LeadRollover rollover = newRollover(handover.toString());
McpSchema.CallToolResult openResult = FleetMcp.handover(rollover, LEAD, Map.of("action", "open"));
assertFalse(openResult.isError(), textOf(openResult));
String token = extractToken(textOf(openResult));
// Rewrite the handover file so its mtime is measurably after open()'s requestedAtMillis —
// LeadRollover#confirm's freshness check (HANDOVER_STALE) requires this.
Thread.sleep(50);
Files.writeString(handover, "the real handover content");
McpSchema.CallToolResult wrongCaller = FleetMcp.handover(rollover, OTHER_LEAD,
Map.of("action", "confirm", "token", token));
assertFalse(wrongCaller.isError(), "a refusal is a legitimate outcome, not a protocol error");
assertTrue(textOf(wrongCaller).contains("NOT_YOUR_ROLLOVER"),
"a different lead terminal confirming must surface NOT_YOUR_ROLLOVER: " + textOf(wrongCaller));
McpSchema.CallToolResult confirmed = FleetMcp.handover(rollover, LEAD,
Map.of("action", "confirm", "token", token));
assertFalse(confirmed.isError(), textOf(confirmed));
assertTrue(textOf(confirmed).contains("\"accepted\":true"),
"the SAME terminal that opened the request must be able to confirm it: " + textOf(confirmed));
}
// --- acceptance 6: cancel on an unknown token is clean, not a failure ----------------------
@Test
@DisplayName("cancel on an unknown token reports no pending request, rather than failing")
void cancelUnknownTokenIsCleanNotAFailure() {
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
Map.of("action", "cancel", "token", "does-not-exist"));
assertFalse(r.isError());
assertTrue(textOf(r).contains("\"cancelled\":false"), textOf(r));
assertTrue(textOf(r).contains("no pending request"), textOf(r));
}
@Test
@DisplayName("cancel on a token actually opened reports cancelled:true")
void cancelKnownTokenSucceeds() {
LeadRollover rollover = newRollover(tmp.resolve("h.md").toString());
String token = extractToken(textOf(FleetMcp.handover(rollover, LEAD, Map.of("action", "open"))));
McpSchema.CallToolResult r = FleetMcp.handover(rollover, LEAD,
Map.of("action", "cancel", "token", token));
assertFalse(r.isError());
assertTrue(textOf(r).contains("\"cancelled\":true"), textOf(r));
}
// --- acceptance 7 (wiring) is covered by FleetdLeadRolloverWiringTest, unchanged -----------
}
@@ -574,6 +574,38 @@ class FleetMcpTest {
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
}
/**
* fleetd #466 scope item 2: {@code fleet_profiles} must carry the repeat count beside the
* remaining seconds, and the two must come off the one {@link BackendQuarantine#status} call so
* they can never disagree about which streak this is (see {@code profilesView}'s javadoc).
*/
@Test
void profilesReportsQuarantineAttemptBesideRemainingSeconds() {
FakeHerdr h = new FakeHerdr();
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai"); // attempt 1: 1800s
now.set(TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai"); // attempt 2: 3600s
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
McpSchema.CallToolResult res = FleetMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
String out = textOf(res);
assertTrue(out.contains("\"quarantinedForSeconds\":3600"), out);
assertTrue(out.contains("\"quarantineAttempt\":2"), out);
}
/** A never-quarantined profile must not carry {@code quarantineAttempt} either. */
@Test
void profilesOmitsQuarantineAttemptWhenNotQuarantined() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = FleetMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), FleetMcp.QuarantineSource.none());
String out = textOf(res);
assertFalse(out.contains("quarantineAttempt"), out);
}
/** fleetd #201 Unit 5: {@code coolingOff} is a SEPARATE map from {@code quarantined}. */
@Test
void profilesReportsACoolingOffCredentialInASeparateMap() {
@@ -1121,6 +1153,35 @@ class FleetMcpTest {
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
}
/**
* fleetd #466 scope item 2: {@code fleet_list}'s capacity rows (CB-583: they reuse quarantine)
* must carry {@code quarantineAttempt} beside {@code quarantinedForSeconds}, off the same
* {@link BackendQuarantine#status} call as {@code fleet_profiles} -- see {@code capacityView}'s
* comment pointing back to {@code profilesView}.
*/
@Test
void capacityRowReportsQuarantineAttemptBesideRemainingSeconds() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
java.util.concurrent.atomic.AtomicLong now = new java.util.concurrent.atomic.AtomicLong(0L);
BackendQuarantine quarantine = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai"); // attempt 1
now.set(TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai"); // attempt 2
now.set(TimeUnit.MINUTES.toNanos(60));
quarantine.quarantine("shared-openai"); // attempt 3
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
assertTrue(out.contains("\"quarantineAttempt\":3"), out);
}
@Test
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
FakeHerdr h = new FakeHerdr();
@@ -27,16 +27,21 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
* safe to write.
*
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
* doc's own header carries that caveat for its readers.
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown, and it only catches a name
* in the doc that the server does not register. It cannot catch a flow that describes the wrong
* order, or a parameter name in prose — those are not name-shaped. The doc's own header carries
* that caveat for its readers.
*
* <p>fleetd #469: the registered side used to be its own scrape of {@code FleetMcp.java}'s {@code
* tool("…")} calls — a third copy of the same list {@code CharterToolSurfaceTest} and {@code
* FleetMcpAuthzTest} each kept their own copy of too. All three now read {@link
* FleetTool#wireNames()}, the one canonical set {@code FleetMcp} itself derives its tool schemas and
* authorization switch from.
*/
class McpContractDocTest {
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(Path file, String regex) throws Exception {
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
@@ -52,15 +57,10 @@ class McpContractDocTest {
return matches(DOC, "(fleet_[a-z_]+)");
}
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
}
@Test
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
void theDocNamesNoToolThatDoesNotExist() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> registered = FleetTool.wireNames();
Set<String> named = toolsNamedInTheDoc();
Set<String> unknown = new LinkedHashSet<>(named);
@@ -84,13 +84,13 @@ class McpContractDocTest {
@Test
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
void theCheckActuallyHasSomethingToCheck() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> registered = FleetTool.wireNames();
Set<String> named = toolsNamedInTheDoc();
assertTrue(registered.size() >= 10,
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
+ "and the check above is now vacuous");
"FleetTool.wireNames() returned only " + registered.size() + " tool(s) (" + registered
+ "); the server registers eleven, so FleetTool has stopped listing the real "
+ "tool surface and the check above is now vacuous");
assertTrue(named.size() >= 4,
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
+ "The flows describe delegation, clarification, detached delivery and the "
@@ -1,17 +1,17 @@
package dev.ltms.fleet.msg;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.impl.DefaultExceptionHandler;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.lang.reflect.Proxy;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.assertEquals;
@@ -32,21 +32,17 @@ class AmqpConnectionFailureLoggerTest {
assertEquals(AmqpConnectionFailureLogger.REPLY_INBOX, inboxHandler.connectionName());
assertEquals(AmqpConnectionFailureLogger.LEAD_MAILBOX, mailboxHandler.connectionName());
ListAppender<ILoggingEvent> inboxEvents = attach(AmqpReplyInbox.class);
ListAppender<ILoggingEvent> mailboxEvents = attach(LeadMailbox.class);
IllegalStateException inboxFailure = new IllegalStateException("inbox failure");
IllegalStateException mailboxFailure = new IllegalStateException("mailbox failure");
try {
try (CapturedLog inboxLog = attach(AmqpReplyInbox.class);
CapturedLog mailboxLog = attach(LeadMailbox.class)) {
inboxHandler.handleUnexpectedConnectionDriverException(null, inboxFailure);
mailboxHandler.handleConnectionRecoveryException(null, mailboxFailure);
assertError(inboxEvents, "AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred",
assertError(inboxLog.events(), "AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred",
inboxFailure, "inbox failure line");
assertError(mailboxEvents, "AMQP connection fleetd-lead-mailbox: Caught an exception during connection recovery!",
assertError(mailboxLog.events(), "AMQP connection fleetd-lead-mailbox: Caught an exception during connection recovery!",
mailboxFailure, "mailbox recovery line");
} finally {
detach(AmqpReplyInbox.class, inboxEvents);
detach(LeadMailbox.class, mailboxEvents);
}
}
@@ -54,17 +50,14 @@ class AmqpConnectionFailureLoggerTest {
void connectionResetKeepsForgivingHandlerWarningSemantics() {
AmqpConnectionFailureLogger handler = new AmqpConnectionFailureLogger(
AmqpConnectionFailureLogger.REPLY_INBOX, LoggerFactory.getLogger(AmqpReplyInbox.class));
ListAppender<ILoggingEvent> events = attach(AmqpReplyInbox.class);
try {
try (CapturedLog log = attach(AmqpReplyInbox.class)) {
handler.handleUnexpectedConnectionDriverException(null, new IOException("Connection reset"));
assertEquals(1, events.list.size(), "the handler must still log a reset");
ILoggingEvent event = events.list.getFirst();
assertEquals(1, log.events().size(), "the handler must still log a reset");
ILoggingEvent event = log.events().getFirst();
assertEquals(Level.WARN, event.getLevel(), "ForgivingExceptionHandler logs connection resets at WARN");
assertEquals("AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred "
+ "(Exception message: Connection reset)", event.getFormattedMessage());
assertTrue(event.getThrowableProxy() == null, "ForgivingExceptionHandler does not attach a reset stack trace");
} finally {
detach(AmqpReplyInbox.class, events);
}
}
@@ -103,17 +96,8 @@ class AmqpConnectionFailureLoggerTest {
"all exception-handling methods must remain inherited from DefaultExceptionHandler");
}
private static ListAppender<ILoggingEvent> attach(Class<?> owner) {
Logger logger = (Logger) LoggerFactory.getLogger(owner);
logger.setLevel(Level.DEBUG);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(Class<?> owner, ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(owner)).detachAppender(appender);
private static CapturedLog attach(Class<?> owner) {
return CapturedLog.at(owner, Level.DEBUG);
}
private static AmqpConnectionFailureLogger installedStrictHandler(ConnectionFactory factory, String connection) {
@@ -123,9 +107,9 @@ class AmqpConnectionFailureLoggerTest {
return assertInstanceOf(AmqpConnectionFailureLogger.class, factory.getExceptionHandler());
}
private static void assertError(ListAppender<ILoggingEvent> events, String message, Throwable cause, String name) {
assertEquals(1, events.list.size(), name);
ILoggingEvent event = events.list.getFirst();
private static void assertError(List<ILoggingEvent> events, String message, Throwable cause, String name) {
assertEquals(1, events.size(), name);
ILoggingEvent event = events.getFirst();
assertEquals(Level.ERROR, event.getLevel(), name);
assertEquals(message, event.getFormattedMessage(), name);
assertEquals(cause.toString(), event.getThrowableProxy().getClassName() + ": "
@@ -116,4 +116,262 @@ class BackendQuarantineTest {
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
}
// --- fleetd #466: escalating cooldown -----------------------------------------------------
//
// Base cooldown 600s (10 min), multiplier 2.0, ceiling 2400s (4x base) — small round numbers
// chosen so every deadline is an exact assertion, not just "greater than before". Each call
// below lands at or before the previous deadline (a zero or negative gap), which is always
// "no more than one base cooldown after the previous deadline" — i.e. every call continues the
// same streak, matching a credential that keeps reporting exhausted with no lull.
@Test
void anInvalidBackoffMultiplierIsRejected() {
assertThrows(IllegalArgumentException.class,
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 0.5, TimeUnit.HOURS.toNanos(6)));
}
@Test
void aCeilingBelowTheBaseCooldownIsRejected() {
assertThrows(IllegalArgumentException.class,
() -> new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0, TimeUnit.MINUTES.toNanos(10)));
}
@Test
void repeatedExhaustionEscalatesTheCooldownByExactAmounts() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: base cooldown
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"));
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine (deadline 600s)
q.quarantine("shared-openai"); // 2nd: 600 * 2^1 = 1200
assertEquals(OptionalLong.of(1200L), q.remainingSeconds("shared-openai"),
"a second consecutive exhaustion must double the cooldown, not just increase it");
now.set(TimeUnit.SECONDS.toNanos(1300)); // exactly the 2nd deadline (100 + 1200)
q.quarantine("shared-openai"); // 3rd: 600 * 2^2 = 2400 (exactly at the ceiling)
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
}
@Test
void escalationStopsAtTheCeiling() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 600 * 4 = 2400, at the ceiling, deadline 4200
now.set(TimeUnit.SECONDS.toNanos(4200));
q.quarantine("shared-openai"); // 4th: 600 * 8 = 4800 uncapped, must stay capped at 2400
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
"the cooldown must never exceed the configured ceiling, however long the streak gets");
now.set(TimeUnit.SECONDS.toNanos(6600)); // 4th deadline
q.quarantine("shared-openai"); // 5th: still capped
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"),
"pushing well past the ceiling must not budge it");
}
@Test
void aQuietGapLongerThanTheBaseCooldownResetsToTheBaseCooldown() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600, deadline 600
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200, deadline 1800
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 2400, deadline 4200
assertEquals(OptionalLong.of(2400L), q.remainingSeconds("shared-openai"));
// Quiet for well over one base cooldown (600s) past the 3rd deadline (4200s).
now.set(TimeUnit.SECONDS.toNanos(20_000));
q.quarantine("shared-openai"); // treated as a fresh occurrence
assertEquals(OptionalLong.of(600L), q.remainingSeconds("shared-openai"),
"a long quiet gap must reset the streak back to the base cooldown");
}
@Test
void escalatingOneCredentialDoesNotSlowAnother() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 2400 — three-in-a-row streak on this credential only
q.quarantine("another-credential"); // its first and only exhaustion
assertEquals(OptionalLong.of(600L), q.remainingSeconds("another-credential"),
"an unrelated credential's cooldown must stay at the base rate, unaffected by a sibling's streak");
}
@Test
void withEscalationDefaultsToDoublingCappedAtTwelveTimesTheBase() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = BackendQuarantine.withEscalation(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"),
"the first occurrence must still use the base cooldown");
now.set(TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertEquals(OptionalLong.of(3600L), q.remainingSeconds("shared-openai"),
"the default multiplier must be 2.0");
}
// --- fleetd #466 scope item 2: status() reports repeatCount beside remainingSeconds ------------
@Test
void statusReportsAttemptOneForAFirstOccurrenceNeverAbsentOrZero() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
TimeUnit.HOURS.toNanos(6));
q.quarantine("shared-openai");
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
assertEquals(1800L, status.remainingSeconds());
assertEquals(1, status.repeatCount(),
"a first-ever occurrence must report attempt 1, not 0 or absent -- 1 means unambiguously "
+ "'the first time', where 0 would be indistinguishable from a bug that forgot to count");
}
@Test
void statusIsAbsentWhenTheCredentialIsNotQuarantined() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30), 2.0,
TimeUnit.HOURS.toNanos(6));
assertTrue(q.status("shared-openai").isEmpty());
}
/**
* The acceptance criterion's strong form: the reported {@code repeatCount} must match the exact
* step the cooldown's own growth implies, read off {@link BackendQuarantine#status}'s single
* call -- not two independent reads that happen to agree in this easy case.
*/
@Test
void statusReportsTheGrowingAttemptCountAlongsideTheEscalatingCooldown() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai"); // 1st: 600s, attempt 1
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
assertEquals(600L, first.remainingSeconds());
assertEquals(1, first.repeatCount());
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai"); // 2nd: 1200s, attempt 2
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
assertEquals(1200L, second.remainingSeconds());
assertEquals(2, second.repeatCount());
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3rd: 2400s (at the ceiling), attempt 3
BackendQuarantine.Status third = q.status("shared-openai").orElseThrow();
assertEquals(2400L, third.remainingSeconds());
assertEquals(3, third.repeatCount());
}
/**
* Once the cooldown hits its ceiling, every further consecutive exhaustion reports the SAME
* {@code remainingSeconds} -- so a {@code repeatCount} re-derived from the cooldown value (e.g.
* inverting {@code cooldownNanos * multiplier^(n-1)}) could not tell attempt 4 from attempt 9;
* only the stored counter can. This is the scenario that makes "read the count off a second,
* independent computation" provably wrong rather than just risky.
*/
@Test
void repeatCountKeepsGrowingPastTheCeilingEvenThoughTheCooldownStaysFlat() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
long t = 0L;
for (int attempt = 1; attempt <= 5; attempt++) {
now.set(t);
q.quarantine("shared-openai");
BackendQuarantine.Status status = q.status("shared-openai").orElseThrow();
assertEquals(attempt, status.repeatCount(),
"attempt " + attempt + " must be reported as exactly " + attempt
+ ", not collapsed to whatever attempt first reached the ceiling");
if (attempt >= 3) {
assertEquals(2400L, status.remainingSeconds(), "attempt " + attempt + " must be capped");
}
t += status.remainingSeconds(); // land exactly on the next deadline: still the same streak
}
}
@Test
void aQuietGapResetsTheReportedAttemptCountToOneToo() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // attempt 3
assertEquals(3, q.status("shared-openai").orElseThrow().repeatCount());
now.set(TimeUnit.SECONDS.toNanos(20_000)); // long quiet gap
q.quarantine("shared-openai");
assertEquals(1, q.status("shared-openai").orElseThrow().repeatCount(),
"a reset streak must report attempt 1 again, matching the reset base cooldown");
}
@Test
void escalatingOneCredentialsAttemptCountDoesNotAffectAnother() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600), 2.0,
TimeUnit.SECONDS.toNanos(2400));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai");
now.set(TimeUnit.SECONDS.toNanos(1800));
q.quarantine("shared-openai"); // 3-in-a-row streak on this credential only
q.quarantine("another-credential");
assertEquals(1, q.status("another-credential").orElseThrow().repeatCount(),
"an unrelated credential's attempt count must stay at 1, unaffected by a sibling's streak");
}
@Test
void noneReportsNoStatusForAnything() {
BackendQuarantine q = BackendQuarantine.none();
q.quarantine("shared-openai"); // no-op on none(), same as every other mutator
assertTrue(q.status("shared-openai").isEmpty(),
"none() quarantines nothing, so it must report no status at all -- never a fabricated "
+ "attempt count for a credential that was never actually quarantined");
}
/** The flat (non-escalating) two-argument constructor must still report a real, growing count. */
@Test
void aFlatTwoArgumentInstanceStillReportsAGrowingAttemptCountEvenThoughTheCooldownStaysFlat() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.SECONDS.toNanos(600));
q.quarantine("shared-openai");
BackendQuarantine.Status first = q.status("shared-openai").orElseThrow();
assertEquals(600L, first.remainingSeconds());
assertEquals(1, first.repeatCount());
now.set(TimeUnit.SECONDS.toNanos(100)); // still inside the 1st quarantine
q.quarantine("shared-openai");
BackendQuarantine.Status second = q.status("shared-openai").orElseThrow();
assertEquals(600L, second.remainingSeconds(),
"the flat constructor's cooldown must stay exactly the base length regardless of the streak");
assertEquals(2, second.repeatCount(),
"the flat constructor still counts the real streak -- it just does not scale the cooldown by it");
}
}
@@ -1,16 +1,13 @@
package dev.ltms.fleet.session;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.IThrowableProxy;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
@@ -489,36 +486,33 @@ class GitWorktreesTest {
// ---- CB-189: broader remote-URL coverage — every remote, both fetch and push URLs, any
// non-SSH scheme. Reporting only, additive to the origin/https strip-and-refuse tests above. ----
private Logger reportingLogger;
private ListAppender<ILoggingEvent> reportingAppender;
private CapturedLog reportingLog;
/** {@link GitWorktrees}'s own logger, captured fresh for each test so assertions never see a
* message left over from a previous test. */
* message left over from a previous test. fleetd #529: {@link CapturedLog#close} restores the
* level it captured here (whatever it truly was before this test, not just WARN), so a test
* below that further lowers the level to INFO for its own assertion (via {@link
* #reportingLog}'s {@link CapturedLog#setLevel}) can never leak that INFO pin past its own
* {@code @AfterEach} — every test's window is self-contained. */
@BeforeEach
void attachReportingLogCapture() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
reportingLogger = ctx.getLogger(GitWorktrees.class);
reportingLogger.setLevel(Level.WARN);
reportingAppender = new ListAppender<>();
reportingAppender.setContext(ctx);
reportingAppender.start();
reportingLogger.addAppender(reportingAppender);
reportingLog = CapturedLog.at(GitWorktrees.class, Level.WARN);
}
@AfterEach
void detachReportingLogCapture() {
reportingLogger.detachAppender(reportingAppender);
reportingLog.close();
}
private List<String> capturedMessages() {
return reportingAppender.list.stream().map(ILoggingEvent::getFormattedMessage).toList();
return reportingLog.events().stream().map(ILoggingEvent::getFormattedMessage).toList();
}
/** Asserts {@code secret} appears in no captured message, and in no attached exception's
* message either — the constraint is that a credential must never reach a log, however it
* would have gotten there. */
private void assertNoLeak(String secret) {
for (ILoggingEvent event : reportingAppender.list) {
for (ILoggingEvent event : reportingLog.events()) {
assertFalse(event.getFormattedMessage().contains(secret),
"log message leaked a credential (" + secret + "): " + event.getFormattedMessage());
IThrowableProxy thrown = event.getThrowableProxy();
@@ -596,7 +590,7 @@ class GitWorktreesTest {
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-189-d", "HEAD");
assertTrue(reportingAppender.list.isEmpty(),
assertTrue(reportingLog.events().isEmpty(),
"expected no report for an ssh remote and a credential-free https remote, got:\n"
+ capturedMessages());
}
@@ -812,7 +806,7 @@ class GitWorktreesTest {
/** Criterion 1, all three present: the summary names the denominator and every neutralized file. */
@Test
void isolateToolSurfaceLogsAllThreeConfigsNeutralized(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = initRepoWithAllThreeConfigs(tmp.resolve("repo"));
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-134-log-all", "HEAD");
@@ -827,7 +821,7 @@ class GitWorktreesTest {
/** Criterion 1, two absent: the summary must still name the denominator and say why. */
@Test
void isolateToolSurfaceLogsAbsentConfigsWithReason(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = initRepo(tmp.resolve("repo")); // only .mcp.json + README committed
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-134-log-partial", "HEAD");
@@ -1435,7 +1429,7 @@ class GitWorktreesTest {
/** Criterion 2: both candidates present — the summary line names both and the denominator. */
@Test
void overlayParityLogsBothCopiedWhenBothCandidatesArePresent(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = initRepo(tmp.resolve("repo"));
Files.writeString(repo.resolve(".env"), "A=1\n");
Files.writeString(repo.resolve(".envrc"), "export A=1\n");
@@ -1452,7 +1446,7 @@ class GitWorktreesTest {
* denominator, and why the other candidate was not copied. */
@Test
void overlayParityLogsOneCopiedOneAbsent(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = initRepo(tmp.resolve("repo"));
Files.writeString(repo.resolve(".env"), "A=1\n");
// .envrc deliberately not created — the absent candidate.
@@ -1472,7 +1466,7 @@ class GitWorktreesTest {
* change, silently, unless this line told it so beforehand. */
@Test
void overlayParityLogsSkipWorktreeConsequenceForATrackedFile(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = initRepo(tmp.resolve("repo"));
Files.writeString(repo.resolve(".env"), "A=1\n");
git(repo, "add", ".env");
@@ -1495,7 +1489,7 @@ class GitWorktreesTest {
/** Criterion 5: null and empty overlay lists return quietly — no exception, no log noise. */
@Test
void overlayParityWithNoCandidatesLogsNothing(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = initRepo(tmp.resolve("repo"));
Path wt = bareWorktree(repo, tmp.resolve("wt"), "cb134-empty");
GitWorktrees worktrees = new GitWorktrees(tmp.resolve("wts").toString());
@@ -1503,7 +1497,7 @@ class GitWorktreesTest {
worktrees.overlayParity(repo.toString(), wt.toString(), null);
worktrees.overlayParity(repo.toString(), wt.toString(), List.of());
assertTrue(reportingAppender.list.isEmpty(),
assertTrue(reportingLog.events().isEmpty(),
"a null/empty overlay must log nothing, got:\n" + capturedMessages());
}
@@ -1683,7 +1677,7 @@ class GitWorktreesTest {
* denominator, what was seeded, and what was kept because the repo already had it. */
@Test
void seedSkillsLogsSeededAndKept(@TempDir Path tmp) throws Exception {
reportingLogger.setLevel(Level.INFO);
reportingLog.setLevel(Level.INFO);
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
@@ -1,9 +1,7 @@
package dev.ltms.fleet.session;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.MemberLifecycle;
import dev.ltms.fleet.config.ConfigRef;
@@ -25,7 +23,13 @@ import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.AfterAll;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.MethodOrderer;
import org.junit.jupiter.api.Order;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.TestMethodOrder;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
@@ -50,9 +54,46 @@ import static org.junit.jupiter.api.Assertions.*;
* CB-301 / CB-303 acceptance tests for the authoritative session registry, one-shot lifecycle FSM,
* and configurable lifecycle limits (idle TTL, context cap, drain).
* No live herdr — everything runs against the same {@link FakeHerdr} the rest of the project uses.
*
* <p>fleetd #525: only {@link #onTurnFailedIsLoggedAtWarnWithThePriorState} (explicitly
* {@link Order#value() @Order(1)}) and the proving test right after it
* ({@link #sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn}, {@code @Order(2)})
* care about method order — every other test here has no {@code @Order} and so runs after both of
* these (JUnit 5's {@link MethodOrderer.OrderAnnotation} gives an unannotated method the lowest
* priority), in whatever relative order it already ran in.
*/
@TestMethodOrder(MethodOrderer.OrderAnnotation.class)
class SessionManagerTest {
/**
* fleetd #525: the level {@link SessionManager}'s logger had when this class started, captured
* before any test here — including the leak this ticket fixes — can touch it. {@code
* pinSessionManagerLoggerToAKnownBaseline} then forces a distinctive, known value (DEBUG) so
* {@link #sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn} can tell "the
* level came back to what it was" apart from "the level happens to already be WARN because
* some earlier test class in this JVM fork (surefire reuses forks by default) left it there" —
* a real risk, since {@code ch.qos.logback.classic.Logger} instances are cached per class and
* shared across the whole JVM, and this exact logger is also touched by
* {@code WorktreeSessionManagerTest#releasePreservesDirtyWorktreeAndLogsWarn}, which has the
* same unfixed leak (reported, not fixed — out of this ticket's scope).
*/
private static Level sessionManagerLevelBeforeThisClass;
@BeforeAll
static void pinSessionManagerLoggerToAKnownBaseline() {
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
sessionManagerLevelBeforeThisClass = sessionLog.getLevel();
sessionLog.setLevel(Level.DEBUG);
}
@AfterAll
static void restoreSessionManagerLoggerLevel() {
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
sessionLog.setLevel(sessionManagerLevelBeforeThisClass);
}
private SessionManager sessionManager(FakeHerdr herdr) {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
@@ -81,6 +122,9 @@ class SessionManagerTest {
return new SessionManager(workers, worktrees, clock);
}
// fleetd #529: CapturedLog moved to dev.ltms.fleet.testing.CapturedLog (imported above) so
// every test file shares one implementation instead of each hand-rolling its own capture.
/**
* CB-581: a {@link Worktrees} test double whose {@code hasUncommitted} and {@code remove} can
* be told to throw, so {@link SessionManager#release} can be exercised against exactly the
@@ -466,24 +510,15 @@ class SessionManagerTest {
@Test
void backendErrorForUnknownTargetIsWarnedAndDoesNotCreateASession() {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
try {
try (CapturedLog log = CapturedLog.of(SessionManager.class)) {
SessionManager sessions = sessionManager(new FakeHerdr());
assertFalse(sessions.onBackendError("term_missing", "backend exited"));
assertTrue(sessions.roster().isEmpty(), "unknown target must not create a session");
assertTrue(appender.list.stream().anyMatch(e -> e.getLevel().equals(Level.WARN)
assertTrue(log.events().stream().anyMatch(e -> e.getLevel().equals(Level.WARN)
&& e.getFormattedMessage().contains("term_missing")),
"unknown target is logged at WARN");
} finally {
sessionLog.detachAppender(appender);
}
}
@@ -516,19 +551,12 @@ class SessionManagerTest {
}
@Test
@Order(1)
void onTurnFailedIsLoggedAtWarnWithThePriorState() {
// CB-564: this transition used to be a bare DEBUG "session marked failed" — a symptom with no
// cause. A member that can no longer be delegated to must be at least WARN, and should name
// what stage it failed at (here: BUSY, i.e. a turn was in flight and never resolved).
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.WARN)) {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
@@ -538,18 +566,37 @@ class SessionManagerTest {
sessions.onTurnFailed(terminal);
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no turn-failed WARN logged");
assertTrue(warn.contains(terminal), "the log names the member: " + warn);
assertTrue(warn.contains("BUSY"), "the log names the stage it failed at: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
/**
* fleetd #525: proves the leak in {@link #onTurnFailedIsLoggedAtWarnWithThePriorState} above
* (which runs immediately before this, via {@code @Order}) is closed. That test pins the
* shared {@link SessionManager} logger to WARN through a {@link CapturedLog}; if {@link
* CapturedLog#close} only detached the appender — the original bug, before this ticket's fix —
* the level would still read WARN here instead of the {@code DEBUG} baseline this class's
* {@code @BeforeAll} set. Runs at {@code @Order(2)}, guaranteed after {@code @Order(1)} and
* before every other (unannotated) test in this class.
*/
@Test
@Order(2)
void sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn() {
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
assertEquals(Level.DEBUG, sessionLog.getLevel(),
"onTurnFailedIsLoggedAtWarnWithThePriorState pins the shared SessionManager logger "
+ "to WARN; its cleanup must restore the level it captured (DEBUG, set by "
+ "this class's @BeforeAll) rather than leaving WARN pinned for every test "
+ "that runs after it");
}
/**
* fleetd #226: a contended slot is refused through the real {@link SessionManager#acquire}
* path before the real launcher can hand an architect charter to a process.
@@ -576,28 +623,18 @@ class SessionManagerTest {
SessionManager sessions = sessionManager(herdr);
MemberRegistry members = architectRegistry();
sessions.setMemberLifecycle(bindFailureAfterReservation(members));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger registryLog = (ch.qos.logback.classic.Logger)
LoggerFactory.getLogger(MemberRegistry.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
registryLog.addAppender(appender);
registryLog.setLevel(Level.WARN);
try {
try (CapturedLog log = CapturedLog.at(MemberRegistry.class, Level.WARN)) {
MemberSession session = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.DEV, session.role(), "a failed reservation bind must use the fallback");
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no slot-exhaustion WARN logged");
assertTrue(warn.contains("ltms-local"), "the WARN names the profile: " + warn);
assertTrue(warn.contains(session.terminalId()), "the WARN names the terminal: " + warn);
} finally {
registryLog.detachAppender(appender);
}
}
@@ -902,6 +939,82 @@ class SessionManagerTest {
.count();
}
/**
* fleetd #512: a drain that releases every session cleanly used to log nothing at all — the
* two log calls in {@code drainAll}/{@code drainSnapshot} both sit on abnormal paths, so
* "clean drain" and "died on the first session" were indistinguishable. This asserts the new
* {@code log.info} line fires on the ordinary, nothing-went-wrong path, and that its numbers
* are the real counts (two released, zero abandoned) rather than just a non-empty string.
*/
@Test
void drainAllLogsACompletionLineWithTheRealCountsOnACleanDrain() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession first = sessions.acquire("ltms-local", "/one", "/caller", "ownerOne");
MemberSession second = sessions.acquire("ltms-local", "/two", "/caller", "ownerTwo");
sessions.asPresence().markPresent(first.terminalId());
sessions.asPresence().markPresent(second.terminalId());
// Both stay READY — neither is delivered a turn, so neither is BUSY and the drain below
// has nothing abnormal to hit.
// Pin INFO explicitly: fleetd #525 made CapturedLog itself restore the level it pins, but
// this pin stays anyway as belt-and-braces — a later change to the sweep must not be able
// to make this INFO assertion vacuous again by leaving some other test's WARN pin in place.
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.INFO)) {
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(sessions.roster().isEmpty(), "precondition: the drain actually ran");
String info = log.events().stream()
.filter(e -> e.getLevel().equals(Level.INFO))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("drain complete"))
.findFirst()
.orElse("no drain-complete INFO logged");
assertTrue(info.contains("released=2"),
"both released sessions must be counted: " + info);
assertTrue(info.contains("abandoned=0"),
"neither session was BUSY, so nothing was abandoned mid-turn: " + info);
}
}
/**
* fleetd #512: the same completion line must also report a non-zero abandoned count when a
* session is still {@code BUSY} once the whole-drain deadline passes — the case the ticket
* calls out as the one a script needs to be able to see. Reuses the same BUSY/READY mix as
* {@link #drainAllReleasesBusyAndReadySessionsAndWaitsForBusy}, which already forces the busy
* session to spin until the real-time deadline expires (its state never leaves BUSY on its
* own), and adds the log assertion that test does not make.
*/
@Test
void drainAllLogsANonZeroAbandonedCountForASessionStillBusyAtTheDeadline() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId(), TestTurnTokens.inert(busy.terminalId()));
// busy never leaves BUSY — no completion is delivered — so the drain below must spin the
// full timeout and then release it anyway, counting it abandoned.
// Pin INFO explicitly — see the comment in drainAllLogsACompletionLineWithTheRealCountsOnACleanDrain.
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.INFO)) {
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(sessions.roster().isEmpty(), "precondition: the drain actually ran");
String info = log.events().stream()
.filter(e -> e.getLevel().equals(Level.INFO))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("drain complete"))
.findFirst()
.orElse("no drain-complete INFO logged");
assertTrue(info.contains("released=2"),
"both the ready and the busy session are released: " + info);
assertTrue(info.contains("abandoned=1"),
"the busy session hit the deadline still BUSY and must be counted: " + info);
}
}
// --- fleetd #308: a spawn accepted while the shutdown drain is running must not orphan ---
@Test
@@ -1202,21 +1315,13 @@ class SessionManagerTest {
new WorktreeRequest("cb-581a", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.WARN)) {
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a throwing dirty check must not abort the release");
assertTrue(worktrees.removeCalls().isEmpty(),
"the worktree is preserved when its dirty state cannot be determined");
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains(s.worktree()))
@@ -1224,8 +1329,6 @@ class SessionManagerTest {
.orElse("no warn logged naming the worktree");
assertTrue(warn.contains(s.paneId()), "the WARN names the pane: " + warn);
assertTrue(warn.contains(s.terminalId()), "the WARN names the terminal: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
@@ -1392,20 +1495,12 @@ class SessionManagerTest {
// this itself, so it no longer propagates out of release() at all.
worktrees.failRemoveFor(b.worktree());
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
int reaped;
try {
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.WARN)) {
clock[0] = 100;
reaped = sessions.reapIdle(10);
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains(b.paneId()))
@@ -1413,8 +1508,6 @@ class SessionManagerTest {
.orElse("no worktree-removal-failure WARN logged");
assertTrue(warn.contains(b.terminalId()), "the WARN names the failed session's terminal: " + warn);
assertTrue(warn.contains(b.worktree()), "the WARN names the failed session's worktree: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
assertEquals(3, reaped,
@@ -1460,28 +1553,18 @@ class SessionManagerTest {
// trigger reapIdle's own guard is for, now that #283 closed the worktree-removal trigger.
herdr.paneCloseFailsForPane("w9:pRoot_2", "internal_error");
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
int reaped;
try {
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.WARN)) {
clock[0] = 100;
reaped = sessions.reapIdle(10);
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("reap failed") && m.contains(b.paneId()))
.findFirst()
.orElse("no reap-failed WARN logged for the failing session");
assertTrue(warn.contains(b.terminalId()), "the WARN names the failed session's terminal: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
assertEquals(2, reaped,
@@ -7,13 +7,17 @@ import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.TestTurnTokens;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.AfterAll;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.MethodOrderer;
import org.junit.jupiter.api.Order;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.TestMethodOrder;
import org.slf4j.LoggerFactory;
import java.util.List;
@@ -27,9 +31,41 @@ import static org.junit.jupiter.api.Assertions.*;
* CB-301-ext acceptance tests for worktree provisioning and config-parity overlay.
* No live git — every Worktrees call is handled by {@link FakeWorktrees} and every herdr
* call by {@link FakeHerdr}, matching the project's fake-based test style.
*
* <p>fleetd #529: only {@link #releasePreservesDirtyWorktreeAndLogsWarn} (explicitly
* {@link Order#value() @Order(1)}) and the proving test right after it
* ({@link #sharedSessionManagerLoggerLevelIsRestoredAfterDirtyWorktreeReleasePinsWarn},
* {@code @Order(2)}) care about method order — every other test here has no {@code @Order} and so
* runs after both (JUnit 5's {@link MethodOrderer.OrderAnnotation} gives an unannotated method the
* lowest priority), in whatever relative order it already ran in.
*/
@TestMethodOrder(MethodOrderer.OrderAnnotation.class)
class WorktreeSessionManagerTest {
/**
* fleetd #529: the level {@link SessionManager}'s logger had when this class started, captured
* before any test here touches it, then forced to a distinctive, known value (TRACE) so {@link
* #sharedSessionManagerLoggerLevelIsRestoredAfterDirtyWorktreeReleasePinsWarn} can tell "the
* level came back to what it was" apart from "the level happens to already be WARN because
* some other test class in this JVM fork (surefire reuses forks by default) left it there".
*/
private static ch.qos.logback.classic.Level sessionManagerLevelBeforeThisClass;
@BeforeAll
static void pinSessionManagerLoggerToAKnownBaseline() {
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
sessionManagerLevelBeforeThisClass = sessionLog.getLevel();
sessionLog.setLevel(ch.qos.logback.classic.Level.TRACE);
}
@AfterAll
static void restoreSessionManagerLoggerLevel() {
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
sessionLog.setLevel(sessionManagerLevelBeforeThisClass);
}
private static MemberRegistry members() {
return new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("architect", new FleetConfig.Slot("ltms-local")),
@@ -254,6 +290,7 @@ class WorktreeSessionManagerTest {
* path, the session, and the cause an operator needs to find the work.
*/
@Test
@Order(1)
void releasePreservesDirtyWorktreeAndLogsWarn() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
@@ -262,21 +299,13 @@ class WorktreeSessionManagerTest {
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-576", null));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
try (CapturedLog log = CapturedLog.at(SessionManager.class, Level.WARN)) {
sessions.release(s.paneId());
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
assertTrue(worktrees.removeCalls().isEmpty(),
"a dirty worktree is never removed — it holds the only copy of the work");
String warn = appender.list.stream()
String warn = log.events().stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("dirty worktree"))
@@ -285,11 +314,38 @@ class WorktreeSessionManagerTest {
assertTrue(warn.contains(s.worktree()), "the WARN names the worktree path: " + warn);
assertTrue(warn.contains(s.terminalId()), "the WARN names the session: " + warn);
assertTrue(warn.contains("COMPLETED"), "the WARN names the release cause: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
/**
* fleetd #529 proving test: pins that the leak this ticket fixes stays fixed. {@link
* #releasePreservesDirtyWorktreeAndLogsWarn} above (which runs immediately before this, via
* {@code @Order}) pins the shared {@link SessionManager} logger to WARN through a {@link
* CapturedLog}; if {@link CapturedLog#close} only detached the appender — the original bug —
* the level would still read WARN here instead of the {@code TRACE} baseline this class's
* {@code @BeforeAll} set. Runs at {@code @Order(2)}, guaranteed after {@code @Order(1)} and
* before every other (unannotated) test in this class.
*
* <p>This proves only the WITHIN-CLASS case: JUnit 5's {@code @TestMethodOrder} orders methods
* inside one class, not test classes relative to each other, and surefire's default class order
* is not something a single test can force. The cross-class leak fleetd #525 measured — this
* class's {@code releasePreservesDirtyWorktreeAndLogsWarn} pinning WARN and bleeding into a
* later-running {@code SessionManagerTest} in the same fork — is fixed by the same {@link
* CapturedLog} mechanism proven here, but that cross-class ordering itself is NOT asserted by
* any test and remains unproven by construction.
*/
@Test
@Order(2)
void sharedSessionManagerLoggerLevelIsRestoredAfterDirtyWorktreeReleasePinsWarn() {
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
assertEquals(ch.qos.logback.classic.Level.TRACE, sessionLog.getLevel(),
"releasePreservesDirtyWorktreeAndLogsWarn pins the shared SessionManager logger to "
+ "WARN; its cleanup must restore the level it captured (TRACE, set by this "
+ "class's @BeforeAll) rather than leaving WARN pinned for every test that "
+ "runs after it");
}
/**
* CB-576 review (fleetd #116). A worktree that is already gone (operator cleanup,
* {@code git worktree prune}, an earlier half-completed release) must not break teardown.
@@ -0,0 +1,112 @@
package dev.ltms.fleet.testing;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import org.slf4j.LoggerFactory;
import java.util.List;
/**
* fleetd #529 (promoted from {@code SessionManagerTest}, merged in #527 for fleetd #525): captures
* a logger's output and, on {@link #close}, restores <em>both</em> the appender and the level to
* what they were before.
*
* <p>{@code ch.qos.logback.classic.Logger} instances are cached per class and shared across the
* whole JVM — and surefire reuses forks by default — so a bare {@code addAppender}/{@code
* setLevel} pair whose {@code finally} only detaches the appender leaves the level pinned for
* every test that runs after it, in the same class or in a completely unrelated one sharing the
* fork. try-with-resources makes "restored the appender but not the level" impossible to write,
* because there is only one thing to close.
*
* <p>New code must use this rather than hand-rolling the {@code ListAppender} + {@code setLevel} +
* {@code finally detachAppender} pattern: use {@link #at} or {@link #of}. It is <em>not</em> yet
* the only instance of the pattern in this test tree, and the earlier wording here said it was —
* which would leave a reader who greps unable to tell a leftover from a violation.
*
* <p>Measured on main at af95897 (2026-09-12): nine test files still hand-roll it, with 42
* {@code setLevel} calls on a raw logback {@code Logger} between them — {@code
* FleetdStartupReportTest}, {@code GitHostShapeReportTest}, {@code MemberCredentialsGapReportTest},
* {@code MemberTrustModelReportTest}, {@code FleetHealthMonitorTest}, {@code
* ClaudeCodeLauncherTest}, {@code HerdrPeerLauncherAllowListWiringTest}, {@code
* HerdrPeerLauncherCharterTest} and {@code OpenCodeLauncherTest}. Every one of them pairs its pin
* with a restore, so none is the fleetd #525 leak and none was in fleetd #529's scope, which was
* the 19 <em>unrestored</em> pins only. They are unmigrated, not broken.
*
* <p>Re-measure with the two commands below, from the repo root. A file that appears in the first
* list and not the second still hand-rolls the pattern. When the first list comes back empty, this
* paragraph is spent and the sentence above can go back to saying "the one way" — delete the
* paragraph then rather than updating the count.
*
* <pre>{@code
* grep -rlE '\.setLevel\(' fleetd/src/test/java --include='*.java' | grep -v CapturedLog.java
* grep -rl 'CapturedLog' fleetd/src/test/java --include='*.java'
* }</pre>
*/
public final class CapturedLog implements AutoCloseable {
private final Logger logger;
private final Level originalLevel;
private final ListAppender<ILoggingEvent> appender;
private CapturedLog(Logger logger, Level pinnedLevel) {
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
this.logger = logger;
this.originalLevel = logger.getLevel();
this.appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
logger.addAppender(appender);
if (pinnedLevel != null) {
logger.setLevel(pinnedLevel);
}
}
/** Capture {@code loggerClass}'s output, pinning its level to {@code pinnedLevel} for the
* duration of the try-with-resources block. */
public static CapturedLog at(Class<?> loggerClass, Level pinnedLevel) {
return new CapturedLog((Logger) LoggerFactory.getLogger(loggerClass), pinnedLevel);
}
/** Capture {@code loggerClass}'s output without changing its level. */
public static CapturedLog of(Class<?> loggerClass) {
return new CapturedLog((Logger) LoggerFactory.getLogger(loggerClass), null);
}
/**
* Capture the named logger's output, pinning its level to {@code pinnedLevel}. For a logger
* obtained in production code via {@code LoggerFactory.getLogger("some-name")} rather than a
* class — e.g. {@code AuditLog}'s {@code "audit"} logger — where {@link #at(Class, Level)}
* would capture the wrong {@code Logger} instance (class name and logger name are different
* strings and resolve to different cached loggers).
*/
public static CapturedLog at(String loggerName, Level pinnedLevel) {
return new CapturedLog((Logger) LoggerFactory.getLogger(loggerName), pinnedLevel);
}
/** Capture the named logger's output without changing its level. See {@link #at(String, Level)}. */
public static CapturedLog of(String loggerName) {
return new CapturedLog((Logger) LoggerFactory.getLogger(loggerName), null);
}
public List<ILoggingEvent> events() {
return appender.list;
}
/**
* Re-pin the level while this capture is still open — for example to lower it further for one
* assertion inside a test whose fixture already pinned a coarser baseline in {@code @BeforeEach}.
* This does not change what {@link #close} restores: that is always the level captured when
* this instance was created, never a value set through this method.
*/
public void setLevel(Level level) {
logger.setLevel(level);
}
@Override
public void close() {
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
}
+118 -33
View File
@@ -64,7 +64,106 @@
# The two outputs side by side are the finding: any name whose hash matches between them is a
# credential the member holds in full.
#
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
#
# This used to feed the parser straight into `mapfile -t _FIELDS < <(producer)`. That form cannot
# see the producer fail: `<` `<(...)` is a process substitution, not a pipeline, so `set -o
# pipefail` does not reach inside it, and mapfile's own exit status reports whether the BUILTIN
# ran, not whether the command substituted into it succeeded — a failing jq or python3 there still
# leaves mapfile at rc=0 with an empty array, read as a parse that genuinely found nothing (fleetd
# #500). Capturing the parser's output with command substitution first, and checking ITS exit
# status, reports the producer's real failure while the fact still exists — before it is handed to
# mapfile at all.
#
# mapfile then reads from that captured string with `<<<` (a herestring), not `< <(...)`: `<<<`
# materialises the whole string in memory first, where `< <(...)` would stream it. That only
# matters for a large producer; this one is a short credential-name policy response, so the
# tradeoff is irrelevant here — noted because it would not be for every producer.
parse_policy_fields() {
if command -v jq >/dev/null 2>&1; then
_FIELDS_RAW="$(printf '%s' "$POLICY_JSON" | jq -r '
(.present | tostring),
(.policy // ""),
(.knownCount // 0 | tostring),
(.allowedCount // 0 | tostring),
(.blockedCount // 0 | tostring),
(.known[]? // empty)')"
_PARSE_STATUS=$?
_PARSER_NAME="jq"
else
_FIELDS_RAW="$(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
import json, sys
data = json.load(sys.stdin)
print(str(data.get("present")))
print(data.get("policy") or "")
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
for n in (data.get("known") or []):
print(n)
PY
)"
_PARSE_STATUS=$?
_PARSER_NAME="python3"
fi
if [ "$_PARSE_STATUS" -ne 0 ]; then
echo "refusing to run: could not parse the policy fetched from $POLICY_URL — $_PARSER_NAME exited" \
"non-zero (status $_PARSE_STATUS). That is a parser failure, not a claim about the policy" \
"itself; the policy response has not been read." >&2
return 4
fi
# A herestring adds a newline, so mapfile would turn an empty parser result into one empty field.
# Keep that case separate so the refusal reports what the parser actually returned: zero fields.
if [ -z "$_FIELDS_RAW" ]; then
_FIELDS=()
else
mapfile -t _FIELDS <<< "$_FIELDS_RAW"
fi
# Arity check — the CORRECTNESS fix (fleetd #500). A parser that exits 0 can still return fewer
# than the 5 fixed fields (present, policy mode, 3 counts) that every fixed-field read in main() expects,
# whatever the reason: a producer that printed nothing, malformed JSON that jq/python3 still
# accepted, or a schema change upstream. main()'s slice (`_FIELDS[@]:5`) does not fire
# `set -u` on an unset OR a short array, and every fixed-field read there used a `:-` default, so
# without this check a short `_FIELDS` reaches the "0 known names" guard further down with the
# same look as a policy that genuinely has 0 names. Check the count here, at the one point the
# fact is still present, before the slice consumes it.
if (( ${#_FIELDS[@]} < 5 )); then
echo "refusing to run: the policy parser ($_PARSER_NAME) returned ${#_FIELDS[@]} field(s); at" \
"least 5 are required (present, policy mode, knownCount, allowedCount, blockedCount). The" \
"parse ran but its shape is wrong — this is not a claim about how many names the policy" \
"knows." >&2
return 5
fi
}
main() {
set -uo pipefail
# `pipefail` is not what catches the parser failure handled in parse_policy_fields() above (fleetd #500): in
# `printf '%s' "$POLICY_JSON" | jq -r '...'`, jq is the LAST element of the pipe, so the pipeline's
# own exit status is already jq's status, with or without pipefail. It is kept as insurance for if
# a post-processing stage is ever appended after the parser (e.g. `| tail -n +2`) — at that point
# the parser would sit upstream and pipefail becomes the only thing that still reports its status.
# --- refuse on an interpreter that cannot run this script (fleetd #500) -------------------------
#
# mapfile, used below to parse the policy response, was added in bash 4.0. macOS ships bash 3.2.57
# at /bin/bash, which predates it. This script's own `set -uo pipefail` does not catch a missing
# mapfile: the builtin just fails with "command not found" on stderr, and every line below that
# reads the array it would have filled uses a `:-` default or a slice, neither of which `set -u`
# catches on an unset array. Left unguarded, that chain ends in the "0 known names" refusal further
# down — a claim about the POLICY, for a failure that is actually about the INTERPRETER. So the
# interpreter is checked once, explicitly, before it is asked to do anything mapfile depends on.
if (( ${BASH_VERSINFO[0]} < 4 )); then
echo "refusing to run: this script uses mapfile, which needs bash 4 or newer. This shell is bash" \
"${BASH_VERSION:-<unknown, no \$BASH_VERSION>}. Re-run it under a newer bash, for example:" \
"\"\$(command -v bash)\" \"$0\"" "$@" >&2
exit 3
fi
FLEETD_HOST="${FLEETD_HOST:-http://127.0.0.1:8765}"
POLICY_URL="${FLEETD_HOST%/}/member-credentials"
@@ -121,31 +220,10 @@ EOF
exit 1
fi
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
if command -v jq >/dev/null 2>&1; then
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | jq -r '
(.present | tostring),
(.policy // ""),
(.knownCount // 0 | tostring),
(.allowedCount // 0 | tostring),
(.blockedCount // 0 | tostring),
(.known[]? // empty)')
else
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
import json, sys
data = json.load(sys.stdin)
print(str(data.get("present")))
print(data.get("policy") or "")
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
for n in (data.get("known") or []):
print(n)
PY
)
fi
# Parse the policy in one pass and refuse on any of the three failure causes. The decision, the
# three refusals and the reasoning behind each live in parse_policy_fields() above — kept there
# with the code rather than here, so the explanation cannot drift away from what it explains.
parse_policy_fields || exit $?
PRESENT="${_FIELDS[0]:-null}"
POLICY_MODE="${_FIELDS[1]:-}"
@@ -164,19 +242,21 @@ case "$KNOWN_COUNT_REPORTED" in
;;
esac
# --- guard the denominator explicitly — never proceed on a zero/short count ---------------------
# --- guard the denominator explicitly — never proceed on a zero count ---------------------------
#
# This is the exact trap named in the ticket: an empty (or truncated) NAMES array passes every
# subsequent "is it set" check vacuously and prints a table that LOOKS complete. So this is checked
# before anything else runs, with a message that says why, not just that it failed.
# This is the exact trap named in the ticket: an empty NAMES array passes every subsequent "is it
# set" check vacuously and prints a table that LOOKS complete. By this point the interpreter gate,
# the parser-exit-status check, and the arity check above have already ruled out "the interpreter
# couldn't run mapfile", "the parser failed", and "the parser returned the wrong shape" — so a zero
# count reaching here really does mean the policy itself reports 0 known names, not a swallowed
# failure upstream. That is still checked before anything else runs, with a message that says so.
if [ "${#NAMES[@]}" -eq 0 ] || [ "$KNOWN_COUNT_REPORTED" -eq 0 ]; then
cat >&2 <<EOF
refusing to run: the policy fetched from $POLICY_URL contains 0 known names (present=${PRESENT:-unknown}).
Either memberCredentials: is absent/empty on the running daemon (nothing is protected — see fleetd's
own startup warning), or the response could not be parsed. Either way, checking zero names would
print a clean-looking table for a policy that protects nothing, or for a probe that read nothing.
This is refused rather than reported as a pass.
memberCredentials: is absent or empty on the running daemon — nothing is protected (see fleetd's own
startup warning). Checking zero names would print a clean-looking table for a policy that protects
nothing. This is refused rather than reported as a pass.
EOF
exit 1
fi
@@ -251,3 +331,8 @@ How to read this:
hardcoded list did. If the daemon's policy changes, the next run of this script reflects it
with no edit to this file.
EOF
}
if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then
main "$@"
fi
+741 -66
View File
@@ -5,7 +5,7 @@
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
#
# This script exists to turn six remembered traps into one auditable command:
# This script exists to turn eight remembered traps into one auditable command:
#
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
@@ -27,6 +27,29 @@
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
# at any moment is whichever one you asked to act, never both.
# 7. fleetd #492 — a systemd --user unit is a THIRD possible supervisor (seen on a second host):
# Restart=on-failure treats this JVM's SIGTERM exit code (143, per CB-594 above) as a failure
# too, so a bare `kill` there would race systemd's own restart of the OLD jar exactly like
# launchd would. This script now tells launchd, systemd, and "genuinely unsupervised" apart as
# three different answers, drives whichever one it finds through its own control plane
# (`launchctl` / `systemctl --user`), and REFUSES outright — never falls back to `kill` — when
# it finds a supervision signal it cannot map to exactly one of the two it knows how to drive.
# A wrong guess here is how two daemons end up running against one herdr session. Follow-up:
# "not currently loaded" is not the same fact as "unsupervised" — a unit that is installed but
# activating/failed/pending-restart, or a `systemctl` call that could not answer at all (e.g.
# no user-bus access), both now read as a fifth answer, "unclear", and REFUSE the same way
# "ambiguous" does, rather than silently falling through to "none".
# 8. fleetd #492 — a post-restart check counts running fleetd processes and fails the whole run if
# more than one is alive. That is the one thing none of the checks above (healthz 200, jar id,
# the fresh "listening" line) can see: every one of them is satisfied by EITHER daemon.
# 9. fleetd #512 — the ERROR-line count above is blind by construction to the exact failure #493
# is about: an uncaught exception in a shutdown thread never passes through the logger, so it
# never carries an ERROR (or SEVERE) token that any count could see. This script now also
# greps the previous daemon's shutdown window for that exception's real shape, and separately
# asserts that SessionManager's drain-complete line (fleetd #522) is present there — its
# absence is the real signal, because a drain that dies on its first session prints nothing
# else either. Warns loudly; never fails the redeploy, because by the time this is detectable
# the new daemon is already up and healthy.
#
# Usage:
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
@@ -41,6 +64,13 @@ set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MODULE="$REPO/fleetd"
JAR="$MODULE/target/fleetd.jar"
# fleetd #493: never build into the path a running process holds. The build writes here first
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
# stage_built_jar/swap_staged_jar below.
JAR_STAGED="$MODULE/target/fleetd-new.jar"
OUT="$MODULE/fleetd.out"
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
@@ -55,13 +85,42 @@ HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
LAUNCHD_LABEL='dev.ltms.fleetd'
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
# fleetd #492: the systemd --user unit this script must not fight with either (see trap 7 above).
# Measured on the second host: `systemctl --user cat fleetd` names the unit "fleetd" (not
# "dev.ltms.fleetd" — systemd user units here are not namespaced the way the launchd label is).
SYSTEMD_UNIT='fleetd'
# fleetd #492 follow-up: detect_supervisor packs TWO values (kind, detail) onto the one stdout
# line that survives its $(...) call — see the constraints comment above that function. This is
# the separator between them: the ASCII "unit separator" byte, chosen because it never occurs in
# any of the prose detail strings and needs no escaping in a `case`/glob pattern.
SUPERVISOR_DETAIL_SEP=$'\x1f'
# fleetd #492 follow-up: set by systemd_loaded/systemd_installed when the underlying `systemctl`
# call could not answer cleanly — it exited non-zero AND wrote something to stderr, which is a real
# tool failure (e.g. it cannot reach the user bus over a non-lingering ssh session), never the same
# fact as a clean negative answer ("not active", no stderr). Initialized here, not just inside the
# probes, so detect_supervisor can read them under `set -u` even before either probe has ever run,
# and so a test that stubs a probe with a plain `return 0`/`return 1` body (leaving these untouched)
# reads a deterministic 0 rather than whatever a previous probe call left behind.
SYSTEMD_LOADED_ERRORED=0
SYSTEMD_INSTALLED_ERRORED=0
# fleetd #492 follow-up: SUPERVISOR_UNCLEAR_DETAIL is the specific supervisor/reason that
# require_drivable_supervisor's die() names on an "unclear" answer. Deliberately NOT pre-declared
# here (unlike the two flags above): it is set only by the real call site, right after it unpacks
# detect_supervisor's stdout (see the constraints comment above detect_supervisor). If that call
# site is ever skipped or broken, a bare `set -u` reference to this variable in
# require_drivable_supervisor must fail loudly with "unbound variable" — a pre-declared empty
# default would instead silently print an empty reason, hiding exactly the value this ticket
# exists to surface.
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
for arg in "$@"; do
case "$arg" in
--yes|-y) ASSUME_YES=1 ;;
--no-build) DO_BUILD=0 ;;
--check) CHECK_ONLY=1 ;;
-h|--help) sed -n '3,37p' "${BASH_SOURCE[0]}"; exit 0 ;;
-h|--help) sed -n '3,48p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
esac
done
@@ -71,14 +130,338 @@ ok() { printf ' ok %s\n' "$*"; }
warn() { printf ' WARN %s\n' "$*"; }
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; }
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
# jar right after a build (before it has been swapped in) without ever changing what a bare
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && shasum -a 256 "$f" | cut -c1-12 || echo "absent"; }
running_pid() { pgrep -f "$PATTERN" || true; }
# fleetd #493 — three small, independently testable pieces of "never build into the path a
# running process holds":
#
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
# path, immediately after a successful build. Dies (leaving the OLD daemon
# untouched — this runs before the stop step) if Maven reported success but
# left no jar behind, or if the move itself fails.
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
# jar already sitting at the live path from an earlier successful run, and
# die with the same truthful message this script has always used if not.
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
# daemon actually exited — extracted to its own function so the main flow
# can be relied on to call swap_staged_jar only AFTER this returns success,
# and so a test can prove that ordering by reading the script's own source.
# swap_staged_jar the actual swap: a plain `mv` of the staged jar onto the live path. Called
# only once the OLD daemon is confirmed gone (see wait_for_daemon_exit above),
# so this is never a write into a path a running process holds — by the time
# it runs, nothing holds that path anymore. If it fails, the caller must not
# start a new daemon: die() below already refuses that by exiting the script.
stage_built_jar() {
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
The running daemon was NOT touched."
mv -f "$JAR" "$JAR_STAGED" \
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
The running daemon was NOT touched."
}
require_no_build_jar() {
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
}
wait_for_daemon_exit() {
local timeout="$1" _i
for _i in $(seq "$timeout"); do
[ -z "$(running_pid)" ] && return 0
sleep 1
done
[ -z "$(running_pid)" ]
}
swap_staged_jar() {
local staged="$1" live="$2"
[ -f "$staged" ] || die "no staged jar at $staged to swap in — the daemon was NOT started."
mv -f "$staged" "$live" \
|| die "could not move the staged jar from $staged into place at $live — the daemon was NOT
started. The built jar is still sitting at $staged; a manual 'mv \"$staged\" \"$live\"'
may recover this once you find out why the move failed."
}
# fleetd #521 — the swap decision, and the step that acts on it.
#
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
# the freshly built one to $JAR_STAGED by then) or on a stale one.
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
# and compares line positions, and a same-line edit moves no line.
#
# Why these are TWO functions, and why the second one exists at all. Extracting only the predicate
# — `should_swap`, which is what #521 asked for — is not enough, and this was measured, not guessed:
# with the main flow calling `if should_swap "$DO_BUILD"; then`, changing THAT to `if false; then`
# still left the whole suite at exit 0 with no failures. Tests that call a predicate directly prove
# the predicate is right; nothing makes the code that does the work consult it. Extraction had moved
# the untested decision one level up rather than removing it.
#
# So the decision and the action live together in swap_if_built, and the main flow has no guard of
# its own to get wrong — it calls one function unconditionally. A test then calls swap_if_built with
# both values of do_build and checks whether the swap actually happened, which fails if the guard is
# removed, inverted, or stops being consulted. should_swap stays a separate predicate because it is
# the decision itself and is worth naming and testing on its own.
#
# What this still does not pin: deleting the swap_if_built call from the main flow altogether. That
# is the ordering test's job — its needle is that call site — and no test in this file can do better,
# because sourcing stops before the main flow ever runs (see the SOURCED guard below).
should_swap() {
local do_build="$1"
[ "$do_build" = 1 ]
}
swap_if_built() {
local do_build="$1"
should_swap "$do_build" || return 0
say "swap"
swap_staged_jar "$JAR_STAGED" "$JAR"
ok "jar in place: $(jar_id)"
}
# `launchctl list <label>` exits 0 iff the label is loaded (registered with launchd) — true whether
# or not it is currently running, which is exactly "supervision is active" for our purposes. Read-
# only: neither helper below changes anything, so both are also safe under --check.
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
# fleetd #492: same two questions for systemd --user. Kept as separate, overridable functions
# (never an inline `systemctl` call at each use site) so a test on a box with no systemd at all
# (this repo is developed on macOS) can substitute each one independently — the same seam
# launchd_installed/launchd_loaded above already use.
#
# fleetd #492 follow-up: both functions used to throw `systemctl`'s stderr straight into
# /dev/null, which meant "systemctl answered no" and "systemctl could not answer at all" (e.g. it
# cannot reach the user bus over a non-lingering ssh session) looked identical — both a plain
# nonzero exit. They now capture stderr separately and set their own *_ERRORED flag ONLY when the
# call exited non-zero AND wrote something to stderr — a real tool failure, never a clean "not
# installed"/"not active" answer (which exits non-zero with empty stderr). detect_supervisor reads
# the flag right after calling the probe, so a probe that could not answer routes to "unclear",
# never silently becomes "none".
#
# "installed": a unit FILE by this name exists, regardless of its current state — the systemd
# analogue of the plist file existing on disk. `list-unit-files` reads unit definitions without
# depending on runtime state, so this stays read-only and safe under --check.
systemd_installed() {
SYSTEMD_INSTALLED_ERRORED=0
command -v systemctl >/dev/null 2>&1 || return 1
local err_file out rc=0
if ! err_file="$(mktemp -t systemd-installed-err)"; then
SYSTEMD_INSTALLED_ERRORED=1
return 1
fi
out="$(systemctl --user list-unit-files "$SYSTEMD_UNIT.service" --no-legend 2>"$err_file")" || rc=$?
if [ "$rc" -ne 0 ]; then
if [ -s "$err_file" ]; then
SYSTEMD_INSTALLED_ERRORED=1
fi
rm -f "$err_file"
return "$rc"
fi
rm -f "$err_file"
printf '%s' "$out" | grep -q .
}
# "loaded": systemd currently supervises this unit as an active job — the systemd analogue of
# `launchctl list <label>` succeeding. Measured on the second host: `systemctl --user is-active
# fleetd` -> "active". A clean "no" (inactive/failed/activating/deactivating) exits non-zero with
# nothing on stderr; a probe that could not reach systemd at all exits non-zero WITH a stderr
# message — see the fleetd #492 follow-up note above.
systemd_loaded() {
SYSTEMD_LOADED_ERRORED=0
command -v systemctl >/dev/null 2>&1 || return 1
local err_file rc=0
if ! err_file="$(mktemp -t systemd-loaded-err)"; then
SYSTEMD_LOADED_ERRORED=1
return 1
fi
systemctl --user is-active "$SYSTEMD_UNIT" >/dev/null 2>"$err_file" || rc=$?
if [ "$rc" -ne 0 ] && [ -s "$err_file" ]; then
SYSTEMD_LOADED_ERRORED=1
fi
rm -f "$err_file"
return "$rc"
}
# fleetd #504: the "loaded but not currently running" branches in the main stop step (case
# launchd/systemd, reached when $OLD_PID is empty) used to run `launchctl unload`/`systemctl --user
# stop` with `2>/dev/null || true` and then print `ok` unconditionally — the exact conflation
# systemd_loaded/systemd_installed above already fixed on the READ side (fleetd #492 follow-up): a
# genuine "already stopped" answer (nonzero exit, nothing on stderr) is harmless, but a real tool
# failure (nonzero exit WITH a stderr message — e.g. launchd or the systemd user bus is
# unreachable) is not, and reporting `ok` on THAT means the start step below can register a fresh
# load on top of a supervisor that never actually let go: the exact two-daemons failure fleetd #492
# exists to prevent, reached from the one state (already odd) where a false `ok` is least
# affordable. These two functions apply the same "capture stderr separately, flag only a nonzero
# exit WITH stderr as a real failure" pattern to the WRITE side. No ${VAR:-default} anywhere here —
# see the #492 follow-up constraints comment above detect_supervisor for why a default would hide a
# lost value instead of surfacing it (fleetd #497's defect class).
unload_launchd_if_loaded() {
local err_file rc=0
if ! err_file="$(mktemp -t launchd-unload-err)"; then
die "could not create a temp file to capture 'launchctl unload' stderr — cannot tell a real
failure from a clean already-unloaded answer, so refusing to guess. The daemon's
supervision state was NOT touched."
fi
launchctl unload -w "$LAUNCHD_PLIST" >/dev/null 2>"$err_file" || rc=$?
if [ "$rc" -ne 0 ] && [ -s "$err_file" ]; then
die "'launchctl unload -w $LAUNCHD_PLIST' failed: $(cat "$err_file")
The daemon may still be under supervision; investigate before retrying."
fi
rm -f "$err_file"
}
stop_systemd_if_loaded() {
local err_file rc=0
if ! err_file="$(mktemp -t systemd-stop-err)"; then
die "could not create a temp file to capture 'systemctl --user stop' stderr — cannot tell a
real failure from a clean already-stopped answer, so refusing to guess. The daemon's
supervision state was NOT touched."
fi
systemctl --user stop "$SYSTEMD_UNIT" >/dev/null 2>"$err_file" || rc=$?
if [ "$rc" -ne 0 ] && [ -s "$err_file" ]; then
die "'systemctl --user stop $SYSTEMD_UNIT' failed: $(cat "$err_file")
The daemon may still be under supervision; investigate before retrying."
fi
rm -f "$err_file"
}
# fleetd #492: three real answers, not two — launchd, systemd, or genuinely unsupervised — plus a
# fourth, "ambiguous", for the one case this script cannot tell apart: both signals firing at once.
# That is exactly "I cannot tell who supervises this process", and guessing wrong here is how two
# daemons end up running against one herdr session (see trap 7 in the header).
#
# fleetd #492 follow-up: a fifth answer, "unclear", for two more situations that must NEVER be read
# as "none" (measured — see the report this ticket is a follow-up to):
# - installed-but-not-loaded, on EITHER supervisor. `systemctl --user is-active` answers "no" for
# `activating`, `deactivating`, `failed`, and while an auto-restart is pending — every one of
# those is a host that IS under systemd (or launchd) and whose supervisor is about to act again.
# `*_installed` already knows the unit/agent exists; this is the first place that fact is
# actually consulted in the decision, not just printed as a warning.
# - a probe that could not answer at all. systemd_loaded/systemd_installed set their own
# *_ERRORED flag (see the comment above them) when `systemctl` exits non-zero WITH a stderr
# message — a real tool failure, e.g. it cannot reach the user bus over a non-lingering ssh
# session — never conflated with a clean negative answer.
# "none" now means only: neither supervisor is installed, neither is loaded, and neither probe
# errored.
#
# fleetd #492 follow-up — constraints every caller of this function depends on (learned the hard
# way: an earlier version of this fix set a SUPERVISOR_UNCLEAR_DETAIL global from inside here and
# it was silently lost, because every real call site invokes this as `$(detect_supervisor)`):
# 1. It is called as `$(detect_supervisor)`, so ONLY STDOUT crosses back to the caller. Anything
# this function needs to tell its caller — the "unclear" detail included — must be printed,
# never assigned to a global: a global set inside a `$( )` subshell dies with that subshell.
# This function packs BOTH values (kind and detail) onto that one stdout line, joined by
# $SUPERVISOR_DETAIL_SEP, and the caller unpacks them on its own side of the subshell boundary.
# 2. This script runs under `set -u` (part of the `set -euo pipefail` at the top of the file), so
# an unset variable is a loud failure. Do not add a `${VAR:-default}` anywhere downstream to
# paper over a value that should always be there — that hides a lost value instead of
# surfacing it (fleetd #497's defect class). The mechanism is named rather than cited by line
# number on purpose: a line number in a comment goes stale on the next insert above it, and
# this one already had — it said line 50 while the `set` line was at 54.
# 3. Every `case` on this function's return value needs an explicit final `*)` arm, chosen by
# whether that caller ACTS on the value (`die` — an unrecognised value must never be silently
# driven) or only DISPLAYS it (`echo`/`warn` and continue — a diagnostic must not go silent on
# exactly the value it most needs to report).
#
# Pure and side-effect-free besides the two *_ERRORED flags (read back within this same call, never
# by the caller — see the constraints above): reads the four probes and decides — never mutates
# anything, so it is safe under --check and testable by overriding
# launchd_installed/launchd_loaded/systemd_installed/systemd_loaded after sourcing.
detect_supervisor() {
local ld=0 sd=0 li=0 si=0 kind detail=""
SYSTEMD_LOADED_ERRORED=0
SYSTEMD_INSTALLED_ERRORED=0
launchd_loaded && ld=1
systemd_loaded && sd=1
launchd_installed && li=1
systemd_installed && si=1
if [ "$SYSTEMD_LOADED_ERRORED" = 1 ] || [ "$SYSTEMD_INSTALLED_ERRORED" = 1 ]; then
detail="the systemd --user probe for '$SYSTEMD_UNIT' could not answer cleanly (systemctl exited non-zero and reported an error on stderr, not a clean negative — e.g. it cannot reach the user bus)"
kind="unclear"
elif [ "$ld" = 1 ] && [ "$sd" = 1 ]; then
kind="ambiguous"
elif [ "$li" = 1 ] && [ "$ld" = 0 ]; then
detail="the launchd agent ($LAUNCHD_LABEL) is installed ($LAUNCHD_PLIST exists) but is not currently loaded"
kind="unclear"
elif [ "$si" = 1 ] && [ "$sd" = 0 ]; then
detail="the systemd --user unit ($SYSTEMD_UNIT) is installed but not currently active — it may be activating, deactivating, failed, or waiting on an auto-restart"
kind="unclear"
elif [ "$ld" = 1 ]; then
kind="launchd"
elif [ "$sd" = 1 ]; then
kind="systemd"
else
kind="none"
fi
printf '%s%s%s' "$kind" "$SUPERVISOR_DETAIL_SEP" "$detail"
}
# fleetd #492: turns anything detect_supervisor returns that is NOT exactly one of the two
# supervisors this script knows how to drive into a die() — never a fall-through to the `kill`
# path. Kept as its own function so a test can call it directly (in a subshell, since it die()s)
# without running the whole report-state flow or needing a real launchd/systemd.
require_drivable_supervisor() {
local kind="$1"
case "$kind" in
launchd|systemd|none) ;;
ambiguous)
die "both launchd ($LAUNCHD_LABEL) and systemd --user ($SYSTEMD_UNIT) report themselves as
loaded for this daemon at the same time. This script cannot tell which one actually
supervises the running process, and driving either alone risks the OTHER reviving the
OLD jar out from under it — the exact failure this ticket (fleetd #492) exists to
prevent. Stop one of the two supervisors by hand, confirm only one remains loaded, then
rerun." ;;
unclear)
# fleetd #492 follow-up: SUPERVISOR_UNCLEAR_DETAIL crosses back from detect_supervisor's
# subshell via its stdout, unpacked by the caller BEFORE it calls this function (see the
# constraints comment above detect_supervisor). No ${VAR:-default} here on purpose: if the
# detail is somehow missing, `set -u` makes this reference fail loudly instead of silently
# naming nothing — a default that hides a lost value is the same defect class as fleetd
# #497.
die "a supervisor looks present but this script cannot tell whether it actually drives this
daemon: $SUPERVISOR_UNCLEAR_DETAIL. Guessing wrong here is the same failure 'ambiguous'
above exists to prevent: driving the daemon while an unseen supervisor revives the OLD
jar out from under it (fleetd #492). Check 'launchctl list $LAUNCHD_LABEL' and
'systemctl --user status $SYSTEMD_UNIT' by hand, resolve whichever looks unclear, then
rerun." ;;
*)
die "detect_supervisor returned an unrecognized value '$kind' — refusing to guess which
supervisor, if any, controls this daemon." ;;
esac
}
# fleetd #492: the exact symptom a racing supervisor produces — count how many fleetd processes are
# alive right now. Takes the pid list as a parameter (rather than calling running_pid() itself) so a
# test can pass a canned two-line string without a real second process running. Pure except for the
# die() in assert_single_daemon below.
count_daemon_pids() {
local pids="$1"
if [ -z "$pids" ]; then
echo 0
else
printf '%s\n' "$pids" | grep -c .
fi
}
assert_single_daemon() {
local pids="$1" count
count="$(count_daemon_pids "$pids")"
if [ "$count" -gt 1 ]; then
die "more than one fleetd process is running after this restart (pids: $(printf '%s' "$pids" | tr '\n' ' ')).
This is the exact failure a racing supervisor produces: the OLD jar was revived by its
supervisor while this script started a NEW copy. Two daemons on one herdr session kill
each other's members. Investigate with 'pgrep -f \"$PATTERN\"' and stop the wrong one by
hand — do not assume either pid is the one you want."
fi
}
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
@@ -160,6 +543,169 @@ classify_amqp_connection_errors() {
REDEPLOY_UNEXPLAINED_ERRORS=$((REDEPLOY_UNEXPLAINED_ERRORS + pending_inbox + pending_lead_mailbox))
}
# fleetd #512 part 2 — the negative check. #493's failure (an uncaught exception in a shutdown
# thread) never passes through the logger: the JVM's default uncaught-exception handler prints
# straight to stderr, so the line never carries a level, so classify_amqp_connection_errors's
# ERROR/SEVERE token scan is structurally blind to it — measured on two hosts, including one where
# even a syslog PRIORITY filter is blind to it too (fd 1 and fd 2 collapse to one socket there, so
# every uncaught-exception line lands at priority 6/info). The fix is to grep the shape instead of
# the level: `Exception in thread` at the start of a line (the handler's own banner) or
# `NoClassDefFoundError` anywhere in it (the one real instance seen so far, but not the only shape
# this could take). Kept as its own function, never folded into classify_amqp_connection_errors —
# this is not an AMQP concern, and the two must stay independently readable and independently
# testable.
#
# Sets REDEPLOY_UNCAUGHT_EXCEPTION_COUNT (lines matched) and REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE
# (the first matching line, "" if none) so a caller can report both a count and a concrete quote
# without re-reading the file. Pure: reads $1, sets globals, no side effects.
scan_uncaught_exceptions() {
local log_file="$1" line
REDEPLOY_UNCAUGHT_EXCEPTION_COUNT=0
REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE=""
while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
'Exception in thread'*|*NoClassDefFoundError*)
REDEPLOY_UNCAUGHT_EXCEPTION_COUNT=$((REDEPLOY_UNCAUGHT_EXCEPTION_COUNT + 1))
[ -n "$REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE" ] || REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE="$line"
;;
esac
done < "$log_file"
}
# fleetd #512 part 2 — the positive check. fleetd #522 added a `log.info` at the very end of
# SessionManager.drainAll's normal path (never in a `finally` — see the ticket discussion for why
# that distinction matters): "drain complete: released=N abandoned=M (still BUSY at the shutdown
# deadline)", printed once, on every successful drain, including the all-zero case. A drain that
# dies partway through never reaches that statement, so the line's ABSENCE is a real signal — unlike
# the ERROR-count check above, this one does not depend on the failure happening to throw.
#
# Sets REDEPLOY_DRAIN_COMPLETE_LINE to the matching line (last one, though drainAll runs at most
# once per shutdown so there should never be more than one) or "" if absent. Pure, same shape as
# scan_uncaught_exceptions above.
find_drain_complete_line() {
local log_file="$1"
REDEPLOY_DRAIN_COMPLETE_LINE="$(grep -F 'drain complete: released=' "$log_file" | tail -1 || true)"
}
# fleetd #512 part 2 — THE TRAP, and the reason this is one function instead of two independent
# checks the caller ORs together. Absence of the drain-complete line has TWO causes that need
# OPPOSITE handling, and a naive "line absent -> the drain died" reading collapses them exactly the
# way this whole ticket exists to stop: the line is emitted by the daemon being STOPPED, which is
# running the OLD jar. Until a redeploy has landed fleetd #522 once, every previous daemon predates
# the line and cannot emit it no matter how cleanly it drained — so on the very first redeploy after
# #522 merged, "absent" means "too old to know how", not "died". Only once BOTH signals — this
# line's absence AND scan_uncaught_exceptions' result — have been read together can the three real
# outcomes be told apart:
#
# complete -> the line is present: the drain finished. Name the counts it reported.
# died -> the line is absent AND an uncaught-exception shape was found: the drain died. Name
# what was found.
# unknown -> the line is absent AND no exception shape either: cannot tell. Say so, and say why
# (predates the line, or failed without throwing) — never worded as a pass or a
# failure, and never reassuring: "ok, no ERROR lines" one level up is the exact mistake
# this ticket exists to fix, and this outcome must not reproduce it.
#
# A fourth case, n/a, covers a cold start or a "loaded but wasn't running" restart: no previous
# daemon was actually stopped THIS run, so there is no shutdown window in $log_file to have an
# opinion about at all — scanning it anyway would read the NEW daemon's own startup lines and could
# misreport "cannot tell" on every clean cold start. had_previous_daemon carries that fact in from
# the caller (it already knows $OLD_PID) rather than this function re-deriving it from log content.
#
# Same shape as swap_if_built/refuse_drain_gate (fleetd #521/#528): the decision (which of the four
# outcomes applies) and the action (which ok/warn line to print, and setting REDEPLOY_DRAIN_STATE
# for the "result" section below to consult) live together in ONE function that the main flow calls
# unconditionally — there is no guard left in the main flow to remove, invert, or bypass
# independently of this function. Never calls die(): #512's own decision is to warn loudly and let
# the redeploy stand, because by the time this is detectable the new daemon is already up and
# healthy and failing here would give the operator nothing to do differently.
report_shutdown_drain() {
local log_file="$1" had_previous_daemon="$2"
REDEPLOY_UNCAUGHT_EXCEPTION_COUNT=0
REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE=""
REDEPLOY_DRAIN_COMPLETE_LINE=""
if [ "$had_previous_daemon" != 1 ]; then
REDEPLOY_DRAIN_STATE="n/a"
ok "no previous daemon was running before this restart — nothing to check for a died shutdown drain"
return 0
fi
find_drain_complete_line "$log_file"
scan_uncaught_exceptions "$log_file"
if [ -n "$REDEPLOY_DRAIN_COMPLETE_LINE" ]; then
REDEPLOY_DRAIN_STATE="complete"
ok "previous daemon's shutdown drain finished: $REDEPLOY_DRAIN_COMPLETE_LINE"
elif [ "$REDEPLOY_UNCAUGHT_EXCEPTION_COUNT" -gt 0 ]; then
REDEPLOY_DRAIN_STATE="died"
warn "previous daemon's shutdown drain DIED — no drain-complete line, and an uncaught exception"
warn "was found in its shutdown window ($REDEPLOY_UNCAUGHT_EXCEPTION_COUNT line(s)):"
warn " $REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE"
warn "Some sessions from the PREVIOUS daemon may not have been released."
else
REDEPLOY_DRAIN_STATE="unknown"
warn "cannot tell whether the previous daemon's shutdown drain finished — no drain-complete line"
warn "and no uncaught-exception shape either. This is NOT a pass and NOT a failure: it means"
warn "either that daemon predates fleetd #522's drain-complete log line, or its drain failed"
warn "without throwing (hung, or returned early)."
fi
}
# fleetd #517: extracted so the suite can call this decision directly, the same way #510 extracted
# wait_for_daemon_exit so its ordering became checkable. Before this, the only test of the drain-gate
# abort message was a grep of this script's own source for the wording — so mutating the `if` below
# to `if false` (making the branch unreachable) left every test green, because the wording was still
# sitting in the file. Pure: only decides which message applies and prints it, no side effects, so a
# test can call it directly with an in-memory staged path instead of driving the real drain-gate flow
# (which needs a live $OLD_PID and an interactive prompt neither test can supply).
#
# The four cases:
# build ran, staged jar present -> names the staged jar and how to finish or discard it
# build ran, staged jar absent -> "nothing changed" (nothing was staged this run either)
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
# there is nothing here for the operator to lose track of.
# --no-build, staged jar absent -> "nothing changed"
drain_gate_refusal() {
local do_build="$1" staged_path="$2"
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
the freshly built jar is no longer at the live path that --no-build requires — or
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
else
printf 'aborted — nothing changed'
fi
}
# fleetd #528 — drain_gate_refusal above is well tested (four cases, all direct), but nothing made
# the MAIN FLOW's abort actually consult it. Before this, the main flow read
# `die "$(drain_gate_refusal "$DO_BUILD" "$JAR_STAGED")"` directly, and mutating that one line to a
# flat `die "aborted — nothing changed"` left the whole suite at exit 0 with zero FAIL lines and
# byte-identical output to a clean run — every one of drain_gate_refusal's own tests still passed,
# because they call the predicate directly and never touch this call site. That silently reinstated
# the exact defect #517 was filed to fix. Same shape as #521/#526's should_swap/swap_if_built: a
# predicate alone is not enough, because a test proving the predicate is right cannot also prove the
# main flow consults it. So the decision (drain_gate_refusal) and the action (die) now live together
# in ONE function, and the main flow calls it unconditionally instead of building the die() call
# itself — there is no guard left in the main flow to remove, invert, or bypass independently of this
# function. drain_gate_refusal stays separate and separately tested because the message-selection
# logic is worth naming and testing on its own; refuse_drain_gate is the only thing that ever dies.
#
# What the behavioural tests above still cannot pin on their own: deleting the call to this function
# from the main flow altogether — they call refuse_drain_gate directly, never through the main flow,
# because sourcing stops before the main flow ever runs (see the SOURCED guard below). That gap is
# closed the same way swap_if_built's is: test_refuse_drain_gate_call_site_present greps this script
# for the real invocation, the same shape test_swap_ordered_after_wait_and_before_start already uses
# for the swap call. Deliberately NOT written out here as a literal quoted string, so this comment
# itself can never become a second match for that test's needle.
refuse_drain_gate() {
local do_build="$1" staged_path="$2"
die "$(drain_gate_refusal "$do_build" "$staged_path")"
}
# CB-600: sourceable for testing. When this file is SOURCED (not executed) it stops here — nothing
# below runs — so a test harness can `source` it to call check_log_path_matches_plist (or the
# other pure helpers above) against a throwaway plist fixture without ever reaching the mutating
@@ -182,25 +728,57 @@ fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594: supervision state. Installed and loaded are different facts — a copied-but-never-loaded
# plist supervises nothing, and a loaded label with no file backing it (rare, but possible after an
# edited/moved plist) is still what launchd will act on.
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
# copied-but-never-loaded plist (or an unloaded systemd unit) supervises nothing, and a loaded
# label/unit with no file backing it is still what its supervisor will act on.
if launchd_installed; then
ok "launchd agent installed: $LAUNCHD_PLIST"
else
warn "launchd agent NOT installed (no supervision — a crash will not restart the daemon)."
warn "launchd agent NOT installed."
fi
SUPERVISED=0
if launchd_loaded; then
SUPERVISED=1
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
# would read different log files — every check after this point is worthless otherwise.
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
if systemd_installed; then
ok "systemd --user unit installed: $SYSTEMD_UNIT"
else
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
warn "systemd --user unit NOT installed ($SYSTEMD_UNIT)."
fi
# fleetd #492: decide which of the two (if either) actually supervises this daemon, and refuse
# outright — before touching anything — if that cannot be told apart (see require_drivable_
# supervisor above). --check reaches this same line, so a host with an undrivable supervisor is
# reported as a failure even in --check, without ever reaching the build/stop/start steps.
# fleetd #492 follow-up: detect_supervisor runs as $(...), so only the printed line survives —
# unpack kind and detail from it HERE, in this shell, before calling anything downstream. See the
# constraints comment above detect_supervisor for why this cannot be done any other way.
SUPERVISOR_RAW="$(detect_supervisor)"
SUPERVISOR_KIND="${SUPERVISOR_RAW%%"$SUPERVISOR_DETAIL_SEP"*}"
SUPERVISOR_UNCLEAR_DETAIL="${SUPERVISOR_RAW#*"$SUPERVISOR_DETAIL_SEP"}"
require_drivable_supervisor "$SUPERVISOR_KIND"
ok "supervisor detected: $SUPERVISOR_KIND"
SUPERVISED=0
case "$SUPERVISOR_KIND" in
launchd)
SUPERVISED=1
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
# would read different log files — every check after this point is worthless otherwise.
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
;;
systemd)
SUPERVISED=1
ok "systemd --user unit active ($SYSTEMD_UNIT) — systemd supervises this daemon"
;;
none)
warn "no supervisor loaded — this script is the only thing that will restart the daemon."
;;
*)
# fleetd #492 follow-up: this block only DISPLAYS state, it changes nothing yet — so a value
# it doesn't recognise gets reported, not an abort that goes silent on exactly the state most
# worth seeing. (Unreachable today: require_drivable_supervisor above already died on
# "ambiguous"/"unclear" before this case runs. Guards the value nobody has invented yet.)
warn "unrecognised supervisor kind: '$SUPERVISOR_KIND' — detect_supervisor returned a value this block does not know; continuing to report the rest of the state."
;;
esac
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
# below. Never prints the value — only whether it resolved.
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
@@ -249,6 +827,9 @@ fi
if [ "$DO_BUILD" = 1 ]; then
say "build"
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
# anything else, so that run's leftovers can never be mistaken for this run's output.
rm -f "$JAR_STAGED"
BUILD_LOG="$(mktemp -t fleetd-build)"
echo " log: $BUILD_LOG"
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
@@ -258,13 +839,19 @@ if [ "$DO_BUILD" = 1 ]; then
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
ok "jar now: $(jar_id)"
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
# daemon, if any, is still up at this point (build always runs before stop). From here until the
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
stage_built_jar
ok "jar now: $(jar_id "$JAR_STAGED")"
else
say "build skipped (--no-build)"
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
# sitting at the live path from an earlier successful run. Same check, same message as before.
require_no_build_jar
fi
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
# ----------------------------------------------------------------- drain gate
if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
@@ -276,82 +863,148 @@ if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
echo " you still want, BEFORE continuing."
echo
read -r -p " Fleet drained? type yes to restart: " reply
[ "$reply" = "yes" ] || die "aborted — nothing changed"
if [ "$reply" != "yes" ]; then
# fleetd #493 / #517 / #528: "nothing changed" would be a lie once a build has run and staged a
# jar — see drain_gate_refusal above for the full decision and why each of its four cases reads
# the way it does. refuse_drain_gate composes that message AND calls die itself, so this guard
# has nothing left of its own to get wrong beyond whether it calls refuse_drain_gate at all.
refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"
fi
fi
# ------------------------------------------------------------------ stop
#
# CB-594: when SUPERVISED, launchd owns the stop — never a raw `kill` here. A bare SIGTERM makes
# this JVM exit 143 even with its shutdown hook running to completion (verified separately: a
# throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login shell that
# could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
# KeepAlive.SuccessfulExit=false treats any nonzero exit as a crash and restarts the OLD jar,
# which would race this script's own restart of the NEW one. `launchctl unload` avoids that race
# by deregistering the job first, so no KeepAlive is left armed when the process actually stops.
# CB-594 / fleetd #492: when SUPERVISED, the supervisor owns the stop — never a raw `kill` here. A
# bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to completion (verified
# separately: a throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login
# shell that could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
# KeepAlive.SuccessfulExit=false and systemd's Restart=on-failure both treat any nonzero exit as a
# crash and restart the OLD jar, which would race this script's own restart of the NEW one.
# `launchctl unload` avoids that race by deregistering the job first, so no KeepAlive is left
# armed when the process actually stops. `systemctl --user stop` needs no such dance: unlike
# KeepAlive, systemd's Restart= does not fire on a deliberate stop, only on an unexpected exit of
# an active unit.
if [ -n "$OLD_PID" ]; then
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
if [ "$SUPERVISED" = 1 ]; then
echo " supervision is ON: using 'launchctl unload' (not kill) so launchd's own KeepAlive"
echo " cannot restart the OLD jar out from under this script — see the CB-594 comment above."
launchctl unload -w "$LAUNCHD_PLIST" \
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
else
kill "$OLD_PID"
fi
for _ in $(seq "$STOP_WAIT"); do
[ -z "$(running_pid)" ] && break
sleep 1
done
if [ -n "$(running_pid)" ]; then
case "$SUPERVISOR_KIND" in
launchd)
echo " supervision is ON (launchd): using 'launchctl unload' (not kill) so launchd's own"
echo " KeepAlive cannot restart the OLD jar out from under this script — see the CB-594"
echo " comment above."
launchctl unload -w "$LAUNCHD_PLIST" \
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
;;
systemd)
echo " supervision is ON (systemd --user): using 'systemctl --user stop' (not kill) so"
echo " systemd's own Restart=on-failure cannot restart the OLD jar out from under this"
echo " script — see the fleetd #492 comment above."
systemctl --user stop "$SYSTEMD_UNIT" \
|| die "'systemctl --user stop $SYSTEMD_UNIT' failed — the daemon may still be under supervision; investigate before retrying"
;;
none)
kill "$OLD_PID"
;;
*)
# fleetd #492 follow-up: this block ACTS (stops the daemon one specific way per kind) — an
# unrecognised value must never fall through to a default action, silently picking the wrong
# one (or none at all) while reporting success. (Unreachable today: require_drivable_
# supervisor already died before this runs. Guards the value nobody has invented yet.)
die "detect_supervisor returned an unrecognized value '$SUPERVISOR_KIND' at the stop step —
refusing to guess how to stop a daemon under an unknown supervisor. The daemon was NOT
stopped." ;;
esac
if ! wait_for_daemon_exit "$STOP_WAIT"; then
die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically:
the shutdown hook releases sessions and worktrees in order, and killing it hard can
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
fi
ok "pid $OLD_PID exited"
elif [ "$SUPERVISED" = 1 ]; then
elif [ "$SUPERVISOR_KIND" = "launchd" ]; then
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
# start step below does a clean load, never a load stacked on an already-loaded label.
# fleetd #504: unload_launchd_if_loaded (above) tolerates a genuine already-unloaded answer but
# dies on a real `launchctl` failure — never a bare `|| true` that would print `ok` either way.
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true
unload_launchd_if_loaded
ok "launchd agent unloaded (was already not running)"
elif [ "$SUPERVISOR_KIND" = "systemd" ]; then
# Same case for systemd: the unit is known/active-capable but not currently running. `stop` on an
# already-stopped unit is a harmless no-op — kept for symmetry with the launchd branch above.
# fleetd #504: stop_systemd_if_loaded (above) tolerates that genuine no-op but dies on a real
# `systemctl` failure — never a bare `|| true` that would print `ok` either way.
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
stop_systemd_if_loaded
ok "systemd --user unit stopped (was already not running)"
else
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
fi
# ------------------------------------------------------------------ swap
#
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
swap_if_built "$DO_BUILD"
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
# Supervised: launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins fleetd/.
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# above) and WorkingDirectory is already pinned to fleetd/.
say "start"
if [ "$SUPERVISED" = 1 ]; then
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
echo " process, instead of a manual nohup that launchd would know nothing about."
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
# than before this script ran, because a later reboot or login will not bring it back either. One
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
# die with the exact recovery command rather than a bare "failed".
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
warn "launchctl load failed on the first attempt — retrying once after a short pause"
sleep 2
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
never got the chance to clear it. Recover with:
launchctl load -w \"$LAUNCHD_PLIST\"
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
fi
else
# Absolute jar path so `ps` names which checkout is running.
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
fi
case "$SUPERVISOR_KIND" in
launchd)
echo " supervision is ON (launchd): using 'launchctl load' so launchd starts and keeps"
echo " supervising this process, instead of a manual nohup that launchd would know nothing"
echo " about."
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
# than before this script ran, because a later reboot or login will not bring it back either. One
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
# die with the exact recovery command rather than a bare "failed".
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
warn "launchctl load failed on the first attempt — retrying once after a short pause"
sleep 2
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
never got the chance to clear it. Recover with:
launchctl load -w \"$LAUNCHD_PLIST\"
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
fi
;;
systemd)
echo " supervision is ON (systemd --user): using 'systemctl --user start' so systemd starts"
echo " and keeps supervising this process, instead of a manual nohup it would know nothing"
echo " about."
systemctl --user start "$SYSTEMD_UNIT" || die "'systemctl --user start $SYSTEMD_UNIT' failed.
Check 'systemctl --user status $SYSTEMD_UNIT' and $OUT before assuming a retry will succeed."
;;
none)
# Absolute jar path so `ps` names which checkout is running.
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
;;
*)
# fleetd #492 follow-up: this block ACTS (starts the daemon one specific way per kind) — an
# unrecognised value must never fall through to a default action, silently picking the wrong
# one (or none at all) while reporting success. (Unreachable today: require_drivable_
# supervisor already died before this runs. Guards the value nobody has invented yet.)
die "detect_supervisor returned an unrecognized value '$SUPERVISOR_KIND' at the start step —
refusing to guess how to start a daemon under an unknown supervisor. The daemon was NOT
started." ;;
esac
for _ in $(seq 10); do
NEW_PID="$(running_pid)"
@@ -411,10 +1064,32 @@ FRESH_LOG="$(mktemp -t fleetd-fresh-log)"
trap 'rm -f "$FRESH_LOG"' EXIT
tail -n "+$((RESTART_MARK + 1))" "$OUT" > "$FRESH_LOG" 2>/dev/null || true
classify_amqp_connection_errors "$FRESH_LOG"
# fleetd #512 part 2: the previous daemon's shutdown drain, checked in the same fresh-log region —
# see report_shutdown_drain above for the full decision (four outcomes, one of them a deliberate
# "cannot tell"). HAD_OLD_PID crosses in whether a previous daemon was actually stopped this run;
# see the function's own comment for why that matters.
say "previous daemon's shutdown drain"
HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1
report_shutdown_drain "$FRESH_LOG" "$HAD_OLD_PID"
# fleetd #492: checked here, after healthz and the fresh-log check have both had time to run, so a
# supervisor that revives the OLD jar a few seconds late is caught too. Every check above (healthz
# 200, jar id, the fresh 'listening' line) is satisfied by EITHER daemon if two are alive — this is
# the only one that can tell.
assert_single_daemon "$(running_pid)"
say "result"
ok "pid $NEW_PID, jar $(jar_id)"
if [ "$REDEPLOY_ERROR_COUNT" -eq 0 ]; then
ok "no ERROR lines since restart"
# fleetd #512 item 4: this line must not print when the shutdown-drain check above found the
# previous daemon's drain died, or could not tell — either would make "no ERROR lines" read as a
# clean bill of health it is not (an uncaught exception never carries an ERROR token to begin
# with, so this count alone cannot see that failure). "complete" and "n/a" are the only two
# outcomes report_shutdown_drain sets that mean nothing is wrong there.
if [ "$REDEPLOY_DRAIN_STATE" = "complete" ] || [ "$REDEPLOY_DRAIN_STATE" = "n/a" ]; then
ok "no ERROR lines since restart"
fi
elif [ "$REDEPLOY_UNEXPLAINED_ERRORS" -eq 0 ]; then
ok "$REDEPLOY_RECOVERED_AMQP_ERRORS AMQP connection reset ERROR lines recovered since restart"
else
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Self-contained checks for the policy parsing guards in probe-member-credentials.sh.
set -euo pipefail
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
PROBE="$ROOT/scripts/probe-member-credentials.sh"
TMP="$(mktemp -d "$ROOT/.probe-member-credentials-test.XXXXXX")"
trap 'rm -rf "$TMP"' EXIT
# The SOURCED guard exposes this pure parser without contacting POLICY_URL.
source "$PROBE"
fail() {
printf 'FAIL: %s\n' "$*" >&2
return 1
}
assert_equals() {
local expected="$1" actual="$2" description="$3"
[ "$expected" = "$actual" ] || fail "$description: expected $expected, got $actual"
}
assert_contains() {
local needle="$1" text="$2" description="$3"
printf '%s' "$text" | grep -qF "$needle" || fail "$description: missing $needle"
}
make_jq() {
local body="$1"
mkdir -p "$TMP/bin"
printf '%s\n' '#!/usr/bin/env bash' "$body" > "$TMP/bin/jq"
chmod +x "$TMP/bin/jq"
}
run_parser() {
local output rc=0
POLICY_JSON="$(< "$TMP/policy.json")"
POLICY_URL="fixture://member-credentials"
output="$(PATH="$TMP/bin:$PATH" parse_policy_fields 2>&1)" || rc=$?
PARSER_OUTPUT="$output"
PARSER_RC="$rc"
}
test_bash_older_than_four_refuses() {
local output rc=0 version
version="$(/bin/bash -c 'printf %s "$BASH_VERSION"')"
output="$(/bin/bash "$PROBE" 2>&1)" || rc=$?
assert_equals 3 "$rc" "bash 3 refusal status"
assert_contains 'This shell is bash' "$output" "bash 3 refusal"
assert_contains "$version" "$output" "bash 3 refusal version"
}
test_parser_non_zero_refuses() {
make_jq 'exit 17'
run_parser
assert_equals 4 "$PARSER_RC" "parser failure status"
assert_contains 'jq exited non-zero (status 17)' "$PARSER_OUTPUT" "parser failure message"
}
test_short_parser_output_refuses() {
make_jq "printf '%s\\n' true enforce 3 2"
run_parser
assert_equals 5 "$PARSER_RC" "short parser output status"
assert_contains 'policy parser (jq) returned 4 field(s)' "$PARSER_OUTPUT" "short parser output count"
}
test_empty_parser_output_reports_zero_fields() {
make_jq ':'
run_parser
assert_equals 5 "$PARSER_RC" "empty parser output status"
assert_contains 'policy parser (jq) returned 0 field(s)' "$PARSER_OUTPUT" "empty parser output count"
}
test_well_formed_policy_prints_name_table() {
local output rc=0
make_jq "cat '$TMP/policy.fields'"
# Shell functions cannot be passed in an environment assignment. Run the executable through bash.
output="$(BRIDGED_MEMBER=1 FIXTURE="$TMP/policy.json" PROBE="$PROBE" PATH="$TMP/bin:$PATH" bash -c '
curl() { cat "$FIXTURE"; }
export -f curl
exec "$PROBE"
' 2>&1)" || rc=$?
assert_equals 0 "$rc" "well-formed policy status"
assert_contains 'ALPHA_TOKEN' "$output" "name table"
assert_contains 'BETA_TOKEN' "$output" "name table"
assert_contains 'GAMMA_TOKEN' "$output" "name table"
}
cat > "$TMP/policy.json" <<'JSON'
{"present":true,"policy":"enforce","knownCount":3,"allowedCount":2,"blockedCount":1,"known":["ALPHA_TOKEN","BETA_TOKEN","GAMMA_TOKEN"]}
JSON
cat > "$TMP/policy.fields" <<'FIELDS'
true
enforce
3
2
1
ALPHA_TOKEN
BETA_TOKEN
GAMMA_TOKEN
FIELDS
test_bash_older_than_four_refuses
test_parser_non_zero_refuses
test_short_parser_output_refuses
test_empty_parser_output_reports_zero_fields
test_well_formed_policy_prints_name_table
printf 'PASS: probe member credentials guards\n'
+912
View File
@@ -25,6 +25,706 @@ classify_fixture() {
classify_amqp_connection_errors "$TMP/$name"
}
# fleetd #492 follow-up: detect_supervisor's stdout is now "kind<SEP>detail" (see the constraints
# comment above detect_supervisor in redeploy-fleetd.sh) — every test below that only cares about
# the kind must split it out with the SAME in-shell parameter expansion the real call site (:438)
# uses, never a bare string comparison against the raw output.
supervisor_kind_of() {
printf '%s' "${1%%"$SUPERVISOR_DETAIL_SEP"*}"
}
supervisor_detail_of() {
printf '%s' "${1#*"$SUPERVISOR_DETAIL_SEP"}"
}
# fleetd #492 — supervisor detection. Detect_supervisor() reads launchd_loaded/systemd_loaded, so
# each test overrides BOTH pairs (installed + loaded) explicitly, rather than relying on either
# being naturally absent: this machine may itself be running a real fleetd under launchd right now
# (see CLAUDE.md/MEMORY.md — launchd supervision has been live here since 2026-08-26), so leaving
# launchd_loaded unmocked in a "systemd only" test would silently read this host's own live state
# instead of the fixture.
test_detect_supervisor_launchd_only() {
launchd_installed() { return 0; }
launchd_loaded() { return 0; }
systemd_installed() { return 1; }
systemd_loaded() { return 1; }
assert_equals "launchd" "$(supervisor_kind_of "$(detect_supervisor)")" "launchd-only detection"
}
test_detect_supervisor_systemd_only() {
launchd_installed() { return 1; }
launchd_loaded() { return 1; }
systemd_installed() { return 0; }
systemd_loaded() { return 0; }
assert_equals "systemd" "$(supervisor_kind_of "$(detect_supervisor)")" "systemd-only detection"
}
test_detect_supervisor_none() {
launchd_installed() { return 1; }
launchd_loaded() { return 1; }
systemd_installed() { return 1; }
systemd_loaded() { return 1; }
assert_equals "none" "$(supervisor_kind_of "$(detect_supervisor)")" "unsupervised detection"
}
# fleetd #492 follow-up — detect_supervisor must never answer "none" when the truth is "could not
# tell". `systemd_installed`/`systemd_loaded` already know a unit file exists; this proves that
# fact is now actually consulted, not just printed as a warning: an installed-but-not-loaded unit
# reads as unclear, because is-active answers "no" for activating/deactivating/failed/pending
# auto-restart too, and every one of those is a host that IS under systemd.
test_detect_supervisor_systemd_installed_not_loaded_is_unclear() {
launchd_installed() { return 1; }
launchd_loaded() { return 1; }
systemd_installed() { return 0; } # the unit file IS there
systemd_loaded() { return 1; } # is-active says no — could be activating/failed/pending restart
local raw
raw="$(detect_supervisor)"
assert_equals "unclear" "$(supervisor_kind_of "$raw")" "systemd installed-but-not-loaded must read as unclear, not none"
printf '%s' "$(supervisor_detail_of "$raw")" | grep -qF "$SYSTEMD_UNIT" \
|| fail "detail does not name the systemd unit it found installed-but-not-loaded"
}
# Same fact, the launchd side: a plist on disk that is not currently loaded (unloaded without being
# removed, or about to be reloaded) must not read as "no supervisor" either.
test_detect_supervisor_launchd_installed_not_loaded_is_unclear() {
launchd_installed() { return 0; } # the plist IS there
launchd_loaded() { return 1; } # launchctl list says not loaded
systemd_installed() { return 1; }
systemd_loaded() { return 1; }
local raw
raw="$(detect_supervisor)"
assert_equals "unclear" "$(supervisor_kind_of "$raw")" "launchd installed-but-not-loaded must read as unclear, not none"
printf '%s' "$(supervisor_detail_of "$raw")" | grep -qF "$LAUNCHD_LABEL" \
|| fail "detail does not name the launchd label it found installed-but-not-loaded"
}
# Drives the REAL systemd_loaded/systemd_installed bodies (never stubbed) through a `systemctl`
# stub placed first on PATH that exits non-zero AND writes to stderr — the shape of a systemctl
# that runs but cannot reach the user bus (measured elsewhere as a headless ssh session with no
# lingering). This must read as unclear, never none: a probe that could not answer at all is not
# the same fact as "no supervisor is loaded".
test_detect_supervisor_systemd_probe_error_is_unclear() {
# Re-source first to restore the REAL launchd_*/systemd_* probe bodies. Earlier tests in this
# file permanently override them with stub `return 0`/`return 1` bodies (that is the whole point
# of those tests), and a bash function definition is global for the rest of the process — without
# this, systemd_loaded here would still be whatever the previous test left it as, never touching
# a real `systemctl` call at all.
source "$ROOT/scripts/redeploy-fleetd.sh"
local bin_dir result rc=0
bin_dir="$TMP/stub-bin-systemctl-errors"
mkdir -p "$bin_dir"
cat > "$bin_dir/systemctl" <<'STUB'
#!/usr/bin/env bash
echo "Failed to connect to bus: No such file or directory" >&2
exit 1
STUB
chmod +x "$bin_dir/systemctl"
PATH="$bin_dir:$PATH" systemd_loaded && rc=0 || rc=$?
[ "$rc" -ne 0 ] || fail "systemd_loaded must not report loaded=true when systemctl only errored"
assert_equals "1" "$SYSTEMD_LOADED_ERRORED" "systemd_loaded must flag a probe error, not a clean negative"
launchd_installed() { return 1; }
launchd_loaded() { return 1; }
result="$(PATH="$bin_dir:$PATH" detect_supervisor)"
assert_equals "unclear" "$(supervisor_kind_of "$result")" "a systemd probe error must read as unclear, not none"
printf '%s' "$(supervisor_detail_of "$result")" | grep -qF "$SYSTEMD_UNIT" \
|| fail "detail does not name the systemd unit whose probe errored"
}
# fleetd #492 follow-up (Item 1): this must go through the REAL call-site shape at :437-440, not a
# hand-constructed "unclear" value — a test that builds "unclear" directly proves the switch, not
# the handoff, and that is exactly the gap that let SUPERVISOR_UNCLEAR_DETAIL never reach the real
# caller in b17f37a. detect_supervisor runs as $(detect_supervisor): a subshell. Only stdout
# survives that boundary, so kind AND detail must both cross on it — this test proves they do.
test_require_drivable_supervisor_refuses_unclear() {
launchd_installed() { return 1; }
launchd_loaded() { return 1; }
systemd_installed() { return 0; }
systemd_loaded() { return 1; }
local SUPERVISOR_RAW SUPERVISOR_KIND SUPERVISOR_UNCLEAR_DETAIL output rc=0
# Exactly what :437-439 does — do not shortcut this by constructing "unclear" by hand.
SUPERVISOR_RAW="$(detect_supervisor)"
SUPERVISOR_KIND="${SUPERVISOR_RAW%%"$SUPERVISOR_DETAIL_SEP"*}"
SUPERVISOR_UNCLEAR_DETAIL="${SUPERVISOR_RAW#*"$SUPERVISOR_DETAIL_SEP"}"
assert_equals "unclear" "$SUPERVISOR_KIND" "setup: expected unclear before testing the refusal"
[ -n "$SUPERVISOR_UNCLEAR_DETAIL" ] \
|| fail "detail did not survive the \$(...) call-site boundary — SUPERVISOR_UNCLEAR_DETAIL is empty in the parent shell"
printf '%s' "$SUPERVISOR_UNCLEAR_DETAIL" | grep -qF "$SYSTEMD_UNIT" \
|| fail "detail that crossed the subshell boundary does not name the systemd unit it found installed-but-not-loaded"
output="$(require_drivable_supervisor "$SUPERVISOR_KIND" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] || fail "require_drivable_supervisor accepted an unclear (undrivable) supervisor"
# Check for the ACTUAL DETAIL TEXT, not just "$SYSTEMD_UNIT" — the die() message's boilerplate
# recovery instructions name the unit unconditionally either way ("systemctl --user status
# $SYSTEMD_UNIT"), so a bare unit-name grep here would pass even on a lost/fallback detail. Only
# the specific detail string proves the crossed value, not the boilerplate, reached the message.
printf '%s' "$output" | grep -qF "$SUPERVISOR_UNCLEAR_DETAIL" \
|| fail "refusal message does not contain the specific detail that crossed the subshell boundary"
}
# The heart of the ticket: a supervisor this script cannot drive must refuse, never fall through to
# `kill`. require_drivable_supervisor die()s, so it is invoked inside a command substitution — that
# forks a subshell, so its exit() only ends the subshell and this test script keeps running under
# `set -e`.
test_require_drivable_supervisor_refuses_ambiguous() {
local output rc=0
output="$(require_drivable_supervisor "ambiguous" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] || fail "require_drivable_supervisor accepted an ambiguous (undrivable) supervisor"
printf '%s' "$output" | grep -qF "$LAUNCHD_LABEL" \
|| fail "refusal message does not name the launchd label it found"
printf '%s' "$output" | grep -qF "$SYSTEMD_UNIT" \
|| fail "refusal message does not name the systemd unit it found"
}
test_require_drivable_supervisor_accepts_known_kinds() {
require_drivable_supervisor "launchd" || fail "refused a drivable launchd supervisor"
require_drivable_supervisor "systemd" || fail "refused a drivable systemd supervisor"
require_drivable_supervisor "none" || fail "refused the unsupervised case"
}
# fleetd #492 — the one-daemon check. Two live pids is the exact symptom a racing supervisor
# produces, and none of the other post-restart checks (healthz, jar id, the fresh log line) can see
# it because either daemon alone satisfies them.
test_count_daemon_pids() {
assert_equals 0 "$(count_daemon_pids "")" "count of an empty pid list"
assert_equals 1 "$(count_daemon_pids "4242")" "count of a single pid"
assert_equals 2 "$(count_daemon_pids "$(printf '4242\n4343\n')")" "count of two pids"
}
test_assert_single_daemon_accepts_one_pid() {
assert_single_daemon "4242" || fail "assert_single_daemon rejected a single running pid"
}
test_assert_single_daemon_rejects_two_pids() {
local output rc=0
output="$(assert_single_daemon "$(printf '4242\n4343\n')" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] || fail "assert_single_daemon accepted two simultaneously running pids"
printf '%s' "$output" | grep -qF '4242' || fail "refusal message does not list the pids it found"
printf '%s' "$output" | grep -qF '4343' || fail "refusal message does not list the pids it found"
}
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
# rather than accidentally matching.
test_jar_id_defaults_to_live_and_reports_explicit_path() {
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
local live_hash staged_hash default_result explicit_result
dir="$TMP/jar-id"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
printf 'live jar bytes' > "$JAR"
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
live_hash="$(shasum -a 256 "$JAR" | cut -c1-12)"
staged_hash="$(shasum -a 256 "$JAR_STAGED" | cut -c1-12)"
default_result="$(jar_id)"
explicit_result="$(jar_id "$JAR_STAGED")"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
}
# fleetd #517 — jar_id()'s "absent" branch was unpinned by any test: the existing test above (#511)
# proves both halves of the present-file contract but never exercises the missing-file path. This
# word matters more than a string usually would: "absent" is the #413 signal that a `mvn clean`
# deleted the running daemon's jar out from under it, and the `redeploy-fleetd` skill points
# operators at `--check` for exactly this. Covers both the no-argument default and an explicit path,
# since the mutation (`absent` -> `present`) sits on the single shared `|| echo` and would flip both.
test_jar_id_reports_absent_for_missing_file() {
local saved_jar="$JAR" dir default_result explicit_result
dir="$TMP/jar-id-absent"; mkdir -p "$dir"
JAR="$dir/does-not-exist.jar"
[ ! -f "$JAR" ] || fail "test fixture error: \$JAR unexpectedly exists at $JAR"
default_result="$(jar_id)"
explicit_result="$(jar_id "$dir/also-does-not-exist.jar")"
JAR="$saved_jar"
assert_equals "absent" "$default_result" "jar_id with no arguments must report absent when \$JAR does not exist"
assert_equals "absent" "$explicit_result" "jar_id with an explicit missing path must report absent"
}
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
# are exercised directly against real files on disk (not stubs), because the whole point is file
# behavior (does the content move, does the source disappear, does a failure leave both sides
# intact) that a stubbed function cannot prove.
test_stage_built_jar_moves_off_live_path() {
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-ok"; mkdir -p "$dir"
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
printf 'built jar bytes' > "$jar"
JAR="$jar"; JAR_STAGED="$staged"
stage_built_jar || fail "stage_built_jar rejected a real build output"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
}
test_stage_built_jar_dies_when_build_produced_nothing() {
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-missing"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
output="$(stage_built_jar 2>&1)" || rc=$?
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|| fail "refusal message does not name the missing jar path"
}
test_swap_staged_jar_moves_staged_onto_live() {
local dir staged live
dir="$TMP/swap-ok"; mkdir -p "$dir"
staged="$dir/fleetd-new.jar"; live="$dir/fleetd.jar"
printf 'swapped jar bytes' > "$staged"
swap_staged_jar "$staged" "$live" || fail "swap_staged_jar rejected a real staged jar"
[ ! -f "$staged" ] || fail "swap_staged_jar left the staged file behind at $staged"
[ -f "$live" ] || fail "swap_staged_jar did not create the live jar at $live"
grep -qF 'swapped jar bytes' "$live" || fail "live jar does not carry the staged content"
}
# The heart of the ticket's item 3: a failed swap must refuse to start. This function dies on
# failure, and die() exits — so like the require_drivable_supervisor tests above, the call goes
# inside a command substitution to contain that exit to a subshell.
test_swap_staged_jar_dies_without_staged_file() {
local dir output rc=0
dir="$TMP/swap-missing"; mkdir -p "$dir"
output="$(swap_staged_jar "$dir/fleetd-new.jar" "$dir/fleetd.jar" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] || fail "swap_staged_jar accepted a missing staged jar"
[ ! -f "$dir/fleetd.jar" ] || fail "swap_staged_jar must not create the live jar when nothing was staged"
printf '%s' "$output" | grep -qF "$dir/fleetd-new.jar" \
|| fail "refusal message does not name the missing staged path"
}
test_swap_staged_jar_dies_when_mv_fails() {
local dir staged live output rc=0
dir="$TMP/swap-fail"; mkdir -p "$dir/src"
staged="$dir/src/fleetd-new.jar"
printf 'fake jar bytes' > "$staged"
live="$dir/no-such-dir/fleetd.jar" # parent directory does not exist -> mv fails
output="$(swap_staged_jar "$staged" "$live" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] || fail "swap_staged_jar accepted a failing mv"
[ -f "$staged" ] || fail "swap_staged_jar must leave the staged jar in place when the move fails"
[ ! -f "$live" ] || fail "swap_staged_jar must not report success when the move failed"
printf '%s' "$output" | grep -qF "$staged" \
|| fail "refusal message does not name the staged path that could not be moved"
}
# --no-build must still resolve $JAR (never the staged path — there is nothing to stage on this
# path) and must still die with the exact wording documented in the script's own header comment.
test_require_no_build_jar_dies_when_absent() {
local saved_jar="$JAR" output rc=0 missing="$TMP/no-build-absent/fleetd.jar"
JAR="$missing"
output="$(require_no_build_jar 2>&1)" || rc=$?
JAR="$saved_jar"
[ "$rc" -ne 0 ] || fail "require_no_build_jar accepted a missing jar"
printf '%s' "$output" | grep -qF "no jar at $missing — run without --no-build" \
|| fail "refusal message does not match the documented --no-build wording"
}
test_require_no_build_jar_accepts_present_jar() {
local saved_jar="$JAR" dir
dir="$TMP/no-build-present"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"
printf 'existing jar' > "$JAR"
require_no_build_jar || fail "require_no_build_jar rejected an existing jar"
JAR="$saved_jar"
}
# wait_for_daemon_exit is the seam the swap ordering depends on: it must not report success while
# running_pid() still answers, and must report success the moment it clears. `sleep` is shadowed so
# the timeout-loop test does not actually wait out its budget.
test_wait_for_daemon_exit_returns_true_once_pid_clears() {
# running_pid() runs inside a $(...) — a subshell — every time wait_for_daemon_exit calls it, so
# a plain shell variable it increments would reset on each call instead of accumulating. Count in
# a file instead, which is the one thing that actually survives across those subshells.
local counter_file="$TMP/wait-exit-calls" final_calls
printf '0' > "$counter_file"
running_pid() {
local n
n="$(cat "$counter_file")"
n=$((n + 1))
printf '%s' "$n" > "$counter_file"
if [ "$n" -lt 3 ]; then printf '4242'; else printf ''; fi
}
sleep() { :; }
wait_for_daemon_exit 10 || fail "wait_for_daemon_exit did not report success once the pid cleared"
final_calls="$(cat "$counter_file")"
[ "$final_calls" -ge 3 ] || fail "wait_for_daemon_exit returned before actually re-checking running_pid"
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
}
test_wait_for_daemon_exit_times_out_if_pid_never_clears() {
local rc=0
running_pid() { printf '4242'; }
sleep() { :; }
wait_for_daemon_exit 3 || rc=$?
[ "$rc" -ne 0 ] || fail "wait_for_daemon_exit reported success while the pid never cleared"
source "$ROOT/scripts/redeploy-fleetd.sh" # restore the real running_pid/sleep for later tests
}
# fleetd #521 — the swap step's guard, at two levels.
#
# The first two tests call the predicate should_swap() directly. They pin its logic, and that is all
# they pin. On their own they did NOT close #521, and this was measured rather than argued: with the
# main flow reading `if should_swap "$DO_BUILD"; then`, changing that line to `if false; then` left
# this whole suite at exit 0 with zero FAIL lines, because nothing here made the code that performs
# the swap consult the predicate at all. Extracting the decision had moved the untested decision up
# a level, not removed it.
#
# So the last two tests call swap_if_built() — the function the main flow actually calls, holding the
# guard and the swap together — with a recording stub in place of the real `mv`. Those fail if the
# guard is removed, inverted, or stops being consulted.
#
# What none of these four can catch: deleting the `swap_if_built "$DO_BUILD"` line from the main flow
# altogether. That is test_swap_ordered_after_wait_and_before_start's job below, because sourcing
# stops before the main flow runs, so no test in this file can invoke it.
test_should_swap_true_when_build_ran() {
should_swap 1 || fail "should_swap 1 (a build ran and staged a jar) must return true"
}
test_should_swap_false_when_build_skipped() {
if should_swap 0; then
fail "should_swap 0 (--no-build; nothing was staged this run) must return false"
fi
}
# Both of these re-source redeploy-fleetd.sh at the START, because a bash function definition is
# global for the rest of the process and an earlier test may have left swap_staged_jar or jar_id
# overridden (see the longer note on this at test_detect_supervisor_systemd_probe_error_is_unclear),
# and again at the END, so their own stubs do not leak into every test that runs after them.
test_swap_if_built_performs_the_swap_when_build_ran() {
source "$ROOT/scripts/redeploy-fleetd.sh"
local marker="$TMP/swap-if-built-ran"
rm -f "$marker"
swap_staged_jar() { printf '%s -> %s\n' "$1" "$2" > "$marker"; }
jar_id() { printf 'stubbed\n'; }
swap_if_built 1 > /dev/null
[ -f "$marker" ] \
|| fail "swap_if_built 1 (a build ran and staged a jar) must perform the swap, and did not"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
test_swap_if_built_skips_the_swap_when_build_skipped() {
source "$ROOT/scripts/redeploy-fleetd.sh"
local marker="$TMP/swap-if-built-skipped"
rm -f "$marker"
swap_staged_jar() { printf 'swapped\n' > "$marker"; }
jar_id() { printf 'stubbed\n'; }
swap_if_built 0 > /dev/null
if [ -f "$marker" ]; then
fail "swap_if_built 0 (--no-build; nothing was staged this run) must not swap, but it did"
fi
source "$ROOT/scripts/redeploy-fleetd.sh"
}
# fleetd #493 item 2: "put the swap after that wait, before the start." Sourcing stops before the
# main flow ever runs (see the SOURCED guard in redeploy-fleetd.sh), so the ordering guarantee
# itself — as opposed to the pure functions it's built from — can only be checked by reading the
# script's own call sites, the same way test_recovery_patterns_match_source below checks Java
# source shape instead of behavior it cannot invoke directly.
#
# Two details about the three greps below, both of which have already gone wrong here.
#
# The needle for the swap is the MAIN FLOW's call site, `swap_if_built "$DO_BUILD"` — not
# `swap_staged_jar "$JAR_STAGED" "$JAR"`. Since fleetd #521 that second string lives inside
# swap_if_built's body, which is defined near the top of the script, far ABOVE the stop step. Using
# it made this test report "swap_staged_jar (line 215) is not after wait_for_daemon_exit (line 730)"
# — a true statement about a function definition, and nothing at all about the order of the steps.
#
# Each grep ends in `|| true`. This file runs under `set -euo pipefail`, and `pipefail` makes the
# pipeline's status grep's status, so a needle that is simply ABSENT failed the assignment and `set
# -e` killed the whole suite on the spot — before reaching the `[ -n ... ] || fail` line written to
# report exactly that. Measured: the suite exited 1 having printed zero bytes, no FAIL line and no
# name of the missing call site. `|| true` lets the assignment succeed empty so the guard can speak.
test_swap_ordered_after_wait_and_before_start() {
local src="$ROOT/scripts/redeploy-fleetd.sh" wait_line swap_line start_line
wait_line="$(grep -Fn 'wait_for_daemon_exit "$STOP_WAIT"' "$src" | head -1 | cut -d: -f1 || true)"
swap_line="$(grep -Fn 'swap_if_built "$DO_BUILD"' "$src" | head -1 | cut -d: -f1 || true)"
start_line="$(grep -Fn 'say "start"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$wait_line" ] || fail "could not find the wait-for-exit call site in redeploy-fleetd.sh"
[ -n "$swap_line" ] || fail "could not find the swap call site in redeploy-fleetd.sh"
[ -n "$start_line" ] || fail "could not find the start section in redeploy-fleetd.sh"
[ "$swap_line" -gt "$wait_line" ] \
|| fail "swap_if_built (line $swap_line) is not after wait_for_daemon_exit (line $wait_line)"
[ "$swap_line" -lt "$start_line" ] \
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
}
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
# so the only way to pin its exact wording is to read the source.
test_drain_gate_abort_message_says_no_no_build() {
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
fail "abort message still claims a rerun WITH --no-build can finish the restart"
fi
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|| fail "abort message does not tell the operator to rerun without --no-build"
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|| fail "abort message does not say why --no-build cannot finish the restart"
}
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
# a source-text grep (test_drain_gate_abort_message_says_no_no_build, below): it greps this script's
# own file for the wording, which stays in the file even if the `if` guarding it is mutated to
# `if false` and the branch can never run. These four tests call drain_gate_refusal directly instead,
# so they fail if the branch is unreachable OR if its wording regresses — the grep test is KEPT
# alongside these, not replaced, because it catches a different regression (a re-wording that still
# reaches the right branch would not change which case fires here, but would still be worth pinning).
test_drain_gate_refusal_build_ran_staged_present() {
local dir staged result
dir="$TMP/drain-refusal-build-staged"; mkdir -p "$dir"
staged="$dir/fleetd-new.jar"
printf 'staged jar bytes' > "$staged"
result="$(drain_gate_refusal 1 "$staged")"
printf '%s' "$result" | grep -qF "$staged" \
|| fail "build-ran+staged-present refusal does not name the staged jar path"
printf '%s' "$result" | grep -qF 'Rerun WITHOUT --no-build' \
|| fail "build-ran+staged-present refusal does not tell the operator how to finish the restart"
if printf '%s' "$result" | grep -qF 'nothing changed'; then
fail "build-ran+staged-present refusal must not claim nothing changed — the jar already moved"
fi
}
test_drain_gate_refusal_build_ran_staged_absent() {
local dir result
dir="$TMP/drain-refusal-build-no-staged"; mkdir -p "$dir"
result="$(drain_gate_refusal 1 "$dir/fleetd-new.jar")"
assert_equals "aborted — nothing changed" "$result" "build-ran+staged-absent refusal wording"
}
# --no-build itself never builds or stages anything (require_no_build_jar, above), so a staged jar
# found here is a leftover from an earlier, unrelated run — THIS run truly changed nothing. See the
# comment above drain_gate_refusal in redeploy-fleetd.sh for the full reasoning.
test_drain_gate_refusal_no_build_staged_present() {
local dir staged result
dir="$TMP/drain-refusal-no-build-staged"; mkdir -p "$dir"
staged="$dir/fleetd-new.jar"
printf 'leftover staged jar bytes' > "$staged"
result="$(drain_gate_refusal 0 "$staged")"
assert_equals "aborted — nothing changed" "$result" "no-build+staged-present refusal must deliberately say nothing changed"
}
test_drain_gate_refusal_no_build_staged_absent() {
local dir result
dir="$TMP/drain-refusal-no-build-no-staged"; mkdir -p "$dir"
result="$(drain_gate_refusal 0 "$dir/fleetd-new.jar")"
assert_equals "aborted — nothing changed" "$result" "no-build+staged-absent refusal wording"
}
# fleetd #528 — the four tests above pin drain_gate_refusal(), and that is ALL they pin: they call
# the predicate directly and never touch the main flow's call site. That was measured to be not
# enough, the same way test_should_swap_true_when_build_ran/test_should_swap_false_when_build_skipped
# were not enough for #521: with the main flow reading `die "$(drain_gate_refusal "$DO_BUILD"
# "$JAR_STAGED")"`, replacing that whole line with a flat `die "aborted — nothing changed"` left this
# suite at exit 0 with zero FAIL lines and byte-identical output to a clean run. Nothing above could
# tell the difference, because none of it calls anything at or above the call site itself.
#
# So these four call refuse_drain_gate() — the function the main flow actually calls, holding the
# composed message and the die() together — with die() stubbed to RECORD whether it was called and
# with what message, instead of exiting the process. That fails if refuse_drain_gate stops consulting
# drain_gate_refusal, mangles what it passes it, or simply never calls die.
#
# What none of these four can catch: deleting the `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` line
# from the main flow altogether — see the comment above refuse_drain_gate in redeploy-fleetd.sh for
# why no test in this file can do better than that (sourcing stops before the main flow runs).
DIED_CALLED=0
DIED_MESSAGE=""
stub_die_recorder() {
DIED_CALLED=0
DIED_MESSAGE=""
die() { DIED_CALLED=1; DIED_MESSAGE="$*"; }
}
test_refuse_drain_gate_build_ran_staged_present() {
source "$ROOT/scripts/redeploy-fleetd.sh"
local dir staged
dir="$TMP/refuse-drain-build-staged"; mkdir -p "$dir"
staged="$dir/fleetd-new.jar"
printf 'staged jar bytes' > "$staged"
stub_die_recorder
refuse_drain_gate 1 "$staged"
[ "$DIED_CALLED" = 1 ] \
|| fail "refuse_drain_gate build-ran+staged-present must call die, and did not"
printf '%s' "$DIED_MESSAGE" | grep -qF "$staged" \
|| fail "refuse_drain_gate build-ran+staged-present die message does not name the staged jar"
printf '%s' "$DIED_MESSAGE" | grep -qF 'Rerun WITHOUT --no-build' \
|| fail "refuse_drain_gate build-ran+staged-present die message is missing the rerun instruction"
if printf '%s' "$DIED_MESSAGE" | grep -qF 'nothing changed'; then
fail "refuse_drain_gate build-ran+staged-present must not claim nothing changed — the jar already moved"
fi
source "$ROOT/scripts/redeploy-fleetd.sh"
}
test_refuse_drain_gate_build_ran_staged_absent() {
source "$ROOT/scripts/redeploy-fleetd.sh"
local dir
dir="$TMP/refuse-drain-build-no-staged"; mkdir -p "$dir"
stub_die_recorder
refuse_drain_gate 1 "$dir/fleetd-new.jar"
[ "$DIED_CALLED" = 1 ] \
|| fail "refuse_drain_gate build-ran+staged-absent must call die, and did not"
assert_equals "aborted — nothing changed" "$DIED_MESSAGE" "refuse_drain_gate build-ran+staged-absent die message"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
test_refuse_drain_gate_no_build_staged_present() {
source "$ROOT/scripts/redeploy-fleetd.sh"
local dir staged
dir="$TMP/refuse-drain-no-build-staged"; mkdir -p "$dir"
staged="$dir/fleetd-new.jar"
printf 'leftover staged jar bytes' > "$staged"
stub_die_recorder
refuse_drain_gate 0 "$staged"
[ "$DIED_CALLED" = 1 ] \
|| fail "refuse_drain_gate no-build+staged-present must call die, and did not"
assert_equals "aborted — nothing changed" "$DIED_MESSAGE" "refuse_drain_gate no-build+staged-present die message"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
test_refuse_drain_gate_no_build_staged_absent() {
source "$ROOT/scripts/redeploy-fleetd.sh"
local dir
dir="$TMP/refuse-drain-no-build-no-staged"; mkdir -p "$dir"
stub_die_recorder
refuse_drain_gate 0 "$dir/fleetd-new.jar"
[ "$DIED_CALLED" = 1 ] \
|| fail "refuse_drain_gate no-build+staged-absent must call die, and did not"
assert_equals "aborted — nothing changed" "$DIED_MESSAGE" "refuse_drain_gate no-build+staged-absent die message"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
# fleetd #528 — closes the one gap the four behavioural tests above cannot: they call
# refuse_drain_gate directly, and sourcing stops before the main flow ever runs (the SOURCED guard),
# so none of them can prove the main flow still CALLS refuse_drain_gate at all. Same shape as
# test_swap_ordered_after_wait_and_before_start: a source-text grep for the real call site. This is
# what actually kills the item-1 mutation from the ticket — replacing the main flow's call with a
# flat `die "aborted — nothing changed"` removes this exact needle, where none of the behavioural
# tests above would even notice.
#
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
test_refuse_drain_gate_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's refuse_drain_gate call site in redeploy-fleetd.sh"
}
# fleetd #504 — the "loaded but not currently running" branches for launchd/systemd used to run
# `launchctl unload`/`systemctl --user stop` with `2>/dev/null || true` and print `ok`
# unconditionally, so a real supervisor failure (e.g. it cannot reach launchd/the systemd user bus)
# read exactly like a harmless already-stopped answer. unload_launchd_if_loaded/
# stop_systemd_if_loaded (redeploy-fleetd.sh, right after systemd_loaded) apply systemd_loaded's own
# "capture stderr separately — only a non-zero exit WITH stderr is a real failure" pattern to the
# WRITE side. Both call the real `launchctl`/`systemctl` binaries directly (they are not overridable
# wrapper functions the way launchd_loaded/systemd_loaded are), so these tests put a stub binary
# first on PATH — the same technique test_detect_supervisor_systemd_probe_error_is_unclear above
# already uses for `systemctl`.
test_unload_launchd_if_loaded_dies_on_real_failure() {
local bin_dir output rc=0 saved_plist="$LAUNCHD_PLIST"
bin_dir="$TMP/stub-bin-launchctl-error"; mkdir -p "$bin_dir"
cat > "$bin_dir/launchctl" <<'STUB'
#!/usr/bin/env bash
echo "Could not find specified service" >&2
exit 1
STUB
chmod +x "$bin_dir/launchctl"
LAUNCHD_PLIST="$TMP/fake-fail.plist"
output="$(PATH="$bin_dir:$PATH" unload_launchd_if_loaded 2>&1)" || rc=$?
LAUNCHD_PLIST="$saved_plist"
[ "$rc" -ne 0 ] \
|| fail "unload_launchd_if_loaded must die when launchctl exits non-zero AND writes to stderr"
printf '%s' "$output" | grep -qF 'launchctl unload' \
|| fail "die message does not name the failing launchctl unload command"
}
# Captured via $(...) rather than called bare: unload_launchd_if_loaded's own die() does a hard
# `exit`, and calling it directly at this level would let a regression that makes it die on this
# clean-negative case kill the WHOLE suite before the `|| fail` below ever ran — printing die's own
# message instead of this test's. Inside a command substitution, that `exit` only ends the subshell
# (a-guard-is-defeated-by-its-calling-context: the same reason the *_dies_on_real_failure tests
# above capture this way), so this test's own message is what actually reaches the report.
test_unload_launchd_if_loaded_tolerates_clean_negative() {
local bin_dir saved_plist="$LAUNCHD_PLIST" output rc=0
bin_dir="$TMP/stub-bin-launchctl-noop"; mkdir -p "$bin_dir"
cat > "$bin_dir/launchctl" <<'STUB'
#!/usr/bin/env bash
exit 1
STUB
chmod +x "$bin_dir/launchctl"
LAUNCHD_PLIST="$TMP/fake-noop.plist"
output="$(PATH="$bin_dir:$PATH" unload_launchd_if_loaded 2>&1)" || rc=$?
LAUNCHD_PLIST="$saved_plist"
[ "$rc" -eq 0 ] \
|| fail "unload_launchd_if_loaded must tolerate a clean already-unloaded answer (non-zero exit, empty stderr): $output"
}
test_stop_systemd_if_loaded_dies_on_real_failure() {
local bin_dir output rc=0
bin_dir="$TMP/stub-bin-systemctl-stop-error"; mkdir -p "$bin_dir"
cat > "$bin_dir/systemctl" <<'STUB'
#!/usr/bin/env bash
echo "Failed to connect to bus: No such file or directory" >&2
exit 1
STUB
chmod +x "$bin_dir/systemctl"
output="$(PATH="$bin_dir:$PATH" stop_systemd_if_loaded 2>&1)" || rc=$?
[ "$rc" -ne 0 ] \
|| fail "stop_systemd_if_loaded must die when systemctl exits non-zero AND writes to stderr"
printf '%s' "$output" | grep -qF 'systemctl --user stop' \
|| fail "die message does not name the failing systemctl --user stop command"
}
# Same subshell-capture reasoning as test_unload_launchd_if_loaded_tolerates_clean_negative above:
# stop_systemd_if_loaded's own die() does a hard `exit`, so this must run inside $(...) or a
# regression here would kill the whole suite with die's message instead of this test's.
test_stop_systemd_if_loaded_tolerates_clean_negative() {
local bin_dir output rc=0
bin_dir="$TMP/stub-bin-systemctl-stop-noop"; mkdir -p "$bin_dir"
cat > "$bin_dir/systemctl" <<'STUB'
#!/usr/bin/env bash
exit 1
STUB
chmod +x "$bin_dir/systemctl"
output="$(PATH="$bin_dir:$PATH" stop_systemd_if_loaded 2>&1)" || rc=$?
[ "$rc" -eq 0 ] \
|| fail "stop_systemd_if_loaded must tolerate a clean already-stopped answer (non-zero exit, empty stderr): $output"
}
# Closes the same gap test_refuse_drain_gate_call_site_present closes for the drain gate: the four
# tests above call unload_launchd_if_loaded/stop_systemd_if_loaded directly, and sourcing stops
# before the main flow ever runs (the SOURCED guard), so none of them can prove the main flow still
# CALLS these two functions instead of the original bare `2>/dev/null || true`. A source-text check,
# like test_swap_ordered_after_wait_and_before_start. The call-site needle is anchored (`^ name$`)
# so it cannot be satisfied by the comment lines above each call site that merely mention the
# function by name.
test_stop_branches_call_tolerant_helpers_not_bare_or_true() {
local src="$ROOT/scripts/redeploy-fleetd.sh" unload_call_line stop_call_line
unload_call_line="$(grep -n '^ unload_launchd_if_loaded$' "$src" | head -1 | cut -d: -f1 || true)"
stop_call_line="$(grep -n '^ stop_systemd_if_loaded$' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$unload_call_line" ] \
|| fail "could not find the main flow's call to unload_launchd_if_loaded in redeploy-fleetd.sh"
[ -n "$stop_call_line" ] \
|| fail "could not find the main flow's call to stop_systemd_if_loaded in redeploy-fleetd.sh"
if grep -qF 'launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true' "$src"; then
fail "the bare 'launchctl unload ... 2>/dev/null || true' defect (fleetd #504) is back in redeploy-fleetd.sh"
fi
if grep -qF 'systemctl --user stop "$SYSTEMD_UNIT" 2>/dev/null || true' "$src"; then
fail "the bare 'systemctl --user stop ... 2>/dev/null || true' defect (fleetd #504) is back in redeploy-fleetd.sh"
fi
}
test_no_errors() {
cat > "$TMP/no-errors.log" <<'LOG'
2026-09-05 12:00:00 INFO fleetd listening
@@ -224,6 +924,207 @@ test_unattributable_quiet_mutation_is_caught() {
printf 'Unattributable mutation: FAIL: cross-unattributable recovered: expected 0, got 2\n'
}
# fleetd #512 part 2 — the negative check (scan_uncaught_exceptions). The heart of this half of the
# ticket: a fixture with the uncaught-exception shape and NO line carrying an ERROR token at all,
# proving the scan finds it without one. A fixture that also carried an ERROR line would pass for
# the wrong reason.
test_scan_uncaught_exceptions_finds_shape_without_error_token() {
cat > "$TMP/scan-died.log" <<'LOG'
2026-09-12 10:15:00 INFO fleetd listening on 127.0.0.1:8765
Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions
at dev.ltms.fleet.session.SessionManager.drainAll(SessionManager.java:1081)
LOG
local error_count
error_count="$(grep -c ' ERROR ' "$TMP/scan-died.log" || true)"
[ "$error_count" = "0" ] \
|| fail "test fixture error: scan-died.log unexpectedly carries an ERROR token"
scan_uncaught_exceptions "$TMP/scan-died.log"
assert_equals 1 "$REDEPLOY_UNCAUGHT_EXCEPTION_COUNT" "scan must find the exception without an ERROR token"
printf '%s' "$REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE" | grep -qF 'NoClassDefFoundError' \
|| fail "scan did not capture the matching line as the sample"
}
test_scan_uncaught_exceptions_clean_control() {
cat > "$TMP/scan-clean.log" <<'LOG'
2026-09-12 10:15:00 INFO fleetd listening on 127.0.0.1:8765
2026-09-12 10:15:05 INFO dev.ltms.fleet.session.SessionManager - drain complete: released=0 abandoned=0 (still BUSY at the shutdown deadline)
LOG
scan_uncaught_exceptions "$TMP/scan-clean.log"
assert_equals 0 "$REDEPLOY_UNCAUGHT_EXCEPTION_COUNT" "clean control must find no uncaught exception"
assert_equals "" "$REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE" "clean control sample must be empty"
}
# fleetd #512 part 2 — the positive check (find_drain_complete_line). Both halves of #522's line:
# present, and absent.
test_find_drain_complete_line_present() {
cat > "$TMP/drain-line-present.log" <<'LOG'
2026-09-12 10:15:05 INFO dev.ltms.fleet.session.SessionManager - drain complete: released=2 abandoned=1 (still BUSY at the shutdown deadline)
LOG
find_drain_complete_line "$TMP/drain-line-present.log"
printf '%s' "$REDEPLOY_DRAIN_COMPLETE_LINE" | grep -qF 'released=2 abandoned=1' \
|| fail "find_drain_complete_line did not capture the present line"
}
test_find_drain_complete_line_absent() {
cat > "$TMP/drain-line-absent.log" <<'LOG'
2026-09-12 10:15:00 INFO fleetd listening on 127.0.0.1:8765
LOG
find_drain_complete_line "$TMP/drain-line-absent.log"
assert_equals "" "$REDEPLOY_DRAIN_COMPLETE_LINE" "find_drain_complete_line must report empty when absent"
}
# fleetd #512 part 2 — report_shutdown_drain, the composite decision+action function the main flow
# calls unconditionally (same shape as swap_if_built/refuse_drain_gate, #521/#528). These four cover
# the four outcomes named in the ticket's "trap": complete, died, unknown ("cannot tell" — neither a
# pass nor a failure), and n/a (no previous daemon was actually stopped this run).
#
# Deliberately NOT run inside `$(...)`: report_shutdown_drain sets REDEPLOY_DRAIN_STATE as a global
# side effect that these tests need to read back afterward, and a command substitution forks a
# subshell that global assignment would not survive (the exact trap documented above
# detect_supervisor in redeploy-fleetd.sh, for the same reason). Plain output redirection to a file
# does not fork a subshell, so it is used to capture what was printed instead.
test_report_shutdown_drain_died_without_error_token() {
cat > "$TMP/drain-died.log" <<'LOG'
2026-09-12 10:15:00 INFO fleetd listening on 127.0.0.1:8765
2026-09-12 10:15:05 INFO dev.ltms.fleet.Fleetd - shutting down
Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions
at dev.ltms.fleet.session.SessionManager.drainAll(SessionManager.java:1081)
LOG
local error_count
error_count="$(grep -c ' ERROR ' "$TMP/drain-died.log" || true)"
[ "$error_count" = "0" ] \
|| fail "test fixture error: drain-died.log unexpectedly carries an ERROR token"
report_shutdown_drain "$TMP/drain-died.log" 1 > "$TMP/drain-died-output" 2>&1
assert_equals "died" "$REDEPLOY_DRAIN_STATE" "died fixture must set REDEPLOY_DRAIN_STATE=died"
assert_equals 1 "$REDEPLOY_UNCAUGHT_EXCEPTION_COUNT" "died fixture uncaught-exception count"
grep -qF 'NoClassDefFoundError' "$TMP/drain-died-output" \
|| fail "report_shutdown_drain did not report the uncaught-exception shape it found"
grep -qF 'DIED' "$TMP/drain-died-output" \
|| fail "report_shutdown_drain did not report the drain as DIED"
}
test_report_shutdown_drain_complete_control() {
cat > "$TMP/drain-complete.log" <<'LOG'
2026-09-12 10:15:00 INFO fleetd listening on 127.0.0.1:8765
2026-09-12 10:15:05 INFO dev.ltms.fleet.Fleetd - shutting down
2026-09-12 10:15:05 INFO dev.ltms.fleet.session.SessionManager - drain complete: released=3 abandoned=0 (still BUSY at the shutdown deadline)
LOG
report_shutdown_drain "$TMP/drain-complete.log" 1 > "$TMP/drain-complete-output" 2>&1
assert_equals "complete" "$REDEPLOY_DRAIN_STATE" "complete-control fixture must set REDEPLOY_DRAIN_STATE=complete"
assert_equals 0 "$REDEPLOY_UNCAUGHT_EXCEPTION_COUNT" "complete-control fixture must find no uncaught exception"
grep -qF 'released=3 abandoned=0' "$TMP/drain-complete-output" \
|| fail "report_shutdown_drain did not report the drain-complete counts"
}
test_report_shutdown_drain_unknown_cannot_tell() {
cat > "$TMP/drain-unknown.log" <<'LOG'
2026-09-12 10:15:00 INFO fleetd listening on 127.0.0.1:8765
2026-09-12 10:15:05 INFO dev.ltms.fleet.Fleetd - shutting down
LOG
report_shutdown_drain "$TMP/drain-unknown.log" 1 > "$TMP/drain-unknown-output" 2>&1
assert_equals "unknown" "$REDEPLOY_DRAIN_STATE" "cannot-tell fixture must set REDEPLOY_DRAIN_STATE=unknown"
grep -qF 'cannot tell' "$TMP/drain-unknown-output" \
|| fail "report_shutdown_drain did not say it could not tell"
if grep -qF ' ok' "$TMP/drain-unknown-output"; then
fail "cannot-tell outcome must not be printed via ok() — it is neither a pass nor a failure"
fi
}
# A cold start (or a restart where nothing was actually stopped) has no previous-daemon shutdown
# window to have an opinion about at all. This fixture's log content looks exactly like a died drain
# — proving the had_previous_daemon=0 gate is actually consulted, not merely documented: without it,
# this would misreport "died" or "unknown" on every clean cold start.
test_report_shutdown_drain_no_previous_daemon_is_na() {
cat > "$TMP/drain-na.log" <<'LOG'
Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions
LOG
report_shutdown_drain "$TMP/drain-na.log" 0 > "$TMP/drain-na-output" 2>&1
assert_equals "n/a" "$REDEPLOY_DRAIN_STATE" "no-previous-daemon fixture must set REDEPLOY_DRAIN_STATE=n/a even though the log content looks like a died drain"
grep -qF 'nothing to check' "$TMP/drain-na-output" \
|| fail "report_shutdown_drain did not report that there was nothing to check"
}
# fleetd #512 — closes the gap none of the seven tests above can: they call report_shutdown_drain
# directly, and sourcing stops before the main flow ever runs (the SOURCED guard), so none of them
# can prove the main flow still calls it at all. Same shape as test_refuse_drain_gate_call_site_present
# and test_swap_ordered_after_wait_and_before_start: a source-text grep for the real call site, plus
# an ordering check against its neighbours in the verify/result flow.
test_report_shutdown_drain_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'report_shutdown_drain "$FRESH_LOG" "$HAD_OLD_PID"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's report_shutdown_drain call site in redeploy-fleetd.sh"
}
test_report_shutdown_drain_ordered_after_classify_and_before_result() {
local src="$ROOT/scripts/redeploy-fleetd.sh" classify_line drain_line result_line
classify_line="$(grep -Fn 'classify_amqp_connection_errors "$FRESH_LOG"' "$src" | tail -1 | cut -d: -f1 || true)"
drain_line="$(grep -Fn 'report_shutdown_drain "$FRESH_LOG" "$HAD_OLD_PID"' "$src" | head -1 | cut -d: -f1 || true)"
result_line="$(grep -Fn 'say "result"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$classify_line" ] || fail "could not find the classify_amqp_connection_errors call site"
[ -n "$drain_line" ] || fail "could not find the report_shutdown_drain call site"
[ -n "$result_line" ] || fail "could not find the result section"
[ "$drain_line" -gt "$classify_line" ] \
|| fail "report_shutdown_drain (line $drain_line) is not after classify_amqp_connection_errors (line $classify_line)"
[ "$drain_line" -lt "$result_line" ] \
|| fail "report_shutdown_drain (line $drain_line) is not before the result section (line $result_line)"
}
# fleetd #512 item 4 — the summary line must not read as reassurance when the shutdown-drain check
# found something wrong (or could not tell). Sourcing stops before the main flow runs, so this is a
# source-text check like test_drain_gate_abort_message_says_no_no_build above.
test_no_error_lines_message_gated_by_drain_state() {
local src="$ROOT/scripts/redeploy-fleetd.sh" block
block="$(grep -B2 -F 'ok "no ERROR lines since restart"' "$src")"
[ -n "$block" ] || fail "could not find the 'no ERROR lines since restart' line in redeploy-fleetd.sh"
printf '%s' "$block" | grep -qF 'REDEPLOY_DRAIN_STATE' \
|| fail "'no ERROR lines since restart' is not guarded by the shutdown-drain outcome (fleetd #512 item 4)"
}
test_detect_supervisor_launchd_only
test_detect_supervisor_systemd_only
test_detect_supervisor_none
test_detect_supervisor_systemd_installed_not_loaded_is_unclear
test_detect_supervisor_launchd_installed_not_loaded_is_unclear
test_detect_supervisor_systemd_probe_error_is_unclear
test_require_drivable_supervisor_refuses_ambiguous
test_require_drivable_supervisor_refuses_unclear
test_require_drivable_supervisor_accepts_known_kinds
test_count_daemon_pids
test_assert_single_daemon_accepts_one_pid
test_assert_single_daemon_rejects_two_pids
test_jar_id_defaults_to_live_and_reports_explicit_path
test_jar_id_reports_absent_for_missing_file
test_stage_built_jar_moves_off_live_path
test_stage_built_jar_dies_when_build_produced_nothing
test_swap_staged_jar_moves_staged_onto_live
test_swap_staged_jar_dies_without_staged_file
test_swap_staged_jar_dies_when_mv_fails
test_should_swap_true_when_build_ran
test_should_swap_false_when_build_skipped
test_swap_if_built_performs_the_swap_when_build_ran
test_swap_if_built_skips_the_swap_when_build_skipped
test_require_no_build_jar_dies_when_absent
test_require_no_build_jar_accepts_present_jar
test_wait_for_daemon_exit_returns_true_once_pid_clears
test_wait_for_daemon_exit_times_out_if_pid_never_clears
test_swap_ordered_after_wait_and_before_start
test_drain_gate_abort_message_says_no_no_build
test_drain_gate_refusal_build_ran_staged_present
test_drain_gate_refusal_build_ran_staged_absent
test_drain_gate_refusal_no_build_staged_present
test_drain_gate_refusal_no_build_staged_absent
test_refuse_drain_gate_build_ran_staged_present
test_refuse_drain_gate_build_ran_staged_absent
test_refuse_drain_gate_no_build_staged_present
test_refuse_drain_gate_no_build_staged_absent
test_refuse_drain_gate_call_site_present
test_unload_launchd_if_loaded_dies_on_real_failure
test_unload_launchd_if_loaded_tolerates_clean_negative
test_stop_systemd_if_loaded_dies_on_real_failure
test_stop_systemd_if_loaded_tolerates_clean_negative
test_stop_branches_call_tolerant_helpers_not_bare_or_true
test_no_errors
test_recovery_patterns_match_source
test_attributed_recovered_connection_error
@@ -235,4 +1136,15 @@ test_other_error_is_unexplained
test_recovery_requirement_mutation_is_caught
test_shared_counter_mutation_is_caught
test_unattributable_quiet_mutation_is_caught
test_scan_uncaught_exceptions_finds_shape_without_error_token
test_scan_uncaught_exceptions_clean_control
test_find_drain_complete_line_present
test_find_drain_complete_line_absent
test_report_shutdown_drain_died_without_error_token
test_report_shutdown_drain_complete_control
test_report_shutdown_drain_unknown_cannot_tell
test_report_shutdown_drain_no_previous_daemon_is_na
test_report_shutdown_drain_call_site_present
test_report_shutdown_drain_ordered_after_classify_and_before_result
test_no_error_lines_message_gated_by_drain_state
printf 'PASS: redeploy log classifier\n'