Compare commits

..

50 Commits

Author SHA1 Message Date
Dai Ha d105da978d fleetd #680: pin the JAR/BUILD_JAR split, check the plist's jar path, and widen the daemon-locator pattern
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Failing after 2m28s
Part 1: asserts JAR and BUILD_JAR as the script sources them (no test assigns
them first), so reverting JAR to a path under target/ now fails the suite.

Part 2: check_jar_path_matches_plist reads the installed launchd plist's
ProgramArguments and refuses when its jar path does not resolve to $JAR,
mirroring check_log_path_matches_plist. Wired into report_supervisor_state's
launchd branch; unaffected when no plist is installed.

Part 3 (added to the ticket after the brief, by comment): PATTERN narrowed to
'fleetd.jar' so running_pid()/assert_single_daemon see a daemon regardless of
which build layout (run/ or target/) its jar sits under. The comm=java
allowlist still excludes a self-matching shell. Brought
.claude/skills/fleets-status/SKILL.md's pgrep pattern into agreement.

Verified: bash scripts/test-redeploy-fleetd.sh exits 0. Mutation both
directions for part 1 (JAR under target/ -> suite fails; JAR elsewhere ->
suite passes), a positive control for part 2 (neutralizing the mismatch
check makes the new test fail), and a positive control for part 3 (narrowing
PATTERN back to run/fleetd.jar makes the new target/-dir test fail). mvn -o
clean install: Tests run: 1929, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 172 surefire report files.
2026-10-03 21:28:59 +02:00
Dai Ha 5051a06443 Merge PR #679: fleetd #664 — run the daemon from fleetd/run/fleetd.jar
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m52s
Verified by the lead before merge:
- three-dot diff: 1 commit, 9 files, nothing dragged in
- trial merge in a throwaway worktree: no conflicts
- mvn clean install from fleetd/: Tests run: 1929, Failures: 0, 172 report files
- scripts/test-redeploy-fleetd.sh: exit 0; its 3 FAIL and 5 mktemp lines are
  deliberate self-test output, confirmed identical on unmerged origin/main
- run/ is correctly ignored by the new fleetd/.gitignore rule
- the build writes target/fleetd.jar and does not create run/
2026-10-03 21:13:33 +02:00
Dai Ha a134eccc57 fleetd #664: run the daemon from fleetd/run/fleetd.jar, not target/
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 2m58s
Separate Maven's output path (target/fleetd.jar) from the daemon's runtime
path (run/fleetd.jar), so a clean/install in the main clone can no longer
reach the jar a running daemon holds open. redeploy-fleetd.sh now swaps the
built jar onto the runtime path with a same-filesystem mv, only after the
old daemon is confirmed gone; --check reports the built and running jars
as two separately labelled hash+mtime facts. Updates every process-locator
pattern and runtime-path reference found by git grep, with a positive
control added for both running_pid()'s PATTERN and fleets-status's pgrep
pattern. fleetd/run/ is gitignored.
2026-10-03 21:06:42 +02:00
Dai Ha 6f275227d2 fleetd #668: drop counts from the javadoc that the new case made wrong
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m3s
CI / build (push) Failing after 1m35s
The claim-2 banner said validateAll reaches seven of eight validators. With
the validateLeadRollover case added it reaches all of them, so the banner was
false. Two more places said 'today's six' and were already stale at eight.

Removed the '1491 tests, 0 failures' parenthetical: a measurement in a comment
goes stale on its own, and the suite is far past that number. The count that
readers must keep correct still lives in the Set.of, which is where the
surrounding javadoc already points them.
2026-10-03 20:52:29 +02:00
Dai Ha 283ccf8423 Merge PR #674: fleetd #668 — add the validateLeadRollover reachability case 2026-10-03 20:51:41 +02:00
Dai Ha b96fba4a03 fleetd #672: keep the test's comments to the contract
The class javadoc carried a ticket key, a file:line reference, a comparison
with FleetMcpAuthzTest and pointers at two other assembly tests. That is
history and review justification, which belong in the commit message and the
PR, not in the code. The javadoc now names the behaviour the test protects.

The reflection note keeps the constraint a maintainer needs (the package
boundary, and that no catch can hide a renamed denyFor) and drops the rest.
The assertion message no longer names a line number that will drift.
2026-10-03 20:50:16 +02:00
Dai Ha 6d97d210b4 Merge PR #673: fleetd #672 — pin AuthorizationMode.ENFORCED at FleetdAssembly.java:481 2026-10-03 20:49:34 +02:00
Dai Ha 37b23cd704 fleetd #668: add the missing validateLeadRollover reachability case
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 1m43s
validateAllReachesEveryOneOfTodaysRealValidators now exercises all
eight FleetConfig validators through validateAll(), not seven -
adding a minimal leadRollover: block with no handoverPath as the
eighth fixture. The canary's failure message in
fleetConfigDeclaresExactlyTheseValidatorsToday now also points the
reader at the reachability enumeration, since updating the expected
set alone does not prove validateAll() reaches a newly added
validator.
2026-10-03 20:48:03 +02:00
Dai Ha 367facf6a6 fleetd #672: pin AuthorizationMode.ENFORCED at FleetdAssembly.java:481
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 1m41s
Adds FleetdAssemblyAuthorizationModeTest: drives FleetMcp#denyFor (via reflection, since it is package-private to dev.ltms.fleet.mcp) on the real FleetMcp FleetdAssembly#assembleAndStart builds, asserting an unauthorized worker is refused SPAWN and the primary is still allowed. Mutating line 481 to UNENFORCED turns this test red and no other test.
2026-10-03 20:43:21 +02:00
Dai Ha 7f9a9c09f9 Merge PR #671: fleetd #670 — pin excludedWorkspaceLabels at FleetdAssembly.java:265
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 1m49s
2026-10-03 20:27:15 +02:00
Dai Ha e854957247 fleetd #670: pin excludedWorkspaceLabels at FleetdAssembly.java:265
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 1m58s
Adds a test that reaches the real LeadTabScanner FleetdAssembly's
production boot path builds (via the LeadCoordLoop field that stores
the same leads supplier instance) and asserts, by reflection, that
the excludedWorkspaceLabels field is empty. Mutating line 265 to any
non-empty set now turns this test red.
2026-10-03 20:23:06 +02:00
Dai Ha b4b7cf5155 fleetd #661: drop the validator counts from validateAll's javadoc
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m46s
FleetConfig declares eight public no-arg void validate* methods, at lines
2671, 2708, 2755, 2790, 2827, 2856, 2896 and 2958. The sweep's javadoc still
named a count. The wording is now count-free, so it cannot drift again.
2026-10-03 19:59:07 +02:00
Dai Ha cbb35ad947 Merge PR #667: fleetd #661 — refuse pane placement when a lead tab is configured 2026-10-03 19:54:10 +02:00
Dai Ha 656588f597 fleetd #661: add the pane-placement case to the validateAll reachability enumeration
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m40s
validateAllReachesEveryOneOfTodaysSixValidators only covered six of
the eight real validators; the new validator was reachability-tested
only from FleetConfigTest, in a different file from the one whose job
is to enumerate every validateAll-reachability case.

Add the pane-placement case to the enumeration, rename the method to
drop the hardcoded count (validateAllReachesEveryOneOfTodaysRealValidators),
and correct the surrounding claims to say seven of eight, naming
validateLeadRollover as the one case still missing (fleetd #668, not
fixed here).
2026-10-03 19:50:16 +02:00
Dai Ha e33377b2ca fleetd #661: fix LeadCount javadoc, dangling @link, and the validator-count word
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 1m50s
LeadLauncher.LeadCount's javadoc carried the same false member-space-
exclusion claim as the three comments fixed earlier in this ticket;
its neighbouring body comment in countLeads was already correct and
is unchanged.

FleetConfigValidateAllTest's canary test is renamed to drop the
number from its name (the count now lives only in the Set.of literal
and the javadoc, so the two cannot drift), which also fixes the
dangling {@link} to the old name and the stale 'seventh' wording.
2026-10-03 19:42:50 +02:00
Dai Ha 4b4a8688c2 Merge PR #666: fleetd #664 — correct the main-clone build guidance
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 1m44s
The old text warned only that a 'mvn clean' deletes the running daemon's jar.
Replacement is enough: the shutdown drain loads its classes lazily, at
shutdown, from the jar file the JVM opened at boot. So any build that writes
fleetd/target/fleetd.jar under a live daemon breaks its drain, nothing warns
at the time, and the damage surfaces at the next restart where it looks like
the restart's fault.

Both files now state the real rule and the positive one: verify a merge by
building in a throwaway git worktree, and let only
scripts/redeploy-fleetd.sh touch the main clone's jar.

The canonical block is untouched (20938 bytes, identical to main) and the
wiki sync check passes against wiki/7-Use-Cases.md.
2026-10-03 19:40:35 +02:00
Dai Ha 7e48d4b86c Merge PR #665: fleetd #663 — remove LeadContextGauge.read's 3-arg overload
CI / build (push) Failing after 1m35s
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 1m0s
The 3-arg form delegated to the 4-arg one with a null window, so any caller
reaching for it silently got the fixed HIGH_THRESHOLD_TOKENS back instead of
the profile's effective auto-compact window. It had no production callers.
All 15 test call sites move to the 4-arg form.

Verified at 60fa86a in a throwaway worktree: Tests run: 1923, Failures: 0,
Errors: 0, Skipped: 0, BUILD SUCCESS, 170 surefire report files. Mutating
HIGH_THRESHOLD_TOKENS to 200_000 * 2 turns
noEffectiveWindowFallsBackToTheFixed200000Default red at line 81.
2026-10-03 19:38:13 +02:00
Dai Ha 41cc785534 fleetd #664: explain delayed jar failure
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m43s
2026-10-03 19:36:08 +02:00
Dai Ha 60fa86a107 fleetd #663: remove the dead legacy assertion review flagged on PR 665
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m6s
noEffectiveWindowFallsBackToTheFixed200000Default's third assertion
(legacyConfigDir/legacyGauge) used to call the 3-arg read() to prove
it behaved like a null window. With the 3-arg form gone, it is the
same call, same input (200_000) and same expectation as the atGauge
assertion above it, so it cannot fail unless that one already failed,
and its message named a method that no longer exists. Deleted it; the
first two assertions (199_999 -> OK, 200_000 -> HIGH) are unchanged.
2026-10-03 19:34:47 +02:00
Dai Ha f288cee2bb fleetd #661: refuse pane placement when a lead tab is configured
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m55s
A pane-placed member lands inside the focused tab rather than its own,
so it can land inside a lead's labelled tab and be read back as that
lead by LeadTabScanner, which does not exclude the member space in
production. Add FleetConfig.validatePanePlacementAgainstLeadTabs(),
wired automatically into validateAll() by the existing reflective
sweep, to refuse that combination at startup.

Also correct three stale comments that claimed a member-space
exclusion already blocked this path, in LeadTabScanner, FleetConfig's
validateLeadTabPrefixes javadoc, and LeadTabScannerTest.
2026-10-03 19:34:34 +02:00
Dai Ha a5d6ce1a37 fleetd #664: protect the running daemon jar
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Failing after 1m53s
2026-10-03 19:29:25 +02:00
Dai Ha a52ca35d34 fleetd #663: remove LeadContextGauge.read's 3-arg overload
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m34s
The 3-arg read(configDir, sessionId, agentType) delegated to the 4-arg
form with a null window, silently restoring the fixed 200_000 fallback
that #637 moved away from. It had zero production callers; both
production call sites already use the 4-arg form.

Migrate all test call sites to the 4-arg form. For the generic property
tests in LeadContextGaugeTest, the window is irrelevant and null is
filler. In LeadContextGaugeHighThresholdTest's
noEffectiveWindowFallsBackToTheFixed200000Default, null is the
meaningful value under test, not filler; the assertion is unchanged.
2026-10-03 19:26:11 +02:00
Dai Ha 9417de1123 Merge PR #662: fleetd #659 — remove the dead back-compat forms around LeadConfigDirSource/leadContextLookup
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m33s
CI / build (push) Failing after 1m58s
2026-10-03 19:11:18 +02:00
Dai Ha 03d92be751 fleetd #659: remove the dead back-compat forms around LeadConfigDirSource/leadContextLookup
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m41s
CI / build (pull_request) Failing after 2m31s
FleetMcp.LeadConfigDirSource's 1-arg constructor, Fleetd.leadContextLookup's
4-arg overload, and Fleetd.leadContextSource's 4-arg overload each existed only
to keep old call sites compiling, and each silently resolved no auto-compact
window — reverting any caller that picked one up to LeadContextGauge's fixed
200,000 HIGH threshold, the exact defect #637 fixed. Removed all three and
updated the 8 call sites across 4 test files to pass the window lookup
explicitly.

Added FleetdLeadContextSourceWindowAssemblyTest: no existing test called the
real assembled LeadHeartbeatLoop far enough to prove FleetdAssembly's
window-lookup argument into Fleetd.leadContextSource actually reaches the
gauge. Mutating that argument to `_ -> null` compiled clean and left the whole
suite green; the new test fails against that mutation. Needed a small FakeHerdr
addition (agentSessionId(..)) since its default agent.get response carries no
session id.
2026-10-03 18:55:24 +02:00
Dai Ha 136bec8e28 Merge PR #660: fleetd #637 — scale the lead context HIGH threshold with the effective auto-compact window
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m4s
CI / build (push) Failing after 2m35s
2026-10-03 16:27:55 +02:00
Dai Ha ae94d511d7 fleetd #637: pin the fleet_list window wiring and fold the threshold into the gauge's cache key
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 2m22s
Fleetd.leadConfigDirSource built its window argument with nothing calling the
factory itself to prove it, so a mutation to a no-op lookup left the whole
suite green. Add a wiring test that calls the factory directly, modeled on
FleetdLeadConfigDirSourceWiringTest's shape for the configDir half of the
same factory.

LeadContextGauge's result cache keyed only on (configDir, sessionId), but the
cached Reading's state now depends on the caller's resolved effective window.
Fold the derived threshold into the cache key so two reads of the same
session with different windows inside the TTL each report against their own
window.
2026-10-03 16:18:57 +02:00
Dai Ha 436b026696 fleetd #637: scale LeadContextGauge's HIGH threshold with the effective auto-compact window
HIGH_THRESHOLD_TOKENS was a hardcoded 200_000, while the event it warns about
(auto-compaction) is configured per profile via autoCompactWindow and can legally go
as low as 100_000 — making HIGH unreachable before a compaction on such a profile.

LeadContextGauge.read now takes an optional effective window and fires HIGH at 2/3 of
it, falling back to the fixed 200_000 when no window is resolvable (unresolved callers,
including the pre-existing 3-arg read(), keep today's behaviour exactly).

FleetConfig.Profile.effectiveAutoCompactWindow() resolves that window the way a
launched Claude Code session actually reads it: env.CLAUDE_CODE_AUTO_COMPACT_WINDOW
wins over the autoCompactWindow launch flag when both are set.

Wired into both real consumers: fleet_list's context row (FleetMcp.LeadConfigDirSource,
widened with a back-compat constructor so no unrelated call site changes) and the lead
heartbeat's context-high nudge (Fleetd.leadContextLookup/leadContextSource, widened the
same way).
2026-10-03 16:11:41 +02:00
Dai Ha 905fa3a454 Merge PR #658: fleetd #656 — regression tests for both redact() leaks
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m54s
2026-10-03 16:10:38 +02:00
Dai Ha 011ee80067 fleetd #656: add regression tests for the two cases #639's redact() fix covers
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Failing after 1m52s
Criterion 19 covers a block-scalar body whose key line falls outside
diff -u's default 3-line context (an 8-line body with only the 6th
line changed). Criterion 20 covers a blank line inside the value,
which used to reset the old indentation-anchored mask.

Both are RED against the pre-#639 redact() (git show 28ea0de) and
GREEN against the current one; each asserts both the secret's
absence and a non-secret control line's presence.
2026-10-03 16:02:50 +02:00
Dai Ha 3fab743152 Merge PR #655: fleetd #639 — mask a masked key's value by file line number, not by indentation anchor
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m4s
CI / build (push) Failing after 2m4s
2026-10-03 15:50:21 +02:00
Dai Ha 52eb9c2277 Merge PR #653: fleetd #641 — warn when --set reformats the whole fleetd.yaml 2026-10-03 15:47:30 +02:00
Dai Ha 6794fd8200 Merge PR #654: fleetd #642 — scope the herdr-control source-text guard and add vacuity anchors 2026-10-03 15:47:30 +02:00
Dai Ha 9b5c1cdcff Merge PR #652: fleetd #650 — scope FleetdAssemblyFleetAppTest's dead-end javadoc to loopback-trust 2026-10-03 15:46:42 +02:00
Dai Ha 3b69e0103b fleetd #639: redact block scalars by line number
CI / shell-tests (pull_request) Failing after 6s
CI / build (pull_request) Failing after 1m53s
CI / contract (pull_request) Successful in 1m21s
2026-10-03 15:46:20 +02:00
Dai Ha c97b1bba5a fleetd #641: warn on --set reformat churn
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m30s
2026-10-03 15:45:14 +02:00
Dai Ha a42253f597 fleetd #642: widen FleetdHerdrControlConstructionTest to cover FleetdAssembly.java
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Failing after 1m53s
Add a positive anchor per watched file (Fleetd.java and FleetdAssembly.java),
matching FleetdConfigRefWiringTest's [SOURCE TEXT] style, so a broken read
fails loudly instead of passing the negative check vacuously. Fix dangling
javadoc @link references in FleetdConfigRefWiringTest to the *AssemblyTest
names those classes were renamed to.
2026-10-03 15:45:14 +02:00
Dai Ha e69eafcc9f fleetd #650: scope the READ-unreachable javadoc to loopback-trust
CI / shell-tests (pull_request) Failing after 9s
CI / build (pull_request) Failing after 2m25s
CI / contract (pull_request) Successful in 2m34s
FleetdAssemblyFleetAppTest's class javadoc stated that /sessions'
Authz.Action.READ gate is always refused, and that READ always needs
Caller.resolved(). That is true only under auth.mode: loopback-trust,
the mode this test runs under because it configures no auth: block.
Under auth.mode: token, CallerResolver.resolve() returns before ever
consulting Caller.resolved()/scanComplete(), so a valid bearer token
resolves to PRIMARY with no pid lookup on that path. Names the two
tests that already exercise that path against a real assembly.
2026-10-03 15:40:52 +02:00
Dai Ha a42b12440c Merge PR #649: fleetd #612 Shape A r9+r11 — pin capacitySource, healthCoverageSource, coordinator peers via a real fleet_list round-trip
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 1m52s
Test-only, 243 lines, one new file. ZERO production change — this is the point
of the PR, and it replaces PR #647, which added three public accessors to
FleetMcp on a false premise.

#647 argued a same-JVM test caller can never reach fleet_list's READ gate,
because FleetdAssembly hardcodes a real LsofPeerPidLookup, so the caller
resolves ANONYMOUS. The premise about lsof is true; the conclusion is not. In
CallerResolver.resolve, once c.terminal() == null the next branch is:

    if (tokenMode) {
        return presentedTokenMatches(authorizationHeader)
                ? Principal.primary(c.pid()) : Principal.anonymous();
    }

c.resolved() and c.scanComplete() guard only the LATER loopback-trust branch,
which token mode returns before reaching. So under auth.mode: token a bearer
token resolves to PRIMARY with no pid lookup on the path. This PR does exactly
that: real McpSyncClient callTool("fleet_list") over a real transport with an
Authorization: Bearer header, a faked leadMailboxOpener so no broker is touched,
and a real coordinator: block with a non-empty peers list. No reflection.

Lead verification, measured myself (not taken from the worker's report):

  git diff --stat origin/main -- fleetd/src/main/java/   -> EMPTY
  mutate :481 Fleetd.capacitySource(...) -> CapacitySource.none()
      -> Tests run: 3, Failures: 1
         fleetListReportsTheAssembledCapacitySource:222
         the other two tests stayed GREEN

The failure message carries the live fleet_list JSON body, showing
healthCoverage, loopHealth and coordinator.peers present with capacity absent —
so the round-trip, the PRIMARY resolution and the coordinator visibility are all
real, and the mutation removed exactly one thing.

Worker also reported, each with a grep -c anchor of 1 restored: health mutation
-> 1 failure named fleetListReportsTheAssembledHealthCoverageSource; peers
mutation -> 1 failure named fleetListReportsTheAssembledCoordinatorPeers; final
mvn clean install Tests run: 1899, Failures: 0 / BUILD SUCCESS. I reproduced the
capacity cycle only; the other two are the worker's measurement, not mine.

Carries forward #647's genuine find: the peers ternary at :489 is
inert-equals-absent, so the test configures a real coordinator: block with peers.

Detail on #612 and #647.
2026-10-02 04:53:56 +02:00
Dai Ha a06426c33c Merge PR #648: fleetd #612 Shape A r4 — pin quarantineSource + outageSource through BOTH operator windows
Test-only, 368 lines, one new file. No production change.

Lead verification, measured myself in a throwaway detached worktree (not taken
from the worker's report):

  starve FleetApp pass site :529   -> Tests run: 4, Failures: 2
                                      (...ByTheRealAssembledFleetApp:333, :363)
                                      both ...FleetMcp tests stayed GREEN
  starve FleetMcp pass sites :484/:486 -> Tests run: 4, Failures: 2
                                      (...ByTheRealAssembledFleetMcp:319, :349)
                                      both ...FleetApp tests stayed GREEN

So the two windows are pinned independently. Failure messages carry the live
JSON response body from a real round-trip, not a source-text match.

Why this is worth merging — the REST window was completely unpinned. With this
PR's test parked and :529 starved, the FULL suite reported:

  Tests run: 1892, Failures: 0, Errors: 0, Skipped: 0 / BUILD SUCCESS

Zero pre-existing tests notice the REST window losing its sources. An operator
reads GET /profiles exactly when the MCP mount is down. Arithmetic control:
1892 + this PR's 4 = 1896 = main at 6539efe.

Known caveat, recorded not fixed: the MCP-side quarantine assertion (:313) is
DUPLICATE coverage. A reviewer measured on clean origin/main that starving the
MCP quarantine arg already fails three pre-existing tests
(FleetdBackendQuarantineAssemblyTest:167, FleetdExhaustedPatternAssemblyTest:202,
FleetdOpenCodeExhaustionForwardingAssemblyTest:182), so that site was already
pinned and the PR's claim otherwise is wrong. Low severity, left in place as a
valid fleet_profiles output assertion. Detail on #648 and #612.

Reviewed by two reviewers against the diff (soundness: no issue; coupling: the
duplicate-coverage finding above).
2026-10-02 04:53:39 +02:00
Dai Ha 28ea0de575 fleetd #612: cover fleet list reporting sources
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Failing after 2m10s
2026-10-02 04:42:07 +02:00
Dai Ha 91792e11fc fleetd #612 Shape A unit r4: pin quarantineSource + outageSource through both operator windows
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Failing after 1m56s
FleetdAssembly.java's quarantineSource (:471-472) and outageSource (:473-476)
each feed two consumers: FleetMcp (fleet_profiles, :484/:486) and FleetApp
(GET /profiles, :529). No existing test distinguished the two windows for
either source.

New test drives the real FleetdAssembly.assembleAndStart, classifies a real
exhaustion/outage through the real CompletionResolver, and reads the result
back through a real McpSyncClient (fleet_profiles) and a real HttpClient
(GET /profiles), both authenticated via token-mode auth (sidesteps the
in-JVM pid-resolution dead end). No source text is read; no production code
changed.

Verified with six mutation cycles (3 per site x 2 sites: FleetMcp starved,
FleetApp starved, mis-wire with a disconnected collaborator), each run
against the full unfiltered suite and reverted after confirming the
expected test(s) alone went red. The quarantineSource FleetMcp-starve and
mis-wire cycles also trip three pre-existing tests that read the live
BackendQuarantine via FleetMcp#quarantineSource() for their own unrelated
assertions - a pre-existing incidental coupling, not newly introduced here.

Out of scope, noted per the ticket's dispatch comment: loopHealthSource
(FleetdAssembly.java:478) shares this same two-consumer shape and is
already assigned to a separate unit, r10.
2026-10-02 04:20:53 +02:00
ltms 6539efe9fa Merge PR #645: fleetd #612 Shape A r10 — pin loopHealthSource at BOTH consumers
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m34s
CI / build (push) Failing after 2m11s
Pins FleetdAssembly's loopHealth local at both of its pass sites: FleetMcp (:483, the fleet_list source) and FleetApp (:529, the real /healthz body). Two independent assertions, so starving one site leaves the other green.

Lead verification, independent of the implementer's own proof, in a throwaway detached worktree at 68397f7:
- :483 starved only -> Tests run: 2, Failures: 1. RED: fleetListUsesTheRunningLoopsInTheRealAssembledMcp, expected: <RUNNING> but was: <STOPPED>. REST test stayed GREEN.
- :529 starved only -> Tests run: 2, Failures: 1. RED: healthzUsesTheRunningLoopsInTheRealAssembledApp, with the real body {"loopHealth":{"sessionReaper":"STOPPED","statusPoller":"STOPPED"}}. MCP test stayed GREEN.
- full suite with the PR: Tests run: 1896, Failures: 0, Errors: 0 — BUILD SUCCESS (1894 baseline + 2).

The REST half binds port 0 (ephemeral) via runtime.app().start("127.0.0.1", 0), not the configured 8765, so it cannot clash with the live daemon. Teardown stops the bound app, closes the runtime, and asserts the shutdown hook was registered and the herdr client actually closed.

I chased the implementer's honestly-reported anomaly (an unrelated ClaudeCodeLauncherTest.noFixtureSeededTheDefaultClaudeJson failure in one cycle, and a suite total that moved between cycles). It is NOT caused by this PR: a baseline run at 68397f7 with this PR absent added the same 3 temp-dir project entries to the real ~/.claude.json as the run with it applied (65->68 without, 68->71 with), and ClaudeCodeLauncherTest passed in both. Filed separately.

Test-only diff, no production code touched.
2026-10-02 04:15:04 +02:00
ltms 68397f78d5 Merge PR #644: fleetd #612 Shape A r12 — pin the assembled turn registrar
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 1m32s
CI / build (push) Failing after 1m57s
Pins FleetdAssembly.java:349 behaviourally. The test installs a throwing wrapper on the real assembled Injector's listener to construct the narrow between-delivery-and-completion window, then asserts runtime.completion() still resolves the registered waiter. No sleep: it advances a controllable nano clock. Teardown runs the captured shutdown hook and asserts herdr actually closed.

Lead verification, independent of the implementer's own proof, in a throwaway detached worktree at dac5f88:
- unmutated: Tests run: 1, Failures: 0 — BUILD SUCCESS
- :349 registrar -> (_, _) -> { }: Failures: 1 — "must wire the Injector registrar to this runtime's real CompletionResolver" expected: <true> but was: <false>
- mis-wire I built myself (differs from the implementer's): Fleetd.turnRegistrar(new CompletionResolver(agents, new Rendezvous(), ...)) — a fresh Rendezvous so the registry genuinely differs: Failures: 1, same assertion
- reverted, full suite: Tests run: 1894, Failures: 0, Errors: 0 — BUILD SUCCESS, 58s (1892 baseline + r5 + r12)

Test-only diff, no production code touched. Reflection is used only to install the throwing listener; that is how the failure window is constructed, not how the assertion is made. No reviewer fan-out: member capacity is committed to the remaining Shape A implementers.
2026-10-02 04:09:29 +02:00
Dai Ha af901ff1d2 fleetd #612: pin assembled loop health
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 2m23s
2026-10-02 04:08:02 +02:00
ltms dac5f88812 Merge PR #643: fleetd #612 Shape A r5 — pin FleetdAssembly's leadConfigDirSource call site
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m14s
Pins the #602/#606 call site behaviourally: drives the real FleetdAssembly.assembleAndStart and asserts the assembled LeadConfigDirSource resolves a real configured configDir, which none() cannot produce.

Lead verification, run independently of the implementer's own proof, in a throwaway detached worktree at 141ae3b:
- unmutated: Tests run: 1, Failures: 0 — BUILD SUCCESS
- FleetdAssembly.java:488 -> FleetMcp.LeadConfigDirSource.none(): Tests run: 1, Failures: 1 — expected: </mnt/fake-lead-configdir> but was: <null>
- reverted, full suite: Tests run: 1893, Failures: 0, Errors: 0 — BUILD SUCCESS, 59s (baseline 1892)

Test-only diff, no production code touched. No reviewer fan-out was run: all member capacity is committed to the five Shape A implementers.
2026-10-02 04:05:50 +02:00
Dai Ha 2e663e5968 fleetd #612: pin assembled turn registrar
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Failing after 1m51s
2026-10-02 04:04:54 +02:00
Dai Ha 8c14ed2846 fleetd #612 Shape A r5: pin FleetdAssembly's leadConfigDirSource call site
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 1m0s
CI / build (pull_request) Failing after 2m4s
FleetdLeadConfigDirSourceWiringTest already pins Fleetd.leadConfigDirSource
itself, but by its own javadoc cannot cover whether FleetdAssembly.java:488
still calls it -- that call site could be swapped for a bare
FleetMcp.LeadConfigDirSource.none() (the literal fleetd #602/#606 defect)
and the whole suite would stay green.

Add FleetdLeadConfigDirSourceAssemblyTest: assembles the real FleetdRuntime
via FleetdAssembly.assembleAndStart, reads the leadConfigDirs field off the
real FleetMcp via reflection (no public accessor exists), and asserts it
resolves a configured lead's real configDir rather than none()'s hardcoded
null.

Verified: loud control (flip expected value) goes RED, reverts green;
mutation (i) inert none() at the call site goes RED; mutation (ii) mis-wire
(empty profile map, symbols otherwise intact) goes RED; both mutations
revert to an empty git diff. Full mvn clean install: 1893 tests, 0
failures, 0 errors, BUILD SUCCESS.
2026-10-02 04:02:18 +02:00
ltms 141ae3b04d Merge PR #640: fleetd #639 — redact()'s comment claimed more than the code does
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m32s
CI / build (push) Failing after 2m29s
Comment text only, no behaviour change. Corrects the paragraph added in d7f94ca so it no longer
claims the continuation masking holds for any masked key "present or future". It holds only while
the masked key's own line is inside the printed hunk; diff -u's three lines of context routinely
leave it out, and a blank line inside a block scalar drops the anchor too.

Verified: suite green in a clean copy of the edited tree (16 criteria + 3 extras, exit 0), bash -n
clean, and a control on the edit itself (old claim gone, #639 reference present). Evidence in #639.
2026-10-01 19:00:56 +02:00
Dai Ha faefea14c4 fleetd #639: redact()'s comment claimed more than the code does
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Failing after 2m11s
The paragraph added in d7f94ca ended with "This needs no knowledge of the key's
name and so protects a block scalar under any masked key, present or future."
The continuation masking is real and it is an improvement, but that sentence is
too strong: the masking only holds while the masked key's own line is inside the
hunk being printed.

redact() is fed `diff -u` output, which prints three lines of context. A block
scalar's body therefore often arrives with its key line left out. With no key
line, `masked` is never set and the body prints in full, with no "<redacted>"
anywhere. A blank line inside a block scalar loses the anchor the same way: a
blank diff line measures as indent 0, so `indent > masked_indent` is false and
the mask ends early — this time directly under a "<redacted>" marker.

Both were reproduced through the real script with --dry-run, each with a
positive control run first to prove the secret's lines actually reached the
output (without that control, "the secret never entered the diff" and "it
entered and was redacted" are indistinguishable). Filed as fleetd #639, which
also records that this is latent rather than live: today's fleetd.yaml holds 5
block scalars and all 5 sit under non-secret keys.

Comment text only. No change to redact() or to any other function, and no
change to the test suite. scripts/test-config-edit.sh still passes in a clean
copy (16 criteria + 3 extras, exit 0); bash -n clean.

The reason this is worth its own commit: #635 exists because an incomplete
redactor that looks complete is worse than one that visibly does nothing. A
comment that overstates the guarantee is the same defect in prose, and the next
session to read it has no other source.
2026-10-01 19:00:29 +02:00
ltms ea6896f2ef Merge PR #636: fleetd #635 — config-edit.sh, the one auditable way to edit fleetd.yaml
CI / shell-tests (push) Failing after 14s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m5s
Closes fleetd #635.

Verified by the lead before merge:
- scripts/test-config-edit.sh run in a PRISTINE copy of d7f94ca (git archive + git init, no
  worktree, no daemon): 16 criteria + 3 extras, exit 0.
- Diff read in full. 4 files, 1325 additions, 0 deletions, no .java/pom.xml/.yaml (checked with a
  positive control, so the negative is real) — no Maven gate needed.
- Defects 1-6 were verified by the previous lead by mutation; defects 7 and 8 (criteria 15a/15b/16)
  are fixed in d7f94ca and the fix was re-measured here independently.

Known limitation, filed as #639 and NOT a regression: redact()'s continuation masking only holds
while the masked key line is itself inside the printed diff hunk. diff -u prints three lines of
context, so a block scalar's body can appear without its key, and then nothing is masked; a blank
line inside a block scalar loses the anchor the same way. Reproduced through the real script with
a positive control proving the body reached the output. Latent, not live: today's fleetd.yaml has
5 block scalars and all 5 sit under non-secret keys. The pre-fix code leaked these cases too, so
this commit is a strict improvement.

Two sentences in config-edit.sh's redact() comment and in this PR's body claim the masking holds
regardless of the key's name. That is too strong; see #639. Being corrected in a follow-up.
2026-10-01 18:58:27 +02:00
42 changed files with 3222 additions and 273 deletions
+1 -1
View File
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
PIDS="$(pgrep -f 'fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
+16 -7
View File
@@ -22,13 +22,22 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
early. The script still checks the file is there and dies with `no jar at … — run without
--no-build` if it is not, but it cannot tell you the jar is old.
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
main clone harmless now: neither can reach the file the running daemon holds open, because that
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
runtime jar — is visible before you decide anything.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
+7
View File
@@ -323,6 +323,13 @@ must obey belongs in the charter, not here.
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
throwaway git worktree anyway: a build in the main clone still ships nothing until
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
### Redeploying the daemon — the lead may do this (primary only)
+1 -1
View File
@@ -47,7 +47,7 @@
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
<string>fleetd.yaml</string>
</array>
+1 -1
View File
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
+4
View File
@@ -2,6 +2,10 @@
target/
dependency-reduced-pom.xml
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
run/
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
fleetd.yaml
+4 -2
View File
@@ -189,8 +189,10 @@
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
fleetd/run/fleetd.jar, not this plugin's own output path — see
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
armed, so this name, the plist, and the wrapper must still move as one. -->
<finalName>fleetd</finalName>
<plugins>
<plugin>
@@ -1075,7 +1075,33 @@ public final class Fleetd {
*/
static FleetMcp.LeadConfigDirSource leadConfigDirSource(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders) {
return new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(profiles, leaders));
return new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(profiles, leaders),
leadContextWindowLookup(profiles, leaders));
}
/**
* Per-lead-name factory for the effective auto-compact window {@link LeadContextGauge} scales
* its HIGH threshold against — the same {@code fleet.leaders.<name>.profile} link {@link
* #leadConfigDirLookup} already follows, one step further to that profile's own {@link
* FleetConfig.Profile#effectiveAutoCompactWindow()}. A lead entry that names no
* {@code profile:}, or whose named profile is not configured, or whose profile resolves no
* window at all, returns {@code null} — {@link LeadContextGauge} then falls back to its own
* fixed HIGH threshold.
*/
static Function<String, Long> leadContextWindowLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders) {
return leadName -> {
FleetConfig.Leader lead = leaders.get(leadName);
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
return null;
}
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
if (leadProfile == null) {
return null;
}
Integer window = leadProfile.effectiveAutoCompactWindow();
return window == null ? null : window.longValue();
};
}
/**
@@ -1097,22 +1123,30 @@ public final class Fleetd {
* @param liveLeadTerminals terminal id → lead name for every CURRENTLY recognised lead
* @param configDirForLeadName lead name → {@code configDir}, normally {@link
* #leadConfigDirLookup}'s return
* @param windowForLeadName lead name → that lead's profile's effective auto-compact window,
* normally {@link #leadContextWindowLookup}'s return, and passed
* through to {@link LeadContextGauge#read} so the heartbeat's own
* HIGH reading scales with that lead's real window, or {@code null}
* when it cannot be resolved — either way {@link LeadContextGauge}
* falls back to its own fixed HIGH threshold
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
Function<String, Long> windowForLeadName) {
return terminal -> {
String leadName = liveLeadTerminals.get().get(terminal);
if (leadName == null) {
return LeadContextGauge.Reading.unknown();
}
String configDir = configDirForLeadName.apply(leadName);
Long effectiveWindow = windowForLeadName.apply(leadName);
Agent live;
try {
live = agents.get(terminal);
} catch (RuntimeException e) {
return LeadContextGauge.Reading.unknown();
}
return gauge.read(configDir, live.sessionId(), live.agentType());
return gauge.read(configDir, live.sessionId(), live.agentType(), effectiveWindow);
};
}
@@ -1124,9 +1158,10 @@ public final class Fleetd {
* (see {@code FleetdLeadConfigDirSourceWiringTest}'s javadoc for the measured gap this shape closes).
*/
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
Function<String, Long> windowForLeadName) {
return new LeadHeartbeatLoop.LeadContextSource(
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName));
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName, windowForLeadName));
}
/**
@@ -405,7 +405,8 @@ final class FleetdAssembly {
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics,
Fleetd.leadContextSource(leadContextGauge, router.leadAgents(), leads,
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders)),
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders),
Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)),
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
heartbeat.start();
} else {
@@ -798,6 +798,25 @@ public record FleetConfig(
return isSubscription() ? SUBSCRIPTION_CREDENTIAL_ID : profile;
}
/**
* The auto-compaction window a launched Claude Code session actually runs on: {@code env:
* CLAUDE_CODE_AUTO_COMPACT_WINDOW} when it parses as an integer, since that environment
* variable wins over the {@code --autocompact} flag {@link #autoCompactWindow} produces (see
* {@code ClaudeCodeArguments}); {@link #autoCompactWindow} otherwise. {@code null} when
* neither resolves to a usable number.
*/
public Integer effectiveAutoCompactWindow() {
String envValue = env.get(CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV);
if (envValue == null) {
return autoCompactWindow;
}
try {
return Integer.valueOf(envValue.trim());
} catch (NumberFormatException e) {
return autoCompactWindow;
}
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
@@ -2184,6 +2203,12 @@ public record FleetConfig(
static final int AUTO_COMPACT_WINDOW_MIN = 100_000;
/** Highest {@code autoCompactWindow} Claude Code's {@code --autocompact <tokens>} flag accepts. */
static final int AUTO_COMPACT_WINDOW_MAX = 1_000_000;
/**
* The {@code env:} key a launched Claude Code session reads for its auto-compaction window,
* ahead of the {@code --autocompact} launch flag {@code autoCompactWindow} produces (see
* {@link Profile#effectiveAutoCompactWindow()}).
*/
static final String CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV = "CLAUDE_CODE_AUTO_COMPACT_WINDOW";
/**
* Reject a profile whose {@code autoCompactWindow:} is set but outside the token band Claude
@@ -2265,13 +2290,13 @@ public record FleetConfig(
if (!(entry.getValue() instanceof Map<?, ?> profile)
|| !(profile.get("autoCompactWindow") instanceof Number window)
|| !(profile.get("env") instanceof Map<?, ?> env)
|| !env.containsKey("CLAUDE_CODE_AUTO_COMPACT_WINDOW")) {
|| !env.containsKey(CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV)) {
continue;
}
Object kind = profile.get("kind");
boolean claudeCode = kind == null || String.valueOf(kind).isBlank()
|| Profile.KIND_CLAUDE_CODE.equalsIgnoreCase(String.valueOf(kind));
Object envValue = env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW");
Object envValue = env.get(CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV);
if (claudeCode && !String.valueOf(window).equals(String.valueOf(envValue))) {
String name = String.valueOf(entry.getKey());
names.add(name);
@@ -2663,9 +2688,10 @@ public record FleetConfig(
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
* a member template matches and the daemon starts labelling its own members as leads, promoting
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
* realistic path, but defence that depends on one workspace label holding is not defence enough
* for a privilege boundary.
* {@link #validatePanePlacementAgainstLeadTabs()} is the check that stops a pane-placed member
* from landing inside a lead's tab in the first place; this check is a second, independent
* guard that catches the hazard even when every profile places members correctly, by refusing
* a label that a scan would still misread as a lead.
*
* <p>CB-557 shrank this check rather than removing it. The default template is
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
@@ -2712,6 +2738,45 @@ public record FleetConfig(
+ "lead tabs cannot be confused.");
}
/**
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
* entry names a {@code tab}. A pane-placed member lands inside the focused tab rather than its
* own, so it can land inside a lead's own labelled tab. {@link
* dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead purely by that tab's label — it does
* not exclude the member space — so a member that ends up there would be read back as the lead
* and granted spawn/stop/send on the whole fleet.
*
* <p>Only a leader with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
*
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
* {@code fleet.leaders} entry names a non-blank {@code tab}
*/
public void validatePanePlacementAgainstLeadTabs() {
if (fleet == null || fleet.leaders().isEmpty()) {
return;
}
boolean anyLeaderHasTab = fleet.leaders().values().stream()
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
if (!anyLeaderHasTab) {
return;
}
List<String> bad = new ArrayList<>();
profiles().entrySet().stream()
.filter(e -> !e.getValue().tabPlacement())
.map(Map.Entry::getKey)
.sorted()
.forEach(bad::add);
if (bad.isEmpty()) {
return;
}
throw new IllegalStateException("refusing to start: profile(s) " + bad
+ " use placement: pane while fleet.leaders names a tab. A pane-placed member can "
+ "land inside a lead's labelled tab and be read back as the lead, granted "
+ "spawn/stop/send on the whole fleet. Set placement: tab for each named profile, "
+ "or remove the tab from every fleet.leaders entry.");
}
/**
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
@@ -2912,11 +2977,11 @@ public record FleetConfig(
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each of the six validators above was well pinned on its own, nothing proved
* that although each validator above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* invoked it — deleting a call site left the full suite green. The fix is not one more test
* per caller; a hand-maintained list of names here would have the exact same defect its
* own javadoc would warn against: the next validator someone adds has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
@@ -2925,7 +2990,7 @@ public record FleetConfig(
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* the six individually — see the comments at those two call sites for why
* each validator individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
@@ -2942,9 +3007,9 @@ public record FleetConfig(
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* would sweep any new {@code validateXxx()} method added to any class, not just something
* special-cased to the set {@link FleetConfig} declares today) without needing to add a real,
* unwanted extra validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
@@ -37,8 +37,11 @@ import java.util.function.Supplier;
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>{@code excludedWorkspaceLabels} can filter a workspace out of the scan, but this class does
* not by itself stop a worker from landing inside a matching tab — a caller may pass an empty
* set, and the daemon does. The guard against that is {@code
* FleetConfig.validatePanePlacementAgainstLeadTabs}: it refuses, at startup, any profile that
* places members by pane while a lead names a tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
@@ -68,8 +68,10 @@ import java.util.function.LongSupplier;
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
* {@code fleet_list} call reports on.
* keyed by {@code (configDir, sessionId, highThreshold)}, so it is safe to share across every lead
* a single {@code fleet_list} call reports on, and a call that resolves a different effective
* window for the same lead never reads back a state computed against the other window's
* threshold.
*/
public final class LeadContextGauge {
@@ -97,14 +99,19 @@ public final class LeadContextGauge {
static final long DEFAULT_CACHE_TTL_MILLIS = 5_000;
/**
* Live tokens at or above this count report {@link State#HIGH}. On the host this was measured
* on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH
* state is to warn before that happens, not at it — 200,000 is the standard Claude context
* window size and a sensible built-in default: no config key is required to pick it, and a
* lead crossing it is already deep enough into its window that a compaction is foreseeable.
* Fallback HIGH threshold used when a caller resolves no effective auto-compact window for the
* lead being read (see {@link #read(String, String, String, Long)}) — the built-in default so
* no config key is required to get a warning at all.
*/
static final long HIGH_THRESHOLD_TOKENS = 200_000;
/**
* The fraction of a resolved effective auto-compact window that HIGH warns at, so the warning
* margin scales with the window instead of only ever meaning something against the fixed
* {@link #HIGH_THRESHOLD_TOKENS} fallback.
*/
static final double HIGH_THRESHOLD_FRACTION = 2.0 / 3.0;
/** The only peer kind this reader understands ({@code Agent.agentType()}'s wire value). */
private static final String CLAUDE_AGENT_TYPE = "claude";
@@ -165,8 +172,13 @@ public final class LeadContextGauge {
* {@code "claude"} (including {@code null}, meaning undetected) reports
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
* transcript format
* @param effectiveWindowTokens the caller's resolved effective auto-compact window for this
* lead's own profile, or {@code null} when it cannot be resolved.
* HIGH fires at {@link #HIGH_THRESHOLD_FRACTION} of this value;
* {@code null} (or a non-positive value) falls back to the fixed
* {@link #HIGH_THRESHOLD_TOKENS}
*/
public Reading read(String configDir, String sessionId, String agentType) {
public Reading read(String configDir, String sessionId, String agentType, Long effectiveWindowTokens) {
if (sessionId == null || sessionId.isBlank()) {
return Reading.unknown();
}
@@ -176,18 +188,27 @@ public final class LeadContextGauge {
String base = (configDir == null || configDir.isBlank())
? System.getProperty("user.home") + "/.claude"
: configDir;
String cacheKey = base + '\u0000' + sessionId;
long highThreshold = highThreshold(effectiveWindowTokens);
String cacheKey = base + '\u0000' + sessionId + '\u0000' + highThreshold;
long now = clock.getAsLong();
CacheEntry cached = cache.get(cacheKey);
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
return cached.reading();
}
Reading fresh = readUncached(base, sessionId);
Reading fresh = readUncached(base, sessionId, highThreshold);
cache.put(cacheKey, new CacheEntry(fresh, now));
return fresh;
}
private Reading readUncached(String base, String sessionId) {
/** {@link #HIGH_THRESHOLD_FRACTION} of {@code effectiveWindowTokens}, or the fixed fallback. */
private static long highThreshold(Long effectiveWindowTokens) {
if (effectiveWindowTokens == null || effectiveWindowTokens <= 0) {
return HIGH_THRESHOLD_TOKENS;
}
return (long) (effectiveWindowTokens * HIGH_THRESHOLD_FRACTION);
}
private Reading readUncached(String base, String sessionId, long highThreshold) {
diskReads.incrementAndGet();
Path file = findTranscript(base, sessionId);
if (file == null) {
@@ -199,7 +220,7 @@ public final class LeadContextGauge {
} catch (IOException e) {
return Reading.unknown();
}
return parse(tail);
return parse(tail, highThreshold);
}
/**
@@ -258,7 +279,7 @@ public final class LeadContextGauge {
* read raced — see the "torn final line" section of the class javadoc) is skipped, not fatal.
* Only when none of the remaining lines parse does this report {@link State#UNKNOWN}.
*/
private Reading parse(TailRead tail) {
private Reading parse(TailRead tail, long highThreshold) {
String text = new String(tail.bytes(), StandardCharsets.UTF_8);
List<String> lines = new ArrayList<>(List.of(text.split("\n", -1)));
if (!lines.isEmpty() && lines.get(lines.size() - 1).isEmpty()) {
@@ -299,7 +320,7 @@ public final class LeadContextGauge {
if (tokens == null) {
return new Reading(State.UNKNOWN, null, compactions);
}
State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
State state = tokens >= highThreshold ? State.HIGH : State.OK;
return new Reading(state, tokens, compactions);
}
@@ -201,8 +201,8 @@ public final class LeadLauncher {
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
* tab carries a different label, not because any workspace is excluded from this count.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
@@ -284,10 +284,15 @@ public final class FleetMcp {
* no profile, or that profile sets no {@code configDir} override — either
* way {@link LeadContextGauge} then falls back to its own built-in default,
* exactly as before this ticket
* @param windowFor lead name → that lead's profile's effective auto-compact window (see
* {@code dev.ltms.fleet.config.FleetConfig.Profile
* #effectiveAutoCompactWindow()}), or {@code null} when it cannot be
* resolved — either way {@link LeadContextGauge} falls back to its own
* fixed HIGH threshold
*/
public record LeadConfigDirSource(Function<String, String> configDirFor) {
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in default {@code configDir}. */
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null); }
public record LeadConfigDirSource(Function<String, String> configDirFor, Function<String, Long> windowFor) {
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in defaults. */
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null, _ -> null); }
}
/**
@@ -2153,7 +2158,8 @@ public final class FleetMcp {
m.put("self", true);
}
String configDir = leadConfigDirs.configDirFor().apply(name);
m.put("context", contextView(contextGauge, live, configDir));
Long effectiveWindow = leadConfigDirs.windowFor().apply(name);
m.put("context", contextView(contextGauge, live, configDir, effectiveWindow));
return m;
}
@@ -2163,12 +2169,15 @@ public final class FleetMcp {
* {@code fleet.leaders.<name>.profile} → that profile's own {@code configDir:} — or {@code null}
* when the lead's entry names no profile, or that profile sets no override, in which case
* {@link LeadContextGauge#read} falls back to its own built-in default
* ({@code <user.home>/.claude}).
* ({@code <user.home>/.claude}). {@code effectiveWindowTokens} is the same lead's resolved
* auto-compact window, or {@code null} when it cannot be resolved, in which case the gauge
* falls back to its own fixed HIGH threshold instead.
*/
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir) {
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir,
Long effectiveWindowTokens) {
String sessionId = live == null ? null : live.sessionId();
String agentType = live == null ? null : live.agentType();
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType);
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType, effectiveWindowTokens);
Map<String, Object> c = new LinkedHashMap<>();
c.put("state", reading.state().name().toLowerCase());
if (reading.tokens() != null) {
@@ -461,9 +461,9 @@ public final class LeadHeartbeatLoop {
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
} else {
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
// lives in another class and nothing asserts it, so this branch does not rely on it —
// it drops the token clause rather than printing "null tokens".
// reaches HIGH by comparing a number against a threshold. That invariant lives in
// another class and nothing asserts it, so this branch does not rely on it — it drops
// the token clause rather than printing "null tokens".
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
.append(" so far).");
}
@@ -0,0 +1,166 @@
package dev.ltms.fleet;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Method;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Asserts that the {@link FleetMcp} built by {@link FleetdAssembly#assembleAndStart} applies the
* authorization table: a worker is refused {@code SPAWN}, and the primary is allowed it.
*/
class FleetdAssemblyAuthorizationModeTest {
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Binding a real port would clash with any daemon already listening on it.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private TestResourcePorts ports;
@AfterEach
void tearDown() {
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""");
return FleetConfig.load(file);
}
private FleetMcp assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
return runtime.mcp();
}
/**
* Invokes {@code FleetMcp#denyFor}, which is package-private to {@code dev.ltms.fleet.mcp}
* while this test is in {@code dev.ltms.fleet}. Nothing here catches a missing method: if
* {@code denyFor} is renamed or removed, {@link NoSuchMethodException} propagates and the
* test fails.
*/
private static McpSchema.CallToolResult denyFor(FleetMcp mcp, Principal caller, Authz.Action action,
String target) throws Exception {
Method m = FleetMcp.class.getDeclaredMethod("denyFor", Principal.class, Authz.Action.class, String.class);
m.setAccessible(true);
return (McpSchema.CallToolResult) m.invoke(mcp, caller, action, target);
}
@Test
void productionBootPathRefusesAnUnauthorizedCallerThroughTheAssembledFleetMcp(@TempDir Path dir)
throws Exception {
FleetMcp mcp = assemble(dir);
McpSchema.CallToolResult deniedForWorker = denyFor(mcp, Principal.worker("term_a", 200),
Authz.Action.SPAWN, "term_a");
assertNotNull(deniedForWorker,
"a worker must not be able to fleet_spawn through the assembled FleetMcp");
assertTrue(deniedForWorker.isError(), "a refusal is returned as an MCP tool error");
McpSchema.CallToolResult allowedForPrimary = denyFor(mcp, Principal.primary(100),
Authz.Action.SPAWN, "term_a");
assertNull(allowedForPrimary,
"control: the primary must still be allowed to fleet_spawn — otherwise the worker "
+ "refusal above would pass even with the gate wired backwards");
}
}
@@ -58,20 +58,31 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
* {@code HttpClient} — no accessor needed for this half.
*
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
* not pin the merge half of the deleted test's javadoc. {@code /sessions} requires
* {@code Authz.Action.READ}, which — through the REAL assembly's real {@code
* CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded {@code
* new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid from
* {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a JUnit
* test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is always
* {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's fail-closed
* rule) before the route handler — and its {@code memberHerdr} merge — is ever reached. Verified
* directly: driving {@code GET /sessions} here returns {@code 401 unauthenticated}, not the
* merged body. {@code FleetAppTwoDaemonTest} avoids this because it builds {@code FleetApp} with
* {@code callers: null}, which is not what the real assembly passes. The {@code /healthz} pin
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
* {@code /sessions} correctly once handed two clients.
* not pin the merge half of the deleted test's javadoc. This class configures no {@code auth:}
* block, so it runs under the default {@code loopback-trust} mode ({@code FleetConfig}). Under
* that mode, {@code /sessions} requires {@code Authz.Action.READ}, which — through the REAL
* assembly's real {@code CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded
* {@code new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid
* from {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a
* JUnit test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is
* always {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's
* fail-closed rule) before the route handler — and its {@code memberHerdr} merge — is ever
* reached. Verified directly: driving {@code GET /sessions} here returns {@code 401
* unauthenticated}, not the merged body. {@code FleetAppTwoDaemonTest} avoids this because it
* builds {@code FleetApp} with {@code callers: null}, which is not what the real assembly
* passes. The {@code /healthz} pin below is what this class relies on for CB-185's {@code
* FleetApp} half; {@code FleetAppTwoDaemonTest} remains the full behavioural proof that
* {@code FleetApp} itself merges {@code /sessions} correctly once handed two clients.
*
* <p><strong>This refusal is {@code loopback-trust}-specific, not a property of {@code
* CallerResolver} in general.</strong> Under {@code auth.mode: token}, {@code
* CallerResolver#resolve} returns before ever consulting {@code Caller.resolved()} or {@code
* Caller.scanComplete()}: a request carrying a valid bearer token in its {@code Authorization}
* header resolves to {@code Role#PRIMARY} with no pid lookup at all, so the same-JVM-pid
* exclusion above never comes into play. {@code FleetdQuarantineOutageDualWindowAssemblyTest}
* and {@code FleetdListReportingSourcesAssemblyTest} both drive {@code Authz.Action.READ} this
* way, over a real {@code McpSyncClient}/{@code HttpClient} against a real {@code
* FleetdAssembly#assembleAndStart}, and both get the real response rather than a refusal.
*
* <p><strong>fleetd #629 follow-up.</strong> The fix below (see {@link TwoHerdrResourcePorts})
* makes {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}'s fake {@code
@@ -0,0 +1,201 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
* production boot path passes to {@link LeadTabScanner} at {@code FleetdAssembly.java:265}
* ({@code Set.of()}).
*
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
* producer. This test instead reaches the exact object {@link FleetdAssembly#assembleAndStart}
* builds: a {@code fleet.leaders:} block makes the assembly construct a real
* {@link LeadTabScanner} for its local {@code leads} supplier, and a {@code coordinator:} block
* makes it hand that same supplier instance to {@link LeadCoordLoop} (fleetd #637), which stores
* it as a field. Reflection recovers it from there, and then from the scanner itself, so the
* assertion is against the real production argument rather than a copy built for this test.
*/
class FleetdAssemblyLeadTabScannerExclusionTest {
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
fleet:
leaders:
primary:
tab: "lead: primary"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(file);
}
@Test
void productionBootPathPassesNoExcludedWorkspaceLabels(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
TestResourcePorts ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
LeadCoordLoop coordLoop = runtime.leadCoordLoop();
assertNotNull(coordLoop, "control: a configured coordinator: block must build LeadCoordLoop");
Field leadsField = LeadCoordLoop.class.getDeclaredField("leads");
leadsField.setAccessible(true);
@SuppressWarnings("unchecked")
Supplier<Map<String, String>> leads = (Supplier<Map<String, String>>) leadsField.get(coordLoop);
assertInstanceOf(LeadTabScanner.class, leads,
"control: a non-empty fleet.leaders: block must make FleetdAssembly build a real "
+ "LeadTabScanner for its `leads` supplier, not the Map::of fallback — "
+ "otherwise this test would pass for the wrong reason");
Field excludedField = LeadTabScanner.class.getDeclaredField("excludedWorkspaceLabels");
excludedField.setAccessible(true);
Set<?> excluded = (Set<?>) excludedField.get(leads);
assertTrue(excluded.isEmpty(),
"FleetdAssembly.java:265 must pass an empty excludedWorkspaceLabels to "
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
}
}
}
@@ -0,0 +1,180 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.mcp.FleetMcp;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Shape A, rank 10: {@link FleetdAssembly} creates one {@link FleetMcp.LoopHealthSource}
* from the real started {@code StatusPoller} and {@code SessionReaper}, then gives it to two operator
* windows. These tests reach the real assembled objects through {@link FleetdRuntime}, rather than
* building a second source beside them. A hardcoded {@code RUNNING} source would pass a simple
* "running" test, so the mutation proof also mis-wires the source to never-started loops: both
* windows must then report {@code STOPPED} and these assertions go red.
*/
class FleetdAssemblyLoopHealthTest {
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("no coordinator is configured");
};
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("healthy FakeHerdr must not be polled");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Bind runtime.app() to an ephemeral port only in the REST assertion below.
}
}
private final HttpClient http = HttpClient.newHttpClient();
private RecordingResourcePorts ports;
private FleetdRuntime runtime;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (runtime != null) {
runtime.close();
}
if (ports != null) {
assertNotNull(ports.shutdownHook, "the assembly must register its shutdown hook");
assertTrue(ports.herdr.closed, "FleetdRuntime.close must close the real assembled herdr client");
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
lifecycle:
idleTtlSeconds: 600
idleSleepGuard:
enabled: false
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""");
return FleetConfig.load(file);
}
private void assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new RecordingResourcePorts();
runtime = FleetdAssembly.assembleAndStart(
new AssemblyInputs(cfg, new ConfigRef(dir.resolve("fleetd.yaml"), cfg),
new SubscriptionGuard(cfg.guard().hostSet())), ports);
}
@Test
void fleetListUsesTheRunningLoopsInTheRealAssembledMcp(@TempDir Path dir) throws Exception {
assemble(dir);
// The field is the exact source captured by FleetMcp's fleet_list handler. Reflection is
// necessary because FleetMcp has no public source accessor; it is not a source-text check.
FleetMcp.LoopHealthSource loopHealth = loopHealthOf(runtime.mcp());
assertRunning(loopHealth, "FleetMcp's real fleet_list source");
}
@Test
void healthzUsesTheRunningLoopsInTheRealAssembledApp(@TempDir Path dir) throws Exception {
assemble(dir);
boundApp = runtime.app().start("127.0.0.1", 0);
HttpRequest request = HttpRequest.newBuilder(
URI.create("http://127.0.0.1:" + boundApp.port() + "/healthz")).GET().build();
HttpResponse<String> response = http.send(request, HttpResponse.BodyHandlers.ofString());
assertEquals(200, response.statusCode(), response.body());
assertTrue(response.body().contains("\"statusPoller\":\"RUNNING\""), response.body());
assertTrue(response.body().contains("\"sessionReaper\":\"RUNNING\""), response.body());
}
private static FleetMcp.LoopHealthSource loopHealthOf(FleetMcp mcp) throws Exception {
Field field = FleetMcp.class.getDeclaredField("loopHealth");
field.setAccessible(true);
return (FleetMcp.LoopHealthSource) field.get(mcp);
}
private static void assertRunning(FleetMcp.LoopHealthSource source, String consumer) {
assertEquals(LoopWatchdog.State.RUNNING, source.statusPoller().get(),
consumer + " must report the started real StatusPoller as RUNNING");
assertEquals(LoopWatchdog.State.RUNNING, source.sessionReaper().get(),
consumer + " must report the started real SessionReaper as RUNNING");
}
}
@@ -0,0 +1,184 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 rank 12 — {@code FleetdAssembly} supplies the {@link Injector}'s {@code TurnRegistrar}
* with {@code Fleetd.turnRegistrar(completion)}. The normal delivery path cannot distinguish that
* registrar from {@code TurnRegistrar.NOOP}: {@link dev.ltms.fleet.inject.CompletionResolver#onDelivered}
* registers the same turn shortly afterwards. The distinction matters when a delivery listener throws
* after the pane received the message but before normal completion runs. The registrar must already have
* registered the waiter with the REAL assembled resolver, so a later completion can still resolve it.
*
* <p>The test installs a throwing wrapper around the real assembled listener. It does not replace the
* registrar or resolver. This constructs the narrow failure condition without a sleep, then observes the
* real resolver through {@link FleetdRuntime#completion()}. An inert registrar, or a registrar wired to a
* throwaway resolver, leaves this waiter's turn absent from the real resolver and makes the final assertion
* fail.
*/
class FleetdAssemblyTurnRegistrarBehaviouralTest {
private static final String TARGET = "term_a";
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L);
Runnable shutdownHook;
void advanceSeconds(long seconds) {
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
}
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> {
throw new UnsupportedOperationException("no broker: block is configured");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
fleet:
leaders:
primary:
tab: "lead: primary"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(file);
}
@Test
void realAssembledResolverStillResolvesAfterDeliveredListenerThrows(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ControllableResourcePorts ports = new ControllableResourcePorts();
ports.herdr.withTab("w2", "w2:t7", "lead: primary");
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
Injector injector = runtime.injector();
installThrowingDeliveredListener(injector);
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
injector.enqueue(TARGET, "brief", new TurnToken(TARGET, waiter));
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> injector.onStatus(TARGET, AgentStatus.IDLE),
"control: delivery must reach the installed listener and it must throw after delivery");
assertTrue(thrown.getMessage().contains("listener failure"), thrown::getMessage);
ports.herdr.readText("worker report after the listener failure");
ports.advanceSeconds(3);
runtime.completion().resolveBeforePostAction(TARGET);
assertTrue(waiter.isDone(),
"FleetdAssembly must wire the Injector registrar to this runtime's real CompletionResolver: "
+ "after a delivered listener throws, resolveBeforePostAction must still find and "
+ "resolve the registered waiter");
} finally {
assertTrue(ports.shutdownHook != null, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
assertTrue(ports.herdr.closed, "teardown control: the captured shutdown hook must close herdr");
}
}
private static void installThrowingDeliveredListener(Injector injector) throws Exception {
Field field = Injector.class.getDeclaredField("turnListener");
field.setAccessible(true);
field.set(injector, new TurnListener() {
@Override
public void onTurnComplete(String target) {
}
@Override
public void onDelivered(String target, TurnToken token) {
throw new IllegalStateException("listener failure after delivery");
}
});
}
}
@@ -24,8 +24,8 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* {@link FleetdBackendQuarantineAssemblyTest}, {@link FleetdLeadSeatAssemblyTest} and {@link
* FleetdCompletionResolverAssemblyTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
@@ -2,16 +2,67 @@ package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* {@code AgentControl} caches {@code paneByTerminal}, so {@code HerdrRouter} must be its only
* production factory — a second instance means a second cache; the same reasoning applies to
* {@code WorkspaceControl}. {@code HerdrRouter}'s constructor is the one place both are built.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* HerdrRouter} and never runs {@code FleetdAssembly.assembleAndStart} — a green result proves only
* that neither watched file's text contains {@code new AgentControl(} or {@code new
* WorkspaceControl(}. It does not prove the instances {@code HerdrRouter} does build are the ones
* actually wired through the rest of the daemon, and it does not cover a bypass written into a
* production file other than the two this test reads.
*/
class FleetdHerdrControlConstructionTest {
private static String source(String relativePath) throws Exception {
return Files.readString(Path.of(relativePath));
}
@Test
@DisplayName("[SOURCE TEXT] Fleetd.java never constructs AgentControl or WorkspaceControl directly")
void fleetdDelegatesStatefulControlsToTheRouter() throws Exception {
// AgentControl caches paneByTerminal, so the router must be its only production factory.
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new AgentControl("));
assertFalse(source.contains("new WorkspaceControl("));
String source = source("src/main/java/dev/ltms/fleet/Fleetd.java");
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse checks below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"the read of Fleetd.java did not come back containing its own class declaration — "
+ "the assertFalse checks below would pass vacuously on a broken read; fix the "
+ "read before trusting this test.");
assertFalse(source.contains("new AgentControl("),
"Fleetd.java must not construct AgentControl directly — HerdrRouter is its only "
+ "production factory");
assertFalse(source.contains("new WorkspaceControl("),
"Fleetd.java must not construct WorkspaceControl directly — HerdrRouter is its only "
+ "production factory");
}
@Test
@DisplayName("[SOURCE TEXT] FleetdAssembly.java never constructs AgentControl or WorkspaceControl directly")
void fleetdAssemblyDelegatesStatefulControlsToTheRouter() throws Exception {
String source = source("src/main/java/dev/ltms/fleet/FleetdAssembly.java");
assertTrue(source.contains("final class FleetdAssembly"),
"the read of FleetdAssembly.java did not come back containing its own class "
+ "declaration — the assertFalse checks below would pass vacuously on a broken "
+ "read; fix the read before trusting this test.");
assertFalse(source.contains("new AgentControl("),
"FleetdAssembly.java must not construct AgentControl directly — HerdrRouter is its "
+ "only production factory");
assertFalse(source.contains("new WorkspaceControl("),
"FleetdAssembly.java must not construct WorkspaceControl directly — HerdrRouter is "
+ "its only production factory");
}
}
@@ -0,0 +1,213 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #612 Shape A, unit r5 — {@code FleetdAssembly.java:488} wires {@link
* FleetMcp.LeadConfigDirSource} with {@code Fleetd.leadConfigDirSource(() -> config.get().profiles(),
* leaders)}. {@link FleetdLeadConfigDirSourceWiringTest} already pins that the FACTORY itself
* delegates to the real {@link Fleetd#leadConfigDirLookup} — but, by its own javadoc, it "does not
* and structurally cannot cover" whether the real call site in {@code FleetdAssembly} still calls
* that factory at all. Measured there: swapping that one-line call for a bare {@code
* FleetMcp.LeadConfigDirSource.none()} compiles with 0 errors and leaves the full suite green.
*
* <p>This is the literal fleetd #602/#606 defect, one call site away from its own fix: {@code main}
* (now {@code FleetdAssembly}) used to build {@code LeadConfigDirSource.none()} inline, the whole
* suite passed, and the live daemon reported {@code "state":"unknown"} for every lead's context,
* forever, with no test noticing. The fix extracted the factory; this test is the one that proves
* {@code FleetdAssembly}'s own call site still reaches it.
*
* <p>This test drives the REAL {@link FleetMcp} the real {@link FleetdAssembly#assembleAndStart}
* builds, reached through {@link FleetdRuntime#mcp()}, and reads the {@code leadConfigDirs} field it
* was constructed with via reflection — {@code FleetMcp} exposes no public accessor for it (unlike
* {@code quarantineSource()}/{@code leadSeatSource()}), so there is no non-reflective route to the
* live instance. The assertion resolves a REAL lead name against a REAL configured {@code
* configDir:}: {@link FleetMcp.LeadConfigDirSource#none()} (the historical defect, and the
* mis-wire this test's mutation cycles reintroduce) always returns {@code null} regardless of the
* input, so a non-null, config-matching answer is a property {@code none()} can never produce by
* accident.
*/
class FleetdLeadConfigDirSourceAssemblyTest {
private static final String LEAD_NAME = "opus";
private static final String LEAD_TAB = "lead: opus";
private static final String LEAD_PROFILE = "sonnet";
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
fleet:
leaders:
%s:
tab: "%s"
profile: %s
profiles:
%s:
subscription: true
argv: ["ccs", "sonnet"]
configDir: "%s"
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
private static FleetMcp.LeadConfigDirSource leadConfigDirSourceOf(FleetMcp mcp) throws Exception {
Field field = FleetMcp.class.getDeclaredField("leadConfigDirs");
field.setAccessible(true);
return (FleetMcp.LeadConfigDirSource) field.get(mcp);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadConfigDirSource resolves a lead's REAL "
+ "configured configDir, not the none() stand-in's hardcoded null")
void assembledLeadConfigDirSourceResolvesTheRealConfiguredConfigDir(@TempDir Path dir) throws Exception {
String configuredConfigDir = "/mnt/fake-lead-configdir";
FleetConfig cfg = writeConfig(dir, configuredConfigDir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
// agent) to match fleet.leaders.opus.tab exactly, so LeadLauncher.ensureLeads() sees the
// lead as already live and does not try to auto-launch a second one.
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
FleetMcp.LeadConfigDirSource source = leadConfigDirSourceOf(runtime.mcp());
assertEquals(configuredConfigDir, source.configDirFor().apply(LEAD_NAME),
"fleet.leaders." + LEAD_NAME + ".profile (" + LEAD_PROFILE + ") configures "
+ "configDir: " + configuredConfigDir + " — the real assembled source must "
+ "resolve it. FleetMcp.LeadConfigDirSource.none() (the inert stand-in "
+ "this test's mutation cycles swap the call site for, and the historical "
+ "fleetd #602/#606 defect) always reports null here, whatever the input");
// A lead name the config does not recognise still resolves to null, not a crash — the
// same source, applied to an input that must stay at the inert answer even on the real,
// non-inert instance.
assertNull(source.configDirFor().apply("no-such-lead"));
} finally {
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
// failure path too — hence try/finally rather than a bare statement at the end.
runtime.close();
}
}
}
@@ -0,0 +1,63 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* Pins {@link Fleetd#leadConfigDirSource}'s own wiring of the window lookup into the returned
* {@link FleetMcp.LeadConfigDirSource}, not only the detached {@link Fleetd#leadContextWindowLookup}
* factory it delegates to. Calls the producer directly, with real {@link FleetConfig.Profile}/
* {@link FleetConfig.Leader} fixtures, and asserts on {@code windowFor()} — the companion of
* {@link FleetdLeadConfigDirSourceWiringTest}, which pins the same factory's {@code configDirFor()}.
*/
class FleetdLeadConfigDirSourceWindowWiringTest {
private static FleetConfig.Profile profileWithWindow(String name, Integer autoCompactWindow) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, null, true, null, null, null, null, null, autoCompactWindow, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("the returned source resolves the lead's REAL configured effective window, not a hardcoded null")
void resolvesTheRealConfiguredWindow() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", 250_000));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertEquals(250_000L, source.windowFor().apply("primary"),
"windowFor must delegate to the real leadContextWindowLookup, not a stub that always "
+ "returns null");
}
@Test
@DisplayName("a lead on a profile with no window configured still resolves to null, not a crash")
void leadWithNoWindowConfiguredResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertNull(source.windowFor().apply("primary"));
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
assertNull(source.windowFor().apply("ghost-lead"));
}
}
@@ -99,7 +99,7 @@ class FleetdLeadContextLookupTest {
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
@@ -112,7 +112,7 @@ class FleetdLeadContextLookupTest {
void agentsGetThrowingDegradesToUnknown() {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
@@ -128,7 +128,7 @@ class FleetdLeadContextLookupTest {
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
@@ -146,7 +146,7 @@ class FleetdLeadContextLookupTest {
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
@@ -0,0 +1,223 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.lang.reflect.Field;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #659: {@code FleetdAssembly.java} wires {@code Fleetd.leadContextSource}'s window-lookup
* argument with {@code Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)} —
* but nothing called the real assembled {@link LeadHeartbeatLoop} far enough to prove that
* argument is the one the live heartbeat reads through. Measured: swapping that one call-site
* argument for {@code _ -> null} compiles with 0 errors and leaves the full suite green.
*
* <p>This test drives the REAL {@link LeadHeartbeatLoop} the real {@link
* FleetdAssembly#assembleAndStart} builds, reached through {@link FleetdRuntime#heartbeat()}, and
* reads its private {@code contextSource} field via reflection — the loop exposes no public
* accessor for it, the same reason {@link FleetdLeadConfigDirSourceAssemblyTest} reflects on
* {@code FleetMcp.leadConfigDirs}. The configured profile's {@code autoCompactWindow: 100000}
* resolves a HIGH threshold of {@code 66666} ({@link LeadContextGauge}'s {@code 2/3} fraction) —
* far below the fixed {@code 200000} fallback a lost window argument would silently revert to.
* {@code 90000} live tokens sits between the two: HIGH under the real window, OK under the
* fallback — a property the fallback can never produce by accident.
*/
class FleetdLeadContextSourceWindowAssemblyTest {
private static final String LEAD_NAME = "opus";
private static final String LEAD_TAB = "lead: opus";
private static final String LEAD_PROFILE = "sonnet";
/** {@code FakeHerdr}'s own default {@code agent.list} entry: terminal {@code term_a}, session {@code sess-1111}. */
private static final String LEAD_TERMINAL = "term_a";
private static final String LEAD_SESSION_ID = "sess-1111";
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
leadHeartbeat:
idleAfterSeconds: 600
backoffMs: 15000
quietNudgeCap: 5
fleet:
leaders:
%s:
tab: "%s"
profile: %s
profiles:
%s:
subscription: true
argv: ["ccs", "sonnet"]
configDir: "%s"
autoCompactWindow: 100000
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
private static LeadHeartbeatLoop.LeadContextSource contextSourceOf(LeadHeartbeatLoop heartbeat) throws Exception {
Field field = LeadHeartbeatLoop.class.getDeclaredField("contextSource");
field.setAccessible(true);
return (LeadHeartbeatLoop.LeadContextSource) field.get(heartbeat);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled heartbeat loop resolves HIGH against the lead's "
+ "REAL configured window, not the fixed 200000 fallback a lost window argument reverts to")
void assembledHeartbeatContextSourceResolvesTheRealConfiguredWindow(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir, dir.toString());
writeTranscript(dir, LEAD_SESSION_ID, 90_000);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
// agent on session sess-1111) to match fleet.leaders.opus.tab exactly, so LeadTabScanner
// recognises it as the live "opus" lead without a second auto-launched pane.
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB).agentSessionId(LEAD_SESSION_ID);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
LeadHeartbeatLoop.LeadContextSource source = contextSourceOf(runtime.heartbeat());
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.HIGH, reading.state(),
"profiles." + LEAD_PROFILE + ".autoCompactWindow: 100000 resolves a HIGH threshold "
+ "of 66666 tokens — 90000 live tokens must read HIGH against it. Mutating "
+ "FleetdAssembly's window-lookup argument to `_ -> null` falls back to the "
+ "fixed 200000 threshold, under which 90000 reads OK instead: " + reading);
} finally {
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
// failure path too — hence try/finally rather than a bare statement at the end.
runtime.close();
}
}
}
@@ -87,7 +87,8 @@ class FleetdLeadContextSourceWiringTest {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
agentControlStub(), () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
@@ -101,7 +102,7 @@ class FleetdLeadContextSourceWiringTest {
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), Map::of, name -> null);
agentControlStub(), Map::of, name -> null, name -> null);
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
}
@@ -0,0 +1,96 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* {@link Fleetd#leadContextWindowLookup} is the factory wired into {@code
* FleetMcp.LeadConfigDirSource} and {@code LeadHeartbeatLoop.LeadContextSource} so {@link
* dev.ltms.fleet.lead.LeadContextGauge} scales its HIGH threshold against a lead's own profile's
* effective auto-compact window instead of always the gauge's fixed fallback — the same {@code
* fleet.leaders.<name>.profile} link {@link Fleetd#leadConfigDirLookup} already follows, one step
* further to {@link FleetConfig.Profile#effectiveAutoCompactWindow()}.
*/
class FleetdLeadContextWindowLookupTest {
private static FleetConfig.Profile profileWithWindow(String name, Integer autoCompactWindow,
Map<String, String> env) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, env,
null, null, true, null, null, null, null, null, autoCompactWindow, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("a lead on a profile that sets autoCompactWindow resolves to that window")
void leadOnAProfileWithAutoCompactWindowResolvesToIt() {
Map<String, FleetConfig.Profile> profiles =
Map.of("opus", profileWithWindow("opus", 250_000, Map.of()));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
assertEquals(250_000L, lookup.apply("primary"));
}
@Test
@DisplayName("yaml and env disagree: the lookup resolves the env value, not the yaml one")
void yamlAndEnvDisagreeLookupResolvesTheEnvValue() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus",
profileWithWindow("opus", 250_000, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "150000")));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
assertEquals(150_000L, lookup.apply("primary"));
}
@Test
@DisplayName("a lead entry with no `profile:` resolves to null, not a thrown exception")
void recogniseOnlyLeadWithNoProfileResolvesToNull() {
Map<String, FleetConfig.Profile> profiles =
Map.of("opus", profileWithWindow("opus", 250_000, Map.of()));
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
"claude", "claude-sonnet-5");
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("a lead naming a profile that is not configured resolves to null, not a thrown exception")
void leadOnAnUnconfiguredProfileResolvesToNull() {
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("ghost-profile"));
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(Map::of, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("a lead on a profile that resolves no window at all resolves to null")
void leadOnAProfileWithNoWindowResolvesToNull() {
Map<String, FleetConfig.Profile> profiles =
Map.of("opus", profileWithWindow("opus", null, Map.of()));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
assertNull(lookup.apply("primary"));
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(Map::of, Map.of());
assertNull(lookup.apply("ghost-lead"));
}
}
@@ -0,0 +1,243 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import io.modelcontextprotocol.client.McpClient;
import io.modelcontextprotocol.client.McpSyncClient;
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
import io.modelcontextprotocol.spec.McpClientTransport;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.Timeout;
import org.junit.jupiter.api.io.TempDir;
import java.net.http.HttpRequest;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Shape A ranks 9 and 11: drives the real {@code fleet_list} MCP route through a
* token-authenticated primary caller. The assertions read the response from the exact {@link
* dev.ltms.fleet.mcp.FleetMcp} instance assembled by {@link FleetdAssembly}, rather than a source
* scrape or a separately built reporting source.
*
* <p>The coordinator fixture contains a peer on purpose. Both a missing coordinator and an
* incorrectly wired peers argument can render as an empty list, so the non-empty peer assertion
* distinguishes the real call-site value from that inert result. Its fake mailbox never contacts a
* broker.
*/
class FleetdListReportingSourcesAssemblyTest {
private static final String TOKEN = "fleetd-r9-r11-test-token";
private static final String TOKEN_ENV = "FLEETD_R9_TEST_TOKEN";
private static final String PROFILE = "capacity-profile";
private static final String PEER = "peer-fleet";
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
private final FakeHerdr herdr = new FakeHerdr();
private Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of(TOKEN_ENV, TOKEN);
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// The test binds runtime.app() itself, below.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("healthy FakeHerdr must not poll");
};
}
}
private FleetdRuntime runtime;
private TestResourcePorts ports;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
auth:
mode: token
tokenEnv: %s
idleSleepGuard:
enabled: false
health:
enabled: true
notifications:
mode: webhook
profiles:
%s:
baseUrl: http://capacity.test:8000
model: test-model
maxLoad: 7
coordinator:
uri: amqp://fake-coordinator/vh
selfId: test-lead
peers:
- %s
""".formatted(TOKEN_ENV, PROFILE, PEER));
return FleetConfig.load(file);
}
private int assembleAndBind(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
ports = new TestResourcePorts();
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config,
new SubscriptionGuard(cfg.guard().hostSet())), ports);
boundApp = runtime.app().start("127.0.0.1", 0);
return boundApp.port();
}
private static String fleetList(int port) {
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder()
.header("Authorization", "Bearer " + TOKEN);
McpClientTransport transport = HttpClientStreamableHttpTransport.builder("http://127.0.0.1:" + port)
.endpoint("/mcp")
.requestBuilder(requestTemplate)
.build();
try (McpSyncClient client = McpClient.sync(transport).build()) {
client.initialize();
McpSchema.CallToolResult response = client.callTool(McpSchema.CallToolRequest.builder("fleet_list")
.arguments(Map.of()).build());
return ((McpSchema.TextContent) response.content().getFirst()).text();
}
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void fleetListReportsTheAssembledCapacitySource(@TempDir Path dir) throws Exception {
String response = fleetList(assembleAndBind(dir));
assertTrue(response.contains("\"capacity\""), response);
assertTrue(response.contains("\"" + PROFILE + "\""), response);
assertTrue(response.contains("\"maxLoad\":7"), response);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void fleetListReportsTheAssembledHealthCoverageSource(@TempDir Path dir) throws Exception {
String response = fleetList(assembleAndBind(dir));
assertTrue(response.contains("\"healthCoverage\":\"full\""), response);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void fleetListReportsTheAssembledCoordinatorPeers(@TempDir Path dir) throws Exception {
String response = fleetList(assembleAndBind(dir));
assertTrue(response.contains("\"coordinator\""), response);
assertTrue(response.contains("\"coordId\":\"" + PEER + "\""), response);
}
}
@@ -0,0 +1,368 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.fleet.session.MemberSession;
import io.javalin.Javalin;
import io.modelcontextprotocol.client.McpClient;
import io.modelcontextprotocol.client.McpSyncClient;
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
import io.modelcontextprotocol.spec.McpClientTransport;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.Timeout;
import org.junit.jupiter.api.io.TempDir;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Shape A, unit r4 — {@link FleetdAssembly}'s {@code quarantineSource} (lines 471-472)
* and {@code outageSource} (lines 473-476 at {@code main} = {@code 141ae3b}), EACH of which feeds
* two separate consumers: {@code FleetMcp} ({@code fleet_profiles}) and {@code FleetApp}
* ({@code GET /profiles}), one call site ({@code :484}/{@code :486}) into the MCP constructor and
* the SAME shared local again ({@code :529}) into the REST constructor.
*
* <p><strong>The property pinned here</strong> (from the ticket): when the real assembly has built
* a real {@link dev.ltms.fleet.placement.BackendQuarantine} holding a quarantined credential, and a
* real {@link dev.ltms.fleet.placement.BackendOutagePolicy} holding a cooling-off credential, BOTH
* operator windows must report that state — the real assembled {@code FleetMcp} (reached through
* {@link FleetdRuntime#mcp()}) and the real assembled REST surface (reached through {@link
* FleetdRuntime#app()}). A mutation that starves one consumer while leaving the other wired must
* make only that consumer's assertion go red.
*
* <p><strong>No source-text assertion anywhere in this file.</strong> Both windows are read off the
* REAL running objects: {@code fleet_profiles} is called through a real MCP client over a real
* HTTP connection to the servlet {@link FleetdAssembly} actually mounted, and {@code GET /profiles}
* is called through a real {@link java.net.http.HttpClient} against the real bound {@link
* FleetdRuntime#app()}. Neither is a copy built alongside the assembly for this test's benefit.
*
* <p><strong>How this gets past CB-185's own pid-resolution dead end.</strong> {@code
* FleetdAssemblyFleetAppTest}'s class javadoc explains that {@code GET /sessions} cannot be driven
* over real HTTP here because {@code LsofPeerPidLookup} excludes its own pid and an in-process test
* client/server share one JVM pid — every such request resolves {@code ANONYMOUS} and is refused
* before the handler runs. {@code GET /profiles} and {@code fleet_profiles} sit behind the exact
* same {@code Authz.Action.READ} gate. This test sidesteps the dead end instead of hitting it:
* {@code auth.mode: token} (see {@link dev.ltms.fleet.auth.CallerResolver#resolve}) resolves a
* caller to {@code PRIMARY} from a valid {@code Authorization: Bearer} header ALONE, with no pid
* resolution involved at all — the same technique {@code FleetMcpContextExtractorTest} already uses
* to drive a real {@code fleet_whoami} call through the real transport.
*
* <p><strong>How the quarantined/cooling-off state is set up.</strong> Both {@code BackendQuarantine}
* and {@code BackendOutagePolicy} are private to the collaborators the assembly wires them into, and
* (unlike {@code quarantineSource()}) neither {@code FleetMcp} nor {@code FleetApp} exposes a public
* accessor for the live {@code BackendOutagePolicy} instance. Rather than add one (acceptance
* criterion 1: no production change), this test drives the REAL production classification path —
* exactly the recipe {@code FleetdExhaustedPatternAssemblyTest} (quarantine) and {@code
* FleetdCompletionResolverAssemblyTest} (cool-off) already proved works end to end against this same
* {@link FleetdAssembly#assembleAndStart}: acquire a real {@link MemberSession}, feed the real {@link
* CompletionResolver} a pane scrape matching the profile's configured {@code exhaustedPattern} /
* {@code errorPattern}, and let the real {@code exhaustionSink}/{@code backendErrorSink} write into
* the real, shared tracker. Each setup step asserts its own {@link Rendezvous.Kind} as a CONTROL —
* if the resolver were never actually exercised, the setup itself fails loudly before either window
* is ever read.
*/
class FleetdQuarantineOutageDualWindowAssemblyTest {
private static final String TOKEN = "s3cret-r4-token";
private static final String TOKEN_ENV = "FLEETD_R4_TEST_TOKEN";
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr;
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
Runnable shutdownHook;
ControllableResourcePorts(FakeHerdr herdr) {
this.herdr = herdr;
}
void advanceSeconds(long seconds) {
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
}
@Override
public Map<String, String> environment() {
return Map.of(TOKEN_ENV, TOKEN);
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
// Never invoked: this test's config has no `broker:` block.
return (uri, prefetch) -> {
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
// Never invoked: this test's config has no `coordinator:` block.
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private FleetdRuntime runtime;
private ControllableResourcePorts ports;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
auth:
mode: token
tokenEnv: %s
idleSleepGuard:
enabled: false
quarantineCooldownSeconds: %d
profiles:
exhaustprofile:
baseUrl: http://exhausthost.local:8000
model: sonnet
exhaustedPattern: "usage limit reached"
coolprofile:
baseUrl: http://coolhost.local:8000
model: sonnet
errorPattern: "credential outage"
guard:
offSubscriptionHosts:
- exhausthost.local
- coolhost.local
""".formatted(TOKEN_ENV, cooldownSeconds));
return FleetConfig.load(f);
}
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
private int assembleAndBind(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir, 120);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ports = new ControllableResourcePorts(new FakeHerdr());
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
boundApp = runtime.app().start("127.0.0.1", 0);
return boundApp.port();
}
/**
* Drives the real assembled {@link CompletionResolver} through a scrape matching {@code
* exhaustprofile}'s configured {@code exhaustedPattern}, exactly {@code
* FleetdExhaustedPatternAssemblyTest}'s own recipe, so the real {@code exhaustionSink} quarantines
* the credential ({@code effectiveCredentialId() == "exhaustprofile"}, no explicit credentialId
* configured).
*/
private void quarantineExhaustProfile(Path dir) {
MemberSession session = runtime.sessions().acquire("exhaustprofile", null, dir.toString(), null);
String target = session.terminalId();
CompletionResolver completion = runtime.completion();
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
ports.herdr.readText("idle, nothing yet");
completion.onDelivered(target, new TurnToken(target, waiter, null));
ports.herdr.readText("usage limit reached: try again in a few hours");
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s), no real sleep
completion.resolveBeforePostAction(target);
Rendezvous.Resolution resolution = waiter.getNow(null);
assertTrue(resolution != null && resolution.kind() == Rendezvous.Kind.BACKEND_EXHAUSTED,
"CONTROL: setup must classify as BACKEND_EXHAUSTED before either window is read — "
+ "if this fails, the assembled resolver was never actually exercised: " + resolution);
}
/**
* Drives the real assembled {@link CompletionResolver} with TWO distinct targets on {@code
* coolprofile}, each matching its configured {@code errorPattern}, exactly {@code
* FleetdCompletionResolverAssemblyTest}'s own recipe, so the real {@code backendErrorSink} cools
* the credential off ({@code effectiveCredentialId() == "coolprofile"}).
*/
private void coolOffCoolProfile(Path dir) {
MemberSession s1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
MemberSession s2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
CompletionResolver completion = runtime.completion();
String t1 = s1.terminalId();
CompletableFuture<Rendezvous.Resolution> w1 = new CompletableFuture<>();
ports.herdr.readText("idle 1");
completion.onDelivered(t1, new TurnToken(t1, w1, null));
ports.herdr.readText("credential outage: upstream 503");
ports.advanceSeconds(3);
completion.resolveBeforePostAction(t1);
Rendezvous.Resolution r1 = w1.getNow(null);
assertTrue(r1 != null && r1.kind() == Rendezvous.Kind.FAILED,
"CONTROL: target1's setup must classify FAILED (backend error): " + r1);
String t2 = s2.terminalId();
CompletableFuture<Rendezvous.Resolution> w2 = new CompletableFuture<>();
ports.herdr.readText("idle 2");
completion.onDelivered(t2, new TurnToken(t2, w2, null));
ports.herdr.readText("credential outage: upstream 503 again");
ports.advanceSeconds(3);
completion.resolveBeforePostAction(t2);
Rendezvous.Resolution r2 = w2.getNow(null);
assertTrue(r2 != null && r2.kind() == Rendezvous.Kind.FAILED,
"CONTROL: target2's setup must classify FAILED (backend error) — two distinct "
+ "targets are required to cross BackendOutagePolicy.THRESHOLD: " + r2);
}
/** Calls the real {@code fleet_profiles} tool over a real MCP client, token-authenticated as PRIMARY. */
private static McpSchema.CallToolResult callProfilesViaMcp(int port) {
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder().header("Authorization", "Bearer " + TOKEN);
McpClientTransport transport = HttpClientStreamableHttpTransport.builder("http://127.0.0.1:" + port)
.endpoint("/mcp")
.requestBuilder(requestTemplate)
.build();
try (McpSyncClient client = McpClient.sync(transport).build()) {
client.initialize();
return client.callTool(McpSchema.CallToolRequest.builder("fleet_profiles").arguments(Map.of()).build());
}
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
/** Calls the real {@code GET /profiles} route over a real {@link HttpClient}, same token. */
private static String getProfilesViaRest(int port) throws Exception {
HttpClient http = HttpClient.newHttpClient();
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + "/profiles"))
.header("Authorization", "Bearer " + TOKEN)
.GET().build();
HttpResponse<String> res = http.send(req, HttpResponse.BodyHandlers.ofString());
assertEquals(200, res.statusCode(), "GET /profiles must succeed with the real token: " + res.body());
return res.body();
}
// --- quarantineSource (FleetdAssembly.java :471-472) --------------------------------------
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void quarantinedCredentialIsReportedByTheRealAssembledFleetMcp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
quarantineExhaustProfile(dir);
String out = textOf(callProfilesViaMcp(port));
assertTrue(out.contains("\"quarantined\""), "fleet_profiles must report a quarantined "
+ "section once the real BackendQuarantine holds a quarantined credential: " + out);
assertTrue(out.contains("\"exhaustprofile\""), out);
assertTrue(out.contains("\"quarantinedForSeconds\""), out);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void quarantinedCredentialIsReportedByTheRealAssembledFleetApp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
quarantineExhaustProfile(dir);
String out = getProfilesViaRest(port);
assertTrue(out.contains("\"quarantined\""), "GET /profiles must report a quarantined "
+ "section once the real BackendQuarantine holds a quarantined credential: " + out);
assertTrue(out.contains("\"exhaustprofile\""), out);
assertTrue(out.contains("\"quarantinedForSeconds\""), out);
}
// --- outageSource (FleetdAssembly.java :473-476) -------------------------------------------
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void coolingOffCredentialIsReportedByTheRealAssembledFleetMcp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
coolOffCoolProfile(dir);
String out = textOf(callProfilesViaMcp(port));
assertTrue(out.contains("\"coolingOff\""), "fleet_profiles must report a coolingOff "
+ "section once the real BackendOutagePolicy holds a cooling-off credential: " + out);
assertTrue(out.contains("\"coolprofile\""), out);
assertTrue(out.contains("\"coolingOffForSeconds\""), out);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void coolingOffCredentialIsReportedByTheRealAssembledFleetApp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
coolOffCoolProfile(dir);
String out = getProfilesViaRest(port);
assertTrue(out.contains("\"coolingOff\""), "GET /profiles must report a coolingOff "
+ "section once the real BackendOutagePolicy holds a cooling-off credential: " + out);
assertTrue(out.contains("\"coolprofile\""), out);
assertTrue(out.contains("\"coolingOffForSeconds\""), out);
}
}
@@ -0,0 +1,65 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* {@link FleetConfig.Profile#effectiveAutoCompactWindow()} resolves the window a launched Claude
* Code session actually runs on, not just the {@code autoCompactWindow:} launch flag — {@code env:
* CLAUDE_CODE_AUTO_COMPACT_WINDOW} overrides that flag, so a profile setting both resolves from the
* environment variable.
*/
class FleetConfigProfileEffectiveAutoCompactWindowTest {
private static FleetConfig.Profile profile(Integer autoCompactWindow, Map<String, String> env) {
return new FleetConfig.Profile("sonnet", null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, env,
null, null, true, null, null, null, null, null, autoCompactWindow, null);
}
@Test
@DisplayName("yaml and env disagree: the env var wins, not the yaml key")
void yamlAndEnvDisagreeEnvWins() {
FleetConfig.Profile p = profile(250_000, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "150000"));
assertEquals(150_000, p.effectiveAutoCompactWindow(),
"the two inputs must give DIFFERENT thresholds (150,000 vs 250,000) and the env value must win");
}
@Test
@DisplayName("only autoCompactWindow set: that value resolves")
void onlyAutoCompactWindowSetResolvesToIt() {
FleetConfig.Profile p = profile(250_000, Map.of());
assertEquals(250_000, p.effectiveAutoCompactWindow());
}
@Test
@DisplayName("only the env var set: that value resolves")
void onlyEnvVarSetResolvesToIt() {
FleetConfig.Profile p = profile(null, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "150000"));
assertEquals(150_000, p.effectiveAutoCompactWindow());
}
@Test
@DisplayName("neither set: resolves to null")
void neitherSetResolvesToNull() {
FleetConfig.Profile p = profile(null, Map.of());
assertNull(p.effectiveAutoCompactWindow());
}
@Test
@DisplayName("an unparseable env value falls back to autoCompactWindow rather than throwing")
void unparseableEnvValueFallsBackToAutoCompactWindow() {
FleetConfig.Profile p = profile(250_000, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "not-a-number"));
assertEquals(250_000, p.effectiveAutoCompactWindow());
}
}
@@ -765,6 +765,92 @@ class FleetConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
/**
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
* its own, so it can land inside a lead's labelled tab and be read back as that lead.
*/
@Test
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
@Test
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
profile: gx10
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
}
@Test
void aTabPlacedProfileWithALeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("tab-safe.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: tab
fleet:
leaders:
opus:
tab: "lead: opus"
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a member in its own tab cannot land inside a lead's tab");
}
/** Proves the reflective sweep behind {@code validateAll} really reaches this validator. */
@Test
void validateAllAlsoRefusesPanePlacementAgainstALeadTab(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard-sweep.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
@Test
@@ -42,12 +42,13 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
* <li>{@link #validateAllReachesEveryOneOfTodaysRealValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches each of today's six real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those six
* failures.</li>
* reaches every one of today's real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, plus a
* dedicated fixture for {@link FleetConfig#validateLeadRollover()}, which no other test
* drives through {@code validateAll()} — so a single call to {@code validateAll()} is shown
* to reproduce every one of those failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
@@ -57,13 +58,13 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* a test in this module.
*
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
* FleetConfig#validateAll()} to a hardcoded list of today's method calls leaves the whole
* suite green. Nothing ties {@code validateAll()} to
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* 2 proves {@code validateAll()} reaches today's validators, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
* is declared, which forces whoever adds it to look at this file.
* #fleetConfigDeclaresExactlyTheseValidatorsToday()}: it fails the moment any validator is added
* or removed, which forces whoever changes the set to look at this file.
*/
class FleetConfigValidateAllTest {
@@ -209,18 +210,24 @@ class FleetConfigValidateAllTest {
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches every validator ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
* above, on an unrelated class) — it is a visible denominator: today there are six, named
* here, so a reader adding a seventh sees this assertion name the new count rather than a
* silent pass at the old one.
* above, on an unrelated class) — it is a visible denominator: today there are eight, named in
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
* count rather than a silent pass at the old one. The count lives only in that set, not in this
* method's name, so the two cannot drift apart.
*
* <p>This assertion alone proves only that the validator exists with the right shape — it
* cannot prove {@code validateAll()} actually reaches it. Only {@link
* #validateAllReachesEveryOneOfTodaysRealValidators()} proves reachability, which is why this
* method's failure message sends the reader there too.
*/
@Test
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
Set<String> names = new TreeSet<>();
for (Method m : FleetConfig.class.getMethods()) {
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
@@ -233,12 +240,18 @@ class FleetConfigValidateAllTest {
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels", "validateLeadRollover")), names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
names,
"FleetConfig's public validate*() methods changed. Do THREE things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
+ "test in this class, so this assertion is the only place that will ever "
+ "make you check. Only then update the expected set to match.");
+ "make you check. Second, update the expected set below to match. Third, "
+ "add or remove a case for that validator in "
+ "validateAllReachesEveryOneOfTodaysRealValidators() below — this "
+ "assertion proves only that the validator exists with the right shape, "
+ "never that validateAll() reaches it; that enumeration is the test that "
+ "does.");
}
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
@@ -262,14 +275,21 @@ class FleetConfigValidateAllTest {
}
/**
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
* filter), exactly one of these six would start passing when it must not.
* The heart of claim 2: for every one of today's real validators, a minimal file that
* fails ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly, or a dedicated minimal fixture where no other test drives that validator through
* {@code validateAll()} — must also fail through {@link FleetConfig#validateAll()}. If a
* future edit to {@code validateAll()} silently dropped one of these from the sweep
* (e.g. a typo'd name filter), exactly one of them would start passing when it must not.
*
* <p>This is the single place that proves {@code validateAll()} reaches a given validator.
* Adding or removing a validator on {@link FleetConfig} must add or remove a case here, not
* only an updated name in {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}'s expected
* set — that assertion proves the validator's shape, never that {@code validateAll()} reaches
* it.
*/
@Test
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
// validateAuthExposure: a non-loopback bind without token mode.
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
bind:
@@ -342,6 +362,29 @@ class FleetConfigValidateAllTest {
allow:
- model: claude-sonnet-5
""", "rogue");
// validatePanePlacementAgainstLeadTabs: a pane-placed profile while a lead names a tab.
assertValidateAllRefuses(dir, "pane-placement.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""", "gx10");
// validateLeadRollover: a leadRollover: block present with no handoverPath.
assertValidateAllRefuses(dir, "lead-rollover.yaml", """
bind:
host: 127.0.0.1
port: 8765
leadRollover:
requireOperatorConfirm: false
""", "handoverPath");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
@@ -51,6 +51,7 @@ public final class FakeHerdr implements HerdrClient {
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String agentSessionId = null; // agent_session.value on agent.get; null = omitted
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
@@ -173,6 +174,16 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the {@code agent_session.value} that {@code agent.get} reports for {@code term_a} — the
* default omits the field entirely (herdr not yet having resolved one), matching the real
* daemon's own "not resolved yet" shape.
*/
public FakeHerdr agentSessionId(String sessionId) {
this.agentSessionId = sessionId;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
@@ -311,10 +322,12 @@ public final class FakeHerdr implements HerdrClient {
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
String sessionField = agentSessionId == null ? ""
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + agentSessionId + "\"}";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentField, agentStatus));
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"%s}}""")
.formatted(agentField, agentStatus, sessionField));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
@@ -232,8 +232,9 @@ class LeadTabScannerTest {
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing fleetd placed can land here (see the worker-space test above).
// A human may split their own lead tab. Both panes are theirs, so both are that lead.
// A pane-placed member landing here instead is refused at startup by
// FleetConfig.validatePanePlacementAgainstLeadTabs, not by this scanner.
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0",
@@ -0,0 +1,104 @@
package dev.ltms.fleet.lead;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* {@code LeadContextGauge}'s HIGH threshold scales with the caller's resolved effective
* auto-compact window, so the margin it warns at makes sense against a window that can legally sit
* as low as {@code 100_000}, not only against the fixed fallback. These properties pin that
* scaling, its boundary, and the fallback used when no window is resolvable.
*/
class LeadContextGaugeHighThresholdTest {
private static final String SESSION_ID = "55555555-5555-5555-5555-555555555555";
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
private static String writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
return configDir.toString();
}
@Test
@DisplayName("an effective window of 100000 reports HIGH strictly below 100000")
void effectiveWindowOf100000ReportsHighStrictlyBelow100000(@TempDir Path tmp) throws IOException {
// 90,000 is below the 100,000 window itself, but above a fixed 200,000 fallback would ever
// reach — only a threshold that scales with the window can report HIGH here.
String configDir = writeTranscript(tmp, SESSION_ID, 90_000);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", 100_000L);
assertEquals(LeadContextGauge.State.HIGH, reading.state(),
"90,000 tokens against a 100,000 effective window must already be HIGH, with margin to spare "
+ "before a compaction at the window itself");
}
@Test
@DisplayName("the boundary sits strictly between OK and HIGH on both sides")
void boundarySitsStrictlyBetweenOkAndHighOnBothSides(@TempDir Path tmp) throws IOException {
long window = 100_000L;
long threshold = (long) (window * (2.0 / 3.0)); // 66,666
String belowConfigDir = writeTranscript(tmp.resolve("below"), SESSION_ID, threshold - 1);
LeadContextGauge belowGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.OK,
belowGauge.read(belowConfigDir, SESSION_ID, "claude", window).state(),
"one token short of the threshold must stay OK");
String atConfigDir = writeTranscript(tmp.resolve("at"), SESSION_ID, threshold);
LeadContextGauge atGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.HIGH,
atGauge.read(atConfigDir, SESSION_ID, "claude", window).state(),
"exactly at the threshold must already be HIGH");
}
@Test
@DisplayName("with no effective window resolvable, the threshold is still the fixed 200000 default")
void noEffectiveWindowFallsBackToTheFixed200000Default(@TempDir Path tmp) throws IOException {
String belowConfigDir = writeTranscript(tmp.resolve("below"), SESSION_ID, 199_999);
LeadContextGauge belowGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.OK,
belowGauge.read(belowConfigDir, SESSION_ID, "claude", null).state(),
"one token short of the fixed default must stay OK when no window is resolvable");
String atConfigDir = writeTranscript(tmp.resolve("at"), SESSION_ID, 200_000);
LeadContextGauge atGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.HIGH,
atGauge.read(atConfigDir, SESSION_ID, "claude", null).state(),
"the fixed default must still be 200,000 when no window is resolvable");
}
@Test
@DisplayName("a second read with a different window, within the TTL, reports against its own window, not the first call's cached state")
void aSecondReadWithADifferentWindowWithinTheTtlReportsAgainstItsOwnWindow(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, 90_000);
long[] now = {0L};
LeadContextGauge gauge = new LeadContextGauge(() -> now[0], 5_000);
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", 100_000L);
assertEquals(LeadContextGauge.State.HIGH, first.state(),
"90,000 tokens against a 100,000 window is HIGH");
now[0] += 1_000; // stays inside the 5,000ms TTL — the cache key must still vary with the window
LeadContextGauge.Reading second = gauge.read(configDir, SESSION_ID, "claude", 1_000_000L);
assertEquals(LeadContextGauge.State.OK, second.state(),
"90,000 tokens against a 1,000,000 window must report OK regardless of the previous call's "
+ "window, even while that call's cache entry is still within its TTL");
}
}
@@ -85,7 +85,7 @@ class LeadContextGaugeTest {
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
assertEquals(LeadContextGauge.State.OK, first.state());
@@ -93,7 +93,7 @@ class LeadContextGaugeTest {
// must change with it, not stay pinned to the first fixture's total.
String otherSession = "22222222-2222-2222-2222-222222222222";
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude", null);
assertEquals(200_000L, second.tokens());
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
}
@@ -111,12 +111,12 @@ class LeadContextGaugeTest {
compactionLine(),
usageLine(3_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude", null);
assertEquals(2, twoCompactions.compactions());
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude", null);
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
}
@@ -126,7 +126,7 @@ class LeadContextGaugeTest {
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
void missingFileIsUnknown(@TempDir Path tmp) {
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@@ -148,7 +148,7 @@ class LeadContextGaugeTest {
assumeFalse(Files.isReadable(file),
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
} finally {
@@ -171,7 +171,7 @@ class LeadContextGaugeTest {
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.OK, reading.state(),
"a torn final line must not turn a good earlier reading into UNKNOWN");
assertEquals(6_000L, reading.tokens(),
@@ -187,7 +187,7 @@ class LeadContextGaugeTest {
"{this is not json at all",
"neither is this{{{");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"every line unparseable is the real format-change signal and must still report UNKNOWN");
assertNull(reading.tokens());
@@ -224,12 +224,12 @@ class LeadContextGaugeTest {
AtomicLong now = new AtomicLong(0);
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
gauge.read(configDir, SESSION_ID, "claude", null);
gauge.read(configDir, SESSION_ID, "claude", null); // still inside the TTL window
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
now.set(6_000); // past the TTL
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
}
@@ -241,8 +241,8 @@ class LeadContextGaugeTest {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode", null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude", null).state());
}
}
@@ -86,7 +86,7 @@ class FleetMcpLeadContextGaugeWiringTest {
LeadContextGauge contextGauge = new LeadContextGauge();
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
LEAD_NAME.equals(name) ? configuredDir.get() : null);
LEAD_NAME.equals(name) ? configuredDir.get() : null, _ -> null);
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(firstRead.contains("\"tokens\":11000"),
+139 -16
View File
@@ -2,6 +2,10 @@
#
# The one auditable way to edit the live fleetd.yaml.
#
# `--set` uses yq and rewrites the whole YAML document in yq's output style. Use `--from` for a
# candidate whose comment alignment or other formatting carries meaning: it copies that file
# verbatim while keeping this script's backup, parse check, atomic install, and verdict read-back.
#
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
@@ -46,6 +50,11 @@
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
# scripts/config-edit.sh --restore
#
# `--set` rewrites the whole file in yq's output style, not only the requested keys. The script
# warns before installation when the candidate changes more lines than its number of --set pairs.
# Use `--from <candidate.yaml>` when comment alignment or other formatting is meaningful: --from
# copies the candidate verbatim, with no yq round-trip.
#
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
# field": a null value falls back to its default rather than erroring, which is silent, not
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
@@ -154,12 +163,10 @@ done
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
# the key, each indented deeper than it. The key-name match above only ever sees the key line
# itself, so those continuation lines used to flow straight through unredacted while the key line
# right above them printed a reassuring "<redacted>" — an incomplete redactor that looks complete
# is worse than one that visibly does nothing, because it stops a reviewer from looking further.
# The fix is structural, not another name to match: once a key line is masked, every following
# line indented STRICTLY DEEPER than that key is masked too, by indentation alone, until the
# indentation returns to the key's own level or shallower. This needs no knowledge of the key's
# name and so protects a block scalar under any masked key, present or future.
# right above them printed a reassuring "<redacted>". The redactor maps masked continuation lines
# from each complete file before it reads the diff. It then masks a printed line when that file
# line is inside a masked key's value. This covers block-scalar bodies even when the key line is
# outside the printed hunk, and it keeps blank lines inside the value masked.
#
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
@@ -168,37 +175,102 @@ done
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
# fix, not a YAML parser.
map_masked_lines() {
local file="$1" side="$2" line content indent lead key line_number=0
local masked=0 masked_indent=0
case "$side" in
old) OLD_MASKED_LINES=() ;;
new) NEW_MASKED_LINES=() ;;
*) die "internal error: unknown redaction map side $side" ;;
esac
while IFS= read -r line || [ -n "$line" ]; do
line_number=$((line_number + 1))
content="$line"
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$masked" = 1 ]; then
if [ -z "${content// /}" ] || [ "$indent" -gt "$masked_indent" ]; then
case "$side" in
old) OLD_MASKED_LINES[$line_number]=1 ;;
new) NEW_MASKED_LINES[$line_number]=1 ;;
esac
continue
fi
masked=0
fi
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
masked=1
masked_indent="$indent"
fi
fi
done < "$file"
}
redact() {
local line prefix content indent lead key
local masked=0 masked_indent=0 saved_nocasematch=0
local old_file="$1" new_file="$2"
local line prefix content indent lead key old_line=0 new_line=0 in_hunk=0
local old_masked new_masked saved_nocasematch=0
shopt -q nocasematch && saved_nocasematch=1
shopt -s nocasematch
sed -E 's#://[^@]*@#://<redacted>@#g' | while IFS= read -r line || [ -n "$line" ]; do
map_masked_lines "$old_file" old
map_masked_lines "$new_file" new
while IFS= read -r line || [ -n "$line" ]; do
if [[ "$line" =~ ^@@\ -([0-9]+)(,([0-9]+))?\ \+([0-9]+)(,([0-9]+))?\ @@ ]]; then
old_line="${BASH_REMATCH[1]}"
new_line="${BASH_REMATCH[4]}"
in_hunk=1
printf '%s\n' "$line"
continue
fi
case "$line" in
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
*) prefix=""; content="$line" ;;
esac
old_masked=0
new_masked=0
if [ "$in_hunk" = 1 ]; then
case "$prefix" in
' ')
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
old_line=$((old_line + 1)); new_line=$((new_line + 1)) ;;
-)
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
old_line=$((old_line + 1)) ;;
+)
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
new_line=$((new_line + 1)) ;;
esac
fi
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$masked" = 1 ] && [ "$indent" -gt "$masked_indent" ]; then
if [ "$old_masked" = 1 ] || [ "$new_masked" = 1 ]; then
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
continue
fi
masked=0
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
masked=1
masked_indent="$indent"
continue
fi
fi
printf '%s\n' "$line"
printf '%s\n' "$line" | sed -E 's#://[^@]*@#://<redacted>@#g'
done
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
}
@@ -457,6 +529,50 @@ parse_check() {
yq eval '.' "$1" >/dev/null 2>&1
}
# Count logical changed lines in a unified diff. A replacement counts once, while an added or
# deleted line also counts once. One changed `--set` value normally produces one changed line.
changed_line_count() {
local before="$1" after="$2" line count=0 old_count=0 new_count=0
while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
---\ *|+++\ *|@@\ *)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
old_count=0
new_count=0
;;
-*) old_count=$((old_count + 1)) ;;
+*) new_count=$((new_count + 1)) ;;
*)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
old_count=0
new_count=0
;;
esac
done < <(diff -u "$before" "$after" || true)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
printf '%s' "$count"
}
warn_set_reformat() {
local before="$1" after="$2" changed
changed="$(changed_line_count "$before" "$after")"
if [ "$changed" -gt "${#SETS[@]}" ]; then
warn "--set changed $changed candidate lines for ${#SETS[@]} pair(s); yq reformatted the whole file. Use --from for meaningful comment alignment or formatting."
fi
}
install_candidate() {
local cand="$1" live="$2"
mv -f "$cand" "$live"
@@ -618,10 +734,14 @@ run_edit() {
fi
ok "candidate parses"
if [ "$MODE" = "set" ]; then
warn_set_reformat "$backup" "$cand"
fi
apply_mode "$cand" "$orig_mode"
say "change (redacted)"
diff -u "$backup" "$cand" | redact || true
diff -u "$backup" "$cand" | redact "$backup" "$cand" || true
say "install"
install_candidate "$cand" "$CONFIG" \
@@ -649,8 +769,11 @@ dry_run_diff() {
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
fi
if [ "$MODE" = "set" ]; then
warn_set_reformat "$CONFIG" "$cand"
fi
say "dry run — diff (redacted), nothing installed"
diff -u "$CONFIG" "$cand" | redact || true
diff -u "$CONFIG" "$cand" | redact "$CONFIG" "$cand" || true
rm -f "$cand"; CAND=""
return 0
}
+131 -66
View File
@@ -71,20 +71,22 @@ set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MODULE="$REPO/fleetd"
JAR="$MODULE/target/fleetd.jar"
# fleetd #493: never build into the path a running process holds. The build writes here first
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
# stage_built_jar/swap_staged_jar below.
JAR_STAGED="$MODULE/target/fleetd-new.jar"
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
# Maven's own output directory and this script does not change it — but the daemon is launched
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
BUILD_JAR="$MODULE/target/fleetd.jar"
JAR="$MODULE/run/fleetd.jar"
OUT="$MODULE/fleetd.out"
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='target/fleetd.jar'
# Matches a fleetd daemon's command line wherever its jar sits — absolute or relative, under
# run/, under target/, or anywhere else a build or a hand-start might point it. Detecting a
# daemon this script did not start, including one running from a jar outside $JAR's own
# directory, is this pattern's whole job; running_pid()'s `comm = java` allowlist below is what
# keeps that breadth from counting a shell that merely types the pattern as literal text.
PATTERN='fleetd.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
@@ -169,10 +171,10 @@ hash256() {
fi
}
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
# jar right after a build (before it has been swapped in) without ever changing what a bare
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
@@ -180,15 +182,35 @@ hash256() {
# was the whole defect this ticket fixes.
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
report_jar_state() {
local built_hash running_hash built_mtime running_mtime
built_hash="$(jar_id "$BUILD_JAR")"
running_hash="$(jar_id "$JAR")"
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
fi
}
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
# runs on BSD.
@@ -210,7 +232,7 @@ jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# though: `PATTERN='fleetd.jar'` two lines up already assumes the daemon is a jar, which
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
@@ -229,16 +251,12 @@ running_pid() {
printf '%s' "$out"
}
# fleetd #493 — three small, independently testable pieces of "never build into the path a
# running process holds":
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
# process holds":
#
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
# path, immediately after a successful build. Dies (leaving the OLD daemon
# untouched — this runs before the stop step) if Maven reported success but
# left no jar behind, or if the move itself fails.
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
# jar already sitting at the live path from an earlier successful run, and
# die with the same truthful message this script has always used if not.
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
# sitting at the live path ($JAR, under run/) from an earlier successful run,
# and die with the same truthful message this script has always used if not.
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
# daemon actually exited — extracted to its own function so the main flow
# can be relied on to call swap_staged_jar only AFTER this returns success,
@@ -248,14 +266,10 @@ running_pid() {
# so this is never a write into a path a running process holds — by the time
# it runs, nothing holds that path anymore. If it fails, the caller must not
# start a new daemon: die() below already refuses that by exiting the script.
stage_built_jar() {
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
The running daemon was NOT touched."
mv -f "$JAR" "$JAR_STAGED" \
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
The running daemon was NOT touched."
}
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
# be moved off the live path right after compiling, because target/ was never
# the live path to begin with.
require_no_build_jar() {
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
}
@@ -282,8 +296,7 @@ swap_staged_jar() {
#
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
# the freshly built one to $JAR_STAGED by then) or on a stale one.
# every step succeeding while the daemon started on no jar at all or on a stale one.
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
# and compares line positions, and a same-line edit moves no line.
#
@@ -312,7 +325,12 @@ swap_if_built() {
local do_build="$1"
should_swap "$do_build" || return 0
say "swap"
swap_staged_jar "$JAR_STAGED" "$JAR"
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
# parent directory (a failing mv), and widening it to auto-create one would change what that
# test proves.
mkdir -p "$(dirname "$JAR")"
swap_staged_jar "$BUILD_JAR" "$JAR"
ok "jar in place: $(jar_id)"
}
@@ -610,6 +628,45 @@ check_log_path_matches_plist() {
ok "log path check: script and plist agree ($resolved_out)"
}
# Reads the launchd plist's ProgramArguments for the argument that follows "-jar", resolves it
# alongside $jar_path, and dies when the two differ. Call it only when the agent is loaded; it
# never touches launchd or the daemon itself.
check_jar_path_matches_plist() {
local jar_path="$1" plist_path="$2"
local plist_args plist_jar resolved_jar resolved_plist_jar
if ! plist_args="$(/usr/libexec/PlistBuddy -c 'Print :ProgramArguments' "$plist_path" 2>/dev/null)"; then
die "launchd agent is loaded but PlistBuddy could not read ProgramArguments from
$plist_path
— cannot verify which jar the supervised daemon launches. Fix the plist before
redeploying supervised."
fi
plist_jar="$(printf '%s\n' "$plist_args" | awk '
{ gsub(/^[ \t]+|[ \t]+$/, "") }
prev == "-jar" { print; exit }
{ prev = $0 }
')"
if [ -z "$plist_jar" ]; then
die "launchd agent is loaded but its ProgramArguments at
$plist_path
do not contain a '-jar <path>' pair — cannot verify which jar the supervised daemon
launches. Fix the plist before redeploying supervised."
fi
resolved_jar="$(cd "$(dirname "$jar_path")" 2>/dev/null && pwd -P)/$(basename "$jar_path")" || true
resolved_plist_jar="$(cd "$(dirname "$plist_jar")" 2>/dev/null && pwd -P)/$(basename "$plist_jar")" || true
if [ -z "$resolved_jar" ] || [ -z "$resolved_plist_jar" ] || [ "$resolved_jar" != "$resolved_plist_jar" ]; then
die "jar path mismatch — this script deploys to
$jar_path (resolved: ${resolved_jar:-<directory does not exist>})
but the loaded plist's ProgramArguments names
$plist_jar (resolved: ${resolved_plist_jar:-<directory does not exist>})
The swap renames the built jar into place, so the old path stops existing after a redeploy;
a launchd-initiated start from this plist (a reboot, or KeepAlive after a crash) would then
run java against a missing file. Reinstall the plist at
$plist_path
so its ProgramArguments names $jar_path before redeploying supervised."
fi
ok "jar path check: script and plist agree ($resolved_jar)"
}
# fleetd #552: the post-restart fresh-log capture, pulled out of the main flow so it is testable by
# sourcing (the same reason systemd_installed/systemd_loaded above guard their OWN mktemp inline
# instead of leaving it bare) even though its only caller sits below the SOURCED guard. By the time
@@ -826,16 +883,25 @@ report_shutdown_drain() {
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
# there is nothing here for the operator to lose track of.
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
# before the build even starts — so there is nothing here for the operator to lose track of.
# --no-build, staged jar absent -> "nothing changed"
drain_gate_refusal() {
local do_build="$1" staged_path="$2"
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
# already running before this gate fired (reaching this message at all requires OLD_PID to
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
# would NOT refuse, it would just restart that same old jar and silently throw away the one
# sitting at %s.
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
the freshly built jar is no longer at the live path that --no-build requires — or
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
would restart the daemon on that OLD jar and silently discard the one you just built — or
remove %s by hand if you want to discard this build instead.' \
"$staged_path" "$JAR" "$JAR" "$staged_path"
else
printf 'aborted — nothing changed'
fi
@@ -934,6 +1000,7 @@ report_supervisor_state() {
SUPERVISED=1
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
check_jar_path_matches_plist "$JAR" "$LAUNCHD_PLIST"
;;
systemd)
SUPERVISED=1
@@ -1170,7 +1237,7 @@ if [ -n "$OLD_PID" ]; then
else
warn "no daemon running — this will be a cold start"
fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
report_jar_state
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
@@ -1252,9 +1319,6 @@ stop_if_check_only "$CHECK_ONLY"
if [ "$DO_BUILD" = 1 ]; then
say "build"
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
# anything else, so that run's leftovers can never be mistaken for this run's output.
rm -f "$JAR_STAGED"
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
echo " log: $BUILD_LOG"
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
@@ -1264,16 +1328,16 @@ if [ "$DO_BUILD" = 1 ]; then
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
# daemon, if any, is still up at this point (build always runs before stop). From here until the
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
stage_built_jar
ok "jar now: $(jar_id "$JAR_STAGED")"
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
# again until the swap.
ok "jar now: $(jar_id "$BUILD_JAR")"
else
say "build skipped (--no-build)"
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
# sitting at the live path from an earlier successful run. Same check, same message as before.
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
# the live path from an earlier successful run. Same check, same message as before.
require_no_build_jar
fi
@@ -1282,8 +1346,8 @@ fi
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
# changed" would be a lie once a build has staged a jar.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
# changed" would be a lie once a build has produced a jar not yet swapped in.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
# ------------------------------------------------------------------ stop
#
@@ -1334,21 +1398,22 @@ fi
# ------------------------------------------------------------------ swap
#
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
swap_if_built "$DO_BUILD"
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins fleetd/.
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# above) and WorkingDirectory is already pinned to fleetd/.
say "start"
+143
View File
@@ -474,6 +474,90 @@ test_passphrase_key_is_redacted() {
assert_not_contains "FAKELEAK-PASSPHRASE" "$RUN_OUTPUT" "passphrase case: the passphrase VALUE must never leak"
}
# ------------------- acceptance criterion 19: the key line falls outside the printed hunk
# fleetd #656 — criteria 15a/15b both put the edit right next to the key line, so the key line is
# always inside diff -u's default 3-line context. Neither covers the actual case #639 fixed: an
# 8-line block-scalar body with only its SIXTH line changed, so the printed hunk (3 lines of
# context on each side of the change) covers body lines 3-8 and never includes the "token:" key
# line at all. The old, line-by-line redact() only ever masks after it has SEEN the key line go
# past; with the key line outside the hunk it never sets its mask, and the whole body — the
# changed line included — passes through raw. The control key sits right after the body, inside
# the same hunk, so the positive control below proves the fix is not simply printing nothing.
new_fixture_hunk_without_key_line() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
auth:
token: |
SECRET-LINE-1
SECRET-LINE-2
SECRET-LINE-3
SECRET-LINE-4
SECRET-LINE-5
SECRET-LINE-6
SECRET-LINE-7
SECRET-LINE-8
control: CTRL-MUST-APPEAR
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_key_line_outside_hunk_is_still_redacted() {
local dir
dir="$(new_fixture_hunk_without_key_line)"
sed 's/SECRET-LINE-6$/SECRET-LINE-6-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
start_run "$dir" 5 --from "$dir/candidate.yaml"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "hunk-without-key-line case reload exit code"
# Positive control FIRST: without this, a diff that printed nothing at all would pass the
# negative assertion right below identically to a correctly redacted one.
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "hunk-without-key-line case: the non-secret control line must still print unmasked"
assert_not_contains "SECRET-LINE-6-CHANGED" "$RUN_OUTPUT" "hunk-without-key-line case: the changed body line must never leak, even with the key line outside the printed hunk"
}
# ------------------------------- acceptance criterion 20: a blank line inside the value
# fleetd #656 — the old, line-by-line redact() reset its mask on any line whose indentation was
# not STRICTLY greater than the key's, and a wholly blank line has indentation 0, so it reset the
# mask exactly like the "control:" line that legitimately ends the block scalar. Everything after
# the blank line then printed raw. The current fix tracks masked lines by FILE line number instead
# of by indentation seen so far, so a blank line inside the value stays masked.
new_fixture_blank_line_in_value() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
printf 'bind:\n host: 127.0.0.1\n port: 19999\nbroker:\n uri: amqp://user:hunter2@host/vhost\nauth:\n token: |\n LEAK-BEFORE-BLANK\n\n LEAK-AFTER-BLANK\n control: CTRL-MUST-APPEAR\nprofiles:\n sonnet:\n weight: 3\n maxLoad: 5\n' > "$dir/fleetd.yaml"
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_blank_line_inside_value_is_still_redacted() {
local dir
dir="$(new_fixture_blank_line_in_value)"
sed 's/LEAK-AFTER-BLANK$/LEAK-AFTER-BLANK-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
start_run "$dir" 5 --from "$dir/candidate.yaml"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "blank-line-in-value case reload exit code"
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "blank-line-in-value case: the non-secret control line must still print unmasked"
assert_not_contains "LEAK-AFTER-BLANK-CHANGED" "$RUN_OUTPUT" "blank-line-in-value case: the line after the blank must never leak"
}
# ----------------------------------- acceptance criterion 16: a failing --set must not echo value
# fleetd #635 follow-up (ticket comment 17673, defect 8) — apply_set_pairs used to echo the FULL
# "$kv" (path=value, exactly as typed) in its yq-failure messages, so a broken --set with a
@@ -545,6 +629,57 @@ test_refusal_shape_from_parse_failure_wording_is_recognised() {
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
}
# --set runs yq over the whole candidate. It warns when that changes more lines than the requested
# pairs, but a simple file with only the intended changed line must stay quiet.
new_fixture_reformat_sensitive() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
# A section comment that documents the next block.
bind:
host: 127.0.0.1 # Keep this aligned with the port note.
port: 19999 # A fixture port.
# These comments use their placement as documentation.
profiles:
sonnet:
weight: 3
bootstrapText: >-
First line.
Second line.
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_set_warns_when_yq_reformats_extra_lines() {
local dir
dir="$(new_fixture_reformat_sensitive)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "reformat warning case reload exit code"
assert_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
"a --set that changes extra candidate lines must warn before installation"
}
test_set_stays_quiet_without_formatting_churn() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "no-reformat warning case reload exit code"
assert_not_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
"a --set that changes only its requested candidate line must not warn"
}
echo "== acceptance criterion 1: refusal restores byte for byte =="
test_refusal_restores_byte_for_byte
echo "== acceptance criterion 2: clean reload keeps the edit =="
@@ -573,6 +708,10 @@ echo "== acceptance criterion 15a: a block scalar's continuation lines are redac
test_block_scalar_continuation_is_redacted
echo "== acceptance criterion 15b: a passphrase key is also recognised =="
test_passphrase_key_is_redacted
echo "== acceptance criterion 19: the key line falls outside the printed hunk =="
test_key_line_outside_hunk_is_still_redacted
echo "== acceptance criterion 20: a blank line inside the value =="
test_blank_line_inside_value_is_still_redacted
echo "== acceptance criterion 16: a failing --set must not echo its value =="
test_failing_set_does_not_echo_its_value
echo "== extra: dry-run never installs, and redacts =="
@@ -581,5 +720,9 @@ echo "== extra: --check is read-only and always exits 0 =="
test_check_is_read_only_and_exits_zero
echo "== extra: the parse-failure refusal shape is also recognised =="
test_refusal_shape_from_parse_failure_wording_is_recognised
echo "== acceptance criterion 17: --set warns about yq formatting churn =="
test_set_warns_when_yq_reformats_extra_lines
echo "== acceptance criterion 18: --set stays quiet without formatting churn =="
test_set_stays_quiet_without_formatting_churn
printf 'PASS: config-edit acceptance criteria\n'
+213 -63
View File
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
test_running_pid_excludes_self_matching_wrapper_shell() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -475,11 +475,26 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
# `ps`, which is identical bash on every platform and carries no such platform question.
test_running_pid_finds_a_real_java_named_second_process() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
standin_pid=$!
sleep 0.3
after="$(running_pid)"
kill "$standin_pid" 2>/dev/null || true
wait "$standin_pid" 2>/dev/null || true
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
}
# PATTERN matches a fleetd jar in either build layout, not only the run/ one: a process whose
# argv names a jar under target/ must be found too, the same way the run/ case above is.
test_running_pid_finds_a_real_java_named_process_from_target_dir() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
@@ -489,7 +504,24 @@ test_running_pid_finds_a_real_java_named_second_process() {
kill "$standin_pid" 2>/dev/null || true
wait "$standin_pid" 2>/dev/null || true
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) naming a jar under target/: before=[$before] after=[$after]"
}
# Broadening PATTERN to match both build layouts must not also broaden it into matching a
# non-exec'ing shell that merely holds the target/ text as a literal argument, the same
# self-matching shape test_running_pid_excludes_self_matching_wrapper_shell above excludes for
# the run/ text.
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
kill "$wrapper_pid" 2>/dev/null || true
wait "$wrapper_pid" 2>/dev/null || true
[ "$after" = "$before" ] \
|| fail "running_pid() counted a self-matching wrapper shell (pid $wrapper_pid, holding 'target/fleetd.jar' as literal text in its own argv, not the daemon): before=[$before] after=[$after]"
}
# fleetd #593 CORRECTION 1, hole 2 — the round-1 filter denied known shell names (sh/bash/zsh/
@@ -557,29 +589,30 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
}
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
# rather than accidentally matching.
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
# matching.
test_jar_id_defaults_to_live_and_reports_explicit_path() {
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
local live_hash staged_hash default_result explicit_result
local dir saved_jar="$JAR"
local live_hash other_hash default_result explicit_result other_path
dir="$TMP/jar-id"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
printf 'live jar bytes' > "$JAR"
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
printf 'other jar bytes, not the same content' > "$other_path"
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
live_hash="$(hash256 "$JAR")"
staged_hash="$(hash256 "$JAR_STAGED")"
other_hash="$(hash256 "$other_path")"
default_result="$(jar_id)"
explicit_result="$(jar_id "$JAR_STAGED")"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
explicit_result="$(jar_id "$other_path")"
JAR="$saved_jar"
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
}
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
@@ -651,34 +684,77 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
}
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
# are exercised directly against real files on disk (not stubs), because the whole point is file
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
# content must print both labels and never warn.
test_report_jar_state_agrees_when_hashes_match() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'identical jar bytes' > "$BUILD_JAR"
printf 'identical jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'built jar' \
|| fail "report_jar_state did not label the built jar"
printf '%s' "$output" | grep -qF 'running jar' \
|| fail "report_jar_state did not label the running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars have identical content"
fi
}
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
test_report_jar_state_warns_when_hashes_differ() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'freshly built jar bytes' > "$BUILD_JAR"
printf 'older running jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'differ' \
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
}
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
test_report_jar_state_both_absent_is_not_a_mismatch() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'absent' \
|| fail "report_jar_state did not report absent for a missing built/running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars are simply absent"
fi
}
# Reads JAR and BUILD_JAR exactly as the script sources them, with nothing here assigning
# either first. JAR must resolve outside $MODULE/target/, and JAR must differ from BUILD_JAR:
# the daemon's live path and Maven's own build output are never the same file.
test_jar_and_build_jar_are_sourced_outside_target_and_differ() {
source "$ROOT/scripts/redeploy-fleetd.sh"
case "$JAR" in
"$MODULE"/target/*)
fail "\$JAR must not live under \$MODULE/target/ — got $JAR" ;;
esac
[ "$JAR" != "$BUILD_JAR" ] \
|| fail "\$JAR and \$BUILD_JAR must not be the same path — got $JAR"
}
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
# exercised directly against real files on disk (not stubs), because the whole point is file
# behavior (does the content move, does the source disappear, does a failure leave both sides
# intact) that a stubbed function cannot prove.
test_stage_built_jar_moves_off_live_path() {
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-ok"; mkdir -p "$dir"
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
printf 'built jar bytes' > "$jar"
JAR="$jar"; JAR_STAGED="$staged"
stage_built_jar || fail "stage_built_jar rejected a real build output"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
}
test_stage_built_jar_dies_when_build_produced_nothing() {
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-missing"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
output="$(stage_built_jar 2>&1)" || rc=$?
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|| fail "refusal message does not name the missing jar path"
}
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
# of that seam).
test_swap_staged_jar_moves_staged_onto_live() {
local dir staged live
dir="$TMP/swap-ok"; mkdir -p "$dir"
@@ -989,24 +1065,26 @@ test_swap_ordered_after_wait_and_before_start() {
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
}
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
# so the only way to pin its exact wording is to read the source.
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
# read the source.
test_drain_gate_abort_message_says_no_no_build() {
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
fail "abort message still claims a rerun WITH --no-build can finish the restart"
fi
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|| fail "abort message does not tell the operator to rerun without --no-build"
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|| fail "abort message does not say why --no-build cannot finish the restart"
printf '%s' "$msg" | grep -qF 'silently discard' \
|| fail "abort message does not say that --no-build would silently discard the build just made"
}
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
@@ -1149,19 +1227,22 @@ test_refuse_drain_gate_no_build_staged_absent() {
#
# fleetd #555 — the main flow's own call site moved: it used to read
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
# can do better than reading the source for this one specific gap).
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
# flow runs, so no test in this file can do better than reading the source for this one specific
# gap).
#
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
test_run_drain_gate_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
}
@@ -1269,12 +1350,71 @@ test_run_drain_gate_declined_reply_refuses() {
source "$ROOT/scripts/redeploy-fleetd.sh"
}
# Writes a launchd-plist fixture naming jar_path as the ProgramArguments entry after "-jar", so
# check_jar_path_matches_plist has something real to read back.
write_launchd_plist_fixture() {
local path="$1" jar_path="$2"
cat > "$path" <<PLIST
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>test.fixture</string>
<key>ProgramArguments</key>
<array>
<string>/usr/bin/java</string>
<string>-jar</string>
<string>$jar_path</string>
<string>fleetd.yaml</string>
</array>
</dict>
</plist>
PLIST
}
# Agreeing case: a plist whose ProgramArguments names the same jar, resolved, must proceed and
# say so, never die.
test_check_jar_path_matches_plist_agrees_ok() {
local dir jar plist output rc=0
dir="$TMP/jar-path-agree"; mkdir -p "$dir/run"
jar="$dir/run/fleetd.jar"
plist="$dir/agree.plist"
write_launchd_plist_fixture "$plist" "$jar"
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
[ "$rc" -eq 0 ] \
|| fail "check_jar_path_matches_plist must succeed when the plist names the same jar: $output"
printf '%s' "$output" | grep -qF 'jar path check' \
|| fail "check_jar_path_matches_plist did not print the agreement line: $output"
}
# The disagreeing case, and the positive control this check exists for: a plist naming a
# different jar must die, naming both paths.
test_check_jar_path_matches_plist_mismatch_dies() {
local dir jar other_jar plist output rc=0
dir="$TMP/jar-path-mismatch"; mkdir -p "$dir/run" "$dir/target"
jar="$dir/run/fleetd.jar"
other_jar="$dir/target/fleetd.jar"
plist="$dir/mismatch.plist"
write_launchd_plist_fixture "$plist" "$other_jar"
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] \
|| fail "check_jar_path_matches_plist must die when the plist names a different jar"
printf '%s' "$output" | grep -qF "$jar" \
|| fail "die message does not name this script's jar path: $output"
printf '%s' "$output" | grep -qF "$other_jar" \
|| fail "die message does not name the plist's jar path: $output"
}
# fleetd #555 item 4 — the report-state dispatch on $SUPERVISOR_KIND. Inverting this used to report
# the wrong supervisor and, for the launchd arm specifically, skip check_log_path_matches_plist.
CHECK_LOG_PATH_CALLED=0
CHECK_JAR_PATH_CALLED=0
stub_check_log_path_recorder() {
CHECK_LOG_PATH_CALLED=0
CHECK_JAR_PATH_CALLED=0
check_log_path_matches_plist() { CHECK_LOG_PATH_CALLED=1; }
check_jar_path_matches_plist() { CHECK_JAR_PATH_CALLED=1; }
}
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
@@ -1285,6 +1425,8 @@ test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state launchd must set SUPERVISED=1"
[ "$CHECK_LOG_PATH_CALLED" = 1 ] \
|| fail "report_supervisor_state launchd must call check_log_path_matches_plist"
[ "$CHECK_JAR_PATH_CALLED" = 1 ] \
|| fail "report_supervisor_state launchd must call check_jar_path_matches_plist"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
@@ -1296,6 +1438,8 @@ test_report_supervisor_state_systemd_sets_supervised_without_log_path_check() {
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state systemd must set SUPERVISED=1"
[ "$CHECK_LOG_PATH_CALLED" = 0 ] \
|| fail "report_supervisor_state systemd must NOT call check_log_path_matches_plist"
[ "$CHECK_JAR_PATH_CALLED" = 0 ] \
|| fail "report_supervisor_state systemd must NOT call check_jar_path_matches_plist"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
@@ -2153,6 +2297,8 @@ test_assert_single_daemon_accepts_one_pid
test_assert_single_daemon_rejects_two_pids
test_running_pid_excludes_self_matching_wrapper_shell
test_running_pid_finds_a_real_java_named_second_process
test_running_pid_finds_a_real_java_named_process_from_target_dir
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir
test_running_pid_drops_a_pid_whose_comm_is_not_java
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup
test_running_pid_counts_a_pid_whose_comm_is_java
@@ -2161,8 +2307,10 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
test_hash256_computes_a_real_sha256
test_jar_id_reports_absent_for_missing_file
test_jar_id_reports_unhashable_when_no_hasher_on_path
test_stage_built_jar_moves_off_live_path
test_stage_built_jar_dies_when_build_produced_nothing
test_report_jar_state_agrees_when_hashes_match
test_report_jar_state_warns_when_hashes_differ
test_report_jar_state_both_absent_is_not_a_mismatch
test_jar_and_build_jar_are_sourced_outside_target_and_differ
test_swap_staged_jar_moves_staged_onto_live
test_swap_staged_jar_dies_without_staged_file
test_swap_staged_jar_dies_when_mv_fails
@@ -2204,6 +2352,8 @@ test_drain_confirmed_false_on_anything_else
test_run_drain_gate_skips_prompt_when_not_required
test_run_drain_gate_confirmed_reply_does_not_refuse
test_run_drain_gate_declined_reply_refuses
test_check_jar_path_matches_plist_agrees_ok
test_check_jar_path_matches_plist_mismatch_dies
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path
test_report_supervisor_state_systemd_sets_supervised_without_log_path_check
test_report_supervisor_state_none_leaves_supervised_zero