Compare commits

...

26 Commits

Author SHA1 Message Date
Dai Ha a42b12440c Merge PR #649: fleetd #612 Shape A r9+r11 — pin capacitySource, healthCoverageSource, coordinator peers via a real fleet_list round-trip
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 1m52s
Test-only, 243 lines, one new file. ZERO production change — this is the point
of the PR, and it replaces PR #647, which added three public accessors to
FleetMcp on a false premise.

#647 argued a same-JVM test caller can never reach fleet_list's READ gate,
because FleetdAssembly hardcodes a real LsofPeerPidLookup, so the caller
resolves ANONYMOUS. The premise about lsof is true; the conclusion is not. In
CallerResolver.resolve, once c.terminal() == null the next branch is:

    if (tokenMode) {
        return presentedTokenMatches(authorizationHeader)
                ? Principal.primary(c.pid()) : Principal.anonymous();
    }

c.resolved() and c.scanComplete() guard only the LATER loopback-trust branch,
which token mode returns before reaching. So under auth.mode: token a bearer
token resolves to PRIMARY with no pid lookup on the path. This PR does exactly
that: real McpSyncClient callTool("fleet_list") over a real transport with an
Authorization: Bearer header, a faked leadMailboxOpener so no broker is touched,
and a real coordinator: block with a non-empty peers list. No reflection.

Lead verification, measured myself (not taken from the worker's report):

  git diff --stat origin/main -- fleetd/src/main/java/   -> EMPTY
  mutate :481 Fleetd.capacitySource(...) -> CapacitySource.none()
      -> Tests run: 3, Failures: 1
         fleetListReportsTheAssembledCapacitySource:222
         the other two tests stayed GREEN

The failure message carries the live fleet_list JSON body, showing
healthCoverage, loopHealth and coordinator.peers present with capacity absent —
so the round-trip, the PRIMARY resolution and the coordinator visibility are all
real, and the mutation removed exactly one thing.

Worker also reported, each with a grep -c anchor of 1 restored: health mutation
-> 1 failure named fleetListReportsTheAssembledHealthCoverageSource; peers
mutation -> 1 failure named fleetListReportsTheAssembledCoordinatorPeers; final
mvn clean install Tests run: 1899, Failures: 0 / BUILD SUCCESS. I reproduced the
capacity cycle only; the other two are the worker's measurement, not mine.

Carries forward #647's genuine find: the peers ternary at :489 is
inert-equals-absent, so the test configures a real coordinator: block with peers.

Detail on #612 and #647.
2026-10-02 04:53:56 +02:00
Dai Ha a06426c33c Merge PR #648: fleetd #612 Shape A r4 — pin quarantineSource + outageSource through BOTH operator windows
Test-only, 368 lines, one new file. No production change.

Lead verification, measured myself in a throwaway detached worktree (not taken
from the worker's report):

  starve FleetApp pass site :529   -> Tests run: 4, Failures: 2
                                      (...ByTheRealAssembledFleetApp:333, :363)
                                      both ...FleetMcp tests stayed GREEN
  starve FleetMcp pass sites :484/:486 -> Tests run: 4, Failures: 2
                                      (...ByTheRealAssembledFleetMcp:319, :349)
                                      both ...FleetApp tests stayed GREEN

So the two windows are pinned independently. Failure messages carry the live
JSON response body from a real round-trip, not a source-text match.

Why this is worth merging — the REST window was completely unpinned. With this
PR's test parked and :529 starved, the FULL suite reported:

  Tests run: 1892, Failures: 0, Errors: 0, Skipped: 0 / BUILD SUCCESS

Zero pre-existing tests notice the REST window losing its sources. An operator
reads GET /profiles exactly when the MCP mount is down. Arithmetic control:
1892 + this PR's 4 = 1896 = main at 6539efe.

Known caveat, recorded not fixed: the MCP-side quarantine assertion (:313) is
DUPLICATE coverage. A reviewer measured on clean origin/main that starving the
MCP quarantine arg already fails three pre-existing tests
(FleetdBackendQuarantineAssemblyTest:167, FleetdExhaustedPatternAssemblyTest:202,
FleetdOpenCodeExhaustionForwardingAssemblyTest:182), so that site was already
pinned and the PR's claim otherwise is wrong. Low severity, left in place as a
valid fleet_profiles output assertion. Detail on #648 and #612.

Reviewed by two reviewers against the diff (soundness: no issue; coupling: the
duplicate-coverage finding above).
2026-10-02 04:53:39 +02:00
Dai Ha 28ea0de575 fleetd #612: cover fleet list reporting sources
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Failing after 2m10s
2026-10-02 04:42:07 +02:00
Dai Ha 91792e11fc fleetd #612 Shape A unit r4: pin quarantineSource + outageSource through both operator windows
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Failing after 1m56s
FleetdAssembly.java's quarantineSource (:471-472) and outageSource (:473-476)
each feed two consumers: FleetMcp (fleet_profiles, :484/:486) and FleetApp
(GET /profiles, :529). No existing test distinguished the two windows for
either source.

New test drives the real FleetdAssembly.assembleAndStart, classifies a real
exhaustion/outage through the real CompletionResolver, and reads the result
back through a real McpSyncClient (fleet_profiles) and a real HttpClient
(GET /profiles), both authenticated via token-mode auth (sidesteps the
in-JVM pid-resolution dead end). No source text is read; no production code
changed.

Verified with six mutation cycles (3 per site x 2 sites: FleetMcp starved,
FleetApp starved, mis-wire with a disconnected collaborator), each run
against the full unfiltered suite and reverted after confirming the
expected test(s) alone went red. The quarantineSource FleetMcp-starve and
mis-wire cycles also trip three pre-existing tests that read the live
BackendQuarantine via FleetMcp#quarantineSource() for their own unrelated
assertions - a pre-existing incidental coupling, not newly introduced here.

Out of scope, noted per the ticket's dispatch comment: loopHealthSource
(FleetdAssembly.java:478) shares this same two-consumer shape and is
already assigned to a separate unit, r10.
2026-10-02 04:20:53 +02:00
ltms 6539efe9fa Merge PR #645: fleetd #612 Shape A r10 — pin loopHealthSource at BOTH consumers
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m34s
CI / build (push) Failing after 2m11s
Pins FleetdAssembly's loopHealth local at both of its pass sites: FleetMcp (:483, the fleet_list source) and FleetApp (:529, the real /healthz body). Two independent assertions, so starving one site leaves the other green.

Lead verification, independent of the implementer's own proof, in a throwaway detached worktree at 68397f7:
- :483 starved only -> Tests run: 2, Failures: 1. RED: fleetListUsesTheRunningLoopsInTheRealAssembledMcp, expected: <RUNNING> but was: <STOPPED>. REST test stayed GREEN.
- :529 starved only -> Tests run: 2, Failures: 1. RED: healthzUsesTheRunningLoopsInTheRealAssembledApp, with the real body {"loopHealth":{"sessionReaper":"STOPPED","statusPoller":"STOPPED"}}. MCP test stayed GREEN.
- full suite with the PR: Tests run: 1896, Failures: 0, Errors: 0 — BUILD SUCCESS (1894 baseline + 2).

The REST half binds port 0 (ephemeral) via runtime.app().start("127.0.0.1", 0), not the configured 8765, so it cannot clash with the live daemon. Teardown stops the bound app, closes the runtime, and asserts the shutdown hook was registered and the herdr client actually closed.

I chased the implementer's honestly-reported anomaly (an unrelated ClaudeCodeLauncherTest.noFixtureSeededTheDefaultClaudeJson failure in one cycle, and a suite total that moved between cycles). It is NOT caused by this PR: a baseline run at 68397f7 with this PR absent added the same 3 temp-dir project entries to the real ~/.claude.json as the run with it applied (65->68 without, 68->71 with), and ClaudeCodeLauncherTest passed in both. Filed separately.

Test-only diff, no production code touched.
2026-10-02 04:15:04 +02:00
ltms 68397f78d5 Merge PR #644: fleetd #612 Shape A r12 — pin the assembled turn registrar
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 1m32s
CI / build (push) Failing after 1m57s
Pins FleetdAssembly.java:349 behaviourally. The test installs a throwing wrapper on the real assembled Injector's listener to construct the narrow between-delivery-and-completion window, then asserts runtime.completion() still resolves the registered waiter. No sleep: it advances a controllable nano clock. Teardown runs the captured shutdown hook and asserts herdr actually closed.

Lead verification, independent of the implementer's own proof, in a throwaway detached worktree at dac5f88:
- unmutated: Tests run: 1, Failures: 0 — BUILD SUCCESS
- :349 registrar -> (_, _) -> { }: Failures: 1 — "must wire the Injector registrar to this runtime's real CompletionResolver" expected: <true> but was: <false>
- mis-wire I built myself (differs from the implementer's): Fleetd.turnRegistrar(new CompletionResolver(agents, new Rendezvous(), ...)) — a fresh Rendezvous so the registry genuinely differs: Failures: 1, same assertion
- reverted, full suite: Tests run: 1894, Failures: 0, Errors: 0 — BUILD SUCCESS, 58s (1892 baseline + r5 + r12)

Test-only diff, no production code touched. Reflection is used only to install the throwing listener; that is how the failure window is constructed, not how the assertion is made. No reviewer fan-out: member capacity is committed to the remaining Shape A implementers.
2026-10-02 04:09:29 +02:00
Dai Ha af901ff1d2 fleetd #612: pin assembled loop health
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 2m23s
2026-10-02 04:08:02 +02:00
ltms dac5f88812 Merge PR #643: fleetd #612 Shape A r5 — pin FleetdAssembly's leadConfigDirSource call site
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m14s
Pins the #602/#606 call site behaviourally: drives the real FleetdAssembly.assembleAndStart and asserts the assembled LeadConfigDirSource resolves a real configured configDir, which none() cannot produce.

Lead verification, run independently of the implementer's own proof, in a throwaway detached worktree at 141ae3b:
- unmutated: Tests run: 1, Failures: 0 — BUILD SUCCESS
- FleetdAssembly.java:488 -> FleetMcp.LeadConfigDirSource.none(): Tests run: 1, Failures: 1 — expected: </mnt/fake-lead-configdir> but was: <null>
- reverted, full suite: Tests run: 1893, Failures: 0, Errors: 0 — BUILD SUCCESS, 59s (baseline 1892)

Test-only diff, no production code touched. No reviewer fan-out was run: all member capacity is committed to the five Shape A implementers.
2026-10-02 04:05:50 +02:00
Dai Ha 2e663e5968 fleetd #612: pin assembled turn registrar
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Failing after 1m51s
2026-10-02 04:04:54 +02:00
Dai Ha 8c14ed2846 fleetd #612 Shape A r5: pin FleetdAssembly's leadConfigDirSource call site
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 1m0s
CI / build (pull_request) Failing after 2m4s
FleetdLeadConfigDirSourceWiringTest already pins Fleetd.leadConfigDirSource
itself, but by its own javadoc cannot cover whether FleetdAssembly.java:488
still calls it -- that call site could be swapped for a bare
FleetMcp.LeadConfigDirSource.none() (the literal fleetd #602/#606 defect)
and the whole suite would stay green.

Add FleetdLeadConfigDirSourceAssemblyTest: assembles the real FleetdRuntime
via FleetdAssembly.assembleAndStart, reads the leadConfigDirs field off the
real FleetMcp via reflection (no public accessor exists), and asserts it
resolves a configured lead's real configDir rather than none()'s hardcoded
null.

Verified: loud control (flip expected value) goes RED, reverts green;
mutation (i) inert none() at the call site goes RED; mutation (ii) mis-wire
(empty profile map, symbols otherwise intact) goes RED; both mutations
revert to an empty git diff. Full mvn clean install: 1893 tests, 0
failures, 0 errors, BUILD SUCCESS.
2026-10-02 04:02:18 +02:00
ltms 141ae3b04d Merge PR #640: fleetd #639 — redact()'s comment claimed more than the code does
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m32s
CI / build (push) Failing after 2m29s
Comment text only, no behaviour change. Corrects the paragraph added in d7f94ca so it no longer
claims the continuation masking holds for any masked key "present or future". It holds only while
the masked key's own line is inside the printed hunk; diff -u's three lines of context routinely
leave it out, and a blank line inside a block scalar drops the anchor too.

Verified: suite green in a clean copy of the edited tree (16 criteria + 3 extras, exit 0), bash -n
clean, and a control on the edit itself (old claim gone, #639 reference present). Evidence in #639.
2026-10-01 19:00:56 +02:00
Dai Ha faefea14c4 fleetd #639: redact()'s comment claimed more than the code does
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Failing after 2m11s
The paragraph added in d7f94ca ended with "This needs no knowledge of the key's
name and so protects a block scalar under any masked key, present or future."
The continuation masking is real and it is an improvement, but that sentence is
too strong: the masking only holds while the masked key's own line is inside the
hunk being printed.

redact() is fed `diff -u` output, which prints three lines of context. A block
scalar's body therefore often arrives with its key line left out. With no key
line, `masked` is never set and the body prints in full, with no "<redacted>"
anywhere. A blank line inside a block scalar loses the anchor the same way: a
blank diff line measures as indent 0, so `indent > masked_indent` is false and
the mask ends early — this time directly under a "<redacted>" marker.

Both were reproduced through the real script with --dry-run, each with a
positive control run first to prove the secret's lines actually reached the
output (without that control, "the secret never entered the diff" and "it
entered and was redacted" are indistinguishable). Filed as fleetd #639, which
also records that this is latent rather than live: today's fleetd.yaml holds 5
block scalars and all 5 sit under non-secret keys.

Comment text only. No change to redact() or to any other function, and no
change to the test suite. scripts/test-config-edit.sh still passes in a clean
copy (16 criteria + 3 extras, exit 0); bash -n clean.

The reason this is worth its own commit: #635 exists because an incomplete
redactor that looks complete is worse than one that visibly does nothing. A
comment that overstates the guarantee is the same defect in prose, and the next
session to read it has no other source.
2026-10-01 19:00:29 +02:00
ltms ea6896f2ef Merge PR #636: fleetd #635 — config-edit.sh, the one auditable way to edit fleetd.yaml
CI / shell-tests (push) Failing after 14s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m5s
Closes fleetd #635.

Verified by the lead before merge:
- scripts/test-config-edit.sh run in a PRISTINE copy of d7f94ca (git archive + git init, no
  worktree, no daemon): 16 criteria + 3 extras, exit 0.
- Diff read in full. 4 files, 1325 additions, 0 deletions, no .java/pom.xml/.yaml (checked with a
  positive control, so the negative is real) — no Maven gate needed.
- Defects 1-6 were verified by the previous lead by mutation; defects 7 and 8 (criteria 15a/15b/16)
  are fixed in d7f94ca and the fix was re-measured here independently.

Known limitation, filed as #639 and NOT a regression: redact()'s continuation masking only holds
while the masked key line is itself inside the printed diff hunk. diff -u prints three lines of
context, so a block scalar's body can appear without its key, and then nothing is masked; a blank
line inside a block scalar loses the anchor the same way. Reproduced through the real script with
a positive control proving the body reached the output. Latent, not live: today's fleetd.yaml has
5 block scalars and all 5 sit under non-secret keys. The pre-fix code leaked these cases too, so
this commit is a strict improvement.

Two sentences in config-edit.sh's redact() comment and in this PR's body claim the masking holds
regardless of the key's name. That is too strong; see #639. Being corrected in a follow-up.
2026-10-01 18:58:27 +02:00
Dai Ha d7f94cafa2 fleetd #635 follow-up: redact() masks block-scalar continuations + passphrase; --set failures stop echoing the value (defects 7 and 8, criteria 15a/15b/16)
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m42s
CI / build (pull_request) Failing after 1m50s
Defect 7 (comment 17670): redact() only masked a line that itself started with a
secret-looking key, so a YAML block scalar's value leaked on the lines that
followed the key while the key line right above it printed a reassuring
"<redacted>". Fixed by tracking the masked key's own indentation and masking
every following line indented deeper than it, stopping once indentation returns
to the key's level or shallower; the diff's leading +/-/space marker is stripped
before indentation is measured, per the comment's own pitfall. "passphrase" is
now also in the key-name backstop.

Defect 8 (comment 17673): apply_set_pairs echoed the operator's full
"path=value" input, unredacted, in both of its yq-failure die messages — a
failing --set with a secret-looking value printed that value right back. Fixed
to print only the path; deliberately not routed through redact, which would
pass a non-"key: value"-shaped string straight through.

Adds acceptance criteria 15a (block-scalar continuation), 15b (passphrase key),
and 16 (failing --set never echoes its value) to scripts/test-config-edit.sh,
each with a positive control proving the relevant line really was in the
printed output before asserting the secret is absent. All three confirmed RED
against the pre-fix code and GREEN after, in isolation, before being folded
into the full suite (16 criteria + 3 extras, exit 0).

Also updates PR #636's description per comment 17671: the redact() sentence now
names the continuation-masking rule and says plainly that the key-name list is
a backstop, never a complete list.
2026-10-01 18:47:00 +02:00
Dai Ha 096f08c866 fleetd #635 follow-up: fix the stale --restore not-found message (defect 6, criterion 14)
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 2m48s
The --restore "no backup found" message still printed the old beside-the-config glob
(${CONFIG}.bak.*) even though newest_backup had already moved to searching the managed
.config-backups/ directory. The message was left behind when the search moved — the search
itself was already correct (ticket comment 17664). Fix is reporting-only: the message now
names the directory actually searched (via backup_dir_for), and separately says that a
backup written the old way, directly beside the config, is not searched any more, with the
one-line cp to recover one by hand. No search fallback was added — reading backups from
outside the managed directory stays unsupported, as instructed.

Acceptance criterion 14 proves both directions: the not-found message names the real
directory (confirmed red on the pre-fix code, green after), and a restore with a real backup
present in .config-backups/ still succeeds (confirmed this catches an "always not-found"
regression that direction 1 alone would miss).

All 14 criteria plus 3 extras pass in scripts/test-config-edit.sh.
2026-10-01 18:25:08 +02:00
Dai Ha 4eb720029c fleetd #635 follow-up: refuse empty --set values, gitignore backups, preserve file mode
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m43s
CI / build (pull_request) Failing after 1m52s
Five fixes against PR #636, all verified by the lead's own review and reproduced here:

1. --set .a.b= (a forgotten value) is now refused outright instead of silently nulling the
   field — a null numeric config value falls back to its default rather than erroring, which
   widens capacity silently instead of failing loudly. A deliberate clear gets its own spelling,
   --set .a.b=null, which writes a literal YAML null via yq, never through strenv(). (criteria
   9, 10)

2. Backups move from beside fleetd.yaml to a dedicated fleetd/.config-backups/ directory,
   gitignored at the repo root (so it also covers scripts/test-config-edit.sh's own throwaway
   fixtures) and in fleetd/.gitignore, plus a fleetd.yaml.bak.* glob backstop for any stray
   backup written the old way. A backup of a file that must never be committed inherits that
   requirement. (criterion 11)

3. The live config's file mode now survives both an edit and a restore. mv from a mktemp
   candidate used to carry mktemp's 0600 onto the live path forever, and cp onto an existing
   file keeps the destination's mode, so a restore did not undo it either. (criterion 12)

4. A global CAND + single EXIT/INT/TERM trap prevents an uninstalled .config-edit.XXXXXX
   candidate from leaking if the script is interrupted mid-run. No acceptance criterion is
   gated on this — a reproducible leak could not be made to happen on demand — but it is cheap
   and obviously right.

5. Acceptance criterion 7's redaction check gained a positive control: it now asserts the
   output actually CONTAINS the redaction marker and the changed key, not only that it lacks
   the secret. The prior two assertions were negative-only and passed just as happily when the
   diff was never printed at all — confirmed by reproducing the lead's own mutation (deleting
   the redacted diff print on the edit path) and watching it survive the old test and get
   caught by the new one. (criterion 13)

All 13 acceptance criteria plus 3 extras pass in scripts/test-config-edit.sh. Criteria 9, 10,
11, 12 and 13 were each proven non-vacuous: criteria 9/10 by mutating the test's own expected
value and watching it fail by name, then reverting; criteria 11/12/13 by reverting or mutating
the corresponding fix in config-edit.sh and watching the matching criterion fail by name, then
restoring the fix and re-confirming a clean pass.
2026-10-01 18:07:12 +02:00
Dai Ha 0db6d31dc2 fleetd #635: add scripts/config-edit.sh, the one auditable way to edit fleetd.yaml
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 2m12s
CI / build (pull_request) Failing after 2m48s
Backs up, builds a candidate off the live file, parse-checks it with yq before
install, installs atomically, then reads the daemon's own ConfigRef reload
verdict back out of fleetd.out (marked from before the edit, so a stale line
can never be mistaken for this edit's result). Four exit codes: 0 clean, 3
needs a restart, 4 refused (backup restored), 5 cannot tell (nothing
restored, printed --restore command). Every diff is redacted.

scripts/test-config-edit.sh drives it end to end against fixtures in a
throwaway temp dir, with no daemon involved.
2026-10-01 17:29:04 +02:00
Dai Ha 158a2a84b5 Merge PR #632: fleetd #612 ranks 6+7 + #630 — behavioural pins for lifecycle wirings
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m13s
CI / build (push) Failing after 2m10s
Pins three call sites the assembly owns and nothing observed:
- healthFailTarget (FleetdAssembly:429) — inert, a dead member's waiting ticket sits
  PENDING for the full 30-minute async timeout instead of failing immediately.
- releaseCleanup (:447) — inert, every teardown leaks three things: a stuck rendezvous
  waiter, an unreleased reply-inbox consumer, and a stale lead binding.
- requireOperatorConfirm (:402/:409, fleetd #630) — dropping the 14th constructor
  argument selects #621's 13-arg overload, which hardcodes true, silently reverting
  the operator's fix on a host that set requireOperatorConfirm: false. Pinned in both
  directions, plus an assertion that the two notice strings differ, so no constant
  satisfies both.

Verified by the lead beyond the worker's proof: its releaseCleanup mutation killed all
three cleanups at once and so proved only the first assertion had teeth. Starving them
one at a time — abandon kept, release starved; then abandon and release kept, forget
starved — each fails its own named assertion. All three leaks are pinned independently.

A review pass found these three tests assembled real schedulers and never tore them
down, by any route: no close(), no shutdownHook, no @AfterEach, no finally. Surefire
runs one JVM fork for the whole suite, so those loops outlived their tests. Fixed by
capturing the hook and running it in a finally, with an assertion on FakeHerdr.closed
so the teardown itself is pinned rather than assumed.

MERGE RESOLUTION BY THE LEAD, the same collision as #633. These 3 test files each add
a ResourcePorts fake, and #633 landed first making herdrPollWait() abstract with no
default. Git reported a clean merge that did not compile. Added the override to all 3,
matching the established convention for an always-healthy fake — a Runnable that
throws, verified first that none of the three uses healthy(false), so the tripwire can
only fire if the test's herdr behaviour changes.

Full suite on the resolved merge: 1892 tests, 0 failures (1889 + this branch's 3).
2026-10-01 16:52:39 +02:00
Dai Ha 41ebc9cf69 Merge PR #633: fleetd #629 + #625 — the two ResourcePorts seams
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 2m57s
#629: the herdr boot wait went through ports.nanoClock() but hardcoded
Fleetd::sleepHerdrPoll, so a test assembling against an unhealthy lead herdr burned
30 real seconds whatever clock it injected. The poll sleep now goes through
ResourcePorts.herdrPollWait(). FleetdAssemblyFleetAppTest's lead-down test drops from
30.276s to 0.062s. It also gains @Timeout(10, SEPARATE_THREAD) at class level: the
pin's failure mode is otherwise an infinite hang, because the test's fake clock only
advances when herdrPollWait() is called. SAME_THREAD cannot interrupt a real
Thread.sleep, so the thread mode is load-bearing, not decoration.

#625: guard.assertPrimaryClean(System.getenv()) could be deleted with a fully green
suite — the check behind charter invariant 1, which keeps the primary on the
operator's subscription. main() now delegates to main(String[], ResourcePorts) and
the guard reads ports.environment(), so a test can taint the environment without
touching the real process env. The guard still runs before cfg.validateAll() and
before any socket, broker or HTTP work; two tests with different fixtures pin that
ordering as two independently falsifiable claims, not one.

MERGE RESOLUTION BY THE LEAD. ResourcePorts.herdrPollWait() is abstract with no
default, by design (#629 keeps ResourcePorts free of a none() default). This branch
patched the 8 implementations that existed when it forked. PRs #631 and #634 merged
ahead of it and added 3 more fakes, so git reported a clean 14-file merge that did
not compile:

  FleetdAssemblyAmqpOpenersTest.RecordingPorts is not abstract and does not
  override abstract method herdrPollWait() in dev.ltms.fleet.ResourcePorts

(plus the two ControllableResourcePorts in #634's tests). I added the override to
those 3, matching this branch's own convention for an always-healthy fake: return a
Runnable that throws, so if one of those assemblies ever does start polling herdr it
fails loudly instead of sleeping quietly. All 13 implementations now carry it.

Full suite on the resolved merge: 1889 tests, 0 failures (1886 + this branch's 3).
2026-10-01 16:50:10 +02:00
Dai Ha 180c953c42 Merge PR #634: fleetd #612 ranks 1+2 — behavioural pins for the exhaustion wirings
CI / shell-tests (push) Failing after 8s
CI / build (push) Failing after 2m20s
CI / contract (push) Successful in 2m22s
Pins all four exhaustion call sites in FleetdAssembly: liveExhaustedPatterns (:299),
exhaustedPatternLookup (:300), publishExhaustionSink (:327) and the independent
OpenCode forwardingExhaustionSink (:164). Rank 1 is the worst consequence in the
#612 sweep — an inert lookup hands a genuine usage-limit refusal back to a waiting
caller as real completed work instead of BACKEND_EXHAUSTED.

Two separate tests, so rank 2's OpenCode half is pinned independently: mutating
:164 fails only the forwarding test, which is the independence the ticket asserts.

Verified by the lead beyond the worker's own proof: wiring publishExhaustionSink to
a throwaway BackendQuarantine — every symbol kept at the call site, only the
collaborator identity changed — is caught by both tests. So these pins survive
mis-wiring, not just deletion.

Test-only; no production change. Tears down each assembly via the captured
shutdown hook in a finally.
2026-10-01 16:45:20 +02:00
Dai Ha 4af919af0c fleetd #629 follow-up: bound FleetdAssemblyFleetAppTest with a SEPARATE_THREAD @Timeout
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 2m25s
The mutation cycle for #629 proved deleting the fix (reverting FleetdAssembly's awaitHerdr
call back to a hardcoded Fleetd::sleepHerdrPoll) doesn't just make a test fail — it hangs
forever, because the test's fake nanoClock() only advances when ports.herdrPollWait() is
actually called. SAME_THREAD @Timeout (JUnit's default) can't catch that: it only measures
elapsed time after the test method returns on its own, which never happens here. A class-level
@Timeout(10s, SEPARATE_THREAD) does, since it runs the test on its own thread and interrupts it
on timeout. Verified by re-running the same mutation: the suite now fails fast with a named
TimeoutException instead of hanging indefinitely.
2026-10-01 16:41:40 +02:00
Dai Ha 09c37061c1 Merge PR #631: fleetd #612 ranks 3+8 — behavioural pins for the AMQP assembly openers
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 59s
CI / build (push) Failing after 3m14s
Pins FleetdAssembly's replyInboxOpener (:357) and leadMailboxOpener (:363) so an
inert opener can no longer pass a green suite. Both directions covered: a durable
opener must leave the runtime owning the exact object it returned, and a failing
one must leave the in-memory fallback with coordination off — with the startup
report agreeing with the real state in each case.

Verified by the lead beyond the worker's own proof: a mutation that still CALLS
ports.replyInboxOpener() and discards the result is caught by assertSame, which is
rank 3's live defect (opener called, result thrown away, log still printing
'reply inbox: AMQP broker (durable)').

Test-only; no production change.
2026-10-01 16:38:16 +02:00
Dai Ha 2be287ea03 fleetd #612 step 4 (ranks 1&2): behavioural assembly tests for the exhaustion wirings
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 3m5s
Replaces nothing (no existing test covered these call sites through the real
assembly); adds two new FleetdAssembly-driven tests, matching Unit A's pattern
of inspecting FleetdRuntime's real, assembled objects rather than a copy.

- FleetdExhaustedPatternAssemblyTest pins the liveExhaustedPatterns/
  exhaustedPatterns call sites (rank 1 — "the worst consequence in the whole
  sweep": a genuine usage-limit refusal handed back as real completed work)
  together with publishExhaustionSink (rank 2, non-OpenCode half): drives a
  real CompletionResolver through a pane scrape matching a configured
  exhaustedPattern and asserts BACKEND_EXHAUSTED classification plus a real
  BackendQuarantine credential quarantine.

- FleetdOpenCodeExhaustionForwardingAssemblyTest pins forwardingExhaustionSink
  (rank 2, OpenCode half — independent of publishExhaustionSink per the
  ticket) by reflectively reaching the real, assembled OpenCodeLauncher's
  exhaustionSink field (SessionManager.launcher -> CompositePeerLauncher.
  byProfile -> OpenCodeLauncher.exhaustionSink) and proving it forwards into
  the same production BackendQuarantine.

All four call sites (FleetdAssembly.java:164,299,300,327) were each put
through grep-anchor -> line-anchored sed mutation -> mvn -o compile -> full
mvn -o test (named test RED) -> restore -> full mvn -o test (1885/0 GREEN).
Mutating line 327 alone also fails the OpenCode test, confirming the
documented construction-order dependency (forwardingExhaustionSink reads
exhaustionSinkRef, which publishExhaustionSink sets) without weakening either
site's independent pin.
2026-10-01 16:38:00 +02:00
Dai Ha c3b0406826 fleetd #612: isolate fallback AMQP reports
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Failing after 2m20s
2026-10-01 16:33:45 +02:00
Dai Ha e76fa1660b fleetd #629/#625: inject the herdr poll wait and the startup env read through ResourcePorts
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Failing after 2m42s
#629: ResourcePorts gains herdrPollWait() (SystemResourcePorts: Fleetd::sleepHerdrPoll).
FleetdAssembly's awaitHerdr call now takes its poll wait from ports instead of a hardcoded
Thread.sleep, so a test's fake clock can actually reach the deadline without burning real
wall-clock time. FleetdAssemblyFleetAppTest's down-lead case drops from ~30s to well under 1s.
All other ResourcePorts fakes get a trivial "must not be called" override since their herdr is
always healthy and never polls.

#625: Fleetd.main(String[]) now delegates to a new package-private
main(String[], ResourcePorts) overload that reads the startup environment via
ports.environment() instead of System.getenv(), reusing the same ports instance for
FleetdAssembly.assembleAndStart. The guard call stays at its original point, before
cfg.validateAll() and before any assembly/socket/broker/HTTP work. New
FleetdSubscriptionGuardOrderingTest drives the real main() with a tainted fake environment and
pins both presence and ordering against validateAll() and against assembly's first port call,
plus a clean-env control proving the guard only blocks on an actual taint.
2026-10-01 16:33:15 +02:00
Dai Ha 91be5079d6 fleetd #612: pin AMQP assembly openers
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m44s
CI / build (pull_request) Failing after 3m41s
2026-10-01 16:26:02 +02:00
29 changed files with 3479 additions and 4 deletions
+7
View File
@@ -20,6 +20,13 @@ fleetd.out
fleetd/fleetd.out
logs/
# fleetd #635 follow-up — scripts/config-edit.sh's backup directory. No leading slash, so this is
# ignored at every depth: the real one lives under fleetd/ (also named in fleetd/.gitignore, next
# to the config it backs up), and scripts/test-config-edit.sh's own throwaway fixtures build one
# under the repo root while the suite runs. --config can point anywhere, so the directory name is
# ignored everywhere rather than only where the live daemon happens to use it.
.config-backups/
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
+7
View File
@@ -7,6 +7,13 @@ dependency-reduced-pom.xml
fleetd.yaml
bridged.yaml
# fleetd #635 follow-up — scripts/config-edit.sh's backups of fleetd.yaml. A backup of a file
# that must never be committed inherits that requirement. The directory is the real protection
# (it keeps working even if the backup naming changes); the glob is a backstop for a stray
# backup written the old way, directly beside fleetd.yaml, or by an older copy of the script.
.config-backups/
fleetd.yaml.bak.*
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
@@ -131,6 +131,27 @@ public final class Fleetd {
}
static void main(String[] args) {
main(args, ResourcePorts.system());
}
/**
* fleetd #625: package-private so a test can drive the literal startup sequence — not a copy
* of it — against a non-production {@link ResourcePorts} whose {@link
* ResourcePorts#environment()} is tainted, without ever touching the real process environment.
* The real process environment is exactly what a test cannot taint from inside the JVM, which
* is why nothing could pin {@link SubscriptionGuard#assertPrimaryClean} running at this call
* site before this ticket.
*
* <p>{@link #main(String[])} is the one production caller, passing {@link
* ResourcePorts#system()}. Every statement below, in the same order, is otherwise unchanged
* from before this ticket — in particular the guard still runs here, before {@code
* cfg.validateAll()} and before {@link FleetdAssembly#assembleAndStart} ever touches a socket,
* a broker, or HTTP, exactly as it always has. {@code ports} is reused for the guard check and
* then handed on to the assembly, rather than a second instance being constructed there, so a
* test's fake backs the whole boot path with one consistent view — see {@code
* FleetdSubscriptionGuardOrderingTest}.
*/
static void main(String[] args, ResourcePorts ports) {
Path configPath = args.length > 0 ? Path.of(args[0]) : chooseDefaultConfigFile(Path.of(""));
// The config file was renamed bridged.yaml -> fleetd.yaml. Name the file we actually
// loaded, whichever of the two names it carries.
@@ -166,7 +187,7 @@ public final class Fleetd {
// The primary/host env that launched fleetd must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
guard.assertPrimaryClean(ports.environment());
// Every FleetConfig.validateXxx() the operator's config can fail — CB-501's auth-exposure
// check, CB-531's lead-tab-prefix check, CB-542's subscription-profile check, the charter
@@ -189,7 +210,9 @@ public final class Fleetd {
// after validateAll() (this line) puts both back under test, in the same relative order,
// before either one does any I/O — see FleetdAssembly's javadoc for the full boot-order
// contract this preserves exactly.
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ResourcePorts.system());
// fleetd #625: the same `ports` the guard check above just used, not a second
// ResourcePorts.system() instance — see this method's own javadoc.
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
}
/**
@@ -194,7 +194,7 @@ final class FleetdAssembly {
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// exists. Wait, then degrade rather than die: serving with /healthz reporting "degraded" is
// strictly more useful than exiting.
Fleetd.HerdrAwaitOutcome herdrOutcome = Fleetd.awaitHerdr(herdr, ports.nanoClock(), Fleetd::sleepHerdrPoll);
Fleetd.HerdrAwaitOutcome herdrOutcome = Fleetd.awaitHerdr(herdr, ports.nanoClock(), ports.herdrPollWait());
boolean herdrUp = Fleetd.logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
@@ -43,6 +43,19 @@ public interface ResourcePorts {
/** A monotonic elapsed-time clock. Production: {@link System#nanoTime()}. */
LongSupplier nanoClock();
/**
* fleetd #629: the per-poll wait {@code FleetdAssembly#assembleAndStart} passes to {@code
* Fleetd#awaitHerdr} while polling for herdr's socket. Production: {@link
* Fleetd#sleepHerdrPoll()} — a real {@code Thread.sleep}. {@link #nanoClock()} alone is not
* enough to make {@code awaitHerdr}'s deadline controllable: the old call site passed {@code
* Fleetd::sleepHerdrPoll} directly, hardcoded, so a test that injected a fake clock still had
* to wait out the real sleep between each poll to ever reach the deadline — the clock looked
* injected and was not actually controllable. A test supplies a no-op that advances its own
* injected {@link #nanoClock()} instead, so the deadline becomes reachable without any real
* wall-clock time passing.
*/
Runnable herdrPollWait();
/**
* A wall-clock reading, in nanoseconds. Production: {@code System.currentTimeMillis()}
* converted to nanoseconds. Kept separate from {@link #nanoClock()} because {@link
@@ -44,6 +44,11 @@ final class SystemResourcePorts implements ResourcePorts {
return System::nanoTime;
}
@Override
public Runnable herdrPollWait() {
return Fleetd::sleepHerdrPoll;
}
@Override
public LongSupplier wallClockNanos() {
return () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
@@ -0,0 +1,183 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.LeadChannel;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 step 4 ranks 3 and 8: the assembled daemon must use the AMQP openers from {@link
* ResourcePorts}, and its startup report must describe the object the runtime actually owns. These
* fakes never open a socket.
*/
class FleetdAssemblyAmqpOpenersTest {
private static final String COORD_ID = "assembly-test";
private static final class DurableReplyInbox implements ReplyInbox {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
}
private static final class DurableLeadMailbox implements LeadChannelHandle {
@Override public void publish(String toCoordId, LeadMessage message) { }
@Override public List<LeadMessage> peek() { return List.of(); }
@Override public void ack(String msgId) { }
@Override public String selfCoordId() { return COORD_ID; }
@Override public boolean heldDurable() { return true; }
@Override public MailboxState inspect(String coordId) { return MailboxState.unknown(coordId); }
@Override public void close() { }
}
private static final class RecordingPorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final DurableReplyInbox replyInbox = new DurableReplyInbox();
final DurableLeadMailbox leadMailbox = new DurableLeadMailbox();
final AtomicInteger replyOpenCalls = new AtomicInteger();
final AtomicInteger mailboxOpenCalls = new AtomicInteger();
final boolean openSucceeds;
RecordingPorts(boolean openSucceeds) {
this.openSucceeds = openSucceeds;
}
@Override public Map<String, String> environment() { return Map.of(); }
@Override public HerdrClient connectHerdr(Path socketPath) { return herdr; }
@Override public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> {
replyOpenCalls.incrementAndGet();
if (!openSucceeds) throw new IllegalStateException("fake reply broker is down");
return replyInbox;
};
}
@Override public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfId, prefetch) -> {
mailboxOpenCalls.incrementAndGet();
if (!openSucceeds) throw new IllegalStateException("fake coordination broker is down");
return leadMailbox;
};
}
@Override public LongSupplier nanoClock() { return System::nanoTime; }
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
@Override public LongSupplier wallClockNanos() { return System::nanoTime; }
@Override public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override public void addShutdownHook(Runnable hook) { }
@Override public void startHttp(Javalin app, String host, int port) { }
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Files.createDirectories(dir);
Path config = dir.resolve("fleetd.yaml");
Files.writeString(config, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-reply-broker/vh"
coordinator:
uri: "amqp://fake-coordination-broker/vh"
selfId: "assembly-test"
""");
return FleetConfig.load(config);
}
private static FleetdRuntime assemble(Path dir, RecordingPorts ports) throws Exception {
FleetConfig cfg = writeConfig(dir);
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, new ConfigRef(dir.resolve("fleetd.yaml"), cfg),
new SubscriptionGuard(cfg.guard().hostSet())), ports);
}
private static boolean reportContains(ListAppender<ILoggingEvent> appender, String text) {
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).anyMatch(message -> message.contains(text));
}
@Test
void assembledAmqpOpenersAndTheirReportsAgreeOnDurableAndFallbackStates(@TempDir Path dir) throws Exception {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
Level oldLevel = logger.getLevel();
ListAppender<ILoggingEvent> reports = new ListAppender<>();
reports.start();
logger.setLevel(Level.INFO);
logger.addAppender(reports);
try {
RecordingPorts durablePorts = new RecordingPorts(true);
FleetdRuntime durable = assemble(dir.resolve("durable"), durablePorts);
try {
// Control: this fails loudly if the assembly did not run or used an inert opener.
assertEquals(1, durablePorts.replyOpenCalls.get(), "assembly must call replyInboxOpener once");
assertEquals(1, durablePorts.mailboxOpenCalls.get(), "assembly must call leadMailboxOpener once");
assertSame(durablePorts.replyInbox, durable.replyInbox(),
"the durable reply report must describe the exact inbox the runtime owns");
assertSame(durablePorts.leadMailbox, durable.leadMailbox(),
"the coordination-on report must describe the exact mailbox the runtime owns");
assertNotNull(durable.leadCoordLoop(), "a durable mailbox must start lead coordination");
assertTrue(reportContains(reports, "reply inbox: AMQP broker (durable)"));
assertTrue(reportContains(reports, "lead coordination: ON as coord-id " + COORD_ID));
} finally {
durable.close();
}
reports.list.clear();
RecordingPorts fallbackPorts = new RecordingPorts(false);
FleetdRuntime fallback = assemble(dir.resolve("fallback"), fallbackPorts);
try {
assertEquals(1, fallbackPorts.replyOpenCalls.get(), "assembly must call the failing reply opener once");
assertEquals(1, fallbackPorts.mailboxOpenCalls.get(), "assembly must call the failing mailbox opener once");
assertTrue(fallback.replyInbox() instanceof InMemoryReplyInbox,
"a failed reply opener must make the runtime own the in-memory fallback");
assertNull(fallback.leadMailbox(), "a failed mailbox opener must leave coordination off");
assertNull(fallback.leadCoordLoop(), "coordination must not start without a mailbox");
assertTrue(reportContains(reports, "reply inbox: in-memory (soft-state)"));
assertTrue(reportContains(reports, "lead-to-lead messaging is OFF"));
} finally {
fallback.close();
}
} finally {
logger.detachAppender(reports);
logger.setLevel(oldLevel);
}
}
}
@@ -117,6 +117,14 @@ class FleetdAssemblyConnectionIdentityTest {
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind — this test never issues a real HTTP request.
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private FleetdRuntime runtime;
@@ -140,6 +140,14 @@ class FleetdAssemblyCoordinatorLifecycleTest {
public void startHttp(Javalin app, String host, int port) {
// No real HTTP bind in a unit test.
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox {
@@ -8,6 +8,7 @@ import dev.ltms.fleet.herdr.HerdrClient;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.Timeout;
import org.junit.jupiter.api.io.TempDir;
import java.net.URI;
@@ -21,6 +22,8 @@ import java.util.Map;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
@@ -69,13 +72,33 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
* {@code /sessions} correctly once handed two clients.
*
* <p><strong>fleetd #629 follow-up.</strong> The fix below (see {@link TwoHerdrResourcePorts})
* makes {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}'s fake {@code
* nanoClock()} frozen unless {@code herdrPollWait()} itself advances it. That is a sharper pin
* than an assertion — if a future edit to {@code FleetdAssembly} ever bypasses {@code
* ports.herdrPollWait()} again (e.g. reverting to a hardcoded {@code Thread.sleep}), the clock
* never advances, {@code Fleetd#awaitHerdr}'s deadline is never reached, and this test hangs
* forever instead of failing — proven by deliberately reintroducing that exact regression while
* fixing this ticket. {@code @Timeout} turns that silent hang into a bounded, named test failure:
* {@code SEPARATE_THREAD} so JUnit's timeout governor can actually interrupt a thread stuck in a
* real {@code Thread.sleep} loop (the default {@code SAME_THREAD} mode cannot — it only measures
* elapsed time after the test method returns on its own, which never happens here). 10 seconds is
* roughly 150x the real passing times measured here (~0.06s), so a slow CI machine has no reason
* to flake, and it is still 3x faster than discovering the regression by burning a CI job's whole
* wall-clock budget.
*/
@Timeout(value = 10, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
class FleetdAssemblyFleetAppTest {
private static final class TwoHerdrResourcePorts implements ResourcePorts {
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
// fleetd #629: a fake, advanceable clock — NOT System::nanoTime. awaitHerdr's poll wait
// (herdrPollWait() below) advances this on every poll instead of sleeping for real, so the
// down-lead test below reaches awaitHerdr's deadline without burning real wall-clock time.
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
Runnable shutdownHook;
@Override
@@ -107,7 +130,7 @@ class FleetdAssemblyFleetAppTest {
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
return nowNanos::get;
}
@Override
@@ -115,6 +138,13 @@ class FleetdAssemblyFleetAppTest {
return System::nanoTime;
}
@Override
public Runnable herdrPollWait() {
// fleetd #629: advance the fake clock instead of a real Thread.sleep, so awaitHerdr's
// deadline is reached in real time regardless of the configured poll interval.
return () -> nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(1));
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
@@ -78,6 +78,13 @@ class FleetdAssemblyHealthFailTargetBehaviouralTest {
};
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
@@ -126,6 +126,14 @@ class FleetdAssemblyLifecycleTest {
ledger.add("startHttp");
this.startedApp = app;
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
/** A fake {@link ReplyInbox} that is also {@link AutoCloseable}, so the ledger can prove it closes. */
@@ -0,0 +1,180 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.mcp.FleetMcp;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Shape A, rank 10: {@link FleetdAssembly} creates one {@link FleetMcp.LoopHealthSource}
* from the real started {@code StatusPoller} and {@code SessionReaper}, then gives it to two operator
* windows. These tests reach the real assembled objects through {@link FleetdRuntime}, rather than
* building a second source beside them. A hardcoded {@code RUNNING} source would pass a simple
* "running" test, so the mutation proof also mis-wires the source to never-started loops: both
* windows must then report {@code STOPPED} and these assertions go red.
*/
class FleetdAssemblyLoopHealthTest {
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("no coordinator is configured");
};
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("healthy FakeHerdr must not be polled");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Bind runtime.app() to an ephemeral port only in the REST assertion below.
}
}
private final HttpClient http = HttpClient.newHttpClient();
private RecordingResourcePorts ports;
private FleetdRuntime runtime;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (runtime != null) {
runtime.close();
}
if (ports != null) {
assertNotNull(ports.shutdownHook, "the assembly must register its shutdown hook");
assertTrue(ports.herdr.closed, "FleetdRuntime.close must close the real assembled herdr client");
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
lifecycle:
idleTtlSeconds: 600
idleSleepGuard:
enabled: false
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""");
return FleetConfig.load(file);
}
private void assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new RecordingResourcePorts();
runtime = FleetdAssembly.assembleAndStart(
new AssemblyInputs(cfg, new ConfigRef(dir.resolve("fleetd.yaml"), cfg),
new SubscriptionGuard(cfg.guard().hostSet())), ports);
}
@Test
void fleetListUsesTheRunningLoopsInTheRealAssembledMcp(@TempDir Path dir) throws Exception {
assemble(dir);
// The field is the exact source captured by FleetMcp's fleet_list handler. Reflection is
// necessary because FleetMcp has no public source accessor; it is not a source-text check.
FleetMcp.LoopHealthSource loopHealth = loopHealthOf(runtime.mcp());
assertRunning(loopHealth, "FleetMcp's real fleet_list source");
}
@Test
void healthzUsesTheRunningLoopsInTheRealAssembledApp(@TempDir Path dir) throws Exception {
assemble(dir);
boundApp = runtime.app().start("127.0.0.1", 0);
HttpRequest request = HttpRequest.newBuilder(
URI.create("http://127.0.0.1:" + boundApp.port() + "/healthz")).GET().build();
HttpResponse<String> response = http.send(request, HttpResponse.BodyHandlers.ofString());
assertEquals(200, response.statusCode(), response.body());
assertTrue(response.body().contains("\"statusPoller\":\"RUNNING\""), response.body());
assertTrue(response.body().contains("\"sessionReaper\":\"RUNNING\""), response.body());
}
private static FleetMcp.LoopHealthSource loopHealthOf(FleetMcp mcp) throws Exception {
Field field = FleetMcp.class.getDeclaredField("loopHealth");
field.setAccessible(true);
return (FleetMcp.LoopHealthSource) field.get(mcp);
}
private static void assertRunning(FleetMcp.LoopHealthSource source, String consumer) {
assertEquals(LoopWatchdog.State.RUNNING, source.statusPoller().get(),
consumer + " must report the started real StatusPoller as RUNNING");
assertEquals(LoopWatchdog.State.RUNNING, source.sessionReaper().get(),
consumer + " must report the started real SessionReaper as RUNNING");
}
}
@@ -86,6 +86,13 @@ class FleetdAssemblyReleaseCleanupBehaviouralTest {
};
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
@@ -82,6 +82,13 @@ class FleetdAssemblyRequireOperatorConfirmBehaviouralTest {
};
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
@@ -98,6 +98,14 @@ class FleetdAssemblyRoleFallbackBoundaryTest {
public void startHttp(Javalin app, String host, int port) {
// No real HTTP bind in a unit test.
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox {
@@ -0,0 +1,184 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 rank 12 — {@code FleetdAssembly} supplies the {@link Injector}'s {@code TurnRegistrar}
* with {@code Fleetd.turnRegistrar(completion)}. The normal delivery path cannot distinguish that
* registrar from {@code TurnRegistrar.NOOP}: {@link dev.ltms.fleet.inject.CompletionResolver#onDelivered}
* registers the same turn shortly afterwards. The distinction matters when a delivery listener throws
* after the pane received the message but before normal completion runs. The registrar must already have
* registered the waiter with the REAL assembled resolver, so a later completion can still resolve it.
*
* <p>The test installs a throwing wrapper around the real assembled listener. It does not replace the
* registrar or resolver. This constructs the narrow failure condition without a sleep, then observes the
* real resolver through {@link FleetdRuntime#completion()}. An inert registrar, or a registrar wired to a
* throwaway resolver, leaves this waiter's turn absent from the real resolver and makes the final assertion
* fail.
*/
class FleetdAssemblyTurnRegistrarBehaviouralTest {
private static final String TARGET = "term_a";
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L);
Runnable shutdownHook;
void advanceSeconds(long seconds) {
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
}
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> {
throw new UnsupportedOperationException("no broker: block is configured");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
fleet:
leaders:
primary:
tab: "lead: primary"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(file);
}
@Test
void realAssembledResolverStillResolvesAfterDeliveredListenerThrows(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ControllableResourcePorts ports = new ControllableResourcePorts();
ports.herdr.withTab("w2", "w2:t7", "lead: primary");
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
Injector injector = runtime.injector();
installThrowingDeliveredListener(injector);
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
injector.enqueue(TARGET, "brief", new TurnToken(TARGET, waiter));
IllegalStateException thrown = assertThrows(IllegalStateException.class,
() -> injector.onStatus(TARGET, AgentStatus.IDLE),
"control: delivery must reach the installed listener and it must throw after delivery");
assertTrue(thrown.getMessage().contains("listener failure"), thrown::getMessage);
ports.herdr.readText("worker report after the listener failure");
ports.advanceSeconds(3);
runtime.completion().resolveBeforePostAction(TARGET);
assertTrue(waiter.isDone(),
"FleetdAssembly must wire the Injector registrar to this runtime's real CompletionResolver: "
+ "after a delivered listener throws, resolveBeforePostAction must still find and "
+ "resolve the registered waiter");
} finally {
assertTrue(ports.shutdownHook != null, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
assertTrue(ports.herdr.closed, "teardown control: the captured shutdown hook must close herdr");
}
}
private static void installThrowingDeliveredListener(Injector injector) throws Exception {
Field field = Injector.class.getDeclaredField("turnListener");
field.setAccessible(true);
field.set(injector, new TurnListener() {
@Override
public void onTurnComplete(String target) {
}
@Override
public void onDelivered(String target, TurnToken token) {
throw new IllegalStateException("listener failure after delivery");
}
});
}
}
@@ -100,6 +100,14 @@ class FleetdBackendQuarantineAssemblyTest {
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@@ -126,6 +126,14 @@ class FleetdCompletionResolverAssemblyTest {
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind a real port.
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static FleetConfig writeConfig(Path dir, String profilesYaml, String extraGuardHost,
@@ -0,0 +1,218 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.session.MemberSession;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.OptionalLong;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 step 4, ranks 1 and 2 (publish side) — {@link FleetdAssembly} lines
* {@code liveExhaustedPatterns}/{@code exhaustedPatterns} (CB-578 stage A, the ticket's own "worst
* consequence in the whole sweep": a genuine usage-limit refusal handed back to a waiting caller
* AS REAL COMPLETED WORK) and {@code Fleetd.publishExhaustionSink(...)} (CB-578 stage B: the
* credential that hit the limit is never quarantined). None of these three lines is driven by an
* existing test through the real assembly: {@code FleetdExhaustedPatternLookupWiringTest} and
* {@code FleetdLiveExhaustedPatternsWiringTest} (fleetd #589) call {@code Fleetd.liveExhaustedPatterns}
* / {@code Fleetd.exhaustedPatternLookup} directly as factories, never through {@link
* FleetdAssembly#assembleAndStart} — they prove the FACTORY classifies correctly, never that THIS
* call site is the one that actually got wired into the running {@link CompletionResolver}. {@link
* FleetdBackendQuarantineAssemblyTest} drives {@code BackendQuarantine.withEscalation(...)}
* directly, a different call site from {@code publishExhaustionSink} here.
*
* <p>This test drives the REAL assembled {@link CompletionResolver} ({@link
* FleetdRuntime#completion()}) with a profile carrying a configured {@code exhaustedPattern},
* through a pane scrape that matches it, and asserts both halves of the production consequence:
* (1) the resolution is {@link Rendezvous.Kind#BACKEND_EXHAUSTED}, never a plain completion handed
* back as real work, and (2) the profile's credential is actually quarantined afterward, through
* the REAL {@link BackendQuarantine} the same assembly built ({@link
* FleetdRuntime#mcp()}{@code .quarantineSource().quarantine()}) — never a copy.
*
* <p>Same {@code ControllableResourcePorts} shape as {@code FleetdCompletionResolverAssemblyTest}:
* a fake, advanceable {@code nanoClock} so {@code CompletionResolver.MIN_TURN_NANOS} clears without
* a real sleep, and {@link FakeHerdr#readText} to drive the pane scrape.
*/
class FleetdExhaustedPatternAssemblyTest {
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr;
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
Runnable shutdownHook;
ControllableResourcePorts(FakeHerdr herdr) {
this.herdr = herdr;
}
void advanceSeconds(long seconds) {
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
}
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> {
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
};
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind a real port.
}
}
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
quarantineCooldownSeconds: %d
profiles:
exhaustprofile:
baseUrl: http://exhausthost.local:8000
model: sonnet
exhaustedPattern: "usage limit reached"
guard:
offSubscriptionHosts:
- exhausthost.local
""".formatted(cooldownSeconds));
return FleetConfig.load(f);
}
/**
* fleetd #589's own description of this gap ({@code Fleetd#exhaustedPatternLookup}'s javadoc):
* "the worst consequence in the whole #589 sweep" — a genuine usage-limit refusal stops being
* classified as {@code BACKEND_EXHAUSTED} and is handed back to a waiting {@code fleet_send} as
* if it were real completed work. Pins {@code FleetdAssembly}'s {@code liveExhaustedPatterns}
* AND {@code exhaustedPatterns} lines (rank 1) together with {@code publishExhaustionSink}
* (rank 2, the non-OpenCode half) in one flow: classify, then quarantine.
*/
@Test
@DisplayName("[BEHAVIOURAL] a scrape matching the profile's exhaustedPattern resolves "
+ "BACKEND_EXHAUSTED (never a plain completion) and quarantines the credential")
void assembledResolverClassifiesExhaustionAndQuarantinesTheCredential(@TempDir Path dir) throws Exception {
int cooldownSeconds = 120;
FleetConfig cfg = writeConfig(dir, cooldownSeconds);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
MemberSession session = runtime.sessions().acquire("exhaustprofile", null, dir.toString(), null);
String target = session.terminalId();
CompletionResolver completion = runtime.completion();
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
ports.herdr.readText("idle, nothing yet");
completion.onDelivered(target, new TurnToken(target, waiter, null));
// The matched text must START the pane line (CompletionResolver.startsWithExhaustion) —
// no preceding sentence — for the quarantine side-effect to fire, same as production.
ports.herdr.readText("usage limit reached: try again in a few hours");
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s), no real sleep
completion.resolveBeforePostAction(target);
// CONTROL: the waiter must have resolved synchronously at all — if the assembled
// CompletionResolver were never actually driven (e.g. a wiring break upstream silently
// left the resolver unreachable), this fails loudly before the real assertions below
// ever run, rather than passing on an untouched waiter.
Rendezvous.Resolution resolution = waiter.getNow(null);
assertTrue(resolution != null, "CONTROL: the waiter must have resolved synchronously — "
+ "if this is null, the assembled resolver was never actually exercised");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, resolution.kind(),
"a scrape matching the profile's configured exhaustedPattern must classify as "
+ "BACKEND_EXHAUSTED, not a plain completion handed back as real work — "
+ "replacing FleetdAssembly's liveExhaustedPatterns/exhaustedPatterns "
+ "lines with their inert forms (Map.of() / target -> null) must fail "
+ "this assertion; got: " + resolution);
assertTrue(resolution.text().contains("usage limit reached"), resolution.text());
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
assertTrue(quarantine.isQuarantined("exhaustprofile"),
"the real publishExhaustionSink-built sink must have quarantined the profile's "
+ "credential (effectiveCredentialId() == the profile name here, no "
+ "credentialId configured) — replacing FleetdAssembly's "
+ "publishExhaustionSink call site with a hardcoded ExhaustionSink.none() "
+ "must fail this assertion, since nothing would ever call "
+ "quarantine.quarantine(...)");
OptionalLong remaining = quarantine.remainingSeconds("exhaustprofile");
assertTrue(remaining.isPresent() && remaining.getAsLong() > 0
&& remaining.getAsLong() <= cooldownSeconds,
"a fresh quarantine must block for at most the configured base cooldown: " + remaining);
} finally {
if (ports.shutdownHook != null) ports.shutdownHook.run();
}
}
}
@@ -0,0 +1,213 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* fleetd #612 Shape A, unit r5 — {@code FleetdAssembly.java:488} wires {@link
* FleetMcp.LeadConfigDirSource} with {@code Fleetd.leadConfigDirSource(() -> config.get().profiles(),
* leaders)}. {@link FleetdLeadConfigDirSourceWiringTest} already pins that the FACTORY itself
* delegates to the real {@link Fleetd#leadConfigDirLookup} — but, by its own javadoc, it "does not
* and structurally cannot cover" whether the real call site in {@code FleetdAssembly} still calls
* that factory at all. Measured there: swapping that one-line call for a bare {@code
* FleetMcp.LeadConfigDirSource.none()} compiles with 0 errors and leaves the full suite green.
*
* <p>This is the literal fleetd #602/#606 defect, one call site away from its own fix: {@code main}
* (now {@code FleetdAssembly}) used to build {@code LeadConfigDirSource.none()} inline, the whole
* suite passed, and the live daemon reported {@code "state":"unknown"} for every lead's context,
* forever, with no test noticing. The fix extracted the factory; this test is the one that proves
* {@code FleetdAssembly}'s own call site still reaches it.
*
* <p>This test drives the REAL {@link FleetMcp} the real {@link FleetdAssembly#assembleAndStart}
* builds, reached through {@link FleetdRuntime#mcp()}, and reads the {@code leadConfigDirs} field it
* was constructed with via reflection — {@code FleetMcp} exposes no public accessor for it (unlike
* {@code quarantineSource()}/{@code leadSeatSource()}), so there is no non-reflective route to the
* live instance. The assertion resolves a REAL lead name against a REAL configured {@code
* configDir:}: {@link FleetMcp.LeadConfigDirSource#none()} (the historical defect, and the
* mis-wire this test's mutation cycles reintroduce) always returns {@code null} regardless of the
* input, so a non-null, config-matching answer is a property {@code none()} can never produce by
* accident.
*/
class FleetdLeadConfigDirSourceAssemblyTest {
private static final String LEAD_NAME = "opus";
private static final String LEAD_TAB = "lead: opus";
private static final String LEAD_PROFILE = "sonnet";
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
fleet:
leaders:
%s:
tab: "%s"
profile: %s
profiles:
%s:
subscription: true
argv: ["ccs", "sonnet"]
configDir: "%s"
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
private static FleetMcp.LeadConfigDirSource leadConfigDirSourceOf(FleetMcp mcp) throws Exception {
Field field = FleetMcp.class.getDeclaredField("leadConfigDirs");
field.setAccessible(true);
return (FleetMcp.LeadConfigDirSource) field.get(mcp);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadConfigDirSource resolves a lead's REAL "
+ "configured configDir, not the none() stand-in's hardcoded null")
void assembledLeadConfigDirSourceResolvesTheRealConfiguredConfigDir(@TempDir Path dir) throws Exception {
String configuredConfigDir = "/mnt/fake-lead-configdir";
FleetConfig cfg = writeConfig(dir, configuredConfigDir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
// agent) to match fleet.leaders.opus.tab exactly, so LeadLauncher.ensureLeads() sees the
// lead as already live and does not try to auto-launch a second one.
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
FleetMcp.LeadConfigDirSource source = leadConfigDirSourceOf(runtime.mcp());
assertEquals(configuredConfigDir, source.configDirFor().apply(LEAD_NAME),
"fleet.leaders." + LEAD_NAME + ".profile (" + LEAD_PROFILE + ") configures "
+ "configDir: " + configuredConfigDir + " — the real assembled source must "
+ "resolve it. FleetMcp.LeadConfigDirSource.none() (the inert stand-in "
+ "this test's mutation cycles swap the call site for, and the historical "
+ "fleetd #602/#606 defect) always reports null here, whatever the input");
// A lead name the config does not recognise still resolves to null, not a crash — the
// same source, applied to an input that must stay at the inert answer even on the real,
// non-inert instance.
assertNull(source.configDirFor().apply("no-such-lead"));
} finally {
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
// failure path too — hence try/finally rather than a bare statement at the end.
runtime.close();
}
}
}
@@ -115,6 +115,14 @@ class FleetdLeadRolloverAssemblyTest {
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@@ -90,6 +90,14 @@ class FleetdLeadSeatAssemblyTest {
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@@ -0,0 +1,243 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import io.modelcontextprotocol.client.McpClient;
import io.modelcontextprotocol.client.McpSyncClient;
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
import io.modelcontextprotocol.spec.McpClientTransport;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.Timeout;
import org.junit.jupiter.api.io.TempDir;
import java.net.http.HttpRequest;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Shape A ranks 9 and 11: drives the real {@code fleet_list} MCP route through a
* token-authenticated primary caller. The assertions read the response from the exact {@link
* dev.ltms.fleet.mcp.FleetMcp} instance assembled by {@link FleetdAssembly}, rather than a source
* scrape or a separately built reporting source.
*
* <p>The coordinator fixture contains a peer on purpose. Both a missing coordinator and an
* incorrectly wired peers argument can render as an empty list, so the non-empty peer assertion
* distinguishes the real call-site value from that inert result. Its fake mailbox never contacts a
* broker.
*/
class FleetdListReportingSourcesAssemblyTest {
private static final String TOKEN = "fleetd-r9-r11-test-token";
private static final String TOKEN_ENV = "FLEETD_R9_TEST_TOKEN";
private static final String PROFILE = "capacity-profile";
private static final String PEER = "peer-fleet";
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
private final FakeHerdr herdr = new FakeHerdr();
private Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of(TOKEN_ENV, TOKEN);
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// The test binds runtime.app() itself, below.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("healthy FakeHerdr must not poll");
};
}
}
private FleetdRuntime runtime;
private TestResourcePorts ports;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
auth:
mode: token
tokenEnv: %s
idleSleepGuard:
enabled: false
health:
enabled: true
notifications:
mode: webhook
profiles:
%s:
baseUrl: http://capacity.test:8000
model: test-model
maxLoad: 7
coordinator:
uri: amqp://fake-coordinator/vh
selfId: test-lead
peers:
- %s
""".formatted(TOKEN_ENV, PROFILE, PEER));
return FleetConfig.load(file);
}
private int assembleAndBind(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
ports = new TestResourcePorts();
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config,
new SubscriptionGuard(cfg.guard().hostSet())), ports);
boundApp = runtime.app().start("127.0.0.1", 0);
return boundApp.port();
}
private static String fleetList(int port) {
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder()
.header("Authorization", "Bearer " + TOKEN);
McpClientTransport transport = HttpClientStreamableHttpTransport.builder("http://127.0.0.1:" + port)
.endpoint("/mcp")
.requestBuilder(requestTemplate)
.build();
try (McpSyncClient client = McpClient.sync(transport).build()) {
client.initialize();
McpSchema.CallToolResult response = client.callTool(McpSchema.CallToolRequest.builder("fleet_list")
.arguments(Map.of()).build());
return ((McpSchema.TextContent) response.content().getFirst()).text();
}
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void fleetListReportsTheAssembledCapacitySource(@TempDir Path dir) throws Exception {
String response = fleetList(assembleAndBind(dir));
assertTrue(response.contains("\"capacity\""), response);
assertTrue(response.contains("\"" + PROFILE + "\""), response);
assertTrue(response.contains("\"maxLoad\":7"), response);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void fleetListReportsTheAssembledHealthCoverageSource(@TempDir Path dir) throws Exception {
String response = fleetList(assembleAndBind(dir));
assertTrue(response.contains("\"healthCoverage\":\"full\""), response);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void fleetListReportsTheAssembledCoordinatorPeers(@TempDir Path dir) throws Exception {
String response = fleetList(assembleAndBind(dir));
assertTrue(response.contains("\"coordinator\""), response);
assertTrue(response.contains("\"coordId\":\"" + PEER + "\""), response);
}
}
@@ -0,0 +1,197 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 step 4, rank 2 (OpenCode half) — {@link FleetdAssembly}'s {@code
* forwardingExhaustionSink} line ({@code Fleetd.forwardingExhaustionSink(exhaustionSinkRef)}),
* handed to {@link dev.ltms.fleet.member.OpenCodeLauncher} so its fleetd #175 model-mismatch check
* can quarantine a credential before {@code sessions} exists to build the real sink (the
* construction-order cycle documented at that call site). The ticket calls this independent from
* {@code publishExhaustionSink} (pinned by {@link FleetdExhaustedPatternAssemblyTest}): a credential
* that hits a usage limit through THIS path is never quarantined if {@code forwardingExhaustionSink}
* is swapped for a hardcoded {@link ExhaustionSink#none()} at that call site — the OpenCode
* launcher's own quarantine check keeps compiling and keeps "running", but it permanently talks to
* a sink that does nothing, independent of whatever {@code publishExhaustionSink} does later.
*
* <p>{@code FleetdExhaustionSinkForwardingWiringTest} (fleetd #589) already proves {@code
* Fleetd.forwardingExhaustionSink(ref)} forwards to whatever {@code ref} holds — as a bare factory
* call, never through {@link FleetdAssembly#assembleAndStart}. It proves nothing about whether
* THIS call site is the one FleetdAssembly actually wires into the real {@code OpenCodeLauncher}
* it builds, which is exactly the #602/#606-shaped gap this ticket exists to close.
*
* <p>No accessor on {@link FleetdRuntime} reaches the adapter instances (by design — see that
* class's own javadoc: only the final collaborators it owns directly are exposed), so this test
* reaches the REAL, assembled {@code OpenCodeLauncher}'s {@code exhaustionSink} field the same way
* {@code SessionManager}/{@code CompositePeerLauncher} wire it internally: a short, targeted
* reflective walk ({@code SessionManager.launcher} → {@code CompositePeerLauncher.byProfile} →
* {@code OpenCodeLauncher.exhaustionSink}) onto the exact object the assembly built — never a copy,
* and never a read of the source text. Reflection is used the same way elsewhere in this suite
* (e.g. {@code StatusPollerWatchdogTest}) to reach a private collaborator a production constructor
* intentionally does not expose a public accessor for.
*/
class FleetdOpenCodeExhaustionForwardingAssemblyTest {
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L);
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> {
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
};
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind a real port.
}
}
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
quarantineCooldownSeconds: %d
profiles:
gemini:
kind: opencode
model: google/gemini-2.5-pro
""".formatted(cooldownSeconds));
return FleetConfig.load(f);
}
/** Reach a declared field by name on {@code target}'s runtime class, bypassing the access check. */
private static Object readField(Object target, Class<?> declaringClass, String fieldName) throws Exception {
Field field = declaringClass.getDeclaredField(fieldName);
field.setAccessible(true);
return field.get(target);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled OpenCodeLauncher's exhaustionSink field forwards "
+ "an onExhausted call into the real daemon's BackendQuarantine")
void assembledOpenCodeLauncherExhaustionSinkQuarantinesTheCredential(@TempDir Path dir) throws Exception {
int cooldownSeconds = 90;
FleetConfig cfg = writeConfig(dir, cooldownSeconds);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ControllableResourcePorts ports = new ControllableResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
PeerLauncher launcherField = (PeerLauncher) readField(runtime.sessions(),
runtime.sessions().getClass(), "launcher");
// CONTROL: the composite launcher must actually be the real production type with a
// "gemini" -> OpenCodeLauncher entry — if this fails, nothing below exercised the real
// assembly at all, rather than silently passing on an empty/wrong object.
assertTrue(launcherField instanceof CompositePeerLauncher,
"CONTROL: SessionManager.launcher must be the real CompositePeerLauncher the "
+ "assembly built, got: " + launcherField);
@SuppressWarnings("unchecked")
Map<String, HerdrPeerLauncher> byProfile = (Map<String, HerdrPeerLauncher>)
readField(launcherField, CompositePeerLauncher.class, "byProfile");
HerdrPeerLauncher adapter = byProfile.get("gemini");
assertTrue(adapter != null && adapter.getClass().getSimpleName().equals("OpenCodeLauncher"),
"CONTROL: the 'gemini' profile must resolve to a real OpenCodeLauncher adapter, "
+ "got: " + adapter);
ExhaustionSink sink = (ExhaustionSink) readField(adapter, adapter.getClass(), "exhaustionSink");
assertTrue(sink != null, "CONTROL: OpenCodeLauncher.exhaustionSink must never be null");
// The exact call OpenCodeLauncher.SessionAwareHandle#checkModelMatch makes on a real
// model mismatch (fleetd #175): target, reason, and its own already-known profile name.
sink.onExhausted("term_gemini_1", "opencode model mismatch (test)", "gemini");
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
assertTrue(quarantine.isQuarantined("gemini"),
"the real forwardingExhaustionSink-wired field must have delegated into the "
+ "published production sink, which quarantines the profile's credential "
+ "('gemini' here — no credentialId configured) — replacing "
+ "FleetdAssembly's forwardingExhaustionSink call site with a hardcoded "
+ "ExhaustionSink.none() must fail this assertion, since the field read "
+ "above would then BE the inert no-op and nothing would ever reach "
+ "quarantine.quarantine(...)");
assertEquals(cooldownSeconds, quarantine.remainingSeconds("gemini").orElseThrow(
() -> new AssertionError("credential must report a remaining cooldown")));
} finally {
if (ports.shutdownHook != null) ports.shutdownHook.run();
}
}
}
@@ -0,0 +1,368 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.CompletionResolver;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.TurnToken;
import dev.ltms.fleet.session.MemberSession;
import io.javalin.Javalin;
import io.modelcontextprotocol.client.McpClient;
import io.modelcontextprotocol.client.McpSyncClient;
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
import io.modelcontextprotocol.spec.McpClientTransport;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.Timeout;
import org.junit.jupiter.api.io.TempDir;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #612 Shape A, unit r4 — {@link FleetdAssembly}'s {@code quarantineSource} (lines 471-472)
* and {@code outageSource} (lines 473-476 at {@code main} = {@code 141ae3b}), EACH of which feeds
* two separate consumers: {@code FleetMcp} ({@code fleet_profiles}) and {@code FleetApp}
* ({@code GET /profiles}), one call site ({@code :484}/{@code :486}) into the MCP constructor and
* the SAME shared local again ({@code :529}) into the REST constructor.
*
* <p><strong>The property pinned here</strong> (from the ticket): when the real assembly has built
* a real {@link dev.ltms.fleet.placement.BackendQuarantine} holding a quarantined credential, and a
* real {@link dev.ltms.fleet.placement.BackendOutagePolicy} holding a cooling-off credential, BOTH
* operator windows must report that state — the real assembled {@code FleetMcp} (reached through
* {@link FleetdRuntime#mcp()}) and the real assembled REST surface (reached through {@link
* FleetdRuntime#app()}). A mutation that starves one consumer while leaving the other wired must
* make only that consumer's assertion go red.
*
* <p><strong>No source-text assertion anywhere in this file.</strong> Both windows are read off the
* REAL running objects: {@code fleet_profiles} is called through a real MCP client over a real
* HTTP connection to the servlet {@link FleetdAssembly} actually mounted, and {@code GET /profiles}
* is called through a real {@link java.net.http.HttpClient} against the real bound {@link
* FleetdRuntime#app()}. Neither is a copy built alongside the assembly for this test's benefit.
*
* <p><strong>How this gets past CB-185's own pid-resolution dead end.</strong> {@code
* FleetdAssemblyFleetAppTest}'s class javadoc explains that {@code GET /sessions} cannot be driven
* over real HTTP here because {@code LsofPeerPidLookup} excludes its own pid and an in-process test
* client/server share one JVM pid — every such request resolves {@code ANONYMOUS} and is refused
* before the handler runs. {@code GET /profiles} and {@code fleet_profiles} sit behind the exact
* same {@code Authz.Action.READ} gate. This test sidesteps the dead end instead of hitting it:
* {@code auth.mode: token} (see {@link dev.ltms.fleet.auth.CallerResolver#resolve}) resolves a
* caller to {@code PRIMARY} from a valid {@code Authorization: Bearer} header ALONE, with no pid
* resolution involved at all — the same technique {@code FleetMcpContextExtractorTest} already uses
* to drive a real {@code fleet_whoami} call through the real transport.
*
* <p><strong>How the quarantined/cooling-off state is set up.</strong> Both {@code BackendQuarantine}
* and {@code BackendOutagePolicy} are private to the collaborators the assembly wires them into, and
* (unlike {@code quarantineSource()}) neither {@code FleetMcp} nor {@code FleetApp} exposes a public
* accessor for the live {@code BackendOutagePolicy} instance. Rather than add one (acceptance
* criterion 1: no production change), this test drives the REAL production classification path —
* exactly the recipe {@code FleetdExhaustedPatternAssemblyTest} (quarantine) and {@code
* FleetdCompletionResolverAssemblyTest} (cool-off) already proved works end to end against this same
* {@link FleetdAssembly#assembleAndStart}: acquire a real {@link MemberSession}, feed the real {@link
* CompletionResolver} a pane scrape matching the profile's configured {@code exhaustedPattern} /
* {@code errorPattern}, and let the real {@code exhaustionSink}/{@code backendErrorSink} write into
* the real, shared tracker. Each setup step asserts its own {@link Rendezvous.Kind} as a CONTROL —
* if the resolver were never actually exercised, the setup itself fails loudly before either window
* is ever read.
*/
class FleetdQuarantineOutageDualWindowAssemblyTest {
private static final String TOKEN = "s3cret-r4-token";
private static final String TOKEN_ENV = "FLEETD_R4_TEST_TOKEN";
private static final class ControllableResourcePorts implements ResourcePorts {
final FakeHerdr herdr;
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
Runnable shutdownHook;
ControllableResourcePorts(FakeHerdr herdr) {
this.herdr = herdr;
}
void advanceSeconds(long seconds) {
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
}
@Override
public Map<String, String> environment() {
return Map.of(TOKEN_ENV, TOKEN);
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
// Never invoked: this test's config has no `broker:` block.
return (uri, prefetch) -> {
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
// Never invoked: this test's config has no `coordinator:` block.
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return nowNanos::get;
}
@Override
public LongSupplier wallClockNanos() {
return nowNanos::get;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
this.shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private FleetdRuntime runtime;
private ControllableResourcePorts ports;
private Javalin boundApp;
@AfterEach
void tearDown() {
if (boundApp != null) {
boundApp.stop();
}
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
auth:
mode: token
tokenEnv: %s
idleSleepGuard:
enabled: false
quarantineCooldownSeconds: %d
profiles:
exhaustprofile:
baseUrl: http://exhausthost.local:8000
model: sonnet
exhaustedPattern: "usage limit reached"
coolprofile:
baseUrl: http://coolhost.local:8000
model: sonnet
errorPattern: "credential outage"
guard:
offSubscriptionHosts:
- exhausthost.local
- coolhost.local
""".formatted(TOKEN_ENV, cooldownSeconds));
return FleetConfig.load(f);
}
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
private int assembleAndBind(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir, 120);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
ports = new ControllableResourcePorts(new FakeHerdr());
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
boundApp = runtime.app().start("127.0.0.1", 0);
return boundApp.port();
}
/**
* Drives the real assembled {@link CompletionResolver} through a scrape matching {@code
* exhaustprofile}'s configured {@code exhaustedPattern}, exactly {@code
* FleetdExhaustedPatternAssemblyTest}'s own recipe, so the real {@code exhaustionSink} quarantines
* the credential ({@code effectiveCredentialId() == "exhaustprofile"}, no explicit credentialId
* configured).
*/
private void quarantineExhaustProfile(Path dir) {
MemberSession session = runtime.sessions().acquire("exhaustprofile", null, dir.toString(), null);
String target = session.terminalId();
CompletionResolver completion = runtime.completion();
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
ports.herdr.readText("idle, nothing yet");
completion.onDelivered(target, new TurnToken(target, waiter, null));
ports.herdr.readText("usage limit reached: try again in a few hours");
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s), no real sleep
completion.resolveBeforePostAction(target);
Rendezvous.Resolution resolution = waiter.getNow(null);
assertTrue(resolution != null && resolution.kind() == Rendezvous.Kind.BACKEND_EXHAUSTED,
"CONTROL: setup must classify as BACKEND_EXHAUSTED before either window is read — "
+ "if this fails, the assembled resolver was never actually exercised: " + resolution);
}
/**
* Drives the real assembled {@link CompletionResolver} with TWO distinct targets on {@code
* coolprofile}, each matching its configured {@code errorPattern}, exactly {@code
* FleetdCompletionResolverAssemblyTest}'s own recipe, so the real {@code backendErrorSink} cools
* the credential off ({@code effectiveCredentialId() == "coolprofile"}).
*/
private void coolOffCoolProfile(Path dir) {
MemberSession s1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
MemberSession s2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
CompletionResolver completion = runtime.completion();
String t1 = s1.terminalId();
CompletableFuture<Rendezvous.Resolution> w1 = new CompletableFuture<>();
ports.herdr.readText("idle 1");
completion.onDelivered(t1, new TurnToken(t1, w1, null));
ports.herdr.readText("credential outage: upstream 503");
ports.advanceSeconds(3);
completion.resolveBeforePostAction(t1);
Rendezvous.Resolution r1 = w1.getNow(null);
assertTrue(r1 != null && r1.kind() == Rendezvous.Kind.FAILED,
"CONTROL: target1's setup must classify FAILED (backend error): " + r1);
String t2 = s2.terminalId();
CompletableFuture<Rendezvous.Resolution> w2 = new CompletableFuture<>();
ports.herdr.readText("idle 2");
completion.onDelivered(t2, new TurnToken(t2, w2, null));
ports.herdr.readText("credential outage: upstream 503 again");
ports.advanceSeconds(3);
completion.resolveBeforePostAction(t2);
Rendezvous.Resolution r2 = w2.getNow(null);
assertTrue(r2 != null && r2.kind() == Rendezvous.Kind.FAILED,
"CONTROL: target2's setup must classify FAILED (backend error) — two distinct "
+ "targets are required to cross BackendOutagePolicy.THRESHOLD: " + r2);
}
/** Calls the real {@code fleet_profiles} tool over a real MCP client, token-authenticated as PRIMARY. */
private static McpSchema.CallToolResult callProfilesViaMcp(int port) {
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder().header("Authorization", "Bearer " + TOKEN);
McpClientTransport transport = HttpClientStreamableHttpTransport.builder("http://127.0.0.1:" + port)
.endpoint("/mcp")
.requestBuilder(requestTemplate)
.build();
try (McpSyncClient client = McpClient.sync(transport).build()) {
client.initialize();
return client.callTool(McpSchema.CallToolRequest.builder("fleet_profiles").arguments(Map.of()).build());
}
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
/** Calls the real {@code GET /profiles} route over a real {@link HttpClient}, same token. */
private static String getProfilesViaRest(int port) throws Exception {
HttpClient http = HttpClient.newHttpClient();
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + "/profiles"))
.header("Authorization", "Bearer " + TOKEN)
.GET().build();
HttpResponse<String> res = http.send(req, HttpResponse.BodyHandlers.ofString());
assertEquals(200, res.statusCode(), "GET /profiles must succeed with the real token: " + res.body());
return res.body();
}
// --- quarantineSource (FleetdAssembly.java :471-472) --------------------------------------
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void quarantinedCredentialIsReportedByTheRealAssembledFleetMcp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
quarantineExhaustProfile(dir);
String out = textOf(callProfilesViaMcp(port));
assertTrue(out.contains("\"quarantined\""), "fleet_profiles must report a quarantined "
+ "section once the real BackendQuarantine holds a quarantined credential: " + out);
assertTrue(out.contains("\"exhaustprofile\""), out);
assertTrue(out.contains("\"quarantinedForSeconds\""), out);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void quarantinedCredentialIsReportedByTheRealAssembledFleetApp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
quarantineExhaustProfile(dir);
String out = getProfilesViaRest(port);
assertTrue(out.contains("\"quarantined\""), "GET /profiles must report a quarantined "
+ "section once the real BackendQuarantine holds a quarantined credential: " + out);
assertTrue(out.contains("\"exhaustprofile\""), out);
assertTrue(out.contains("\"quarantinedForSeconds\""), out);
}
// --- outageSource (FleetdAssembly.java :473-476) -------------------------------------------
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void coolingOffCredentialIsReportedByTheRealAssembledFleetMcp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
coolOffCoolProfile(dir);
String out = textOf(callProfilesViaMcp(port));
assertTrue(out.contains("\"coolingOff\""), "fleet_profiles must report a coolingOff "
+ "section once the real BackendOutagePolicy holds a cooling-off credential: " + out);
assertTrue(out.contains("\"coolprofile\""), out);
assertTrue(out.contains("\"coolingOffForSeconds\""), out);
}
@Test
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
void coolingOffCredentialIsReportedByTheRealAssembledFleetApp(@TempDir Path dir) throws Exception {
int port = assembleAndBind(dir);
coolOffCoolProfile(dir);
String out = getProfilesViaRest(port);
assertTrue(out.contains("\"coolingOff\""), "GET /profiles must report a coolingOff "
+ "section once the real BackendOutagePolicy holds a cooling-off credential: " + out);
assertTrue(out.contains("\"coolprofile\""), out);
assertTrue(out.contains("\"coolingOffForSeconds\""), out);
}
}
@@ -0,0 +1,202 @@
package dev.ltms.fleet;
import dev.ltms.fleet.guard.GuardException;
import dev.ltms.fleet.herdr.HerdrClient;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #625: pins {@link dev.ltms.fleet.guard.SubscriptionGuard#assertPrimaryClean}'s call site
* in {@link Fleetd#main(String[])} — the ONE place it runs at startup, and the check behind the
* bridge charter's invariant 1 (never let the primary carry {@code ANTHROPIC_BASE_URL}). Nothing
* pinned it before this ticket: deleting {@code guard.assertPrimaryClean(...)} from {@code main}
* left the full suite green, because the call site read the real process environment ({@code
* System.getenv()}), which a test cannot taint from inside the JVM.
*
* <p>{@link Fleetd#main(String[], ResourcePorts)} (added by this ticket) is the literal production
* sequence — not a copy of it — driven here with a {@link ResourcePorts} whose {@link
* ResourcePorts#environment()} is a plain {@code Map} a test controls. The guard itself was
* already pinned by {@code SubscriptionGuardTest}, directly, with a {@code Map} — that proves the
* method's behaviour, not that {@code main} still calls it at the right point. This class pins the
* call site and, separately, the ORDER: the guard must still run before {@code cfg.validateAll()}
* and before {@code FleetdAssembly.assembleAndStart} touches a socket, a broker, or HTTP — not just
* be present somewhere in {@code main}.
*
* <p>Presence alone is not enough (a fix that pins only presence trades an invisible deletion for
* an invisible reordering), so each test below is built so that EITHER deleting the guard call OR
* moving it later makes the <em>same</em> test fail — with a different exception type than the one
* asserted, not a vacuous pass. See each test's own javadoc for how.
*/
class FleetdSubscriptionGuardOrderingTest {
private static final Map<String, String> TAINTED_ENV =
Map.of("ANTHROPIC_BASE_URL", "http://tainted.example");
private static final Map<String, String> CLEAN_ENV = Map.of("PATH", "/usr/bin");
/**
* A {@link ResourcePorts} whose {@link #environment()} is fixed to whatever the test hands it,
* and whose every other method refuses to be called at all. That refusal is the ordering pin:
* if {@code main} ever reaches {@link FleetdAssembly#assembleAndStart} before the guard has had
* a chance to throw, the very first thing the assembly does with {@code ports} is {@link
* #connectHerdr} — so a test that expects {@link GuardException} and instead observes {@link
* UnsupportedOperationException} has just caught the guard running too late (or not at all).
*/
private static final class FixedEnvPorts implements ResourcePorts {
private final Map<String, String> env;
FixedEnvPorts(Map<String, String> env) {
this.env = env;
}
@Override
public Map<String, String> environment() {
return env;
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
throw new UnsupportedOperationException("connectHerdr must not be called before the guard runs");
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
throw new UnsupportedOperationException("replyInboxOpener must not be called before the guard runs");
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
throw new UnsupportedOperationException("leadMailboxOpener must not be called before the guard runs");
}
@Override
public LongSupplier nanoClock() {
throw new UnsupportedOperationException("nanoClock must not be called before the guard runs");
}
@Override
public LongSupplier wallClockNanos() {
throw new UnsupportedOperationException("wallClockNanos must not be called before the guard runs");
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
throw new UnsupportedOperationException("newScheduler must not be called before the guard runs");
}
@Override
public void addShutdownHook(Runnable hook) {
throw new UnsupportedOperationException("addShutdownHook must not be called before the guard runs");
}
@Override
public void startHttp(Javalin app, String host, int port) {
throw new UnsupportedOperationException("startHttp must not be called before the guard runs");
}
@Override
public Runnable herdrPollWait() {
throw new UnsupportedOperationException("herdrPollWait must not be called before the guard runs");
}
}
/**
* Otherwise-invalid: {@code bind.host: 0.0.0.0} with no {@code auth.mode: token} fails {@code
* cfg.validateAll()} (CB-501's auth-exposure check — the same fixture {@code
* FleetdStartupValidationTest#mainRefusesANonLoopbackBindWithoutTokenMode} uses), with an
* {@link IllegalStateException}. That is deliberate: it is what {@code main} would throw INSTEAD
* of {@link GuardException} if the guard call were deleted, or moved to run after {@code
* validateAll()} — a different, distinguishable exception type.
*/
private static Path invalidConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 0.0.0.0
port: 8765
""");
return f;
}
/** Passes {@code cfg.validateAll()} cleanly — nothing here trips any of its checks. */
private static Path validConfig(Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
""");
return f;
}
/**
* Pins the order against {@code cfg.validateAll()}. The environment is tainted and the config
* is otherwise invalid (see {@link #invalidConfig}). If the guard runs first (the required
* order), {@code main} throws {@link GuardException} before {@code validateAll()} is ever
* reached. If the guard were deleted, or reordered to run after {@code validateAll()}, {@code
* validateAll()} throws {@link IllegalStateException} instead and this assertion fails on the
* wrong exception type.
*/
@Test
void mainRefusesATaintedEnvironmentBeforeValidatingTheConfig(@TempDir Path dir) throws Exception {
Path config = invalidConfig(dir);
FixedEnvPorts ports = new FixedEnvPorts(TAINTED_ENV);
GuardException ex = assertThrows(GuardException.class,
() -> Fleetd.main(new String[]{config.toString()}, ports));
assertTrue(ex.getMessage().contains("tainted"),
"expected the primary-taint message, got: " + ex.getMessage());
}
/**
* Pins the order against {@code FleetdAssembly.assembleAndStart}. The environment is tainted
* and the config is otherwise VALID (see {@link #validConfig}), so {@code cfg.validateAll()}
* passes silently and the next thing that could possibly run is the assembly's first socket
* call. If the guard runs first (the required order), {@code main} throws {@link
* GuardException} before assembly starts. If the guard were deleted, or reordered to run after
* assembly begins touching {@code ports}, {@link FixedEnvPorts#connectHerdr} throws {@link
* UnsupportedOperationException} instead and this assertion fails on the wrong exception type.
*/
@Test
void mainRefusesATaintedEnvironmentBeforeAssemblyTouchesAnyPort(@TempDir Path dir) throws Exception {
Path config = validConfig(dir);
FixedEnvPorts ports = new FixedEnvPorts(TAINTED_ENV);
GuardException ex = assertThrows(GuardException.class,
() -> Fleetd.main(new String[]{config.toString()}, ports));
assertTrue(ex.getMessage().contains("tainted"),
"expected the primary-taint message, got: " + ex.getMessage());
}
/**
* The CONTROL for the two tests above. Same otherwise-valid config, same {@link FixedEnvPorts}
* whose every method but {@code environment()} refuses to be called — but a CLEAN environment.
* Without this, a guard that always threw {@link GuardException} regardless of input (the
* opposite bug — e.g. the check inverted) would make the two tests above pass for the wrong
* reason: not because they actually drove a real taint through a real guard, but because
* anything would have thrown {@code GuardException}. Here, with nothing to taint, the guard
* must let {@code main} proceed into {@code cfg.validateAll()} and on into the real assembly,
* which reaches {@code ports.connectHerdr} — and THAT throws. A loud, positive assertion: if
* the boot path never actually ran this far, there is no {@link UnsupportedOperationException}
* to catch, only a quiet, unexpected hang or an unrelated early failure.
*/
@Test
void mainProceedsPastTheGuardOnACleanEnvironment(@TempDir Path dir) throws Exception {
Path config = validConfig(dir);
FixedEnvPorts ports = new FixedEnvPorts(CLEAN_ENV);
UnsupportedOperationException ex = assertThrows(UnsupportedOperationException.class,
() -> Fleetd.main(new String[]{config.toString()}, ports));
assertTrue(ex.getMessage().contains("connectHerdr"),
"expected forward progress to reach the assembly's first port call, got: " + ex.getMessage());
}
}
+732
View File
@@ -0,0 +1,732 @@
#!/usr/bin/env bash
#
# The one auditable way to edit the live fleetd.yaml.
#
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
# convenience — it is the thing a raw file write can never give you: a backup, a parse check
# BEFORE the file is installed, and the daemon's own reload verdict read back afterwards. An
# edit to a live config is not finished when the bytes are written. It is finished when the
# daemon has said what it did with them.
#
# What the daemon says, and how this script finds it — measured against `ConfigRef.java` on
# fleetd commit 158a2a8, 2026-10-01:
#
# 1. `ConfigRef` re-reads fleetd.yaml only when the WATCHER sees the mtime move (every 10s by
# default — read the real interval out of the daemon's own startup line, "config watch: ...
# re-read when it changes (every Ns)"). So a verdict never appears before the next tick.
# 2. `ConfigRef.Outcome.summary()` logs exactly one of five strings (ConfigRef.java:371-391):
# config reload refused — <error message>
# config reload refused — these keys cannot change under a running daemon: <keys>. ...
# config reloaded
# config reloaded; these changes need a restart to take effect: <keys>
# config reloaded; partially live — <key: detail | ...>
# A parse/validation failure logs a DIFFERENT line instead, before any summary ever runs
# (ConfigRef.java:425): "config reload from <path> refused, keeping the running config:
# <message>". This script recognises both shapes of refusal.
# 3. The em dash in those strings is a real multi-byte character — match the stable prefix
# "config reload refused" (or "...refused, keeping the running config" for the parse-failure
# shape), never the dash itself.
# 4. A cold-key change (bind/herdrSocket/memberHerdrSocket/broker/auth) throws away the WHOLE
# reload — the running config keeps every old value, not only the cold one.
# 5. A deferred/split change IS applied (current.set(fresh) runs) — "needs a restart" is a
# SUCCESS with a follow-up, never a failure.
#
# Four outcomes, and they stay four (see the exit code table below). The one most likely to be
# gotten wrong is "cannot tell" (exit 5): the daemon may be down, or the watcher may be stalled,
# and folding that into either "refused" or "applied" is worse than never checking at all,
# because a caller then acts on a verdict nobody actually read. So exit 5 never restores — a
# visible, recoverable edit beats an invisible revert of a GOOD edit.
#
# Usage:
# scripts/config-edit.sh --check
# scripts/config-edit.sh --set <yq-path>=<value> [--set ...]
# scripts/config-edit.sh --from <candidate.yaml>
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
# scripts/config-edit.sh --restore
#
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
# field": a null value falls back to its default rather than erroring, which is silent, not
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
# is always written as a YAML string (via yq's strenv(), never spliced into the expression), so
# there is currently no --set spelling for the literal three-character STRING "null" itself — use
# --from for that rare case.
#
# Overrides (so this is drivable with no daemon — see scripts/test-config-edit.sh):
# --config <path> default: fleetd/fleetd.yaml
# --log <path> default: fleetd/fleetd.out
# --wait-seconds <n> default: 4x the watch interval this script reads out of --log (10 -> 40)
#
# Exit codes (the --check/--restore/usage-error paths are reported separately, see below):
# 0 applied; verdict read; clean
# 3 applied; verdict read; needs a restart (deferred or split keys named)
# 4 REFUSED by the daemon; backup restored (and the restore's own verdict reported if seen)
# 5 CANNOT TELL — no verdict line inside the wait window. Nothing is restored.
#
# Never prints a secret. fleetd.yaml keeps credentials out by indirection (broker.uriEnv,
# gitTokenEnv) but this script does not rely on that staying true: every diff it prints is piped
# through `redact`, which (a) blanks the userinfo of any `scheme://user:pass@host` and (b) masks
# the whole value on any line whose key looks like a credential. See `redact` below.
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SELF="$REPO/scripts/config-edit.sh"
CONFIG="$REPO/fleetd/fleetd.yaml"
LOG="$REPO/fleetd/fleetd.out"
WAIT_SECONDS_OVERRIDE=""
FALLBACK_PORT=8765
MODE=""
DRY_RUN=0
SETS=()
FROM_FILE=""
# fleetd #635 follow-up — a signal (or any early exit while a candidate is still uninstalled) must
# not leave a `.config-edit.XXXXXX` file sitting beside the live config forever. CAND is global
# (never a function-local) on purpose: this ONE trap, set once, covers every path that ever
# creates a candidate — run_edit and dry_run_diff both assign it, and clear it back to "" once the
# file is consumed (installed, or explicitly removed), so a later, unrelated exit never retries a
# path that already served its purpose.
CAND=""
cleanup_candidate() { [ -n "$CAND" ] && rm -f "$CAND" 2>/dev/null; return 0; }
trap cleanup_candidate EXIT INT TERM
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
ok() { printf ' ok %s\n' "$*"; }
warn() { printf ' WARN %s\n' "$*"; }
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
set_mode() {
local new="$1"
if [ -n "$MODE" ] && [ "$MODE" != "$new" ]; then
die "cannot combine --$MODE and --$new in one invocation"
fi
MODE="$new"
}
while [ $# -gt 0 ]; do
case "$1" in
--check) set_mode check; shift ;;
--restore) set_mode restore; shift ;;
--set)
[ $# -ge 2 ] || die "--set requires <yq-path>=<value>"
set_mode set
SETS+=("$2")
shift 2 ;;
--from)
[ $# -ge 2 ] || die "--from requires a candidate file path"
set_mode from
FROM_FILE="$2"
shift 2 ;;
--dry-run) DRY_RUN=1; shift ;;
--config)
[ $# -ge 2 ] || die "--config requires a path"
CONFIG="$2"; shift 2 ;;
--log)
[ $# -ge 2 ] || die "--log requires a path"
LOG="$2"; shift 2 ;;
--wait-seconds)
[ $# -ge 2 ] || die "--wait-seconds requires a number of seconds"
WAIT_SECONDS_OVERRIDE="$2"; shift 2 ;;
-h|--help) sed -n '3,70p' "$SELF"; exit 0 ;;
*) echo "unknown option: $1 (try --help)" >&2; exit 2 ;;
esac
done
[ -n "$MODE" ] || die "no action given — use --check, --set, --from, or --restore (see --help)"
# ------------------------------------------------------------------------------------- redaction
#
# Two independent passes, applied to every diff this script ever prints:
# 1. `scheme://user:pass@host` -> `scheme://<redacted>@host`, globally (the `g` flag matters —
# a line can carry more than one URI).
# 2. Any line whose key looks like TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY,
# matched case-insensitively against the key text (uriEnv, gitTokenEnv, ... are camelCase,
# not SCREAMING_CASE) has its whole value blanked, diff marker and indentation kept so the
# shape of the change is still visible. Deliberately conservative: a false-positive
# redaction on an unrelated line costs nothing, an unredacted secret is a security defect
# (acceptance criterion 7).
#
# fleetd #635 follow-up (ticket comment 17670, defect 7) — a masked key line is not the whole
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
# the key, each indented deeper than it. The key-name match above only ever sees the key line
# itself, so those continuation lines used to flow straight through unredacted while the key line
# right above them printed a reassuring "<redacted>" — an incomplete redactor that looks complete
# is worse than one that visibly does nothing, because it stops a reviewer from looking further.
# The fix is structural, not another name to match: once a key line is masked, every following
# line indented STRICTLY DEEPER than that key is masked too, by indentation alone, until the
# indentation returns to the key's own level or shallower. This needs no knowledge of the key's
# name, so it covers a block scalar under any masked key — but ONLY while that key's own line is
# itself inside the hunk being printed. `diff -u` prints just three lines of context, so a block
# scalar's body often reaches this function with its key line left out; there is then nothing to
# anchor to, `masked` is never set, and the body prints in full. A blank line inside a block
# scalar loses the anchor the same way, because a blank diff line measures as indent 0. Both are
# measured and filed as fleetd #639 — do not read this paragraph as a guarantee that a masked
# key's value can never be printed.
#
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
# That one leading character is NOT part of the YAML indentation, and must be stripped before
# indentation is measured or a key is matched — otherwise a changed ('+' or '-') line reads one
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
# fix, not a YAML parser.
redact() {
local line prefix content indent lead key
local masked=0 masked_indent=0 saved_nocasematch=0
shopt -q nocasematch && saved_nocasematch=1
shopt -s nocasematch
sed -E 's#://[^@]*@#://<redacted>@#g' | while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
*) prefix=""; content="$line" ;;
esac
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$masked" = 1 ] && [ "$indent" -gt "$masked_indent" ]; then
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
continue
fi
masked=0
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
masked=1
masked_indent="$indent"
continue
fi
fi
printf '%s\n' "$line"
done
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
}
# ------------------------------------------------------------------------------------- the probe
#
# Probe the SOCKET, never `pgrep`/`ps -f` — both print argv, and argv holds `NAME=value`, making
# either a credential channel. The port comes from the config's own `bind.port`; 8765 is only a
# fallback when that key is absent or the file does not parse yet.
resolve_port() {
local file="$1" port
if [ -f "$file" ] && command -v yq >/dev/null 2>&1; then
port="$(yq eval '.bind.port' "$file" 2>/dev/null || true)"
else
port=""
fi
case "$port" in
''|null) echo "$FALLBACK_PORT" ;;
*) echo "$port" ;;
esac
}
daemon_listening() {
local port="$1"
if command -v nc >/dev/null 2>&1; then
nc -z -w1 127.0.0.1 "$port" 2>/dev/null
else
( exec 3<>"/dev/tcp/127.0.0.1/$port" ) 2>/dev/null
fi
}
# --------------------------------------------------------------------------------- the log marker
#
# Take the log's line count BEFORE touching anything. Every later read of "what did the daemon
# say" starts strictly after this mark, so a refusal from hours ago can never be mistaken for
# this edit's verdict. Same approach as scripts/redeploy-fleetd.sh's RESTART_MARK.
log_mark() {
local file="$1"
if [ -f "$file" ]; then
wc -l < "$file" 2>/dev/null || echo 0
else
echo 0
fi
}
read_verdict_after_marker() {
local file="$1" mark="$2"
[ -f "$file" ] || return 0
tail -n "+$((mark + 1))" "$file" 2>/dev/null || true
}
# Classifies one log LINE. Echoes one of: refused | clean | needs-restart | none. Always
# succeeds (every branch ends in `echo`), so it is safe to call from inside `$( )`.
classify_verdict_line() {
local line="$1"
case "$line" in
*'config reload refused'*) echo refused ;;
*'config reload from '*'refused, keeping the running config'*) echo refused ;;
*'config reloaded'*)
case "$line" in
*'need a restart'*|*'partially live'*) echo needs-restart ;;
*) echo clean ;;
esac ;;
*) echo none ;;
esac
return 0
}
scan_region_for_verdict() {
local region="$1" line kind
[ -n "$region" ] || return 1
while IFS= read -r line || [ -n "$line" ]; do
kind="$(classify_verdict_line "$line")"
if [ "$kind" != "none" ]; then
VERDICT_KIND="$kind"
VERDICT_LINE="$line"
return 0
fi
done <<< "$region"
return 1
}
# Sets VERDICT_KIND/VERDICT_LINE and returns 0 on the first verdict line found after $mark;
# returns 1 (VERDICT_KIND=none) if none appeared inside $wait_s seconds. Checks once before each
# sleep AND once more after the last sleep, the same boundary idiom
# scripts/redeploy-fleetd.sh's wait_for_daemon_exit/wait_for_new_pid already use.
wait_for_verdict() {
local log="$1" mark="$2" wait_s="$3" _i region
VERDICT_KIND="none"
VERDICT_LINE=""
for _i in $(seq "$wait_s"); do
region="$(read_verdict_after_marker "$log" "$mark")"
scan_region_for_verdict "$region" && return 0
sleep 1
done
region="$(read_verdict_after_marker "$log" "$mark")"
scan_region_for_verdict "$region" && return 0
return 1
}
last_verdict_line() {
local file="$1" line out=""
[ -f "$file" ] || return 0
while IFS= read -r line || [ -n "$line" ]; do
if [ "$(classify_verdict_line "$line")" != "none" ]; then
out="$line"
fi
done < "$file"
printf '%s' "$out"
}
default_wait_seconds() {
local log="$1" interval=""
if [ -f "$log" ]; then
interval="$(grep -F 'config watch:' "$log" 2>/dev/null | tail -1 \
| sed -E 's/.*\(every ([0-9]+)s\).*/\1/' || true)"
fi
case "$interval" in
''|*[!0-9]*) interval=10 ;;
esac
echo $((interval * 4))
}
# ----------------------------------------------------------------------------------- the backup
#
# Timestamped, never pruned — "keep backups" per the ticket. A pid suffix avoids a same-second
# collision between two invocations.
#
# fleetd #635 follow-up — lands under a DEDICATED, gitignored directory beside the config
# (<dir>/.config-backups/), never beside the config file itself. The whole reason fleetd.yaml is
# gitignored is that it must never be committed, and a backup of it inherits that requirement — a
# bare `fleetd.yaml.bak.*` next to a tracked directory is one `git add -A`/`git add .` away from
# committing the live config. A directory beats a glob on its own: the glob only protects today's
# naming, a location keeps working even if the naming changes later. (See .gitignore for the glob
# kept anyway, as a backstop for a stray backup written the old way.)
BACKUP_DIRNAME=".config-backups"
backup_dir_for() {
local src="$1"
printf '%s/%s' "$(dirname "$src")" "$BACKUP_DIRNAME"
}
backup_config() {
local src="$1" ts backup dir base
dir="$(backup_dir_for "$src")"
mkdir -p "$dir" \
|| die "could not create the backup directory $dir — refusing to edit without a backup. The live config at $src was NOT touched."
ts="$(date -u +%Y%m%dT%H%M%S)Z"
base="$(basename "$src")"
backup="${dir}/${base}.bak.${ts}.$$"
cp "$src" "$backup" \
|| die "could not create a backup at $backup — refusing to edit without one. The live config at $src was NOT touched."
printf '%s' "$backup"
}
newest_backup() {
local cfg="$1" dir base
dir="$(backup_dir_for "$cfg")"
base="$(basename "$cfg")"
ls -t "${dir}/${base}".bak.* 2>/dev/null | head -1 || true
}
# ------------------------------------------------------------------------------------ the file mode
#
# fleetd #635 follow-up — `mv` from a mktemp candidate carries mktemp's 0600 onto the live path
# forever (measured: 644 -> 600 after one --set), and a restore does not undo it either, because
# `cp` onto an EXISTING file keeps the DESTINATION's mode, not the source's. Capture the live
# file's mode before anything touches it, and reapply it to whatever lands on that path
# afterwards — the candidate before install, and the config again after a restore — so an edit
# changes the file's CONTENT only, never its permissions. BSD `stat -f '%Lp'` first (matches this
# project's dev machine), GNU `stat -c '%a'` as the fallback. Prints nothing when the file does
# not exist yet, so apply_mode then does nothing and a first-ever edit falls back to the normal
# umask default rather than inventing a number.
file_mode() {
local file="$1"
[ -f "$file" ] || return 0
stat -f '%Lp' "$file" 2>/dev/null || stat -c '%a' "$file" 2>/dev/null || true
}
apply_mode() {
local file="$1" mode="$2"
[ -n "$mode" ] || return 0
chmod "$mode" "$file" 2>/dev/null || true
}
# --------------------------------------------------------------------------- candidate builders
#
# Never edit the live file in place. Each builder fills $1 (a temp file already sitting in the
# SAME directory as the live config, so the later `mv` install is a rename, not a cross-device
# copy — see run_edit).
# fleetd #635 follow-up — a forgotten value (`--set .a.b=`, a plausible typo) must never be
# accepted as "clear the field". `*=*` alone cannot tell "--set .a.b=" from "--set .a.b=7" apart
# — both contain an `=` — so the guard has to look at the VALUE, not the shape of the argument.
# An empty value refuses outright: nothing is installed, and the message names the likely cause
# AND the two ways to actually mean it (clear on purpose, or an intentional empty string via
# --from). Measured against the real daemon loader: a quoted empty string reads back as a null
# field (`quoted empty -> OK int=null`), and a null numeric field FALLS BACK TO ITS DEFAULT rather
# than erroring — so this is not a cosmetic nit, it is the one shape of edit that widens capacity
# silently instead of failing loudly, which is exactly what this script exists to catch.
#
# A deliberate clear needs its own spelling, because `""` and YAML `null` are NOT the same value
# to the loader (`""` is a valid empty String; `null` means absent, and an Integer field reads
# either the same way — null — but a String field would keep `""` as a real value). `--set
# .a.b=null` is that spelling: it writes a literal, unquoted `null` via yq, never the string
# "null" through strenv(). One consequence worth knowing: there is currently no --set spelling
# for the three-character STRING "null" itself (it collides with the clear spelling) — use
# --from for that rare case.
apply_set_pairs() {
local cand="$1" kv path value
shift
for kv in "$@"; do
case "$kv" in
*=*) : ;;
*) die "--set expects <yq-path>=<value>, got: '$kv'" ;;
esac
path="${kv%%=*}"
path="${path#.}"
value="${kv#*=}"
if [ -z "$value" ]; then
die "--set '$kv' has an EMPTY value — refusing. Nothing was installed. A forgotten value
would NULL the field, and a null value falls back to its default rather than erroring —
silent, not safe. Did you mean --set .${path}=null to clear it on purpose, or --from a
file if you need a genuinely empty string?"
fi
# fleetd #635 follow-up (ticket comment 17673, defect 8) — these two failure messages used to
# echo the full "$kv" (path=value, exactly as the operator typed it), unredacted. The operator
# already has the value, so a terminal is not where this leaks — the risk is where the output
# goes NEXT: this fleet pastes command output into tickets, PRs and fleet_reply bodies, and a
# failure is exactly when someone copies it to ask for help. Print the PATH, which is what's
# needed to fix the command, and never the value. $kv is not key:value-shaped YAML, so piping
# it through redact would just pass it straight through — a false sense of coverage, the same
# mistake as defect 7.
if [ "$value" = "null" ]; then
yq eval -i ".${path} = null" "$cand" \
|| die "yq could not clear --set '.${path}=null' — nothing was installed. The live config is unchanged."
continue
fi
CONFIG_EDIT_SET_VALUE="$value" yq eval -i ".${path} = strenv(CONFIG_EDIT_SET_VALUE)" "$cand" \
|| die "yq could not apply --set '.${path}=<value>' — nothing was installed. The live config is unchanged."
done
}
build_from_set() {
local cand="$1"
cp "$CONFIG" "$cand"
apply_set_pairs "$cand" "${SETS[@]}"
}
build_from_file() {
local cand="$1"
[ -f "$FROM_FILE" ] || die "--from file not found: $FROM_FILE"
cp "$FROM_FILE" "$cand"
}
parse_check() {
yq eval '.' "$1" >/dev/null 2>&1
}
install_candidate() {
local cand="$1" live="$2"
mv -f "$cand" "$live"
}
# -------------------------------------------------------------------------------- the report path
#
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
restore_command_line() {
printf '%q --restore --config %q --log %q --wait-seconds %q' "$SELF" "$CONFIG" "$LOG" "$WAIT_SECONDS"
}
# State 4 only: restore the pre-edit backup, then wait for a SECOND verdict confirming the
# restore itself reloaded cleanly. Never claims a restore it did not observe — if the second wait
# also times out, it says so plainly rather than reporting "restored" as though confirmed.
restore_and_confirm() {
local backup="$1" mark2 orig_mode
orig_mode="$(file_mode "$CONFIG")"
mark2="$(log_mark "$LOG")"
cp "$backup" "$CONFIG" \
|| die "could not restore $backup onto $CONFIG — the live config is left as the REFUSED edit. Fix this by hand immediately: cp \"$backup\" \"$CONFIG\""
apply_mode "$CONFIG" "$orig_mode"
ok "restored from $backup"
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
case "$VERDICT_KIND" in
refused) warn "the RESTORE was also refused by the daemon: $VERDICT_LINE" ;;
*) ok "restore confirmed: $VERDICT_LINE" ;;
esac
else
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
warn "cannot confirm the restore reloaded cleanly — check $LOG by hand"
fi
return 0
}
# The four-outcome decision. Echoed as a function so run_edit/restore_mode share one place that
# can return 0/3/4/5 — never duplicated, never re-worded between the two callers.
report_outcome() {
local mark="$1" backup="$2" kind line
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
kind="$VERDICT_KIND"; line="$VERDICT_LINE"
else
kind="none"
fi
case "$kind" in
clean)
ok "daemon verdict: $line"
say "result: applied cleanly"
return 0 ;;
needs-restart)
ok "daemon verdict: $line"
say "result: applied — a restart is needed for the change(s) named above"
return 3 ;;
refused)
warn "daemon verdict: $line"
say "result: REFUSED — restoring the backup"
restore_and_confirm "$backup"
return 4 ;;
none)
warn "no verdict line appeared within ${WAIT_SECONDS}s after $LOG line $mark"
warn "CANNOT TELL whether the daemon applied this edit, refused it, or is simply down."
warn "Nothing was restored — the edit is still on disk at $CONFIG."
echo
echo " backup: $backup"
echo " to restore it by hand:"
echo " $(restore_command_line)"
return 5 ;;
esac
}
# ------------------------------------------------------------------------------------- the modes
check_mode() {
say "config-edit --check"
if [ -f "$CONFIG" ]; then
if parse_check "$CONFIG"; then
ok "config parses: $CONFIG"
else
warn "config does NOT parse as valid YAML: $CONFIG"
fi
else
warn "no config file at $CONFIG"
fi
local port
port="$(resolve_port "$CONFIG")"
if daemon_listening "$port"; then
ok "daemon is listening on 127.0.0.1:$port"
else
warn "no daemon detected listening on 127.0.0.1:$port"
fi
ok "watch interval assumed: $(( $(default_wait_seconds "$LOG") / 4 ))s (derives --wait-seconds default of $(default_wait_seconds "$LOG")s)"
local verdict
verdict="$(last_verdict_line "$LOG")"
if [ -n "$verdict" ]; then
ok "last verdict in log: $verdict"
else
warn "no reload verdict line found in $LOG"
fi
local backup
backup="$(newest_backup "$CONFIG")"
if [ -n "$backup" ]; then
ok "newest backup: $backup"
else
warn "no backups found for $CONFIG"
fi
if command -v yq >/dev/null 2>&1; then
ok "yq: $(yq --version 2>&1)"
else
warn "yq not found on PATH"
fi
return 0
}
# Shared by --set and --from: backup, build, parse-check, redacted diff, install, await verdict.
run_edit() {
local builder="$1"
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to edit"
local mark orig_mode
mark="$(log_mark "$LOG")"
orig_mode="$(file_mode "$CONFIG")"
say "probe"
local port
port="$(resolve_port "$CONFIG")"
if daemon_listening "$port"; then
ok "daemon appears to be listening on 127.0.0.1:$port"
else
warn "no daemon detected listening on 127.0.0.1:$port — a verdict may never appear"
fi
say "backup"
local backup
backup="$(backup_config "$CONFIG")"
ok "backup: $backup"
say "candidate"
local cand
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|| die "could not create a candidate temp file next to $CONFIG"
cand="$CAND"
if ! "$builder" "$cand"; then
rm -f "$cand"; CAND=""
die "could not build the candidate — nothing was installed. The live config at $CONFIG is unchanged."
fi
if ! parse_check "$cand"; then
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — nothing was installed. The live config at $CONFIG is unchanged."
fi
ok "candidate parses"
apply_mode "$cand" "$orig_mode"
say "change (redacted)"
diff -u "$backup" "$cand" | redact || true
say "install"
install_candidate "$cand" "$CONFIG" \
|| die "could not install the candidate onto $CONFIG — the live config was NOT changed. The validated candidate is sitting at $cand; investigate before retrying."
CAND=""
ok "installed: $CONFIG"
local rc=0
report_outcome "$mark" "$backup" || rc=$?
return "$rc"
}
dry_run_diff() {
local builder="$1"
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to diff against"
local cand
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|| die "could not create a candidate temp file next to $CONFIG"
cand="$CAND"
if ! "$builder" "$cand"; then
rm -f "$cand"; CAND=""
die "could not build the candidate — this was a --dry-run, nothing would have been installed either"
fi
if ! parse_check "$cand"; then
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
fi
say "dry run — diff (redacted), nothing installed"
diff -u "$CONFIG" "$cand" | redact || true
rm -f "$cand"; CAND=""
return 0
}
restore_mode() {
[ -f "$CONFIG" ] || die "no config at $CONFIG to restore onto"
local backup dir base
backup="$(newest_backup "$CONFIG")"
if [ -z "$backup" ]; then
# fleetd #635 follow-up (ticket comment 17664) — this message must name the directory the
# code actually searches (backup_dir_for, same as newest_backup), not the old beside-the-
# config glob. A backup written the OLD way is real and NOT searched any more — say so and
# give the one-line recovery command — but do NOT make the search itself look there; that
# would be a behaviour change nobody asked for. The message is the only thing being fixed.
dir="$(backup_dir_for "$CONFIG")"
base="$(basename "$CONFIG")"
die "no backup found matching ${dir}/${base}.bak.* — nothing to restore.
A backup written the OLD way, directly beside the config (${CONFIG}.bak.*), is NOT
searched — that location was retired so a backup of a file that must never be committed
cannot sit next to a tracked directory. If one exists there, recover it by hand:
cp ${CONFIG}.bak.<timestamp>.<pid> $CONFIG"
fi
[ -f "$backup" ] || die "backup candidate $backup vanished"
say "restore"
ok "restoring $backup onto $CONFIG"
local mark orig_mode
mark="$(log_mark "$LOG")"
orig_mode="$(file_mode "$CONFIG")"
cp "$backup" "$CONFIG" || die "could not copy $backup onto $CONFIG"
apply_mode "$CONFIG" "$orig_mode"
ok "installed: $CONFIG"
local rc=0
report_outcome "$mark" "$backup" || rc=$?
return "$rc"
}
# -------------------------------------------------------------------------------------- dispatch
if [ -n "$WAIT_SECONDS_OVERRIDE" ]; then
WAIT_SECONDS="$WAIT_SECONDS_OVERRIDE"
else
WAIT_SECONDS="$(default_wait_seconds "$LOG")"
fi
RC=0
case "$MODE" in
check)
check_mode || RC=$?
;;
set)
[ "${#SETS[@]}" -gt 0 ] || die "--set requires at least one <yq-path>=<value>"
if [ "$DRY_RUN" = 1 ]; then
dry_run_diff build_from_set || RC=$?
else
run_edit build_from_set || RC=$?
fi
;;
from)
[ -n "$FROM_FILE" ] || die "--from requires a candidate file path"
if [ "$DRY_RUN" = 1 ]; then
dry_run_diff build_from_file || RC=$?
else
run_edit build_from_file || RC=$?
fi
;;
restore)
restore_mode || RC=$?
;;
esac
exit "$RC"
+585
View File
@@ -0,0 +1,585 @@
#!/usr/bin/env bash
# Self-contained checks for scripts/config-edit.sh — fleetd ticket #635.
#
# Drives the REAL config-edit.sh as a subprocess against a FIXTURE config and a FIXTURE log in a
# throwaway temp directory this file creates and removes. Never touches fleetd/fleetd.yaml or
# fleetd/fleetd.out, and never starts, stops, or contacts a daemon — there is no daemon here, so
# each test PLAYS the daemon: it starts config-edit.sh in the background (it is waiting on the
# log), appends the verdict line it wants, then collects the real exit code.
set -euo pipefail
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
EDIT="$ROOT/scripts/config-edit.sh"
TMP="$(mktemp -d "$ROOT/.config-edit-test.XXXXXX")"
trap 'rm -rf "$TMP"' EXIT
fail() {
printf 'FAIL: %s\n' "$*" >&2
return 1
}
assert_equals() {
local expected="$1" actual="$2" description="$3"
[ "$expected" = "$actual" ] || fail "$description: expected $expected, got $actual"
}
assert_contains() {
local needle="$1" text="$2" description="$3"
printf '%s' "$text" | grep -qF -- "$needle" || fail "$description: missing [$needle]"
}
assert_not_contains() {
local needle="$1" text="$2" description="$3"
if printf '%s' "$text" | grep -qF -- "$needle"; then
fail "$description: must NOT contain [$needle], but it does"
fi
return 0
}
# A fresh fixture pair per test: $1/fleetd.yaml (the config) and $1/fleetd.out (the log), plus a
# small wait-seconds budget so no test takes long. Returns the fixture dir via stdout.
new_fixture() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
# Runs config-edit.sh in the background against $dir's fixtures, with the given extra args, and
# a short --wait-seconds. Sets RUN_PID. Caller appends to $dir/fleetd.out (or not, for the
# silence test) and then calls collect_run to block for the exit code.
start_run() {
local dir="$1" wait_s="$2"; shift 2
(
# config-edit.sh deliberately exits 3/4/5 on several of these tests. This subshell inherits
# the parent's `set -e`, and without disabling it here the FIRST nonzero exit would kill the
# subshell before the `echo $? > rc` line ever ran — the real code would never reach the file.
set +e
"$EDIT" --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds "$wait_s" "$@" \
> "$dir/stdout.log" 2>&1
echo $? > "$dir/rc"
) &
RUN_PID=$!
}
collect_run() {
local dir="$1"
# wait echoes back the backgrounded subshell's own exit status (here, deliberately 3/4/5 on
# several tests) — under `set -e` a bare nonzero `wait` would abort this whole test script, so
# it is neutralized with `|| true`; the real code is read from $dir/rc right after.
wait "$RUN_PID" || true
RUN_OUTPUT="$(cat "$dir/stdout.log")"
RUN_RC="$(cat "$dir/rc")"
}
# -------------------------------------------------------------- acceptance criterion 1: refusal
test_refusal_restores_byte_for_byte() {
local dir
dir="$(new_fixture)"
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
start_run "$dir" 5 --set '.broker.uri=amqp://changed@host/x'
sleep 1
printf 'config reload refused — these keys cannot change under a running daemon: broker. Restart fleetd to apply them.\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "refusal exit code"
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|| fail "refusal must restore the config byte for byte onto the pre-edit backup"
}
# -------------------------------------------------------------- acceptance criterion 2: clean
test_clean_reload_keeps_the_edit() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=7'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "clean reload exit code"
assert_equals "7" "$(yq eval '.profiles.sonnet.weight' "$dir/fleetd.yaml")" "clean reload live value"
}
# ----------------------------------------------------- acceptance criterion 3: deferred != clean
test_deferred_reload_is_told_apart_from_clean() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=9'
sleep 1
printf 'config reloaded; these changes need a restart to take effect: profiles\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 3 "$RUN_RC" "deferred reload exit code"
[ "$RUN_RC" != 0 ] || fail "deferred reload must not report exit 0"
assert_contains "profiles" "$RUN_OUTPUT" "deferred reload names the key"
assert_contains "restart" "$RUN_OUTPUT" "deferred reload says a restart is needed"
}
# -------------------------------------------------------------- acceptance criterion 4: silence
test_silence_is_its_own_answer() {
local dir
dir="$(new_fixture)"
start_run "$dir" 2 --set '.profiles.sonnet.weight=11'
# Feed the log nothing.
collect_run "$dir"
assert_equals 5 "$RUN_RC" "silence exit code"
assert_equals "11" "$(yq eval '.profiles.sonnet.weight' "$dir/fleetd.yaml")" "the edited value must still be on disk"
assert_contains '--restore' "$RUN_OUTPUT" "silence prints the --restore command"
local backup restore_cmd
backup="$(ls -t "$dir"/.config-backups/fleetd.yaml.bak.* | head -1)"
[ -n "$backup" ] || fail "silence must still have taken a backup"
restore_cmd="$(printf '%s\n' "$RUN_OUTPUT" | grep -F -- '--restore --config' | sed -E 's/^[[:space:]]*//')"
[ -n "$restore_cmd" ] || fail "could not find the printed --restore invocation in the output"
# Running this --restore invocation installs the backup, then itself waits for a confirming
# verdict that this fixture never feeds — so it legitimately exits 5 ("cannot tell") here, same
# as any edit with no daemon on the other end. Only a usage/internal error (1 or 2) is a real
# failure of the command itself; the actual assertion is the byte-for-byte cmp below.
local restore_rc=0
eval "$restore_cmd" > "$dir/restore.log" 2>&1 || restore_rc=$?
case "$restore_rc" in
0|3|4|5) : ;;
*) fail "the printed --restore command errored out (exit $restore_rc): $(cat "$dir/restore.log")" ;;
esac
cmp -s "$dir/fleetd.yaml" "$backup" \
|| fail "running the printed --restore command must put the file back to the original backup"
}
# --------------------------------------------------------- acceptance criterion 5: bad candidate
test_broken_candidate_never_reaches_live_path() {
local dir rc=0
dir="$(new_fixture)"
printf 'foo: [unclosed\n' > "$dir/broken.yaml"
"$EDIT" --from "$dir/broken.yaml" --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
> "$dir/stdout.log" 2>&1 || rc=$?
[ "$rc" -ne 0 ] || fail "a broken --from candidate must exit non-zero"
cmp -s "$dir/fleetd.yaml" <(new_fixture_yaml) \
|| fail "the broken candidate must never reach the live fixture config"
}
new_fixture_yaml() {
cat <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
}
# -------------------------------------------------------------------- acceptance criterion 6
test_marker_skips_lines_before_it() {
local dir
dir="$(new_fixture)"
printf 'config reload refused — something ancient\n' > "$dir/fleetd.out"
start_run "$dir" 5 --set '.profiles.sonnet.weight=5'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "a stale refusal before the marker must not be read as this edit's verdict"
}
# ------------------------------------------------------------------- acceptance criterion 7 (+13)
# fleetd #635 follow-up (ticket comment 17659) — the two assertions below this comment were the
# WHOLE test before the follow-up, and both are negative-only: they pass just as happily when the
# diff is never printed at all as when it is printed and correctly redacted. A mutant that deletes
# `diff -u "$backup" "$cand" | redact` from the edit path survives them, because an absent output
# contains neither "hunter2" nor "user:" either — see the mutation-and-revert proof in the reply.
# Criterion 13 is the fix: a LOUD positive control that only passes when a diff was demonstrably
# printed AND the redaction demonstrably ran on real content, not merely that nothing leaked.
test_redaction_holds() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "redaction-case reload exit code"
assert_not_contains "hunter2" "$RUN_OUTPUT" "full output must never contain the password"
assert_not_contains "user:" "$RUN_OUTPUT" "full output must never contain the userinfo"
# acceptance criterion 13 — positive control: the diff's default 3-line context around the
# changed "weight" key also covers the fixture's "uri:" line, so a genuinely-printed, genuinely-
# redacted diff must contain BOTH the redaction marker and the changed key's name. A test that
# only ever asserts absence cannot tell "redacted" from "never printed" apart; this can.
assert_contains "<redacted>" "$RUN_OUTPUT" "the redaction must be PROVEN to have run on real content, not merely absent"
assert_contains "weight" "$RUN_OUTPUT" "a diff must have been demonstrably printed at all"
}
# ------------------------------------------------------- acceptance criterion 9: forgotten value
# `--set .a.b=` is a plausible typo (the value simply forgotten), and it must be refused outright
# rather than silently nulling the field — a null numeric field falls back to its default, which
# widens capacity instead of failing loudly. No background verdict feeder here: a refused --set
# must never even reach the daemon, so this never starts a background run at all.
test_forgotten_value_refuses_and_installs_nothing() {
local dir rc=0
dir="$(new_fixture)"
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
"$EDIT" --set '.profiles.sonnet.maxLoad=' \
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
> "$dir/stdout.log" 2>&1 || rc=$?
RUN_OUTPUT="$(cat "$dir/stdout.log")"
[ "$rc" -ne 0 ] || fail "an empty --set value must exit non-zero, got 0"
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|| fail "an empty --set value must install nothing — the live fixture changed"
assert_contains "EMPTY value" "$RUN_OUTPUT" "the refusal must name the empty value"
}
# ---------------------------------------------------------- acceptance criterion 10: explicit null
# `--set .a.b=null` is the deliberate-clear spelling, and it must write a REAL yaml null, never
# the string "''" — those are different values to the daemon's loader (fleetd ticket #635's
# follow-up comment measured `""` reading back as a null field anyway, which is exactly why the
# two forms must not collapse onto each other: `--set path=` refuses instead of silently reaching
# this same null outcome through the back door). Read the RAW line with grep, never only through
# `yq` — `yq eval` reports `null` for both an actual null and a missing/absent key, so it cannot
# tell "wrote null" apart from "wrote nothing"; only the literal line on disk can.
test_explicit_null_writes_bare_null_not_empty_string() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.maxLoad=null'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "explicit null clear exit code"
local raw_line
raw_line="$(grep -E 'maxLoad' "$dir/fleetd.yaml")"
assert_contains "null" "$raw_line" "the installed line must spell a bare null"
assert_not_contains '""' "$raw_line" "the installed line must NOT be a quoted empty string"
}
# ------------------------------------------------------- acceptance criterion 11: backup never committable
# A backup of fleetd.yaml inherits fleetd.yaml's own "never commit this" requirement (fleetd #635
# follow-up, ticket comment 17655). Proves two things: the backup lands somewhere `git
# check-ignore` reports as ignored (equivalently, a path `git status --porcelain` never lists as
# untracked), AND that --restore still finds and uses it from that location.
test_backup_is_never_committable() {
local dir backup
dir="$(new_fixture)"
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
start_run "$dir" 5 --set '.profiles.sonnet.weight=55'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "setup edit exit code"
backup="$(ls -t "$dir"/.config-backups/fleetd.yaml.bak.* 2>/dev/null | head -1)"
[ -n "$backup" ] || fail "no backup found under .config-backups/ — did the location change?"
git -C "$ROOT" check-ignore -q -- "$backup" \
|| fail "the backup at $backup is NOT gitignored — it would survive a git add -A"
if git -C "$ROOT" status --porcelain -- "$backup" 2>/dev/null | grep -q '^??'; then
fail "git status still lists the backup as untracked: $backup"
fi
start_run "$dir" 5 --restore
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "--restore after the backup-location change exit code"
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|| fail "--restore from the new backup location must still put the file back byte for byte"
}
# --------------------------------------------------------- acceptance criterion 12: file mode
# `mv` from a mktemp candidate carries mktemp's 0600 forever, and a plain `cp` onto an existing
# file keeps the DESTINATION's mode rather than the source's, so a restore does not undo the
# narrowing either (fleetd #635 follow-up, ticket comment 17657). Proves the mode survives an edit
# AND a subsequent restore, from two different starting points — 644 is the common case, 600
# proves the fix PRESERVES whatever mode was there rather than hardcoding 644.
test_file_mode_survives_edit_and_restore() {
local dir want got
for want in 644 600; do
dir="$(new_fixture)"
chmod "$want" "$dir/fleetd.yaml"
start_run "$dir" 5 --set ".profiles.sonnet.weight=${want}"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "mode-preservation setup edit exit code ($want)"
got="$(stat -f '%Lp' "$dir/fleetd.yaml" 2>/dev/null || stat -c '%a' "$dir/fleetd.yaml")"
assert_equals "$want" "$got" "mode must survive a --set ($want)"
start_run "$dir" 5 --restore
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "mode-preservation restore exit code ($want)"
got="$(stat -f '%Lp' "$dir/fleetd.yaml" 2>/dev/null || stat -c '%a' "$dir/fleetd.yaml")"
assert_equals "$want" "$got" "mode must survive a --restore ($want)"
done
}
# ----------------------------------------- acceptance criterion 14: restore message names the real directory
# fleetd #635 follow-up (ticket comment 17664, defect 6) — the --restore "no backup found"
# message used to print the OLD beside-the-config glob even though newest_backup had already
# moved to searching the managed directory. Proves BOTH directions: the not-found message names
# the directory actually searched (not merely that it says SOMETHING), and that a real backup
# sitting in that directory still lets --restore succeed — otherwise the fix could regress into
# a message that is always printed regardless of whether a backup exists.
test_restore_message_names_the_searched_directory() {
local dir rc=0
# Direction 1: no backup anywhere — the message must name .config-backups/, not the bare
# beside-the-config glob the OLD code printed.
dir="$(new_fixture)"
"$EDIT" --restore --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
> "$dir/stdout.log" 2>&1 || rc=$?
RUN_OUTPUT="$(cat "$dir/stdout.log")"
assert_equals 1 "$rc" "--restore with no backup anywhere exit code"
assert_contains ".config-backups/fleetd.yaml.bak.*" "$RUN_OUTPUT" \
"the not-found message must name the directory actually searched, not the old beside-the-config glob"
# Direction 2: a real backup IS present in .config-backups/ — --restore must still succeed, so
# the message fix cannot have turned into one that prints regardless of whether a backup exists.
dir="$(new_fixture)"
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
start_run "$dir" 5 --set '.profiles.sonnet.weight=77'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "setup edit exit code for criterion 14's second half"
start_run "$dir" 5 --restore
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "--restore with a real backup present must still succeed"
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|| fail "--restore with a real backup present must put the file back byte for byte"
}
# ------------------------- acceptance criterion 15a: block-scalar continuation lines are redacted
# fleetd #635 follow-up (ticket comment 17670, defect 7) — redact() used to look only AT the key
# line. A YAML block scalar (`|`) puts its value on the lines that FOLLOW the key, each indented
# deeper than it, so the real secret flowed through untouched while the key line right above it
# printed a reassuring "<redacted>" — worse than no redaction, because the marker stops a reader
# from looking further. The edited key here ("retries") sits directly next to the block scalar,
# well inside diff -u's default 3-line context window, so the printed hunk is GUARANTEED to
# include the secret's lines — placing the edit further away would let this pass today even
# without the fix, proving nothing (the ticket comment's own warning, from the lead's first
# reproduction attempt). The positive control runs FIRST: without it, "the secret never entered
# the diff at all" would pass identically to "it entered and was correctly redacted".
new_fixture_block_scalar() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
auth:
token: |
FAKELEAK-BLOCK-SCALAR
retries: 1
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_block_scalar_continuation_is_redacted() {
local dir
dir="$(new_fixture_block_scalar)"
start_run "$dir" 5 --set '.auth.retries=2'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "block-scalar case reload exit code"
# Positive control FIRST: the key's own (masked) line must really be in the printed diff, or the
# negative assertion right after proves nothing — see the comment above this test.
assert_contains "token:" "$RUN_OUTPUT" "block-scalar case: the key's line must be in the printed diff"
assert_contains "<redacted>" "$RUN_OUTPUT" "block-scalar case: redaction must be proven to have run on real content"
assert_not_contains "FAKELEAK-BLOCK-SCALAR" "$RUN_OUTPUT" "block-scalar case: the block scalar's VALUE must never leak"
}
# ------------------------------------- acceptance criterion 15b: "passphrase" is also recognised
# "passphrase" was in none of TOKEN|SECRET|PASSWORD|PASSWD|CREDENTIAL|URI|_KEY (ticket comment
# 17670). This is a plain key:value line, not a block scalar — kept in its OWN fixture and OWN
# function, separate from criterion 15a, so that a failure in one case can never mask a failure in
# the other (a single combined test would abort under `set -e` at its first failing assertion,
# and the second case would then never even run).
new_fixture_passphrase() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
auth:
passphrase: FAKELEAK-PASSPHRASE
retries: 1
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_passphrase_key_is_redacted() {
local dir
dir="$(new_fixture_passphrase)"
start_run "$dir" 5 --set '.auth.retries=2'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "passphrase case reload exit code"
assert_contains "passphrase:" "$RUN_OUTPUT" "passphrase case: the key's line must be in the printed diff"
assert_contains "<redacted>" "$RUN_OUTPUT" "passphrase case: redaction must be proven to have run on real content"
assert_not_contains "FAKELEAK-PASSPHRASE" "$RUN_OUTPUT" "passphrase case: the passphrase VALUE must never leak"
}
# ----------------------------------- acceptance criterion 16: a failing --set must not echo value
# fleetd #635 follow-up (ticket comment 17673, defect 8) — apply_set_pairs used to echo the FULL
# "$kv" (path=value, exactly as typed) in its yq-failure messages, so a broken --set with a
# secret-looking value printed that value right back out. The path alone is what the positive
# control proves is still there — it is what the operator needs to fix their command — and the
# negative assertion proves the value itself never appears. Kept to exactly this one failure
# shape (an invalid yq path/expression), matching the ticket's own reproduction.
test_failing_set_does_not_echo_its_value() {
local dir rc=0
dir="$(new_fixture)"
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
"$EDIT" --dry-run --set '.broker.["bad=FAKELEAK-SETVALUE' \
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
> "$dir/stdout.log" 2>&1 || rc=$?
RUN_OUTPUT="$(cat "$dir/stdout.log")"
[ "$rc" -ne 0 ] || fail "a --set with an invalid yq expression must exit non-zero, got 0"
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|| fail "a failing --set must install nothing — the live fixture changed"
# Positive control FIRST: the path must still be in the message, or the negative assertion right
# after proves nothing (the message could simply have disappeared entirely).
assert_contains '.broker.["bad' "$RUN_OUTPUT" "the failure message must still name the PATH"
assert_not_contains "FAKELEAK-SETVALUE" "$RUN_OUTPUT" "the failure message must NEVER echo the VALUE"
}
# dry-run must never touch the live file and must still redact.
test_dry_run_never_installs_and_redacts() {
local dir before
dir="$(new_fixture)"
before="$(cat "$dir/fleetd.yaml")"
"$EDIT" --dry-run --set '.profiles.sonnet.weight=99' \
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
> "$dir/stdout.log" 2>&1
local rc=$?
RUN_OUTPUT="$(cat "$dir/stdout.log")"
assert_equals 0 "$rc" "dry-run exit code"
assert_equals "$before" "$(cat "$dir/fleetd.yaml")" "dry-run must never write the live config"
assert_not_contains "hunter2" "$RUN_OUTPUT" "dry-run diff must also be redacted"
assert_contains "99" "$RUN_OUTPUT" "dry-run diff must show the candidate value"
# Same positive-control reasoning as acceptance criterion 13, applied to the dry-run diff path.
assert_contains "<redacted>" "$RUN_OUTPUT" "the dry-run diff's redaction must be PROVEN to have run, not merely absent"
}
# --check is read-only and always exits 0, even against a dead "daemon".
test_check_is_read_only_and_exits_zero() {
local dir before rc=0
dir="$(new_fixture)"
before="$(cat "$dir/fleetd.yaml")"
"$EDIT" --check --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" \
> "$dir/stdout.log" 2>&1 || rc=$?
assert_equals 0 "$rc" "--check exit code"
assert_equals "$before" "$(cat "$dir/fleetd.yaml")" "--check must never modify the config"
}
test_refusal_shape_from_parse_failure_wording_is_recognised() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=6'
sleep 1
printf 'config reload from %s refused, keeping the running config: boom\n' "$dir/fleetd.yaml" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
}
echo "== acceptance criterion 1: refusal restores byte for byte =="
test_refusal_restores_byte_for_byte
echo "== acceptance criterion 2: clean reload keeps the edit =="
test_clean_reload_keeps_the_edit
echo "== acceptance criterion 3: deferred reload told apart from clean =="
test_deferred_reload_is_told_apart_from_clean
echo "== acceptance criterion 4: silence is its own answer =="
test_silence_is_its_own_answer
echo "== acceptance criterion 5: broken candidate never reaches the live path =="
test_broken_candidate_never_reaches_live_path
echo "== acceptance criterion 6: the marker works =="
test_marker_skips_lines_before_it
echo "== acceptance criterion 7 (+13: redaction is proven to have run) =="
test_redaction_holds
echo "== acceptance criterion 9: a forgotten value refuses and installs nothing =="
test_forgotten_value_refuses_and_installs_nothing
echo "== acceptance criterion 10: an explicit clear writes a bare null =="
test_explicit_null_writes_bare_null_not_empty_string
echo "== acceptance criterion 11: a backup is never committable =="
test_backup_is_never_committable
echo "== acceptance criterion 12: the file mode survives an edit and a restore =="
test_file_mode_survives_edit_and_restore
echo "== acceptance criterion 14: the restore message names the directory actually searched =="
test_restore_message_names_the_searched_directory
echo "== acceptance criterion 15a: a block scalar's continuation lines are redacted =="
test_block_scalar_continuation_is_redacted
echo "== acceptance criterion 15b: a passphrase key is also recognised =="
test_passphrase_key_is_redacted
echo "== acceptance criterion 16: a failing --set must not echo its value =="
test_failing_set_does_not_echo_its_value
echo "== extra: dry-run never installs, and redacts =="
test_dry_run_never_installs_and_redacts
echo "== extra: --check is read-only and always exits 0 =="
test_check_is_read_only_and_exits_zero
echo "== extra: the parse-failure refusal shape is also recognised =="
test_refusal_shape_from_parse_failure_wording_is_recognised
printf 'PASS: config-edit acceptance criteria\n'