Commit Graph

1074 Commits

Author SHA1 Message Date
ltms 141ae3b04d Merge PR #640: fleetd #639 — redact()'s comment claimed more than the code does
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m32s
CI / build (push) Failing after 2m29s
Comment text only, no behaviour change. Corrects the paragraph added in d7f94ca so it no longer
claims the continuation masking holds for any masked key "present or future". It holds only while
the masked key's own line is inside the printed hunk; diff -u's three lines of context routinely
leave it out, and a blank line inside a block scalar drops the anchor too.

Verified: suite green in a clean copy of the edited tree (16 criteria + 3 extras, exit 0), bash -n
clean, and a control on the edit itself (old claim gone, #639 reference present). Evidence in #639.
2026-10-01 19:00:56 +02:00
Dai Ha faefea14c4 fleetd #639: redact()'s comment claimed more than the code does
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Failing after 2m11s
The paragraph added in d7f94ca ended with "This needs no knowledge of the key's
name and so protects a block scalar under any masked key, present or future."
The continuation masking is real and it is an improvement, but that sentence is
too strong: the masking only holds while the masked key's own line is inside the
hunk being printed.

redact() is fed `diff -u` output, which prints three lines of context. A block
scalar's body therefore often arrives with its key line left out. With no key
line, `masked` is never set and the body prints in full, with no "<redacted>"
anywhere. A blank line inside a block scalar loses the anchor the same way: a
blank diff line measures as indent 0, so `indent > masked_indent` is false and
the mask ends early — this time directly under a "<redacted>" marker.

Both were reproduced through the real script with --dry-run, each with a
positive control run first to prove the secret's lines actually reached the
output (without that control, "the secret never entered the diff" and "it
entered and was redacted" are indistinguishable). Filed as fleetd #639, which
also records that this is latent rather than live: today's fleetd.yaml holds 5
block scalars and all 5 sit under non-secret keys.

Comment text only. No change to redact() or to any other function, and no
change to the test suite. scripts/test-config-edit.sh still passes in a clean
copy (16 criteria + 3 extras, exit 0); bash -n clean.

The reason this is worth its own commit: #635 exists because an incomplete
redactor that looks complete is worse than one that visibly does nothing. A
comment that overstates the guarantee is the same defect in prose, and the next
session to read it has no other source.
2026-10-01 19:00:29 +02:00
ltms ea6896f2ef Merge PR #636: fleetd #635 — config-edit.sh, the one auditable way to edit fleetd.yaml
CI / shell-tests (push) Failing after 14s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m5s
Closes fleetd #635.

Verified by the lead before merge:
- scripts/test-config-edit.sh run in a PRISTINE copy of d7f94ca (git archive + git init, no
  worktree, no daemon): 16 criteria + 3 extras, exit 0.
- Diff read in full. 4 files, 1325 additions, 0 deletions, no .java/pom.xml/.yaml (checked with a
  positive control, so the negative is real) — no Maven gate needed.
- Defects 1-6 were verified by the previous lead by mutation; defects 7 and 8 (criteria 15a/15b/16)
  are fixed in d7f94ca and the fix was re-measured here independently.

Known limitation, filed as #639 and NOT a regression: redact()'s continuation masking only holds
while the masked key line is itself inside the printed diff hunk. diff -u prints three lines of
context, so a block scalar's body can appear without its key, and then nothing is masked; a blank
line inside a block scalar loses the anchor the same way. Reproduced through the real script with
a positive control proving the body reached the output. Latent, not live: today's fleetd.yaml has
5 block scalars and all 5 sit under non-secret keys. The pre-fix code leaked these cases too, so
this commit is a strict improvement.

Two sentences in config-edit.sh's redact() comment and in this PR's body claim the masking holds
regardless of the key's name. That is too strong; see #639. Being corrected in a follow-up.
2026-10-01 18:58:27 +02:00
Dai Ha d7f94cafa2 fleetd #635 follow-up: redact() masks block-scalar continuations + passphrase; --set failures stop echoing the value (defects 7 and 8, criteria 15a/15b/16)
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m42s
CI / build (pull_request) Failing after 1m50s
Defect 7 (comment 17670): redact() only masked a line that itself started with a
secret-looking key, so a YAML block scalar's value leaked on the lines that
followed the key while the key line right above it printed a reassuring
"<redacted>". Fixed by tracking the masked key's own indentation and masking
every following line indented deeper than it, stopping once indentation returns
to the key's level or shallower; the diff's leading +/-/space marker is stripped
before indentation is measured, per the comment's own pitfall. "passphrase" is
now also in the key-name backstop.

Defect 8 (comment 17673): apply_set_pairs echoed the operator's full
"path=value" input, unredacted, in both of its yq-failure die messages — a
failing --set with a secret-looking value printed that value right back. Fixed
to print only the path; deliberately not routed through redact, which would
pass a non-"key: value"-shaped string straight through.

Adds acceptance criteria 15a (block-scalar continuation), 15b (passphrase key),
and 16 (failing --set never echoes its value) to scripts/test-config-edit.sh,
each with a positive control proving the relevant line really was in the
printed output before asserting the secret is absent. All three confirmed RED
against the pre-fix code and GREEN after, in isolation, before being folded
into the full suite (16 criteria + 3 extras, exit 0).

Also updates PR #636's description per comment 17671: the redact() sentence now
names the continuation-masking rule and says plainly that the key-name list is
a backstop, never a complete list.
2026-10-01 18:47:00 +02:00
Dai Ha 096f08c866 fleetd #635 follow-up: fix the stale --restore not-found message (defect 6, criterion 14)
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 2m48s
The --restore "no backup found" message still printed the old beside-the-config glob
(${CONFIG}.bak.*) even though newest_backup had already moved to searching the managed
.config-backups/ directory. The message was left behind when the search moved — the search
itself was already correct (ticket comment 17664). Fix is reporting-only: the message now
names the directory actually searched (via backup_dir_for), and separately says that a
backup written the old way, directly beside the config, is not searched any more, with the
one-line cp to recover one by hand. No search fallback was added — reading backups from
outside the managed directory stays unsupported, as instructed.

Acceptance criterion 14 proves both directions: the not-found message names the real
directory (confirmed red on the pre-fix code, green after), and a restore with a real backup
present in .config-backups/ still succeeds (confirmed this catches an "always not-found"
regression that direction 1 alone would miss).

All 14 criteria plus 3 extras pass in scripts/test-config-edit.sh.
2026-10-01 18:25:08 +02:00
Dai Ha 4eb720029c fleetd #635 follow-up: refuse empty --set values, gitignore backups, preserve file mode
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m43s
CI / build (pull_request) Failing after 1m52s
Five fixes against PR #636, all verified by the lead's own review and reproduced here:

1. --set .a.b= (a forgotten value) is now refused outright instead of silently nulling the
   field — a null numeric config value falls back to its default rather than erroring, which
   widens capacity silently instead of failing loudly. A deliberate clear gets its own spelling,
   --set .a.b=null, which writes a literal YAML null via yq, never through strenv(). (criteria
   9, 10)

2. Backups move from beside fleetd.yaml to a dedicated fleetd/.config-backups/ directory,
   gitignored at the repo root (so it also covers scripts/test-config-edit.sh's own throwaway
   fixtures) and in fleetd/.gitignore, plus a fleetd.yaml.bak.* glob backstop for any stray
   backup written the old way. A backup of a file that must never be committed inherits that
   requirement. (criterion 11)

3. The live config's file mode now survives both an edit and a restore. mv from a mktemp
   candidate used to carry mktemp's 0600 onto the live path forever, and cp onto an existing
   file keeps the destination's mode, so a restore did not undo it either. (criterion 12)

4. A global CAND + single EXIT/INT/TERM trap prevents an uninstalled .config-edit.XXXXXX
   candidate from leaking if the script is interrupted mid-run. No acceptance criterion is
   gated on this — a reproducible leak could not be made to happen on demand — but it is cheap
   and obviously right.

5. Acceptance criterion 7's redaction check gained a positive control: it now asserts the
   output actually CONTAINS the redaction marker and the changed key, not only that it lacks
   the secret. The prior two assertions were negative-only and passed just as happily when the
   diff was never printed at all — confirmed by reproducing the lead's own mutation (deleting
   the redacted diff print on the edit path) and watching it survive the old test and get
   caught by the new one. (criterion 13)

All 13 acceptance criteria plus 3 extras pass in scripts/test-config-edit.sh. Criteria 9, 10,
11, 12 and 13 were each proven non-vacuous: criteria 9/10 by mutating the test's own expected
value and watching it fail by name, then reverting; criteria 11/12/13 by reverting or mutating
the corresponding fix in config-edit.sh and watching the matching criterion fail by name, then
restoring the fix and re-confirming a clean pass.
2026-10-01 18:07:12 +02:00
Dai Ha 0db6d31dc2 fleetd #635: add scripts/config-edit.sh, the one auditable way to edit fleetd.yaml
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 2m12s
CI / build (pull_request) Failing after 2m48s
Backs up, builds a candidate off the live file, parse-checks it with yq before
install, installs atomically, then reads the daemon's own ConfigRef reload
verdict back out of fleetd.out (marked from before the edit, so a stale line
can never be mistaken for this edit's result). Four exit codes: 0 clean, 3
needs a restart, 4 refused (backup restored), 5 cannot tell (nothing
restored, printed --restore command). Every diff is redacted.

scripts/test-config-edit.sh drives it end to end against fixtures in a
throwaway temp dir, with no daemon involved.
2026-10-01 17:29:04 +02:00
Dai Ha 158a2a84b5 Merge PR #632: fleetd #612 ranks 6+7 + #630 — behavioural pins for lifecycle wirings
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m13s
CI / build (push) Failing after 2m10s
Pins three call sites the assembly owns and nothing observed:
- healthFailTarget (FleetdAssembly:429) — inert, a dead member's waiting ticket sits
  PENDING for the full 30-minute async timeout instead of failing immediately.
- releaseCleanup (:447) — inert, every teardown leaks three things: a stuck rendezvous
  waiter, an unreleased reply-inbox consumer, and a stale lead binding.
- requireOperatorConfirm (:402/:409, fleetd #630) — dropping the 14th constructor
  argument selects #621's 13-arg overload, which hardcodes true, silently reverting
  the operator's fix on a host that set requireOperatorConfirm: false. Pinned in both
  directions, plus an assertion that the two notice strings differ, so no constant
  satisfies both.

Verified by the lead beyond the worker's proof: its releaseCleanup mutation killed all
three cleanups at once and so proved only the first assertion had teeth. Starving them
one at a time — abandon kept, release starved; then abandon and release kept, forget
starved — each fails its own named assertion. All three leaks are pinned independently.

A review pass found these three tests assembled real schedulers and never tore them
down, by any route: no close(), no shutdownHook, no @AfterEach, no finally. Surefire
runs one JVM fork for the whole suite, so those loops outlived their tests. Fixed by
capturing the hook and running it in a finally, with an assertion on FakeHerdr.closed
so the teardown itself is pinned rather than assumed.

MERGE RESOLUTION BY THE LEAD, the same collision as #633. These 3 test files each add
a ResourcePorts fake, and #633 landed first making herdrPollWait() abstract with no
default. Git reported a clean merge that did not compile. Added the override to all 3,
matching the established convention for an always-healthy fake — a Runnable that
throws, verified first that none of the three uses healthy(false), so the tripwire can
only fire if the test's herdr behaviour changes.

Full suite on the resolved merge: 1892 tests, 0 failures (1889 + this branch's 3).
2026-10-01 16:52:39 +02:00
Dai Ha 41ebc9cf69 Merge PR #633: fleetd #629 + #625 — the two ResourcePorts seams
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 2m57s
#629: the herdr boot wait went through ports.nanoClock() but hardcoded
Fleetd::sleepHerdrPoll, so a test assembling against an unhealthy lead herdr burned
30 real seconds whatever clock it injected. The poll sleep now goes through
ResourcePorts.herdrPollWait(). FleetdAssemblyFleetAppTest's lead-down test drops from
30.276s to 0.062s. It also gains @Timeout(10, SEPARATE_THREAD) at class level: the
pin's failure mode is otherwise an infinite hang, because the test's fake clock only
advances when herdrPollWait() is called. SAME_THREAD cannot interrupt a real
Thread.sleep, so the thread mode is load-bearing, not decoration.

#625: guard.assertPrimaryClean(System.getenv()) could be deleted with a fully green
suite — the check behind charter invariant 1, which keeps the primary on the
operator's subscription. main() now delegates to main(String[], ResourcePorts) and
the guard reads ports.environment(), so a test can taint the environment without
touching the real process env. The guard still runs before cfg.validateAll() and
before any socket, broker or HTTP work; two tests with different fixtures pin that
ordering as two independently falsifiable claims, not one.

MERGE RESOLUTION BY THE LEAD. ResourcePorts.herdrPollWait() is abstract with no
default, by design (#629 keeps ResourcePorts free of a none() default). This branch
patched the 8 implementations that existed when it forked. PRs #631 and #634 merged
ahead of it and added 3 more fakes, so git reported a clean 14-file merge that did
not compile:

  FleetdAssemblyAmqpOpenersTest.RecordingPorts is not abstract and does not
  override abstract method herdrPollWait() in dev.ltms.fleet.ResourcePorts

(plus the two ControllableResourcePorts in #634's tests). I added the override to
those 3, matching this branch's own convention for an always-healthy fake: return a
Runnable that throws, so if one of those assemblies ever does start polling herdr it
fails loudly instead of sleeping quietly. All 13 implementations now carry it.

Full suite on the resolved merge: 1889 tests, 0 failures (1886 + this branch's 3).
2026-10-01 16:50:10 +02:00
Dai Ha 619769bb52 fleetd #632: tear down the three assembly behavioural tests' background loops
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 2m39s
Each of FleetdAssemblyHealthFailTargetBehaviouralTest,
FleetdAssemblyReleaseCleanupBehaviouralTest and
FleetdAssemblyRequireOperatorConfirmBehaviouralTest called
FleetdAssembly.assembleAndStart without ever tearing it down: no close(),
no shutdownHook, no @AfterEach, no finally. Surefire runs the whole suite
in one JVM fork, so every scheduler/loop these tests started kept running
for the rest of the suite.

Capture the shutdown hook in each test's fake ResourcePorts (the existing
pattern from FleetdAssemblyLifecycleTest et al.) and run it in a finally
block, on the failure path too. RequireOperatorConfirmBehaviouralTest
assembles twice in one method, so assembleHeartbeat now returns both the
loop and its ResourcePorts so each assembly gets its own teardown.

Proof the teardown actually runs: assert ports.herdr.closed after running
the hook (FakeHerdr.close() only flips that flag from inside the real
close chain). Verified the assertion is load-bearing by temporarily
removing one shutdownHook.run() call and confirming the test then fails.

No assertion, test name or reflection changed. Diff is test-only.
2026-10-01 16:49:14 +02:00
Dai Ha 180c953c42 Merge PR #634: fleetd #612 ranks 1+2 — behavioural pins for the exhaustion wirings
CI / shell-tests (push) Failing after 8s
CI / build (push) Failing after 2m20s
CI / contract (push) Successful in 2m22s
Pins all four exhaustion call sites in FleetdAssembly: liveExhaustedPatterns (:299),
exhaustedPatternLookup (:300), publishExhaustionSink (:327) and the independent
OpenCode forwardingExhaustionSink (:164). Rank 1 is the worst consequence in the
#612 sweep — an inert lookup hands a genuine usage-limit refusal back to a waiting
caller as real completed work instead of BACKEND_EXHAUSTED.

Two separate tests, so rank 2's OpenCode half is pinned independently: mutating
:164 fails only the forwarding test, which is the independence the ticket asserts.

Verified by the lead beyond the worker's own proof: wiring publishExhaustionSink to
a throwaway BackendQuarantine — every symbol kept at the call site, only the
collaborator identity changed — is caught by both tests. So these pins survive
mis-wiring, not just deletion.

Test-only; no production change. Tears down each assembly via the captured
shutdown hook in a finally.
2026-10-01 16:45:20 +02:00
Dai Ha 4af919af0c fleetd #629 follow-up: bound FleetdAssemblyFleetAppTest with a SEPARATE_THREAD @Timeout
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 2m25s
The mutation cycle for #629 proved deleting the fix (reverting FleetdAssembly's awaitHerdr
call back to a hardcoded Fleetd::sleepHerdrPoll) doesn't just make a test fail — it hangs
forever, because the test's fake nanoClock() only advances when ports.herdrPollWait() is
actually called. SAME_THREAD @Timeout (JUnit's default) can't catch that: it only measures
elapsed time after the test method returns on its own, which never happens here. A class-level
@Timeout(10s, SEPARATE_THREAD) does, since it runs the test on its own thread and interrupts it
on timeout. Verified by re-running the same mutation: the suite now fails fast with a named
TimeoutException instead of hanging indefinitely.
2026-10-01 16:41:40 +02:00
Dai Ha 09c37061c1 Merge PR #631: fleetd #612 ranks 3+8 — behavioural pins for the AMQP assembly openers
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 59s
CI / build (push) Failing after 3m14s
Pins FleetdAssembly's replyInboxOpener (:357) and leadMailboxOpener (:363) so an
inert opener can no longer pass a green suite. Both directions covered: a durable
opener must leave the runtime owning the exact object it returned, and a failing
one must leave the in-memory fallback with coordination off — with the startup
report agreeing with the real state in each case.

Verified by the lead beyond the worker's own proof: a mutation that still CALLS
ports.replyInboxOpener() and discards the result is caught by assertSame, which is
rank 3's live defect (opener called, result thrown away, log still printing
'reply inbox: AMQP broker (durable)').

Test-only; no production change.
2026-10-01 16:38:16 +02:00
Dai Ha 2be287ea03 fleetd #612 step 4 (ranks 1&2): behavioural assembly tests for the exhaustion wirings
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 3m5s
Replaces nothing (no existing test covered these call sites through the real
assembly); adds two new FleetdAssembly-driven tests, matching Unit A's pattern
of inspecting FleetdRuntime's real, assembled objects rather than a copy.

- FleetdExhaustedPatternAssemblyTest pins the liveExhaustedPatterns/
  exhaustedPatterns call sites (rank 1 — "the worst consequence in the whole
  sweep": a genuine usage-limit refusal handed back as real completed work)
  together with publishExhaustionSink (rank 2, non-OpenCode half): drives a
  real CompletionResolver through a pane scrape matching a configured
  exhaustedPattern and asserts BACKEND_EXHAUSTED classification plus a real
  BackendQuarantine credential quarantine.

- FleetdOpenCodeExhaustionForwardingAssemblyTest pins forwardingExhaustionSink
  (rank 2, OpenCode half — independent of publishExhaustionSink per the
  ticket) by reflectively reaching the real, assembled OpenCodeLauncher's
  exhaustionSink field (SessionManager.launcher -> CompositePeerLauncher.
  byProfile -> OpenCodeLauncher.exhaustionSink) and proving it forwards into
  the same production BackendQuarantine.

All four call sites (FleetdAssembly.java:164,299,300,327) were each put
through grep-anchor -> line-anchored sed mutation -> mvn -o compile -> full
mvn -o test (named test RED) -> restore -> full mvn -o test (1885/0 GREEN).
Mutating line 327 alone also fails the OpenCode test, confirming the
documented construction-order dependency (forwardingExhaustionSink reads
exhaustionSinkRef, which publishExhaustionSink sets) without weakening either
site's independent pin.
2026-10-01 16:38:00 +02:00
Dai Ha c3b0406826 fleetd #612: isolate fallback AMQP reports
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Failing after 2m20s
2026-10-01 16:33:45 +02:00
Dai Ha e76fa1660b fleetd #629/#625: inject the herdr poll wait and the startup env read through ResourcePorts
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Failing after 2m42s
#629: ResourcePorts gains herdrPollWait() (SystemResourcePorts: Fleetd::sleepHerdrPoll).
FleetdAssembly's awaitHerdr call now takes its poll wait from ports instead of a hardcoded
Thread.sleep, so a test's fake clock can actually reach the deadline without burning real
wall-clock time. FleetdAssemblyFleetAppTest's down-lead case drops from ~30s to well under 1s.
All other ResourcePorts fakes get a trivial "must not be called" override since their herdr is
always healthy and never polls.

#625: Fleetd.main(String[]) now delegates to a new package-private
main(String[], ResourcePorts) overload that reads the startup environment via
ports.environment() instead of System.getenv(), reusing the same ports instance for
FleetdAssembly.assembleAndStart. The guard call stays at its original point, before
cfg.validateAll() and before any assembly/socket/broker/HTTP work. New
FleetdSubscriptionGuardOrderingTest drives the real main() with a tainted fake environment and
pins both presence and ordering against validateAll() and against assembly's first port call,
plus a clean-env control proving the guard only blocks on an actual taint.
2026-10-01 16:33:15 +02:00
Dai Ha 9dca376604 fleetd #612 ranks 6/7 + #630: behavioural pins for healthFailTarget, releaseCleanup, requireOperatorConfirm
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 2m58s
Three FleetdAssembly.java call sites had no test that drives the real
assembly and inspects what it actually built, so each could be swapped
for an inert variant and the suite would stay green:

- FleetdAssembly.java:429 (healthFailTarget, #612 rank 7): a no-op
  BiConsumer leaves a dead member's waiting ticket PENDING for the full
  30-minute async timeout instead of failing it immediately.
- FleetdAssembly.java:447 (releaseCleanup, #612 rank 6): a no-op
  onRelease listener leaks a stuck rendezvous waiter, an unreleased
  reply-inbox consumer, and a stale lead binding on every teardown.
- FleetdAssembly.java:402/:409 (requireOperatorConfirm, #630): dropping
  the 14th LeadHeartbeatLoop constructor argument selects the
  13-argument overload, which hardcodes true (fleetd #621) regardless
  of leadRollover.requireOperatorConfirm — silently reverting an
  operator's own config choice.

Each existing "wiring" test for these sites (FleetdHealthFailTargetWiringTest,
FleetdReleaseCleanupWiringTest) calls the Fleetd.* factory method directly
and never drives FleetdAssembly.assembleAndStart, so none of them can see
whether the real call site still passes the real, assembled collaborators.

The three new tests here assemble the real daemon via
FleetdAssembly.assembleAndStart, reach the real wired object (reflection,
same technique StatusPollerResilienceTest already uses — the fields and
the requireOperatorConfirm overload resolution point are package-private),
and assert the real behavioural effect against the real collaborators
FleetdRuntime exposes. Each is mutation-tested: a line-anchored sed to the
inert form compiles clean and turns the new test RED; restoring the
original line turns it GREEN again. requireOperatorConfirm is proven in
both directions (false and true produce genuinely different notice text).

Full suite: 1886 tests, 0 failures, 0 errors (baseline 1883 + 3 new).
2026-10-01 16:30:23 +02:00
Dai Ha 91be5079d6 fleetd #612: pin AMQP assembly openers
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m44s
CI / build (pull_request) Failing after 3m41s
2026-10-01 16:26:02 +02:00
ltms 26f198675c Merge pull request 'fleetd #612 Unit A: extract main's boot composition into FleetdAssembly/FleetdRuntime' (#620) from worker/fleetd-612-unita-87807e-1 into main
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m30s
2026-09-22 07:43:39 +02:00
lead 72f46d7c0c fleetd #612: merge main into Unit A, carrying #621's requireOperatorConfirm into the assembly
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m59s
Resolves the one conflict in Fleetd.java. main's side is the inline boot block
that Unit A had already moved into FleetdAssembly.assembleAndStart, so the
resolution keeps Unit A's single assembly call.

That resolution is not purely mechanical. #622 (fleetd #621) landed on main
AFTER Unit A forked, and it added

    boolean requireOperatorConfirm = cfg.leadRollover() == null
            || cfg.leadRollover().requireOperatorConfirm();

plus a 14th argument to the LeadHeartbeatLoop constructor, inside the very block
Unit A moved. Taking Unit A's side alone would have dropped both and silently
reverted the operator's #621 fix: the 13-argument overload still exists and
delegates with `true`, so the daemon would go back to telling every lead to ask
the operator before a context roll. Both are carried into FleetdAssembly here.

Measured: with the carried line removed, the full suite is
`Tests run: 1883, Failures: 0, Errors: 0` — nothing pins it. That is the fleetd
#612 defect shape applied to #621's own wiring, and it is filed separately
rather than fixed here, because this commit is a merge resolution and must not
also introduce new tests.

Full suite on this resolved tree: Tests run: 1883, Failures: 0, Errors: 0.
2026-09-22 12:43:22 +07:00
ltms 640f4d5f23 Merge pull request 'fleetd #612 B3: behavioural replacements for lead-seat, quarantine, lead-rollover guards' (#628) from worker/612-b3-mcpwirings-da2b58-3 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 12s
CI / contract (pull_request) Successful in 1m33s
CI / build (pull_request) Failing after 2m25s
2026-09-22 07:36:05 +02:00
Dai Ha 6edeb70bc4 fleetd #612 B3 correction: distinguish lead vs member herdr in FleetdLeadRolloverAssemblyTest
Ticket comment 17553 on fleetd #612 found that the test's single shared
FakeHerdr made router.leadAgents() and router.memberAgents() collapse to
the identical client (FleetdAssembly.java:140-142's no-distinct-socket
fallback), so a mutation swapping leadAgents() for memberAgents() at the
FleetdAssembly.java:408 call site was invisible to this test even though
the two are genuinely different daemons in production.

Configure two distinct herdr sockets and two distinct FakeHerdr instances
(the same TwoHerdrResourcePorts shape B2's FleetdAssemblyConnectionIdentityTest
uses) and assert the roll's /clear + bootstrap sends land on the LEAD fake
and never on the MEMBER one.

Proven red against the router.memberAgents() mutation, reverted, touched,
and re-run green — both outputs recorded in the PR.
2026-09-22 12:31:39 +07:00
Dai Ha bc49d87cb8 Merge remote-tracking branch 'origin/worker/fleetd-612-unita-87807e-1' into worker/612-b3-mcpwirings-da2b58-3 2026-09-22 12:27:47 +07:00
ltms 14b169c410 Merge pull request 'fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair' (#626) from worker/612-b2-cb185-176d3a-2 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 2m6s
2026-09-22 07:19:28 +02:00
Dai Ha 2b52324d9a fleetd #612 step 2 unit B3: behavioural replacements for lead-seat, quarantine
and lead-rollover source-text guards

Replaces three FleetdAssembly.java source-text guards (each scraped
Fleetd.java for a call site that fleetd #612 Unit A moved into
FleetdAssembly.java) with tests that drive the real assembled objects
through FleetdAssembly.assembleAndStart(...) -> FleetdRuntime.mcp(),
per the step-2 B-unit split (issue #612 comment 17513).

- Deleted FleetdLeadSeatWiringTest (fleetd #176): pinned that
  FleetMcp's LeadSeatSource construction still wires
  Fleetd.leadSeatLookup(...) by scraping the constructor call's text.
  Replaced by FleetdLeadSeatAssemblyTest, which seeds one FakeHerdr
  tab labelled to match a configured fleet.leaders.opus.tab and
  asserts the REAL assembled LeadSeatSource (via
  runtime.mcp().leadSeatSource()) reports the live lead's seat against
  its own subscription profile -- 1, not the 0 LeadSeatSource.none()
  (the inert stand-in) could ever report.

- Deleted FleetdBackendQuarantineWiringTest (fleetd #466): pinned that
  the escalating BackendQuarantine.withEscalation(...) text was
  present and the flat two-argument constructor's text was absent.
  Replaced by FleetdBackendQuarantineAssemblyTest, which quarantines
  the same credential twice through the REAL assembled
  BackendQuarantine (via runtime.mcp().quarantineSource().quarantine())
  at controlled fake-clock offsets and asserts the second cooldown
  doubles (200s vs 100s) -- the one behavioural difference escalation
  and the flat constructor actually produce.

- Deleted FleetdLeadRolloverWiringTest (fleetd #480), all three
  methods: unrelatedAnchorStillPresent was a scaffold anchor with no
  independent claim, needing no replacement.
  mainStillCallsTheLeadRolloverFactory pinned the leadRollover
  assignment's call-site text. factoryGatesOnConfigPresence pinned
  that an absent leadRollover: config yields no LeadRollover.
  Replaced by FleetdLeadRolloverAssemblyTest's two tests:
  assembledLeadRolloverRunsTheRealClearAndBootstrapSequence drives the
  REAL assembled LeadRollover (via runtime.mcp().leadRollover())
  through open()/confirm() end to end and asserts /clear then
  bootstrapText were actually sent through the real herdr router,
  reaching ROLLED. absentLeadRolloverConfigMeansNoRolloverIsBuilt
  calls Fleetd.leadRollover(...) directly with no leadRollover: block
  and asserts null -- this claim was found uncovered elsewhere
  (LeadRolloverTest's only related assertion is vacuous, assertNull
  (null), and never calls the real factory).

Each of the three FleetdAssembly.java call sites (quarantine
line 179-180, leadRollover line 408, lead seats line 479) was mutated
to its named inert variant, run against ONLY its new test (RED),
reverted, touch'd (Maven mtime trap) and re-run (GREEN) -- six proven
runs, pasted in the PR body.

FleetMcp.java: adds three accessors (quarantineSource(),
leadSeatSource(), leadRollover()) alongside the existing
registeredTools() -- but public, not package-private, and this is a
deliberate deviation from that precedent, not an oversight: these new
assembly tests cannot live in package dev.ltms.fleet.mcp the way
registeredTools()'s callers do, because they also build the
ResourcePorts FleetdAssembly.assembleAndStart(...) needs, and
ResourcePorts' methods return Fleetd-nested types visible only from
package dev.ltms.fleet. Package-private would compile but be
unreachable from there.

Full mvn -o test in fleetd/: Tests run: 1879, Failures: 6 (down from
the branch baseline's 1880/9 by exactly the 3 guards this unit
deletes) -- the remaining 6 are FleetdCompletionResolverWiringTest (4)
and FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest (1 each), all out of this unit's scope
(B1/B2).
2026-09-22 12:18:48 +07:00
ltms b6147a39f6 Merge pull request 'fleetd #612 Unit B1: CompletionResolver behavioural test (replaces source-text guard)' (#627) from worker/612-b1-completion-457459-1 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 1m37s
2026-09-22 07:14:25 +02:00
Dai Ha dbfc34cb6d fleetd #612 B2 fixup: cover the symmetric FleetApp daemon-drop mutation
Per ticket comment 17525: my second independent mutation on
FleetdAssembly.java:518 survived — new FleetApp(memberHerdr, memberHerdr, ...)
(dropping the LEAD client instead of the member one) left both existing
FleetdAssemblyFleetAppTest cases green. That is the symmetric form of the
CB-185 defect (/healthz green while the LEAD daemon is down), and the guard
this PR deletes would have caught it: its positive assertion required the
exact pair "new FleetApp(herdr, memberHerdr, workers,", which does not
survive either daemon being dropped.

Adds healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp, symmetric
to the existing member-down case.

Proven red today: reverting FleetdAssembly.java:518 to
"new FleetApp(memberHerdr, memberHerdr, workers, ..." and running only
FleetdAssemblyFleetAppTest gives Tests run: 3, Failures: 1 — the new case
fails ("expected: <503> but was: <200>", body has no "member" key); the other
two cases stay green. Reverted, git diff --stat empty, file touched, re-ran:
Tests run: 3, Failures: 0.

Full mvn -o test: Tests run: 1884, Failures: 7 (same 7 B1/B3-scope failures
as before this fixup; 1884 = 1883 + 1 new case).
2026-09-22 12:11:28 +07:00
Dai Ha a20cb96730 fleetd #612 Unit B1: replace CompletionResolver source-text guards with behavioural tests
FleetdCompletionResolverWiringTest read Fleetd.java's literal source text and
asserted the CompletionResolver construction call still named the right
arguments — proof of spelling, not behaviour. Unit A (FleetdAssembly) moved
that call site out of Fleetd.main, breaking all four of its tests on a
harmless relocation.

FleetdCompletionResolverAssemblyTest replaces it, driving the real
FleetdAssembly.assembleAndStart(...) and reading FleetdRuntime.completion() —
the exact CompletionResolver instance production uses — through its public
onDelivered/resolveBeforePostAction API, with a controllable ResourcePorts
nanoClock in place of real sleeps.

Deleted test -> what it pinned -> replacement:
- worktreeBranchLookupIsStillPassedAtTheCallSite (8th constructor arg) ->
  assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport: a
  real git-worktree-provisioned MemberSession's branch must appear in a
  noReportMessage fallback (fleetd #241).
- backendErrorArgumentsAreStillNamedAtTheCallSite,
  backendErrorPatternsComesFromTheFactory, backendErrorSinkComesFromTheFactory
  (5th/6th args) -> assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern:
  a configured errorPattern the built-in fallback never matches must classify
  as FAILED (not COMPLETION), transition the session to BACKEND_ERROR, and
  cool the credential off after two distinct targets within the window
  (fleetd #201 Unit 5).

Both behaviours were proven red today by mutating FleetdAssembly.java's real
call site to its inert variant (_ -> null; BackendErrorPatternLookup.legacy();
BackendErrorSink.none()), confirming the new test failed with the expected
message, then reverting (touching the file to defeat Maven's stale-mtime
skip) and confirming green again. FleetdAssembly.java itself is unchanged in
this commit.

Full suite: Tests run: 1878, Failures: 5 (down from the baseline 9 — the
remaining 5 are the other in-flight workers' own guard files:
FleetdBackendQuarantineWiringTest, FleetdConnectionIdentityConstructionTest,
FleetdFleetAppConstructionTest, FleetdLeadRolloverWiringTest,
FleetdLeadSeatWiringTest), Errors: 0, Skipped: 0.
2026-09-22 12:10:38 +07:00
Dai Ha cda1a6a917 fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair
Deletes the two source-text guards fleetd #612 Unit A broke by moving their
scraped call sites from Fleetd.java into FleetdAssembly.java, replacing each
with a behavioural test that drives the real assembled graph instead.

- FleetdConnectionIdentityConstructionTest pinned that Fleetd.java contained
  "new PaneLocator(herdr, memberHerdr)". Replaced by
  FleetdAssemblyConnectionIdentityTest, which drives the real PaneLocator a
  real FleetdAssembly.assembleAndStart(...) built (reached via
  runtime.mcp().identity().panes(), never a copy) with two distinct FakeHerdr
  daemons, and proves it finds a pane that exists on only one of them —
  first the lead-only case (the CB-185 bug: a lead's own connection going
  unresolvable), then the member-only case, plus a no-match control.

- FleetdFleetAppConstructionTest pinned that Fleetd.java contained
  "new FleetApp(herdr, memberHerdr, workers,". Replaced by
  FleetdAssemblyFleetAppTest, which binds the real Javalin app
  FleetdAssembly built (runtime.app()) to a real ephemeral port and proves
  GET /healthz goes 503 when only the member daemon is down — the consequence
  named in the deleted test's javadoc (a down member daemon invisible behind
  a healthy lead). The GET /sessions merge half of that javadoc could not be
  driven the same way: it requires Authz.Action.READ, which needs a real
  positive pid from the real assembly's hardcoded LsofPeerPidLookup, and an
  in-process test's HTTP client and server share one JVM pid so that pid is
  always -1 (fleetd #317's fail-closed rule then refuses the request before
  the route, and its merge, is ever reached) — documented in the new test's
  javadoc; FleetAppTwoDaemonTest remains the full behavioural proof that
  FleetApp itself merges /sessions correctly given two clients.

Both replacements were proven red today: reverting FleetdAssembly.java:443-444
to "new PaneLocator(memberHerdr)" turns the identity test red (expected
"term_a", got null); reverting :518 to "new FleetApp(herdr, herdr, ..." turns
the app test red (expected 503, got 200 with no "member" key). Both mutations
reverted, tree confirmed clean, and the mutated file touched afterward so
Maven does not skip recompiling a stale .class.

Adds two small production accessors needed to reach the real objects rather
than a copy, since FleetdRuntime may not gain a field (three workers touch
that file): ConnectionIdentity#panes() exposes the PaneLocator it resolves
against, and FleetMcp#identity() exposes the ConnectionIdentity it was built
with (now also stored as a field).

Baseline on 608e449: 1880 tests, 9 failures (the known source-text set).
After this change: 1883 tests, 7 failures — the remaining #248/#176/#466/#480
guards, which are B1's and B3's scope, not this one's.
2026-09-22 12:03:38 +07:00
ltms 608e4496be Merge pull request 'fleetd #612 A-gaps: coordinator path + reportRoleFallbackGaps inside the assembly boundary' (#624) from worker/612-agaps-73a926-2 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m34s
CI / build (pull_request) Failing after 1m42s
2026-09-22 06:41:42 +02:00
ltms fa61dc587c Merge pull request '#608 replace MessageService timing sleeps' (#623) from worker/608-sleeps-3a64ff-3 into main
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 3h14m55s
2026-09-22 06:38:29 +02:00
Dai Ha 526e3b7459 fleetd #612 A-gaps: exercise the coordinator path and pull reportRoleFallbackGaps inside the assembly boundary
Gap 1: generalise Fleetd.LeadMailboxOpener (and FleetdRuntime's field) from the concrete
LeadMailbox to a new closeable LeadChannelHandle (LeadChannel + AutoCloseable), so a test can
fake the configured-coordinator path without a real broker. FleetdAssemblyCoordinatorLifecycleTest
drives FleetdAssembly.assembleAndStart with a coordinator: block and a fake channel, proving the
assembly builds it and shutdown closes it.

Gap 2: move reportRoleFallbackGaps(cfg) and assertChartersNameOnlyRegisteredTools(cfg) out of
Fleetd.main and into FleetdAssembly.assembleAndStart, immediately after cfg.validateAll(), so
both run inside the tested assembly boundary before any I/O. FleetdAssemblyRoleFallbackBoundaryTest
pins the moved call site behaviourally (captured log output), not by reading source text.

Both new tests were verified red under a targeted mutation (a non-closing/null coordinator for
gap 1; deleting the moved call for gap 2) and restored.
2026-09-22 11:38:19 +07:00
Dai Ha b6b006c651 #608 replace MessageService timing sleeps
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Failing after 2m27s
2026-09-22 11:33:48 +07:00
ltms 63eec8a0da Merge pull request 'fleetd #621: make the context-roll notice obey requireOperatorConfirm' (#622) from worker/621-b4520b-1 into main
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m11s
CI / build (push) Failing after 1m52s
2026-09-22 06:31:02 +02:00
Dai Ha cbe872b538 fleetd #621: make the context-roll notice obey requireOperatorConfirm
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m0s
CI / build (pull_request) Failing after 1m53s
contextNotice hardcoded 'ask the operator' and 'Only the operator can
approve the roll', so setting leadRollover.requireOperatorConfirm to
false stopped the daemon refusing the roll but never stopped the lead
being told to ask. Thread the effective config value into
contextNotice: when true the text stays byte-identical, when false it
tells the lead to confirm on its own judgement against the three
handover-file checks instead.

LeadRollover.confirm's own enforcement is untouched — this is the
message only.
2026-09-22 11:27:17 +07:00
Dai Ha 7d9a807243 handover skill: the rollover bootstrap is proven, and requireOperatorConfirm is per-host
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 2m33s
Two bullets in the handover skill were telling every outgoing lead something
that is no longer true.

1. The skill said the bootstrap prompt "has never yet landed, and the fix is
   unproven (fleetd #489)", and told the lead to warn the operator it may fail.
   Measured today from fleetd/fleetd.out:

     grep -c "lead-rollover: rolled" -> 4
     grep -c "lead-rollover:"        -> 16   (positive control)
     grep -c "Unknown command"       -> 0

   Three of the four rolls ran on 2026-09-22 (10:01:43, 10:38:28, 11:15:47).
   Each cleared the old lead and bootstrapped a fresh one against the handover
   file. The old "Unknown command: /clearFresh" failure does not appear at all.
   The paragraph now carries the measured result and the three re-measure
   commands, including the control line, because a broken grep pattern returns
   a clean 0 that reads like good news.

   It also records that the "/clear was never observed as WORKING ... releasing
   rather than wedging the roll" WARN accompanies every successful roll. That
   is the safe branch, not a failure, and it was being misread as one.

2. The skill said requireOperatorConfirm "defaults to true and this is the only
   thing standing between a judgement call and a wiped session", which reads as
   if asking is always required. The default is still true
   (FleetConfig.java:1426), but this host set it to false on 2026-09-22 on the
   operator's explicit grant. The bullet now says to read the live value rather
   than assume, and notes the key is deferred, not hot.

   It also warns that until fleetd #621 merges, LeadHeartbeatLoop.contextNotice()
   still hardcodes "ask the operator" and takes no config, so the nudge text and
   the config disagree. Trust the config. That warning names the ticket that
   removes it.

Documentation only. No code or test changes.
2026-09-22 11:24:03 +07:00
ltms 8915e40c7d Merge pull request 'fleetd #618: state the measured auto-compact precedence' (#619) from worker/618-b83894-2 into main
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 2m12s
2026-09-22 05:53:35 +02:00
Dai Ha 3f7bc3815e fleetd #612 Unit A: extract Fleetd.main's boot composition into FleetdAssembly/FleetdRuntime
CI / shell-tests (pull_request) Failing after 12s
CI / build (pull_request) Failing after 1m29s
CI / contract (pull_request) Successful in 1m28s
Fleetd.main kept config loading, startup reports and validation. Everything from
the herdr socket connect onward moved verbatim, same order, into
FleetdAssembly.assembleAndStart(AssemblyInputs, ResourcePorts), which returns a
FleetdRuntime owning the real objects (package-private accessors, never a copy)
and their single ordered close(). ResourcePorts/SystemResourcePorts abstract every
boot-time side effect (env, herdr connect, broker openers, clocks, schedulers,
shutdown-hook registration, HTTP start) with no inert production variant, per the
architect proposal on the ticket.

sleepHerdrPoll widened from private to package-private so FleetdAssembly can pass
a method reference to it; no other signature changed.

FleetdAssemblyLifecycleTest drives the real assembly with FakeHerdr, a temp
FleetConfig and a fake ResourcePorts recording a start/close ledger, asserting it
against the order recorded from the pre-move main() and shutdown hook, and proving
every resource the ledger can observe (herdr client, three schedulers, the AMQP
reply inbox) closes via FleetdRuntime.close(). FakeHerdr gained a closed flag for
this.
2026-09-22 10:52:35 +07:00
Dai Ha 6cb31a10e4 fleetd #618: fix the third stale spot the brief missed
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Failing after 2m6s
The method-level javadoc on FleetConfig.warnConflictingAutoCompactWindows
(above the log.warn call) still claimed the autoCompactWindow vs
CLAUDE_CODE_AUTO_COMPACT_WINDOW precedence was 'intentionally not
asserted' and cited fleetd.yaml's now-corrected comment as evidence the
question was open. Replace it with the measured answer from #618: the
env var wins, so autoCompactWindow is inert on a profile that sets both.
Kept the WARN-not-throw rationale paragraph above it untouched (#601)
and kept the ClaudeCodeArguments cross-reference, which now points to an
agreeing claim instead of a contradicting one. No behaviour change.
2026-09-22 10:50:59 +07:00
Dai Ha 8368a274a0 fleetd #618: state the measured auto-compact precedence, not 'unverified'
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 1m57s
ClaudeCodeArguments.withAutoCompactWindow's javadoc and FleetConfig's
warnConflictingAutoCompactWindows WARN text both used to say the
precedence between --autocompact and CLAUDE_CODE_AUTO_COMPACT_WINDOW was
not verified. fleetd #618 measured it: the env var wins, so the flag has
no effect when both are set. Update both texts to say so, name #618, and
warn that deleting the env var to resolve the conflict LOWERS the live
window rather than fixing anything. No behaviour change; the WARN still
fires on the same condition and stays a WARN (per #601).
2026-09-22 10:46:31 +07:00
ltms 17127efb88 Merge #601: pass auto-compact window to leads; warn instead of refusing on a conflict (CB-617)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m16s
CI / build (push) Failing after 2m33s
2026-09-22 05:22:56 +02:00
ltms 203f034528 Merge #617: write FAILED instead of leaving a dead roll stuck at IN_PROGRESS (fleetd #615)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m14s
CI / build (push) Failing after 1m47s
2026-09-22 05:21:21 +02:00
ltms 9ee16f5b85 Merge #616: report role-fallback gaps at boot, name contextHighNudge (fleetd #613)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m15s
CI / build (push) Failing after 1m43s
2026-09-22 05:17:52 +02:00
Dai Ha 388ef5a3c3 fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m25s
LeadRollover.runRollover made two unwrapped agents.send calls. HerdrException
is unchecked, and the production continuationRunner is a bare virtual thread
with no uncaught-exception handler, so a throw from either send call killed
the continuation silently — confirm() had already written IN_PROGRESS into
outcomes before scheduling it, and nothing ever overwrote that entry with a
terminal state.

Wrap the whole continuation body in one try/catch(RuntimeException), matching
the local convention already used around agents.status in
waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED
entry naming the exception, in the same diagnostic style as
TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED.

Two new tests make send() throw on the /clear call and on the bootstrap-text
call respectively, each asserting status(token) reports FAILED, not
IN_PROGRESS. Reverting only the production catch (keeping the tests) turns
both red; restoring it turns them green again.
2026-09-22 10:17:12 +07:00
Dai Ha be6c45ff78 CB-617 review: warn instead of refuse on conflicting autoCompactWindow
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m9s
rejectConflictingAutoCompactWindows threw and stopped fleetd from starting when a
Claude Code profile's autoCompactWindow flag and CLAUDE_CODE_AUTO_COMPACT_WINDOW
env var disagreed. Under launchd that is a restart loop, and the config that
would fix it (fleetd.yaml) is gitignored, so the cause is invisible on the host
where it bites (measured live: 4 profiles on this host trip it, including the
lead's own profile and the one every worker spawns on).

Renamed to warnConflictingAutoCompactWindows: it now logs a WARN naming each
offending profile with BOTH values (autoCompactWindow=... and
env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=...) instead of throwing, so the daemon
starts and an operator can fix the config without reading the source. Equal
values still load silently.

Also reworded ClaudeCodeArguments' javadoc, which stated as fact that the env
var takes precedence over the flag. That was never measured, and this host's
own fleetd.yaml comment asserts the opposite — the javadoc no longer picks a
side.
2026-09-22 10:13:14 +07:00
Dai Ha e99cb70a8b CB-617: pass auto-compact window to leads 2026-09-22 10:12:56 +07:00
Dai Ha 987ccef4c7 fleetd #613: log role-fallback gaps at boot, name contextHighNudge in the heartbeat line
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Failing after 1m48s
- reportRoleFallbackGaps(cfg), called right after cfg.validateAll() in Fleetd.main, logs every
  MemberRole with no fleet.<role>s: pool (naming the profile count and the resolved
  defaultProfileFor(role) first choice) and, separately, every role with no
  fleet.charters.<role>: entry. Log only — the deliberate 'unconstrained' fallback in
  FleetConfig#candidateProfiles / CompositePeerLauncher#poolFor is unchanged, and a config with
  profiles: and no fleet: block still starts and still spawns.
- LeadHeartbeatLoop#start()'s boot line now also names contextHighNudge (fleetd #609), alongside
  the three settings it already logged.
- RoleFallbackGapReportTest (new) and two new LeadHeartbeatLoopTest cases pin both lines' content
  via a ListAppender, raising the dev.ltms.fleet logger past logback-test.xml's WARN override for
  the INFO-level lines.
2026-09-22 10:12:09 +07:00
ltms 076cc43f7b Merge #614: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 2m20s
CI has been red on main itself since #602/#606, on this one test, so the build has been giving no second opinion on any PR. Cause: the CI job runs in a container as root. setReadable(false) really does clear the read bit, so the test's own setup guard passes, but root opens the file anyway and the gauge correctly returns OK. The test was asserting on a condition the environment never created.

The fix adds an assumeFalse(Files.isReadable(file), ...) after the chmod and before the gauge is built, inside the existing try, so the finally still restores the bit on a skip.

Verified by me on a scratch worktree merging this onto 955b9ea:
- 1864 tests, 0 failures, 0 errors, 0 skipped, 149 surefire reports, mvn exit 0. The suite-wide skipped=0 is the point: the fix did not quietly turn the test into a permanent skip.
- LeadContextGaugeTest on this non-root Mac: 9 tests, 0 skipped, and unreadableFileIsUnknown present in the report. The assumption does not fire here, so developers keep the coverage.
- Mutation: made the IOException path return OK instead of UNKNOWN. unreadableFileIsUnknown failed with "expected: <UNKNOWN> but was: <OK>". The test still has teeth. Production file reverted, git diff clean before merge.

Known trade, recorded rather than hidden: under root this case is now covered by nothing at all. A skip is honest about that, which an assertion on an unreachable state was not. The durable fix is to run the CI build as a non-root user; that is a CI configuration change and out of scope here.
2026-09-20 12:43:43 +02:00
Dai Ha bad47a8444 fleetd CI: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m27s
The test set the file's read bit off via setReadable(false), but on the
Gitea CI runner (root inside the container) the OS ignores that bit and
opens the file anyway, so the test asserted on a condition it never
actually created (LeadContextGaugeTest.java:142 UNKNOWN vs OK, CI run
1887 job 3104, commit fa62e99 on main).

Add Files.isReadable(file) after setReadable(false) and before the gauge
runs, and assumeFalse on it: a skip means "could not set up the case",
never "the behaviour is fine". Restores the read bit either way so
@TempDir cleanup still works.
2026-09-20 17:40:47 +07:00
ltms 955b9ea013 Merge #610: nudge an idle lead to hand over when its own context reads HIGH (fleetd #609)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m51s
Closes fleetd #609. Completes the second half of the context work: #602/#606 could detect a full lead context, and nothing acted on it. LeadHeartbeatLoop now offers a handover when the lead's own gauge reads HIGH.

Never rolls a pane by itself. The nudge is text only; the lead still has to call fleet_handover, and that still needs operatorConfirmed.

Verified by me on a scratch worktree merging 89cb8ff onto fa62e99:
- 1864 tests, 0 failures, 0 errors, 149 surefire reports, mvn exit 0.
- Four mutations, all killed: the two the implementer ran (latch back in applyDecision: 4 failures; call site drops the latch argument: 2 failures), one of my own at the line the logic moved TO (latch set regardless of send outcome: 1 failure), and a control on an untouched line (quiet-cap boundary < to <=: 4 failures). The control is what makes the other kills evidence.
- The new tests assert on herdr.sentTexts() — what actually reached the fake pane — not on source text. That is the right observable for a defect whose essence was "the latch says told, the pane got nothing".

The review blocker from the first round is fixed: the latch used to be committed by applyDecision before injectNudge tried to send, and injectNudge swallows its own RuntimeException. On quietNudgeCap: 0, which is this host's configuration, that was the normal path and not an edge case. The latch is now set only when a notice was actually included and the send returned.

Known and deliberately not blocked: Fleetd.main's own one-line call to leadContextSource is not pinned by a test. That is pre-existing and class-wide, tracked in #612.
2026-09-20 12:37:43 +02:00