33720c42b3c97d6ac4fb0b6e20c5023bf638f8fc
902 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
33720c42b3 |
fleetd #512 (part 1): log a positive completion line when drainAll finishes
drainAll used to log nothing on a clean drain — both existing log calls (drainSnapshot's per-session failure, drainAll's straggler-sweep warning) sit on abnormal paths, so "drained fine" and "died on the first session" looked identical: no log line either way. Add one log.info at the end of drainAll: "drain complete: released=N abandoned=M (still BUSY at the shutdown deadline)". It fires on the normal path, including the all-zero case, and folds both drainSnapshot passes (main snapshot + straggler sweep) into one line. drainSnapshot now returns a private DrainTally(released, abandoned) record instead of void, and the private release(paneId, cause) overload now returns the removed MemberSession (previously void) so drainSnapshot can read its state at the moment of removal — the same check logPreservedForShutdown already makes. Both signature changes are private with a single call site, so the blast radius stays small. |
||
|
|
b37def9238 | Merge #516: the probe refuses with three distinct messages, each naming its own cause (fleetd #500) | ||
|
|
8f02576df6 | Merge #515: pin the two-client completeness fold, and legacyPrincipal earns no authority (fleetd #509) | ||
|
|
525bc1c5f4 | Merge #514: the drain-gate abort message names a recovery that works, and jar_id()'s default is pinned (fleetd #511) | ||
|
|
d59ece6dec |
fleetd #500: stop a wrong-interpreter or failed-parse reading a policy as empty
probe-member-credentials.sh used mapfile < <(producer) to parse the fetched policy. That
hides a producer failure three ways: mapfile is bash 4+ and missing on macOS's /bin/bash
3.2, a process substitution's exit status is never propagated to mapfile, and the
downstream reads (":-" defaults and a slice) never fire set -u on a short or unset array.
All three converge on the same "0 known names" refusal, which blames the policy for a
failure that is actually the interpreter or the parser.
Three distinct guards, each closing one cause with its own message:
- a BASH_VERSINFO gate at the top refuses outright on bash < 4 (exit 3)
- the parser's output is captured via command substitution instead of mapfile < <(...),
so a non-zero jq/python3 exit is caught at the call while the fact still exists (exit 4)
- an arity check before the field slice refuses a parse that exits 0 but returns fewer
than 5 fields (exit 5)
The existing "0 known names" guard is now honest: by the time it fires, the three causes
above are already ruled out, so it really does mean the policy has 0 known names.
|
||
|
|
32408d1e64 |
fleetd #509: pin the pane-scan completeness fold, and stop legacyPrincipal handing out primary
Unit 1 — PaneLocator.terminalForPid's completeness fold across herdr clients (PaneLocator.java:117) had no test that varied the number of clients, so a mutation that keeps only the last client's Lookup.complete() instead of ANDing every client's outcome survived: 14 of 15 existing tests agree with the mutant on a single client. Added a two-client test where the lead client errors on the pane that would have owned the pid (an incomplete, negative scan) and the member client cleanly finds no panes (a complete, negative scan) — the real fold ANDs these to false, a last-wins fold reads it as true. Proved against MUTANTC (complete = outcome.complete();): the new test fails with "expected: <false> but was: <true>", the file was restored byte-identical (sha256 unchanged), and the control run is green. Unit 2 — FleetMcp.legacyPrincipal's else-branch returned Principal.primary for ANY caller the connection did not resolve to a worker pane, with none of CallerResolver.java:254's isLoopback/scanComplete guards. Measured that no production caller passes null callers (Fleetd.java:696 always constructs a real CallerResolver) but FleetMcpAuthzTest.mcp(false) legitimately does, for its "legacy constructor leaves the gate open" test — so the null-callers path is not dead code to delete (option a), it is a documented legacy mode (option b). Changed the else-branch to Principal.anonymous() and widened legacyPrincipal to package-private (like denyFor) so a new test pins the behavior directly, since it only ever ran inside a contextExtractor closure no existing test triggers. |
||
|
|
6e23bf8309 |
fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
The drain-gate abort message told the operator a rerun "with or without --no-build" would finish the restart. That is wrong: by the time this message can fire, stage_built_jar has already moved the jar off $JAR, so --no-build hits require_no_build_jar's own refusal. Reworded to say the rerun must NOT use --no-build, and why: the built jar is no longer at the live path that --no-build requires. Also added a test pinning jar_id()'s no-argument default (reports $JAR, the live path) and its explicit-argument behavior (reports that path instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm -f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn clean install deletes target/ on the next line regardless). |
||
|
|
aa4c0b84c3 | Merge #510: never build into the path a running daemon holds (fleetd #493) | ||
|
|
40c593cd09 |
Merge #508: a herdr error during the pane scan is refused, not promoted to primary (fleetd #505)
Closes the worker->primary escalation that fleetd #317 left open on the other side of the same
ternary. #317 stopped an UNRESOLVABLE pid being promoted. This stops a RESOLVED pid whose pane scan
failed being promoted: the scan now reports completeness, and an incomplete scan resolves to
anonymous.
Shape chosen: PaneLocator.terminalForPid returns Lookup(terminal, complete); paneOwnsAnyOf returns a
private Ownership enum OWNS/DOES_NOT_OWN/UNKNOWN, so a HerdrException is UNKNOWN rather than a clean
negative. ConnectionIdentity.Caller gains scanComplete as a third field -- resolved() was NOT
widened, correctly: it is documented as testing the lsof sentinel and #505 is a different axis. A
definite match still short-circuits, so a genuinely vanished non-owning pane does not become a
refusal.
Verified by me, not taken from the report.
Trial-merged onto main (
|
||
|
|
979adf82eb |
fleetd #493: never build into the path a running daemon holds
redeploy-fleetd.sh's build step wrote straight into fleetd/target/fleetd.jar
via `mvn clean install` while the OLD daemon was still running from that
exact path. A JVM loads classes lazily, so a class the daemon had not
touched yet could be read from a jar already replaced or removed -- the
failure landed on the shutdown drain (NoClassDefFoundError, exit 143,
looks clean).
Stage the freshly built jar at target/fleetd-new.jar (stage_built_jar),
confirm the OLD pid has actually exited (wait_for_daemon_exit, extracted
from the existing wait loop), and only then swap it into the live path
(swap_staged_jar) -- strictly after the wait, strictly before start. A
failed swap dies without starting. --no-build and --check keep their
existing, truthful behavior (require_no_build_jar; jar_id still reads
the live path by default). A leftover staged jar from an interrupted
run is wiped before the next build. The build still runs before
anything is stopped, so a failed build still never takes the fleet down.
Also fixed: the drain-gate abort message ("aborted -- nothing changed")
now names the staged jar when one exists, since staging already moves
the freshly built jar off the live path before that prompt runs.
Adds unit tests for stage_built_jar, swap_staged_jar, require_no_build_jar,
wait_for_daemon_exit, and a source-order test proving swap sits after the
wait and before start (sourcing stops before the main flow ever runs, so
the ordering itself can only be checked by reading the script's own call
sites).
|
||
|
|
36870836aa |
fleetd #505: a herdr error during the pane scan must not read as a clean negative
A transient herdr error on pane.process_info during PaneLocator's pid→pane scan used to be swallowed into a plain "does not own it", so a real worker whose owning pane errored mid-scan resolved with a null terminal but a resolved (real) pid — exactly what CallerResolver's loopback-trust fallback reads as the primary. That is a worker→primary privilege escalation through the door fleetd #317 did not close: #317 guards a failed lsof lookup (c.resolved()), not a failed herdr pane scan. Fix: add a third state to the scan instead of widening Caller.resolved() (which stays centralised next to the lsof sentinel it tests, per #505's explicit instruction not to reopen that decision). PaneLocator.terminalForPid now returns a Lookup(terminal, complete) record: a HerdrException on one pane marks that pane's ownership UNKNOWN, not DOES_NOT_OWN, and the scan is complete only if every pane was either matched or confirmed not to own the pid. A definite match found elsewhere in the same scan still short-circuits as complete — a pane that genuinely vanished mid-scan without being the caller's own does not turn into a refusal. ConnectionIdentity.Caller carries the new scanComplete flag alongside the unchanged resolved(). CallerResolver's loopback-trust fallback now requires both resolved() and scanComplete() before promoting to Principal.primary(); an incomplete scan resolves anonymous, which fails toward the recoverable error (a refused primary retries loudly; a promoted worker would not). Logs a warning naming the pane and which herdr client (of how many) failed, so the incomplete-scan path is diagnosable rather than silent (fleetd #317's own lesson). |
||
|
|
136312fb11 |
Merge #503: Injector's readiness-grace warn prints measured elapsed time, never arithmetic on constants (fleetd #501)
Verified by the lead, not taken from the worker's report. Trial-merged onto main ( |
||
|
|
4f28da62a3 |
Merge #499: redeploy-fleetd.sh gains an "unclear" supervisor state that cannot reach kill, and the detail survives the subshell (fleetd #492, #497)
Two commits, |
||
|
|
708f1795ad |
Merge #502: awaitHerdr reports three outcomes with measured elapsed time, not one boolean (fleetd #498)
Verified by the lead, not taken from the worker's report. Trial-merged onto main ( |
||
|
|
599419f9e6 |
fleetd #492 follow-up: carry the unclear detail across detect_supervisor's subshell boundary
|
||
|
|
ac351ee1de |
fleetd #501: readiness-grace expiry logs measured elapsed time and the loop's own poll counter, never the configured budget
The line at Injector.java:383-387 printed two numbers that read as measurements but were both compile-time constants: READINESS_GRACE_POLLS for the poll count (the loop's own Target.notReadySincePoll counter was in scope at the same call site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed time — arithmetic on two constants, never a measurement, and wrong in the direction that says everything ran on schedule. Fix, copying the LongSupplier-clock shape LeadRollover already uses: - print t.notReadySincePoll instead of the constant for the poll count (the two agree by construction on this branch, so no test can tell them apart — the comment says so honestly). - add Target.notReadySinceMillis, stamped at the first non-ready sample and reset at all three sites notReadySincePoll already resets (:304, :355, :389 pre-fix line numbers), to compute a real elapsed time at expiry. - inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis) through new package-private constructor overloads so a test can supply a clock whose advance does not track POLL_INTERVAL_MILLIS. Tests use a ListAppender to assert on the log message contents, per LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub clock returning two literal, non-derived values so it can fail if the fix regresses to the constant-arithmetic line — proved by mutation: reverting the elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that one test red (1 failure); reverting the poll-count print to the constant is an equivalent mutant (0 failures), because the counter and the constant are identical at that exact call site by construction. fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0. |
||
|
|
274afafde6 |
fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are two distinct, named states instead of the same false (fleetd #497's shape). - awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required parameters, with no defaulted overload (fleetd #415), so a test can drive it. - The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself cannot be driven from a unit test; it logs a distinct message per outcome, always printing the measured elapsed time next to the configured budget, never the budget alone. - Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt flag) and the call site (the three distinct log messages), using ListAppender. |
||
|
|
8f59019305 |
Merge #496: lead-rollover logs measured elapsed time and measured counts, never the configured budget (fleetd #494)
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report. |
||
|
|
e966cbadf9 |
fleetd #494 follow-up (2nd pass): the grace-release line still had one constant, and its test could not tell the difference
Two more fixes on the same line, LeadRollover.java:600-624:
1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
loop's own idlePollsAwaitingPickup counter instead of the constant. Same
defect shape as the nudges fix from
|
||
|
|
b17f37a683 |
fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
Two situations were silently landing in the "none" answer, which require_drivable_supervisor accepts and the script then falls back to a raw kill + nohup — exactly the wrong move when a supervisor actually IS present: - installed-but-not-loaded, on either supervisor. `systemctl --user is-active` answers "no" for activating/deactivating/failed and while an auto-restart is pending too, and every one of those is a host that IS under systemd (or launchd) and about to act again. `*_installed` already knew this; it was only ever consulted for a warning line, never by the decision itself. - a systemd probe that could not answer at all (e.g. systemctl cannot reach the user bus over a non-lingering ssh session) looked identical to a clean negative, because both probes redirected stderr straight to /dev/null. detect_supervisor now returns a fifth answer, "unclear", for both cases. systemd_loaded/systemd_installed capture systemctl's exit status and stderr separately and set their own *_ERRORED flag only on a real tool failure (non-zero exit WITH stderr), never on a clean negative. "none" now means only: neither supervisor installed, neither loaded, neither probe errored. require_drivable_supervisor die()s on "unclear" exactly like it already does on "ambiguous", naming the specific supervisor and reason via the new SUPERVISOR_UNCLEAR_DETAIL global. Tests: 4 new cases (systemd/launchd installed-but-not-loaded, a real systemd_loaded run through a systemctl stub that errors on stderr, and the die() refusal for "unclear" naming the unit). All 3 new guards were verified by mutation: each was removed from the real script, the suite caught it (a new FAIL line naming the exact broken assertion), then the file was restored byte-identically and the suite went green again. |
||
|
|
c87cc25aa6 |
fleetd #494 follow-up: print the measured nudge count, and fix the sibling turn-settle timeout line
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8 nudges where 7 went out) and a code read did not catch it — only a mutation test did. Printing the counter cannot drift from the loop's real behaviour. - waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value printed as if measured' defect as the three lines fixed in the original #494 commit, but was out of scope because the brief named specific lines instead of the shape. Fixed the same way: it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds() as the measured wait. - No behaviour change: same sends, same order, same release/refuse decisions. - Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line. - Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via two greps that the mutant applied and the original constant was gone, ran the grace-release test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8 and reran the full LeadRolloverTest suite as a control (33/33 green). |
||
|
|
3fb331145a |
fleetd #494: log measured elapsed time, never the configured budget, on a lead-rollover failure/success
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
prints the measured elapsed time next to the target pane — this is the exact path that
reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
self-advancing fake clock so the measured elapsed/nudge values are deterministic and
provably distinct from the configured budget.
|
||
|
|
dcd505286f |
fleetd #492: teach redeploy-fleetd.sh systemd --user as a third supervisor
launchd, systemd --user, and unsupervised are three different answers, not two. Refuse (die) rather than fall through to kill+nohup when a supervisor is detected that this script cannot drive (e.g. both signals fire at once), and add a post-restart check that fails the run if more than one fleetd process is alive. detect_supervisor()/require_drivable_supervisor()/ count_daemon_pids()/assert_single_daemon() are pure, overridable functions so scripts/test-redeploy-fleetd.sh can exercise them without a real launchd or systemd. |
||
|
|
60fa958da2 |
docs: a relative handoverPath lands in the LEAD's repo, not fleetd's (#491)
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore protects the handover file and the distinction is invisible. On fleet01 the lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory, and that repo has no .handover rule — measured 2026-09-12. Tell the lead to check its own workspace's .gitignore before writing, and to report it rather than committing the file or editing someone's .gitignore. Refs fleetd #491, #487, #480. |
||
|
|
9950361bc9 |
docs: #489 is fixed and deployed, but criterion 6 is still unmet
The skill said the roll 'does nothing until #489 is merged and redeployed'. Both have happened, so that sentence now reads as a green light. It is not one: no roll has bootstrapped a fresh session end to end yet. Say that plainly, and tell the lead to warn the operator before confirming. Refs fleetd #489, #480. |
||
|
|
9f74b6619a |
Merge #490: nudge the /clear submit keystroke before bootstrapText (fleetd #489)
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".
Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on
|
||
|
|
f687046450 |
fleetd #489 follow-up: fix stale class javadoc, off-by-one nudge count, weak test
Three review corrections on top of the previous commit:
1. The class javadoc's four-step continuation list (lines 47-58) was stale.
Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
described the old plain re-check ("the original, pre-correction wait...
still here") - fixed to describe what waitForClearPickupAndSettle
actually does: nudge while unpicked-up, then wait for a real WORKING ->
IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.
2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
polls, not the number of nudges - the 8th poll releases instead of
nudging again, so 8 polls produce 7 nudges. The log.info in the release
branch and two javadoc spots said "8 nudges"; fixed all three to state
the poll count and the nudge count separately and correctly. Behavior
and the constant are unchanged.
3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
and sendKeysCallCount, both of which a return-true stub also satisfies.
Added an assertion on the already-tracked postClearGetCalls counter
(>= 2), which only a real post-/clear poll loop can produce - this is
what makes the test fail against a return-true mutant.
|
||
|
|
c6058652be |
fleetd #489: nudge the /clear submit keystroke before bootstrapText
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled for IDLE/DONE, which /clear itself never leaves since it starts no real turn, so it always returned true on the first poll. Combined with a direct agents.send bypassing Injector (deliberate, to avoid wedging the pane), the submit Enter that accompanies /clear could race the paste and leave it unsubmitted — bootstrapText then landed concatenated onto the same input line, exactly as measured live on 2026-09-12. Replace that second wait with waitForClearPickupAndSettle, which copies the pickup-nudge pattern Injector already ships for its own post-turn /clear housekeeping (fleetd #306): nudge agents.submit while the pane hasn't reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is observed. BLOCKED stays excluded from both the nudge and the boundary check, same as the (unchanged) first wait — a paused live turn is not settled, and nudging Enter into an open prompt could wrongly answer it. Adds four tests to LeadRolloverTest covering the paste-race regression (nudge ordered between /clear and bootstrapText), a confirmed pickup, a deadline expiry with no boundary ever reached, and a throwing submit(). |
||
|
|
2a95da2ff4 |
docs: the handover skill must warn that the automatic roll is broken (#489)
Section 11 said the bootstrap prompt landing was 'not yet proven end-to-end'. It is now measured failing: the first real roll joined /clear and bootstrapText into one line. Nothing was cleared, so the failure is safe, but a lead that reads the old wording would reach for fleet_handover expecting it to work. Refs fleetd #489, #480. |
||
|
|
7f9137fcb6 |
fleetd #480: ignore the handover file, and note the absolute-path guarantee in the handover skill
The lead rollover handover file now lives inside the workspace, at the relative path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one moment's live state, so it must never enter git history. The handover skill also now says the handoverPath fleetd hands back is always absolute, even when the configured value is relative — a lead that resolves it itself can pick a different file from the one the daemon checks. |
||
|
|
3bf3968bc7 | Merge #487: resolve a relative leadRollover.handoverPath against the calling lead's workspace (fleetd #480 follow-up) | ||
|
|
261aa056f9 |
fleetd #480 follow-up correction 2: guarantee resolveHandoverPath is always absolute
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.
Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.
Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.
Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.
Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
|
||
|
|
042b8c99dd |
fleetd #480 follow-up correction: cover Fleetd.leadRollover(...)'s own wiring behaviourally
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...) factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at construction time (a lead discovered by the tab scan after leadRollover(...) was built still resolves correctly). Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always "no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full 1669-test suite green before this commit. With the new test added, the same one-line mutation now fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during verification and is not part of this commit (git diff on Fleetd.java is empty). FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test could catch this wiring dropping out, which was true only before this commit and only covered the factory's own body, not its call site. Restated what each test class actually covers: the source- text pin covers the call site's argument list; the new behavioural test covers the lambda's body. |
||
|
|
4bfab6b718 |
fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed back to the lead in the fleet_handover open response, and the default bootstrapText sentence all see the same absolute location instead of a value resolved against whatever directory the daemon process happened to start in. FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor (resolvedHandoverPath) method builds the default sentence from the resolved path instead. Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on every call, never off a startup snapshot. |
||
|
|
008a457557 |
docs: fleet_handover is shipped — update the charter table and the handover skill
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the canonical block is an incomplete change. CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover. Both copies edited identically; the byte-identical sync check prints True. The row carries the trap rather than just the call: open FIRST, then write the file, then confirm — because confirm refuses the file as stale unless its modified time is later than the open request, so the obvious order fails. .claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover tool: it does not exist." That was true when it was written this morning and is now false, which is the worst state for an instruction file to be in. Replaced with the two real paths (by hand, or with the tool when leadRollover: is configured), plus a new section 11 giving the three-step order and the five things that surprise a caller — chiefly that `accepted` does not mean the pane has been cleared, and that operatorConfirmed is a report of what a human said, not a confidence level. Section 11 also states plainly that the bootstrap prompt landing in a freshly cleared pane is not yet proven end-to-end, and that the recovery is the manual path — which is why the file is written before confirm, never after. |
||
|
|
9494a6b99a |
Merge #485: fleetd #480 Unit C — the fleet_handover MCP tool
fleet_handover{action: "open"|"confirm"|"cancel"} drives LeadRollover, which #483 and
#484 landed with nothing calling it. Primary-only via a new Authz.Action.HANDOVER, on
the same case line as SPAWN/STOP/DRAIN.
The tool has NO terminal, session or leadTerminal parameter of any kind — the pane is
always callerTerminal(exchange), resolved from the connection. A lead can therefore only
ever roll itself, never another lead. That is charter invariant 3, and it is the second
of the two corrections recorded in LeadRollover's class javadoc.
Registered unconditionally, so the tool surface does not vary with config: with
leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal instead of
failing, and open()'s IllegalStateException (config removed by a hot reload after
construction) is caught and turned into the same refusal. A config-dependent tool set
would have collided with #474's charter tool-surface gate and McpContractDocTest.
Correction round applied before merge, and it is the reason this took two passes.
The unit first shipped with a defaulted 15-argument FleetMcp constructor delegating to
the new 16-argument one with leadRollover = null. I mutated the wiring rather than
reasoning about it: deleting just the leadRollover argument from Fleetd.main's FleetMcp
call compiled with 0 errors and passed all 1659 tests, BUILD SUCCESS — while the live
daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
Neither FleetMcpHandoverTest (it builds its own FleetMcp) nor FleetdLeadRolloverWiringTest
(it pins that LeadRollover is constructed, not that it is passed on) could see it.
The worker then found the defect was wider than I had named: all five shorter
constructors (11/12/13/14/15-arg) formed one defaulting chain into the 16-arg one, each
silently supplying another feature's "off" value — leadChannel, outage, leadSeats, peers,
and finally leadRollover. All five are deleted. FleetMcp now has exactly one public
constructor, so every one of those features is compile-enforced at its call site, not
just this one.
Verified by the lead before merge, on PR head merged with current main (0176378):
- CI run 1728 green on
|
||
|
|
eb0557621e |
fleetd #480 correction round: collapse FleetMcp to one required constructor
FleetMcp had a defaulted 15-argument constructor that delegated to the new 16-argument one with an implicit null for leadRollover. Dropping the leadRollover argument from Fleetd.main's FleetMcp(...) call fell back to that shorter overload, compiled fine, and left all 1659 tests green — the live daemon would then answer NOT_CONFIGURED to fleet_handover forever with nothing going red. Delete every overload that could reach the 16-arg constructor with a silently-defaulted leadRollover (11/12/13/14/15-arg forms all chained to it), leaving the 16-arg constructor as FleetMcp's sole public constructor. Update FleetMcpAuthzTest's call site to pass every parameter explicitly (leadChannel null, OutageSource.none(), LeadSeatSource.none(), List.of(), leadRollover null) — Fleetd.java and FleetMcpHandoverTest already called the full form. Proved with mvn -o -q compile: removing the leadRollover argument from Fleetd.main now fails to compile instead of silently defaulting. No behaviour changes — NOT_CONFIGURED refusals are unchanged. |
||
|
|
a0eed6f01b |
Merge #484: fleetd #480 Unit E — a BLOCKED pane is not a settled pane
LeadRollover's two settle waits used AgentStatus#injectable(), which is
IDLE || BLOCKED || DONE. That is the right rule for Injector ("may I deliver a message
without stepping on a live turn") and the wrong one here ("has the turn actually
ended"), because BLOCKED is a live turn that is merely paused — a pane sitting on an
approval prompt.
Path in: the lead calls confirm(); its turn carries on and hits anything needing
approval; herdr reports blocked; within 250ms the deferred continuation reads that as
settled; /clear is typed into an open prompt. That destroys the lead's live context
mid-turn, which is exactly what the turnSettleSeconds gate added in #483 exists to
prevent. The window is turnSettleSeconds, default 20s.
Both waits now require a real turn boundary — IDLE or DONE. AgentStatus#injectable() is
untouched: it is correct for Injector, LeadHeartbeatLoop, LeadCoordLoop, ReplyPushLoop
and HerdrPeerLauncher's two readiness checks, all of which are delivery gates.
Verified by the lead before merge:
- CI run 1726 green on head
|
||
|
|
e2a91e883e |
fleetd #480 Unit E correction: retire "injectable" wording from the log lines
Both waitUntilAtTurnBoundary guard messages still said "never went idle" / "did not become injectable" — the old mental model the rename was meant to retire. Made both say what the code now actually waits for: a turn boundary (IDLE or DONE). |
||
|
|
62646957ea |
fleetd #480 Unit C: wire fleet_handover MCP tool onto LeadRollover
Adds the fleet_handover tool (open/confirm/cancel) as a thin adapter over LeadRollover, registered unconditionally so the charter tool-surface gate sees a stable set regardless of whether leadRollover: is configured. With a null LeadRollover every action degrades to a clean NOT_CONFIGURED refusal instead of throwing. Gated on a new Authz.Action.HANDOVER (primary-only, same as SPAWN/STOP/DRAIN). The caller's own connection-resolved terminal is the only lead identity ever used — the tool's input schema carries no terminal/session/leadTerminal parameter, so a lead can only ever roll itself. Fleetd.main now passes its existing leadRollover local into FleetMcp via a new trailing constructor parameter. |
||
|
|
802c0ab701 |
fleetd #480 Unit E: BLOCKED is not a settled turn boundary
LeadRollover's waitUntilInjectable used AgentStatus#injectable(), which accepts BLOCKED. A BLOCKED pane is paused mid-turn on a prompt, not settled — reusing injectable() let /clear (or the bootstrap text after it) fire into an open approval prompt within the 20s settle window, destroying the lead's live context. Renamed the helper to waitUntilAtTurnBoundary and restricted both waits to IDLE or DONE only, with a comment explaining why this class does not reuse injectable() (it answers "may I deliver", not "has the turn ended"). Added tests for BLOCKED-forever on both waits (zero sends / exactly one send) and for DONE still completing the full roll. |
||
|
|
4bab23e241 |
Merge #483: fleetd #480 Unit A — lead rollover core (config + executor)
Two corrections were applied to the original unit before this merge, and the PR description above still describes the pre-correction shape: 1. confirm() no longer rolls inline. It validates every gate, then hands a one-shot continuation to continuationRunner and returns. confirm() is called BY the lead FROM its own turn, so the pane is WORKING and cannot report injectable until confirm() returns; the old code sent /clear first and then timed out waiting, which destroyed the lead's context and started no fresh session. The continuation waits for the calling turn to settle FIRST (turnSettleSeconds, a new knob), and if that wait times out it sends no /clear at all. 2. The pane to roll comes from the caller's terminal, not PrimaryRegistry. A single-slot lookup let lead X's confirm() clear lead Y's pane. open() records the caller's terminal; confirm() refuses with NOT_YOUR_ROLLOVER on a mismatch. Verified by the lead before merge: - CI run 1722 green on head |
||
|
|
a94262271b |
fleetd #480 correction round: defer the roll, and gate it on caller identity
Two defects found after the fact, both from the original brief, both fixed here. 1. confirm() is called FROM the calling lead's own turn, so its pane is still WORKING and can never report injectable inside that same call. The old confirm() sent /clear before polling for that — the poll always timed out, but only after /clear had already fired and queued, destroying the lead's context with no fresh session ever started and a refusal return that lied about what had happened. Fix: confirm() now only validates and, if every gate passes, hands a one-shot continuation to a new continuationRunner (a real virtual thread in production, Runnable::run in tests) and returns RollDecision.approved() immediately - "scheduled", not "rolled". The continuation itself does the actual work, once the calling turn has ended: wait for the SAME pane to report injectable again (new turnSettleSeconds config key, default 20) - if this never happens, /clear is NEVER sent, at all - then /clear, then wait again (clearSettleSeconds, as before), then bootstrapText. The "no timer/scheduler, only confirm() can roll" invariant is restated precisely in LeadRollover's class javadoc: it is about initiative, not synchronicity - a single-shot continuation of an already-approved confirm() call still satisfies it; a recurring background loop would not. 2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(), a single-slot lookup that is correct for a background loop with no caller but wrong here: on a daemon with more than one labelled lead tab, lead X's confirm() could clear lead Y's pane, violating the charter's "identity comes from the connection, never an argument" invariant. Fix: open() and confirm() now take the caller's terminal id as a parameter (resolved by the MCP layer from the connection - the later MCP-tool unit must pass it in, never accept it as a request field). confirm() refuses with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open() recorded. LeadRollover no longer depends on PrimaryRegistry at all. Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the new meaning - approved and scheduled, not necessarily cleared yet. Added turnSettleSeconds to the leadRollover: config block (documented in fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's leadRollover(...) factory to drop the primaryRegistry parameter, with FleetdLeadRolloverWiringTest's source-text pin updated to match. New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters most - a turn that never ends means /clear is never sent) and aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER), plus a settle-after-clear timeout test and an open() input-validation test. LeadRolloverTest: 11 -> 14 tests. |
||
|
|
5c12865c25 |
fleetd #480 Unit A: lead rollover core (config block + executor)
Adds the opt-in leadRollover: config block and LeadRollover, the executor a later unit's MCP tool will call. A lead writes a handover file, then open() records a token and confirm() verifies it (exists, non-empty, fresh) and an operator confirmation before clearing the lead's own pane via /clear (sent directly through AgentControl, bypassing Injector, same as ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing but an explicit confirm() call can ever roll a pane - no timer, no heartbeat, no background thread anywhere in this class. Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when leadRollover: is present at startup, and nothing calls it yet - the MCP tool is a separate, later unit. Classified leadRollover: as HOT in ConfigRef (joins placement/ memberCredentials/memberLoginShell/models): the executor holds Supplier<FleetConfig.LeadRollover> and reads every field fresh per call, unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's construction is still gated on presence in the startup config snapshot, so a freshly-added block needs a restart before anything exists to call. Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no object without the config block, only confirm() can roll, missing/empty/ stale handover file each refuse by name, requireOperatorConfirm gating, and the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source- text pin on Fleetd.main's construction call, mirroring FleetdCompletionResolverWiringTest). Also updated the existing FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest, ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to account for the new record component. |
||
|
|
3f8c956325 |
docs: the handover skill must not promise a tool that does not exist
The skill shipped in #481 described fleetd #480's automation in the present tense: "fleetd ticket #480 lets a lead session hand off to a fresh one". The tool is not built. A session loading the skill would look for a fleet_handover tool it cannot call, and this repo already knows what a confidently wrong instruction costs — a wrong comment stops the reader investigating with a false conclusion, which is worse than no comment. Now says plainly: today the handoff is manual, #480 will automate the same cycle, and do not call fleet_handover because it does not exist. The nine procedure rules are unchanged — only who performs the swap changes when #480 ships, not what the file must contain. My wording to fix, not the worker's: the brief handed them the present tense. |
||
|
|
2051ceeb26 |
Merge #481: the handover skill (fleetd #480 Unit B)
Adds .claude/skills/handover/SKILL.md — the outgoing lead's procedure for writing
the file a fresh lead session inherits.
Reviewed by the lead against the eight requirements in the brief: every number
carries its command, measured-vs-reported is marked, open decisions name their
owner, live hazards, what is not owed, a first section of three re-measurement
commands, a time-and-commit stamp, and a plain statement that the file goes
stale. All eight are present, plus a leave-out section and a template.
Verified by the lead, not taken from the worker's report:
- CLAUDE.md diff is the addendum's skill list only, in the primary-side group.
- The canonical block is byte-identical with the wiki template (sync check True).
- CI run 1718 green on
|
||
|
|
325d0771a4 |
fleetd #480 Unit B: add handover skill for lead session handoff
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead follows to write the handover file a fresh lead session inherits when fleetd clears the pane. Registers the new skill in CLAUDE.md's primary-side skills list; no other change to CLAUDE.md. |
||
|
|
7b97aae85b |
docs: move the redeploy procedure out of CLAUDE.md into a skill
CLAUDE.md loads into every session. The daemon-redeploy procedure is needed only when someone redeploys, so it paid for context it did not use: 58 lines, about 918 est. tokens, every session. The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md, which loads only when invoked. The moved text is byte-identical to what was removed, plus one new paragraph documenting --no-build (scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki already had that flag, CLAUDE.md never did. CLAUDE.md keeps an 11-line pointer, because two rules must stay resident: a merge is not a deployment, and workers must never redeploy. A session learns it needs the procedure before it needs the skill, then the next line names the skill. Also updates the addendum's primary-side skill list, as this file's own rule for .claude/skills/** changes requires. CLAUDE.md: 34,442 -> 31,591 chars. Canonical block untouched — the wiki sync check still prints True. |
||
|
|
49a404ddf3 |
Merge #474 follow-up: pin main's ConfigRef wiring against the surviving mutation
My battery on the #474 merge found one survivor: reverting Fleetd.java:154 from the three-argument ConfigRef constructor to the plain two-argument one turns the live reload gate off and leaves all 1633 tests green. Both new #474 tests build their own ConfigRef with the method reference, so neither reads what main chose. This adds FleetdConfigRefWiringTest, following the three source-text precedents already in the tree (FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather than the weaker sibling pattern that builds the wiring itself. No production change. The worker branched fresh off |
||
|
|
72d6a6878b |
fleetd #474 follow-up: pin main's config wiring against M2
Fleetd.main's own choice of the three-argument ConfigRef constructor (with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation) was unpinned. Reverting Fleetd.java:154 to the plain two-argument constructor compiled with 0 errors and left the whole suite green, because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest each build their own ConfigRef directly rather than through main. Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/ FleetdCompletionResolverWiringTest precedent: asserts the exact three-argument construction is present, asserts the plain two-argument form is absent, and guards against a vacuous pass on a broken/empty source read by first asserting an unrelated anchor is present. |