bad47a84444506ac6e23f0f990b6b6b953be022f
1020 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bad47a8444 |
fleetd CI: skip unreadableFileIsUnknown honestly when root ignores the read bit
The test set the file's read bit off via setReadable(false), but on the
Gitea CI runner (root inside the container) the OS ignores that bit and
opens the file anyway, so the test asserted on a condition it never
actually created (LeadContextGaugeTest.java:142 UNKNOWN vs OK, CI run
1887 job 3104, commit
|
||
|
|
9a992d0f70 |
Merge #602: report a lead's live context usage in fleet_list
LeadContextGauge reports OK / HIGH / UNKNOWN for each lead in fleet_list, read
from the transcript Claude Code writes for itself — never from the lead's pane,
so it does not touch the control plane invariant 5 protects.
Why: a lead on this host auto-compacted 30 times in one session, discarding
roughly 250,000 tokens each time, and nothing could see it coming. The handover
feature already existed; what was missing was any way to know when to use it.
Three design points that earned their place:
- finds the transcript by NAME under <configDir>/projects/, never by deriving
the project slug, which is an undocumented Claude Code internal
- three states, not two. Every path that cannot establish a token count says
UNKNOWN with no number, so a lead is never told it is fine when the honest
answer is "I could not look"
- bounded twice: TAIL_BYTES caps bytes read, a 5s cache TTL caps how often
Includes #606, which fixed the two defects this PR's body recorded before it
was ever merged: the null configDir that made the gauge inert on this fleet,
and the torn final line that would have made it flap.
The authoring member died mid-task on a DNS error with the work uncommitted;
the lead recovered it, verified it, and committed it with that stated.
Verified by the lead on the combined state with current main merged in:
mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors,
0 skipped. LeadContextGaugeTest 9/9, FleetMcpLeadContextGaugeWiringTest 2/2,
FleetdLeadConfigDirLookupTest 6/6, FleetdLeadConfigDirSourceWiringTest 3/3.
Not deployed by this merge. A merge is not a deployment; the running daemon
still holds its old jar until redeploy.
|
||
|
|
3762aca307 |
Merge #606: wire the lead context gauge to the real configDir, and stop flapping on a torn line
Fixes the two defects recorded in #602's own body. 1. FleetMcp.contextView passed configDir=null, so the gauge read <user.home>/ .claude while this host's lead profile sets an override. Measured before the fix: 19 transcripts under the real directory, 0 under the fallback. The gauge would have deployed green and reported UNKNOWN forever, for every lead. Now threaded via FleetMcp.LeadConfigDirSource, built by Fleetd .leadConfigDirSource, following fleet.leaders.<name>.profile to that profile's configDir and reading config.get() live inside the lambda. 2. A torn final line no longer means UNKNOWN. fleet_list reads a transcript Claude Code may be mid-write on, so the last line can be cut. The old code treated that as fatal, which would make the gauge flap at random. The stated reason ("a format change should show as UNKNOWN") does not hold: a real format change makes EVERY line unparseable, and that case is still caught. Verified by the lead on the combined state with current main merged in, not on the branch alone: mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors, 0 skipped. Mutation-checked independently by the lead: Fleetd.leadConfigDirSource body -> none() -> KILLED (1 failure) FleetMcp.contextView configDir -> null -> KILLED (worker-measured) KNOWN RESIDUAL, documented rather than overclaimed. main's own one-line call to leadConfigDirSource could be swapped for none() and the suite stays green. Every member of this wiring-test family (loopHealthSource, capacitySource, healthCoverageSource) has the identical gap — no test runs Fleetd.main far enough to observe which factory it called. Filed separately as a class-wide problem rather than patched here. The empirical close is the dogfood check after redeploy. |
||
|
|
4e27bde2d7 |
fleetd #602 gauge-wiring follow-up: pin Fleetd.main's LeadConfigDirSource wiring
Extract the inline new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...)) construction in Fleetd.main into a package-private factory, Fleetd.leadConfigDirSource, mirroring loopHealthSource/capacitySource/ healthCoverageSource. Add FleetdLeadConfigDirSourceWiringTest, which calls the factory directly with real Profile/Leader fixtures and asserts the returned source resolves a real configDir -- a property that is false if the factory's body is mutated to return LeadConfigDirSource.none(). Neither FleetMcpLeadContextGaugeWiringTest nor FleetdLeadConfigDirLookupTest could catch main losing this wiring: each builds its own instance instead of calling what main calls. This closes that gap at the factory level, matching the standard already accepted for loopHealthSource's own wiring test. |
||
|
|
b85d9b0e46 |
Merge #607: redeploy-fleetd.sh no longer fails a deploy that worked
fleetd #603. The pid poll had its own fixed 10s budget and then hard-died,
while the health check right after it was allowed 60s for the same daemon.
Under launchd the java process does not exist yet when launchctl load returns,
so the script reported FAIL on a fully successful deploy — and a false FAIL in
that direction invites the hand-rolled stop/start this script exists to replace.
The pid poll now shares HEALTH_WAIT, and a miss falls through to the health
check rather than killing the run. A genuine failure still dies and still
prints the log tail. Worst-case time-to-fail roughly doubles; that cost lands
only on real failures and is the right trade.
Verified by the lead, not taken from the report. Suite exit 0, unpiped.
Two mutations run independently:
budget back to a hardcoded 10 -> KILLED (suite exit 1)
warn back to die on a pid miss -> KILLED (suite exit 1) after
|
||
|
|
856dfc6318 |
fleetd #603 review: close the untested fall-through path
PR review (comment 17358) found a real gap via mutation testing: replacing the "pid never found" warn with die "no process appeared" left the whole suite green, because neither existing test drove the case the fall-through exists for -- running_pid() never finds anything (as its own doc comment says it eventually will) while /healthz answers anyway. Adds test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives: running_pid always empty, poll_health_body succeeds. Asserts DIED_CALLED=0, the warn line is emitted, and NEW_PID stays empty (the honest "could not establish this" answer, never a guessed pid). Verified both halves myself: reverting the warn to die "no process appeared" turns this one test red (FAIL: await_daemon_started must not die...); restoring it returns the suite to green. |
||
|
|
81c1d8e91c |
fleetd #603: share HEALTH_WAIT between the pid poll and the health check
The start step gave the new process its own short, fixed 10s budget before a hard die, while the health check right after it waits a full HEALTH_WAIT (60s) for the same daemon. Under launchd, launchctl load returns before the java process exists, and on a slow host that took longer than 10s -- so the script reported "no process appeared" on a deploy that had fully succeeded. wait_for_new_pid/await_daemon_started fold the pid poll and the health check into one decision: the pid poll now shares HEALTH_WAIT instead of its own shorter budget, and a miss there falls through to the health check (direct proof the daemon is up) instead of killing the run. A genuine failure still dies, and still prints the log tail. Adds behavioural tests for both acceptance criteria (slow start succeeds, genuine failure still fails and prints the log tail) in one suite run, plus unit tests for wait_for_new_pid and a call-site test for the new function. |
||
|
|
d345b14e43 |
fleetd #602 gauge-wiring: thread a lead's configured configDir into the context gauge
FleetMcp.contextView hardcoded LeadContextGauge.read(null, ...), so a lead whose profile sets its own CLAUDE_CONFIG_DIR always read the wrong transcript directory and reported UNKNOWN forever, with no error anywhere. - Add FleetMcp.LeadConfigDirSource (same idiom as LeadSeatSource) and thread it through the constructor / listFleet overload chain / leadView / contextView. - Add Fleetd.leadConfigDirLookup, wired at construction, following the same fleet.leaders.<name>.profile link leadSeatLookup already uses, one step further to that profile's own configDir. - LeadContextGauge.parse: a single unparseable line (typically the final one, torn by a write this read raced) is now skipped rather than forcing UNKNOWN; only when every line in the read window fails to parse does it report UNKNOWN, which is the real format-change signal. - Tests: FleetdLeadConfigDirLookupTest (lookup logic), FleetMcpLeadContextGaugeWiringTest (end-to-end: config naming directory A vs B decides which is read; a lead with no configured dir degrades without throwing), and two replacement properties in LeadContextGaugeTest for the torn-line fix plus its all-unparseable control. |
||
|
|
9640deeffc |
Merge #605: fleet_list reports charterBytes alongside charterSha256
#604 item 1. A digest answers "same or different" and cannot say how much. The
byte count is the second signal, needed exactly when two hosts find they differ.
Verified by the lead before merging, from fleetd/: mvn -o clean install exit 0,
143 reports, 1821 tests, 0 failures, 0 errors, 0 skipped; SessionManagerTest 74/74.
Includes
|
||
|
|
ad3d81941f |
Correct the invariant claimed in the charterBytes comment
The comment said CharterReceipt never pairs a null digest with a non-zero byte count. It can. compose() derives the digest with digestOf(), which returns null for blank text, while the byte count is getBytes().length, which does not. A whitespace-only role charter on a profile with no MCP produces exactly that pair. No behaviour change. The gate already omits both fields on that path, which is the right answer — a size with no digest would describe an artifact we cannot fingerprint. Only the stated reason was wrong, and a false invariant in a comment is worse than no comment, because the next reader will widen the gate on the strength of it. |
||
|
|
8b986a52e0 |
#604 item 1: fleet_list reports charterBytes alongside charterSha256
CharterReceipt carries a byte count next to its digest, but the roster projection in SessionManager.rosterView only ever copied the digest across. A digest tells a lead whether two members' charters match; it cannot say how far apart they are when they don't. Report charterBytes too, nested in the same conditional as charterSha256 so the two travel together: the receipt's own contract only ever pairs a non-null digest with a real byte count, and a member with no composed charter reports charterSource alone, unchanged from before. Tests: the existing charter-receipt roster test now asserts charterBytes against the receipt's own value (not a literal), plus two new cases — no charter composed (source "none", no digest, no size) and the receipt itself absent (no charter keys at all). |
||
|
|
f336bcef39 |
CLAUDE.md: rewrap the long line my last edit left in the architects paragraph
The fleet01 lead found this. My edit in
|
||
|
|
3ed7bfca67 |
fleetd: report a lead's live context usage in fleet_list
fleetd had no way to see how full a lead's context window is. On this host a lead auto-compacted 30 times in one session, discarding roughly 250,000 tokens and costing 46s to 3m16s each time, and nothing could see it coming. LeadContextGauge reads the transcript Claude Code itself writes, never the lead's pane. It finds <sessionId>.jsonl by NAME under <configDir>/projects/ rather than deriving the project slug, which is an undocumented internal. Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot positively establish a token count reports UNKNOWN with no number, so a lead is never told it is fine when the honest answer is "I could not look". Bounded two ways: TAIL_BYTES caps bytes read off disk, and a 5s cache TTL caps how often that read happens, because fleet_list is polled constantly. Recovered by the lead: the authoring member ended on a backend error (DNS ENOTFOUND) with this work uncommitted and unpushed in its worktree. Verified before committing: mvn -o clean install exit 0, 1813 tests, 0 failures, 0 errors, 0 skipped, 144 reports; LeadContextGaugeTest 8/8. KNOWN INCOMPLETE - see the PR. The fleet_list call site passes configDir=null, which falls back to ~/.claude, but this host's lead profile sets configDir to an override. Measured: 19 transcripts under the real configDir, 0 under the fallback. The gauge is therefore INERT on this fleet until that is wired. |
||
|
|
f5c6a0e4fc |
CLAUDE.md: architects settled the two invented specifics at line 144
Both specifics in the "consult architects" paragraph were mine, not the operator's. The operator declined twice to rule on them and directed the lead to consult architects instead, so two architects on different models settled them over two rounds. "after two rounds" is gone. It was a ceiling nobody had evidence for, and it implied a counter fleetd does not have - nothing in the daemon counts rounds. The bound is now expressed as a shape: form independent positions, then compare. That is a floor of two without naming a number. The three-item operator list read as complete, so a lead hitting anything not on it would conclude it must not ask. It is now explicitly examples, and "granting access" replaces "credentials" - the case that motivated this was a forge merge refusal on a protected branch, which "credentials" covers only awkwardly. Canonical block and the wiki template updated together; sync check passes. |
||
|
|
a7aee5b982 |
Merge #600: fleetd lead-rollover outcomes readable after confirm()
Adds LeadRollover.status() and a 'status' action on fleet_handover, so a lead can find out what happened to its own roll. Every failure past confirm() was a log.warn the lead cannot read. Five states. IN_PROGRESS is written at the confirm hand-off, BEFORE the token leaves 'pending', and status() reads 'outcomes' first — so there is no window in which an in-flight roll reports UNKNOWN. Gated by the lead: mvn -o clean install exit 0, 1819 tests, 0 failures, 0 errors, 0 skipped, 143 reports; LeadRolloverTest 43, FleetMcpHandoverTest 12. Diff read in full. 0 source-text assertions in the new tests. |
||
|
|
bf895616a5 |
fleetd: distinguish an in-flight roll from an unknown token
confirm() removed the token from pending before handing the roll to the continuation, and outcomes was only written when runRollover reached an exit. For the whole duration of the roll the token was in neither map, so status() answered UNKNOWN - documented as "never issued, cancelled, or aged out". A lead polling right after its own confirm was told the roll had never been requested. Adds RollState.IN_PROGRESS, written at the confirm hand-off rather than at the roll's end, so there is no gap. status() now reads outcomes before pending, so the hand-off write cannot race the removal. Also corrects the class javadoc, which still claimed nothing calls this class. Recovered by the lead: the authoring member ended on a backend error (the host slept mid-response) with this work uncommitted in its worktree. Verified before committing: mvn -o clean install exit 0, 1819 tests, 0 failures, 0 errors, 0 skipped, 143 reports; LeadRolloverTest 43 (was 39). |
||
|
|
b874afb0af |
fleetd: make lead-rollover outcomes readable after confirm()
LeadRollover previously logged every post-confirm() failure only — a lead has no way to read the daemon log, so a roll that timed out because its own turn never settled (or /clear never re-settled) was invisible; the lead would carry on believing a fresh session was coming. Add a bounded (cap=200) token -> outcome record, written at each of the three exits in runRollover (ROLLED, TURN_NEVER_SETTLED, CLEAR_NEVER_SETTLED), and a read-only LeadRollover#status(token) accessor. The TURN_NEVER_SETTLED detail names turnSettleSeconds explicitly so a reader knows what to raise. Wire a "status" action onto the fleet_handover MCP tool (handler + schema); it never schedules, cancels, or retries anything — confirm() remains the only path that can ever cause a /clear. Extends LeadRolloverTest (33 -> 39 tests) covering the six acceptance properties, and FleetMcpHandoverTest (8 -> 12) for the new tool action. |
||
|
|
6eb34a654f |
Merge #599: fleetd #589 groups 1+2 — wiring-test 6 sites in Fleetd.main()
Extracts 6 inline constructions in Fleetd.main() (:214-:468) to package-private factories, each pinned by a behavioural wiring test. Production behaviour unchanged. Gate: test-merged onto current main (already carrying #598, which rewrote 117 lines of the same file). No conflict — the two units insert at different anchors, #599 after capacitySource and #598 after loopHealthSource, as their briefs specified. mvn exit=0, 1805 tests / 0 failures / 0 errors from 143 surefire reports; all 6 new classes confirmed to have run with their expected counts. Diff confirmed extraction-only. No source-text assertions in any of the 6 new files; that zero carries a positive control (the same pattern finds 18 such files elsewhere in the repo). Worker self-reported fixing a mutation that threw NullPointerException rather than failing an assertion — the assertNotNull guard is present ahead of the matcher call, confirmed. |
||
|
|
7084d99b89 |
Merge #597: fleetd #593 — running_pid() counts only the daemon
running_pid() filters pgrep -f hits by `comm = java` (allowlist) instead of denying a fixed list of shell names. Closes two false-positive holes: an exited pid (empty comm matched no denied name) and any non-shell wrapper (ssh, perl, python3, ruby) carrying the pattern in its own argv. Gate: merged onto current main, bash suite exit 0 and structurally identical to the baseline run on main (the one "Unattributable mutation" line is pre-existing, confirmed by running the suite on origin/main). All four running_pid tests confirmed defined AND invoked. Mutation check run by me: neutering the allowlist makes the suite exit 1 with a named failure; restore is byte-identical to baseline by git hash-object and green again. Also checked and cleared: the new die-message advice `ps -eo pid,comm,args` does NOT expose process environments on macOS — `-e` with `-o` selects all processes, it does not imply `-E`. Verified with an isolated two-phase probe and a positive control, after three earlier probes gave false positives by self-matching (the grep's own argv, and the probe script's own text). |
||
|
|
ae7845c375 |
fleetd #589 (Groups 1 & 2): pin 6 main() wiring sites with named factories
Extracts 6 inline wiring expressions from Fleetd.main() into named, directly-testable package-private static factories, following the FleetdLoopHealthSourceWiringTest (#584) shape, and adds one wiring test per factory: Group 1 (exhaustion/quarantine): - forwardingExhaustionSink(exhaustionSinkRef) — was inline ExhaustionSink.forwardingTo(exhaustionSinkRef::get) - publishExhaustionSink(...) — was two untested statements building the real sink and .set()-ing it into exhaustionSinkRef - liveExhaustedPatterns(config) — was inline new LiveExhaustedPatterns(() -> config.get().profiles()) - exhaustedPatternLookup(roster, liveExhaustedPatterns) — was an inline lambda resolving a herdr target to its profile's live pattern; silently losing this is the worst regression in the sweep, since a real usage-limit refusal would stop being classified as BACKEND_EXHAUSTED Group 2 (CB-596 credential policy): - claudeCodeLauncher(...) — was an inline `new ClaudeCodeLauncher(...)` whose memberCredentials supplier argument was untestable wiring - openCodeLauncher(...) — same, for OpenCodeLauncher Each new test pins its factory behaviorally (never via source-text assertions): built and confirmed RED by name against the named inert mutation, then confirmed GREEN again after restoring, and separately confirmed GREEN after a behavior-preserving reformat/local-variable extraction of the same call, to rule out a disguised source-text test. Suite: 1789 -> 1799 tests (+10, matching the 10 tests added), 0 failures, mvn -o clean install BUILD SUCCESS. Scope strictly limited to main()'s :214-:468 range per the ticket split with the concurrent worker handling Group 3 at line 500+. |
||
|
|
61115f6f61 |
Merge #598: fleetd #589 group 3 — wiring-test 5 sites in Fleetd.main()
Extracts 5 inline lambdas/method-refs in Fleetd.main() to package-private factories and pins each with a wiring test. Production behaviour unchanged; the releaseCleanup body moved verbatim. Gate: test-merged onto current main in a scratch worktree, mvn exit=0, 1795 tests / 0 failures / 0 errors from 137 surefire reports. Diff read in full. Worker's self-disclosed bare `git stash push` verified as recovered — all 3 surviving stash entries predate today, so no other worktree lost work. |
||
|
|
4b9ebda1b3 |
fleetd #593 CORRECTION 1: allowlist comm=java, not a denylist of shells
The round-1 fix excluded known shell names (sh/bash/zsh/dash/ksh) from running_pid()'s pgrep candidates. Two holes remained, both the same false-positive shape the ticket exists to remove: 1. A pid pgrep lists can exit before the following `ps -o comm=` lookup runs. On a gone pid, ps prints nothing, comm is empty, and an empty string matches no denied shell name -- so a dead pid was still counted. 2. The denylist only knows the shells someone thought to name. ssh, perl, python3, ruby, tail -- anything else carrying the pattern in its own argv -- was still counted alongside the real daemon. The ticket names ssh as a live route. Both close with one change: allowlist comm=java instead of denying shells. The daemon is always `java -jar target/fleetd.jar`, so its comm is always `java`; an empty comm (hole 1) is not `java` either, closing that hole for free. Answers the objection in the code comment: an allowlist can under-count if fleetd ever stops being launched by `java` (a native image, a renamed launcher). That's a false negative, the worse direction for a guard -- but it is not a new assumption: PATTERN='target/fleetd.jar' already assumes a jar run by java, and that pattern breaks before this allowlist would. Replaces the round-1 "real second process" test (which gave its exec -a standin an argv[0] holding the pattern, but not comm=java) with one that forces comm=java via `exec -a java sh -c '...'`. Adds two stubbed pgrep/ps tests pinning the two holes directly (a non-java, non-shell comm such as perl; an empty comm from an already-exited pid) -- deterministic on every platform, unlike a live-process fixture, and immune to the BSD vs Linux difference in how `comm` is derived from a fabricated process. Adds a stubbed positive backstop (comm=java is counted). Confirmed the regression is caught: reverted to the round-1 denylist, reran the suite, watched the new non-shell-comm test fail at `set -e`'s first failure, then isolated the exited-pid test separately and confirmed it also fails against the same broken code. Restored the fix and reran green. Branch merged with origin/main (3 commits: hunter role + CLAUDE.md addendum) before this commit; unrelated, no conflicts. |
||
|
|
1e68d7ee39 | Merge origin/main into worker/593-1a8025-5 | ||
|
|
6cccd458d4 |
#589 Group 3: wiring-test the 5 sites below line 500 in Fleetd.main()
Extracts the inline lambdas/method references at the 5 assigned wiring sites into named package-private factories on Fleetd, following the FleetdLoopHealthSourceWiringTest pattern from #584: - turnRegistrar(CompletionResolver) — was completion::register (Injector) - healthFailTarget(MessageService) — was messages::abandon (FleetHealthMonitor) - releaseCleanup(MessageService, ReplyInbox, PrimaryRegistry) — was the inline sessions.onRelease(detail -> {...}) cleanup lambda - replyInboxOpener() — was AmqpReplyInbox::open passed to selectReplyInbox - leadMailboxOpener() — was LeadMailbox::open passed to openLeadMailbox Each factory has a new runtime test (not source-text) that drives real collaborators through public APIs: MessageService.poll(ticket).phase(), InMemoryReplyInbox.peek(), PrimaryRegistry.nudgeTargetFor(), and the opener tests connect to a guaranteed-closed local port to prove a real network attempt vs. an inert stub. releaseCleanup was done first per the brief: MessageService.abandon's javadoc documents that losing this cleanup leaves a torn-down worker's rendezvous waiter open forever. Tests: 1789 -> 1794 (+5), 0 failures, 0 errors. mvn -q -o test exit 0, no BUILD FAILURE, no piped exit status. Each new test verified RED on the inert form named in the ticket, and GREEN after reformatting the call across lines and extracting the argument into a local/factory. |
||
|
|
42820fbe75 |
fleetd #593 (pid-count half): running_pid() no longer matches the caller
running_pid() was a bare `pgrep -f "$PATTERN"`, which matches ANY process whose full command line contains the pattern text -- including a shell that merely embeds it as literal text (a hand-typed investigation, an ssh-shaped `sh -c '...; ...'`, or a pipeline) rather than being the daemon. That self-match turns a working redeploy into a reported "racing supervisor" failure via assert_single_daemon. pgrep -c does not exist on BSD/macOS, so this can't be fixed by switching flags. running_pid() now keeps pgrep to find candidates (portable), then drops any candidate whose process name (comm) names a shell -- the daemon is always `java`, so a self-matching wrapper of this shape is always excluded while a genuine second daemon-shaped process still counts. assert_single_daemon's die message no longer hands the operator a bare `pgrep -f "$PATTERN"` as remediation -- that was exactly the self-matching invocation -- and now says in words that a pattern can match the caller. Adds three tests: a self-matching wrapper shell must be excluded, a real second daemon-shaped process must still be found, and the die message must not recommend the self-matching command. Verified the first test fails against the pre-fix implementation (confirmed the regression is caught). Leaves instance 1 (the fleetd.out log source, systemd-only) for a Linux host, per the ticket's scope split. |
||
|
|
d91ff886da |
#568 follow-up: fix the text defects the hunter-role merge introduced
Found by reading the diff at the merge gate, not reported by the worker. 1. FleetConfig.java: the operator-facing "unknown key" hint read "'fleet.architects', 'fleet.developers' or 'fleet.hunters' or 'fleet.reviewers'" — a double "or". This is text an operator reads at the moment their config is already wrong, so it should not itself be wrong. 2. MemberLifecycle.java: javadoc continuation asterisk indented 6 spaces, not 5. 3. MemberRegistry.java: javadoc asterisks moved from column 2 to column 4. 4. CallerResolver.java: a // comment indented one space past its block. 2-4 are the worker mangling alignment while widening enum lists to include HUNTER. No behaviour changes. CORRECTION to the #596 merge commit message. It claimed a fifth defect, "two javadoc lines pushed past the 100-column convention". There is no such convention in this repo: no checkstyle, no spotless, no .editorconfig, and 2975 of 32728 lines under fleetd/src/main/java already exceed 100 characters. I asserted the rule before measuring it. Those two lines are untouched. Verified: built in a scratch worktree, 1790 tests, 0 failures, 0 errors, 0 skipped, counted from the surefire XML. |
||
|
|
386e760a5c |
Merge #596: fleetd #568 — add the hunter member role
Verified by the lead before merge, not taken on the worker's report: - branch contains a639969; 1 ahead, 0 behind — clean fast-forward - CLAUDE.md change is +2 lines in the Project addendum, NOT the canonical block - canonical block sync check prints True on main and on this branch - built in a scratch worktree (never mvn clean in the main clone): 1790 tests, 0 failures, 0 errors, 0 skipped, counted from the surefire XML. Baseline 1789. - live fleetd.yaml still loads: the hunters pool is optional Five text defects found by reading the diff, not reported by the worker. They are fixed in a follow-up commit on main rather than a round trip: - FleetConfig.java operator-facing message reads "... 'fleet.developers' or 'fleet.hunters' or 'fleet.reviewers'" — a double "or" - misaligned javadoc continuation asterisks in MemberLifecycle and MemberRegistry - a misaligned // comment in CallerResolver - two javadoc lines pushed past the 100-column convention Known gap, tracked separately: fleet.hunters is absent from the live config, so a hunter cannot spawn on this host until the pool is added after the redeploy. The role ships correct and inert. |
||
|
|
2e349139e9 | #568: add hunter member role | ||
|
|
a639969a9a |
CLAUDE.md: a blocked lead consults architects, not the operator
The operator set this rule on 2026-09-19: when a decision blocks a lead, it consults one or more architect members, who are authorized to agree on one decision and unblock. The operator is not asked. Escalation stays open only for things outside the fleet's authority -- money, credentials, or a promise made to someone else. The paragraph also carries the reason the ticket record is mandatory rather than optional. The operator's old notification channel was the block itself: work stopped, so they found out. Taking the operator out of the loop removes that signal with it, so the decision goes on the ticket, which reaches them whether or not they are at a terminal when it is made. The rule has a second half aimed at architects, which lives in fleet.charters.architect and is applied per daemon -- filed as #591, because a charter can never reach a lead and charters do not travel between hosts. The missing notification event is #592. Block verified byte-identical with wiki 7-Use-Cases.md at 1d9bd1b. |
||
|
|
17c3a69c57 |
docs: CB-591 gateway page had two claims that went stale
The gateway's chat model is served under the stable alias `acoder`, and the model behind that alias changed on 2026-08-28 — it is Qwen3.8-27B now, not DeepSeek-V4-Flash. The old name is still served, so nothing broke, but it names a model this is not. Two claims on the page were wrong as a result, and both were written as current facts rather than dated measurements: - `/v1/models` returns exactly `["deepseek-v4-flash"]` — it returns 6 ids now. This sat under a heading saying it needs no re-testing. - the status banner said `local` and `gx` are both at `weight: 100` — `local` is at 0. Measured today against the live gateway: /v1/models returns acoder, qwen3.8-27b-nvfp4, deepseek-v4-flash and three embedding names; a completion sent as `deepseek-v4-flash` comes back reporting `"model": "acoder"`, which is the alias in plain sight. /v1/deployment reports generation 2026-08-28-qwen3.8-27b-nvfp4. §2 and §3 are left alone. They are the August plan, and rewriting them would destroy the record of the migration. fleetd.yaml moved to `acoder` in the same change. It is not tracked here. |
||
|
|
49a5875586 |
Merge #583: fleetd #582 — assert pending message-id cleanup at every publish cleanup site
All ten assertions proven live: six by the implementer, the last four by the lead. One contract build with four deleted removal lines produced exactly four named failures, one per site, with the total unchanged at 1825. |
||
|
|
634d33b50b | Merge worker/562-loop-health-wiring-test-99611c-5 | ||
|
|
1db79bcaa9 | Merge worker/581-completionresolver-cas-sites-0542b7-6 | ||
|
|
4507bc5a70 | Merge worker/571-attempted-outcome-5739f7-2 | ||
|
|
d7239ed23b |
fleetd #571: pin FleetMcp.formatReply's TIMED_OUT_UNCONFIRMED wording
CORRECTION 5 on the ticket: mutating the new arm's message text to the queued/working arm's text survived every existing test, because nothing asserted the specific wording. This adds one test that asserts the unconfirmed-delivery message and asserts it does NOT carry the queued/working arm's retry invitation — the distinction #571 exists for. No production code changes; formatReply's TIMED_OUT_UNCONFIRMED arm was already correct. |
||
|
|
dfeb9340b4 | fleetd #581: cover completion CAS removals | ||
|
|
1513d4f260 |
fleetd #562 follow-up: extract loopHealthSource factory, pin its wiring
PR #579's inline `new FleetMcp.LoopHealthSource(poller::health, ...)` in Fleetd.main had nothing a test could call directly. Measured: replacing poller::health with a constant () -> RUNNING compiled clean and left all 1771 tests green (see issue #562 comment "HOLD on PR #579"). Extracts the inline construction to a package-private Fleetd.loopHealthSource factory, the same style as the sibling capacitySource/healthCoverageSource factories, and adds FleetdLoopHealthSourceWiringTest with three separate assertions: the statusPoller half, the sessionReaper half, and the reaper == null branch (still STOPPED). |
||
|
|
4ca7d72303 | #582: assert pending message-id cleanup | ||
|
|
c0545d003d |
fleetd #571: make FleetApp.writeReply's inner Outcome switch exhaustive, no default
Ticket comments (17126, 17127) corrected the original acceptance criterion after this unit was already in flight: a hand-listed grep for the enum's constant names goes stale silently the moment a new constant lands, so the compiler must be the enumeration instead. sendOutcomeLabel (MessageService.java) and formatReply (FleetMcp.java) were already default-free switch expressions. The one gap was writeReply's inner "status" switch, which had `default -> "done"` — the exact value that would have lied about TIMED_OUT_UNCONFIRMED. Remove the default and list every Outcome constant explicitly; REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN get an arm too even though the outer switch always dispatches them first, so the inner switch stays exhaustive on its own. The outer switch (a statement, not an expression) keeps its own default — Java does not require exhaustiveness there regardless, and "everything not terminal is a 202" is an intentional catch-all. Verified with the proof the ticket asked for: added a scratch 11th Outcome constant after deleting all default arms and confirmed all three switch-expression sites (and no test file) fail to compile without an arm for it, one at a time, then removed the scratch constant. |
||
|
|
275ac0d251 | fleetd #562: surface loop health | ||
|
|
c1e06c9e12 |
fleetd #571: add TIMED_OUT_UNCONFIRMED so an ATTEMPTED delivery is not reported as never-arriving
MessageService.send's TimeoutException branch collapsed Injector.Cancellation.ATTEMPTED (fleetd #551 — the send call was made but its outcome is unknown) into Outcome.TIMED_OUT_QUEUED, which promises the caller the message will never arrive. On this route agent.prompt may already have pasted and submitted the text, so a caller's natural recovery (resend) risks a double delivery. Add Outcome.TIMED_OUT_UNCONFIRMED and route ATTEMPTED to it. Update the three readers found by searching for the enum's constant names (not `Outcome.`, which misses FleetMcp's unqualified `case REPLIED ->` switches and would false-positive on ConfigRef's unrelated Outcome record): - MessageService.sendOutcomeLabel: add it to the "timeout" metric label group. - FleetMcp.formatReply: its own case, warning against a blind retry (distinct from the generic "retry or poll status" message the other timeouts get). - FleetApp.writeReply: its own "unconfirmed" status and detail text, so it no longer falls through the switch's default -> "done" arm, which would have reported "the delegation completed" for the one case where delivery is unconfirmed. |
||
|
|
204da67d66 |
Merge #576: fleetd #575 — one finally covers answer()'s STALE_TURN exit
Third instance of the #572 shape on this file: one invariant kept at N sites, asserted at fewer than N. Here answer()'s inner try opened AFTER the Task lookup/registration and the STALE_TURN early return, so that return was covered only by a hand-rolled copy of the finally's cleanup pair. The fix widens the try upward and deletes the copy, so every exit runs the one finally exactly once. rendezvous.open(workerSession) correctly stays outside it — nothing to clean if it never opened. This fixed no live leak, and the code comment says so: neither rendezvous.answerAsk nor clearAsyncQuestion(turnId, false) can throw, so nothing ever left through the old gap uncovered. It is a structure fix, the ticket's own fallback case. Verified here, not taken from the worker's report: - baseline on the branch: 1766 tests, 0 failures, from Maven and from an independent sum over target/surefire-reports/*.txt. - my own mutation, located fresh: move the inner `try {` back down below the STALE_TURN return — the exact pre-fix structure, minus the hand-rolled pair. Result: 1766 run, exactly 1 failure, and it is the new test — MessageServiceTest.answerLosingTheRaceToAnAlreadyAnsweredAskStillReturnsStaleTurnAndCleansUpOnce:545 "the forward waiter this answer() call opened must be closed after a STALE_TURN return". One failure, not a crowd: the new test is the only thing holding this path. - the worker's own mutation removed the whole finally and took 8 other tests with it. That proves the finally runs; it does not prove the STALE_TURN path reaches it. Mine does. The new test hook answerAskLapseRaceHookForTest follows the file's existing askTimeoutRaceHookForTest convention. Closes #575. |
||
|
|
b091c51eee |
fleetd #575: widen answer()'s try so one finally covers its STALE_TURN exit
The waiter cleanup pair (asyncTasksByWaiter.remove + rendezvous.close) was duplicated: two sites sit in a finally, the third was hand-rolled inline before answer()'s early STALE_TURN return, structurally outside any finally. Both rendezvous.answerAsk and clearAsyncQuestion(turnId, false) are total (cannot throw), so the gap never leaked in practice. But the duplicate was untested: mutating it away left all 1765 tests green, while the two finally-protected sites are each killed by 8-36 tests. Same shape as #572. Fix: widen the try to wrap the Task registration and the STALE_TURN check, so the single finally covers every exit and the hand-rolled copy is gone. Added a race hook + regression test that deterministically reproduces the 'ask lapsed between the lookup and the unblock' case and proves the fix still returns STALE_TURN and cleans up exactly once. |
||
|
|
84d631b030 |
Merge #572: answer()'s session-lock release is pinned on all four exits
fleetd #572. MessageService releases its per-session lock in a finally at two sites. Removing the
send() one failed 22 tests and errored 1. Removing the answer() one left the whole suite green.
Skip that unlock and the thread holds the session lock forever, so every later send or answer to
that session blocks permanently — no exception, no log line.
Tests only. Against current main the diff is one file, MessageServiceTest.java, +178 lines.
Four new tests, one per exit of answer(): normal REPLIED, TIMED_OUT_WORKING, the ExecutionException
rethrow, and the InterruptedException rethrow. Each proves REACQUISITION rather than the return
value — a bounded follow-up send on the SAME session must not come back BUSY, and send reports BUSY
only when tryLock itself timed out, so a non-BUSY probe is specifically evidence the lock was free.
Lead verification, re-running rather than accepting the worker's numbers, on the branch merged with
main at
|
||
|
|
3f8c38fc54 |
Merge #567: pin LeadMailbox.inspect's probe-channel close
fleetd #567. LeadMailbox.inspect opens a probe channel and closes it in a finally. The production
code was already correct; nothing asserted it, so a future refactor could drop the close and leak an
AMQP channel per inspect() call with the suite green.
Test only. LeadMailbox.java is untouched — sha256 a2cd99be77345b7e... before and after.
The test asserts the CONSEQUENCE rather than the return value: it caps the connection at three
channels (LeadMailbox uses two, consume and publish), runs a successful inspect, then requires a
replacement channel. If the probe is left open, the broker has no channel number left and
createChannel() returns null. A test that only checked inspect()'s MailboxState would pass under the
mutation, which is the whole reason this hole existed.
Lead verification, re-running rather than accepting the worker's numbers, on the branch merged with
main at
|
||
|
|
a4dbc8f8b7 |
fleetd #572: pin answer()'s session-lock release across all four exits
MessageService.answer() releases its per-session lock in an outer finally (MessageService.java:1218) that mutation testing showed was covered but unasserted: removing that line left all 1750 existing tests green, because every existing test on this path checks answer()'s return value, never that the lock it took is actually reacquirable afterward. If it leaked, a session would be wedged forever with no exception and no log line. Adds four tests, one per exit of answer() (normal REPLIED reply, TIMED_OUT_WORKING, ExecutionException rethrow, InterruptedException rethrow), each proving the lock is reacquirable via a bounded (300ms) follow-up send on the same session rather than merely checking answer()'s own outcome. No production change. |
||
|
|
ed2fd6646a |
Merge #561: the completion/session listener fan-out survives either half throwing, and both sites are pinned
fleetd #561. Fleetd composed two TurnListener halves as `completion.X(); sessions.X();`, so a throw from the first half skipped the second. The anonymous class is now a package-private factory, Fleetd.turnListener(completion, sessions), built on two helpers that always attempt both halves and rethrow whatever escaped — a second failure attached with addSuppressed rather than dropped, so it still reaches StatusPoller's catch (Throwable). The wiring at Fleetd.java:498 calls that factory, so the seam under test is the real caller. Lead verification, on a merged tree, re-running the checks rather than accepting the worker's: exit 0, 1761 tests from Maven and from an independent sum over 131 surefire reports. The interesting part is what the first round MISSED. Two helpers maintain one invariant — "the second half always runs" — and the first round's five tests asserted it at only one site. Measured: bothMustRun reverted to the pre-fix bug -> 1 named failure (pinned) bothMustRunKeepingSecondResult the SAME bug -> 1755/1755 GREEN (unpinned) failure.addSuppressed(t) deleted -> 1755/1755 GREEN (unpinned) Both survivors are now killed by new tests, re-verified by the lead after the fix: sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction and bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions, each failing alone under its own mutation. The rule this cost us, worth carrying: COUNT ASSERTIONS PER SITE, NOT PER INVARIANT. The total being non-zero is what hides a zero at one site, and extracting a shared helper makes it worse rather than better — it does not reduce the number of sites, only how many are visible. Credit to the fleet01 lead, who predicted this shape before an instance was found. onDelivered stays deliberately unguarded. Its comment now gives the real reason — CompletionResolver.captureBaseline already catches RuntimeException around its scrape and fails open, so that half does not realistically throw — instead of the previous reason, which was true but about registration rather than about this pair. A correct conclusion resting on a wrong premise reads exactly like a verified one. |
||
|
|
5441a2b321 | fleetd #567: assert inspect closes probe channel | ||
|
|
384867dfa3 |
Merge #551: ATTEMPTED is its own cancellation answer, and the javadoc stops claiming a timed-out send never arrived
fleetd #551. The injector polls a queued entry off the queue and marks it ATTEMPTED BEFORE the irreversible AgentControl#send call, not after. So a Throwable escaping that call can never leave the entry QUEUED at the head (the #546 re-send hazard) and can never be recorded as a confident NOT_DELIVERED for text that may already be in the pane. Cancellation.ATTEMPTED is added as a third answer. NOT_DELIVERED stays reserved for confirmed absence: the readiness grace expiring, drop(), or a herdr *_not_found error, which the rest of this codebase already reads as definitely-absent rather than inconclusive. Also fixes three javadoc/comment sites that claimed a timed-out send definitely did not arrive: TIMED_OUT_QUEUED, hasQueuedDelivery, the queuedDeliveries field, and the comment in send()'s timeout branch. Two claims were wrong, not merely stale: "on every route it will not arrive later" is false on the ATTEMPTED route, and "may already hold a partial paste" understates it — agent.prompt pastes AND SUBMITS in one call, so the target may hold a complete, running turn. Verified by the lead on a merged tree: mvn -o clean install from fleetd/, exit 0, 1754 tests from Maven and from an independent sum over 130 surefire reports. Mutation: folding ATTEMPTED back into NOT_DELIVERED in cancellationOf gives 3 red, each naming the property (anErrorFromSendRemovesTheMessageAndMarksItAttempted:740, aHerdrExceptionFromSendStillSurfacesButNowReportsAttempted:781, aHerdrExceptionAfterThePasteIsNeverRecordedAsConfidentlyNotDelivered:806). The final round is comment-only, proven mechanically rather than by reading: stripping every comment from MessageService.java before and after and collapsing whitespace gives byte-identical code. |
||
|
|
f40c19ecf0 |
fleetd #551 shape sweep: fix stale ATTEMPTED-route javadoc/comments in MessageService
hasQueuedDelivery's javadoc, the queuedDeliveries field javadoc, and a code comment in send()'s timeout path all still made two claims that ATTEMPTED (fleetd #551) falsifies: a blanket "the message will not arrive later" across every route, and "the terminal may already hold a partial paste" — but agent.prompt pastes AND submits in one call, so the target may hold a complete, already-submitted turn. Each now names the three Injector.Cancellation routes (CANCELLED, NOT_DELIVERED, ATTEMPTED) and says plainly that only the first two establish the message will not arrive later. Comment/javadoc only. No behaviour change: Injector.java is untouched (sha256 0c689b6cf36275c0da45497a74b2bd4f5a3d80c4dbda77d46670c66004e53b69) and no test was added. mvn -o clean install: BUILD SUCCESS, Tests run: 1754, Failures: 0, Errors: 0 (unchanged from before this commit). |