6edeb70bc4e7fef3a4aed2003276f4777a3cbb32
1045 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6edeb70bc4 |
fleetd #612 B3 correction: distinguish lead vs member herdr in FleetdLeadRolloverAssemblyTest
Ticket comment 17553 on fleetd #612 found that the test's single shared FakeHerdr made router.leadAgents() and router.memberAgents() collapse to the identical client (FleetdAssembly.java:140-142's no-distinct-socket fallback), so a mutation swapping leadAgents() for memberAgents() at the FleetdAssembly.java:408 call site was invisible to this test even though the two are genuinely different daemons in production. Configure two distinct herdr sockets and two distinct FakeHerdr instances (the same TwoHerdrResourcePorts shape B2's FleetdAssemblyConnectionIdentityTest uses) and assert the roll's /clear + bootstrap sends land on the LEAD fake and never on the MEMBER one. Proven red against the router.memberAgents() mutation, reverted, touched, and re-run green — both outputs recorded in the PR. |
||
|
|
bc49d87cb8 | Merge remote-tracking branch 'origin/worker/fleetd-612-unita-87807e-1' into worker/612-b3-mcpwirings-da2b58-3 | ||
|
|
14b169c410 | Merge pull request 'fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair' (#626) from worker/612-b2-cb185-176d3a-2 into worker/fleetd-612-unita-87807e-1 | ||
|
|
2b52324d9a |
fleetd #612 step 2 unit B3: behavioural replacements for lead-seat, quarantine
and lead-rollover source-text guards Replaces three FleetdAssembly.java source-text guards (each scraped Fleetd.java for a call site that fleetd #612 Unit A moved into FleetdAssembly.java) with tests that drive the real assembled objects through FleetdAssembly.assembleAndStart(...) -> FleetdRuntime.mcp(), per the step-2 B-unit split (issue #612 comment 17513). - Deleted FleetdLeadSeatWiringTest (fleetd #176): pinned that FleetMcp's LeadSeatSource construction still wires Fleetd.leadSeatLookup(...) by scraping the constructor call's text. Replaced by FleetdLeadSeatAssemblyTest, which seeds one FakeHerdr tab labelled to match a configured fleet.leaders.opus.tab and asserts the REAL assembled LeadSeatSource (via runtime.mcp().leadSeatSource()) reports the live lead's seat against its own subscription profile -- 1, not the 0 LeadSeatSource.none() (the inert stand-in) could ever report. - Deleted FleetdBackendQuarantineWiringTest (fleetd #466): pinned that the escalating BackendQuarantine.withEscalation(...) text was present and the flat two-argument constructor's text was absent. Replaced by FleetdBackendQuarantineAssemblyTest, which quarantines the same credential twice through the REAL assembled BackendQuarantine (via runtime.mcp().quarantineSource().quarantine()) at controlled fake-clock offsets and asserts the second cooldown doubles (200s vs 100s) -- the one behavioural difference escalation and the flat constructor actually produce. - Deleted FleetdLeadRolloverWiringTest (fleetd #480), all three methods: unrelatedAnchorStillPresent was a scaffold anchor with no independent claim, needing no replacement. mainStillCallsTheLeadRolloverFactory pinned the leadRollover assignment's call-site text. factoryGatesOnConfigPresence pinned that an absent leadRollover: config yields no LeadRollover. Replaced by FleetdLeadRolloverAssemblyTest's two tests: assembledLeadRolloverRunsTheRealClearAndBootstrapSequence drives the REAL assembled LeadRollover (via runtime.mcp().leadRollover()) through open()/confirm() end to end and asserts /clear then bootstrapText were actually sent through the real herdr router, reaching ROLLED. absentLeadRolloverConfigMeansNoRolloverIsBuilt calls Fleetd.leadRollover(...) directly with no leadRollover: block and asserts null -- this claim was found uncovered elsewhere (LeadRolloverTest's only related assertion is vacuous, assertNull (null), and never calls the real factory). Each of the three FleetdAssembly.java call sites (quarantine line 179-180, leadRollover line 408, lead seats line 479) was mutated to its named inert variant, run against ONLY its new test (RED), reverted, touch'd (Maven mtime trap) and re-run (GREEN) -- six proven runs, pasted in the PR body. FleetMcp.java: adds three accessors (quarantineSource(), leadSeatSource(), leadRollover()) alongside the existing registeredTools() -- but public, not package-private, and this is a deliberate deviation from that precedent, not an oversight: these new assembly tests cannot live in package dev.ltms.fleet.mcp the way registeredTools()'s callers do, because they also build the ResourcePorts FleetdAssembly.assembleAndStart(...) needs, and ResourcePorts' methods return Fleetd-nested types visible only from package dev.ltms.fleet. Package-private would compile but be unreachable from there. Full mvn -o test in fleetd/: Tests run: 1879, Failures: 6 (down from the branch baseline's 1880/9 by exactly the 3 guards this unit deletes) -- the remaining 6 are FleetdCompletionResolverWiringTest (4) and FleetdConnectionIdentityConstructionTest / FleetdFleetAppConstructionTest (1 each), all out of this unit's scope (B1/B2). |
||
|
|
b6147a39f6 | Merge pull request 'fleetd #612 Unit B1: CompletionResolver behavioural test (replaces source-text guard)' (#627) from worker/612-b1-completion-457459-1 into worker/fleetd-612-unita-87807e-1 | ||
|
|
dbfc34cb6d |
fleetd #612 B2 fixup: cover the symmetric FleetApp daemon-drop mutation
Per ticket comment 17525: my second independent mutation on
FleetdAssembly.java:518 survived — new FleetApp(memberHerdr, memberHerdr, ...)
(dropping the LEAD client instead of the member one) left both existing
FleetdAssemblyFleetAppTest cases green. That is the symmetric form of the
CB-185 defect (/healthz green while the LEAD daemon is down), and the guard
this PR deletes would have caught it: its positive assertion required the
exact pair "new FleetApp(herdr, memberHerdr, workers,", which does not
survive either daemon being dropped.
Adds healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp, symmetric
to the existing member-down case.
Proven red today: reverting FleetdAssembly.java:518 to
"new FleetApp(memberHerdr, memberHerdr, workers, ..." and running only
FleetdAssemblyFleetAppTest gives Tests run: 3, Failures: 1 — the new case
fails ("expected: <503> but was: <200>", body has no "member" key); the other
two cases stay green. Reverted, git diff --stat empty, file touched, re-ran:
Tests run: 3, Failures: 0.
Full mvn -o test: Tests run: 1884, Failures: 7 (same 7 B1/B3-scope failures
as before this fixup; 1884 = 1883 + 1 new case).
|
||
|
|
a20cb96730 |
fleetd #612 Unit B1: replace CompletionResolver source-text guards with behavioural tests
FleetdCompletionResolverWiringTest read Fleetd.java's literal source text and asserted the CompletionResolver construction call still named the right arguments — proof of spelling, not behaviour. Unit A (FleetdAssembly) moved that call site out of Fleetd.main, breaking all four of its tests on a harmless relocation. FleetdCompletionResolverAssemblyTest replaces it, driving the real FleetdAssembly.assembleAndStart(...) and reading FleetdRuntime.completion() — the exact CompletionResolver instance production uses — through its public onDelivered/resolveBeforePostAction API, with a controllable ResourcePorts nanoClock in place of real sleeps. Deleted test -> what it pinned -> replacement: - worktreeBranchLookupIsStillPassedAtTheCallSite (8th constructor arg) -> assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport: a real git-worktree-provisioned MemberSession's branch must appear in a noReportMessage fallback (fleetd #241). - backendErrorArgumentsAreStillNamedAtTheCallSite, backendErrorPatternsComesFromTheFactory, backendErrorSinkComesFromTheFactory (5th/6th args) -> assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern: a configured errorPattern the built-in fallback never matches must classify as FAILED (not COMPLETION), transition the session to BACKEND_ERROR, and cool the credential off after two distinct targets within the window (fleetd #201 Unit 5). Both behaviours were proven red today by mutating FleetdAssembly.java's real call site to its inert variant (_ -> null; BackendErrorPatternLookup.legacy(); BackendErrorSink.none()), confirming the new test failed with the expected message, then reverting (touching the file to defeat Maven's stale-mtime skip) and confirming green again. FleetdAssembly.java itself is unchanged in this commit. Full suite: Tests run: 1878, Failures: 5 (down from the baseline 9 — the remaining 5 are the other in-flight workers' own guard files: FleetdBackendQuarantineWiringTest, FleetdConnectionIdentityConstructionTest, FleetdFleetAppConstructionTest, FleetdLeadRolloverWiringTest, FleetdLeadSeatWiringTest), Errors: 0, Skipped: 0. |
||
|
|
cda1a6a917 |
fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair
Deletes the two source-text guards fleetd #612 Unit A broke by moving their
scraped call sites from Fleetd.java into FleetdAssembly.java, replacing each
with a behavioural test that drives the real assembled graph instead.
- FleetdConnectionIdentityConstructionTest pinned that Fleetd.java contained
"new PaneLocator(herdr, memberHerdr)". Replaced by
FleetdAssemblyConnectionIdentityTest, which drives the real PaneLocator a
real FleetdAssembly.assembleAndStart(...) built (reached via
runtime.mcp().identity().panes(), never a copy) with two distinct FakeHerdr
daemons, and proves it finds a pane that exists on only one of them —
first the lead-only case (the CB-185 bug: a lead's own connection going
unresolvable), then the member-only case, plus a no-match control.
- FleetdFleetAppConstructionTest pinned that Fleetd.java contained
"new FleetApp(herdr, memberHerdr, workers,". Replaced by
FleetdAssemblyFleetAppTest, which binds the real Javalin app
FleetdAssembly built (runtime.app()) to a real ephemeral port and proves
GET /healthz goes 503 when only the member daemon is down — the consequence
named in the deleted test's javadoc (a down member daemon invisible behind
a healthy lead). The GET /sessions merge half of that javadoc could not be
driven the same way: it requires Authz.Action.READ, which needs a real
positive pid from the real assembly's hardcoded LsofPeerPidLookup, and an
in-process test's HTTP client and server share one JVM pid so that pid is
always -1 (fleetd #317's fail-closed rule then refuses the request before
the route, and its merge, is ever reached) — documented in the new test's
javadoc; FleetAppTwoDaemonTest remains the full behavioural proof that
FleetApp itself merges /sessions correctly given two clients.
Both replacements were proven red today: reverting FleetdAssembly.java:443-444
to "new PaneLocator(memberHerdr)" turns the identity test red (expected
"term_a", got null); reverting :518 to "new FleetApp(herdr, herdr, ..." turns
the app test red (expected 503, got 200 with no "member" key). Both mutations
reverted, tree confirmed clean, and the mutated file touched afterward so
Maven does not skip recompiling a stale .class.
Adds two small production accessors needed to reach the real objects rather
than a copy, since FleetdRuntime may not gain a field (three workers touch
that file): ConnectionIdentity#panes() exposes the PaneLocator it resolves
against, and FleetMcp#identity() exposes the ConnectionIdentity it was built
with (now also stored as a field).
Baseline on
|
||
|
|
608e4496be | Merge pull request 'fleetd #612 A-gaps: coordinator path + reportRoleFallbackGaps inside the assembly boundary' (#624) from worker/612-agaps-73a926-2 into worker/fleetd-612-unita-87807e-1 | ||
|
|
526e3b7459 |
fleetd #612 A-gaps: exercise the coordinator path and pull reportRoleFallbackGaps inside the assembly boundary
Gap 1: generalise Fleetd.LeadMailboxOpener (and FleetdRuntime's field) from the concrete LeadMailbox to a new closeable LeadChannelHandle (LeadChannel + AutoCloseable), so a test can fake the configured-coordinator path without a real broker. FleetdAssemblyCoordinatorLifecycleTest drives FleetdAssembly.assembleAndStart with a coordinator: block and a fake channel, proving the assembly builds it and shutdown closes it. Gap 2: move reportRoleFallbackGaps(cfg) and assertChartersNameOnlyRegisteredTools(cfg) out of Fleetd.main and into FleetdAssembly.assembleAndStart, immediately after cfg.validateAll(), so both run inside the tested assembly boundary before any I/O. FleetdAssemblyRoleFallbackBoundaryTest pins the moved call site behaviourally (captured log output), not by reading source text. Both new tests were verified red under a targeted mutation (a non-closing/null coordinator for gap 1; deleting the moved call for gap 2) and restored. |
||
|
|
3f7bc3815e |
fleetd #612 Unit A: extract Fleetd.main's boot composition into FleetdAssembly/FleetdRuntime
Fleetd.main kept config loading, startup reports and validation. Everything from the herdr socket connect onward moved verbatim, same order, into FleetdAssembly.assembleAndStart(AssemblyInputs, ResourcePorts), which returns a FleetdRuntime owning the real objects (package-private accessors, never a copy) and their single ordered close(). ResourcePorts/SystemResourcePorts abstract every boot-time side effect (env, herdr connect, broker openers, clocks, schedulers, shutdown-hook registration, HTTP start) with no inert production variant, per the architect proposal on the ticket. sleepHerdrPoll widened from private to package-private so FleetdAssembly can pass a method reference to it; no other signature changed. FleetdAssemblyLifecycleTest drives the real assembly with FakeHerdr, a temp FleetConfig and a fake ResourcePorts recording a start/close ledger, asserting it against the order recorded from the pre-move main() and shutdown hook, and proving every resource the ledger can observe (herdr client, three schedulers, the AMQP reply inbox) closes via FleetdRuntime.close(). FakeHerdr gained a closed flag for this. |
||
|
|
17127efb88 | Merge #601: pass auto-compact window to leads; warn instead of refusing on a conflict (CB-617) | ||
|
|
203f034528 | Merge #617: write FAILED instead of leaving a dead roll stuck at IN_PROGRESS (fleetd #615) | ||
|
|
9ee16f5b85 | Merge #616: report role-fallback gaps at boot, name contextHighNudge (fleetd #613) | ||
|
|
388ef5a3c3 |
fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
LeadRollover.runRollover made two unwrapped agents.send calls. HerdrException is unchecked, and the production continuationRunner is a bare virtual thread with no uncaught-exception handler, so a throw from either send call killed the continuation silently — confirm() had already written IN_PROGRESS into outcomes before scheduling it, and nothing ever overwrote that entry with a terminal state. Wrap the whole continuation body in one try/catch(RuntimeException), matching the local convention already used around agents.status in waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED entry naming the exception, in the same diagnostic style as TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED. Two new tests make send() throw on the /clear call and on the bootstrap-text call respectively, each asserting status(token) reports FAILED, not IN_PROGRESS. Reverting only the production catch (keeping the tests) turns both red; restoring it turns them green again. |
||
|
|
be6c45ff78 |
CB-617 review: warn instead of refuse on conflicting autoCompactWindow
rejectConflictingAutoCompactWindows threw and stopped fleetd from starting when a Claude Code profile's autoCompactWindow flag and CLAUDE_CODE_AUTO_COMPACT_WINDOW env var disagreed. Under launchd that is a restart loop, and the config that would fix it (fleetd.yaml) is gitignored, so the cause is invisible on the host where it bites (measured live: 4 profiles on this host trip it, including the lead's own profile and the one every worker spawns on). Renamed to warnConflictingAutoCompactWindows: it now logs a WARN naming each offending profile with BOTH values (autoCompactWindow=... and env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=...) instead of throwing, so the daemon starts and an operator can fix the config without reading the source. Equal values still load silently. Also reworded ClaudeCodeArguments' javadoc, which stated as fact that the env var takes precedence over the flag. That was never measured, and this host's own fleetd.yaml comment asserts the opposite — the javadoc no longer picks a side. |
||
|
|
e99cb70a8b | CB-617: pass auto-compact window to leads | ||
|
|
987ccef4c7 |
fleetd #613: log role-fallback gaps at boot, name contextHighNudge in the heartbeat line
- reportRoleFallbackGaps(cfg), called right after cfg.validateAll() in Fleetd.main, logs every MemberRole with no fleet.<role>s: pool (naming the profile count and the resolved defaultProfileFor(role) first choice) and, separately, every role with no fleet.charters.<role>: entry. Log only — the deliberate 'unconstrained' fallback in FleetConfig#candidateProfiles / CompositePeerLauncher#poolFor is unchanged, and a config with profiles: and no fleet: block still starts and still spawns. - LeadHeartbeatLoop#start()'s boot line now also names contextHighNudge (fleetd #609), alongside the three settings it already logged. - RoleFallbackGapReportTest (new) and two new LeadHeartbeatLoopTest cases pin both lines' content via a ListAppender, raising the dev.ltms.fleet logger past logback-test.xml's WARN override for the INFO-level lines. |
||
|
|
076cc43f7b |
Merge #614: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI has been red on main itself since #602/#606, on this one test, so the build has been giving no second opinion on any PR. Cause: the CI job runs in a container as root. setReadable(false) really does clear the read bit, so the test's own setup guard passes, but root opens the file anyway and the gauge correctly returns OK. The test was asserting on a condition the environment never created.
The fix adds an assumeFalse(Files.isReadable(file), ...) after the chmod and before the gauge is built, inside the existing try, so the finally still restores the bit on a skip.
Verified by me on a scratch worktree merging this onto
|
||
|
|
bad47a8444 |
fleetd CI: skip unreadableFileIsUnknown honestly when root ignores the read bit
The test set the file's read bit off via setReadable(false), but on the
Gitea CI runner (root inside the container) the OS ignores that bit and
opens the file anyway, so the test asserted on a condition it never
actually created (LeadContextGaugeTest.java:142 UNKNOWN vs OK, CI run
1887 job 3104, commit
|
||
|
|
955b9ea013 |
Merge #610: nudge an idle lead to hand over when its own context reads HIGH (fleetd #609)
Closes fleetd #609. Completes the second half of the context work: #602/#606 could detect a full lead context, and nothing acted on it. LeadHeartbeatLoop now offers a handover when the lead's own gauge reads HIGH. Never rolls a pane by itself. The nudge is text only; the lead still has to call fleet_handover, and that still needs operatorConfirmed. Verified by me on a scratch worktree merging |
||
|
|
89cb8ff79b |
fleetd #609 review: the context latch must mean the notice reached the pane
Fixes the PR #610 review blocker: LeadHeartbeatLoop committed contextNotified before injectNudge attempted the send, so a transient herdr failure marked the lead as told when nothing reached its pane, and contextNotice() carried no latch at all, so a pending-driven INJECT re-appended the notice on every tick while the context stayed HIGH. - injectNudge now reports whether agents.send succeeded and persists contextNotified only when a notice was actually included in the text and the send did not throw. The latch is split out of applyDecision (kept for idleSinceNanos/quietCount, applied unconditionally as before) so it is written on the success path only, once per branch in tick(). - contextNotice gained an overloaded 3-arg form gated on the latch as it stood before the tick's decision; the existing 2-arg form delegates to it with alreadyNotified=false, so all pre-existing callers/tests are unchanged. - tick() is now package-private (mirrors ReplyPushLoop#tick(String)) so tests can drive the real send path with a fake AgentControl instead of only the pure decide() function. - Added tests I-L covering: a failed send does not consume the notice and retries; a successful send does; the text is gated when the latch is already set; and the notice appears exactly once across three differently driven INJECTs. Both required mutations verified red and reverted: 1. Setting the latch from the Decision regardless of send outcome -> test I (iAFailedSendDoesNotConsumeTheNotice) fails. 2. Dropping the latch argument at the contextNotice call site -> tests K (kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice) and L (lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects) fail. |
||
|
|
fa62e9906d |
Merge #611: make the collected-ticket nudge test deterministic (fleetd #608)
Replaces a wall-clock bet with a manually-driven scheduler, so the tick runs only when the test runs it. Also closes a second, smaller race the brief did not name: waiting on Phase.DONE is not enough, because complete() can publish isDone() before every whenComplete dependent has run. Verified by the lead: full suite 1841/1841 green in a clean worktree, and an independent mutation (hasTicketWork forced true) diagnosed as an equivalent mutant — ReplyPushLoop.injectNudge re-reads the pending collections and returns early, so that line cannot reach agent.prompt. The worker's own mutation (ticketCollected made a no-op) is the one that reaches the observable, and it killed. |
||
|
|
d7390ccd37 |
fleetd #609 review: repair a garbled comment carried over from the brief
The brief's sentence about a null token count at HIGH was broken, and the worker copied it into the source verbatim. The code was already right; only the comment was unreadable. Says what is actually true: a HIGH reading always carries a non-null token count today, because LeadContextGauge only reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant lives in another class and nothing asserts it, so the branch stays. |
||
|
|
aa517ae0ec |
fleetd #608: make anAlreadyCollectedTicketProducesNoNudge deterministic
Replace the real ScheduledExecutorService backing ReplyPushLoop in this one test with ManualScheduler, a fake that only runs a tick when the test calls runDueTasks(). The old test bet a 300ms backoff was wide enough that collecting the ticket always won the race against the scheduler's own timer - true on an idle machine, false under a loaded full-suite run, which is exactly the flake reported. The rewritten test also waits on setAfterFinishAsyncTaskCompleteHookForTest (already used elsewhere in this file) instead of polling Phase.DONE, so it does not race CompletableFuture.complete()'s own publish-then-run-dependents gap (fleetd #399) while proving ReplyPushLoop.onTicketTerminal really ran before the ticket is collected. Verified: backoff=1 (the most hostile value) still passes; mutating ReplyPushLoop.ticketCollected to a no-op turns the test red with the same assertion message the original flake reported; three consecutive full-suite runs are green (1841/1841 each). |
||
|
|
60496831c2 |
fleetd #609: nudge an idle lead to hand over when its own context reads HIGH
LeadHeartbeatLoop can now append a text-only notice to its nudge when the lead's own LeadContextGauge reading is HIGH and leadHeartbeat.contextHighNudge is on. Fires once per HIGH stretch (a latch, cleared only by a later OK reading; UNKNOWN neither sets nor clears it), never spends the quietNudgeCap budget, and never rolls a pane itself — only the operator can approve a handover. - LeadContextGauge.Reading.unknown() widened to public for LeadContextSource.none() - FleetConfig.LeadHeartbeat gains contextHighNudge (null/false = off, unchanged default) - LeadHeartbeatLoop.decide gains context/contextNotified; Fleetd wires a new leadContextLookup/leadContextSource factory pair (LeadHeartbeatLoop.LeadContextSource) - fleetd.example.yaml documents the new key |
||
|
|
9a992d0f70 |
Merge #602: report a lead's live context usage in fleet_list
LeadContextGauge reports OK / HIGH / UNKNOWN for each lead in fleet_list, read
from the transcript Claude Code writes for itself — never from the lead's pane,
so it does not touch the control plane invariant 5 protects.
Why: a lead on this host auto-compacted 30 times in one session, discarding
roughly 250,000 tokens each time, and nothing could see it coming. The handover
feature already existed; what was missing was any way to know when to use it.
Three design points that earned their place:
- finds the transcript by NAME under <configDir>/projects/, never by deriving
the project slug, which is an undocumented Claude Code internal
- three states, not two. Every path that cannot establish a token count says
UNKNOWN with no number, so a lead is never told it is fine when the honest
answer is "I could not look"
- bounded twice: TAIL_BYTES caps bytes read, a 5s cache TTL caps how often
Includes #606, which fixed the two defects this PR's body recorded before it
was ever merged: the null configDir that made the gauge inert on this fleet,
and the torn final line that would have made it flap.
The authoring member died mid-task on a DNS error with the work uncommitted;
the lead recovered it, verified it, and committed it with that stated.
Verified by the lead on the combined state with current main merged in:
mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors,
0 skipped. LeadContextGaugeTest 9/9, FleetMcpLeadContextGaugeWiringTest 2/2,
FleetdLeadConfigDirLookupTest 6/6, FleetdLeadConfigDirSourceWiringTest 3/3.
Not deployed by this merge. A merge is not a deployment; the running daemon
still holds its old jar until redeploy.
|
||
|
|
3762aca307 |
Merge #606: wire the lead context gauge to the real configDir, and stop flapping on a torn line
Fixes the two defects recorded in #602's own body. 1. FleetMcp.contextView passed configDir=null, so the gauge read <user.home>/ .claude while this host's lead profile sets an override. Measured before the fix: 19 transcripts under the real directory, 0 under the fallback. The gauge would have deployed green and reported UNKNOWN forever, for every lead. Now threaded via FleetMcp.LeadConfigDirSource, built by Fleetd .leadConfigDirSource, following fleet.leaders.<name>.profile to that profile's configDir and reading config.get() live inside the lambda. 2. A torn final line no longer means UNKNOWN. fleet_list reads a transcript Claude Code may be mid-write on, so the last line can be cut. The old code treated that as fatal, which would make the gauge flap at random. The stated reason ("a format change should show as UNKNOWN") does not hold: a real format change makes EVERY line unparseable, and that case is still caught. Verified by the lead on the combined state with current main merged in, not on the branch alone: mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors, 0 skipped. Mutation-checked independently by the lead: Fleetd.leadConfigDirSource body -> none() -> KILLED (1 failure) FleetMcp.contextView configDir -> null -> KILLED (worker-measured) KNOWN RESIDUAL, documented rather than overclaimed. main's own one-line call to leadConfigDirSource could be swapped for none() and the suite stays green. Every member of this wiring-test family (loopHealthSource, capacitySource, healthCoverageSource) has the identical gap — no test runs Fleetd.main far enough to observe which factory it called. Filed separately as a class-wide problem rather than patched here. The empirical close is the dogfood check after redeploy. |
||
|
|
4e27bde2d7 |
fleetd #602 gauge-wiring follow-up: pin Fleetd.main's LeadConfigDirSource wiring
Extract the inline new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...)) construction in Fleetd.main into a package-private factory, Fleetd.leadConfigDirSource, mirroring loopHealthSource/capacitySource/ healthCoverageSource. Add FleetdLeadConfigDirSourceWiringTest, which calls the factory directly with real Profile/Leader fixtures and asserts the returned source resolves a real configDir -- a property that is false if the factory's body is mutated to return LeadConfigDirSource.none(). Neither FleetMcpLeadContextGaugeWiringTest nor FleetdLeadConfigDirLookupTest could catch main losing this wiring: each builds its own instance instead of calling what main calls. This closes that gap at the factory level, matching the standard already accepted for loopHealthSource's own wiring test. |
||
|
|
b85d9b0e46 |
Merge #607: redeploy-fleetd.sh no longer fails a deploy that worked
fleetd #603. The pid poll had its own fixed 10s budget and then hard-died,
while the health check right after it was allowed 60s for the same daemon.
Under launchd the java process does not exist yet when launchctl load returns,
so the script reported FAIL on a fully successful deploy — and a false FAIL in
that direction invites the hand-rolled stop/start this script exists to replace.
The pid poll now shares HEALTH_WAIT, and a miss falls through to the health
check rather than killing the run. A genuine failure still dies and still
prints the log tail. Worst-case time-to-fail roughly doubles; that cost lands
only on real failures and is the right trade.
Verified by the lead, not taken from the report. Suite exit 0, unpiped.
Two mutations run independently:
budget back to a hardcoded 10 -> KILLED (suite exit 1)
warn back to die on a pid miss -> KILLED (suite exit 1) after
|
||
|
|
856dfc6318 |
fleetd #603 review: close the untested fall-through path
PR review (comment 17358) found a real gap via mutation testing: replacing the "pid never found" warn with die "no process appeared" left the whole suite green, because neither existing test drove the case the fall-through exists for -- running_pid() never finds anything (as its own doc comment says it eventually will) while /healthz answers anyway. Adds test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives: running_pid always empty, poll_health_body succeeds. Asserts DIED_CALLED=0, the warn line is emitted, and NEW_PID stays empty (the honest "could not establish this" answer, never a guessed pid). Verified both halves myself: reverting the warn to die "no process appeared" turns this one test red (FAIL: await_daemon_started must not die...); restoring it returns the suite to green. |
||
|
|
81c1d8e91c |
fleetd #603: share HEALTH_WAIT between the pid poll and the health check
The start step gave the new process its own short, fixed 10s budget before a hard die, while the health check right after it waits a full HEALTH_WAIT (60s) for the same daemon. Under launchd, launchctl load returns before the java process exists, and on a slow host that took longer than 10s -- so the script reported "no process appeared" on a deploy that had fully succeeded. wait_for_new_pid/await_daemon_started fold the pid poll and the health check into one decision: the pid poll now shares HEALTH_WAIT instead of its own shorter budget, and a miss there falls through to the health check (direct proof the daemon is up) instead of killing the run. A genuine failure still dies, and still prints the log tail. Adds behavioural tests for both acceptance criteria (slow start succeeds, genuine failure still fails and prints the log tail) in one suite run, plus unit tests for wait_for_new_pid and a call-site test for the new function. |
||
|
|
d345b14e43 |
fleetd #602 gauge-wiring: thread a lead's configured configDir into the context gauge
FleetMcp.contextView hardcoded LeadContextGauge.read(null, ...), so a lead whose profile sets its own CLAUDE_CONFIG_DIR always read the wrong transcript directory and reported UNKNOWN forever, with no error anywhere. - Add FleetMcp.LeadConfigDirSource (same idiom as LeadSeatSource) and thread it through the constructor / listFleet overload chain / leadView / contextView. - Add Fleetd.leadConfigDirLookup, wired at construction, following the same fleet.leaders.<name>.profile link leadSeatLookup already uses, one step further to that profile's own configDir. - LeadContextGauge.parse: a single unparseable line (typically the final one, torn by a write this read raced) is now skipped rather than forcing UNKNOWN; only when every line in the read window fails to parse does it report UNKNOWN, which is the real format-change signal. - Tests: FleetdLeadConfigDirLookupTest (lookup logic), FleetMcpLeadContextGaugeWiringTest (end-to-end: config naming directory A vs B decides which is read; a lead with no configured dir degrades without throwing), and two replacement properties in LeadContextGaugeTest for the torn-line fix plus its all-unparseable control. |
||
|
|
9640deeffc |
Merge #605: fleet_list reports charterBytes alongside charterSha256
#604 item 1. A digest answers "same or different" and cannot say how much. The
byte count is the second signal, needed exactly when two hosts find they differ.
Verified by the lead before merging, from fleetd/: mvn -o clean install exit 0,
143 reports, 1821 tests, 0 failures, 0 errors, 0 skipped; SessionManagerTest 74/74.
Includes
|
||
|
|
ad3d81941f |
Correct the invariant claimed in the charterBytes comment
The comment said CharterReceipt never pairs a null digest with a non-zero byte count. It can. compose() derives the digest with digestOf(), which returns null for blank text, while the byte count is getBytes().length, which does not. A whitespace-only role charter on a profile with no MCP produces exactly that pair. No behaviour change. The gate already omits both fields on that path, which is the right answer — a size with no digest would describe an artifact we cannot fingerprint. Only the stated reason was wrong, and a false invariant in a comment is worse than no comment, because the next reader will widen the gate on the strength of it. |
||
|
|
8b986a52e0 |
#604 item 1: fleet_list reports charterBytes alongside charterSha256
CharterReceipt carries a byte count next to its digest, but the roster projection in SessionManager.rosterView only ever copied the digest across. A digest tells a lead whether two members' charters match; it cannot say how far apart they are when they don't. Report charterBytes too, nested in the same conditional as charterSha256 so the two travel together: the receipt's own contract only ever pairs a non-null digest with a real byte count, and a member with no composed charter reports charterSource alone, unchanged from before. Tests: the existing charter-receipt roster test now asserts charterBytes against the receipt's own value (not a literal), plus two new cases — no charter composed (source "none", no digest, no size) and the receipt itself absent (no charter keys at all). |
||
|
|
f336bcef39 |
CLAUDE.md: rewrap the long line my last edit left in the architects paragraph
The fleet01 lead found this. My edit in
|
||
|
|
3ed7bfca67 |
fleetd: report a lead's live context usage in fleet_list
fleetd had no way to see how full a lead's context window is. On this host a lead auto-compacted 30 times in one session, discarding roughly 250,000 tokens and costing 46s to 3m16s each time, and nothing could see it coming. LeadContextGauge reads the transcript Claude Code itself writes, never the lead's pane. It finds <sessionId>.jsonl by NAME under <configDir>/projects/ rather than deriving the project slug, which is an undocumented internal. Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot positively establish a token count reports UNKNOWN with no number, so a lead is never told it is fine when the honest answer is "I could not look". Bounded two ways: TAIL_BYTES caps bytes read off disk, and a 5s cache TTL caps how often that read happens, because fleet_list is polled constantly. Recovered by the lead: the authoring member ended on a backend error (DNS ENOTFOUND) with this work uncommitted and unpushed in its worktree. Verified before committing: mvn -o clean install exit 0, 1813 tests, 0 failures, 0 errors, 0 skipped, 144 reports; LeadContextGaugeTest 8/8. KNOWN INCOMPLETE - see the PR. The fleet_list call site passes configDir=null, which falls back to ~/.claude, but this host's lead profile sets configDir to an override. Measured: 19 transcripts under the real configDir, 0 under the fallback. The gauge is therefore INERT on this fleet until that is wired. |
||
|
|
f5c6a0e4fc |
CLAUDE.md: architects settled the two invented specifics at line 144
Both specifics in the "consult architects" paragraph were mine, not the operator's. The operator declined twice to rule on them and directed the lead to consult architects instead, so two architects on different models settled them over two rounds. "after two rounds" is gone. It was a ceiling nobody had evidence for, and it implied a counter fleetd does not have - nothing in the daemon counts rounds. The bound is now expressed as a shape: form independent positions, then compare. That is a floor of two without naming a number. The three-item operator list read as complete, so a lead hitting anything not on it would conclude it must not ask. It is now explicitly examples, and "granting access" replaces "credentials" - the case that motivated this was a forge merge refusal on a protected branch, which "credentials" covers only awkwardly. Canonical block and the wiki template updated together; sync check passes. |
||
|
|
a7aee5b982 |
Merge #600: fleetd lead-rollover outcomes readable after confirm()
Adds LeadRollover.status() and a 'status' action on fleet_handover, so a lead can find out what happened to its own roll. Every failure past confirm() was a log.warn the lead cannot read. Five states. IN_PROGRESS is written at the confirm hand-off, BEFORE the token leaves 'pending', and status() reads 'outcomes' first — so there is no window in which an in-flight roll reports UNKNOWN. Gated by the lead: mvn -o clean install exit 0, 1819 tests, 0 failures, 0 errors, 0 skipped, 143 reports; LeadRolloverTest 43, FleetMcpHandoverTest 12. Diff read in full. 0 source-text assertions in the new tests. |
||
|
|
bf895616a5 |
fleetd: distinguish an in-flight roll from an unknown token
confirm() removed the token from pending before handing the roll to the continuation, and outcomes was only written when runRollover reached an exit. For the whole duration of the roll the token was in neither map, so status() answered UNKNOWN - documented as "never issued, cancelled, or aged out". A lead polling right after its own confirm was told the roll had never been requested. Adds RollState.IN_PROGRESS, written at the confirm hand-off rather than at the roll's end, so there is no gap. status() now reads outcomes before pending, so the hand-off write cannot race the removal. Also corrects the class javadoc, which still claimed nothing calls this class. Recovered by the lead: the authoring member ended on a backend error (the host slept mid-response) with this work uncommitted in its worktree. Verified before committing: mvn -o clean install exit 0, 1819 tests, 0 failures, 0 errors, 0 skipped, 143 reports; LeadRolloverTest 43 (was 39). |
||
|
|
b874afb0af |
fleetd: make lead-rollover outcomes readable after confirm()
LeadRollover previously logged every post-confirm() failure only — a lead has no way to read the daemon log, so a roll that timed out because its own turn never settled (or /clear never re-settled) was invisible; the lead would carry on believing a fresh session was coming. Add a bounded (cap=200) token -> outcome record, written at each of the three exits in runRollover (ROLLED, TURN_NEVER_SETTLED, CLEAR_NEVER_SETTLED), and a read-only LeadRollover#status(token) accessor. The TURN_NEVER_SETTLED detail names turnSettleSeconds explicitly so a reader knows what to raise. Wire a "status" action onto the fleet_handover MCP tool (handler + schema); it never schedules, cancels, or retries anything — confirm() remains the only path that can ever cause a /clear. Extends LeadRolloverTest (33 -> 39 tests) covering the six acceptance properties, and FleetMcpHandoverTest (8 -> 12) for the new tool action. |
||
|
|
6eb34a654f |
Merge #599: fleetd #589 groups 1+2 — wiring-test 6 sites in Fleetd.main()
Extracts 6 inline constructions in Fleetd.main() (:214-:468) to package-private factories, each pinned by a behavioural wiring test. Production behaviour unchanged. Gate: test-merged onto current main (already carrying #598, which rewrote 117 lines of the same file). No conflict — the two units insert at different anchors, #599 after capacitySource and #598 after loopHealthSource, as their briefs specified. mvn exit=0, 1805 tests / 0 failures / 0 errors from 143 surefire reports; all 6 new classes confirmed to have run with their expected counts. Diff confirmed extraction-only. No source-text assertions in any of the 6 new files; that zero carries a positive control (the same pattern finds 18 such files elsewhere in the repo). Worker self-reported fixing a mutation that threw NullPointerException rather than failing an assertion — the assertNotNull guard is present ahead of the matcher call, confirmed. |
||
|
|
7084d99b89 |
Merge #597: fleetd #593 — running_pid() counts only the daemon
running_pid() filters pgrep -f hits by `comm = java` (allowlist) instead of denying a fixed list of shell names. Closes two false-positive holes: an exited pid (empty comm matched no denied name) and any non-shell wrapper (ssh, perl, python3, ruby) carrying the pattern in its own argv. Gate: merged onto current main, bash suite exit 0 and structurally identical to the baseline run on main (the one "Unattributable mutation" line is pre-existing, confirmed by running the suite on origin/main). All four running_pid tests confirmed defined AND invoked. Mutation check run by me: neutering the allowlist makes the suite exit 1 with a named failure; restore is byte-identical to baseline by git hash-object and green again. Also checked and cleared: the new die-message advice `ps -eo pid,comm,args` does NOT expose process environments on macOS — `-e` with `-o` selects all processes, it does not imply `-E`. Verified with an isolated two-phase probe and a positive control, after three earlier probes gave false positives by self-matching (the grep's own argv, and the probe script's own text). |
||
|
|
ae7845c375 |
fleetd #589 (Groups 1 & 2): pin 6 main() wiring sites with named factories
Extracts 6 inline wiring expressions from Fleetd.main() into named, directly-testable package-private static factories, following the FleetdLoopHealthSourceWiringTest (#584) shape, and adds one wiring test per factory: Group 1 (exhaustion/quarantine): - forwardingExhaustionSink(exhaustionSinkRef) — was inline ExhaustionSink.forwardingTo(exhaustionSinkRef::get) - publishExhaustionSink(...) — was two untested statements building the real sink and .set()-ing it into exhaustionSinkRef - liveExhaustedPatterns(config) — was inline new LiveExhaustedPatterns(() -> config.get().profiles()) - exhaustedPatternLookup(roster, liveExhaustedPatterns) — was an inline lambda resolving a herdr target to its profile's live pattern; silently losing this is the worst regression in the sweep, since a real usage-limit refusal would stop being classified as BACKEND_EXHAUSTED Group 2 (CB-596 credential policy): - claudeCodeLauncher(...) — was an inline `new ClaudeCodeLauncher(...)` whose memberCredentials supplier argument was untestable wiring - openCodeLauncher(...) — same, for OpenCodeLauncher Each new test pins its factory behaviorally (never via source-text assertions): built and confirmed RED by name against the named inert mutation, then confirmed GREEN again after restoring, and separately confirmed GREEN after a behavior-preserving reformat/local-variable extraction of the same call, to rule out a disguised source-text test. Suite: 1789 -> 1799 tests (+10, matching the 10 tests added), 0 failures, mvn -o clean install BUILD SUCCESS. Scope strictly limited to main()'s :214-:468 range per the ticket split with the concurrent worker handling Group 3 at line 500+. |
||
|
|
61115f6f61 |
Merge #598: fleetd #589 group 3 — wiring-test 5 sites in Fleetd.main()
Extracts 5 inline lambdas/method-refs in Fleetd.main() to package-private factories and pins each with a wiring test. Production behaviour unchanged; the releaseCleanup body moved verbatim. Gate: test-merged onto current main in a scratch worktree, mvn exit=0, 1795 tests / 0 failures / 0 errors from 137 surefire reports. Diff read in full. Worker's self-disclosed bare `git stash push` verified as recovered — all 3 surviving stash entries predate today, so no other worktree lost work. |
||
|
|
4b9ebda1b3 |
fleetd #593 CORRECTION 1: allowlist comm=java, not a denylist of shells
The round-1 fix excluded known shell names (sh/bash/zsh/dash/ksh) from running_pid()'s pgrep candidates. Two holes remained, both the same false-positive shape the ticket exists to remove: 1. A pid pgrep lists can exit before the following `ps -o comm=` lookup runs. On a gone pid, ps prints nothing, comm is empty, and an empty string matches no denied shell name -- so a dead pid was still counted. 2. The denylist only knows the shells someone thought to name. ssh, perl, python3, ruby, tail -- anything else carrying the pattern in its own argv -- was still counted alongside the real daemon. The ticket names ssh as a live route. Both close with one change: allowlist comm=java instead of denying shells. The daemon is always `java -jar target/fleetd.jar`, so its comm is always `java`; an empty comm (hole 1) is not `java` either, closing that hole for free. Answers the objection in the code comment: an allowlist can under-count if fleetd ever stops being launched by `java` (a native image, a renamed launcher). That's a false negative, the worse direction for a guard -- but it is not a new assumption: PATTERN='target/fleetd.jar' already assumes a jar run by java, and that pattern breaks before this allowlist would. Replaces the round-1 "real second process" test (which gave its exec -a standin an argv[0] holding the pattern, but not comm=java) with one that forces comm=java via `exec -a java sh -c '...'`. Adds two stubbed pgrep/ps tests pinning the two holes directly (a non-java, non-shell comm such as perl; an empty comm from an already-exited pid) -- deterministic on every platform, unlike a live-process fixture, and immune to the BSD vs Linux difference in how `comm` is derived from a fabricated process. Adds a stubbed positive backstop (comm=java is counted). Confirmed the regression is caught: reverted to the round-1 denylist, reran the suite, watched the new non-shell-comm test fail at `set -e`'s first failure, then isolated the exited-pid test separately and confirmed it also fails against the same broken code. Restored the fix and reran green. Branch merged with origin/main (3 commits: hunter role + CLAUDE.md addendum) before this commit; unrelated, no conflicts. |
||
|
|
1e68d7ee39 | Merge origin/main into worker/593-1a8025-5 | ||
|
|
6cccd458d4 |
#589 Group 3: wiring-test the 5 sites below line 500 in Fleetd.main()
Extracts the inline lambdas/method references at the 5 assigned wiring sites into named package-private factories on Fleetd, following the FleetdLoopHealthSourceWiringTest pattern from #584: - turnRegistrar(CompletionResolver) — was completion::register (Injector) - healthFailTarget(MessageService) — was messages::abandon (FleetHealthMonitor) - releaseCleanup(MessageService, ReplyInbox, PrimaryRegistry) — was the inline sessions.onRelease(detail -> {...}) cleanup lambda - replyInboxOpener() — was AmqpReplyInbox::open passed to selectReplyInbox - leadMailboxOpener() — was LeadMailbox::open passed to openLeadMailbox Each factory has a new runtime test (not source-text) that drives real collaborators through public APIs: MessageService.poll(ticket).phase(), InMemoryReplyInbox.peek(), PrimaryRegistry.nudgeTargetFor(), and the opener tests connect to a guaranteed-closed local port to prove a real network attempt vs. an inert stub. releaseCleanup was done first per the brief: MessageService.abandon's javadoc documents that losing this cleanup leaves a torn-down worker's rendezvous waiter open forever. Tests: 1789 -> 1794 (+5), 0 failures, 0 errors. mvn -q -o test exit 0, no BUILD FAILURE, no piped exit status. Each new test verified RED on the inert form named in the ticket, and GREEN after reformatting the call across lines and extracting the argument into a local/factory. |
||
|
|
42820fbe75 |
fleetd #593 (pid-count half): running_pid() no longer matches the caller
running_pid() was a bare `pgrep -f "$PATTERN"`, which matches ANY process whose full command line contains the pattern text -- including a shell that merely embeds it as literal text (a hand-typed investigation, an ssh-shaped `sh -c '...; ...'`, or a pipeline) rather than being the daemon. That self-match turns a working redeploy into a reported "racing supervisor" failure via assert_single_daemon. pgrep -c does not exist on BSD/macOS, so this can't be fixed by switching flags. running_pid() now keeps pgrep to find candidates (portable), then drops any candidate whose process name (comm) names a shell -- the daemon is always `java`, so a self-matching wrapper of this shape is always excluded while a genuine second daemon-shaped process still counts. assert_single_daemon's die message no longer hands the operator a bare `pgrep -f "$PATTERN"` as remediation -- that was exactly the self-matching invocation -- and now says in words that a pattern can match the caller. Adds three tests: a self-matching wrapper shell must be excluded, a real second daemon-shaped process must still be found, and the die message must not recommend the self-matching command. Verified the first test fails against the pre-fix implementation (confirmed the regression is caught). Leaves instance 1 (the fleetd.out log source, systemd-only) for a Linux host, per the ticket's scope split. |