d4a2cd720cf76be79f8a4ae2f2021fc5741fbdf8
840 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d4a2cd720c |
fleetd #393: deliver memberSkills to opencode members, and stop overclaiming seeding success
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as if that were success — but .claude/skills/ is a Claude Code CLI convention. opencode has no such discovery, so a kind: opencode member never actually read a seeded skill even though the log said N of M succeeded. Two changes, both required: 1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans <cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher knows both the kind and the cwd) and appends each to the generated opencode.json's instructions[] array, the same channel already used for the member charter and IDE rules. A skill folder with no SKILL.md is named and skipped rather than silently dropped. 2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills' log now says explicitly that consumption depends on the member's kind and points at the launcher's own log; OpenCodeLauncher logs its own kind-aware "skill delivery: M of N ..." line once the kind is actually known, naming any folder it could not turn into an instructions[] entry. fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code members only; an opencode member reads a different path (.opencode/agent) this key does not touch" — false as of this fix, corrected to name both kinds and how each consumes it. Tests: OpenCodeLauncherTest gains two cases driving the real GitWorktrees#add seeding path (not a hand-built fixture) into an opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in instructions[] plus the honest log line, one covering a skill folder without SKILL.md (delivered skills still flow, the malformed one is named in the log and excluded from instructions[]). ClaudeCodeLauncher is untouched — its native .claude/skills/ discovery already worked and is out of scope. mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS |
||
|
|
1fb6176783 |
Merge #457: exhaustedPattern goes hot, and the warning names the fix (fleetd #446)
Three rounds. Round 1 made exhaustedPattern a hot config key, made the quarantine warning name the fix instead of only the fact, and added model/reason to fleet_profiles' quarantined rows. Round 2 extracted the warning text into usageLimitFixWarning/usageLimitFixWarningNoModel and pinned both. Round 3 extracted the sink itself into a static exhaustionSink(...) factory and pinned what it actually logs, using a ListAppender on this class's own logger. Where the mutation numbers below come from, stated exactly. The battery ran on merge commit 3d2d521, tree |
||
|
|
49df79203c |
Merge #464: a test that charter text names only registered tools (fleetd #464)
CharterToolSurfaceTest extracts every fleet_* / bridge_* token from configured
launch charters and every tool("fleet_...") FleetMcp registers, then asserts the
first set is a subset of the second.
Verified on the merge commit. Its three acceptance criteria are met:
- Catches the ticket's own example. Fixture charter naming bridge_send, the tool
CB-634 renamed away: KILLED.
- Fails loudly with no charter text. Fixture stripped of every tool name: KILLED
by its named.isEmpty() guard, not a silent pass.
- Fails loudly with no registered tools. The tool("...") scrape broken so it
matches nothing: KILLED by its registered.isEmpty() guard.
- And one cell of my own: the server stops registering fleet_reply, which the
fixture names. KILLED. This is what proves the 'registered' half reads real
production source and is not a second fixture.
The scrape finds 11 registered tools: ack, ask, list, poll, profiles, reply,
send, spawn, status, stop, whoami. An independent count of every "fleet_x"
literal in FleetMcp.java is also 11.
WHAT THIS DOES NOT CLOSE, and it is the ticket's actual gap. The charter half is
a @TempDir fixture the test writes itself, so no charter text anyone writes can
make this test fail. Measured: the test mentions fleetd.yaml 0 times, and the
commit changes 0 production files -- FleetConfig.validateCharters() still never
reads charter text (0 lines of its body mention a tool name). So this pins the
comparison logic and acts as a rename tripwire for the two tools the fixture
names. It does not check the live config. That needs a production-side check and
is filed as a follow-up.
That residue is my ticket's fault, not the worker's: the three criteria I wrote
are exactly the three it met.
A note on my own battery, because it nearly published four false kills. The
first run showed rc=1 on all four mutation cells and I would have read that as
four kills. It was zsh: unquoted parameters are NOT word-split, so
'mvn -B $scope test' passed '-Dtest=X -DfailIfNoTests=false' as ONE argument and
surefire ran zero tests. The tell was a missing 'Tests run:' line. The rerun
proves the harness first -- selector alone must report 'Tests run: 1' -- and
every cell now prints its surefire summary count so a void cell cannot pass for
a kill.
|
||
|
|
235644c0f0 |
Merge #467: listFleet's callerIsPrimary default fails closed (fleetd #463)
A wrapper overload that is called with no callerIsPrimary argument used to
default it to true, so a forgotten argument silently handed out the lead's
coordination state. It now defaults to false: a missing identity fails closed.
Verified on the merge commit, not the branch:
- The funnel is real. FleetMcp has 7 listFleet declarations and 7 real calls
(an 8th 'listFleet(' match is a javadoc {@link}). Exactly one call writes a
literal for the new boolean, and it writes false; exactly one writes the real
predicate, coordinatorVisibleTo(principal(exchange)) in the MCP handler. No
call writes true.
- Control battery on the merge: 1584 tests green unmutated.
- M1, the fix reverted at the one line that writes the default (false -> true):
KILLED by listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey.
- M2, the half this round weakened. The worker rewrote
listIsByteForByteUnchangedForThePrimaryCaller and dropped its byte-for-byte
equality assertion, which was correct because that comparison ran against the
implicit-default overload -- now the path #463 closes. So: does anything still
notice if the primary's coordinator row silently loses a field? Dropped
heldCount: KILLED by two tests, that same test and
listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero.
- M3, a regression check on #439's own gate, 'if (callerIsPrimary)' -> 'if (true)':
KILLED by three tests.
What is no longer pinned, stated plainly: the primary's output is now checked
field by field, not as a whole string. A field that no test names could
disappear without failing anything. Every field the tests do name is pinned,
proven by M2. The whole-output answer belongs to fleetd #460.
One control in my own battery was wrong and is worth recording: I labelled
', true);' in FleetMcp.java 'must be 0' and it is 2 -- row.put("configured",
true) and m.put("self", true), neither a listFleet delegation. The pattern was
too wide. The count above comes from a walk over each declaration and call
instead.
|
||
|
|
7772b41993 |
fleetd #446 round 3: extract exhaustionSink and pin its caller (Cell A/B)
Round 2 pinned usageLimitFixWarning/usageLimitFixWarningNoModel's TEXT via FleetdUsageLimitFixWarningTest, but a mutation battery against the merged PR proved two gaps in the caller that builds main()'s real quarantine ExhaustionSink: nothing proved the sink's log.warn actually invokes either method (Cell A), and nothing proved it picks the right one for a profile with vs without a configured model: (Cell B). Extract the inline lambda into a new static Fleetd.exhaustionSink(...) factory (same refactor-for-testability class the lead approved in round 2 for the two static warning methods), behaviourally unchanged from the lambda it replaces. FleetdExhaustionSinkWarningTest drives this factory's return value directly and asserts on the real text a ListAppender attached to Fleetd's own logger captures, covering both cells from one mechanism: - Cell A (ternary result replaced by a literal string): confirmed red, 1594 run / 1 failure, restored, confirmed green (1594/0). - Cell B (ternary's two branches swapped): confirmed red, 1594 run / 2 failures (both test methods independently caught it), restored, confirmed green (1594/0). |
||
|
|
eccd0548ce |
Merge #465: state canonical invariant 5 as a purpose, not a banned tool (fleetd #458)
Verified by me on a local merge of |
||
|
|
c4e23eebad | fleetd #464: guard charter tool names | ||
|
|
7df503dfc2 |
fleetd #463: default listFleet's callerIsPrimary to false, fail closed
The compat overload at FleetMcp.java:1373 defaulted callerIsPrimary to a literal true, so a caller that forgot the argument silently got the coordinator row (this daemon's coord-id, mailbox state, held-mail previews, peer reachability) -- lead-to-lead state fleetd #439 just gated. Flip the default to false: a forgotten argument now yields a missing row instead of a leaked one. Six FleetMcpTest methods relied on the implicit true to see the coordinator row at all; they now pass true explicitly through the canonical overload. listIsByteForByteUnchangedForThePrimaryCaller's own premise (comparing the implicit-default path against an explicit-true path) was the shape of the bug, so it now only exercises the explicit-true path. Added listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey to pin the new default: a compat overload called with no callerIsPrimary argument, against a fully-configured lead channel, must produce a result with the coordinator key absent -- not empty, not redacted, absent. |
||
|
|
f5e02fedd6 |
plans: commit the fleet01 move plan, with today's state measured on the host
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.
Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:
- Phases 1 to 3 are done. Both systemd user units exist, are active and
enabled, linger is on, one java process (so the old restart.sh
double-daemon problem is gone), the live fleetd.yaml is on the host, and
the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at
|
||
|
|
29e7a06c49 |
fleetd #458: restate invariant 5 by purpose, not mechanism
Invariant 5 banned 'driving the terminal multiplexer directly', naming herdr CLI and socket as the banned tool. That bans a mechanism. What it protects is the control plane: nobody may move a fleet session, pane or peer by a route that skips the bridge's policy checks. In this repo herdr is itself the subject under test, so four contract tests must open its socket on purpose (see the addendum's 'Herdr socket tests' note). Under the old wording, a worker assigned to that code reads invariant 5 and finds its only path to finish the task banned. Restate the invariant by purpose: never move fleet state except through the bridge. herdr's CLI and socket stay as the named example of the banned route, not the definition of it. |
||
|
|
92a96fcbd8 |
Merge #462: gate the coordinator row on a named predicate the handler must consult (fleetd #439)
Verified by me on a local merge of |
||
|
|
13482872bb |
contract test: keep the loaded run's passes, discard only its failure
The javadoc threw away the whole loaded run as "not clean evidence". The
fixed cell order makes that too strong in one direction. A cell that runs
last on a climbing load has a free explanation for FAILING. It has no free
explanation for PASSING: surviving a worse condition than a fair order
would have given it is evidence in the safe direction. So the old version's
0 of 3 is still discarded, and the three passes are kept with the load each
one ran at.
Also names the cell that tests the swallow explanation head-on, which the
old text left as "nobody has managed that yet". SHELL_READY_TIMEOUT_MS at 0
types input at once — the worst case for "typed before the prompt" — and it
passed 3 of 3 at load 18.42 to 23.65. The wider read window cannot explain
that away, because a swallowed keystroke is lost, not late: the command
never runs, so no amount of polling makes its output appear.
And a warning not to carry the raw load average to another host. Load
average counts differently per core and per operating system, so only load
per core compares. I broke that rule myself when comparing this run with
another host's numbers.
Both points came from the fleet01 lead reviewing
|
||
|
|
c1ca6273fc |
fleetd #439: pin the caller at the fleet_list call site, not just the gate
Review of PR #462 found M2: the coordinatorVisibleTo gate (then an inline principal(exchange).isPrimary() check) could survive a mutation that replaced the argument with a literal true at the one production call site, because every existing test drove listFleet directly and supplied the boolean itself -- nothing exercised the handler's own call. - Name the decision: FleetMcp.coordinatorVisibleTo(Principal), a small package-private predicate next to denyFor/recordPrimarySingleton. The fleet_list handler now calls coordinatorVisibleTo(principal(exchange)) instead of inlining .isPrimary(). - Pin the predicate's role table in FleetMcpAuthzTest (onlyThePrimaryMaySeeTheCoordinatorRow), covering primary/worker/ architect and, newly, anonymous. - Add a source-reading detector at the boundary (theFleetListHandlerActuallyConsultsCoordinatorVisibleTo), same idiom as toolsTheServerRegisters/everyRegisteredToolHasItsHandlerActionPinned: it reads FleetMcp.java, isolates the listHandler block, asserts (as a control) that the block actually contains a listFleet( call, then asserts the call's trailing boolean argument is exactly coordinatorVisibleTo(principal(exchange)) -- not a literal true/false. Both mutations from the review were reproduced and killed by these tests, then reverted; see the PR body for the full break-and-restore transcript. |
||
|
|
5f1b260c81 |
addendum: the canonical sync check is the lead's, not a member's
I put "the sync script must print in sync: True" in a worker's acceptance criteria for #455. The worker could not run it and said so, honestly, instead of inventing a pass. My brief was the defect. A member's provisioned worktree has wiki/ uninitialized, so the script dies with FileNotFoundError. Measured in three worker worktrees: git submodule status printed a leading '-' and wiki/ held 0 entries. The primary's own clone printed a leading '+' and the file was there. The note is dated, gives the re-measure command, says what each outcome means, and says to delete it once it stops reproducing - as this file requires of any measurement in an addendum. Canonical block untouched: the sync script itself reports in sync: True. |
||
|
|
30d6872779 |
fleetd #446 follow-up: pin the WARNING text and the fleet_profiles model/reason fields
A mutation battery run against merged PR #457 (d02dd1b) proved criteria 2 and 3 shipped without a test that could catch them breaking: deleting either row.put("model", model) or row.put("reason", reason) in FleetMcp.profilesView left all 1586 tests green (rc=0), and renaming the "usage-limit fix:" log tag to something meaningless did too. Criterion 1's own mutation (LiveExhaustedPatterns snapshotting instead of reading live) was correctly killed by the existing FleetdExhaustionDetectionArmedWiringTest/LiveExhaustedPatternsTest — only 2 and 3 were unguarded. - Extracted the exhaustionSink WARNING text out of two inline SLF4J {}-placeholder log.warn calls into two static methods, Fleetd.usageLimitFixWarning(profile, model) and Fleetd.usageLimitFixWarningNoModel(profile, quarantineCooldownSeconds), the same extracted-static-method + dedicated-test idiom as exhaustedPatternCoverageLine/errorPatternCoverageLine. log.warn is now called with each method's return value as a single already-formatted argument, so the string a test asserts on is byte-identical to what fleetd.out receives. No behaviour change: same text, same two branches, same call site. - New FleetdUsageLimitFixWarningTest pins the leading "usage-limit fix:" grep tag and three specific facts (profile name, model name, "enabled: false" under models.allow) rather than the whole sentence, plus the no-model fallback's profile name and cooldown-seconds substitution and its explicit absence of "enabled: false". - New FleetProfilesQuarantineModelReasonFieldsTest exercises FleetMcp.QuarantineSource with a modelFor/reasonFor that actually return values (every existing test used QuarantineSource.none() or a 3-arg form defaulting both to null), asserting the quarantined row's model/reason are present when supplied and absent (not null, not blank) when modelFor/reasonFor return null or a blank string. - Each of the three: implemented, broken by hand (row.put deleted / tag renamed), confirmed the new test goes red, restored, confirmed green again. See PR body and this ticket's fleet_reply for the verbatim failure output of all three. |
||
|
|
9d1306d442 |
Merge #461: project addendum for the herdr socket carve-out (fleetd #455)
Doc-only, one file, 13 added lines, inside §Project addendum. Verified by me:
- Canonical block byte-identical: 17468 chars on main and on the branch.
- The sync script in CLAUDE.md prints "in sync: True" both before and after.
- The note's own re-measure command works. Run verbatim, it returns exactly the 4 files
the note names, and no others:
fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java
fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java
fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java
fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java
Control that the search reaches the tree: 120 test java files, 63 of them mention herdr.
That control matters here — a zero-match re-measure command would tell a future session to
delete a live restriction.
- All four required elements are present: the carve-out, what stays banned (the control plane),
who it applies to, and the perishable half (dated 2026-09-10, the command, what each outcome
means, and delete-when-stale).
- Cross-references #458 for the canonical restatement, which is deliberately not in this change.
Pre-send check applied to the note itself: a worker assigned to AgentControlContractTest can now
name one legal action that finishes its task — let the test open the herdr socket in a throwaway
workspace it tears down.
|
||
|
|
e54e3d87ea |
fleetd #439: omit fleet_list's coordinator key for non-primary callers
The coordinator row is lead-to-lead coordination state (coord-ids, mailbox facts, held-message previews). fleet_list returned it to every caller, including a worker or an architect, because coordinatorView() had no way to know who was asking. Gate at the call site inside listFleet: a new overload takes callerIsPrimary and only assembles/attaches the coordinator row when it is true, so the key is absent (not empty) for a worker or an architect. The MCP handler now passes principal(exchange).isPrimary(); every other listFleet overload keeps passing true, so callers with no caller identity (existing unit tests, the no-op wrappers) are unaffected -- confirmed by a byte-for-byte comparison test against the pre-fix overload. Authz's READ case is untouched: it stays shared by fleet_status, fleet_profiles and fleet_whoami, and the gate here is purely inside fleet_list's own result assembly. |
||
|
|
bf027f10b9 | #455: document herdr test socket exception | ||
|
|
3c5873dfe2 |
fleetd #453: point the new javadoc at the right javadoc
#456's first paragraph said "HerdrPeerLauncher's own {@link
#spawn(SpawnRequest, PlacementDecision)} javadoc". That link resolves to this
interface's own abstract declaration, not to HerdrPeerLauncher's override, so a
reader who follows it lands on the wrong text. The second paragraph already
used the plain {@code HerdrPeerLauncher.spawn(...)} form; both now match.
Also says who is actually forced to read the paragraph, because #456's
reasoning rests on it and the two cases differ. A class that implements this
interface directly must write a body for spawn(SpawnRequest,
PlacementDecision) - it is abstract here - so it reads this javadoc. A
subclass of HerdrPeerLauncher does not: HerdrPeerLauncher already implements
that method (member/HerdrPeerLauncher.java:611) and the subclass inherits the
body. For a subclass the paragraph is advice, not a gate.
javadoc -Ddoclint=reference: 5 "reference not found", the same 5 in the same 5
untouched files as origin/main at
|
||
|
|
e29227d5f4 |
Merge #456: document the override obligation on PeerLauncher.defaultProfileFor/place (fleetd #453)
Doc-only, one file, 17 added lines. Verified by me on a local merge of |
||
|
|
45aca9eb3e |
fleetd #446: make exhaustedPattern hot, name the fix in the warning, report it in fleet_profiles
The model gate can be turned off at runtime (models.allow[].enabled: false, hot since fleetd #422), but the usage-limit detector it's meant to react to was compiled once at Fleetd.main startup into a frozen Map<String,Pattern> — arming or disarming exhaustedPattern needed a daemon restart. Backwards for a feature meant to react live. - New LiveExhaustedPatterns: reads exhaustedPattern off the live config supplier per lookup (matching CompositePeerLauncher#models0's live-supplier pattern), caching compiled Pattern objects by PROFILE NAME (not pattern text — pattern text would grow unboundedly as an operator tunes a regex across reloads; profile names are bounded by the small, restart-gated set of configured profiles). patternFor() backs CompletionResolver's classification; armed() backs fleet_profiles' exhaustionDetectionArmed — both read the same object, the fleetd #404 single-accessor rule CompositePeerLauncher.modelGateState() established for the model gate. - Fleetd.java: on BACKEND_EXHAUSTED, log a WARNING naming the profile's model and the exact fix (enabled: false under models.allow, hot, no restart; remove it again once the window resets) — or, when the profile has no model: configured, say quarantine is the only thing keeping spawns off it. - FleetMcp.QuarantineSource gains modelFor/reasonFor; fleet_profiles' quarantined rows gain model/reason fields so a lead can see why without reading the daemon log. capacityView (fleet_list) intentionally untouched — scoped to fleet_profiles only. - ConfigRef/FleetConfig docs + fleetd.example.yaml updated: exhaustedPattern moves from Deferred to Hot. errorPattern stays deferred on purpose (out of scope for this ticket). - Tests: LiveExhaustedPatternsTest (new, unit-level hotness/caching proof), FleetdExhaustionDetectionArmedWiringTest (rewritten — same QuarantineSource object, read before and after a reload, asserts the answer flips with no restart), ConfigRefTest/ConfigRefProfileCoverageTest updated for the new Hot/Deferred classification. |
||
|
|
2af13ab1ff |
fleetd #449: name the decisive cell and the load the claim was measured at
The previous comment named a cause with no cell behind it. fleet01's review made the point: a confident wrong mechanism gets copied, and a confident under-determined one gets copied the same way. So the comment now names the cell that settles it. Hold the old 1000ms write sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the test goes 0 of 3 to 3 of 3. The read deadline was the whole story. It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The loaded run agreed, but its load climbed from 7 to 50 while the cells ran and the old version ran last, so it is not clean evidence and the comment says so. Comment only. No test or production code changed. |
||
|
|
0788d84be8 |
fleetd #453: document the override obligation on PeerLauncher.defaultProfileFor/place
Decision: leave both as default methods (option 1), not abstract. Neither default is a live defect today — HerdrPeerLauncher is the sole single-profile implementer and the degenerate answer (ignore role, always defaultProfile()) is correct for it. Making them abstract would force ~10 boilerplate one-line overrides across 5 unrelated PeerLauncher test doubles (NeverSpawnsLauncher, RaceLauncher, NoResumeLauncher, ClearContextSpyLauncher, LazyIdLauncher) that never call either method, for a risk that is speculative (no multi-profile HerdrPeerLauncher subclass exists or is planned). Strengthens both javadocs with an explicit MUST-override warning and cross- references HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)'s existing #450 javadoc, which already names both methods as "unoverridden here" and ties that to being a single-profile adapter -- the concrete place a future multi-profile launcher author would read, since #450 made that method abstract and any subclass must write its body. |
||
|
|
20c1094cbf |
fleetd #449: say what the timing fix actually proved, not what it assumed
The polling fix that landed in #452 is right, but its javadoc named a mechanism nobody measured: that input typed before the shell's prompt was swallowed by the shell's own startup. I mutated the settle poll away — SHELL_READY_TIMEOUT_MS = 0, so input is typed at once with no wait — and the test passed 3 of 3. So waitForText is the load-bearing half, and the proven cause is the old 800ms READ deadline, not the 1000ms write delay. The direction is the point: typing at 0ms works where typing at 1000ms failed. If early input were swallowed, 0ms would be worse than 1000ms. It is better, so the swallow explanation is unsupported. waitUntilSettled stays as cheap insurance, now labelled as insurance rather than as the fix. Comment-only; AgentControlContractTest still green. |
||
|
|
9011c59b9f |
Merge #452: run the contract tag in CI, fix the stale protocol 14 assertion (fleetd #449)
Verified on a locally built merge onto |
||
|
|
c11ad71ed0 |
Merge #448: fleet_ack errors instead of claiming success on a miss (fleetd #437)
Verified by the lead on head |
||
|
|
bdcf285265 |
Merge #451: make PeerLauncher.spawn(req, decision) abstract (fleetd #450)
Verified by the lead on head |
||
|
|
d4f93a7b13 |
fleetd #449: fix stale herdr protocol 14 javadocs/assertion, diagnose and fix the timing-raced AgentControlContractTest, select contract tests by tag in CI
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
protocol 19 (CB-521) left the client javadoc and the contract test's own
assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
assertion is real by temporarily changing the expected value to 20 (fails),
then restoring 19 (passes).
- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
failing, not skipping, on a host with a live herdr socket. Diagnosed with a
temporary instrumented run (not committed) that polled the pane every
200ms before and after sending input: the seed shell reliably takes ~2.5s
to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
raced that startup. Input typed too early was swallowed by the shell's own
startup, leaving the typed line followed by the "Restored session" banner
and no command output — indistinguishable at a glance from the env map
never reaching the shell. Once the shell was actually ready, the injected
env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
fixed sleeps with bounded polling on the actual conditions (pane text
settling, then the expected output appearing). Ran the fixed test 3x
standalone, all green.
- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
@Tag("contract") test from CI including the herdr ones above -- which is
how the stale protocol 14 assertion went unnoticed. Changed to
-Dgroups=contract, which selects the whole tagged group and picks up
future contract tests automatically.
|
||
|
|
cfebc575ea |
fleetd #450: make PeerLauncher.spawn(SpawnRequest, PlacementDecision) abstract
The default re-entered the single-argument spawn(SpawnRequest), which re-runs
checks that can refuse the profile place() just chose (#444's window). Only
CompositePeerLauncher overrode it; a future placement-doing launcher could
have inherited the wrong body silently.
Give every current implementer an explicit override, chosen by what it does:
- HerdrPeerLauncher (base of ClaudeCodeLauncher/OpenCodeLauncher, neither of
which overrides spawn(req) or place()) does no placement filtering of its
own, so it gets the re-entering form.
- CompositePeerLauncher's routing-form override is untouched.
- 5 test-fake PeerLauncher implementers (SessionManagerTest, FleetdBackendErrorSinkTest)
get overrides matching their existing spawn(SpawnRequest) shape: delegating
wrappers delegate, unreachable stubs throw, the single-profile fake re-enters.
ConfigRef does not implement PeerLauncher at all (confirmed in this tree at
|
||
|
|
822327eed5 |
Merge #447: pin the place()-to-spawn() window PlacementDecision closes (fleetd #444)
Verified by the lead on the exact tree that lands (head |
||
|
|
5289eb509f |
fleetd #437: pin the ack hit/miss contract in the AMQP contract test
AmqpReplyInboxContractTest is the one contract-group class CI actually runs, and it never asserted on ack()'s return value at all — so the exact defect this ticket fixes (reporting success for an ack that removed nothing) was unpinned in the adapter fleetd runs live. Add ackReportsHitVsMissAgainstARealBroker: a msgId never held for an owned target returns false without throwing, a real held reply returns true and is removed, and acking the same msgId again returns false. Ran against both broker modes the class supports: Testcontainers (AMQP_URI unset) and an external broker via AMQP_URI (the CI shape, using a disposable container — not the shared local LavinMQ instance). |
||
|
|
e4c703a51a |
fleetd #444: separate the adapter's fallback default from the decided profile
Review found the fixture's StubLauncher fell back to 'sol' too — the
same profile place() decides — so an UNSTAMPED request could land on
spawnCount('sol') by coincidence, and the assertion's claim that the
request 'actually carried sol' was unproven. Dropping the stamping
(SpawnRequest routedReq = req) while keeping the routing survived the
test unchanged.
Fix: give the adapter 'b' as its own fallback default instead, so an
unstamped request counts against 'b', not 'sol'. Verified both
mutations against the single test in isolation:
- drop-stamping (routedReq = req): RED, expected <sol> but was <b>
- re-entering (return spawn(req.withProfile(decision.profile())))
i.e. M1 from the first round: still RED, PlacementException
naming the now-quarantined 'sol'
Restored both; full suite green at 1575 tests.
No changes to src/main — PeerLauncher's javadoc from the first round
is unchanged.
|
||
|
|
703a05db41 |
fleetd #437: fleet_ack errors instead of claiming success on a miss
ReplyInbox.ack now returns boolean (true = removed, false = nothing to
remove) instead of void, so FleetMcp.ack can finally tell a hit from a
miss. FleetMcp.ack returns an error when the boolean is false, naming
fleet_poll{coordId} for held peer mail, which has no route through this
call. MessageService.ackReply propagates the boolean; drainReplies keeps
ignoring it (its own javadoc already documents that loss window as
deliberate). Updated the tool schema's target description to match.
Rewrote FleetMcpTest's ack tests to publish a real message before
asserting success, and added tests for a never-queued id and a coord-id
target, both now erroring. Added boolean assertions to
InMemoryReplyInboxTest's existing ack cases.
|
||
|
|
3f036b2a62 |
fleetd #444: pin the place()-to-spawn() window PlacementDecision closes
Add a test that resolves place(role) while nothing is quarantined, then quarantines the resolved profile's credential BEFORE spawning against the held PlacementDecision. CompositePeerLauncher.spawn(req, decision) must still honor the decision and land on the quarantined profile, since it never re-runs the explicit-profile enforce* checks. Verified the test kills the regression: with the override's body replaced by the re-entering spawn(req.withProfile(...)) form, this exact test goes RED with a PlacementException naming the now- quarantined profile; restored, the full suite is green (1575 tests). Also documents on PeerLauncher's default spawn(req, decision) that a launcher routing across more than one profile MUST override it, naming the four enforce* checks the default's re-entry re-applies. |
||
|
|
82fae94c55 |
Merge #445: pin every startup report call in Fleetd.main (fleetd #442)
Test written by a worker whose backend died before it could report; evidence re-run by the lead against the merged tree. Verified: merge of current main clean (0 conflicts); full build 1574 tests, 0 failures, BUILD SUCCESS, 0 compile errors; control green; deleting each of reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and reportExhaustedPatternGap from main() is KILLED by mainReportsEveryStartupGapBeforeValidationAborts. |
||
|
|
b1f34c2e6b |
fleetd #442: drop the unused java.util.List import
The new test never names List — only ListAppender, which has its own import. An unused import is an IDE warning, and this repo treats warnings as gates. No behaviour change: FleetdStartupReportTest still runs 1 test, 0 failures, BUILD SUCCESS, 0 compile errors. |
||
|
|
3f807d9f1b |
Merge #443: derive coordinator.heldDurable from queue durability + ack mode (fleetd #440)
Found by the fleet01 lead reviewing #438 after I had merged it. Verified independently before merging. The implementation choice is the load-bearing part: LeadMailbox.own() now assigns queueDeclare's durable flag and basicConsume's autoAck flag to named locals, passes those SAME locals into the two real AMQP calls (:203/:204), and derives heldDurable from them (:207). So the reported fact cannot drift from a duplicate constant - a mutation to either call's argument moves the behaviour and the report together. LeadChannel.heldDurable() is abstract, so a future implementer gets a compile error rather than a silent default. My own battery, merged tree, control green, tree restored clean: - M1 revert to the literal true -> KILLED by FleetMcpTest.listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable - M3 heldCount forced to 0 (a half the worker did not touch) -> KILLED by FleetMcpTest.listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero - M2 break the derivation itself -> SURVIVED under plain clean install, exactly as the worker reported. LeadMailbox needs a real broker, so the only test that reaches the real queueDeclare/basicConsume is @Tag("contract"), excluded from the default build. The worker ran that arm with -Pcontract and got 3 reds including its own new test. Pre-existing structural limit of this class, not introduced here, and the worker flagged it rather than hiding it. Full build, my own run: Tests run: 1573, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile errors. Not blocking, noted for a possible follow-up: one boolean over two independent facts cannot say WHICH fact was lost. The merged field is still strictly better than the literal it replaces, because it can now go false at all. |
||
|
|
e70263062c | fleetd #442: pin startup report calls | ||
|
|
1a1e586b62 |
Merge #433: carry the PlacementDecision instead of re-resolving the profile (fleetd #425)
Round 4 pins the fix. Verified independently before merging. My own battery in the worker's tree (control green, tree restored clean): - M1 revert SessionManager:677 to launcher.spawn(spawnReq) -> KILLED by SessionManagerTest.acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedThe WorktreeForUnderARotatingPolicy. This is the ticket's own deliverable and it survived 186 tests in round 3. - M3 drop the withProfile stamping in the 2-arg spawn -> KILLED by 3 tests. - M2 make the 2-arg spawn re-enter the refusing branch -> SURVIVED, but it is a near-equivalent mutant, not a gap in this work. Post-#435 all four enforce* conditions are ones the routing branch already filtered on, so the two paths differ only if placement state moves between place() and spawn(). Filed separately. Full build, my own run this turn, whole log redirected and grepped: Tests run: 1572, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile errors. Branch already contains current main. Read the src/main diff. The new test uses roundRobin() (stateful) and asserts agreement between the overlay profile and the spawned profile, rather than a hardcoded name, which is the right shape - select() is stateful, so two calls disagree by design. |
||
|
|
c16d118f09 |
fleetd #440: derive coordinator.heldDurable from queue durability + ack mode
FleetMcp.coordinatorView wrote heldDurable as a literal true, so a change that broke either the durable queue declare or the manual-ack consume in LeadMailbox would leave the field, and the full suite, green. - LeadChannel gets a new heldDurable() method: the conclusion of a durable queue declare AND a manual-ack consumer, derived by the implementation from what it actually did, never asserted. - LeadMailbox.own() captures the exact booleans it passes to queueDeclare/basicConsume and stores their conjunction; heldDurable() returns it. - FleetMcp.coordinatorView now reads channel.heldDurable() instead of a literal; updated the javadoc to say where the fact comes from. - FakeLeadChannel gets a heldDurable field (default true) + withHeldDurable setter so FleetMcpTest can prove the field goes false. - FleetMcpTest: new test asserts heldDurable:false when the channel says so. - LeadMailboxTest (contract, real broker): new test asserts heldDurable() true against a real LeadMailbox. Verified by hand that flipping own()'s autoAck local to true turns this test (and two pre-existing redelivery tests) red, and restoring it turns them green again. |
||
|
|
4b10d02207 |
fleetd #425 rework round 4: mutation-pinning test for the dropped PlacementDecision
SessionManager.acquireWithWorktree's unqualified branch must carry the PlacementDecision it already resolved via launcher.place() into launcher.spawn(spawnReq, decision) rather than re-deriving it through a blank-profile launcher.spawn(spawnReq). Every existing test in this file uses PlacementPolicies.fixed(), which answers select() the same way on every call, so dropping the decision (handle = launcher.spawn(spawnReq);) was invisible: 186 tests stayed green under that mutation. acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy uses PlacementPolicies.roundRobin() instead — deterministic AND stateful, so two select() calls on the same policy instance disagree (index 0 then index 1 across a two-profile pool). It asserts AGREEMENT between the profile the worktree's parity overlay was provisioned for and the profile the member actually spawned on, never a hardcoded expected profile name. Verified as a real mutation, not a no-op: applying the exact mutation (handle = launcher.spawn(spawnReq);) turns it red — expected [b.mcp.json] but was [a.mcp.json] — and reverting turns it green again. Full build: 1572 tests, 0 failures, 0 errors, BUILD SUCCESS. |
||
|
|
c5fbfdbf4a | Merge main into #425 rework branch (brings #438 held-peer-mail read) | ||
|
|
12cff28abb |
Merge #438: let a lead read its own held peer mail, primary-only (fleetd #421)
fleet_poll{coordId} peeks this daemon's own held lead-to-lead mail and
returns full bodies without acking. New Authz.Action COORD_READ, primary
only — not the architect, which holds READ today. pollAction is now
argument-derived over both target and coordId.
heldView and HELD_PREVIEW_MAX_CHARS are untouched: fleet_list stays a
cheap always-safe scan, and the full read is a separately authorized call.
Verified by me, not taken from the report: merged tree builds 1562 green
(main 1555 + 7 new methods), 0 compile errors, no merge conflicts.
Mutation battery on lines the worker did NOT mutate, control 111 green:
- COORD_READ widened to the architect KILLED (AuthzTest + FleetMcpAuthzTest)
- self-coord-id guard removed KILLED (FleetMcpTest)
- full body swapped for the preview KILLED (FleetMcpTest)
The third mutation targets the ticket's own deliverable, and it is pinned.
Authz.permits has no default, so a new action is a compile error rather
than a silently unhandled case.
Follow-up filed as #439: fleet_list's coordinator row is still READ-gated,
so a worker sees peer coord-ids and 80-char previews of lead-to-lead
bodies. Pre-existing; the implementer flagged it and left it alone.
|
||
|
|
77a6a7142e | Merge main into #421 branch (brings #434 model-gate observability and #436 fixed-placement cap) | ||
|
|
9f3671b801 |
fleetd #425 rework round 3: rewrite prose after #435 made fixed honour maxLoad
fleetd #435 (merged to main) made FixedPlacementPolicy evaluate maxLoad during automatic selection, the same way weighted/round-robin already did. Six comments across CompositePeerLauncher.java, PeerLauncher.java, PlacementDecision.java, and SessionManager.java justified round 2's place()/spawn(req, decision) mechanism by saying fixed "deliberately never evaluates maxLoad" — that claim is now false, and needed restating, not just deleting. The honest case after #435: the two-path shape (routing branch falls through an excluded candidate; explicit-profile branch refuses on it) is still real and still deliberate — an operator who names a profile should get a refusal, not a silent substitution. What round 1 got wrong, and what round 2 still needs to prevent, is turning a fall-through into a refusal by accident: resolving a name via place() and then feeding it back to spawn(SpawnRequest) as an explicit profile. Before #435 that accident was reachable through maxLoad specifically, because fixed never evaluated it; #435 closed that specific gap, so a PlacementDecision can no longer be at-cap in the first place. What survives as the justification for spawn(req, decision): it never re-evaluates a condition place() already decided, and it closes the window between that decision and the spawn in which the underlying state could otherwise move — not a failure #435 already prevents. Re-measured the sibling paths this round exists to keep in agreement (one profile at maxLoad: 1, liveCount pinned at 1, PlacementPolicies.fixed(), unqualified spawn): both the with-worktree and without-worktree paths now throw the identical PlacementException — "worker profile 'a' is at maxLoad (1 live >= 1 cap), and no available candidate remains" — closed upstream by #435, at place()/select(), before either path ever reaches a spawn call. The observable asymmetry this PR was filed to fix is gone; what remains is the structural argument above. No behavior change: place()/PlacementDecision/spawn(req, decision) are untouched, and FixedPlacementPolicy/PlacementPolicyUtil are taken wholesale from main's merge. |
||
|
|
84034b34d1 | Merge remote-tracking branch 'origin/main' into worker/425-rework-placement-resolve-c58ba1-9 | ||
|
|
1e9b2c9b7e |
fleetd #421: let a lead peek its own held peer mail, primary-only
fleet_list truncated held lead-to-lead messages to an 80-char preview with
no way to read the full body, and fleet_poll{target} drained the wrong
inbox (a worker's reply queue, not the coordinator mailbox) -- it silently
returned []. fleet_ack would have destroyed the message unread.
Add a non-destructive read: fleet_poll{coordId} peeks (never acks) this
daemon's own held mail via LeadChannel.peek(). The coordId must equal the
caller's own selfCoordId -- passing a peer's id is refused with a reason,
instead of repeating the original silent-[] confusion.
This is authorization-sensitive: mapping it to the existing READ action
would let any worker read every peer lead's mail in full. READ's openness
rests on "the roster carries no secrets" (Authz.java), which does not hold
for lead-to-lead coordination bodies. Added Authz.Action.COORD_READ,
primary-only (not even the architect, which holds READ today), and made
pollAction's signature depend on both target and coordId so every call
site states explicitly what it passes.
Also fixes fleet_list's "pending: 0" trap: mailbox.pending only counts
broker-ready messages, so a healthy held mailbox reads as empty. Added
heldCount/heldDurable beside held[] so the durability fact isn't implied
only by reading the code.
Mutation-tested: pollAction's COORD_READ->READ mapping, the peek->ack
substitution, the 80-char preview cap widened to 81, and the Authz case
widened to include caller.isWorker() -- each breaks exactly its matching
test and nothing else. The first attempt at the preview-cap test used a
homogeneous "x"*200 body, which a widened cap slipped through unnoticed
(contains() found a shifted match); replaced with a sentinel character at
index 80 to actually pin the boundary.
Updates CLAUDE.md's intent->tool table for fleet_poll's new coordId
semantics, per this repo's own "prompt is part of the product" rule.
wiki/ is a submodule and not committable from a worker's worktree --
wiki-bound content is in the PR body instead.
|
||
|
|
5d422f85fa |
Merge #436: make fixed placement honour maxLoad (fleetd #435)
fixed was the one automatic policy that ignored maxLoad, and it is the default for an absent placement: key. It now gates on the same shared atCap predicate weighted and round-robin use, and falls through to the next candidate rather than refusing. Verified here: 1555 tests green (main was 1548, +7 new methods), 0 compile errors. Mutation battery, control 113 green: - atCap boundary >= -> > KILLED (15 tests, across all 3 policies) - drop cap check in fixed walk KILLED (3 tests) - weightExcluded -> false KILLED (2 pre-existing tests) Neither live host is affected today: both run placement: weighted (Mac fleetd.yaml:210, fleet01 fleetd.yaml:172), and weighted already skipped at-cap candidates via PlacementPolicyUtil.available. This aligns fixed with the other policies and with maxLoad's documented contract. |
||
|
|
ed2027b202 |
fleetd #435: make FixedPlacementPolicy honor maxLoad
FixedPlacementPolicy (the default placement policy) never consulted maxLoad, so an at-cap default was chosen anyway on every unqualified spawn -- the cap was advisory, not enforced, for the one policy every config uses by default. weighted/round-robin already gated on it via PlacementPolicyUtil.available(). Extract the "at cap" predicate into PlacementPolicyUtil.atCap(ctx, c) so all three policies share one definition, and consult it at both of FixedPlacementPolicy's filter sites (the default fast path and the candidate walk), mirroring the existing weightExcluded pattern. An at-cap default now falls through to the next candidate instead of refusing the spawn -- only when every candidate is unusable does the policy still throw, naming the cap in the message. Update the class javadoc (five exceptions -> six) and the reason-priority comments to match CompositePeerLauncher's explicit-spawn order (quarantine, cooling off, max load, model-off). |
||
|
|
6b0a99b2b7 |
fleetd #425 rework round 2: stop routedProfileFor's caller re-entering the throwing branch
Round 1 closed quarantine/cool-off/model-off routing for acquireWithWorktree by resolving the profile through routedProfileFor(role) and handing that name back to launcher.spawn(SpawnRequest) as an EXPLICIT profile. That re-resolution has a cost the lead measured directly: naming a profile explicitly makes CompositePeerLauncher.spawn take its THROWING branch (enforceMaxLoad included), while the routing branch a blank spawn takes never calls enforceMaxLoad at all, and FixedPlacementPolicy (the default) deliberately never evaluates maxLoad during automatic selection. So an at-cap pool-first profile that placement itself would have picked for a plain unqualified spawn could die at enforceMaxLoad one call later, purely because the worktree path's route to the spawn passed through an explicit profile name — a new failure a worktree-less unqualified spawn never hits. This closes the two-path shape instead of moving it: PeerLauncher gains place(role), returning an opaque PlacementDecision, and spawn(req, decision), which honors that decision through the SAME routing branch a blank spawn uses — no enforce* check is newly applied. SessionManager. acquireWithWorktree now keeps the PlacementDecision from place() and hands it to spawn(req, decision) for an unqualified request, instead of re-resolving through an explicit profile name. An explicitly-named profile is unaffected: it still goes through spawn(req) and its throwing branch, exactly as before. Also corrects the acquireWithWorktree comment's false claim that round 1 "loses nothing else" — maxLoad was lost too, as a new hard failure, not a retry. The comment now names it explicitly. Kept the four round-1 tests (still pass — routedProfileFor now just delegates to place()). Added one class asserting the invariant itself: an unqualified spawn on a maxLoad-capped profile must land the same outcome with and without a worktree, asserting on the pair rather than a hardcoded direction, so it stays correct however fleetd #435 (not this ticket) resolves whether maxLoad should gate an unqualified spawn at all. |