Compare commits

...

21 Commits

Author SHA1 Message Date
Dai Ha 9e4e423ad6 fleetd #393 follow-up: remove the instructions[] writer-ordering hazard
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 2m11s
OpenCodeLauncher.writeConfig has three writers into the instructions[]
array (charter, seeded skills, IDE rules). The charter writer used
putArray (create-or-REPLACE) instead of withArray (get-or-create), which
"worked" only because it happened to run first against a still-empty
array — an undeclared ordering dependency nothing tested. Found by the
fleet01 lead and verified on this branch's merge: flipping the skills
writer to putArray left the full 1603-test suite green while silently
deleting the charter entry, which would launch an opencode member with
no role contract at all.

Fix: charter's putArray -> withArray (one-word change, behavior-identical
today). Add three tests asserting instructions[] CONTENT as an exact
ordered list (not size) across writer combinations: charter only,
charter + IDE rules, and charter + IDE rules + seeded skills. Mutation
testing (see PR body) shows the skills and IDE-rules writers are each
independently detectable by name; the charter writer's own mutation is
not detectable by any test, because it structurally always runs first
against an empty array, so putArray and withArray are equivalent there.
2026-09-10 19:58:25 +07:00
Dai Ha d4a2cd720c fleetd #393: deliver memberSkills to opencode members, and stop overclaiming seeding success
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m46s
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every
provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as
if that were success — but .claude/skills/ is a Claude Code CLI convention.
opencode has no such discovery, so a kind: opencode member never actually
read a seeded skill even though the log said N of M succeeded.

Two changes, both required:

1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans
   <cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher
   knows both the kind and the cwd) and appends each to the generated
   opencode.json's instructions[] array, the same channel already used for
   the member charter and IDE rules. A skill folder with no SKILL.md is
   named and skipped rather than silently dropped.

2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills'
   log now says explicitly that consumption depends on the member's kind and
   points at the launcher's own log; OpenCodeLauncher logs its own kind-aware
   "skill delivery: M of N ..." line once the kind is actually known, naming
   any folder it could not turn into an instructions[] entry.

fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code
members only; an opencode member reads a different path (.opencode/agent)
this key does not touch" — false as of this fix, corrected to name both
kinds and how each consumes it.

Tests: OpenCodeLauncherTest gains two cases driving the real
GitWorktrees#add seeding path (not a hand-built fixture) into an
opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in
instructions[] plus the honest log line, one covering a skill folder
without SKILL.md (delivered skills still flow, the malformed one is named
in the log and excluded from instructions[]). ClaudeCodeLauncher is
untouched — its native .claude/skills/ discovery already worked and is out
of scope.

mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
2026-09-10 19:37:00 +07:00
Dai Ha 1fb6176783 Merge #457: exhaustedPattern goes hot, and the warning names the fix (fleetd #446)
CI / contract (push) Successful in 1m31s
CI / build (push) Successful in 1m33s
Three rounds. Round 1 made exhaustedPattern a hot config key, made the
quarantine warning name the fix instead of only the fact, and added model/reason
to fleet_profiles' quarantined rows. Round 2 extracted the warning text into
usageLimitFixWarning/usageLimitFixWarningNoModel and pinned both. Round 3
extracted the sink itself into a static exhaustionSink(...) factory and pinned
what it actually logs, using a ListAppender on this class's own logger.

Where the mutation numbers below come from, stated exactly. The battery ran on
merge commit 3d2d521, tree 1dcca23: origin/main at 235644c plus this branch.
This merge commit's tree is 953ce11, which is that tree plus one file --
CharterToolSurfaceTest, from the #464 merge (49df792) that landed on main while
the battery was running. So the battery did not run on this exact tree. That one
added file was verified green on its own merge. The build on THIS tree, tree
953ce11, is: Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 0 compile-error blocks. That is the battery tree's 1600 plus that one
test, which is the arithmetic the two trees predict. Saying this rather than
implying one tree.

On the battery tree: 1600 tests green, 0 compile-error blocks. FleetMcp.java
auto-merged there against #463's change to the same file, so that build was also
the gate on the auto-merge -- a clean auto-merge is not a compiling merge.

- M5, the round-2 survivor reproduced verbatim: the call site stops using either
  pinned method and logs a literal instead. Round 2 left this green across 1592
  tests. Now KILLED by
  FleetdExhaustionSinkWarningTest.profileWithModelGetsTheActionableFix and
  .profileWithNoModelGetsTheFallbackNotAFix.
- M6, the ternary's two branches swapped, so a profile WITH a model gets the
  no-model text and vice versa. Both extracted methods and both their unit tests
  untouched. KILLED by the same two tests, independently.
- M7, THE HALF THIS ROUND DID NOT PIN, and it SURVIVED. main()'s call to the
  factory replaced with an inert lambda: the factory and its test stay perfect,
  the daemon quarantines nothing and logs nothing. 1600 green.

M7 is the same defect shape as round 2, moved one level up, and it is worth
naming plainly. Extracting a thing in order to pin it CREATES the seam the test
then lands on. Round 2 extracted the warning TEXT, pinned the two methods, and
left the call site that selects between them unpinned -- M5. Round 3 extracted
the SINK, pinned the factory's behaviour, and left main()'s wiring of it
unpinned -- M7. The question to ask on any such fix is which of three things a
test now reaches: the value, the call site, or the selection between values.
Extraction only ever answers the first.

M7 is not a reason to hold this merge. Proving main() wires this factory means
driving daemon startup, which nothing here does -- that is fleetd #460. A
cheaper option exists and is recorded there: a source-reading assertion of the
kind #439 used, which would kill M7 without starting the daemon.

Round 3's own caveat, disclosed by the worker and worth keeping: its first Cell A
run was corrupted by a second concurrent mvn against the same module directory.
It killed that run, checked for stray ForkedBooter processes, and re-ran. The
numbers above are mine, from this battery, not that run.
2026-09-10 19:16:24 +07:00
Dai Ha 49df79203c Merge #464: a test that charter text names only registered tools (fleetd #464)
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m56s
CharterToolSurfaceTest extracts every fleet_* / bridge_* token from configured
launch charters and every tool("fleet_...") FleetMcp registers, then asserts the
first set is a subset of the second.

Verified on the merge commit. Its three acceptance criteria are met:

- Catches the ticket's own example. Fixture charter naming bridge_send, the tool
  CB-634 renamed away: KILLED.
- Fails loudly with no charter text. Fixture stripped of every tool name: KILLED
  by its named.isEmpty() guard, not a silent pass.
- Fails loudly with no registered tools. The tool("...") scrape broken so it
  matches nothing: KILLED by its registered.isEmpty() guard.
- And one cell of my own: the server stops registering fleet_reply, which the
  fixture names. KILLED. This is what proves the 'registered' half reads real
  production source and is not a second fixture.

The scrape finds 11 registered tools: ack, ask, list, poll, profiles, reply,
send, spawn, status, stop, whoami. An independent count of every "fleet_x"
literal in FleetMcp.java is also 11.

WHAT THIS DOES NOT CLOSE, and it is the ticket's actual gap. The charter half is
a @TempDir fixture the test writes itself, so no charter text anyone writes can
make this test fail. Measured: the test mentions fleetd.yaml 0 times, and the
commit changes 0 production files -- FleetConfig.validateCharters() still never
reads charter text (0 lines of its body mention a tool name). So this pins the
comparison logic and acts as a rename tripwire for the two tools the fixture
names. It does not check the live config. That needs a production-side check and
is filed as a follow-up.

That residue is my ticket's fault, not the worker's: the three criteria I wrote
are exactly the three it met.

A note on my own battery, because it nearly published four false kills. The
first run showed rc=1 on all four mutation cells and I would have read that as
four kills. It was zsh: unquoted parameters are NOT word-split, so
'mvn -B $scope test' passed '-Dtest=X -DfailIfNoTests=false' as ONE argument and
surefire ran zero tests. The tell was a missing 'Tests run:' line. The rerun
proves the harness first -- selector alone must report 'Tests run: 1' -- and
every cell now prints its surefire summary count so a void cell cannot pass for
a kill.
2026-09-10 19:10:56 +07:00
Dai Ha 235644c0f0 Merge #467: listFleet's callerIsPrimary default fails closed (fleetd #463)
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m38s
A wrapper overload that is called with no callerIsPrimary argument used to
default it to true, so a forgotten argument silently handed out the lead's
coordination state. It now defaults to false: a missing identity fails closed.

Verified on the merge commit, not the branch:

- The funnel is real. FleetMcp has 7 listFleet declarations and 7 real calls
  (an 8th 'listFleet(' match is a javadoc {@link}). Exactly one call writes a
  literal for the new boolean, and it writes false; exactly one writes the real
  predicate, coordinatorVisibleTo(principal(exchange)) in the MCP handler. No
  call writes true.
- Control battery on the merge: 1584 tests green unmutated.
- M1, the fix reverted at the one line that writes the default (false -> true):
  KILLED by listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey.
- M2, the half this round weakened. The worker rewrote
  listIsByteForByteUnchangedForThePrimaryCaller and dropped its byte-for-byte
  equality assertion, which was correct because that comparison ran against the
  implicit-default overload -- now the path #463 closes. So: does anything still
  notice if the primary's coordinator row silently loses a field? Dropped
  heldCount: KILLED by two tests, that same test and
  listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero.
- M3, a regression check on #439's own gate, 'if (callerIsPrimary)' -> 'if (true)':
  KILLED by three tests.

What is no longer pinned, stated plainly: the primary's output is now checked
field by field, not as a whole string. A field that no test names could
disappear without failing anything. Every field the tests do name is pinned,
proven by M2. The whole-output answer belongs to fleetd #460.

One control in my own battery was wrong and is worth recording: I labelled
', true);' in FleetMcp.java 'must be 0' and it is 2 -- row.put("configured",
true) and m.put("self", true), neither a listFleet delegation. The pattern was
too wide. The count above comes from a walk over each declaration and call
instead.
2026-09-10 19:04:51 +07:00
lead eccd0548ce Merge #465: state canonical invariant 5 as a purpose, not a banned tool (fleetd #458)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 2m1s
Verified by me on a local merge of 29e7a06 onto f5e02fe:

- One file, one hunk. 5 changed lines inside the canonical block, all of them
  invariant 5; a line-by-line diff of the block against origin/main shows
  nothing else moved.
- The addendum is untouched: both 'Herdr socket tests (measured 2026-09-10)'
  and 'Only the lead can run that check' appear once on each side.
- Block size: origin/main 17467 chars / 17626 bytes, the merge 17557 chars /
  17718 bytes. The worker's numbers were bytes and were right; I am naming both
  units because len(str) in python is characters and this block is full of em
  dashes, which is a mistake I have made before.
- mvn -B clean test: Tests run: 1583, Failures: 0, Errors: 0, Skipped: 0.
  BUILD SUCCESS, rc=0.

The three things I reserved from the worker, because a provisioned worktree
cannot do them:

- wiki/7-Use-Cases.md updated to match, in the wiki repo's own commit.
- The sync script run in this clone: in sync: True.
- Other projects carrying the block: NONE. 29 CLAUDE.md files under the
  operator's project roots, 18 of them carry the block and 17 of those are this
  repo's own worktrees, which inherit it from git. Only
  LTMS/claude-bridge/CLAUDE.md is a real second copy, and it is this file.

On the worker's criterion-4 question: the addendum note stays. The restated
invariant removes the unsatisfiability, which was the sharp problem, but the
note still names the exact four files the carve-out covers and records why
(fleetd #449, where a fake and the real herdr disagreed for weeks under a green
suite). A rule that is merely satisfiable is not the same as a rule a worker can
apply without re-deriving it.
2026-09-10 18:57:24 +07:00
Dai Ha c4e23eebad fleetd #464: guard charter tool names
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m37s
2026-09-10 18:55:16 +07:00
Dai Ha 7df503dfc2 fleetd #463: default listFleet's callerIsPrimary to false, fail closed
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m9s
The compat overload at FleetMcp.java:1373 defaulted callerIsPrimary to a
literal true, so a caller that forgot the argument silently got the
coordinator row (this daemon's coord-id, mailbox state, held-mail previews,
peer reachability) -- lead-to-lead state fleetd #439 just gated. Flip the
default to false: a forgotten argument now yields a missing row instead of
a leaked one.

Six FleetMcpTest methods relied on the implicit true to see the coordinator
row at all; they now pass true explicitly through the canonical overload.
listIsByteForByteUnchangedForThePrimaryCaller's own premise (comparing the
implicit-default path against an explicit-true path) was the shape of the
bug, so it now only exercises the explicit-true path.

Added listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey
to pin the new default: a compat overload called with no callerIsPrimary
argument, against a fully-configured lead channel, must produce a result
with the coordinator key absent -- not empty, not redacted, absent.
2026-09-10 18:55:13 +07:00
Dai Ha f5e02fedd6 plans: commit the fleet01 move plan, with today's state measured on the host
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m58s
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.

Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:

- Phases 1 to 3 are done. Both systemd user units exist, are active and
  enabled, linger is on, one java process (so the old restart.sh
  double-daemon problem is gone), the live fleetd.yaml is on the host, and
  the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
  origin/main, 0 ahead. So fleet01's daemon runs code from before this
  week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
  uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
  lead is live on the coordination channel, so the headless-login blocker
  it names is solved.

Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.

And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
2026-09-10 18:53:11 +07:00
Dai Ha 29e7a06c49 fleetd #458: restate invariant 5 by purpose, not mechanism
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Failing after 1m56s
Invariant 5 banned 'driving the terminal multiplexer directly', naming
herdr CLI and socket as the banned tool. That bans a mechanism. What it
protects is the control plane: nobody may move a fleet session, pane or
peer by a route that skips the bridge's policy checks.

In this repo herdr is itself the subject under test, so four contract
tests must open its socket on purpose (see the addendum's 'Herdr socket
tests' note). Under the old wording, a worker assigned to that code
reads invariant 5 and finds its only path to finish the task banned.

Restate the invariant by purpose: never move fleet state except through
the bridge. herdr's CLI and socket stay as the named example of the
banned route, not the definition of it.
2026-09-10 18:50:41 +07:00
Dai Ha 92a96fcbd8 Merge #462: gate the coordinator row on a named predicate the handler must consult (fleetd #439)
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m51s
Verified by me on a local merge of c1ca627 onto 1348287:

- mvn -B clean test: see the totals below. Ran in fleetd/, redirected to a file,
  exit code captured on its own line.
- Mutation battery on merge 8402923 (same two parents, earlier base 5f1b260),
  4 cells, each with a proof gate on the occurrence count:
    * CONTROL, unmutated: 1583 tests, 0 failures, rc=0.
    * M2r - the handler stops asking who called: coordinatorVisibleTo(principal(exchange))
      replaced by a literal true. KILLED by
      FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo.
      This is the mutant that survived round 1, so the gap the follow-up was for is closed.
    * M3 - is the new source-reading detector vacuous? Renamed its anchor
      (listHandler -> listHandlerX, 2 sites, behaviour identical). The detector FAILED,
      rc=1, as a source-reading test must: it cannot silently pass on an empty scrape.
    * M4 - the predicate itself always says yes: return caller.isPrimary() replaced by
      return true. KILLED by FleetMcpAuthzTest.onlyThePrimaryMaySeeTheCoordinatorRow.
  Tree verified clean before the battery and restored after each cell.
- ANON is pinned too, not only worker and architect: FleetMcpAuthzTest asserts
  coordinatorVisibleTo(ANON) is false.
- listFleet( appears in exactly one main file, mcp/FleetMcp.java, and no REST class builds
  the coordinator row, so the detector's single-file scope covers every live call site
  today. Control for that sweep: 109 main .java files matched a string they all contain.

Not in this PR, and my call, not the worker's: the six compat overloads of listFleet still
default callerIsPrimary = true, which fails open. Safe today because the one production
call site passes the computed value. Filed separately.
2026-09-10 18:40:48 +07:00
Dai Ha 13482872bb contract test: keep the loaded run's passes, discard only its failure
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 2m6s
The javadoc threw away the whole loaded run as "not clean evidence". The
fixed cell order makes that too strong in one direction. A cell that runs
last on a climbing load has a free explanation for FAILING. It has no free
explanation for PASSING: surviving a worse condition than a fair order
would have given it is evidence in the safe direction. So the old version's
0 of 3 is still discarded, and the three passes are kept with the load each
one ran at.

Also names the cell that tests the swallow explanation head-on, which the
old text left as "nobody has managed that yet". SHELL_READY_TIMEOUT_MS at 0
types input at once — the worst case for "typed before the prompt" — and it
passed 3 of 3 at load 18.42 to 23.65. The wider read window cannot explain
that away, because a swallowed keystroke is lost, not late: the command
never runs, so no amount of polling makes its output appear.

And a warning not to carry the raw load average to another host. Load
average counts differently per core and per operating system, so only load
per core compares. I broke that rule myself when comparing this run with
another host's numbers.

Both points came from the fleet01 lead reviewing 2af13ab.
2026-09-10 18:37:41 +07:00
Dai Ha c1ca6273fc fleetd #439: pin the caller at the fleet_list call site, not just the gate
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m33s
Review of PR #462 found M2: the coordinatorVisibleTo gate (then an inline
principal(exchange).isPrimary() check) could survive a mutation that
replaced the argument with a literal true at the one production call
site, because every existing test drove listFleet directly and supplied
the boolean itself -- nothing exercised the handler's own call.

- Name the decision: FleetMcp.coordinatorVisibleTo(Principal), a small
  package-private predicate next to denyFor/recordPrimarySingleton. The
  fleet_list handler now calls coordinatorVisibleTo(principal(exchange))
  instead of inlining .isPrimary().
- Pin the predicate's role table in FleetMcpAuthzTest
  (onlyThePrimaryMaySeeTheCoordinatorRow), covering primary/worker/
  architect and, newly, anonymous.
- Add a source-reading detector at the boundary
  (theFleetListHandlerActuallyConsultsCoordinatorVisibleTo), same idiom as
  toolsTheServerRegisters/everyRegisteredToolHasItsHandlerActionPinned: it
  reads FleetMcp.java, isolates the listHandler block, asserts (as a
  control) that the block actually contains a listFleet( call, then
  asserts the call's trailing boolean argument is exactly
  coordinatorVisibleTo(principal(exchange)) -- not a literal true/false.

Both mutations from the review were reproduced and killed by these tests,
then reverted; see the PR body for the full break-and-restore transcript.
2026-09-10 18:28:16 +07:00
Dai Ha 5f1b260c81 addendum: the canonical sync check is the lead's, not a member's
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m45s
I put "the sync script must print in sync: True" in a worker's acceptance
criteria for #455. The worker could not run it and said so, honestly, instead
of inventing a pass. My brief was the defect.

A member's provisioned worktree has wiki/ uninitialized, so the script dies
with FileNotFoundError. Measured in three worker worktrees: git submodule
status printed a leading '-' and wiki/ held 0 entries. The primary's own clone
printed a leading '+' and the file was there.

The note is dated, gives the re-measure command, says what each outcome means,
and says to delete it once it stops reproducing - as this file requires of any
measurement in an addendum.

Canonical block untouched: the sync script itself reports in sync: True.
2026-09-10 18:25:32 +07:00
ltms 9d1306d442 Merge #461: project addendum for the herdr socket carve-out (fleetd #455)
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m56s
Doc-only, one file, 13 added lines, inside §Project addendum. Verified by me:

- Canonical block byte-identical: 17468 chars on main and on the branch.
- The sync script in CLAUDE.md prints "in sync: True" both before and after.
- The note's own re-measure command works. Run verbatim, it returns exactly the 4 files
  the note names, and no others:
    fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java
  Control that the search reaches the tree: 120 test java files, 63 of them mention herdr.
  That control matters here — a zero-match re-measure command would tell a future session to
  delete a live restriction.
- All four required elements are present: the carve-out, what stays banned (the control plane),
  who it applies to, and the perishable half (dated 2026-09-10, the command, what each outcome
  means, and delete-when-stale).
- Cross-references #458 for the canonical restatement, which is deliberately not in this change.

Pre-send check applied to the note itself: a worker assigned to AgentControlContractTest can now
name one legal action that finishes its task — let the test open the herdr socket in a throwaway
workspace it tears down.
2026-09-10 13:18:17 +02:00
Dai Ha e54e3d87ea fleetd #439: omit fleet_list's coordinator key for non-primary callers
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
The coordinator row is lead-to-lead coordination state (coord-ids, mailbox
facts, held-message previews). fleet_list returned it to every caller,
including a worker or an architect, because coordinatorView() had no way to
know who was asking.

Gate at the call site inside listFleet: a new overload takes
callerIsPrimary and only assembles/attaches the coordinator row when it is
true, so the key is absent (not empty) for a worker or an architect. The
MCP handler now passes principal(exchange).isPrimary(); every other
listFleet overload keeps passing true, so callers with no caller identity
(existing unit tests, the no-op wrappers) are unaffected -- confirmed by a
byte-for-byte comparison test against the pre-fix overload.

Authz's READ case is untouched: it stays shared by fleet_status,
fleet_profiles and fleet_whoami, and the gate here is purely inside
fleet_list's own result assembly.
2026-09-10 18:16:48 +07:00
Dai Ha bf027f10b9 #455: document herdr test socket exception
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m50s
2026-09-10 18:13:50 +07:00
Dai Ha 3c5873dfe2 fleetd #453: point the new javadoc at the right javadoc
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m52s
#456's first paragraph said "HerdrPeerLauncher's own {@link
#spawn(SpawnRequest, PlacementDecision)} javadoc". That link resolves to this
interface's own abstract declaration, not to HerdrPeerLauncher's override, so a
reader who follows it lands on the wrong text. The second paragraph already
used the plain {@code HerdrPeerLauncher.spawn(...)} form; both now match.

Also says who is actually forced to read the paragraph, because #456's
reasoning rests on it and the two cases differ. A class that implements this
interface directly must write a body for spawn(SpawnRequest,
PlacementDecision) - it is abstract here - so it reads this javadoc. A
subclass of HerdrPeerLauncher does not: HerdrPeerLauncher already implements
that method (member/HerdrPeerLauncher.java:611) and the subclass inherits the
body. For a subclass the paragraph is advice, not a gate.

javadoc -Ddoclint=reference: 5 "reference not found", the same 5 in the same 5
untouched files as origin/main at 2af13ab, and none in PeerLauncher.java. Those
5 are ticket #459. mvn compile rc=0.
2026-09-10 18:05:07 +07:00
ltms e29227d5f4 Merge #456: document the override obligation on PeerLauncher.defaultProfileFor/place (fleetd #453)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m34s
Doc-only, one file, 17 added lines. Verified by me on a local merge of 0788d84
onto 2af13ab (merge commit 5b46538):

- mvn -B clean test: Tests run: 1578, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS, rc=0.
- javadoc reference lint (mvn javadoc:javadoc -Ddoclint=reference): the merge has 5
  "reference not found" errors; origin/main at 2af13ab has the same 5, in the same 5
  untouched files. So this PR adds no broken link. Both runs rc=1 for that pre-existing
  reason, which is now ticket #459.
- The quoted phrase is real: "are unoverridden here and just wrap {@link #defaultProfile()}"
  is at member/HerdrPeerLauncher.java:604. place() is not overloaded, so {@link #place}
  is unambiguous.

One inaccuracy I am fixing in a follow-up commit rather than sending the PR back: the first
new paragraph writes "HerdrPeerLauncher's own {@link #spawn(SpawnRequest, PlacementDecision)}
javadoc", but that link resolves to PeerLauncher's own abstract declaration, not to
HerdrPeerLauncher's override. The second paragraph already gets this right with plain
{@code HerdrPeerLauncher.spawn(...)}.
2026-09-10 13:04:09 +02:00
Dai Ha 2af13ab1ff fleetd #449: name the decisive cell and the load the claim was measured at
CI / contract (push) Successful in 54s
CI / build (push) Successful in 2m8s
The previous comment named a cause with no cell behind it. fleet01's review
made the point: a confident wrong mechanism gets copied, and a confident
under-determined one gets copied the same way.

So the comment now names the cell that settles it. Hold the old 1000ms write
sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the
test goes 0 of 3 to 3 of 3. The read deadline was the whole story.

It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The
loaded run agreed, but its load climbed from 7 to 50 while the cells ran and
the old version ran last, so it is not clean evidence and the comment says so.

Comment only. No test or production code changed.
2026-09-10 17:59:52 +07:00
Dai Ha 0788d84be8 fleetd #453: document the override obligation on PeerLauncher.defaultProfileFor/place
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m34s
Decision: leave both as default methods (option 1), not abstract. Neither
default is a live defect today — HerdrPeerLauncher is the sole single-profile
implementer and the degenerate answer (ignore role, always defaultProfile())
is correct for it. Making them abstract would force ~10 boilerplate one-line
overrides across 5 unrelated PeerLauncher test doubles (NeverSpawnsLauncher,
RaceLauncher, NoResumeLauncher, ClearContextSpyLauncher, LazyIdLauncher) that
never call either method, for a risk that is speculative (no multi-profile
HerdrPeerLauncher subclass exists or is planned).

Strengthens both javadocs with an explicit MUST-override warning and cross-
references HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)'s existing
#450 javadoc, which already names both methods as "unoverridden here" and
ties that to being a single-profile adapter -- the concrete place a future
multi-profile launcher author would read, since #450 made that method
abstract and any subclass must write its body.
2026-09-10 17:49:37 +07:00
12 changed files with 1156 additions and 48 deletions
+26 -2
View File
@@ -58,8 +58,9 @@ and the sender silently receives nothing. Fail toward the recoverable error.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
### Primary (lead) — run this on every task, in order
@@ -223,6 +224,19 @@ must obey belongs in the charter, not here.
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
throwaway workspace that the test tears down. This only covers
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
banned. That is the control plane that invariant 5 protects. Re-measure with
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
result means tests still open the socket and this note still applies. An empty result means nobody
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
is not part of this change.
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
@@ -379,6 +393,16 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
prints a commit with no `-`, delete this paragraph.
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
+15 -6
View File
@@ -793,12 +793,21 @@ guard:
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
# beyond today's behaviour. A skill folder the target repo already carries under
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Claude Code
# members only; an opencode member reads a different path (.opencode/agent) this key does not
# touch. Best-effort like worktreeGroup above: a missing/unreadable directory here is logged and
# skipped, never a failed spawn. Every non-hidden subdirectory of this directory is copied
# wholesale, with no per-file allowlist — don't park scratch files or drafts alongside the real
# skill folders, they will be copied into every provisioned worktree too.
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
# copied into every provisioned worktree too.
#
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
# earlier builds copied the files for every kind but only claude-code could read them, and the
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
# log) to see what a given member actually got.
# memberSkills: /path/to/fleetd/checkout/.claude/skills
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
@@ -455,7 +455,8 @@ public final class FleetMcp {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
leadSeats, callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange),
new CoordinationSource(leadChannel, peers));
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
(exchange, req) -> {
@@ -608,6 +609,20 @@ public final class FleetMcp {
}
}
/**
* fleetd #439: only the primary may read {@code fleet_list}'s {@code coordinator} row —
* lead-to-lead coordination state (coord-ids, mailbox facts, held-message previews), never the
* roster. Split out of the {@code fleet_list} handler, same reason as {@link #denyFor} and
* {@link #recordPrimarySingleton}: the decision must be unit-testable without fabricating an
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather than
* inlining the check, so a future edit cannot silently pass a literal instead of asking who
* called ({@code FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}
* reads the source and asserts the handler calls this method by name, not a literal).
*/
static boolean coordinatorVisibleTo(Principal caller) {
return caller.isPrimary();
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
@@ -1389,12 +1404,48 @@ public final class FleetMcp {
LeadSeatSource.none(), leads, selfTerm, coordination);
}
/** As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}). */
/**
* As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}).
*
* <p>Assumes the caller is <strong>not</strong> the primary (fleetd #463) — every wrapper
* overload above delegates here without carrying a caller identity, which is exactly right for
* them: they exist for call sites (and unit tests) that have no {@link Principal} to hand over,
* and a missing identity should fail closed rather than fail open onto lead-to-lead state. A
* test that wants the {@code coordinator} row must call the canonical overload below with an
* explicit {@code true}. The one call site that has a real caller ({@code fleet_list}'s MCP
* handler) uses {@link #listFleet(PeerLauncher, SessionManager, MessageService, CapacitySource,
* HealthCoverageSource, QuarantineSource, OutageSource, LeadSeatSource, Map, String,
* CoordinationSource, boolean)} instead, so it can pass the true answer.
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
leadSeats, leads, selfTerm, coordination, false);
}
/**
* As above, gated by the caller's role (fleetd #439). The {@code coordinator} row is
* lead-to-lead coordination state — coordination between orchestrators, not roster
* observation — so it is assembled and included only when {@code callerIsPrimary} is
* {@code true}. A worker or an architect gets a result with the {@code coordinator} key
* <strong>absent</strong>, never an empty or redacted one, and never pays the cost of
* {@link #coordinatorView} probing peer mailboxes for a row it will not receive.
*
* @param callerIsPrimary whether the {@code fleet_list} caller is the primary; only the MCP
* handler computes this from the real connection (see
* {@code Principal#isPrimary()}) — every other overload passes
* {@code false} (fleetd #463: a forgotten argument fails closed, not
* open), so a test that wants the {@code coordinator} row must pass
* an explicit {@code true}
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -1416,9 +1467,14 @@ public final class FleetMcp {
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
Map<String, Object> coordinatorRow = coordinatorView(coordination);
if (coordinatorRow != null) {
result.put("coordinator", coordinatorRow);
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
// key is absent rather than present-and-empty.
if (callerIsPrimary) {
Map<String, Object> coordinatorRow = coordinatorView(coordination);
if (coordinatorRow != null) {
result.put("coordinator", coordinatorRow);
}
}
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
@@ -331,11 +331,22 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
+ "form — opencode's per-model context limit could not be applied for this profile",
cfg.profile(), cfg.model());
}
// fleetd #393: memberSkills seeding (GitWorktrees#seedSkills) copies skill folders into
// EVERY provisioned worktree's .claude/skills/ regardless of which kind ultimately spawns
// into it — that copy step cannot know the kind, only the caller of GitWorktrees#add does
// (see that method's own javadoc). .claude/skills/ is a Claude Code CLI convention the CLI
// discovers on its own; opencode has no such discovery, so without this, a seeded skill
// never reaches an opencode member even though GitWorktrees logged it as seeded. Read
// whatever landed under <cwd>/.claude/skills/ here — the one place in this launcher that
// knows both the kind (opencode, by construction: this IS OpenCodeLauncher) and the cwd.
List<Path> skillInstructionFiles = skillInstructionFiles(spec.cwd());
// A config file is needed for the bridge MCP mount, a member charter, the IDE MCP (+ its
// guidance overlay, CB-634), a pinned endpoint (CB-508), or a resolvable autoCompactWindow.
// guidance overlay, CB-634), a pinned endpoint (CB-508), a resolvable autoCompactWindow, or
// at least one seeded skill to deliver via instructions[] (fleetd #393).
if (cfg.hasMcp() || cfg.hasIdeMcp() || spec.charter() != null || hasCustomProvider(cfg)
|| wantsContextLimit) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter(), spec.cwd()).toString());
|| wantsContextLimit || !skillInstructionFiles.isEmpty()) {
workerEnv.put("OPENCODE_CONFIG",
writeConfig(cfg, spec.charter(), spec.cwd(), skillInstructionFiles).toString());
}
applyGitToken(workerEnv, cfg);
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
@@ -358,6 +369,65 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return withAgent;
}
/**
* fleetd #393: the {@code SKILL.md} paths under {@code <cwd>/.claude/skills/} this launcher can
* turn into {@code instructions[]} entries, plus the honest log this ticket asks for — emitted
* here, at the one point a skill's fate for THIS spawn is actually known, rather than trusting
* {@code GitWorktrees#seedSkills}'s kind-blind "N of M" line to mean "and it will be read."
*
* <p>Every non-hidden subdirectory of {@code .claude/skills/} is a candidate, whether it got
* there via {@code memberSkills:} seeding or because the target repo ships its own copy — this
* launcher does not care which; it only cares what it can find at spawn time. A candidate with
* a {@code SKILL.md} at its top level (the same shape {@link #writeConfig} already requires for
* the charter and IDE-rules instructions entries) is delivered; anything else is a directory
* this launcher cannot turn into a flat instructions entry, named explicitly in the log rather
* than silently dropped, so a caller sees a real "cannot consume" reason and not just a smaller
* number than {@code GitWorktrees}' own count.
*
* <p>No candidates at all (directory absent or empty) logs nothing — the same
* no-log-when-nothing-to-say shape {@link #hasCustomProvider} and friends already follow, and
* the shape {@code GitWorktrees#seedSkills} itself uses when {@code memberSkills:} is unset.
* A failure to even list the directory is logged and treated as "nothing delivered" — best
* effort, must never fail the spawn, matching {@code GitWorktrees#seedSkills}'s own contract.
*/
private List<Path> skillInstructionFiles(String cwd) {
if (cwd == null || cwd.isBlank()) {
return List.of();
}
Path skillsDir = Path.of(cwd, ".claude", "skills");
if (!Files.isDirectory(skillsDir)) {
return List.of();
}
List<Path> candidates;
try (var listing = Files.list(skillsDir)) {
candidates = listing.filter(Files::isDirectory)
.filter(p -> !p.getFileName().toString().startsWith("."))
.sorted()
.toList();
} catch (IOException e) {
log.warn("could not scan {} for skill folders to deliver to this opencode member: {}",
skillsDir, e.getMessage());
return List.of();
}
if (candidates.isEmpty()) {
return List.of();
}
List<Path> delivered = candidates.stream()
.map(dir -> dir.resolve("SKILL.md"))
.filter(Files::isRegularFile)
.toList();
List<String> undeliverable = candidates.stream()
.filter(dir -> !Files.isRegularFile(dir.resolve("SKILL.md")))
.map(dir -> dir.getFileName().toString())
.toList();
log.info("skill delivery: {} of {} skill folder(s) under {} reached this opencode member via "
+ "instructions[] (opencode does not read .claude/skills/ natively, unlike "
+ "Claude Code){}",
delivered.size(), candidates.size(), skillsDir,
undeliverable.isEmpty() ? "" : "; no SKILL.md, could not be delivered: " + undeliverable);
return delivered;
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
@@ -418,8 +488,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*
* @param skillInstructionFiles fleetd #393: absolute {@code SKILL.md} paths from
* {@link #skillInstructionFiles(String)}, appended to
* {@code instructions[]} so a {@code memberSkills:}-seeded skill
* reaches this opencode member the same way the charter and IDE
* rules already do.
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd,
List<Path> skillInstructionFiles) {
try {
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
dir.toFile().deleteOnExit();
@@ -447,7 +524,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Files.writeString(charter, charterText);
charter.toFile().deleteOnExit();
root.putArray("instructions").add(charter.toAbsolutePath().toString());
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
// is already at "instructions" — harmless only as long as this block runs first
// against a still-empty root, which is an ordering constraint nothing declared or
// tested. The skills writer just below, and the IDE-rules writer further down,
// both already use withArray (get-or-create) for exactly this reason; this was the
// one straggler. Proven load-bearing on the fleetd #393 merge: flipping this one
// call back to putArray left the whole suite green while silently deleting the
// charter entry whenever skills or IDE rules ran after it — an opencode member
// would launch with no role contract at all, worse than the bug #393 fixed, and
// nothing caught it. See OpenCodeLauncherTest's
// instructionsArrayHoldsCharterThenIdeRulesInOrder and
// instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder.
root.withArray("instructions").add(charter.toAbsolutePath().toString());
}
// fleetd #393: each seeded skill's SKILL.md, delivered as a plain instructions[] entry
// — the only mechanism opencode has for static guidance text. Unlike Claude Code's
// Skill tool, opencode cannot load one of these on demand by name; the content is just
// always part of the system prompt from spawn. That is a real difference in HOW the
// content reaches the member, not a reason to withhold it.
for (Path skillFile : skillInstructionFiles) {
root.withArray("instructions").add(skillFile.toAbsolutePath().toString());
}
if (cfg.hasMcp() || cfg.hasIdeMcp()) {
@@ -168,6 +168,23 @@ public interface PeerLauncher {
* answer for a launcher with no role-pool concept of its own (e.g. a single {@code
* HerdrPeerLauncher} adapter, which is never reached this way in production: {@code
* CompositePeerLauncher} always fronts it and resolves roles itself).
*
* <p>fleetd #453: this default is deliberately <em>not</em> abstract — unlike {@link
* #spawn(SpawnRequest, PlacementDecision)} (fleetd #450), there is no live defect in inheriting
* it today, and the only current single-adapter implementer ({@code HerdrPeerLauncher}) is
* correct to do so. But it stays correct only as long as that holds: <strong>if a launcher ever
* routes more than one profile per role, it MUST override this method</strong>, or every role
* silently resolves to {@link #defaultProfile()} with no error and no log line. {@code
* HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)} — the override in {@code
* dev.ltms.fleet.member}, not the declaration below — names this method and {@link #place}
* explicitly as "unoverridden here" for exactly this reason. Read it before adding role-pool
* routing to any {@code HerdrPeerLauncher} subclass.
*
* <p>Who is forced to read which paragraph, because it is not symmetric. A new class that
* implements this interface directly must write a body for {@link #spawn(SpawnRequest,
* PlacementDecision)}, which is abstract here, so it lands on this javadoc. A subclass of
* {@code HerdrPeerLauncher} does not: that class already implements the method, and the
* subclass inherits the body. So for a subclass this paragraph is advice, not a gate.
*/
default String defaultProfileFor(MemberRole role) {
return defaultProfile();
@@ -228,6 +245,13 @@ public interface PeerLauncher {
* placement condition — the right answer for a launcher with no pool or placement-policy
* concept of its own, matching {@link #defaultProfileFor}'s own default.
*
* <p>fleetd #453: same reasoning as {@link #defaultProfileFor}'s own #453 note — this default
* is deliberately not abstract (no live defect today, correct for the sole single-adapter
* implementer), but <strong>a launcher that ever routes more than one profile per role MUST
* override this method too</strong>, or placement silently ignores {@code role} for it. See
* {@code HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)}'s javadoc, which names this
* method as "unoverridden here" and why that is correct only for a single-profile adapter.
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
@@ -690,14 +690,18 @@ public final class GitWorktrees implements Worktrees {
* {@code fleet.seededSkillsNote}, readable with {@code git config --worktree --get-all
* fleet.seededSkills}.
*
* <p><b>Claude Code specific by construction, not by a backend check here.</b> Only {@code
* .claude/skills/<name>/SKILL.md} is a path any launcher reads today (opencode's equivalent is a
* different shape under {@code .opencode/agent}, out of scope — see issue #362). This method
* only copies files; like {@link #isolateToolSurface} — which neutralizes BOTH {@code .mcp.json}
* and {@code opencode.json} unconditionally — it runs the same for every worktree regardless of
* which backend ultimately spawns into it, because the backend is not yet chosen at {@link #add}
* time. A seeded {@code .claude/skills/} directory in an opencode member's worktree is simply
* never read by that launcher.
* <p><b>Kind-blind by construction, not by a backend check here — this used to be a real gap
* (fleetd #393).</b> This method only copies files; like {@link #isolateToolSurface} — which
* neutralizes BOTH {@code .mcp.json} and {@code opencode.json} unconditionally — it runs the
* same for every worktree regardless of which backend ultimately spawns into it, because no
* caller of {@link #add} hands this class a kind to consult. Before fleetd #393, that made the
* log line below a false claim of success for a {@code kind: opencode} member: opencode has no
* built-in discovery of {@code .claude/skills/}, unlike the Claude Code CLI, so a seeded skill
* never reached one. It now does — {@code OpenCodeLauncher#skillInstructionFiles} reads
* whatever this method copied into {@code .claude/skills/} and appends each {@code SKILL.md} to
* the generated {@code instructions[]} — but that delivery, and the log line that honestly
* claims it (kind-aware, unlike this one), happens at the launcher, once the kind is actually
* known, not here.
*/
private void seedSkills(String worktreePath) {
if (memberSkillsSource == null) {
@@ -739,7 +743,15 @@ public final class GitWorktrees implements Worktrees {
if (detail.isEmpty()) {
detail = "no skill folders found under " + source;
}
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {}",
// fleetd #393: this only claims the copy step, deliberately — it cannot know the member
// kind that will spawn into this worktree (see this method's own javadoc), so it must not
// read as "and the member will act on it." Whether that is true depends on the kind: the
// Claude Code CLI discovers .claude/skills/ on its own; OpenCodeLauncher logs its own
// "skill delivery" line, once the kind is known, naming what it could and could not turn
// into instructions[].
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {} "
+ "(whether the spawned member can act on this depends on its kind — see "
+ "the launcher's own log for that)",
seeded.size(), seeded.size() + kept.size(), source, worktreePath, detail);
if (seeded.isEmpty()) {
return;
@@ -25,17 +25,39 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
* So this polls for a real signal instead of guessing a sleep length.
*
* <p><strong>What was measured, and what was not.</strong> Polling fixes it: 5 standalone runs
* green. The load-bearing half is {@link #waitForText}. With {@link #SHELL_READY_TIMEOUT_MS}
* set to 0 — so input is typed at once, with no settle wait at all — the test still passed 3 of
* 3. So the proven cause is the 800ms READ deadline being too short, not the 1000ms write delay.
* Note the direction, because it matters: typing at 0ms works where typing at 1000ms failed. The
* earlier explanation for this test — that input typed before the prompt is swallowed by the
* shell's startup — is therefore NOT supported by any measurement here. Please do not repeat it
* as the reason; if it were true, 0ms would be worse than 1000ms, and it is better.
* green. The load-bearing half is {@link #waitForText}, and one cell proves it. Keep the old
* 1000ms write sleep and change only the read — the 800ms fixed sleep becomes a 5s poll — and
* the test goes from 0 of 3 passing to 3 of 3. Removing the write wait instead
* ({@link #SHELL_READY_TIMEOUT_MS} set to 0, so input is typed at once) also passes 3 of 3. So
* the cause is the 800ms READ deadline, not the 1000ms write delay. The old version fails every
* time, not sometimes, so "race" is the wrong word for it. The earlier explanation — that input
* typed before the prompt is swallowed by the shell's startup — is not supported by anything
* measured here. Please do not repeat it: if it were true, typing at 0ms would be worse than
* typing at 1000ms, and it is not.
*
* <p><strong>Where this was measured.</strong> A 12-core macOS host, load average 2.6 to 5.9,
* on commit 20c1094. The same four cells were also run under load, with 8 spinners on 12 cores.
* The failure and the passes from that run are not worth the same. The cells ran in a fixed
* order while the load climbed from 7 to 50. The old version ran last, at the top of that climb,
* so it has a free explanation for failing and its 0 of 3 is discarded. A pass has no such free
* explanation: a cell that survives a worse condition than a fair order would have given it is
* evidence in the safe direction. So keep the three passes, each with the load it ran at: the
* fixed version 3 of 3 at load 7.42 to 18.42, the 0ms-write cell 3 of 3 at 18.42 to 23.65, the
* 1000ms-write cell 3 of 3 at 23.65 to 46.40. Above about load 20 everything here is slow for
* reasons that have nothing to do with this seam, so read the positive claim — that the read
* deadline was the whole cause — as "measured near idle on a 12-core host", and nothing
* stronger. Do not carry that raw load average to another host either: load average counts
* differently per core and per operating system, so only load per core compares. If this test
* fails on a smaller or busier machine, raise {@link #OUTPUT_TIMEOUT_MS} before you suspect the
* seam.
*
* <p>{@link #waitUntilSettled} is kept as cheap insurance against that swallow case, not because
* anyone showed it was needed. If you want to delete it, the honest test is whether you can make
* this test fail by typing early. Nobody has managed that yet.
* anyone showed it was needed. The cell that tests the swallow case head-on is the one with
* {@link #SHELL_READY_TIMEOUT_MS} at 0: input is typed at once, which is the worst case for
* "typed before the prompt is ready". It passed 3 of 3 at load 18.42 to 23.65. The wider read
* window cannot explain that pass away, because a swallowed keystroke is LOST, not late — the
* command never runs, so no amount of polling makes its output appear. So the swallow mechanism
* was tested and did not show up. If you want to delete this call, that is the cell to re-run.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -0,0 +1,72 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashSet;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** fleetd #464: launch charters must not name MCP tools the server does not register. */
class CharterToolSurfaceTest {
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(String text, String regex) {
Matcher m = Pattern.compile(regex).matcher(text);
Set<String> found = new LinkedHashSet<>();
while (m.find()) {
found.add(m.group(1));
}
return found;
}
/** Every {@code fleet_*} or legacy {@code bridge_*} token in the configured launch charters. */
private static Set<String> toolsNamedIn(FleetConfig config) {
return matches(String.join("\n", config.fleet().charters().values()),
"(fleet_[a-z_]+|bridge_[a-z_]+)");
}
/** Every tool {@link FleetMcp} registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(Files.readString(MCP_SOURCE), "tool\\(\\\"(fleet_[a-z_]+)\\\"");
}
@Test
@DisplayName("[SOURCE TEXT] every tool named in a configured charter is registered by the server")
void configuredChartersNameOnlyRegisteredTools(@TempDir Path dir) throws Exception {
Path configFile = dir.resolve("charters.yaml");
Files.writeString(configFile, """
fleet:
charters:
dev: |
Send the final handoff through fleet_reply.
reviewer: |
Use fleet_ask only for the lead's decision.
""");
FleetConfig config = FleetConfig.load(configFile);
Set<String> named = toolsNamedIn(config);
Set<String> registered = toolsTheServerRegisters();
assertTrue(!named.isEmpty(),
"the charter fixture named no fleet_* or bridge_* tool. This test would check nothing; "
+ "add charter text that names a tool before changing the extraction.");
assertTrue(!registered.isEmpty(),
"the FleetMcp registration scrape found no tools. This test would check nothing; "
+ "repair the tool(\"…\") extraction before changing the assertion.");
Set<String> unknown = new LinkedHashSet<>(named);
unknown.removeAll(registered);
assertTrue(unknown.isEmpty(),
"configured charter text names " + unknown + ", but FleetMcp does not register it. "
+ "Checked " + named + " against " + registered + ". Fix the charter text or "
+ "register the tool; do NOT weaken this test.");
}
}
@@ -187,6 +187,71 @@ class FleetMcpAuthzTest {
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
}
// --- fleetd #439: who may see fleet_list's coordinator row ----------------------------------
/**
* fleetd #439: {@link FleetMcp#coordinatorVisibleTo} is the whole policy decision for
* {@code fleet_list}'s {@code coordinator} row — lead-to-lead coordination state, not roster
* observation. Only the primary may see it; a worker, an architect, and (the case the previous
* pass of this ticket did not cover) an anonymous caller must all be refused.
*/
@Test
void onlyThePrimaryMaySeeTheCoordinatorRow() {
assertTrue(FleetMcp.coordinatorVisibleTo(PRIMARY), "the primary must see its own coordination state");
assertFalse(FleetMcp.coordinatorVisibleTo(WORKER_A), "a worker must not see lead-to-lead coordination state");
assertFalse(FleetMcp.coordinatorVisibleTo(ARCH_DESIGN),
"an architect holds READ today, but that must not extend to coordinator");
assertFalse(FleetMcp.coordinatorVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* fleetd #439 / PR #462 review finding M2: the predicate above can be perfectly correct while
* the one production call site (the {@code fleet_list} MCP handler) never actually asks it —
* a literal {@code true} compiles, and the whole suite stayed green under that mutation because
* every existing test drives {@link FleetMcp#listFleet} directly and supplies the boolean
* itself. This test reads {@code FleetMcp.java}'s own source (same idiom as {@link
* #toolsTheServerRegisters()} / {@link #everyRegisteredToolHasItsHandlerActionPinned()}) and
* asserts the handler's call passes {@code coordinatorVisibleTo(principal(exchange))} — not a
* literal {@code true} or {@code false} — as {@code listFleet}'s trailing argument.
*
* <p>Anchored on argument position, not a bare substring search: {@code true} appears many
* times elsewhere in this file for unrelated reasons, so a plain {@code contains("true")}
* check would prove nothing. The pattern requires the literal text immediately before the
* closing {@code );} of the {@code listFleet(} call inside the handler block to be exactly
* {@code coordinatorVisibleTo(principal(exchange))}.
*/
@Test
void theFleetListHandlerActuallyConsultsCoordinatorVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
// Isolate the fleet_list handler block: from its declaration up to the next handler's
// declaration. A change to variable naming would break this scrape loudly (see the control
// assertion just below), rather than silently reporting "no violation found".
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
Pattern trailingArg = Pattern.compile(
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;",
Pattern.DOTALL);
Matcher m = trailingArg.matcher(handlerBlock);
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the "
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
String trailing = m.group(1);
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
"the fleet_list handler must ask coordinatorVisibleTo(principal(exchange)) who is "
+ "calling, not pass a literal boolean -- found: " + trailing);
}
// --- which action each tool hands the gate (fleetd #272) ------------------------------------
/**
@@ -626,8 +626,8 @@ class FleetMcpTest {
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
@@ -655,8 +655,8 @@ class FleetMcpTest {
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
@@ -676,6 +676,118 @@ class FleetMcpTest {
"an ordinary fleet's output must be unchanged by this feature");
}
/**
* fleetd #463: a compat overload called with no {@code callerIsPrimary} argument at all must
* fail closed, not open. Before this fix the hidden default was {@code true}, so a caller that
* forgot the argument silently got lead-to-lead coordination state. Lead coordination is fully
* configured here (a real channel, a real mailbox) specifically so this is not conflated with
* {@link #listOmitsTheCoordinatorRowWhenLeadCoordinationIsOff} -- the row is capable of being
* assembled, and the missing argument is the only reason it is not.
*/
@Test
void listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""),
"no callerIsPrimary argument must fail closed (absent), not open (present): " + out);
}
/**
* fleetd #439: a worker calling {@code fleet_list} must get a result with the {@code
* coordinator} key <strong>absent</strong> -- not an empty object, not a redacted one -- even
* though lead coordination is fully configured and would otherwise report a row. This drives
* the same {@code callerIsPrimary} value the MCP handler computes ({@code
* Principal.worker(...).isPrimary()}), so it pins the real production boolean, not a literal.
*/
@Test
void listOmitsTheCoordinatorKeyEntirelyForAWorkerEvenWhenLeadCoordinationIsOn() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary();
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out);
assertTrue(out.contains("\"members\""), out);
assertTrue(out.contains("\"healthCoverage\""), out);
}
/**
* fleetd #439 acceptance criterion 2: an architect gets exactly the same treatment as a worker.
* This is a real, executed test (not just reasoning by analogy) -- it drives the actual
* {@code Principal.architect(...).isPrimary()} value the production handler would compute for
* an architect caller, through the same gate a worker's call goes through.
*/
@Test
void listOmitsTheCoordinatorKeyEntirelyForAnArchitectToo() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary();
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
}
/**
* fleetd #439 acceptance criterion 3: an explicitly-primary caller sees the coordinator row
* fully assembled, with the same content #439 always produced for a primary.
*
* <p>fleetd #463 flipped the compat overloads' hidden default from {@code true} to
* {@code false} (fail closed), so the old "pre-#439 overload" this test used to compare
* against no longer stands in for a primary caller -- it is now exactly the implicit-default
* path #463 closes. Verifying the primary path means calling the canonical overload with an
* explicit {@code callerIsPrimary=true} directly, as the production {@code fleet_list} handler
* does.
*/
@Test
void listIsByteForByteUnchangedForThePrimaryCaller() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
String gatedAsPrimary = textOf(FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true));
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"mailbox\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"heldCount\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
}
@Test
void listReportsHeldMessagesWithATruncatedPreviewNeverTheFullBody() {
FakeHerdr h = new FakeHerdr();
@@ -691,8 +803,8 @@ class FleetMcpTest {
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"msgId\":\"m1\""), out);
@@ -723,8 +835,8 @@ class FleetMcpTest {
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"pending\":0"), out);
@@ -752,8 +864,8 @@ class FleetMcpTest {
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
String out = textOf(res);
assertTrue(out.contains("\"heldDurable\":false"),
@@ -820,8 +932,9 @@ class FleetMcpTest {
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")));
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true);
String out = textOf(res);
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
@@ -762,6 +762,250 @@ class OpenCodeLauncherTest {
"no IDE server when ideMcpUrl is unset");
}
// --- fleetd #393: memberSkills seeding must actually reach an opencode member -------------------
//
// Before this fix, GitWorktrees#seedSkills copied skill folders into EVERY provisioned
// worktree's .claude/skills/ and logged "skill seeding: N of M" regardless of which kind ended
// up spawning into that worktree — a claim that held for kind: claude-code (the CLI discovers
// that directory on its own) but was a guaranteed no-op for kind: opencode, which has no such
// discovery. These tests drive the REAL GitWorktrees#add seeding path (not a hand-built
// .claude/skills/ fixture), then spawn an opencode-kind member against the seeded worktree and
// assert on what the member can actually consume — an instructions[] entry — not on the
// seeding log alone. A minimal, non-hermetic git repo is enough here: unlike
// GitWorktreesTest's own seeding tests, nothing in this file cares about core.excludesFile
// composition, only about what lands in .claude/skills/ and whether OpenCodeLauncher reads it.
private static void git(Path cwd, String... args) throws Exception {
Process p = new ProcessBuilder(prepend("git", args)).directory(cwd.toFile())
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
}
private static List<String> prepend(String head, String... rest) {
List<String> cmd = new ArrayList<>();
cmd.add(head);
cmd.addAll(List.of(rest));
return cmd;
}
private static Path initRepo(Path dir) throws Exception {
Files.createDirectories(dir);
git(dir, "init", "-q", "-b", "main");
git(dir, "config", "user.email", "test@example.invalid");
git(dir, "config", "user.name", "Test");
Files.writeString(dir.resolve("README.md"), "seed\n");
git(dir, "add", "README.md");
git(dir, "commit", "-q", "-m", "seed");
return dir;
}
@Test
void aSeededSkillReachesTheOpencodeMembersInstructionsArray(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path skillsSource = tmp.resolve("skills-src");
Path skillFile = skillsSource.resolve("implementer").resolve("SKILL.md");
Files.createDirectories(skillFile.getParent());
Files.writeString(skillFile, "IMPLEMENTER PROCEDURE\n");
// The real seeding path (fleetd #362), not a hand-built .claude/skills/ fixture — proves
// OpenCodeLauncher reads what GitWorktrees#add actually produced.
dev.ltms.fleet.session.GitWorktrees worktrees =
new dev.ltms.fleet.session.GitWorktrees(tmp.resolve("wts").toString(), null, skillsSource.toString());
String wt = worktrees.add(repo.toString(), "cb-393-opencode", "HEAD");
Path seededSkillMd = Path.of(wt, ".claude", "skills", "implementer", "SKILL.md");
assertTrue(Files.exists(seededSkillMd),
"sanity: the real seeding step must have copied the skill into the worktree");
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
String cfgPath;
try {
// No mcpUrl, no ideUrl, no fleet (no charter): the seeded skill alone must be enough to
// trigger OPENCODE_CONFIG — proves the gate itself was updated, not only writeConfig's body.
FakeHerdr herdr = new FakeHerdr();
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
service(herdr, configRoot, opencodeIdeCfg(null, null, wt)).spawn();
cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
assertNotNull(cfgPath, "a seeded skill with nothing else configured must still write a config");
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
List<String> instructions = new ArrayList<>();
json.path("instructions").forEach(n -> instructions.add(n.asText()));
assertTrue(instructions.contains(seededSkillMd.toAbsolutePath().toString()),
"the seeded skill's SKILL.md must be an instructions[] entry — got: " + instructions);
List<String> infos = appender.list.stream()
.filter(e -> e.getLevel() == Level.INFO)
.map(ILoggingEvent::getFormattedMessage)
.toList();
assertTrue(infos.stream().anyMatch(m -> m.contains("skill delivery") && m.contains("1 of 1")),
"the launcher must log, kind-aware, that it delivered the skill — got:\n" + infos);
}
@Test
void aSkillFolderWithoutSkillMdIsNeverDeliveredAndTheLogNamesIt(@TempDir Path tmp) throws Exception {
Path wt = Files.createDirectories(tmp.resolve("wt"));
Path goodSkill = wt.resolve(".claude").resolve("skills").resolve("implementer");
Files.createDirectories(goodSkill);
Files.writeString(goodSkill.resolve("SKILL.md"), "GOOD\n");
Path halfShipped = wt.resolve(".claude").resolve("skills").resolve("half-shipped");
Files.createDirectories(halfShipped);
Files.writeString(halfShipped.resolve("README.md"), "no SKILL.md here\n");
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
String cfgPath;
try {
FakeHerdr herdr = new FakeHerdr();
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
service(herdr, configRoot, opencodeIdeCfg(null, null, wt.toString())).spawn();
cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
List<String> instructions = new ArrayList<>();
json.path("instructions").forEach(n -> instructions.add(n.asText()));
assertTrue(instructions.contains(goodSkill.resolve("SKILL.md").toAbsolutePath().toString()),
"the well-formed skill is still delivered alongside the malformed one");
assertFalse(instructions.stream().anyMatch(i -> i.contains("half-shipped")),
"a skill folder with no SKILL.md can never become an instructions[] entry");
List<String> infos = appender.list.stream()
.filter(e -> e.getLevel() == Level.INFO)
.map(ILoggingEvent::getFormattedMessage)
.toList();
assertTrue(infos.stream().anyMatch(m -> m.contains("skill delivery") && m.contains("1 of 2")
&& m.contains("half-shipped") && m.contains("could not be delivered")),
"the log must say plainly which folder could not be consumed and why — got:\n" + infos);
}
// --- fleetd #393 follow-up: instructions[] has three writers (charter, seeded skills, IDE
// rules), and no test above ever exercises more than one or two of them together. A writer
// that flips from withArray (get-or-create) to putArray (create-or-REPLACE) silently deletes
// every entry written before it — proven live on this branch's merge: switching just the
// skills writer to putArray left the entire suite (1603 tests) green while deleting the
// charter entry an opencode member needs for its role contract. That hazard was found by the
// fleet01 lead and independently verified against this branch; it is not a defect in the
// skills-delivery or logging tests above, which both hold up under their own mutations — the
// gap is that none of them combine all three writers in one config.
//
// These three tests assert instructions[] CONTENT as an exact, ordered list, not a size or a
// "contains" check: a putArray mutation can replace N entries with a different N entries of
// the same count, so only a content comparison can tell "all three paths present" apart from
// "two paths present that replaced the earlier ones".
@Test
void instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
Path cwd = Files.createDirectory(root.resolve("checkout"));
service(herdr, root, opencodeIdeCfg(null, null, cwd.toString()), () -> fleet).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "a role charter alone still writes a config");
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
assertTrue(Files.exists(charter), "the charter file was written");
List<String> instructions = new ArrayList<>();
json.path("instructions").forEach(n -> instructions.add(n.asText()));
assertEquals(List.of(charter.toAbsolutePath().toString()), instructions,
"with only the charter writer active, instructions[] holds exactly one entry: the "
+ "charter — got: " + instructions);
}
@Test
void instructionsArrayHoldsCharterThenIdeRulesInOrder(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
Path cwd = Files.createDirectory(root.resolve("checkout"));
service(herdr, root, opencodeIdeCfg(null,
"http://127.0.0.1:29170/index-mcp/streamable-http", cwd.toString()), () -> fleet).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "charter + IDE rules still writes a config");
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
Path rules = Path.of(cfgPath).resolveSibling("ide-rules.md");
assertTrue(Files.exists(charter), "the charter file was written");
assertTrue(Files.exists(rules), "the ide-rules file was written");
List<String> instructions = new ArrayList<>();
json.path("instructions").forEach(n -> instructions.add(n.asText()));
assertEquals(List.of(charter.toAbsolutePath().toString(), rules.toAbsolutePath().toString()),
instructions,
"with charter + IDE-rules writers active, instructions[] holds both, charter first — "
+ "got: " + instructions);
}
@Test
void instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path skillsSource = tmp.resolve("skills-src");
Path skillFile = skillsSource.resolve("implementer").resolve("SKILL.md");
Files.createDirectories(skillFile.getParent());
Files.writeString(skillFile, "IMPLEMENTER PROCEDURE\n");
dev.ltms.fleet.session.GitWorktrees worktrees = new dev.ltms.fleet.session.GitWorktrees(
tmp.resolve("wts").toString(), null, skillsSource.toString());
String wt = worktrees.add(repo.toString(), "cb-393-follow-up", "HEAD");
Path seededSkillMd = Path.of(wt, ".claude", "skills", "implementer", "SKILL.md");
assertTrue(Files.exists(seededSkillMd), "sanity: the real seeding step copied the skill");
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
service(herdr, configRoot, opencodeIdeCfg(null,
"http://127.0.0.1:29170/index-mcp/streamable-http", wt), () -> fleet).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "charter + skills + IDE rules still writes a config");
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
Path rules = Path.of(cfgPath).resolveSibling("ide-rules.md");
assertTrue(Files.exists(charter), "the charter file was written");
assertTrue(Files.exists(rules), "the ide-rules file was written");
List<String> instructions = new ArrayList<>();
json.path("instructions").forEach(n -> instructions.add(n.asText()));
// Exact ordered list, not size or "contains": a putArray mutation on any writer after the
// charter replaces every entry written before it, and the replacement can still be a
// plausible-looking array of a different shape. This is the one combination all three
// writers are active for — and per fleet01 the realistic shape on a host where weighted
// placement makes opencode the default for most members.
assertEquals(List.of(charter.toAbsolutePath().toString(),
seededSkillMd.toAbsolutePath().toString(),
rules.toAbsolutePath().toString()),
instructions,
"with all three writers active, instructions[] must hold charter, then the seeded "
+ "skill, then IDE rules — in that order and with nothing replaced. A "
+ "putArray mutation on any writer after the charter would silently drop "
+ "earlier entries here while still producing a same-shaped array — got: "
+ instructions);
}
// --- fleetd #219: config root + discovery root under memberHerdrSocket ------------------------
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
+369
View File
@@ -0,0 +1,369 @@
# Plan — move the fleet back onto fleet01
**Goal.** Stop the fleet depending on a laptop that sleeps.
**Written** 2026-09-05. **Rewritten the same day** after the operator pointed out that fleet01 is a VM
and already has the repo. They were right, and my first draft was wrong in an important way: this is
not a stand-up. **The whole fleet already ran on fleet01 in August.** It was abandoned, not attempted.
Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so.
---
## Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date
This plan is prior art now, not a to-do list. Every number in this section came from a read-only
survey over `ssh fleet01` on 2026-09-10. **Delete this section and the plan once fleet01 is the
fleet's only daemon** — at that point the plan has been executed and stops being useful.
| Plan item | State on 2026-09-10 | Command that re-measures it |
|---|---|---|
| Phase 1, refresh + build | **done, then drifted.** Checkout is on `main` at `4887731`, **77 commits behind** `origin/main`, 0 ahead. `fleetd/target/fleetd.jar` exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. | `git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main` |
| Phase 2, port the config | **done.** `fleetd/fleetd.yaml` exists on the host, with `bind`, `profiles`, `configReload`, `fleet`, `health`, `lifecycle`, `guard`, `memberCredentials`, `broker` and `coordinator` all present. | `test -f ~/LTMS/fleetd/fleetd/fleetd.yaml` |
| Phase 3, supervision | **done.** `~/.config/systemd/user/fleetd.service` and `herdr.service` both exist, both `active` and `enabled`. `loginctl show-user ltms -p Linger` prints `yes`. Exactly 1 java process, so the double-daemon problem in the old `restart.sh` is not present. | `systemctl --user is-active fleetd herdr` |
| Phase 4-5, reachable and a member proven | **partly.** `/healthz` answers 200 on loopback and reports `{"protocol":19,"version":"0.8.0"}`, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. | `curl -s http://127.0.0.1:8765/healthz` on the host |
| Phase 6-7, cutover and reboot proof | **not done.** The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. | `uptime -p` on the host |
| §9 "not moving the lead yet" | **out of date.** `claude` is installed at `/home/ltms/.local/bin/claude` and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. | `ssh fleet01 'command -v claude'` |
### The profiles fleet01 actually offers a member
Measured from the live `fleetd.yaml` on the host, with `placement: weighted`:
| profile | kind | model | weight | maxLoad |
|---|---|---|---|---|
| `gx` | opencode | `gx/deepseek-v4-flash` | 100 | 2 |
| `xf` | opencode | `opencode/mimo-v2.5-free` | 80 | 5 |
| `local` | claude-code | `deepseek-v4-flash` | 10 | 2 |
| `opus` | claude-code | `claude-opus-5` | 0 | 1 |
`weighted` spreads by ratio across every profile with a free slot, so an **unqualified** spawn on
fleet01 lands on `gx` or `xf` almost every time. Pass `profile` explicitly there, as the canonical
block already says.
### The PATH split on fleet01, which is not the one I expected
I went looking for a defect and found the opposite, so this is written down to stop the next
session repeating the search.
`opencode` is installed at `/home/ltms/.opencode/bin/opencode`. Whether a shell can see it depends
on which kind of shell it is:
```
zsh -ic 'command -v opencode' -> /home/ltms/.opencode/bin/opencode (1 PATH entry)
zsh -lc 'command -v opencode' -> nothing (0 PATH entries)
```
So it is on the **interactive** PATH (`.zshrc`), not the login one. Two consequences, and they
point in opposite directions:
- A herdr pane on Linux is a plain non-login interactive zsh, so a pane **can** launch `opencode`.
fleetd types the launch command into the pane rather than exec'ing it, so `gx` and `xf` are not
broken by this. I have not spawned one to confirm, so that is inference from the shell
measurement plus the typing behaviour, not an end-to-end result.
- The daemon itself runs `ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` —
a **login** shell, on purpose, because credentials live in `.zprofile`. Its own PATH has 10
entries and none contains `opencode`.
**The general shape: the login shell and the interactive shell see different PATHs, and which one
matters depends on whether fleetd types a command or execs it.** Credentials live on the login
side; `~/.opencode/bin` lives on the interactive side. Anything fleetd must exec itself is
invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the
login-shell `ExecStart` was added to fix.
---
---
## 1. My first draft was wrong — read this before the rest
I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place
and concluding about all of them. Three times:
| I checked | I concluded | What is actually true |
|---|---|---|
| `~/LTMS/claude-bridge` | "no repo checkout" | the repo is at `~/LTMS/fleetd` — the directory follows the **renamed** repo, and my own notes record that rename |
| `~/.config/herdr/herdr.sock` | "herdr never ran" | the socket lives at `~/.config/herdr/sessions/fleet01/herdr.sock` — a **named session**, and its server log runs to Aug 28 |
| both of the above | "fleetd is not deployed" | `~/LTMS/fleetd/fleetd-run/` holds `restart.sh`, `start-herdr.sh` and a `fleetd.out` from **Aug 24** |
The lesson is the one already written down here: enumerate one channel, conclude about all of them.
A single-path check is not a survey.
---
## 2. What already worked on fleet01, proven from its own log
`~/LTMS/fleetd/fleetd-run/fleetd.out` covers 06:07 to 16:48 on 2026-08-24. It shows:
```
4 distinct panes pane=c2504101-... w1:pC w1:pE
4 worktrees created and removed /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3}
4 allow-list decisions "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables"
AMQP on 127.0.0.1:5672 "AMQP connection recovered; cleared held replies for fresh redelivery"
the full turn machinery SessionManager transitions, ReplyPushLoop, CompletionResolver fallback
```
So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the
credential allow-list fed them their environment, and the broker link was **loopback**.
That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a
loopback connection.
### The two traps I was going to design around are already solved there
My first draft named these as the biggest risks. Both were already handled in August:
**Trap A — systemd sources no login shell.** `restart.sh` already starts through one, and says why in
its own comment:
```sh
# 1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a
# non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that
# stays invisible until a member actually needs it.
setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 &
```
**Trap B — a Linux herdr pane is a plain zsh, so members get no credentials.** Not a problem, and not
for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an
allow-list. The log proves it ran: `allowed 16 of 32 environment variables`, four times. The old
`fleetd.yaml` has a `memberCredentials:` block that configures it.
**The headless pty trap is solved too.** `start-herdr.sh` carries the fix and the explanation:
```sh
# Why the size matters: herdr creates each pane sized to the attached client's view. Started
# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create
# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface.
cat > /tmp/herdr-inner.sh <<'INNER'
stty rows 50 cols 200 2>/dev/null || true
exec herdr --session fleet01
INNER
setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 &
```
---
## 3. Why we are moving
The daemon runs on a Mac laptop. On battery it idle-sleeps after **one minute**:
```
pmset -g custom -> Battery Power: sleep 1
AC Power: sleep 0
```
Since the last restart: 16 `Connection reset` events and 6 ERROR lines, all on the AMQP link. I
matched every one of the 16 to the nearest sleep or wake event in `pmset -g log`. Largest gap **55
seconds**; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network
fault — it sees the client stop sending heartbeats, then closes.
Every reconnect worked, so no message was lost. **The lost messages are not the problem.** The problem
is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly
the case that goes idle.
---
## 4. How the two hosts relate today
This answers "why is fleet01 related to this Mac at all?"
```
Mac laptop fleet01 (KVM/QEMU guest)
┌────────────────────────────┐ ┌──────────────────────┐
│ Claude Code lead │ │ LavinMQ :5672 │
│ fleetd 127.0.0.1:8765 │ AMQP over │ vhost /mac │
│ herdr │──Tailscale────>│ vhost /fleet01 │
│ members + worktrees │ utun4, 1280 │ │
└────────────────────────────┘ │ (fleetd idle since │
│ Aug 24) │
└──────────────────────┘
```
**Today fleet01 runs only the broker.** Everything else — daemon, herdr, members — is on the Mac. The
single link between them is the Mac's fleetd opening AMQP to `10.10.20.13:5672` across Tailscale.
So the errors I reported were **the Mac's client dying when the Mac slept**, not fleet01 failing.
fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat
timeout each time, which is what a broker sees when a client vanishes.
After the move that arrow becomes loopback and the whole class of problem is gone.
---
## 5. The shape we are restoring
```mermaid
flowchart LR
subgraph MAC["Mac laptop (free to sleep)"]
LEAD["Claude Code lead"]
TUN["ssh -N -L"]
end
subgraph F01["fleet01 (KVM guest, always on)"]
FD["fleetd<br/>127.0.0.1:8765"]
HD["herdr --session fleet01"]
WT["members<br/>~/LTMS/.fleet-worktrees"]
MQ["LavinMQ<br/>127.0.0.1:5672"]
end
LEAD --> TUN
TUN -->|"ssh over Tailscale"| FD
FD -->|"unix socket"| HD
HD --> WT
FD -->|"AMQP, loopback"| MQ
WT -->|"AMQP, loopback"| MQ
```
*The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.*
Two constraints fix this shape:
1. **fleetd, herdr and the worktrees must share one filesystem.** The herdr link is a Unix socket plus
absolute path strings. `herdr --remote` is terminal attach, not a transport.
2. **fleetd fails fast on a non-loopback bind without token auth**, by its own design. So we do not
expose `:8765` to the `10.10.20.0/24` LAN. An SSH tunnel keeps the bind on loopback and needs no
new secret.
**Why the lead stays on the Mac for now.** A session not in a herdr pane resolves as `primary`, so a
Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type
into the lead's pane, and fleetd cannot type into a pane on another host — so `wait:false` tickets
stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is
authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block
this.
---
## 6. The actual gap
Everything below is what stands between "it ran in August" and "it runs supervised today".
| # | Gap | Measured state |
|---|---|---|
| 1 | **Checkout is stale** | branch `cb-634-ide-mcp`, HEAD `7655f1b` (2026-08-24), **290 commits behind** `origin/main`, 0 ahead |
| 2 | **Never rebuilt after the rename** | no `target/` anywhere; the scripts still say `bridged.jar` and `cd bridged`, but the tree is now `fleetd/` and `fleetd.jar` |
| 3 | **No live `fleetd.yaml`** | gitignored, so not in git. Three backups exist under `bridged/` — the newest is `fleetd.yaml.bak-cb634-pin`, 5512 bytes, and it is **clean of inline secrets** (0 inline passwords, 6 uses of `uriEnv`/`tokenEnv`) |
| 4 | **No supervision** | `Linger=no`; **zero** systemd user unit files. Only the hand-rolled `restart.sh` / `start-herdr.sh` |
| 5 | **Nothing running now** | no herdr process, no fleetd, no answer on `:8765/healthz` |
Untracked files in the checkout: `.idea/`, `fleetd-run/`, `docs/CB-634-Worker-IDE-Worktree.md`, and the
three yaml backups. All are **untracked, none modified**, and `7655f1b` is already an ancestor of
`origin/main` — so nothing is lost by updating the branch. Keep `fleetd-run/` and the backups; they
are the prior art this plan is built on.
### The old config's keys, which tell us what to port
```
bind: herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock
profiles: placement: weighted
configReload: fleet: health:
lifecycle: guard: worktreeRoot: /home/ltms/LTMS/.fleet-worktrees
memberCredentials: broker:
```
`herdrSocket` already points at the **named-session** path, and `memberCredentials` is already
configured. Those two are what made members work.
### What the repo already has for this
`deploy/fleetd.service` exists and is written for Linux. Three lines need fleet01's real paths:
`ExecStart` names `/usr/lib/jvm/temurin-25-jdk/bin/java` (fleet01 has `/usr/bin/java`), the `PATH`
names `/usr/share/maven/bin` (fleet01 has `/usr/bin/mvn`), and `WorkingDirectory` assumes
`%h/src/claude-bridge`. It also declares `After=herdr.service` — **and no `herdr.service` exists in
`deploy/`**. Writing that unit, from `start-herdr.sh`, is the one genuinely new piece of code here.
---
## 7. Phases
```mermaid
flowchart TD
P1["1. Refresh<br/>update + build"]
P2["2. Config<br/>port fleetd.yaml"]
P3["3. Supervise<br/>linger + 2 units"]
P4["4. Reachability<br/>tunnel, primary"]
P5["5. Prove a member"]
P6["6. Cutover"]
P7["7. Reboot proof"]
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7
```
**Phase 1 — refresh the checkout.** Update to `origin/main` (290 commits). Keep the untracked
`fleetd-run/` and the yaml backups. Then `mvn clean install`, **unpiped** — a pipe hides a failure
behind a zero exit. *Check:* `fleetd/target/fleetd.jar` exists and the suite is green.
**Phase 2 — port the config.** Write `fleetd/fleetd.yaml` from `bridged/fleetd.yaml.bak-cb634-pin`,
updating it for the rename and 290 commits of config changes. Diff its keys against
`fleetd/fleetd.example.yaml` on current main, key by key, and say what changed. *Check:* the daemon
starts and `journalctl ... | grep 'startup secret'` reports **no** `MISSING`.
**Phase 3 — supervision.** This is the part that never existed. `loginctl enable-linger ltms`; write
`deploy/herdr.service` from `start-herdr.sh`, keeping the `stty` sizing; fix the three path lines in
`deploy/fleetd.service` and keep the login-shell `ExecStart`; install both under
`~/.config/systemd/user/`. Secrets go in a `systemctl --user edit` drop-in or a 0600
`EnvironmentFile`, never in the committed unit. *Check:* log out of every ssh session, log back in,
and confirm the socket and the daemon are still there. That is what lingering is for, and it is the
check people skip.
**Phase 4 — reachability.** `ssh -N -L <port>:127.0.0.1:8765 fleet01` from the Mac. Use a **different
local port** for the first test so the Mac's own daemon on `:8765` is untouched and the whole test is
reversible. Point the lead's `.mcp.json` at it — that file is `--skip-worktree` and must never be
committed. Wrap the tunnel in `autossh` or a launchd `KeepAlive`, because the Mac still sleeps.
*Check:* `fleet_whoami` answers `primary`. If it answers `worker`, the `fleet.leaders.*.tab` pin does
not match — a known demotion, not a network fault.
**Phase 5 — prove a member.** `healthz` can be green while every spawn fails, so only a real spawn
proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is
what proves `WORKER_GITEA_TOKEN` resolved. Confirm the allow-list line still appears
(`allowed N of M environment variables`) and that N is what you expect. *Check:* a member on fleet01
opens a PR.
**Phase 6 — cutover.** Drain the Mac fleet properly first: `fleet_list`, `fleet_poll` anything still
wanted, then `fleet_stop` each member — a restart drops in-flight tickets and a member's report is
gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to
the Mac's settings, and it is reversible in one command.
**Phase 7 — reboot proof.** Reboot fleet01. Without touching anything: socket present, healthz
answering, `fleet_whoami` still `primary`, one spawn works. Until this passes, "supervised" is a claim.
---
## 8. Risks
| Risk | Why it bites | What this plan does |
|---|---|---|
| **290 commits of config drift** | the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error | phase 2 diffs key-by-key against current `fleetd.example.yaml` |
| **A new config key gets silently dropped** | `FleetConfig`'s back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away | separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml |
| **Nothing supervises herdr** | `deploy/fleetd.service` depends on a unit that does not exist | phase 3 writes it from the working script; phase 7 proves it |
| **Linger left off** | everything dies at logout and looks fine until then | phase 3, checked by logging out |
| **Wrong JDK/Maven path in the unit** | fleetd propagates its PATH to every member, so a bad PATH means no member can build | three lines fixed in phase 3, proven by phase 5 |
| **Headless pty with no winsize** | `ghostty error -2`, reported three steps later as a spawn failure | the `stty` fix is carried into `herdr.service` |
| **Port 8765 collides during the test** | both daemons want the same local port | phase 4 uses a different local port first |
| **Tunnel dies when the Mac sleeps** | same sleep, far smaller blast radius — it interrupts my session, not members | `autossh`/launchd `KeepAlive` |
| **Lead demoted to worker** | the tab pin no longer matches | phase 4's check is `fleet_whoami` |
| **Upgrading herdr** | 0.8.0 is **protocol 19**, pinned on purpose — 0.8.2 is protocol 20 and fleetd has **no version handshake** | do not upgrade herdr during this work; both hosts measured at 0.8.0 today |
| **`placement: tab` headless** | fails for the same 0x0 reason as `ghostty error -2` | fleet01 must keep `placement: pane`, as its August config did |
| **Profile launch settings are deferred** | editing `placement:` and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot | restart the daemon after those keys, do not wait for the reload |
---
## 9. Deliberately not doing
- **No changes to the Mac's power settings or host config.** The point is to stop depending on it.
- **Not touching the leftover `bridged-lavinmq` container** on the Mac. It is unused and harmless.
- **Not moving the broker.** It is already on fleet01 and already the durable one. After the move its
connection becomes loopback, which is the fix.
- **Not building a second fleet.** This is a move. Two daemons on one herdr session kill each other's
members.
- **Not moving the lead yet.** That needs Claude Code authenticated on a headless Ubuntu box and a
`fleet.leaders.*.tab` pin on its pane. It buys back pane nudges. It has nothing to do with sleep, so
it must not hold up phases 1–7. Note that fleet01 already carries an `opus` profile defined purely so the lead slot resolves; it cannot spawn until someone runs `claude` and completes `/login` on the host.
---
## 10. Open questions for the operator
1. **vhost** — keep `/mac`, or rename now the fleet is not on the Mac? Renaming loses the existing
queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.)
2. **Fallback week** after cutover, or stop the Mac daemon for good?
3. **Delete or keep the stale `cb-634-ide-mcp` branch** on fleet01 once the checkout is updated? Its
tip is already in main, so nothing is lost either way.
The repo path question from the first draft is answered: **`/home/ltms/LTMS/fleetd`**, which already
exists. `deploy/fleetd.service` should be pointed there rather than the reverse.