Compare commits

...

911 Commits

Author SHA1 Message Date
ltms 26f198675c Merge pull request 'fleetd #612 Unit A: extract main's boot composition into FleetdAssembly/FleetdRuntime' (#620) from worker/fleetd-612-unita-87807e-1 into main
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m30s
2026-09-22 07:43:39 +02:00
lead 72f46d7c0c fleetd #612: merge main into Unit A, carrying #621's requireOperatorConfirm into the assembly
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m59s
Resolves the one conflict in Fleetd.java. main's side is the inline boot block
that Unit A had already moved into FleetdAssembly.assembleAndStart, so the
resolution keeps Unit A's single assembly call.

That resolution is not purely mechanical. #622 (fleetd #621) landed on main
AFTER Unit A forked, and it added

    boolean requireOperatorConfirm = cfg.leadRollover() == null
            || cfg.leadRollover().requireOperatorConfirm();

plus a 14th argument to the LeadHeartbeatLoop constructor, inside the very block
Unit A moved. Taking Unit A's side alone would have dropped both and silently
reverted the operator's #621 fix: the 13-argument overload still exists and
delegates with `true`, so the daemon would go back to telling every lead to ask
the operator before a context roll. Both are carried into FleetdAssembly here.

Measured: with the carried line removed, the full suite is
`Tests run: 1883, Failures: 0, Errors: 0` — nothing pins it. That is the fleetd
#612 defect shape applied to #621's own wiring, and it is filed separately
rather than fixed here, because this commit is a merge resolution and must not
also introduce new tests.

Full suite on this resolved tree: Tests run: 1883, Failures: 0, Errors: 0.
2026-09-22 12:43:22 +07:00
ltms 640f4d5f23 Merge pull request 'fleetd #612 B3: behavioural replacements for lead-seat, quarantine, lead-rollover guards' (#628) from worker/612-b3-mcpwirings-da2b58-3 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 12s
CI / contract (pull_request) Successful in 1m33s
CI / build (pull_request) Failing after 2m25s
2026-09-22 07:36:05 +02:00
Dai Ha 6edeb70bc4 fleetd #612 B3 correction: distinguish lead vs member herdr in FleetdLeadRolloverAssemblyTest
Ticket comment 17553 on fleetd #612 found that the test's single shared
FakeHerdr made router.leadAgents() and router.memberAgents() collapse to
the identical client (FleetdAssembly.java:140-142's no-distinct-socket
fallback), so a mutation swapping leadAgents() for memberAgents() at the
FleetdAssembly.java:408 call site was invisible to this test even though
the two are genuinely different daemons in production.

Configure two distinct herdr sockets and two distinct FakeHerdr instances
(the same TwoHerdrResourcePorts shape B2's FleetdAssemblyConnectionIdentityTest
uses) and assert the roll's /clear + bootstrap sends land on the LEAD fake
and never on the MEMBER one.

Proven red against the router.memberAgents() mutation, reverted, touched,
and re-run green — both outputs recorded in the PR.
2026-09-22 12:31:39 +07:00
Dai Ha bc49d87cb8 Merge remote-tracking branch 'origin/worker/fleetd-612-unita-87807e-1' into worker/612-b3-mcpwirings-da2b58-3 2026-09-22 12:27:47 +07:00
ltms 14b169c410 Merge pull request 'fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair' (#626) from worker/612-b2-cb185-176d3a-2 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 2m6s
2026-09-22 07:19:28 +02:00
Dai Ha 2b52324d9a fleetd #612 step 2 unit B3: behavioural replacements for lead-seat, quarantine
and lead-rollover source-text guards

Replaces three FleetdAssembly.java source-text guards (each scraped
Fleetd.java for a call site that fleetd #612 Unit A moved into
FleetdAssembly.java) with tests that drive the real assembled objects
through FleetdAssembly.assembleAndStart(...) -> FleetdRuntime.mcp(),
per the step-2 B-unit split (issue #612 comment 17513).

- Deleted FleetdLeadSeatWiringTest (fleetd #176): pinned that
  FleetMcp's LeadSeatSource construction still wires
  Fleetd.leadSeatLookup(...) by scraping the constructor call's text.
  Replaced by FleetdLeadSeatAssemblyTest, which seeds one FakeHerdr
  tab labelled to match a configured fleet.leaders.opus.tab and
  asserts the REAL assembled LeadSeatSource (via
  runtime.mcp().leadSeatSource()) reports the live lead's seat against
  its own subscription profile -- 1, not the 0 LeadSeatSource.none()
  (the inert stand-in) could ever report.

- Deleted FleetdBackendQuarantineWiringTest (fleetd #466): pinned that
  the escalating BackendQuarantine.withEscalation(...) text was
  present and the flat two-argument constructor's text was absent.
  Replaced by FleetdBackendQuarantineAssemblyTest, which quarantines
  the same credential twice through the REAL assembled
  BackendQuarantine (via runtime.mcp().quarantineSource().quarantine())
  at controlled fake-clock offsets and asserts the second cooldown
  doubles (200s vs 100s) -- the one behavioural difference escalation
  and the flat constructor actually produce.

- Deleted FleetdLeadRolloverWiringTest (fleetd #480), all three
  methods: unrelatedAnchorStillPresent was a scaffold anchor with no
  independent claim, needing no replacement.
  mainStillCallsTheLeadRolloverFactory pinned the leadRollover
  assignment's call-site text. factoryGatesOnConfigPresence pinned
  that an absent leadRollover: config yields no LeadRollover.
  Replaced by FleetdLeadRolloverAssemblyTest's two tests:
  assembledLeadRolloverRunsTheRealClearAndBootstrapSequence drives the
  REAL assembled LeadRollover (via runtime.mcp().leadRollover())
  through open()/confirm() end to end and asserts /clear then
  bootstrapText were actually sent through the real herdr router,
  reaching ROLLED. absentLeadRolloverConfigMeansNoRolloverIsBuilt
  calls Fleetd.leadRollover(...) directly with no leadRollover: block
  and asserts null -- this claim was found uncovered elsewhere
  (LeadRolloverTest's only related assertion is vacuous, assertNull
  (null), and never calls the real factory).

Each of the three FleetdAssembly.java call sites (quarantine
line 179-180, leadRollover line 408, lead seats line 479) was mutated
to its named inert variant, run against ONLY its new test (RED),
reverted, touch'd (Maven mtime trap) and re-run (GREEN) -- six proven
runs, pasted in the PR body.

FleetMcp.java: adds three accessors (quarantineSource(),
leadSeatSource(), leadRollover()) alongside the existing
registeredTools() -- but public, not package-private, and this is a
deliberate deviation from that precedent, not an oversight: these new
assembly tests cannot live in package dev.ltms.fleet.mcp the way
registeredTools()'s callers do, because they also build the
ResourcePorts FleetdAssembly.assembleAndStart(...) needs, and
ResourcePorts' methods return Fleetd-nested types visible only from
package dev.ltms.fleet. Package-private would compile but be
unreachable from there.

Full mvn -o test in fleetd/: Tests run: 1879, Failures: 6 (down from
the branch baseline's 1880/9 by exactly the 3 guards this unit
deletes) -- the remaining 6 are FleetdCompletionResolverWiringTest (4)
and FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest (1 each), all out of this unit's scope
(B1/B2).
2026-09-22 12:18:48 +07:00
ltms b6147a39f6 Merge pull request 'fleetd #612 Unit B1: CompletionResolver behavioural test (replaces source-text guard)' (#627) from worker/612-b1-completion-457459-1 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 1m37s
2026-09-22 07:14:25 +02:00
Dai Ha dbfc34cb6d fleetd #612 B2 fixup: cover the symmetric FleetApp daemon-drop mutation
Per ticket comment 17525: my second independent mutation on
FleetdAssembly.java:518 survived — new FleetApp(memberHerdr, memberHerdr, ...)
(dropping the LEAD client instead of the member one) left both existing
FleetdAssemblyFleetAppTest cases green. That is the symmetric form of the
CB-185 defect (/healthz green while the LEAD daemon is down), and the guard
this PR deletes would have caught it: its positive assertion required the
exact pair "new FleetApp(herdr, memberHerdr, workers,", which does not
survive either daemon being dropped.

Adds healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp, symmetric
to the existing member-down case.

Proven red today: reverting FleetdAssembly.java:518 to
"new FleetApp(memberHerdr, memberHerdr, workers, ..." and running only
FleetdAssemblyFleetAppTest gives Tests run: 3, Failures: 1 — the new case
fails ("expected: <503> but was: <200>", body has no "member" key); the other
two cases stay green. Reverted, git diff --stat empty, file touched, re-ran:
Tests run: 3, Failures: 0.

Full mvn -o test: Tests run: 1884, Failures: 7 (same 7 B1/B3-scope failures
as before this fixup; 1884 = 1883 + 1 new case).
2026-09-22 12:11:28 +07:00
Dai Ha a20cb96730 fleetd #612 Unit B1: replace CompletionResolver source-text guards with behavioural tests
FleetdCompletionResolverWiringTest read Fleetd.java's literal source text and
asserted the CompletionResolver construction call still named the right
arguments — proof of spelling, not behaviour. Unit A (FleetdAssembly) moved
that call site out of Fleetd.main, breaking all four of its tests on a
harmless relocation.

FleetdCompletionResolverAssemblyTest replaces it, driving the real
FleetdAssembly.assembleAndStart(...) and reading FleetdRuntime.completion() —
the exact CompletionResolver instance production uses — through its public
onDelivered/resolveBeforePostAction API, with a controllable ResourcePorts
nanoClock in place of real sleeps.

Deleted test -> what it pinned -> replacement:
- worktreeBranchLookupIsStillPassedAtTheCallSite (8th constructor arg) ->
  assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport: a
  real git-worktree-provisioned MemberSession's branch must appear in a
  noReportMessage fallback (fleetd #241).
- backendErrorArgumentsAreStillNamedAtTheCallSite,
  backendErrorPatternsComesFromTheFactory, backendErrorSinkComesFromTheFactory
  (5th/6th args) -> assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern:
  a configured errorPattern the built-in fallback never matches must classify
  as FAILED (not COMPLETION), transition the session to BACKEND_ERROR, and
  cool the credential off after two distinct targets within the window
  (fleetd #201 Unit 5).

Both behaviours were proven red today by mutating FleetdAssembly.java's real
call site to its inert variant (_ -> null; BackendErrorPatternLookup.legacy();
BackendErrorSink.none()), confirming the new test failed with the expected
message, then reverting (touching the file to defeat Maven's stale-mtime
skip) and confirming green again. FleetdAssembly.java itself is unchanged in
this commit.

Full suite: Tests run: 1878, Failures: 5 (down from the baseline 9 — the
remaining 5 are the other in-flight workers' own guard files:
FleetdBackendQuarantineWiringTest, FleetdConnectionIdentityConstructionTest,
FleetdFleetAppConstructionTest, FleetdLeadRolloverWiringTest,
FleetdLeadSeatWiringTest), Errors: 0, Skipped: 0.
2026-09-22 12:10:38 +07:00
Dai Ha cda1a6a917 fleetd #612 step 2 unit B2: behavioural replacements for the CB-185 pair
Deletes the two source-text guards fleetd #612 Unit A broke by moving their
scraped call sites from Fleetd.java into FleetdAssembly.java, replacing each
with a behavioural test that drives the real assembled graph instead.

- FleetdConnectionIdentityConstructionTest pinned that Fleetd.java contained
  "new PaneLocator(herdr, memberHerdr)". Replaced by
  FleetdAssemblyConnectionIdentityTest, which drives the real PaneLocator a
  real FleetdAssembly.assembleAndStart(...) built (reached via
  runtime.mcp().identity().panes(), never a copy) with two distinct FakeHerdr
  daemons, and proves it finds a pane that exists on only one of them —
  first the lead-only case (the CB-185 bug: a lead's own connection going
  unresolvable), then the member-only case, plus a no-match control.

- FleetdFleetAppConstructionTest pinned that Fleetd.java contained
  "new FleetApp(herdr, memberHerdr, workers,". Replaced by
  FleetdAssemblyFleetAppTest, which binds the real Javalin app
  FleetdAssembly built (runtime.app()) to a real ephemeral port and proves
  GET /healthz goes 503 when only the member daemon is down — the consequence
  named in the deleted test's javadoc (a down member daemon invisible behind
  a healthy lead). The GET /sessions merge half of that javadoc could not be
  driven the same way: it requires Authz.Action.READ, which needs a real
  positive pid from the real assembly's hardcoded LsofPeerPidLookup, and an
  in-process test's HTTP client and server share one JVM pid so that pid is
  always -1 (fleetd #317's fail-closed rule then refuses the request before
  the route, and its merge, is ever reached) — documented in the new test's
  javadoc; FleetAppTwoDaemonTest remains the full behavioural proof that
  FleetApp itself merges /sessions correctly given two clients.

Both replacements were proven red today: reverting FleetdAssembly.java:443-444
to "new PaneLocator(memberHerdr)" turns the identity test red (expected
"term_a", got null); reverting :518 to "new FleetApp(herdr, herdr, ..." turns
the app test red (expected 503, got 200 with no "member" key). Both mutations
reverted, tree confirmed clean, and the mutated file touched afterward so
Maven does not skip recompiling a stale .class.

Adds two small production accessors needed to reach the real objects rather
than a copy, since FleetdRuntime may not gain a field (three workers touch
that file): ConnectionIdentity#panes() exposes the PaneLocator it resolves
against, and FleetMcp#identity() exposes the ConnectionIdentity it was built
with (now also stored as a field).

Baseline on 608e449: 1880 tests, 9 failures (the known source-text set).
After this change: 1883 tests, 7 failures — the remaining #248/#176/#466/#480
guards, which are B1's and B3's scope, not this one's.
2026-09-22 12:03:38 +07:00
ltms 608e4496be Merge pull request 'fleetd #612 A-gaps: coordinator path + reportRoleFallbackGaps inside the assembly boundary' (#624) from worker/612-agaps-73a926-2 into worker/fleetd-612-unita-87807e-1
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m34s
CI / build (pull_request) Failing after 1m42s
2026-09-22 06:41:42 +02:00
ltms fa61dc587c Merge pull request '#608 replace MessageService timing sleeps' (#623) from worker/608-sleeps-3a64ff-3 into main
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 3h14m55s
2026-09-22 06:38:29 +02:00
Dai Ha 526e3b7459 fleetd #612 A-gaps: exercise the coordinator path and pull reportRoleFallbackGaps inside the assembly boundary
Gap 1: generalise Fleetd.LeadMailboxOpener (and FleetdRuntime's field) from the concrete
LeadMailbox to a new closeable LeadChannelHandle (LeadChannel + AutoCloseable), so a test can
fake the configured-coordinator path without a real broker. FleetdAssemblyCoordinatorLifecycleTest
drives FleetdAssembly.assembleAndStart with a coordinator: block and a fake channel, proving the
assembly builds it and shutdown closes it.

Gap 2: move reportRoleFallbackGaps(cfg) and assertChartersNameOnlyRegisteredTools(cfg) out of
Fleetd.main and into FleetdAssembly.assembleAndStart, immediately after cfg.validateAll(), so
both run inside the tested assembly boundary before any I/O. FleetdAssemblyRoleFallbackBoundaryTest
pins the moved call site behaviourally (captured log output), not by reading source text.

Both new tests were verified red under a targeted mutation (a non-closing/null coordinator for
gap 1; deleting the moved call for gap 2) and restored.
2026-09-22 11:38:19 +07:00
Dai Ha b6b006c651 #608 replace MessageService timing sleeps
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Failing after 2m27s
2026-09-22 11:33:48 +07:00
ltms 63eec8a0da Merge pull request 'fleetd #621: make the context-roll notice obey requireOperatorConfirm' (#622) from worker/621-b4520b-1 into main
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m11s
CI / build (push) Failing after 1m52s
2026-09-22 06:31:02 +02:00
Dai Ha cbe872b538 fleetd #621: make the context-roll notice obey requireOperatorConfirm
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m0s
CI / build (pull_request) Failing after 1m53s
contextNotice hardcoded 'ask the operator' and 'Only the operator can
approve the roll', so setting leadRollover.requireOperatorConfirm to
false stopped the daemon refusing the roll but never stopped the lead
being told to ask. Thread the effective config value into
contextNotice: when true the text stays byte-identical, when false it
tells the lead to confirm on its own judgement against the three
handover-file checks instead.

LeadRollover.confirm's own enforcement is untouched — this is the
message only.
2026-09-22 11:27:17 +07:00
Dai Ha 7d9a807243 handover skill: the rollover bootstrap is proven, and requireOperatorConfirm is per-host
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 2m33s
Two bullets in the handover skill were telling every outgoing lead something
that is no longer true.

1. The skill said the bootstrap prompt "has never yet landed, and the fix is
   unproven (fleetd #489)", and told the lead to warn the operator it may fail.
   Measured today from fleetd/fleetd.out:

     grep -c "lead-rollover: rolled" -> 4
     grep -c "lead-rollover:"        -> 16   (positive control)
     grep -c "Unknown command"       -> 0

   Three of the four rolls ran on 2026-09-22 (10:01:43, 10:38:28, 11:15:47).
   Each cleared the old lead and bootstrapped a fresh one against the handover
   file. The old "Unknown command: /clearFresh" failure does not appear at all.
   The paragraph now carries the measured result and the three re-measure
   commands, including the control line, because a broken grep pattern returns
   a clean 0 that reads like good news.

   It also records that the "/clear was never observed as WORKING ... releasing
   rather than wedging the roll" WARN accompanies every successful roll. That
   is the safe branch, not a failure, and it was being misread as one.

2. The skill said requireOperatorConfirm "defaults to true and this is the only
   thing standing between a judgement call and a wiped session", which reads as
   if asking is always required. The default is still true
   (FleetConfig.java:1426), but this host set it to false on 2026-09-22 on the
   operator's explicit grant. The bullet now says to read the live value rather
   than assume, and notes the key is deferred, not hot.

   It also warns that until fleetd #621 merges, LeadHeartbeatLoop.contextNotice()
   still hardcodes "ask the operator" and takes no config, so the nudge text and
   the config disagree. Trust the config. That warning names the ticket that
   removes it.

Documentation only. No code or test changes.
2026-09-22 11:24:03 +07:00
ltms 8915e40c7d Merge pull request 'fleetd #618: state the measured auto-compact precedence' (#619) from worker/618-b83894-2 into main
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 2m12s
2026-09-22 05:53:35 +02:00
Dai Ha 3f7bc3815e fleetd #612 Unit A: extract Fleetd.main's boot composition into FleetdAssembly/FleetdRuntime
CI / shell-tests (pull_request) Failing after 12s
CI / build (pull_request) Failing after 1m29s
CI / contract (pull_request) Successful in 1m28s
Fleetd.main kept config loading, startup reports and validation. Everything from
the herdr socket connect onward moved verbatim, same order, into
FleetdAssembly.assembleAndStart(AssemblyInputs, ResourcePorts), which returns a
FleetdRuntime owning the real objects (package-private accessors, never a copy)
and their single ordered close(). ResourcePorts/SystemResourcePorts abstract every
boot-time side effect (env, herdr connect, broker openers, clocks, schedulers,
shutdown-hook registration, HTTP start) with no inert production variant, per the
architect proposal on the ticket.

sleepHerdrPoll widened from private to package-private so FleetdAssembly can pass
a method reference to it; no other signature changed.

FleetdAssemblyLifecycleTest drives the real assembly with FakeHerdr, a temp
FleetConfig and a fake ResourcePorts recording a start/close ledger, asserting it
against the order recorded from the pre-move main() and shutdown hook, and proving
every resource the ledger can observe (herdr client, three schedulers, the AMQP
reply inbox) closes via FleetdRuntime.close(). FakeHerdr gained a closed flag for
this.
2026-09-22 10:52:35 +07:00
Dai Ha 6cb31a10e4 fleetd #618: fix the third stale spot the brief missed
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Failing after 2m6s
The method-level javadoc on FleetConfig.warnConflictingAutoCompactWindows
(above the log.warn call) still claimed the autoCompactWindow vs
CLAUDE_CODE_AUTO_COMPACT_WINDOW precedence was 'intentionally not
asserted' and cited fleetd.yaml's now-corrected comment as evidence the
question was open. Replace it with the measured answer from #618: the
env var wins, so autoCompactWindow is inert on a profile that sets both.
Kept the WARN-not-throw rationale paragraph above it untouched (#601)
and kept the ClaudeCodeArguments cross-reference, which now points to an
agreeing claim instead of a contradicting one. No behaviour change.
2026-09-22 10:50:59 +07:00
Dai Ha 8368a274a0 fleetd #618: state the measured auto-compact precedence, not 'unverified'
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 1m57s
ClaudeCodeArguments.withAutoCompactWindow's javadoc and FleetConfig's
warnConflictingAutoCompactWindows WARN text both used to say the
precedence between --autocompact and CLAUDE_CODE_AUTO_COMPACT_WINDOW was
not verified. fleetd #618 measured it: the env var wins, so the flag has
no effect when both are set. Update both texts to say so, name #618, and
warn that deleting the env var to resolve the conflict LOWERS the live
window rather than fixing anything. No behaviour change; the WARN still
fires on the same condition and stays a WARN (per #601).
2026-09-22 10:46:31 +07:00
ltms 17127efb88 Merge #601: pass auto-compact window to leads; warn instead of refusing on a conflict (CB-617)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m16s
CI / build (push) Failing after 2m33s
2026-09-22 05:22:56 +02:00
ltms 203f034528 Merge #617: write FAILED instead of leaving a dead roll stuck at IN_PROGRESS (fleetd #615)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m14s
CI / build (push) Failing after 1m47s
2026-09-22 05:21:21 +02:00
ltms 9ee16f5b85 Merge #616: report role-fallback gaps at boot, name contextHighNudge (fleetd #613)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m15s
CI / build (push) Failing after 1m43s
2026-09-22 05:17:52 +02:00
Dai Ha 388ef5a3c3 fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m25s
LeadRollover.runRollover made two unwrapped agents.send calls. HerdrException
is unchecked, and the production continuationRunner is a bare virtual thread
with no uncaught-exception handler, so a throw from either send call killed
the continuation silently — confirm() had already written IN_PROGRESS into
outcomes before scheduling it, and nothing ever overwrote that entry with a
terminal state.

Wrap the whole continuation body in one try/catch(RuntimeException), matching
the local convention already used around agents.status in
waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED
entry naming the exception, in the same diagnostic style as
TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED.

Two new tests make send() throw on the /clear call and on the bootstrap-text
call respectively, each asserting status(token) reports FAILED, not
IN_PROGRESS. Reverting only the production catch (keeping the tests) turns
both red; restoring it turns them green again.
2026-09-22 10:17:12 +07:00
Dai Ha be6c45ff78 CB-617 review: warn instead of refuse on conflicting autoCompactWindow
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m9s
rejectConflictingAutoCompactWindows threw and stopped fleetd from starting when a
Claude Code profile's autoCompactWindow flag and CLAUDE_CODE_AUTO_COMPACT_WINDOW
env var disagreed. Under launchd that is a restart loop, and the config that
would fix it (fleetd.yaml) is gitignored, so the cause is invisible on the host
where it bites (measured live: 4 profiles on this host trip it, including the
lead's own profile and the one every worker spawns on).

Renamed to warnConflictingAutoCompactWindows: it now logs a WARN naming each
offending profile with BOTH values (autoCompactWindow=... and
env.CLAUDE_CODE_AUTO_COMPACT_WINDOW=...) instead of throwing, so the daemon
starts and an operator can fix the config without reading the source. Equal
values still load silently.

Also reworded ClaudeCodeArguments' javadoc, which stated as fact that the env
var takes precedence over the flag. That was never measured, and this host's
own fleetd.yaml comment asserts the opposite — the javadoc no longer picks a
side.
2026-09-22 10:13:14 +07:00
Dai Ha e99cb70a8b CB-617: pass auto-compact window to leads 2026-09-22 10:12:56 +07:00
Dai Ha 987ccef4c7 fleetd #613: log role-fallback gaps at boot, name contextHighNudge in the heartbeat line
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Failing after 1m48s
- reportRoleFallbackGaps(cfg), called right after cfg.validateAll() in Fleetd.main, logs every
  MemberRole with no fleet.<role>s: pool (naming the profile count and the resolved
  defaultProfileFor(role) first choice) and, separately, every role with no
  fleet.charters.<role>: entry. Log only — the deliberate 'unconstrained' fallback in
  FleetConfig#candidateProfiles / CompositePeerLauncher#poolFor is unchanged, and a config with
  profiles: and no fleet: block still starts and still spawns.
- LeadHeartbeatLoop#start()'s boot line now also names contextHighNudge (fleetd #609), alongside
  the three settings it already logged.
- RoleFallbackGapReportTest (new) and two new LeadHeartbeatLoopTest cases pin both lines' content
  via a ListAppender, raising the dev.ltms.fleet logger past logback-test.xml's WARN override for
  the INFO-level lines.
2026-09-22 10:12:09 +07:00
ltms 076cc43f7b Merge #614: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 2m20s
CI has been red on main itself since #602/#606, on this one test, so the build has been giving no second opinion on any PR. Cause: the CI job runs in a container as root. setReadable(false) really does clear the read bit, so the test's own setup guard passes, but root opens the file anyway and the gauge correctly returns OK. The test was asserting on a condition the environment never created.

The fix adds an assumeFalse(Files.isReadable(file), ...) after the chmod and before the gauge is built, inside the existing try, so the finally still restores the bit on a skip.

Verified by me on a scratch worktree merging this onto 955b9ea:
- 1864 tests, 0 failures, 0 errors, 0 skipped, 149 surefire reports, mvn exit 0. The suite-wide skipped=0 is the point: the fix did not quietly turn the test into a permanent skip.
- LeadContextGaugeTest on this non-root Mac: 9 tests, 0 skipped, and unreadableFileIsUnknown present in the report. The assumption does not fire here, so developers keep the coverage.
- Mutation: made the IOException path return OK instead of UNKNOWN. unreadableFileIsUnknown failed with "expected: <UNKNOWN> but was: <OK>". The test still has teeth. Production file reverted, git diff clean before merge.

Known trade, recorded rather than hidden: under root this case is now covered by nothing at all. A skip is honest about that, which an assertion on an unreachable state was not. The durable fix is to run the CI build as a non-root user; that is a CI configuration change and out of scope here.
2026-09-20 12:43:43 +02:00
Dai Ha bad47a8444 fleetd CI: skip unreadableFileIsUnknown honestly when root ignores the read bit
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m27s
The test set the file's read bit off via setReadable(false), but on the
Gitea CI runner (root inside the container) the OS ignores that bit and
opens the file anyway, so the test asserted on a condition it never
actually created (LeadContextGaugeTest.java:142 UNKNOWN vs OK, CI run
1887 job 3104, commit fa62e99 on main).

Add Files.isReadable(file) after setReadable(false) and before the gauge
runs, and assumeFalse on it: a skip means "could not set up the case",
never "the behaviour is fine". Restores the read bit either way so
@TempDir cleanup still works.
2026-09-20 17:40:47 +07:00
ltms 955b9ea013 Merge #610: nudge an idle lead to hand over when its own context reads HIGH (fleetd #609)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m51s
Closes fleetd #609. Completes the second half of the context work: #602/#606 could detect a full lead context, and nothing acted on it. LeadHeartbeatLoop now offers a handover when the lead's own gauge reads HIGH.

Never rolls a pane by itself. The nudge is text only; the lead still has to call fleet_handover, and that still needs operatorConfirmed.

Verified by me on a scratch worktree merging 89cb8ff onto fa62e99:
- 1864 tests, 0 failures, 0 errors, 149 surefire reports, mvn exit 0.
- Four mutations, all killed: the two the implementer ran (latch back in applyDecision: 4 failures; call site drops the latch argument: 2 failures), one of my own at the line the logic moved TO (latch set regardless of send outcome: 1 failure), and a control on an untouched line (quiet-cap boundary < to <=: 4 failures). The control is what makes the other kills evidence.
- The new tests assert on herdr.sentTexts() — what actually reached the fake pane — not on source text. That is the right observable for a defect whose essence was "the latch says told, the pane got nothing".

The review blocker from the first round is fixed: the latch used to be committed by applyDecision before injectNudge tried to send, and injectNudge swallows its own RuntimeException. On quietNudgeCap: 0, which is this host's configuration, that was the normal path and not an edge case. The latch is now set only when a notice was actually included and the send returned.

Known and deliberately not blocked: Fleetd.main's own one-line call to leadContextSource is not pinned by a test. That is pre-existing and class-wide, tracked in #612.
2026-09-20 12:37:43 +02:00
Dai Ha 89cb8ff79b fleetd #609 review: the context latch must mean the notice reached the pane
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m6s
Fixes the PR #610 review blocker: LeadHeartbeatLoop committed contextNotified
before injectNudge attempted the send, so a transient herdr failure marked the
lead as told when nothing reached its pane, and contextNotice() carried no
latch at all, so a pending-driven INJECT re-appended the notice on every tick
while the context stayed HIGH.

- injectNudge now reports whether agents.send succeeded and persists
  contextNotified only when a notice was actually included in the text and
  the send did not throw. The latch is split out of applyDecision (kept for
  idleSinceNanos/quietCount, applied unconditionally as before) so it is
  written on the success path only, once per branch in tick().
- contextNotice gained an overloaded 3-arg form gated on the latch as it
  stood before the tick's decision; the existing 2-arg form delegates to it
  with alreadyNotified=false, so all pre-existing callers/tests are unchanged.
- tick() is now package-private (mirrors ReplyPushLoop#tick(String)) so tests
  can drive the real send path with a fake AgentControl instead of only the
  pure decide() function.
- Added tests I-L covering: a failed send does not consume the notice and
  retries; a successful send does; the text is gated when the latch is
  already set; and the notice appears exactly once across three differently
  driven INJECTs.

Both required mutations verified red and reverted:
1. Setting the latch from the Decision regardless of send outcome -> test I
   (iAFailedSendDoesNotConsumeTheNotice) fails.
2. Dropping the latch argument at the contextNotice call site -> tests K
   (kAPendingDrivenInjectWithTheLatchAlreadySetSendsNoNotice) and L
   (lTheNoticeAppearsExactlyOnceAcrossThreeDifferentlyDrivenInjects) fail.
2026-09-20 17:34:05 +07:00
ltms fa62e9906d Merge #611: make the collected-ticket nudge test deterministic (fleetd #608)
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m48s
Replaces a wall-clock bet with a manually-driven scheduler, so the tick runs
only when the test runs it. Also closes a second, smaller race the brief did
not name: waiting on Phase.DONE is not enough, because complete() can publish
isDone() before every whenComplete dependent has run.

Verified by the lead: full suite 1841/1841 green in a clean worktree, and an
independent mutation (hasTicketWork forced true) diagnosed as an equivalent
mutant — ReplyPushLoop.injectNudge re-reads the pending collections and returns
early, so that line cannot reach agent.prompt. The worker's own mutation
(ticketCollected made a no-op) is the one that reaches the observable, and it
killed.
2026-09-20 12:26:16 +02:00
Dai Ha d7390ccd37 fleetd #609 review: repair a garbled comment carried over from the brief
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Failing after 1m41s
The brief's sentence about a null token count at HIGH was broken, and the
worker copied it into the source verbatim. The code was already right; only
the comment was unreadable.

Says what is actually true: a HIGH reading always carries a non-null token
count today, because LeadContextGauge only reaches HIGH by comparing a number
against HIGH_THRESHOLD_TOKENS. That invariant lives in another class and
nothing asserts it, so the branch stays.
2026-09-20 17:20:48 +07:00
Dai Ha aa517ae0ec fleetd #608: make anAlreadyCollectedTicketProducesNoNudge deterministic
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Failing after 1m34s
Replace the real ScheduledExecutorService backing ReplyPushLoop in this one
test with ManualScheduler, a fake that only runs a tick when the test calls
runDueTasks(). The old test bet a 300ms backoff was wide enough that
collecting the ticket always won the race against the scheduler's own timer
- true on an idle machine, false under a loaded full-suite run, which is
exactly the flake reported.

The rewritten test also waits on setAfterFinishAsyncTaskCompleteHookForTest
(already used elsewhere in this file) instead of polling Phase.DONE, so it
does not race CompletableFuture.complete()'s own publish-then-run-dependents
gap (fleetd #399) while proving ReplyPushLoop.onTicketTerminal really ran
before the ticket is collected.

Verified: backoff=1 (the most hostile value) still passes; mutating
ReplyPushLoop.ticketCollected to a no-op turns the test red with the same
assertion message the original flake reported; three consecutive full-suite
runs are green (1841/1841 each).
2026-09-20 17:18:45 +07:00
Dai Ha 60496831c2 fleetd #609: nudge an idle lead to hand over when its own context reads HIGH
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m5s
LeadHeartbeatLoop can now append a text-only notice to its nudge when the lead's
own LeadContextGauge reading is HIGH and leadHeartbeat.contextHighNudge is on.
Fires once per HIGH stretch (a latch, cleared only by a later OK reading; UNKNOWN
neither sets nor clears it), never spends the quietNudgeCap budget, and never
rolls a pane itself — only the operator can approve a handover.

- LeadContextGauge.Reading.unknown() widened to public for LeadContextSource.none()
- FleetConfig.LeadHeartbeat gains contextHighNudge (null/false = off, unchanged default)
- LeadHeartbeatLoop.decide gains context/contextNotified; Fleetd wires a new
  leadContextLookup/leadContextSource factory pair (LeadHeartbeatLoop.LeadContextSource)
- fleetd.example.yaml documents the new key
2026-09-20 17:11:54 +07:00
ltms 9a992d0f70 Merge #602: report a lead's live context usage in fleet_list
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m14s
CI / build (push) Failing after 1m33s
LeadContextGauge reports OK / HIGH / UNKNOWN for each lead in fleet_list, read
from the transcript Claude Code writes for itself — never from the lead's pane,
so it does not touch the control plane invariant 5 protects.

Why: a lead on this host auto-compacted 30 times in one session, discarding
roughly 250,000 tokens each time, and nothing could see it coming. The handover
feature already existed; what was missing was any way to know when to use it.

Three design points that earned their place:
  - finds the transcript by NAME under <configDir>/projects/, never by deriving
    the project slug, which is an undocumented Claude Code internal
  - three states, not two. Every path that cannot establish a token count says
    UNKNOWN with no number, so a lead is never told it is fine when the honest
    answer is "I could not look"
  - bounded twice: TAIL_BYTES caps bytes read, a 5s cache TTL caps how often

Includes #606, which fixed the two defects this PR's body recorded before it
was ever merged: the null configDir that made the gauge inert on this fleet,
and the torn final line that would have made it flap.

The authoring member died mid-task on a DNS error with the work uncommitted;
the lead recovered it, verified it, and committed it with that stated.

Verified by the lead on the combined state with current main merged in:
mvn -o clean install exit 0, 147 reports, 1841 tests, 0 failures, 0 errors,
0 skipped. LeadContextGaugeTest 9/9, FleetMcpLeadContextGaugeWiringTest 2/2,
FleetdLeadConfigDirLookupTest 6/6, FleetdLeadConfigDirSourceWiringTest 3/3.

Not deployed by this merge. A merge is not a deployment; the running daemon
still holds its old jar until redeploy.
2026-09-20 11:38:31 +02:00
ltms 3762aca307 Merge #606: wire the lead context gauge to the real configDir, and stop flapping on a torn line
CI / shell-tests (pull_request) Failing after 6s
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m50s
Fixes the two defects recorded in #602's own body.

1. FleetMcp.contextView passed configDir=null, so the gauge read <user.home>/
   .claude while this host's lead profile sets an override. Measured before the
   fix: 19 transcripts under the real directory, 0 under the fallback. The gauge
   would have deployed green and reported UNKNOWN forever, for every lead.
   Now threaded via FleetMcp.LeadConfigDirSource, built by Fleetd
   .leadConfigDirSource, following fleet.leaders.<name>.profile to that
   profile's configDir and reading config.get() live inside the lambda.

2. A torn final line no longer means UNKNOWN. fleet_list reads a transcript
   Claude Code may be mid-write on, so the last line can be cut. The old code
   treated that as fatal, which would make the gauge flap at random. The stated
   reason ("a format change should show as UNKNOWN") does not hold: a real
   format change makes EVERY line unparseable, and that case is still caught.

Verified by the lead on the combined state with current main merged in, not on
the branch alone: mvn -o clean install exit 0, 147 reports, 1841 tests, 0
failures, 0 errors, 0 skipped.

Mutation-checked independently by the lead:
  Fleetd.leadConfigDirSource body -> none()   -> KILLED (1 failure)
  FleetMcp.contextView configDir -> null      -> KILLED (worker-measured)

KNOWN RESIDUAL, documented rather than overclaimed. main's own one-line call to
leadConfigDirSource could be swapped for none() and the suite stays green. Every
member of this wiring-test family (loopHealthSource, capacitySource,
healthCoverageSource) has the identical gap — no test runs Fleetd.main far enough
to observe which factory it called. Filed separately as a class-wide problem
rather than patched here. The empirical close is the dogfood check after redeploy.
2026-09-20 11:38:22 +02:00
Dai Ha 4e27bde2d7 fleetd #602 gauge-wiring follow-up: pin Fleetd.main's LeadConfigDirSource wiring
Extract the inline new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(...))
construction in Fleetd.main into a package-private factory,
Fleetd.leadConfigDirSource, mirroring loopHealthSource/capacitySource/
healthCoverageSource. Add FleetdLeadConfigDirSourceWiringTest, which calls the
factory directly with real Profile/Leader fixtures and asserts the returned
source resolves a real configDir -- a property that is false if the factory's
body is mutated to return LeadConfigDirSource.none().

Neither FleetMcpLeadContextGaugeWiringTest nor FleetdLeadConfigDirLookupTest
could catch main losing this wiring: each builds its own instance instead of
calling what main calls. This closes that gap at the factory level, matching
the standard already accepted for loopHealthSource's own wiring test.
2026-09-20 16:34:49 +07:00
ltms b85d9b0e46 Merge #607: redeploy-fleetd.sh no longer fails a deploy that worked
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m49s
fleetd #603. The pid poll had its own fixed 10s budget and then hard-died,
while the health check right after it was allowed 60s for the same daemon.
Under launchd the java process does not exist yet when launchctl load returns,
so the script reported FAIL on a fully successful deploy — and a false FAIL in
that direction invites the hand-rolled stop/start this script exists to replace.

The pid poll now shares HEALTH_WAIT, and a miss falls through to the health
check rather than killing the run. A genuine failure still dies and still
prints the log tail. Worst-case time-to-fail roughly doubles; that cost lands
only on real failures and is the right trade.

Verified by the lead, not taken from the report. Suite exit 0, unpiped.
Two mutations run independently:
  budget back to a hardcoded 10          -> KILLED (suite exit 1)
  warn back to die on a pid miss         -> KILLED (suite exit 1) after 856dfc6
The second survived on the first submission, which is why that commit exists:
nothing covered "pid never appears but healthz answers, so do not die" — the
one case the fall-through is for, and reachable because running_pid() only
recognises a plain java -jar.
2026-09-20 11:30:10 +02:00
Dai Ha 856dfc6318 fleetd #603 review: close the untested fall-through path
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m15s
PR review (comment 17358) found a real gap via mutation testing: replacing
the "pid never found" warn with die "no process appeared" left the whole
suite green, because neither existing test drove the case the fall-through
exists for -- running_pid() never finds anything (as its own doc comment
says it eventually will) while /healthz answers anyway.

Adds test_await_daemon_started_pid_never_found_but_healthy_warns_and_survives:
running_pid always empty, poll_health_body succeeds. Asserts DIED_CALLED=0,
the warn line is emitted, and NEW_PID stays empty (the honest "could not
establish this" answer, never a guessed pid).

Verified both halves myself: reverting the warn to die "no process appeared"
turns this one test red (FAIL: await_daemon_started must not die...);
restoring it returns the suite to green.
2026-09-20 16:28:53 +07:00
Dai Ha 81c1d8e91c fleetd #603: share HEALTH_WAIT between the pid poll and the health check
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m47s
The start step gave the new process its own short, fixed 10s budget before a
hard die, while the health check right after it waits a full HEALTH_WAIT
(60s) for the same daemon. Under launchd, launchctl load returns before the
java process exists, and on a slow host that took longer than 10s -- so the
script reported "no process appeared" on a deploy that had fully succeeded.

wait_for_new_pid/await_daemon_started fold the pid poll and the health check
into one decision: the pid poll now shares HEALTH_WAIT instead of its own
shorter budget, and a miss there falls through to the health check (direct
proof the daemon is up) instead of killing the run. A genuine failure still
dies, and still prints the log tail.

Adds behavioural tests for both acceptance criteria (slow start succeeds,
genuine failure still fails and prints the log tail) in one suite run, plus
unit tests for wait_for_new_pid and a call-site test for the new function.
2026-09-20 16:21:52 +07:00
Dai Ha d345b14e43 fleetd #602 gauge-wiring: thread a lead's configured configDir into the context gauge
FleetMcp.contextView hardcoded LeadContextGauge.read(null, ...), so a lead whose
profile sets its own CLAUDE_CONFIG_DIR always read the wrong transcript directory
and reported UNKNOWN forever, with no error anywhere.

- Add FleetMcp.LeadConfigDirSource (same idiom as LeadSeatSource) and thread it
  through the constructor / listFleet overload chain / leadView / contextView.
- Add Fleetd.leadConfigDirLookup, wired at construction, following the same
  fleet.leaders.<name>.profile link leadSeatLookup already uses, one step
  further to that profile's own configDir.
- LeadContextGauge.parse: a single unparseable line (typically the final one,
  torn by a write this read raced) is now skipped rather than forcing UNKNOWN;
  only when every line in the read window fails to parse does it report
  UNKNOWN, which is the real format-change signal.
- Tests: FleetdLeadConfigDirLookupTest (lookup logic), FleetMcpLeadContextGaugeWiringTest
  (end-to-end: config naming directory A vs B decides which is read; a lead with
  no configured dir degrades without throwing), and two replacement properties in
  LeadContextGaugeTest for the torn-line fix plus its all-unparseable control.
2026-09-20 16:19:37 +07:00
ltms 9640deeffc Merge #605: fleet_list reports charterBytes alongside charterSha256
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m34s
#604 item 1. A digest answers "same or different" and cannot say how much. The
byte count is the second signal, needed exactly when two hosts find they differ.

Verified by the lead before merging, from fleetd/: mvn -o clean install exit 0,
143 reports, 1821 tests, 0 failures, 0 errors, 0 skipped; SessionManagerTest 74/74.

Includes ad3d819, correcting an invariant the added comment claimed but that
CharterReceipt.compose() does not hold: digestOf() returns null for blank text
while getBytes().length does not, so a whitespace-only charter pairs a null
digest with a non-zero size. No behaviour change — the digest gate already
omits both on that path, which is the right answer.
2026-09-20 11:18:17 +02:00
Dai Ha ad3d81941f Correct the invariant claimed in the charterBytes comment
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Successful in 1m36s
CI / contract (pull_request) Successful in 1m42s
The comment said CharterReceipt never pairs a null digest with a non-zero
byte count. It can. compose() derives the digest with digestOf(), which
returns null for blank text, while the byte count is getBytes().length,
which does not. A whitespace-only role charter on a profile with no MCP
produces exactly that pair.

No behaviour change. The gate already omits both fields on that path, which
is the right answer — a size with no digest would describe an artifact we
cannot fingerprint. Only the stated reason was wrong, and a false invariant
in a comment is worse than no comment, because the next reader will widen
the gate on the strength of it.
2026-09-20 16:16:30 +07:00
Dai Ha 8b986a52e0 #604 item 1: fleet_list reports charterBytes alongside charterSha256
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m26s
CharterReceipt carries a byte count next to its digest, but the roster
projection in SessionManager.rosterView only ever copied the digest
across. A digest tells a lead whether two members' charters match; it
cannot say how far apart they are when they don't. Report charterBytes
too, nested in the same conditional as charterSha256 so the two travel
together: the receipt's own contract only ever pairs a non-null digest
with a real byte count, and a member with no composed charter reports
charterSource alone, unchanged from before.

Tests: the existing charter-receipt roster test now asserts charterBytes
against the receipt's own value (not a literal), plus two new cases —
no charter composed (source "none", no digest, no size) and the receipt
itself absent (no charter keys at all).
2026-09-20 16:14:45 +07:00
Dai Ha f336bcef39 CLAUDE.md: rewrap the long line my last edit left in the architects paragraph
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m48s
The fleet01 lead found this. My edit in f5c6a0e left one prose line at 128
columns inside a paragraph wrapped at about 96. It is invisible to any check
that normalises by paragraph, and visible in any raw digest of the block.

No wording changed. Block stays 20937 bytes. The wiki template gets the same
rewrap in its own commit, so the two stay byte-identical.
2026-09-20 16:13:23 +07:00
Dai Ha 3ed7bfca67 fleetd: report a lead's live context usage in fleet_list
CI / shell-tests (pull_request) Failing after 7s
CI / build (pull_request) Failing after 1m37s
CI / contract (pull_request) Successful in 2m7s
fleetd had no way to see how full a lead's context window is. On this host a
lead auto-compacted 30 times in one session, discarding roughly 250,000 tokens
and costing 46s to 3m16s each time, and nothing could see it coming.

LeadContextGauge reads the transcript Claude Code itself writes, never the
lead's pane. It finds <sessionId>.jsonl by NAME under <configDir>/projects/
rather than deriving the project slug, which is an undocumented internal.

Three states, not two: OK, HIGH, UNKNOWN. Every path that cannot positively
establish a token count reports UNKNOWN with no number, so a lead is never
told it is fine when the honest answer is "I could not look".

Bounded two ways: TAIL_BYTES caps bytes read off disk, and a 5s cache TTL caps
how often that read happens, because fleet_list is polled constantly.

Recovered by the lead: the authoring member ended on a backend error (DNS
ENOTFOUND) with this work uncommitted and unpushed in its worktree. Verified
before committing: mvn -o clean install exit 0, 1813 tests, 0 failures,
0 errors, 0 skipped, 144 reports; LeadContextGaugeTest 8/8.

KNOWN INCOMPLETE - see the PR. The fleet_list call site passes configDir=null,
which falls back to ~/.claude, but this host's lead profile sets configDir to
an override. Measured: 19 transcripts under the real configDir, 0 under the
fallback. The gauge is therefore INERT on this fleet until that is wired.
2026-09-20 16:05:15 +07:00
Dai Ha f5c6a0e4fc CLAUDE.md: architects settled the two invented specifics at line 144
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 47s
CI / build (push) Successful in 3m0s
Both specifics in the "consult architects" paragraph were mine, not the
operator's. The operator declined twice to rule on them and directed the lead
to consult architects instead, so two architects on different models settled
them over two rounds.

"after two rounds" is gone. It was a ceiling nobody had evidence for, and it
implied a counter fleetd does not have - nothing in the daemon counts rounds.
The bound is now expressed as a shape: form independent positions, then
compare. That is a floor of two without naming a number.

The three-item operator list read as complete, so a lead hitting anything not
on it would conclude it must not ask. It is now explicitly examples, and
"granting access" replaces "credentials" - the case that motivated this was a
forge merge refusal on a protected branch, which "credentials" covers only
awkwardly.

Canonical block and the wiki template updated together; sync check passes.
2026-09-19 23:32:41 +07:00
Dai Ha a7aee5b982 Merge #600: fleetd lead-rollover outcomes readable after confirm()
Adds LeadRollover.status() and a 'status' action on fleet_handover, so a lead
can find out what happened to its own roll. Every failure past confirm() was a
log.warn the lead cannot read.

Five states. IN_PROGRESS is written at the confirm hand-off, BEFORE the token
leaves 'pending', and status() reads 'outcomes' first — so there is no window
in which an in-flight roll reports UNKNOWN.

Gated by the lead: mvn -o clean install exit 0, 1819 tests, 0 failures,
0 errors, 0 skipped, 143 reports; LeadRolloverTest 43, FleetMcpHandoverTest 12.
Diff read in full. 0 source-text assertions in the new tests.
2026-09-19 23:32:32 +07:00
Dai Ha bf895616a5 fleetd: distinguish an in-flight roll from an unknown token
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m32s
CI / build (pull_request) Successful in 2m8s
confirm() removed the token from pending before handing the roll to the
continuation, and outcomes was only written when runRollover reached an exit.
For the whole duration of the roll the token was in neither map, so status()
answered UNKNOWN - documented as "never issued, cancelled, or aged out". A lead
polling right after its own confirm was told the roll had never been requested.

Adds RollState.IN_PROGRESS, written at the confirm hand-off rather than at the
roll's end, so there is no gap. status() now reads outcomes before pending, so
the hand-off write cannot race the removal.

Also corrects the class javadoc, which still claimed nothing calls this class.

Recovered by the lead: the authoring member ended on a backend error (the host
slept mid-response) with this work uncommitted in its worktree. Verified before
committing: mvn -o clean install exit 0, 1819 tests, 0 failures, 0 errors,
0 skipped, 143 reports; LeadRolloverTest 43 (was 39).
2026-09-19 22:31:47 +07:00
Dai Ha b874afb0af fleetd: make lead-rollover outcomes readable after confirm()
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Successful in 2m14s
LeadRollover previously logged every post-confirm() failure only — a lead
has no way to read the daemon log, so a roll that timed out because its
own turn never settled (or /clear never re-settled) was invisible; the
lead would carry on believing a fresh session was coming.

Add a bounded (cap=200) token -> outcome record, written at each of the
three exits in runRollover (ROLLED, TURN_NEVER_SETTLED, CLEAR_NEVER_SETTLED),
and a read-only LeadRollover#status(token) accessor. The TURN_NEVER_SETTLED
detail names turnSettleSeconds explicitly so a reader knows what to raise.

Wire a "status" action onto the fleet_handover MCP tool (handler + schema);
it never schedules, cancels, or retries anything — confirm() remains the
only path that can ever cause a /clear.

Extends LeadRolloverTest (33 -> 39 tests) covering the six acceptance
properties, and FleetMcpHandoverTest (8 -> 12) for the new tool action.
2026-09-19 16:17:39 +07:00
ltms 6eb34a654f Merge #599: fleetd #589 groups 1+2 — wiring-test 6 sites in Fleetd.main()
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m43s
Extracts 6 inline constructions in Fleetd.main() (:214-:468) to
package-private factories, each pinned by a behavioural wiring test.
Production behaviour unchanged.

Gate: test-merged onto current main (already carrying #598, which rewrote
117 lines of the same file). No conflict — the two units insert at different
anchors, #599 after capacitySource and #598 after loopHealthSource, as their
briefs specified. mvn exit=0, 1805 tests / 0 failures / 0 errors from 143
surefire reports; all 6 new classes confirmed to have run with their
expected counts. Diff confirmed extraction-only.

No source-text assertions in any of the 6 new files; that zero carries a
positive control (the same pattern finds 18 such files elsewhere in the
repo). Worker self-reported fixing a mutation that threw NullPointerException
rather than failing an assertion — the assertNotNull guard is present ahead
of the matcher call, confirmed.
2026-09-19 10:39:47 +02:00
ltms 7084d99b89 Merge #597: fleetd #593 — running_pid() counts only the daemon
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m39s
CI / build (push) Successful in 1m50s
running_pid() filters pgrep -f hits by `comm = java` (allowlist) instead of
denying a fixed list of shell names. Closes two false-positive holes: an
exited pid (empty comm matched no denied name) and any non-shell wrapper
(ssh, perl, python3, ruby) carrying the pattern in its own argv.

Gate: merged onto current main, bash suite exit 0 and structurally identical
to the baseline run on main (the one "Unattributable mutation" line is
pre-existing, confirmed by running the suite on origin/main). All four
running_pid tests confirmed defined AND invoked. Mutation check run by me:
neutering the allowlist makes the suite exit 1 with a named failure; restore
is byte-identical to baseline by git hash-object and green again.

Also checked and cleared: the new die-message advice `ps -eo pid,comm,args`
does NOT expose process environments on macOS — `-e` with `-o` selects all
processes, it does not imply `-E`. Verified with an isolated two-phase probe
and a positive control, after three earlier probes gave false positives by
self-matching (the grep's own argv, and the probe script's own text).
2026-09-19 10:37:50 +02:00
Dai Ha ae7845c375 fleetd #589 (Groups 1 & 2): pin 6 main() wiring sites with named factories
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m52s
Extracts 6 inline wiring expressions from Fleetd.main() into named,
directly-testable package-private static factories, following the
FleetdLoopHealthSourceWiringTest (#584) shape, and adds one wiring test
per factory:

Group 1 (exhaustion/quarantine):
- forwardingExhaustionSink(exhaustionSinkRef) — was inline
  ExhaustionSink.forwardingTo(exhaustionSinkRef::get)
- publishExhaustionSink(...) — was two untested statements building the
  real sink and .set()-ing it into exhaustionSinkRef
- liveExhaustedPatterns(config) — was inline
  new LiveExhaustedPatterns(() -> config.get().profiles())
- exhaustedPatternLookup(roster, liveExhaustedPatterns) — was an inline
  lambda resolving a herdr target to its profile's live pattern; silently
  losing this is the worst regression in the sweep, since a real
  usage-limit refusal would stop being classified as BACKEND_EXHAUSTED

Group 2 (CB-596 credential policy):
- claudeCodeLauncher(...) — was an inline `new ClaudeCodeLauncher(...)`
  whose memberCredentials supplier argument was untestable wiring
- openCodeLauncher(...) — same, for OpenCodeLauncher

Each new test pins its factory behaviorally (never via source-text
assertions): built and confirmed RED by name against the named inert
mutation, then confirmed GREEN again after restoring, and separately
confirmed GREEN after a behavior-preserving reformat/local-variable
extraction of the same call, to rule out a disguised source-text test.

Suite: 1789 -> 1799 tests (+10, matching the 10 tests added), 0
failures, mvn -o clean install BUILD SUCCESS.

Scope strictly limited to main()'s :214-:468 range per the ticket split
with the concurrent worker handling Group 3 at line 500+.
2026-09-19 15:35:32 +07:00
ltms 61115f6f61 Merge #598: fleetd #589 group 3 — wiring-test 5 sites in Fleetd.main()
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 50s
CI / build (push) Successful in 2m9s
Extracts 5 inline lambdas/method-refs in Fleetd.main() to package-private
factories and pins each with a wiring test. Production behaviour unchanged;
the releaseCleanup body moved verbatim.

Gate: test-merged onto current main in a scratch worktree, mvn exit=0,
1795 tests / 0 failures / 0 errors from 137 surefire reports. Diff read in
full. Worker's self-disclosed bare `git stash push` verified as recovered —
all 3 surviving stash entries predate today, so no other worktree lost work.
2026-09-19 10:33:07 +02:00
Dai Ha 4b9ebda1b3 fleetd #593 CORRECTION 1: allowlist comm=java, not a denylist of shells
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m48s
The round-1 fix excluded known shell names (sh/bash/zsh/dash/ksh) from
running_pid()'s pgrep candidates. Two holes remained, both the same
false-positive shape the ticket exists to remove:

1. A pid pgrep lists can exit before the following `ps -o comm=` lookup
   runs. On a gone pid, ps prints nothing, comm is empty, and an empty
   string matches no denied shell name -- so a dead pid was still counted.
2. The denylist only knows the shells someone thought to name. ssh, perl,
   python3, ruby, tail -- anything else carrying the pattern in its own
   argv -- was still counted alongside the real daemon. The ticket names
   ssh as a live route.

Both close with one change: allowlist comm=java instead of denying shells.
The daemon is always `java -jar target/fleetd.jar`, so its comm is always
`java`; an empty comm (hole 1) is not `java` either, closing that hole for
free.

Answers the objection in the code comment: an allowlist can under-count if
fleetd ever stops being launched by `java` (a native image, a renamed
launcher). That's a false negative, the worse direction for a guard -- but
it is not a new assumption: PATTERN='target/fleetd.jar' already assumes a
jar run by java, and that pattern breaks before this allowlist would.

Replaces the round-1 "real second process" test (which gave its exec -a
standin an argv[0] holding the pattern, but not comm=java) with one that
forces comm=java via `exec -a java sh -c '...'`. Adds two stubbed
pgrep/ps tests pinning the two holes directly (a non-java, non-shell comm
such as perl; an empty comm from an already-exited pid) -- deterministic
on every platform, unlike a live-process fixture, and immune to the BSD
vs Linux difference in how `comm` is derived from a fabricated process.
Adds a stubbed positive backstop (comm=java is counted).

Confirmed the regression is caught: reverted to the round-1 denylist,
reran the suite, watched the new non-shell-comm test fail at
`set -e`'s first failure, then isolated the exited-pid test separately
and confirmed it also fails against the same broken code. Restored the
fix and reran green.

Branch merged with origin/main (3 commits: hunter role + CLAUDE.md
addendum) before this commit; unrelated, no conflicts.
2026-09-19 15:31:15 +07:00
Dai Ha 1e68d7ee39 Merge origin/main into worker/593-1a8025-5 2026-09-19 15:27:46 +07:00
Dai Ha 6cccd458d4 #589 Group 3: wiring-test the 5 sites below line 500 in Fleetd.main()
CI / shell-tests (pull_request) Successful in 11s
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 2m11s
Extracts the inline lambdas/method references at the 5 assigned wiring
sites into named package-private factories on Fleetd, following the
FleetdLoopHealthSourceWiringTest pattern from #584:

- turnRegistrar(CompletionResolver) — was completion::register (Injector)
- healthFailTarget(MessageService) — was messages::abandon (FleetHealthMonitor)
- releaseCleanup(MessageService, ReplyInbox, PrimaryRegistry) — was the
  inline sessions.onRelease(detail -> {...}) cleanup lambda
- replyInboxOpener() — was AmqpReplyInbox::open passed to selectReplyInbox
- leadMailboxOpener() — was LeadMailbox::open passed to openLeadMailbox

Each factory has a new runtime test (not source-text) that drives real
collaborators through public APIs: MessageService.poll(ticket).phase(),
InMemoryReplyInbox.peek(), PrimaryRegistry.nudgeTargetFor(), and the
opener tests connect to a guaranteed-closed local port to prove a real
network attempt vs. an inert stub.

releaseCleanup was done first per the brief: MessageService.abandon's
javadoc documents that losing this cleanup leaves a torn-down worker's
rendezvous waiter open forever.

Tests: 1789 -> 1794 (+5), 0 failures, 0 errors. mvn -q -o test exit 0,
no BUILD FAILURE, no piped exit status. Each new test verified RED on
the inert form named in the ticket, and GREEN after reformatting the
call across lines and extracting the argument into a local/factory.
2026-09-19 15:24:11 +07:00
Dai Ha 42820fbe75 fleetd #593 (pid-count half): running_pid() no longer matches the caller
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m45s
running_pid() was a bare `pgrep -f "$PATTERN"`, which matches ANY process whose
full command line contains the pattern text -- including a shell that merely
embeds it as literal text (a hand-typed investigation, an ssh-shaped
`sh -c '...; ...'`, or a pipeline) rather than being the daemon. That self-match
turns a working redeploy into a reported "racing supervisor" failure via
assert_single_daemon.

pgrep -c does not exist on BSD/macOS, so this can't be fixed by switching flags.
running_pid() now keeps pgrep to find candidates (portable), then drops any
candidate whose process name (comm) names a shell -- the daemon is always
`java`, so a self-matching wrapper of this shape is always excluded while a
genuine second daemon-shaped process still counts.

assert_single_daemon's die message no longer hands the operator a bare
`pgrep -f "$PATTERN"` as remediation -- that was exactly the self-matching
invocation -- and now says in words that a pattern can match the caller.

Adds three tests: a self-matching wrapper shell must be excluded, a real
second daemon-shaped process must still be found, and the die message must
not recommend the self-matching command. Verified the first test fails
against the pre-fix implementation (confirmed the regression is caught).

Leaves instance 1 (the fleetd.out log source, systemd-only) for a Linux host,
per the ticket's scope split.
2026-09-19 15:18:32 +07:00
Dai Ha d91ff886da #568 follow-up: fix the text defects the hunter-role merge introduced
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 1m48s
Found by reading the diff at the merge gate, not reported by the worker.

1. FleetConfig.java: the operator-facing "unknown key" hint read
   "'fleet.architects', 'fleet.developers' or 'fleet.hunters' or
   'fleet.reviewers'" — a double "or". This is text an operator reads at the
   moment their config is already wrong, so it should not itself be wrong.
2. MemberLifecycle.java: javadoc continuation asterisk indented 6 spaces, not 5.
3. MemberRegistry.java: javadoc asterisks moved from column 2 to column 4.
4. CallerResolver.java: a // comment indented one space past its block.

2-4 are the worker mangling alignment while widening enum lists to include
HUNTER. No behaviour changes.

CORRECTION to the #596 merge commit message. It claimed a fifth defect, "two
javadoc lines pushed past the 100-column convention". There is no such
convention in this repo: no checkstyle, no spotless, no .editorconfig, and 2975
of 32728 lines under fleetd/src/main/java already exceed 100 characters. I
asserted the rule before measuring it. Those two lines are untouched.

Verified: built in a scratch worktree, 1790 tests, 0 failures, 0 errors,
0 skipped, counted from the surefire XML.
2026-09-19 15:15:54 +07:00
ltms 386e760a5c Merge #596: fleetd #568 — add the hunter member role
CI / shell-tests (push) Successful in 5s
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m37s
Verified by the lead before merge, not taken on the worker's report:

- branch contains a639969; 1 ahead, 0 behind — clean fast-forward
- CLAUDE.md change is +2 lines in the Project addendum, NOT the canonical block
- canonical block sync check prints True on main and on this branch
- built in a scratch worktree (never mvn clean in the main clone): 1790 tests,
  0 failures, 0 errors, 0 skipped, counted from the surefire XML. Baseline 1789.
- live fleetd.yaml still loads: the hunters pool is optional

Five text defects found by reading the diff, not reported by the worker. They are
fixed in a follow-up commit on main rather than a round trip:
- FleetConfig.java operator-facing message reads "... 'fleet.developers' or
  'fleet.hunters' or 'fleet.reviewers'" — a double "or"
- misaligned javadoc continuation asterisks in MemberLifecycle and MemberRegistry
- a misaligned // comment in CallerResolver
- two javadoc lines pushed past the 100-column convention

Known gap, tracked separately: fleet.hunters is absent from the live config, so a
hunter cannot spawn on this host until the pool is added after the redeploy. The
role ships correct and inert.
2026-09-19 10:13:26 +02:00
Dai Ha 2e349139e9 #568: add hunter member role
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m44s
2026-09-19 15:09:07 +07:00
Dai Ha a639969a9a CLAUDE.md: a blocked lead consults architects, not the operator
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 1m32s
CI / build (push) Successful in 1m42s
The operator set this rule on 2026-09-19: when a decision blocks a lead, it
consults one or more architect members, who are authorized to agree on one
decision and unblock. The operator is not asked. Escalation stays open only
for things outside the fleet's authority -- money, credentials, or a promise
made to someone else.

The paragraph also carries the reason the ticket record is mandatory rather
than optional. The operator's old notification channel was the block itself:
work stopped, so they found out. Taking the operator out of the loop removes
that signal with it, so the decision goes on the ticket, which reaches them
whether or not they are at a terminal when it is made.

The rule has a second half aimed at architects, which lives in
fleet.charters.architect and is applied per daemon -- filed as #591, because
a charter can never reach a lead and charters do not travel between hosts.
The missing notification event is #592.

Block verified byte-identical with wiki 7-Use-Cases.md at 1d9bd1b.
2026-09-19 14:57:04 +07:00
Dai Ha 17c3a69c57 docs: CB-591 gateway page had two claims that went stale
CI / shell-tests (push) Successful in 5s
CI / contract (push) Successful in 1m23s
CI / build (push) Failing after 1m37s
The gateway's chat model is served under the stable alias `acoder`, and the
model behind that alias changed on 2026-08-28 — it is Qwen3.8-27B now, not
DeepSeek-V4-Flash. The old name is still served, so nothing broke, but it
names a model this is not.

Two claims on the page were wrong as a result, and both were written as
current facts rather than dated measurements:

- `/v1/models` returns exactly `["deepseek-v4-flash"]` — it returns 6 ids
  now. This sat under a heading saying it needs no re-testing.
- the status banner said `local` and `gx` are both at `weight: 100` —
  `local` is at 0.

Measured today against the live gateway: /v1/models returns acoder,
qwen3.8-27b-nvfp4, deepseek-v4-flash and three embedding names; a completion
sent as `deepseek-v4-flash` comes back reporting `"model": "acoder"`, which
is the alias in plain sight. /v1/deployment reports generation
2026-08-28-qwen3.8-27b-nvfp4.

§2 and §3 are left alone. They are the August plan, and rewriting them would
destroy the record of the migration.

fleetd.yaml moved to `acoder` in the same change. It is not tracked here.
2026-09-13 07:01:15 +07:00
ltms 49a5875586 Merge #583: fleetd #582 — assert pending message-id cleanup at every publish cleanup site
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 2m1s
All ten assertions proven live: six by the implementer, the last four by the lead.
One contract build with four deleted removal lines produced exactly four named failures,
one per site, with the total unchanged at 1825.
2026-09-12 15:52:25 +02:00
Dai Ha 634d33b50b Merge worker/562-loop-health-wiring-test-99611c-5
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 1m45s
2026-09-12 20:28:14 +07:00
Dai Ha 1db79bcaa9 Merge worker/581-completionresolver-cas-sites-0542b7-6 2026-09-12 20:28:14 +07:00
Dai Ha 4507bc5a70 Merge worker/571-attempted-outcome-5739f7-2 2026-09-12 20:28:14 +07:00
Dai Ha d7239ed23b fleetd #571: pin FleetMcp.formatReply's TIMED_OUT_UNCONFIRMED wording
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m42s
CORRECTION 5 on the ticket: mutating the new arm's message text to the
queued/working arm's text survived every existing test, because nothing
asserted the specific wording. This adds one test that asserts the
unconfirmed-delivery message and asserts it does NOT carry the
queued/working arm's retry invitation — the distinction #571 exists for.

No production code changes; formatReply's TIMED_OUT_UNCONFIRMED arm was
already correct.
2026-09-12 20:19:37 +07:00
Dai Ha dfeb9340b4 fleetd #581: cover completion CAS removals
CI / shell-tests (pull_request) Successful in 5s
CI / build (pull_request) Failing after 1m31s
CI / contract (pull_request) Successful in 1m35s
2026-09-12 20:19:20 +07:00
Dai Ha 1513d4f260 fleetd #562 follow-up: extract loopHealthSource factory, pin its wiring
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 1m43s
PR #579's inline `new FleetMcp.LoopHealthSource(poller::health, ...)` in
Fleetd.main had nothing a test could call directly. Measured: replacing
poller::health with a constant () -> RUNNING compiled clean and left all
1771 tests green (see issue #562 comment "HOLD on PR #579").

Extracts the inline construction to a package-private Fleetd.loopHealthSource
factory, the same style as the sibling capacitySource/healthCoverageSource
factories, and adds FleetdLoopHealthSourceWiringTest with three separate
assertions: the statusPoller half, the sessionReaper half, and the
reaper == null branch (still STOPPED).
2026-09-12 20:16:23 +07:00
Dai Ha 4ca7d72303 #582: assert pending message-id cleanup
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m32s
2026-09-12 20:15:48 +07:00
Dai Ha c0545d003d fleetd #571: make FleetApp.writeReply's inner Outcome switch exhaustive, no default
CI / shell-tests (pull_request) Successful in 8s
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m48s
Ticket comments (17126, 17127) corrected the original acceptance criterion after this
unit was already in flight: a hand-listed grep for the enum's constant names goes stale
silently the moment a new constant lands, so the compiler must be the enumeration instead.
sendOutcomeLabel (MessageService.java) and formatReply (FleetMcp.java) were already
default-free switch expressions. The one gap was writeReply's inner "status" switch, which
had `default -> "done"` — the exact value that would have lied about TIMED_OUT_UNCONFIRMED.
Remove the default and list every Outcome constant explicitly; REPLIED, COMPLETED_UNREPLIED,
QUESTION and STALE_TURN get an arm too even though the outer switch always dispatches them
first, so the inner switch stays exhaustive on its own. The outer switch (a statement, not
an expression) keeps its own default — Java does not require exhaustiveness there regardless,
and "everything not terminal is a 202" is an intentional catch-all.

Verified with the proof the ticket asked for: added a scratch 11th Outcome constant after
deleting all default arms and confirmed all three switch-expression sites (and no test file)
fail to compile without an arm for it, one at a time, then removed the scratch constant.
2026-09-12 19:28:59 +07:00
Dai Ha 275ac0d251 fleetd #562: surface loop health
CI / shell-tests (pull_request) Successful in 12s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Successful in 2m11s
2026-09-12 19:26:37 +07:00
Dai Ha c1e06c9e12 fleetd #571: add TIMED_OUT_UNCONFIRMED so an ATTEMPTED delivery is not reported as never-arriving
MessageService.send's TimeoutException branch collapsed Injector.Cancellation.ATTEMPTED
(fleetd #551 — the send call was made but its outcome is unknown) into
Outcome.TIMED_OUT_QUEUED, which promises the caller the message will never arrive. On this
route agent.prompt may already have pasted and submitted the text, so a caller's natural
recovery (resend) risks a double delivery.

Add Outcome.TIMED_OUT_UNCONFIRMED and route ATTEMPTED to it. Update the three readers found
by searching for the enum's constant names (not `Outcome.`, which misses FleetMcp's
unqualified `case REPLIED ->` switches and would false-positive on ConfigRef's unrelated
Outcome record):
 - MessageService.sendOutcomeLabel: add it to the "timeout" metric label group.
 - FleetMcp.formatReply: its own case, warning against a blind retry (distinct from the
   generic "retry or poll status" message the other timeouts get).
 - FleetApp.writeReply: its own "unconfirmed" status and detail text, so it no longer falls
   through the switch's default -> "done" arm, which would have reported "the delegation
   completed" for the one case where delivery is unconfirmed.
2026-09-12 19:21:10 +07:00
Dai Ha 204da67d66 Merge #576: fleetd #575 — one finally covers answer()'s STALE_TURN exit
CI / shell-tests (push) Successful in 16s
CI / build (push) Successful in 1m42s
CI / contract (push) Successful in 1m57s
Third instance of the #572 shape on this file: one invariant kept at N sites,
asserted at fewer than N. Here answer()'s inner try opened AFTER the Task
lookup/registration and the STALE_TURN early return, so that return was covered
only by a hand-rolled copy of the finally's cleanup pair. The fix widens the try
upward and deletes the copy, so every exit runs the one finally exactly once.
rendezvous.open(workerSession) correctly stays outside it — nothing to clean if
it never opened.

This fixed no live leak, and the code comment says so: neither
rendezvous.answerAsk nor clearAsyncQuestion(turnId, false) can throw, so nothing
ever left through the old gap uncovered. It is a structure fix, the ticket's own
fallback case.

Verified here, not taken from the worker's report:
- baseline on the branch: 1766 tests, 0 failures, from Maven and from an
  independent sum over target/surefire-reports/*.txt.
- my own mutation, located fresh: move the inner `try {` back down below the
  STALE_TURN return — the exact pre-fix structure, minus the hand-rolled pair.
  Result: 1766 run, exactly 1 failure, and it is the new test —
  MessageServiceTest.answerLosingTheRaceToAnAlreadyAnsweredAskStillReturnsStaleTurnAndCleansUpOnce:545
  "the forward waiter this answer() call opened must be closed after a STALE_TURN
  return". One failure, not a crowd: the new test is the only thing holding this
  path.
- the worker's own mutation removed the whole finally and took 8 other tests with
  it. That proves the finally runs; it does not prove the STALE_TURN path reaches
  it. Mine does.

The new test hook answerAskLapseRaceHookForTest follows the file's existing
askTimeoutRaceHookForTest convention.

Closes #575.
2026-09-12 19:05:00 +07:00
Dai Ha b091c51eee fleetd #575: widen answer()'s try so one finally covers its STALE_TURN exit
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 2m30s
The waiter cleanup pair (asyncTasksByWaiter.remove + rendezvous.close) was
duplicated: two sites sit in a finally, the third was hand-rolled inline
before answer()'s early STALE_TURN return, structurally outside any finally.

Both rendezvous.answerAsk and clearAsyncQuestion(turnId, false) are total
(cannot throw), so the gap never leaked in practice. But the duplicate was
untested: mutating it away left all 1765 tests green, while the two
finally-protected sites are each killed by 8-36 tests. Same shape as #572.

Fix: widen the try to wrap the Task registration and the STALE_TURN check,
so the single finally covers every exit and the hand-rolled copy is gone.
Added a race hook + regression test that deterministically reproduces the
'ask lapsed between the lookup and the unblock' case and proves the fix
still returns STALE_TURN and cleans up exactly once.
2026-09-12 18:56:02 +07:00
ltms 84d631b030 Merge #572: answer()'s session-lock release is pinned on all four exits
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m44s
fleetd #572. MessageService releases its per-session lock in a finally at two sites. Removing the
send() one failed 22 tests and errored 1. Removing the answer() one left the whole suite green.
Skip that unlock and the thread holds the session lock forever, so every later send or answer to
that session blocks permanently — no exception, no log line.

Tests only. Against current main the diff is one file, MessageServiceTest.java, +178 lines.

Four new tests, one per exit of answer(): normal REPLIED, TIMED_OUT_WORKING, the ExecutionException
rethrow, and the InterruptedException rethrow. Each proves REACQUISITION rather than the return
value — a bounded follow-up send on the SAME session must not come back BUSY, and send reports BUSY
only when tryLock itself timed out, so a non-BUSY probe is specifically evidence the lock was free.

Lead verification, re-running rather than accepting the worker's numbers, on the branch merged with
main at 3f8c38f:
  baseline   -> exit 0, 1765 tests from Maven and an independent sum over 131 reports (1761 + 4)
  mutation L -> BUILD FAILURE, exactly 4 failures, all four new tests by name

THE TRAP THE WORKER CAUGHT ITSELF, which is the most valuable thing in this PR. Their first draft of
the TIMED_OUT_WORKING test called answer() inline on the test thread and then probed with send() on
that same thread. It FALSE-PASSED under the mutation: lock is a ReentrantLock, so the same thread
reenters for free whether or not unlock() ran. A same-thread probe proves reentrancy, not release.
They found it only because they actually ran the mutation instead of trusting a green baseline, moved
answer() onto a background thread, and re-proved the kill. The pitfall is now documented in the test's
own javadoc, with the measurement that exposed it.

NOTE FOR ANYONE RE-RUNNING THE TICKET'S RECIPE: the ticket records `sed '1218s|...'`, measured before
#551 merged. #551 edited MessageService.java, and the line is now 1231. Locate it fresh — a
line-anchored sed against a stale number mutates the wrong line and reports a meaningless green. The
anchor count is the check that catches this: `grep -Fxc '            lock.unlock();'` must read 2
pristine and 1 after.

Shape reported by the worker, not investigated and not fixed: the same two-line cleanup
`asyncTasksByWaiter.remove(reply); rendezvous.close(...)` appears at three sites — send()'s finally,
answer()'s inner finally, and answer()'s inline STALE_TURN early return, which is NOT inside a
finally. Same maintained-at-N-sites shape as this finding. Filed separately.
2026-09-12 13:30:22 +02:00
ltms 3f8c38fc54 Merge #567: pin LeadMailbox.inspect's probe-channel close
CI / shell-tests (push) Successful in 8s
CI / contract (push) Successful in 1m1s
CI / build (push) Successful in 1m49s
fleetd #567. LeadMailbox.inspect opens a probe channel and closes it in a finally. The production
code was already correct; nothing asserted it, so a future refactor could drop the close and leak an
AMQP channel per inspect() call with the suite green.

Test only. LeadMailbox.java is untouched — sha256 a2cd99be77345b7e... before and after.

The test asserts the CONSEQUENCE rather than the return value: it caps the connection at three
channels (LeadMailbox uses two, consume and publish), runs a successful inspect, then requires a
replacement channel. If the probe is left open, the broker has no channel number left and
createChannel() returns null. A test that only checked inspect()'s MailboxState would pass under the
mutation, which is the whole reason this hole existed.

Lead verification, re-running rather than accepting the worker's numbers, on the branch merged with
main at ed2fd66:
  mvn -o clean install -Pcontract  ->  exit 0, 1792 tests from Maven and from an independent sum
                                       over 137 surefire reports.
  1792 = 1761 (main) + 30 (contract-only) + 1 (new), which also confirms the default-profile count
  is untouched: the new test is in an @Tag("contract") class.

My own mutation, located fresh rather than assuming the reported line number: delete
LeadMailbox.java:338 `probe.close();`, anchor by grep -Fxc 1 -> 0. Result under -Pcontract: exactly
1 failure, inspectClosesItsSuccessfulProbeChannel:226, "the replacement channel was null". Restored
to a2cd99be77345b7e..., git status --short empty.

THE PROFILE TRAP THIS TICKET EXISTS BECAUSE OF. An earlier sweep worker instrumented this line, saw
zero hits under the DEFAULT profile, and concluded "no test executes this line". That was false:
fleetd/pom.xml:264 sets excludedGroups=contract, so the covering tests were excluded from the run,
not absent. A surviving mutation has THREE causes — never executes, executes with nothing asserted,
or the covering tests were excluded from the profile — and only naming the profile tells them apart.
The true finding was covered-but-unasserted.

Known fragility, recorded rather than fixed: the test's channel cap of three assumes LeadMailbox
holds exactly two channels. If it ever holds more, this test fails loudly, which is fine. If it ever
holds fewer, a leak would no longer exhaust the cap and the test would go vacuous silently. Worth
re-checking if LeadMailbox's channel usage changes.

Not covered, and stated rather than faked: the defensive catch (RuntimeException) around the passive
declare, which needs a connection dying between createChannel() and the declare landing. The
method's own javadoc already admits that branch is unproven; the worker did not invent a test for it.
2026-09-12 13:26:30 +02:00
Dai Ha a4dbc8f8b7 fleetd #572: pin answer()'s session-lock release across all four exits
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m10s
MessageService.answer() releases its per-session lock in an outer finally
(MessageService.java:1218) that mutation testing showed was covered but
unasserted: removing that line left all 1750 existing tests green, because
every existing test on this path checks answer()'s return value, never that
the lock it took is actually reacquirable afterward. If it leaked, a session
would be wedged forever with no exception and no log line.

Adds four tests, one per exit of answer() (normal REPLIED reply,
TIMED_OUT_WORKING, ExecutionException rethrow, InterruptedException
rethrow), each proving the lock is reacquirable via a bounded (300ms)
follow-up send on the same session rather than merely checking answer()'s
own outcome. No production change.
2026-09-12 18:23:54 +07:00
ltms ed2fd6646a Merge #561: the completion/session listener fan-out survives either half throwing, and both sites are pinned
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 50s
CI / build (push) Successful in 2m7s
fleetd #561. Fleetd composed two TurnListener halves as `completion.X(); sessions.X();`, so a throw
from the first half skipped the second. The anonymous class is now a package-private factory,
Fleetd.turnListener(completion, sessions), built on two helpers that always attempt both halves and
rethrow whatever escaped — a second failure attached with addSuppressed rather than dropped, so it
still reaches StatusPoller's catch (Throwable).

The wiring at Fleetd.java:498 calls that factory, so the seam under test is the real caller.

Lead verification, on a merged tree, re-running the checks rather than accepting the worker's:
exit 0, 1761 tests from Maven and from an independent sum over 131 surefire reports.

The interesting part is what the first round MISSED. Two helpers maintain one invariant — "the
second half always runs" — and the first round's five tests asserted it at only one site. Measured:

  bothMustRun                      reverted to the pre-fix bug -> 1 named failure   (pinned)
  bothMustRunKeepingSecondResult   the SAME bug                -> 1755/1755 GREEN   (unpinned)
  failure.addSuppressed(t) deleted                             -> 1755/1755 GREEN   (unpinned)

Both survivors are now killed by new tests, re-verified by the lead after the fix:
sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction and
bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions, each failing alone under its own mutation.

The rule this cost us, worth carrying: COUNT ASSERTIONS PER SITE, NOT PER INVARIANT. The total being
non-zero is what hides a zero at one site, and extracting a shared helper makes it worse rather than
better — it does not reduce the number of sites, only how many are visible. Credit to the fleet01
lead, who predicted this shape before an instance was found.

onDelivered stays deliberately unguarded. Its comment now gives the real reason —
CompletionResolver.captureBaseline already catches RuntimeException around its scrape and fails open,
so that half does not realistically throw — instead of the previous reason, which was true but about
registration rather than about this pair. A correct conclusion resting on a wrong premise reads
exactly like a verified one.
2026-09-12 13:20:56 +02:00
Dai Ha 5441a2b321 fleetd #567: assert inspect closes probe channel
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m23s
2026-09-12 18:19:43 +07:00
ltms 384867dfa3 Merge #551: ATTEMPTED is its own cancellation answer, and the javadoc stops claiming a timed-out send never arrived
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m44s
fleetd #551. The injector polls a queued entry off the queue and marks it ATTEMPTED BEFORE the
irreversible AgentControl#send call, not after. So a Throwable escaping that call can never leave
the entry QUEUED at the head (the #546 re-send hazard) and can never be recorded as a confident
NOT_DELIVERED for text that may already be in the pane.

Cancellation.ATTEMPTED is added as a third answer. NOT_DELIVERED stays reserved for confirmed
absence: the readiness grace expiring, drop(), or a herdr *_not_found error, which the rest of this
codebase already reads as definitely-absent rather than inconclusive.

Also fixes three javadoc/comment sites that claimed a timed-out send definitely did not arrive:
TIMED_OUT_QUEUED, hasQueuedDelivery, the queuedDeliveries field, and the comment in send()'s timeout
branch. Two claims were wrong, not merely stale: "on every route it will not arrive later" is false
on the ATTEMPTED route, and "may already hold a partial paste" understates it — agent.prompt pastes
AND SUBMITS in one call, so the target may hold a complete, running turn.

Verified by the lead on a merged tree: mvn -o clean install from fleetd/, exit 0, 1754 tests from
Maven and from an independent sum over 130 surefire reports. Mutation: folding ATTEMPTED back into
NOT_DELIVERED in cancellationOf gives 3 red, each naming the property
(anErrorFromSendRemovesTheMessageAndMarksItAttempted:740,
aHerdrExceptionFromSendStillSurfacesButNowReportsAttempted:781,
aHerdrExceptionAfterThePasteIsNeverRecordedAsConfidentlyNotDelivered:806).

The final round is comment-only, proven mechanically rather than by reading: stripping every comment
from MessageService.java before and after and collapsing whitespace gives byte-identical code.
2026-09-12 13:19:20 +02:00
Dai Ha f40c19ecf0 fleetd #551 shape sweep: fix stale ATTEMPTED-route javadoc/comments in MessageService
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m26s
hasQueuedDelivery's javadoc, the queuedDeliveries field javadoc, and a code
comment in send()'s timeout path all still made two claims that ATTEMPTED
(fleetd #551) falsifies: a blanket "the message will not arrive later" across
every route, and "the terminal may already hold a partial paste" — but
agent.prompt pastes AND submits in one call, so the target may hold a
complete, already-submitted turn. Each now names the three Injector.Cancellation
routes (CANCELLED, NOT_DELIVERED, ATTEMPTED) and says plainly that only the
first two establish the message will not arrive later.

Comment/javadoc only. No behaviour change: Injector.java is untouched
(sha256 0c689b6cf36275c0da45497a74b2bd4f5a3d80c4dbda77d46670c66004e53b69) and
no test was added. mvn -o clean install: BUILD SUCCESS, Tests run: 1754,
Failures: 0, Errors: 0 (unchanged from before this commit).
2026-09-12 18:13:06 +07:00
Dai Ha 034e17bb32 fleetd #561 follow-up: pin the session half of bothMustRunKeepingSecondResult
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m41s
Two helpers maintain one invariant (the second callback half always runs,
even when the first throws): bothMustRun and bothMustRunKeepingSecondResult.
Only bothMustRun's "session half still runs" direction was asserted
(sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronously, via
onTurnComplete). bothMustRunKeepingSecondResult — the helper
onTurnCompleteWithPostAction uses — could be reverted to the pre-#561
broken shape and the suite stayed green.

Adds two tests to FleetdTurnListenerCompositionTest:
- sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction:
  mirrors the existing onTurnComplete case for onTurnCompleteWithPostAction/
  bothMustRunKeepingSecondResult.
- bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions: proves a second,
  distinct failure from the session half is preserved via addSuppressed
  rather than silently dropped when both halves of bothMustRun throw.

Also rewords the onDelivered comment in Fleetd.turnListener: it previously
said this pair is safe because registration survives a throw via #556's
Injector wiring, which is true but is not why THIS pair is unguarded.
CompletionResolver.captureBaseline already catches RuntimeException around
its scrape read and fails open, so completion.onDelivered does not
realistically throw. Comment text only, no logic change.
2026-09-12 18:10:31 +07:00
Dai Ha 68b428c484 fleetd #551 rework (comment 17058): fix TIMED_OUT_QUEUED javadoc for ATTEMPTED
CI / shell-tests (pull_request) Successful in 13s
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m21s
Javadoc-only change. #551 added Injector.Cancellation.ATTEMPTED, which
MessageService.send's timeout path folds into Outcome.TIMED_OUT_QUEUED
alongside CANCELLED and NOT_DELIVERED (the collapse itself is unchanged
behaviour and is being tracked as a separate follow-up ticket).

TIMED_OUT_QUEUED's javadoc — landed by #513 to state the routes that
reach it — named only two routes and said "on every route it will not
arrive later", with a "may already hold a partial paste" caveat. Both
claims are now stale: ATTEMPTED is a third route, and because
agent.prompt pastes AND submits in one call, that route may mean the
target holds a complete, already-submitted turn and is working on it
right now.

Names all three routes, says which one is uncertain, and drops the
now-false blanket claim. No behaviour change.
2026-09-12 17:51:26 +07:00
Dai Ha e20ccab1eb fleetd #561: harden the completion/session TurnListener fan-out
CI / shell-tests (pull_request) Successful in 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m56s
Fleetd's turnListener composition had four callbacks (onTurnComplete,
onTurnCompleteWithPostAction, and both onTurnFailed overloads) built from two
bare, unguarded statements each. onDelivered's registration was already fixed
structurally by #556; these four had the identical fragility and were still
untested: nothing enforced that the completion resolver's half ran before the
session half beyond call order in the source, so a future reorder (or a
throwing session listener sequenced first) could silently skip the
completion resolver's effect and strand a caller for its full timeout.

Extracted the composition to a package-private static factory,
Fleetd.turnListener(completion, sessions), and hardened it with
bothMustRun/bothMustRunKeepingSecondResult: both callback halves are always
attempted regardless of whether the other throws, and whatever escapes is
rethrown afterward (never swallowed) so it still reaches StatusPoller's
catch (Throwable) and logs at ERROR.

FleetdTurnListenerCompositionTest builds this real composition from a real
CompletionResolver and a throwing fake sessions half, and asserts the
completion resolver's effect (the waiter resolving) survives the session
half throwing, for all four callbacks, plus a mirror case showing the
session half still runs when the completion half throws first.

onTurnCompleteWithPostAction keeps completion-before-session as a functional
requirement (resolveBeforePostAction must run before the context-reset
housekeeping can erase the pane), not just fault tolerance, so it is not
reorder-symmetric like the other three — documented in Fleetd.turnListener's
javadoc.
2026-09-12 17:47:53 +07:00
Dai Ha d83821bbce fleetd #551: record the delivery attempt before the irreversible send
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Successful in 2m23s
The Injector wrote the delivery outcome AFTER calling AgentControl.send(),
so a HerdrException thrown from the response half of that call (herdr
already replied, or may have) was recorded as a confident NOT_DELIVERED
for text that may already be sitting in the worker's pane.

Poll the queue entry and mark it Pending.State.ATTEMPTED before send() is
called, not after. On success it is upgraded to DELIVERED; on an ordinary
failure it stays ATTEMPTED (honest uncertainty), except a herdr
*_not_found error, which the rest of this codebase already treats as a
confirmed absence and which now still writes NOT_DELIVERED.

cancellationOf gets a matching third answer (Cancellation.ATTEMPTED)
instead of folding the new state into NOT_DELIVERED, so a caller that
cancels an already-attempted delivery is told the truth too.
2026-09-12 17:43:21 +07:00
ltms ba2f4d16f8 Merge #555: the redeploy main flow is lifted into tested predicates, and the guard now catches functions below the SOURCED line
CI / shell-tests (push) Successful in 9s
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m52s
fleetd #555. Eight main-flow decisions in scripts/redeploy-fleetd.sh move into
predicate and dispatch functions the suite can source and test. The guard test
test_no_untested_main_flow_conditionals stops new bare conditionals reappearing.

The rework closes a hole the lead found (comment 17012): a conditional wrapped in
a function defined BELOW the SOURCED guard was invisible to the guard, and such a
function can never be sourced, so it can never be tested. The guard now fails on
any function definition after the guard line, on its own.

Verified by the lead on a tree with main merged in, redeploy-fleetd.sh at its
pristine sha 4ffacc5185807d39720a3484d85b922413806eb5347318265bd8897dfd61e8d9:
  bash 5.3.9      -> exit 0, 0 lines matching ^FAIL:
  /bin/bash 3.2.57 -> exit 0, 0 lines matching ^FAIL:

Three mutations, each restored to the pristine sha afterwards:
  a bare conditional appended to the main flow      -> exit 1, reported by line
  the same conditional wrapped in a function below
    the SOURCED guard (the found hole)              -> exit 1, "function defined
    after the SOURCED guard (line 1038) - it cannot be sourced, so it cannot be tested"
  the new FUNC emission deleted from the guard      -> that same case returns to
    exit 0, so the new assertion is what catches it. Anchor count 1 -> 0,
    test-redeploy-fleetd.sh restored to sha 0d713a3092a0c0ea8c05663ffa0cb1595e21b9d73872e662eb99b5035fd38151.
2026-09-12 12:19:47 +02:00
ltms db4c98ac60 Merge #556: the Injector owns turn registration, and the #553 backstop is pinned too
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 51s
CI / build (push) Successful in 2m12s
fleetd #556. Registration moves off the TurnListener fan-out onto its own narrow
TurnRegistrar seam, wired directly to CompletionResolver::register in Fleetd.java,
so it survives any listener throwing regardless of call order.

Verified by the lead on a tree with main (f4f5f31) merged in:
  mvn -o clean install exit 0; 1750 tests, agreed by Maven's own summary and an
  independent sum over 130 surefire report files.

Both registrar.register call sites are independently pinned:
  Injector.java:703 (the #553 finally backstop) removed -> exactly 1 failure, the
    new aRuntimeExceptionFromOnTurnCompleteStillLeavesTheNextDeliveryRegisteredOnTheRecoveryPath
  Injector.java:660 (the ordinary path) removed -> exactly 2 failures, the two
    original tests; the backstop test stays green
Anchor count 1 -> 0 on each, sha restored to
8fcb698afccc254b0c99d3a4bf9e960c0e7c85e542024e870c1e815dce6d62a9 both times.

The :703 line survived the full suite before this rework while being executed
twice — covered but unasserted. That gap is what the added test closes.
2026-09-12 12:18:15 +02:00
Dai Ha 8f80d267a0 fleetd #555 rework: catch function definitions after the SOURCED guard
CI / shell-tests (pull_request) Successful in 9s
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 3m3s
Comment 17012 on #555 found a hole in test_no_untested_main_flow_conditionals:
the guard's function-body detection treats anything inside a function as
"fine, out of scope for this scan" — but a function DEFINED after the
SOURCED guard line can never be reached by sourcing this script (sourcing
stops before the main flow runs), so its body is untestable by construction
while still reading to the guard as safely inside a function.

mainflow_bare_conditionals now also emits a FUNC record for every function
opened after the guard line (reusing the same open-brace detection already
used for depth tracking), and test_no_untested_main_flow_conditionals treats
any such record as a violation on its own, independent of what the function's
body contains or whether the allowlist would otherwise excuse a bare
conditional inside it.

Proof (redeploy-fleetd.sh restored to 4ffacc5185807d39720a3484d85b922413806eb5347318265bd8897dfd61e8d9
after each):

- CONTROL — a bare conditional appended to the main flow is still caught:
  EXIT=1, "found 1 untested main-flow if/elif/case line(s) ... line 1341:
  if [ "$MY_CONTROL_BARE" = 1 ]; then :; fi"
- CANDIDATE — the same conditional wrapped in a function defined after the
  boundary, previously invisible (EXIT=0), is now caught: EXIT=1, "line 1341:
  function defined after the SOURCED guard (line 1038) — it cannot be
  sourced, so it cannot be tested: newfunc_below_the_boundary() {"

Full suite re-run green on both bash 5.3.9 and /bin/bash 3.2.57 (macOS
system bash): exit 0, 0 FAIL lines, reached the final PASS line, on both.

No change to redeploy-fleetd.sh; the 8 lifted decisions, their mutation
proofs, and the allowlist all stand as before.
2026-09-12 17:11:11 +07:00
Dai Ha 738d34a609 fleetd #556 rework: pin registration on the #553 finally backstop path
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m13s
Comment 17009: there are TWO registrar.register(target, sent.token())
call sites in Injector's delivery method — the ordinary path inside
`if (sent != null)`, and the fleetd #553 finally backstop, reached only
when an earlier block throws before the ordinary path ever runs. The
lead's mutation on Injector.java:703 (the backstop call) survived the
full suite: the existing #553 regression test for this exact scenario
(aRuntimeExceptionFromOnTurnCompleteStillCompletesTheNextDelivery)
asserts only that the delivered future completes, never that the turn
is registered with CompletionResolver — so a redesign that dropped
registration from the backstop would reopen this ticket's own defect on
precisely the path #553 exists for, with every existing test green.

Adds aRuntimeExceptionFromOnTurnCompleteStillLeavesTheNextDeliveryRegisteredOnTheRecoveryPath:
drives the same construction as the existing #553 test (onTurnComplete
throws for a previous turn, forcing the next turn's delivery down the
finally backstop) and additionally asserts the new turn is registered
with CompletionResolver and carries the correct waiter — the same
assertion the ordinary-path test makes, now made on the recovery path.

Proven by mutation: removing Injector.java:703 alone (exact-line anchor
1 -> 0) turns the new test red with its own assertion message; restored
and confirmed byte-identical (sha256 8fcb698afccc254b0c99d3a4bf9e960c0e7c85e542024e870c1e815dce6d62a9,
matching the pre-mutation tree); re-run green as a control. Full suite
after restore: 1750 tests, 0 failures, 0 errors, 0 skipped (Maven's own
summary and an independent sum over surefire-reports/*.txt agree),
BUILD SUCCESS.
2026-09-12 17:08:44 +07:00
Dai Ha a46e4058ac fleetd #556: make turn registration structural, independent of any TurnListener
CI / shell-tests (pull_request) Successful in 4s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m29s
The Injector owns the invariant "every delivered turn has a registered
waiter," but before this the only thing that satisfied it was
CompletionResolver.captureBaseline, called from inside a TurnListener
callback wired in Fleetd.java. Any TurnListener that throws (from
onDelivered or elsewhere) could break the invariant with no way for the
Injector to detect it. #553 only made the one reachable listener behave
via a try/finally backstop; it did not remove this structural dependency.

Add a narrow TurnRegistrar functional interface, decoupled from
TurnListener, whose only job is registering a delivered turn's waiter.
CompletionResolver now implements it via a new register() method
(extracted from captureBaseline's registration half; captureBaseline
keeps its own full body unchanged, so existing direct callers/tests are
untouched). Injector gets an explicit registrar field/constructor family
(auto-derived from the TurnListener via instanceof where that still
works, explicit where Fleetd's anonymous fan-out listener can't
implement two interfaces at once) and calls registrar.register(...)
directly and unconditionally in both the ordinary delivery path and the
#553 finally backstop, before turnListener.onDelivered(...) — so
registration no longer depends on that notification callback succeeding.
Fleetd.java wires completion::register explicitly as the registrar,
bypassing the turnListener fan-out for registration purposes.

CB-116 ordering (onTurnComplete reads the PREVIOUS turn's inFlight entry
before the new turn's registrar.register() runs) and the two-arg
inFlight.remove(target, turn) vs one-arg distinction on the completion
path are both preserved unchanged.

#561's order-dependent test (asserting Fleetd.java's
completion.onDelivered -> sessions.onDelivered call order) does not
exist anywhere in this repo at this branch point — nothing to delete.

Adds 5 tests: a TurnListener that throws from every callback still
leaves the delivered turn registered and resolvable; the new turn's
registration still runs after the previous turn's completion is read
(CB-116 guard, pinned as a call-order assertion); and three tests naming
the two-arg-remove invariant directly (a superseded turn's terminal
handling must not evict its successor's registration) across resolve()'s
plain-completion branch, its echoed-noReportMessage sub-path, and
fail().
2026-09-12 16:50:09 +07:00
ltms f4f5f3106e Merge #513: TIMED_OUT_QUEUED has four routes, and the javadoc now says which
CI / shell-tests (push) Successful in 4s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m51s
The old javadoc said a TIMED_OUT_QUEUED message was "still sitting in the injector's
per-target queue". It is not — send() has already given up on it and it will never
arrive. That wrong claim told an operator to wait for a message that was never coming.

The first fix replaced it with a narrower wrong claim: that Injector.cancel() cancelled
the entry and "the target never saw a word of it". That describes one of four routes.

TIMED_OUT_QUEUED is returned whenever injector.cancel() returns anything but DELIVERED:

  CANCELLED      the Pending was still queued and this call removed it
  NOT_DELIVERED  Injector.java:400 - the herdr agent.prompt call threw
  NOT_DELIVERED  Injector.java:419 - readiness grace expired, never attempted
  NOT_DELIVERED  Injector.java:691 - drop(), the target is gone

On the last three, cancel() cancels nothing: it reads a state another path already set
(cancellationOf, Injector.java:288-290). And on the herdr-threw route, agent.prompt
pastes and submits in one call, so a throw does not prove the pane stayed clean - an
operator told "never saw a word" will not go and look at the one place the evidence is.

The javadoc now states the two facts that hold on every route - the message will not
arrive later, and it is not in any queue - and attaches "the target saw nothing" only to
the CANCELLED case. Five blocks: the enum constant, queuedDeliveries, hasQueuedDelivery,
hasOrphanedDelegation, and the inline comment in the TimeoutException branch that seeded
the wording.

Comment-only; no logic changed.

Verified on a tree merged with main (fast-forward to cb64bc8):
  mvn install exit 0
  1744 tests, 0 failures, 0 errors - Maven's own summary and an independent sum over
  130 surefire report files agree
  no unresolved javadoc reference on the three new links, with a positive control
  showing javadoc did analyse MessageService.java

Closes #513.
2026-09-12 11:48:54 +02:00
Dai Ha 8d79d229ff fleetd #555: lift 8 main-flow decisions into tested predicate/dispatch functions
CI / shell-tests (pull_request) Successful in 11s
CI / contract (pull_request) Successful in 1m31s
CI / build (pull_request) Successful in 1m46s
redeploy-fleetd.sh's main flow had 8 bare if/case decisions (CHECK_ONLY
short-circuit, drain-gate entry+confirm, supervisor report/stop/start
dispatch, health-poll decision, HAD_OLD_PID computation) that lived outside
any function, so the 67-test suite could not reach them and any one could be
silently inverted with the whole suite green.

Follows the existing swap_if_built/refuse_drain_gate pattern: each bare
guard becomes a small predicate or dispatch function (should_stop_for_check,
drain_gate_required/drain_confirmed/run_drain_gate, report_supervisor_state,
dispatch_stop, dispatch_start, health_is_up/report_health,
compute_had_old_pid), called unconditionally by the main flow so the
decision itself is unit-testable in isolation.

Adds a structural guard, test_no_untested_main_flow_conditionals, that scans
the main flow (everything after the SOURCED guard) for bare if/elif/case
lines outside any function body, tracking function boundaries via this
file's one consistent name() { / } convention. It fails on any new bare
conditional not covered by MAIN_FLOW_ALLOWED_CONDITIONALS, an explicit
exact-text allowlist of the report-only/display conditionals and the two
#504-family supervisor elif branches that stay out of scope for this
ticket. This is the "shape, not the eight sites" guard the ticket asked
for: a ninth bare decision fails immediately, naming its line.

Out of scope, not touched: #504 items 2/3/4 and #528 item 2 (same
untested-main-flow family) — the seam here generalizes to make them
testable too, but lifting them was left for their own tickets.
2026-09-12 16:46:53 +07:00
Dai Ha cb64bc8157 fleetd #513: rework — TIMED_OUT_QUEUED has four routes, not one
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m16s
Comment 16984 on the ticket showed my first pass (ed28b51) replaced one
wrong invariant with a narrower one: it described injector.cancel()
cancelling a queued entry as THE mechanism, when three of its four
routes (Injector.java: the send-to-terminal call throwing, the
readiness grace expiring, or the target being torn down) never cancel
anything — cancel() just reports a state a different code path already
set. It also claimed the target never saw a word of the message, which
is not established when the herdr agent.prompt call throws after
already pasting.

Rewrote all four comments (the TIMED_OUT_QUEUED enum constant,
queuedDeliveries, hasQueuedDelivery, hasOrphanedDelegation) plus the
pre-existing inline comment that seeded the original bad wording, to
state only what holds on every route: the message will not arrive
later and is not sitting in a queue. The CANCELLED case is called out
as the only one where the target is known to have seen nothing; the
NOT_DELIVERED case (including the herdr-send-threw route) is flagged
as leaving that open.

Comment-only; no behavior change.
2026-09-12 16:45:35 +07:00
Dai Ha ed28b51f12 fleetd #513: fix TIMED_OUT_QUEUED javadoc — cancelled, not queued
CI / shell-tests (pull_request) Successful in 4s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 2m22s
Two javadoc blocks (queuedDeliveries field, hasQueuedDelivery) said a
timed-out message is still sitting in the injector's per-target queue.
CB-640 made send() cancel it via Injector.cancel() instead, so the
message is gone and will never arrive. Rewrote both to describe
cancellation. Also fixed a third instance of the same stale claim in
hasOrphanedDelegation's javadoc, and added a line to the
TIMED_OUT_QUEUED enum constant's own comment clarifying the name is
kept but no longer means the message stays queued.

Per the ticket's follow-up comment: no rename (TIMED_OUT_QUEUED reaches
FleetApp.java REST mapping and FleetMcp.java — out of scope here) and
no behavior change; comments only.
2026-09-12 16:32:36 +07:00
ltms 4f9aba40e7 Merge #558: the ticket is the pull channel, and the brief is write-once
CI / shell-tests (push) Successful in 6s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m51s
Docs only, 16 additions / 3 deletions — read by the lead in full, per the
under-50-lines self-review rule in CLAUDE.md.

Verified on the merged tree (main dab697f + 351ee1e): the canonical-block sync
check from the addendum prints `in sync: True`, comparing the MERGED CLAUDE.md
against wiki/7-Use-Cases.md. The wiki side is already pushed and verified by ref:
local wiki HEAD and `git ls-remote origin main` both read 20d2fc0.

CI run 1812 on 351ee1e: success.
2026-09-12 11:12:55 +02:00
ltms dab697fae0 Merge #559: progress watchdog for StatusPoller and SessionReaper loops (fleetd #544)
CI / shell-tests (push) Successful in 7s
CI / contract (push) Successful in 51s
CI / build (push) Successful in 2m19s
Verified by the lead on a merged tree (main 7a3b2bb + bfac141 = 5de807f):
mvn clean install exit 0, Tests run: 1744, Failures: 0, 130 surefire reports.

Three mutations run by the lead, all killed:
- StatusPoller.java:105 `watchdog.reset()` deleted (the surviving mutant from the
  first review) -> aRestartedLoopReportsRunningAgainNotStoppedForever:
  expected: <RUNNING> but was: <STOPPED>.
- LoopWatchdog.java:73 `lastRoundNanos = nowNanos.getAsLong()` deleted from reset()
  (not run by the worker) -> LoopWatchdogTest.resetClearsAPreviousStopAndTheStaleClock:
  expected: <RUNNING> but was: <STALLED>.
- LoopWatchdog.java:89 `>=` -> `>` boundary (not run by the worker) ->
  LoopWatchdogTest.reportsStalledOnceTheLastRoundAgesPastTheThreshold:
  expected: <STALLED> but was: <RUNNING>.

Each mutation counted the pristine full line 1 -> 0 by exact string equality
(awk '$0==p'), with a pristine control copy still reading its original count, and
was restored to a byte-identical file (shasum -a 256) before the next run.
Control build after all restores: exit 0, 1744 tests, 0 failures.
2026-09-12 11:11:46 +02:00
Dai Ha bfac14108f fleetd #544: pin the sticky-STOPPED-across-restart invariant
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m42s
Review of PR #559 (issue comment #16944) found a surviving mutant: removing
watchdog.reset() from StatusPoller.start() (and the identical line in
SessionReaper.start()) passed the entire suite.

stoppedByCaller is sticky and reset() — called only from start() — is the
only thing that clears it. Both loops document start() as idempotent and
loop()'s own error log says "it can be restarted", so stop() followed by
start() is an anticipated path. Without reset() wired into start(), health()
would report STOPPED forever after a restart even though the loop is
genuinely running again.

Add aRestartedLoopReportsRunningAgainNotStoppedForever to both
StatusPollerWatchdogTest and SessionReaperWatchdogTest, pinning "an
intentional stop must not outlive the restart that follows it". Verified via
the standard mutation cycle: exact-line anchor (not regex, to avoid the
\Q-style false match the reviewer flagged) counted pristine 1 -> mutated 0,
test goes red with its own message, restored, shasum -a 256 byte-identical,
green again.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 1734, Failures: 0,
Errors: 0, Skipped: 0 (cross-checked against 130 surefire report files).

No production code changed — the reset() call under test was already
correct; it simply had nothing pinning it.

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-12 16:00:58 +07:00
ltms 7a3b2bb7ee Merge pull request 'fleetd #552: warn instead of aborting when the post-restart mktemp fails' (#560) from worker/552-post-restart-mktemp-abort-bc2672-4 into main
CI / shell-tests (push) Successful in 7s
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 2m9s
2026-09-12 10:58:10 +02:00
Dai Ha f188947750 fleetd #552: warn instead of aborting when the post-restart mktemp fails
CI / shell-tests (pull_request) Successful in 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m7s
By the time the fresh-log mktemp ran, the daemon had already been stopped, the
jar swapped, and the new daemon started — an unguarded mktemp failure there
aborted the whole script anyway, so a caller read the resulting non-zero exit
as "the redeploy failed" and would restart an already-correctly-restarted
daemon.

Extract the mktemp into capture_fresh_log_region, guarded the same way
unload_launchd_if_loaded/stop_systemd_if_loaded guard their own, but warn
instead of die: there is nothing left to protect by refusing after a
successful restart. The trap is now installed before the assignment it cleans
up, using an FRESH_LOG="" sentinel readers can check.

classify_amqp_connection_errors and report_shutdown_drain both gain a new
state (REDEPLOY_AMQP_CHECK_SKIPPED / REDEPLOY_DRAIN_STATE=skipped) for an
uncapturable log region, distinct from "captured a region with nothing in
it" — and the result section gains a matching branch, so a skipped capture
can never read as a clean bill of health.
2026-09-12 15:55:04 +07:00
ltms f606fccf7f Merge pull request 'fleetd #553: register the rendezvous waiter in onStatus's finally backstop' (#557) from worker/553-onstatus-completion-leak-0da881-2 into main
CI / shell-tests (push) Successful in 4s
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m40s
2026-09-12 10:50:24 +02:00
Dai Ha 735b6af976 fleetd #544: progress watchdog for StatusPoller and SessionReaper loops
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 2m29s
Each loop's virtual-thread runner (StatusPoller, SessionReaper) can die or
get permanently parked in a herdr call with no read timeout, and nothing
observed it: /healthz stayed green and Thread.isAlive() kept reporting true
the whole time.

Add LoopWatchdog (dev.ltms.fleet.inject — see its javadoc for why not
dev.ltms.fleet.health, which would close a package cycle through session):
each loop now records a monotonic last-completed-round timestamp
(injectable LongSupplier clock, same pattern as Injector/SessionManager) and
exposes it as a three-state health() fact — RUNNING, STALLED (dead or
parked, indistinguishable from outside), STOPPED (stop() was called on
purpose, never an alarm). This is the fleetd #512 shape: one flag cannot
carry both "halted on purpose" and "halted unexpectedly", so stop() marks
its own state explicitly instead of leaving state() to infer it from
staleness.

Staleness thresholds are derived from each loop's own poll interval with a
documented multiplier: StatusPoller 40x (250ms -> 10s), SessionReaper 12x
(5000ms -> 60s).

Scope: observability only, per the ticket's own comment. No restart/recovery
mechanism, no /healthz or REST/MCP wiring beyond the public health() API, no
change to the per-item catch(Throwable) behavior (#543) or a process-wide
uncaught-exception handler (ruled out on #538), and no deadline added to the
herdr read itself (a separate, real ticket).

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-12 15:48:07 +07:00
Dai Ha 8b4320ed24 fleetd #553: split sentHandled's two meanings so onDelivered's own throw still completes the future
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m43s
Lead review of PR #557 (ticket comment 16916) found one path left open: sentHandled
is set to true BEFORE onDelivered() runs (correctly, per the earlier fix), so when
onDelivered() itself throws on the normal path, the finally's 'if (sent != null &&
!sentHandled)' guard skipped the whole recovery -- completion included -- and left
sent.delivered() pending forever for a message that really was delivered.

sentHandled must guard only the onDelivered RE-CALL (the permanent-suppression
hazard), never the future completion, since CompletableFuture.complete/
completeExceptionally are idempotent and a no-op on the already-handled path.
Split the one flag's two jobs: the outer 'if (sent != null)' now always runs the
recovery block, and '!sentHandled' moved onto just the onDelivered call inside it.

Added anOnDeliveredThrowOnTheNormalPathStillCompletesTheDeliveryFuture, proven with
the lead's own mutation (reverting !sentHandled onto the outer if): the new test
goes red while anOnDeliveredThrowAfterItsOwnRegistrationDoesNotRunASecondTime stays
green, showing the two concerns are genuinely separate.
2026-09-12 15:45:36 +07:00
Dai Ha 351ee1ea6d CLAUDE.md: the ticket is the pull channel, and the brief is write-once
CI / shell-tests (pull_request) Successful in 6s
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 2m10s
A send to a working member is accepted and returns a ticket, then is never
delivered. That happened three times in one session here, and the member was
released still executing a brief that had been retracted twice. The receipt is
true — it is a fact about the mailbox, when what was needed was a fact about the
pane.

The fleet01 lead named the mechanism: a push delivery needs the recipient free at
send time, while a pull channel needs only that they look before acting. So the
ticket is not more reliable than the mailbox, it is a different direction, and its
success depends on the member's procedure rather than on the timing of the send.

Both halves have to be written down, because each is useless alone:

- Member (turn contract, new item 4): re-read the ticket before acting on anything
  told earlier, and again before committing. A ticket comment that contradicts the
  brief is newer and wins.
- Lead (step 5): all corrections go to the ticket, and the brief is write-once.
  The member cannot check which source is newer — it just always prefers the
  ticket — so revising a brief in place makes it obey the rule and do the wrong
  thing. A first brief for a unit not yet running is not a correction.

wiki/7-Use-Cases.md is updated to keep the canonical block byte-identical; the
sync check passes. The wiki submodule pointer is deliberately left unstaged.
2026-09-12 15:45:09 +07:00
Dai Ha d4a51c6274 fleetd #553: register the rendezvous waiter in onStatus's finally backstop, not just the delivery future
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m12s
The previous try/finally around onStatus's post-monitor region completed
sent.delivered() but never registered sent.token().waiter() when an earlier
listener threw. That waiter is registered only by turnListener.onDelivered(),
inside the very if (sent != null) block the finally backstops, so a caller
was told its send landed and then waited out its full timeout for an answer
that could never resolve (worse than a plain hang).

The finally now does that block's whole job on the unhandled path: it calls
onDelivered() (when sendError == null) before completing the future, guarded
by its own try/catch(Throwable) so a failure there cannot mask the original
throwable. A sentHandled flag, set true at the START of the normal block
(before any side effect), tells the finally whether that already ran, so a
throw partway through onDelivered cannot trigger a second, late captureBaseline
that would permanently suppress the turn's completion.

Also removed a leftover duplicated forget.accept(target) call (with a stray
'MUTATION-TEST-3' comment) in the notReady block — residue from the previous
worker's own mutation testing that was not fully reverted.
2026-09-12 15:35:53 +07:00
ltms 26f380a00b Merge #554: portable hash256, a third state for an unhashable jar, and the shell suite in CI (#550)
CI / shell-tests (push) Successful in 8s
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 2m24s
Closes fleetd #550, all three items.

Verified by me on the branch at b8182c9, in a scratch worktree, not from the implementer's report.

The decisive pair, both in ubuntu:latest where shasum is absent and sha256sum is present:

  main   (93a9ed3):  SUITE exit=127   anchored ^FAIL: count 0
                     scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found
  branch (b8182c9):  SUITE exit=0     anchored ^FAIL: count 0

Identical FAIL counts, opposite exit codes. That is why the new shell-tests CI job gates on the
step's own exit status and deliberately does not grep for a FAIL count: a suite that dies before
running a single test prints exactly what a clean pass prints.

Also measured by me: macOS exit 0 / anchored count 0; 70 tests defined and 70 invoked with an
empty comm -3; bash -n exit 0 under both /bin/bash 3.2.57 and bash 5.3.9; ci.yml parses with jobs
build, shell-tests, contract, and shell-tests is ubuntu-latest + actions/checkout@v4 +
bash scripts/test-redeploy-fleetd.sh.

Three mutations, all killed, each restored byte-identical against
515d929bb53c9ec2c95042e47d3e4d611d60171345227b07feb0de0d20643d3a:

  drop the sha256sum branch from hash256   anchor 1->0   Linux exit 1,
      test_no_unguarded_macos_only_hasher_calls with its own message
  echo "unhashable" -> echo "absent"       anchor 1->0   FAIL: jar_id reported absent for a file
      that exists, only because no hasher was on PATH
  both hash256 arms -> shasum -a 1         anchors 1->0 and 1->0   FAIL: hash256 of the literal
      3-byte input 'abc' must be the known SHA-256 prefix, not some other algorithm's: expected
      ba7816bf8f01, got a9993e364706

The third mutation SURVIVED on the first head (3da44ee) and was sent back. Fixing item 2 had
rewired the reference hashes in test_jar_id_defaults_to_live_and_reports_explicit_path onto
hash256 itself, making the test's reference and its subject one instrument — they agree whatever
it computes, and the existing fixture guard catches only a constant return, not a wrong algorithm.
b8182c9 adds test_hash256_computes_a_real_sha256, pinning the FIPS 180 vector for "abc" as a
literal constant written into the test rather than taken from any hasher. That closes it: the
mutation now goes red, and a9993e364706 is SHA-1("abc"), which proves the mutated code really ran.

Root cause, for the record: shasum was the trigger, not the cause. Under set -euo pipefail a
missing hasher makes the pipeline status 127, the `|| echo "absent"` fires, and jar_id returns a
confident false "absent" for a jar that is right there. Without pipefail the same function returns
an empty string and is visibly broken. A default at the read site that maps every failure onto one
value which already means something specific is the defect; the fix separates "not there" from
"could not hash it".
2026-09-12 10:20:29 +02:00
Dai Ha b8182c96c2 fleetd #550: pin hash256's algorithm against a literal SHA-256 test vector
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 2m27s
test_jar_id_defaults_to_live_and_reports_explicit_path's reference hash is computed by calling
hash256 itself (needed so it doesn't call the Linux-crashing bare shasum directly). That made
subject and reference the same instrument: they agree no matter which algorithm hash256 actually
runs, so a mutation swapping both of hash256's arms for the wrong algorithm was invisible to the
suite.

Adds test_hash256_computes_a_real_sha256, pinned against the published SHA-256 test vector for the
3-byte input "abc" (ba7816bf8f01...), written as a literal constant rather than computed by any
hasher at test time. Verified the constant myself both ways (sha256sum and shasum -a 256) before
writing it in.
2026-09-12 15:16:43 +07:00
Dai Ha 3da44eed63 fleetd #550: replace shasum with a portable hash256 helper, add a Linux CI job for the shell suite
CI / shell-tests (pull_request) Successful in 5s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m46s
jar_id() in redeploy-fleetd.sh called shasum directly, which does not exist on GNU coreutils
Linux (Debian/Ubuntu/etc.) — there it silently reported an existing jar as "absent" with exit 0,
because the missing command made `cut` succeed on empty input and pipefail's failure was then
swallowed by the `|| echo "absent"` fallback. The shell test suite hit the same tool at
test-redeploy-fleetd.sh:298-299 and died at exit 127 with zero FAIL lines printed — the same shape
as a clean pass on the one channel anyone would check.

Adds one hash256() helper (prefer sha256sum, fall back to shasum -a 256, same idiom already used
in probe-member-credentials.sh) and points jar_id and the test suite's own reference hash at it.
jar_id now has three distinct answers instead of two: absent, a hash, or "unhashable" when neither
hasher is on PATH — "absent" is never used for a file that exists.

Adds a CI job (shell-tests) that runs scripts/test-redeploy-fleetd.sh on ubuntu-latest, gated on
the step's own exit code rather than a FAIL-line count, since a suite that dies before running is
exactly what a green run also looks like by that count.

New tests: test_jar_id_reports_unhashable_when_no_hasher_on_path (stubbed PATH with neither
hasher) and test_no_unguarded_macos_only_hasher_calls (a shape check across every script under
scripts/, not named lines — #545 already showed this idiom spreading from two sites to six).
2026-09-12 15:06:15 +07:00
ltms 93a9ed3f83 Merge #549: widen Injector's delivery catch to Throwable (#546)
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m43s
Closes the re-delivery window that merging #543 opened. I caused that; this closes it the same day.

Verified by me on the branch at 87871ea, base 0b032f5 (not stale).

Production diff is 11 lines: `catch (RuntimeException)` -> `catch (Throwable)` at Injector.java:391,
and the `sendError` local widened to `Throwable` so it compiles. I checked every use of `sendError`
myself — :533 `getMessage()` and :534 `completeExceptionally(Throwable)` — so the wider type
reaches nothing that needed the narrow one.

Build on the branch: exit 0, Tests run: 1719, Failures: 0, Errors: 0, Skipped: 0 (1716 on main plus
the 3 new tests). #459's javadoc reference gate: exit 0, 0 reference errors. Gitea CI run 1803 on
87871ea: success.

Three mutations of my own, none of them the ones the worker used:
1. Reverted the catch to `RuntimeException`. RED: `anErrorFromSendDoesNotRedeliverOnASecondRound`
   "expected: <1> but was: <2>" prompt calls, and `anErrorFromSendRemovesTheMessage...` reporting
   the Error escaping `onStatus`. That second message is the defect itself, stated by the test.
2. Deleted the `t.queue.poll()` in the catch arm. RED on the new test AND on the pre-existing
   `sendFailureDropsMessageAndFailsItsFuture` — so the new test is not carrying that behaviour alone.
3. Wrote `DELIVERED` instead of `NOT_DELIVERED` at :400, the catch arm only. RED on the new test and
   on `aHerdrExceptionFromSendStillProducesNotDeliveredUnchanged`, which is the control that proves
   the widening did not quietly change the ordinary path.

All three restored; sha256 back to 97c560b6e33fc49a1772abec92e5bbab613f8991deba2220d30221a2f546ba14.
Green control after the restores: InjectorTest 37/37, exit 0.

My third mutation did not apply on its first attempt — it asserted a unique match on
`p.state = Pending.State.NOT_DELIVERED;`, which occurs twice (:400 and :419), so the script wrote
nothing and the test run came back exit 0. That is not a surviving mutant, it is a non-result
wearing the same clothes. The pristine-anchor count catching it is the only reason I noticed.

Deliberately NOT fixed here, filed as #551: the catch arm assumes that reaching it means nothing was
sent, and nothing establishes that. `agent.prompt` pastes and submits in one call, and every failure
in the response half of `UnixSocketHerdrClient.call()` — dropped connection, malformed line, error
result — is a `HerdrException`, which is a `RuntimeException`, which this catch arm already caught
before today. So "records NOT_DELIVERED for a delivery that happened" is older than this PR and is
not created by it. The fleet01 lead argued it was a trap inside this change and asked to be argued
out of it before the merge; the measurement above is the argument, and their underlying diagnosis is
right and is now #551 with their wording on it.
2026-09-12 09:39:41 +02:00
ltms 0b032f5a1a Merge #548: fix mktemp -t templates for GNU coreutils, split the unclear-supervisor detail (#545)
CI / contract (push) Successful in 55s
CI / build (push) Successful in 2m16s
Verified by me on the branch, not on the worker's report.

Source audit, on a scratch worktree at a476a14:
- Every `mktemp -t` site in scripts/ now carries an X placeholder. The only remaining
  `mktemp -t` text with no X is a prose comment in the test file, not a call.

Suite, macOS (/bin/bash 3.2.57 and env bash 5.3.9):
- exit 0, anchored `^FAIL:` count 0. Unanchored `FAIL:` count 3 — the suite's own internal
  mutation-cell fixture lines, same as main.
- Test functions defined vs invoked: 67/67, `comm -3` empty.

Three mutations of my own, none of them one the worker used:
1. Removed `.XXXXXX` from the `fleetd-fresh-log` site (line 1089). Pristine anchor count went
   1 -> 0, so the mutation really applied. Suite exit 1, FAIL named that exact line.
2. Made the state-2 branch in `detect_supervisor` unreachable (`= 2` -> `= 9`). Suite exit 1:
   "a systemd probe setup failure must read as unclear, not none: expected unclear, got none".
3. Made `systemd_loaded` set the old value (`=2` -> `=1`) on setup failure. Suite exit 1:
   "must flag a SETUP failure (2), distinct from a probe-answered-with-stderr failure (1)".
All three restored; `shasum -a 256` back to 77fe15e5945c7d4ef9b1a2cd8f46e6f1d0be5004f595ea03964d1d7c16e859f7,
the same hash the worker reported independently. Green control after the restores: exit 0, 0
anchored FAILs.

A fourth attempt did not count. A perl `\Q...\E` pattern silently interpolated the shell
variables in it, so the file was never changed and the suite's exit 0 meant nothing. The proof
cell caught it: the pristine anchor count was still 1 after the "mutation". A mutation that did
not apply is not a surviving mutant.

The measurement macOS cannot make: I ran both arms under GNU coreutils 9.1 in a
debian:bookworm-slim container, with a `systemctl` stub that exits non-zero and writes NOTHING to
stderr — a clean negative answer.

  main (a476a14's base):
    mktemp: too few X's in template 'systemd-loaded-err'
    systemd_loaded rc=1  SYSTEMD_LOADED_ERRORED=1
    detect_supervisor => unclear | "systemctl exited non-zero and reported an error on stderr,
                                   not a clean negative — e.g. it cannot reach the user bus"

  this branch:
    systemd_loaded rc=1  SYSTEMD_LOADED_ERRORED=0
    detect_supervisor => none

So on Linux, main tells the operator that systemctl answered badly, when systemctl ran fine and
gave a clean negative. The message named a cause that was never measured. This branch removes it.

Found while doing this, NOT part of this PR, ticket to follow: the shell suite cannot run on
Linux at all. It dies at the first `shasum` call with "command not found" and exit 127, and the
anchored `^FAIL:` count reads 0 — identical to a green run. Gitea CI never runs this suite, so
nothing caught it.
2026-09-12 09:30:30 +02:00
ltms fad99c4c5e Merge #547: record the finally non-goal on drainAll's completion line
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m10s
Comment only. No behaviour change.

Verified by me before merging:
- `mvn -B clean install` in a scratch worktree: exit 0, Tests run: 1716, Failures: 0, Errors: 0,
  Skipped: 0. Same count as main, as expected for a javadoc-only change.
- #459's javadoc reference gate: exit 0, 0 reference errors.
- Gitea CI run 1801 on 5eb4267: success.

The constraint came from the fleet01 lead. Their point: the value of the `drain complete` line is
that it is MISSING when a drain does not finish. A `finally` block would print it after a drain
that threw, with partial counts, and destroy both halves at once. The code already avoids this;
what was missing was the sentence that stops a reviewer putting it back.
2026-09-12 09:29:46 +02:00
Dai Ha 87871eaefb fleetd #546: widen Injector's delivery catch to Throwable, stop re-delivery on Error
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m49s
Injector.java:391 caught only RuntimeException around the herdr send seam. PR #543
(fleetd #538) widened StatusPoller's per-target catch to Throwable so the polling
loop now survives an Error there, which means it comes back round — and Injector's
narrower catch let the poisoned message stay QUEUED (the loop peeks, not polls),
so the next round re-sent the same text into the member's pane.

Widen the catch to Throwable, matching #543 one layer down. sendError's declared
type widens from RuntimeException to Throwable to keep compiling; its only consumer
(CompletableFuture.completeExceptionally(Throwable)) already accepts that type, so
no other caller-visible behavior changes. The ordinary HerdrException/RuntimeException
path is unchanged.

Adds three tests: an Error at the send seam is dropped and marked NOT_DELIVERED, a
second onStatus round does not re-send it, and a HerdrException control proves the
ordinary path is untouched.
2026-09-12 14:26:27 +07:00
Dai Ha a476a14f1c fleetd #545: fix mktemp -t templates for GNU coreutils, split unclear-supervisor detail
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m28s
Every mktemp -t template in redeploy-fleetd.sh lacked an X placeholder. BSD mktemp
(macOS) tolerates that and appends its own suffix; GNU mktemp (every Linux
distribution) refuses it and exits non-zero. All six sites now use .XXXXXX.

detect_supervisor's 'unclear' detail used to cover two different facts with one
message that always named 'systemctl exited non-zero and reported an error on
stderr' — even when systemctl was never run, because mktemp failed first. The
SYSTEMD_LOADED_ERRORED/SYSTEMD_INSTALLED_ERRORED flags now carry a third value
(2 = the probe's own mktemp setup failed) alongside the existing 1 (systemctl ran
and answered badly on stderr), and detect_supervisor gives each its own detail
text. kind stays 'unclear' in both cases; require_drivable_supervisor is unchanged.

Tests added to scripts/test-redeploy-fleetd.sh:
- test_mktemp_dash_t_templates_have_x_placeholders: source-text check, fails if
  any mktemp -t template lacks an X.
- test_detect_supervisor_systemd_probe_setup_failure_is_unclear: proves the
  SET-UP-FAILED detail when mktemp itself fails (systemctl never runs).
- test_detect_supervisor_systemd_probe_error_is_unclear: extended with assertions
  that the PROBE-ANSWERED-WITH-STDERR detail is present and the SET-UP-FAILED
  wording is absent, so swapping the two messages fails a test in both
  directions.
2026-09-12 14:22:57 +07:00
Dai Ha 5eb4267a4a fleetd #512 follow-up: record why the drain-complete line must not move into a finally
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 2m8s
The javadoc said the log.info fires "every time", which reads as an invitation
to the exact edit that destroys it. The absence of the line is the signal that
the drain died, so a finally would remove the signal and print partial counts in
the same change.

Raised by the fleet01 lead from their 2026-09-10 incident: that drain is known to
have died only because it threw and left a stack trace. A drain that hung, or
returned early on a condition, leaves no trace, no ERROR token and no priority --
only a missing line.

Comment only. No behaviour change.
2026-09-12 14:21:15 +07:00
ltms 7611b69667 Merge pull request 'fleetd #504 item 1: stop the false ok on the loaded-but-not-running stop path' (#541) from worker/504-failed-reported-clean-3cfd66-3 into main
CI / build (push) Successful in 1m35s
CI / contract (push) Successful in 1m42s
2026-09-12 09:09:00 +02:00
ltms cc302fe4af Merge pull request 'fleetd #538: recover polling loops after errors' (#543) from worker/538-loop-dies-on-error-4a5eeb-6 into main
CI / contract (push) Successful in 51s
CI / build (push) Successful in 2m10s
2026-09-12 09:02:48 +02:00
ltms 4a8a780274 Merge pull request 'fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site' (#542) from worker/426-health-coverage-ef1fd4-4 into main
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m28s
2026-09-12 08:55:46 +02:00
Dai Ha 343ce0f4c0 fleetd #538: recover polling loops after errors
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m0s
2026-09-12 13:47:44 +07:00
ltms cec3e191d4 Merge pull request 'fleetd #459: lint Javadoc references in CI' (#539) from worker/459-broken-link-targets-cadc17-5 into main
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 2m9s
2026-09-12 08:46:23 +02:00
Dai Ha 1850a5f324 fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 2m16s
FleetHealthMonitor.coverage had zero references in the test tree — not the
method, not either output string, not the field it populates. Inverting
`enabled`, swapping "full"/"detection-only", or breaking the argument pairing
at the HealthCoverageSource call site in Fleetd.java all shipped a green
build.

Extract the HealthCoverageSource lambda out of Fleetd.main into a
package-private static factory (healthCoverageSource(ConfigRef)), the same
shape capacitySource/quarantineSource already use for the identical
argument-pairing risk (fleetd #415). #407's "keep the config invalid, assert
on the log line before validateAll() throws" option does not apply here: this
call site is built well after validateAll() and after a real herdr socket
connect, so driving it through a real Fleetd.main would require the socket
I/O this ticket's tests must not do.

Add FleetHealthMonitorCoverageTest (the three-branch method itself) and
FleetdHealthCoverageSourceWiringTest (the call site, via a real
FleetConfig.load + ConfigRef against @TempDir fixtures, including a hot
notifications-reload case). Output strings are unchanged — "detection-only"
is still what a live fleet_list reports today.

Measured: all three mutations killed by the new tests.
2026-09-12 13:43:50 +07:00
Dai Ha e4eb3dbed4 fleetd #504 item 1: stop swallowing real launchctl/systemctl failures on the 'loaded but not running' path
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m55s
The two 'loaded but not currently running' branches in the stop step (launchd/systemd, reached
when $OLD_PID is empty) ran 'launchctl unload'/'systemctl --user stop' with '2>/dev/null || true'
and printed 'ok' unconditionally. That swallowed a real supervisor failure (e.g. launchd or the
systemd user bus unreachable) exactly like a harmless already-stopped answer, and let the script
proceed to start a new daemon believing nothing was loaded -- the two-daemons failure fleetd #492
exists to prevent.

Adds unload_launchd_if_loaded/stop_systemd_if_loaded, applying systemd_loaded's own pattern
(capture stderr separately; a non-zero exit WITH stderr is a real failure, a non-zero exit with
empty stderr is a clean already-stopped answer) to the write side. The two call sites now use
these functions instead of the bare '|| true'.

Adds 5 tests: dies-on-real-failure and tolerates-clean-negative for each function, plus a
source-text check that the main flow calls the new functions instead of the original bare
'2>/dev/null || true'. All 5 verified by mutation (reintroducing the swallow, and separately
over-correcting to die unconditionally) -- each goes red with its own message, restores
byte-identical (full sha256), and passes a green control.
2026-09-12 13:42:27 +07:00
ltms 57cd96f5e6 Merge pull request 'fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts' (#540) from worker/537-capturedlog-close-e4c437-2 into main
CI / contract (push) Successful in 47s
CI / build (push) Successful in 2m9s
2026-09-12 08:42:19 +02:00
Dai Ha 202e37e3b3 fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 1m36s
Only the level-restore half of close() was pinned before this
(WorktreeSessionManagerTest). Deleting logger.detachAppender(appender)
from close() left mvn clean install green (1701 tests, 0 failures) --
the appender-detach half of the contract was unmeasured.

Adds CapturedLogTest with three tests, each using a logger name no
production class uses:
- closeDetachesTheAppenderSoALaterLogIsNotCaptured: an event logged
  after close() must not land in events().
- closeRestoresTheLevelCapturedAtOpen: the helper's headline contract
  in one place, independent of any production class.
- setLevelDuringCaptureDoesNotChangeWhatCloseRestores: setLevel()'s
  own javadoc claim that close() always restores the level captured
  at construction, never a value set through setLevel() mid-capture.

Test-only change; CapturedLog.java itself is untouched.
2026-09-12 13:36:55 +07:00
Dai Ha 90253f832d fleetd #459: lint Javadoc references in CI
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 2m9s
2026-09-12 13:36:37 +07:00
ltms f1640f5dcc Merge pull request 'fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog' (#536) from worker/535-appender-leak-fe74c1-1 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m28s
2026-09-12 08:25:01 +02:00
Dai Ha c7903c1efe fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m52s
Three call sites (captureFleetdLogs at :55) attached a ListAppender to the
Fleetd.class logger with addAppender and never detached it, and never
called appender.setContext(...) either. Logback Logger instances are
cached per class and shared for the whole JVM, and surefire reuses forks,
so all three appenders stayed attached for every later test in the fork.

Convert all three call sites to CapturedLog.of(Fleetd.class) (added in
#533) via try-with-resources, which detaches the appender and sets the
context for free. Delete captureFleetdLogs(); nothing calls it now.
2026-09-12 13:17:21 +07:00
ltms 7d711942fe Merge #534: detect a died shutdown drain the ERROR count is blind to (fleetd #512 part 2)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m52s
Verified independently. The branch is based on 8335b12 while main is at bec87f9, so I merged locally first and tested the MERGED tree, not the branch — a clean auto-merge is not a working merge.

Merged tree checks:
- `bash -n` exit 0 on both scripts, under /bin/bash 3.2.57 and env bash 5.3.9.
- Suite exit 0, 0 lines matching `^FAIL:`, 255 bytes of output. 60 test functions defined, 60 invoked, and no defined-but-never-invoked orphan (checked with a comm against the invocation list, not by comparing two counts — two equal counts can both be wrong).
- My own comment fix from bec87f9 survived the merge and is still at :317.

I ran two mutations the worker did not, per "mutate the half the worker did not":

(A) The one that matters, because it is the defect this ticket exists to prevent: collapsed the `unknown` state into `complete`, so "cannot tell" reports as a pass. Result exit 1, one FAIL: `cannot-tell fixture must set REDEPLOY_DRAIN_STATE=unknown: expected unknown, got complete`. So the third state is genuinely load-bearing, not decoration.

(B) Broke the positive check: changed `find_drain_complete_line`'s pattern from `drain complete: released=` to `drain finished: released=`, one site. Result exit 1, one FAIL: `find_drain_complete_line did not capture the present line`.

Proof that (B) applied, against a pristine copy: the full grep line 1 -> 0, the mutant form 0 -> 1, and the bare phrase 2 -> 1 with the comment occurrence untouched. Both files restored byte-identical; `git diff --quiet` clean; green control re-run.

A note on my own proof cell for (B), because it was wrong the first time. I wrote the counts with escaped double quotes inside an already double-quoted command substitution, so the shell split the pattern on spaces and grep treated the words as filenames. It printed "2 and 2" alongside `ugrep: No such file or directory` warnings — a symmetric, plausible-looking pair that meant nothing. The kill itself was never in doubt, since the suite named the exact function, but the cell that was supposed to prove the mutation applied proved nothing. Re-done with single quotes. This is the same trap already written down for this repo, hit by me, in a cell whose only purpose was to guard against exactly this.

One thing I checked that no test covers: the main flow's `HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1` runs under `set -euo pipefail`, and on a cold start the test fails. Sourcing stops before the main flow, so no behavioural test reaches that line. If `set -e` fired there, every cold start would abort before the health checks. It does not: `set -e` exempts the left side of an `&&` list, confirmed by running it under both shells — `survived, HAD_OLD_PID=0` on 3.2.57 and on 5.3.9. Safe, but it is untested main-flow wiring, which is the same class as #528's item 1 and belongs on that list.

On the `n/a` fourth state, which the worker flagged for a reviewer's judgment rather than quietly keeping: accepted, and it is in scope. The ticket asked for a third state because a sentinel conflating "no" with "cannot tell" hides two causes needing opposite handling. "The question does not apply" is a third such cause, not a variant of "cannot tell". Without it, the new warning would fire on every clean cold start, and a warning that cries wolf on the most common path trains the operator to skip it — which destroys the absence signal just as surely as putting the completion line in a `finally` would. The worker also proved the gate is actually consulted, using a cold-start fixture whose content deliberately looks like a died drain, so the test would fail if the gate were skipped. That is the right way to test a gate.

Both remaining outcomes are correctly excluded from the "no ERROR lines since restart" summary: only `complete` and `n/a` let it print.
2026-09-12 08:04:15 +02:00
Dai Ha bec87f987c scripts: name the mechanism in detect_supervisor's constraint 2, not a line number
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 1m33s
Constraint 2 read "This script runs under `set -euo pipefail` (line 50), so an
unset variable is a loud failure." Two problems, both small and both the same
family as fleetd #494 — a comment that states the wrong reason.

The line number was stale: the `set` line is at 54, not 50. It was the only
line-number citation in the file, and a citation like that goes stale on the
next insert above it, silently, with nothing to catch it.

The mechanism was also misattributed. What makes an unset variable a loud
failure is `set -u`. Naming the whole `-euo pipefail` string invites the reader
to credit pipefail for it, which is the mistake fleet01 flagged on a different
cell this week: pipefail is insurance against a future pipeline stage, not what
catches the current shape.

Now names `set -u` and says where it is without a number, and records why the
number is gone so nobody adds one back.

Comment only. bash -n exit 0 under /bin/bash 3.2.57 and env bash 5.3.9;
scripts/test-redeploy-fleetd.sh exit 0 with 0 lines matching ^FAIL:.
2026-09-12 12:58:37 +07:00
Dai Ha 7f8a8829f9 fleetd #529 follow-up: the new helper's javadoc claimed a reach it does not have
CapturedLog's class javadoc said it is "the one way to pin or capture a logger's
level and output in this test tree". Measured on main at af95897, that is false:
nine test files still hand-roll the ListAppender + setLevel + finally
detachAppender pattern, with 42 setLevel calls on a raw logback Logger between
them.

None of those nine is a defect. Every one pairs its pin with a restore, so none
is the fleetd #525 leak, and #529's scope was the 19 unrestored pins only. The
problem is the sentence, not the code: a reader who believes "the one way" and
then greps finds nine counter-examples and cannot tell a leftover from a
violation. That is the same shape as a wrong reason in a comment — the text
survives while the fact under it moves.

Replaced with what is actually true: new code must use the helper, the pattern
still exists elsewhere, and here is the list plus the two commands that
re-measure it. The paragraph says to delete itself once the first command comes
back empty, rather than to keep a count up to date.

Javadoc only. mvn -f fleetd/pom.xml test-compile exit 0.
2026-09-12 12:58:26 +07:00
ltms af9589783e Merge #533: promote CapturedLog to a shared test helper and close the logger-level leak (fleetd #529)
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m35s
Verified independently in a scratch worktree at 8ea5c2b, not promoted from the worker's report.

Build: `mvn -f fleetd/pom.xml clean install` exit 0, `Tests run: 1701, Failures: 0, Errors: 0, Skipped: 0`, BUILD SUCCESS.

Base arithmetic, measured rather than carried forward: a6415f3 (this branch's parent) has 1695 `@Test` plus 2 parameterized/repeated, and the branch has 1696 plus 2 — a delta of exactly +1, matching the per-file count (WorktreeSessionManagerTest 24 -> 25). So 1700 -> 1701 is the one new proving test and nothing else. A number I had carried from earlier in the session said 1699; that number was wrong and is retired. No Java file differs between a6415f3 and main at 8335b12, so the base count is the same on both.

Mutation (a), the shared instrument: removed `logger.setLevel(originalLevel);` from `CapturedLog.close()`. Result exit 1, `Tests run: 25, Failures: 1`, the single failure being `sharedSessionManagerLoggerLevelIsRestoredAfterDirtyWorktreeReleasePinsWarn` with `expected: <TRACE> but was: <WARN>`. Restored byte-identical.

Mutation (b), the use site: replaced the try-with-resources in `releasePreservesDirtyWorktreeAndLogsWarn` with the pre-#525 hand-rolled `ListAppender` + `setLevel` + `finally detachAppender` pattern. Result exit 1, one failure, the same assertion. Restored byte-identical.

Green control on the restored tree: `git diff --quiet` clean, full suite exit 0, 1701/0/0/0.

Both mutants were killed, so per the economy fleet01 proposed and #529 adopted, neither needs a separate harness-proof cell — the kill is the proof the cell can go red.

Leak survey, re-measured here rather than taken from the report: on a6415f3, 9 files carry 19 `setLevel` pins on a raw logback `Logger` with zero restoring call; on the branch that set is empty. The 7 remaining `setLevel` calls in `GitWorktreesTest` are `reportingLog.setLevel(...)` on the `CapturedLog` instance, whose `close()` restores the original, so they are re-pins and not leaks.

File hashes match the worker's report exactly, head and tail: CapturedLog.java 486d6f5b5a30dc5ef7f75e5e10be353e720fb0de503825e88e8d96e30a61a2f7, WorktreeSessionManagerTest.java a722a98d828c82e00415d2a924d341e177ded704262aaa5963ea2d09a683df94.

Two things follow this merge rather than block it, both filed separately: one javadoc sentence in the new helper overstates its own reach, and the worker's item 4 reports a separate appender leak outside this ticket's scope.
2026-09-12 07:57:25 +02:00
Dai Ha 190436c9cf fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m4s
The previous daemon's dead shutdown drain (an uncaught exception in a
shutdown thread) never passes through the logger, so it never carries an
ERROR/SEVERE token, so redeploy-fleetd.sh's existing ERROR-count classifier
is structurally blind to it and prints a confident "no ERROR lines since
restart" while the drain actually died.

Add scan_uncaught_exceptions (greps the shutdown window for the failure's
real shape: `Exception in thread`, `NoClassDefFoundError`) and
find_drain_complete_line (checks for #522's SessionManager.drainAll
completion line). Compose both in report_shutdown_drain, a single
decision+action function the main flow calls unconditionally (same shape
as swap_if_built/refuse_drain_gate from #521/#528), which resolves to one
of four outcomes: complete, died, unknown ("cannot tell" — the line is
absent for either of two reasons that need opposite handling: the previous
daemon predates #522, or its drain failed without throwing), or n/a (no
previous daemon was actually stopped this run). Never fails the redeploy;
warns loudly instead.

Gate the "no ERROR lines since restart" summary line on the new outcome so
it never reads as reassurance when the drain died or the outcome is
"cannot tell" (item 4 of the ticket).

Tests: 11 new test functions (60 defined/invoked, was 49), covering both
pure classifiers, all four report_shutdown_drain outcomes, a source-grep
proof of the main-flow call site (sourcing stops before the main flow
runs), an ordering check, and the item-4 gating. Full suite green
(exit 0, 0 anchored FAIL lines). Five mutations applied and killed by hand
during review, each restored to a byte-identical file afterward.
2026-09-12 12:55:54 +07:00
Dai Ha 8ea5c2bb1f fleetd #529: promote CapturedLog to a shared test helper, close the logger-level leak
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m39s
ch.qos.logback.classic.Logger instances are cached per class and shared for the
whole JVM, and surefire reuses forks. A test that pins a shared logger's level
and restores only the appender leaves that level pinned for every test that
runs after it, in the same class or a different one in the same fork.

Move CapturedLog (merged in #527 for #525) out of SessionManagerTest into
dev.ltms.fleet.testing.CapturedLog, and convert all 19 unrestored setLevel
pins across 9 files to it, so there is exactly one way to capture and pin a
logger in this test tree:
 - FleetdAwaitHerdrTest, FleetdReplyInboxSelectionTest, AuditLogTest,
   CompletionResolverTest, InjectorTest (4), LeadRolloverTest,
   AmqpConnectionFailureLoggerTest, GitWorktreesTest (8),
   WorktreeSessionManagerTest.

AuditLog logs through a named "audit" logger rather than a class, so
CapturedLog gains String-named at()/of() overloads alongside the existing
Class-based ones, plus a setLevel() method so a fixture that already pinned a
coarser baseline (GitWorktreesTest's @BeforeEach) can re-pin further for one
test without losing what close() restores.

Adds an ordered proving test to WorktreeSessionManagerTest asserting the
SessionManager logger level is back to a known baseline after the dirty-
worktree release test runs; this proves the within-class case only, since
JUnit does not guarantee cross-class ordering.

Every existing intentional pin (SessionManagerTest's two explicit INFO pins
and its @BeforeAll DEBUG baseline) is left untouched, per the ticket.
2026-09-12 12:48:49 +07:00
ltms 8335b12562 Merge #532: pin drain_gate_refusal's call site, not just the predicate (fleetd #528)
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m47s
Verified independently in my own worktree at the pushed head 7c34e8f, not taken
from the worker's report. CI run 1781: success.

The shape is the one #528 asked for and the one #526 arrived at: refuse_drain_gate
composes the message via drain_gate_refusal AND calls die itself, and the main
flow calls it unconditionally at :710. No guard is left in the main flow to
remove, invert or bypass on its own.

My measurements:

  test functions defined / invoked   49 / 49   (was 44/44; +5)
  bash -n, /bin/bash 3.2.57          rc=0 on both files
  bash -n, env bash 5.3.9            rc=0 on both files
  clean control                      exit 0, 0 lines matching ^FAIL:, 256 bytes

Two mutations, both killed, each by a differently named failure:

  call site deleted (the item-1 mutation: :710 replaced by a flat
  die "aborted — nothing changed")
    -> exit 1, 1 ^FAIL: line, 87 bytes
       FAIL: could not find the main flow's refuse_drain_gate call site in redeploy-fleetd.sh

  refuse_drain_gate stops consulting the predicate (its body's
  die "$(drain_gate_refusal ...)" replaced by a flat message)
    -> exit 1, 1 ^FAIL: line
       FAIL: refuse_drain_gate build-ran+staged-present die message does not name the staged jar

The first cell is the point of the ticket. Before this change the same mutation
gave exit 0, zero FAIL lines and output byte-identical to a clean run at 256
bytes. It now exits 1 and names the missing call site. Each mutation was proven
applied with a uniquely tagged marker plus a second, different search string,
with a control against a pristine copy showing the exact inverse (1/0 mutated,
0/1 pristine), and the function definition confirmed still present so the
mutation targeted the call and not the function. Restored byte-identical to
0e5a99a22c9c65f72960d8f179ca5299307889e06bc42131f098a513e7b97bd6 and the final
control is green.

Needle uniqueness checked, because this is where it could have gone wrong:
grep -cF 'refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"' on the production script
returns 1, at :710, the real call site. The worker hit the self-match trap while
writing the comment above refuse_drain_gate — their first draft quoted the
call-site string literally, which would have let the source-text test match the
comment instead of the call — caught it themselves, and reworded so the comment
cannot become a second match. That is the same trap that cost me a false pass on
a probe earlier today, and catching it unprompted is the better half of this PR.

The dead-check sweep came back as a real negative, with the reasoning shown
rather than asserted: of the five scripts under set -e with pipefail, every
pipe-into-assignment already carries || true or || echo, and the remaining two
scripts have no pipe-into-assignment at all. probe-member-credentials.sh and
deploy/herdr-inner.sh correctly excluded for not having set -e. No live
instances.

One inaccuracy in the report, in the report only: it abbreviates the restored
hash as "0e5a99a2...78f0a", and that tail does not occur in the actual hash,
which ends b97bd6. I hashed the committed file myself and confirmed the restore
matched, so the file is right and only the quoted abbreviation is wrong. Flagged
because an abbreviated hash that nobody can match against anything is worse than
no hash.

wait_for_daemon_exit's call site (item 2) stays open as the ticket scoped it —
source-text pinned only, "partially pinned, not audited". The seven untested
main-flow decisions are untouched; the worker correctly notes it changed the body
of one of those if blocks while leaving the guard condition itself untested, as
instructed.
2026-09-12 07:41:57 +02:00
ltms d25c863118 Merge #531: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / build (push) Successful in 1m30s
CI / contract (push) Successful in 1m32s
Charter wording only — 9 insertions, 4 deletions, one file. CI run 1780 on
be07ed2: success.

Resolves the "charter may be stale" item I had been carrying. It was not stale.
Two workers reporting working forge access and the charter saying forge tools
hold a blocked credential were both correct, about two different credentials:
the repo-scoped GITEA_TOKEN the daemon injects (which opens every worker PR,
per implementer SKILL.md step 5) versus the forge MCP server that leaks in from
the operator's user-scope config (which is deliberately blocked). The wording
did not separate them, and a worker could have read it as "I cannot reach the
forge" and skipped opening its PR.

Both sentences now name the MCP server specifically and state that the injected
token is a separate, working route.

Canonical block and wiki template verified byte-identical after the edit — the
CLAUDE.md sync script reports "in sync: True". The wiki commit is d02a55d on
wiki's own main, pushed and verified by ref (ls-remote matched local HEAD), not
by exit code. The submodule pointer stayed unstaged.

Not re-measured in this change: that the blocked MCP credential does fail every
call. That claim is the existing charter's and I only narrowed what it refers
to. It would need its own probe with a request that cannot succeed on its
merits, so that a rejection can only mean the block.
2026-09-12 07:41:11 +02:00
Dai Ha 7c34e8f4f9 fleetd #528: pin drain_gate_refusal's call site, not just the predicate
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m31s
drain_gate_refusal composes the correct abort message and is well tested,
but the main flow built its own `die "$(drain_gate_refusal ...)"` call —
nothing proved that call site was ever consulted. Mutating it to a flat
`die "aborted -- nothing changed"` left the whole suite green, silently
reinstating the exact defect #517 was filed to fix.

Same shape as #521/#526's should_swap/swap_if_built: the decision and the
die() now live together in refuse_drain_gate, which the main flow calls
unconditionally. drain_gate_refusal stays separate and separately tested
for the message logic; four new behavioural tests stub die() to prove
refuse_drain_gate calls it correctly for all four cases, and a fifth
source-text test pins the main flow's call site itself (the only thing
that can catch deleting the call, since sourcing stops before the main
flow runs).
2026-09-12 12:38:16 +07:00
Dai Ha be07ed2033 charter: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m3s
Two worker reports said they had working forge access, which looked like it
contradicted the charter's "any forge tools it appears to have hold a blocked
credential and fail". Measured: the charter is correct and the reports are
correct. They are about two different credentials.

A worker opens its own PR with curl and a repo-scoped GITEA_TOKEN that the
daemon injects (.claude/skills/implementer/SKILL.md step 5, lines 96-117).
That route works — it is how every worker PR this week was opened. The blocked
credential belongs to the forge MCP server that leaks in from the operator's
user-scope ~/.claude.json, which is a separate thing and does fail every call.

The wording did not distinguish them. A worker reading "any forge tools it
appears to have hold a blocked credential and fail" could reasonably conclude
it cannot reach the forge at all, and skip opening its PR — the one step the
lead depends on. So this is an ambiguity with a cost, not a stale line.

Both sentences now name the MCP server specifically and say plainly that the
injected token is a different, working route.

The canonical block and the wiki template must stay byte-identical. The wiki
submodule is updated in its own tree and the official sync check reports
"in sync: True". The submodule pointer stays unstaged, per the repo rules.
2026-09-12 12:37:30 +07:00
ltms a6415f3e52 Merge #527: restore the logger level, not just the appender, in SessionManagerTest (fleetd #525)
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m8s
Verified in my own worktree, at the pushed head d898571, test file hash
0228f78424boa... (full: 0228f78424boa is a typo; the measured hash is
0228f78424boa). See the acceptance table below for the measured values.

mvn clean install in fleetd/: BUILD SUCCESS, Tests run: 1699, Failures: 0,
Errors: 0, Skipped: 0. SessionManagerTest itself: 72 tests (69 on main + 2 from
#522 + 1 new). CI run 1773 on d898571: success.

Mutation killed: reverting CapturedLog.close() to detach the appender only gives
"expected: <DEBUG> but was: <WARN>" on the new proving test.

Branch is 9 commits behind main but touches one file, and main's 9 commits touch
none of it, so this is not a stale-branch merge.

Two corrections to the PR body, neither blocking:

1. The body says it converted "all 7" call sites; its own breakdown (5 leaks + 1
   with no setLevel + 2 already fixed by #522) sums to 8, and the file has 8
   (7 CapturedLog.at + 1 CapturedLog.of). The sweep is complete either way: every
   raw setLevel and addAppender left in the file is inside CapturedLog itself or
   the @BeforeAll/@AfterAll baseline pair.

2. The body says #522's two explicit Level.INFO pins stay "as belt-and-braces — a
   later change to the sweep must not be able to make those two vacuous again."
   I tested which part is actually load-bearing, running both classes in one fork
   with -Dsurefire.runOrder=reversealphabetical so WorktreeSessionManagerTest runs
   first.

   Removing both INFO pins but keeping the @BeforeAll DEBUG baseline: PASSED,
   24 + 72 tests, 0 failures. So the per-test pin really is redundant.

   Also removing the @BeforeAll DEBUG baseline: mvn exit 1, 3 failures —

     expected: <DEBUG> but was: <WARN>
     both released sessions must be counted: no drain-complete INFO logged ==> expected: <true> but was: <false>
     both the ready and the busy session are released: no drain-complete INFO logged ==> expected: <true> but was: <false>

   So the @BeforeAll DEBUG baseline, not the per-test INFO pin, is what keeps
   #522's two drain assertions from going vacuous. Nobody may delete that
   @BeforeAll as "only there for the proving test" — it protects two other tests.
   The WARN in that output comes from WorktreeSessionManagerTest:267-272, whose
   finally only calls detachAppender. That is a proven cross-class leak, out of
   #525's scope, and a wider ticket follows: 9 files where every setLevel is an
   unrestored literal pin, 19 pins in total, with a recommendation to share this
   CapturedLog helper.
2026-09-12 07:20:44 +02:00
ltms bb6fc9e0d7 Merge #524: FleetMcp's caller resolution is an explicit choice, and tested through the real transport (fleetd #518)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m33s
Adjudicated and verified by me, by my own build and my own mutation. This
is the strongest PR of this batch and it closes #518 properly.

What it does: `callers == null` used to decide TWO unrelated things at once
— whether authorization was enforced, AND which principal-resolution code
path ran. Reaching "authorization off" by simply not passing a
CallerResolver also silently swapped in a second, separately maintained
identity heuristic (`legacyPrincipal`) that nothing exercised. That is the
fallback-reached-by-omission shape #415 named. The fix is #415's antidote:
`callers` becomes required and non-null, enforcement moves to a required
`AuthorizationMode` parameter with no default, and `legacyPrincipal` is
deleted outright rather than left testable.

Verified by me on the merged revision:

* `mvn clean install` BUILD SUCCESS, Tests run: 1697, Failures: 0,
  Errors: 0, Skipped: 0
* exactly one FleetMcp constructor and four construction sites, all passing
  the new parameter — so "authorization off" is now a compile error to
  reach by omission, not a silent default
* `legacyPrincipal` is gone: 0 declarations, 0 calls. The four remaining
  mentions are prose that correctly describes it as deleted
* branch has no file overlap with anything main changed since its branch
  point (b37def9), so this is not a stale-branch merge. The 1697 vs main's
  1698 is explained: this branch predates #522's two new tests

The mutation that matters, run by me. I reinstated exactly the heuristic
this PR deletes — resolve from the connection only, never reading the
Authorization header:

* `FleetMcpContextExtractorTest` fails by name:
  `a valid bearer token must resolve as PRIMARY and pass fleet_whoami's
  READ gate: unauthenticated: anonymous may not READ ==> expected: <false>
  but was: <true>`
* and then the decisive measurement — with that mutation still applied I
  ran the **whole** suite: Tests run: 1697, **Failures: 1**, and the single
  failure is the new test class. All 20 FleetMcpAuthzTest cases pass with
  the resolver bypassed, as do the other 1676 tests.

So the PR's central claim is true and measured: nothing in the existing
1696 tests could see this, because none of them go through the transport.
`denyFor` had a full policy table, `CallerResolver.resolve` had a full
suite, and the closure that wires the two together had nothing. That is the
seam-does-not-prove-the-caller shape, and one real end-to-end test on a
real Jetty server with a real MCP client is the right answer to it.

Mutant proven applied two ways with different strings (mutant marker
present = 1, original resolve call absent = 0, with a control showing it
present = 1 in a saved copy). Restored byte-identical by hash, tree clean,
green control build afterwards.

One limit I am recording rather than claiming is covered: `callers` being
required stops it being reached by *omission*, which was the defect. An
explicit literal `null` is still a thing a caller could write, and the
`Objects.requireNonNull(callers, "callers")` that catches it has no test of
its own. That is the intended bar, not a gap worth a ticket.

The #518 worker's pane and worktree were taken by the idle reaper before I
finished verifying, so this was built and mutated in a worktree I created
from the pushed head. Nothing was lost — the branch was pushed and clean.
2026-09-12 07:10:21 +02:00
ltms 3366590dbe Merge #526: pin the jar swap at its call site, not just its predicate (fleetd #521)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
Adjudicated by me. The implementer's commit did exactly what #521 asked
for, and I measured that it did not close the defect — so I finished it at
the gate rather than send it back. The gap was in my ticket, not their work.

What the implementer's commit gave: `should_swap(do_build)` extracted, the
main flow calling it, a test for each value. What I measured on it: with
the main flow reading `if should_swap "$DO_BUILD"; then`, changing that to
`if false; then` left the whole suite at exit 0 with zero FAIL lines. The
swap still never ran. Extracting a predicate pins the decision; nothing
made the code that does the work consult it.

Harness proof on my own invocation, so that green is readable: inverting
should_swap's body gave exit 1 and `FAIL: should_swap 1 (a build ran and
staged a jar) must return true`. The suite can fail when I run it.

My fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally, so there is no guard left in the
main flow to get wrong. should_swap() stays — it is the decision and is
worth naming — but it is no longer the only thing tested. Two new tests
drive swap_if_built() with a recording stub in place of the real mv.

Two more things I fixed, both found while verifying:

* The ordering test had to follow the call site to `swap_if_built
  "$DO_BUILD"`. Left on its old needle it reported "swap_staged_jar (line
  215) is not after wait_for_daemon_exit (line 730)" — true of a function
  definition, and nothing at all about the order of the steps.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail`, an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output
  and no FAIL line. Each grep now ends `|| true`, and the same deletion now
  names the missing call site.

Verified by me on the merged revision:

* suite exit 0, 0 `^FAIL:` lines, 44 test functions defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical by hash, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call site"
* CI green on 08771e2 (run 1775)
* main has not touched either file since the branch point, so this is not
  a stale-branch merge

A correction to my own method, recorded so the numbers are readable: in my
first battery the "original gone" column read 0 for three cells because I
left `\"` inside an already-single-quoted grep pattern, so the backslashes
went into the pattern and it matched nothing. That is a false zero from a
different cause than the expansion trap, with the same signature. Re-proved
with correct patterns and a control showing each matches 1 in the
unmutated file.

Not fixed here, filed as #528: drain_gate_refusal has the identical shape.
Replacing `die "$(drain_gate_refusal ...)"` with a flat `die "aborted —
nothing changed"` leaves this suite at exit 0 with output byte-identical to
a clean run, which reinstates the exact wrong message #517 was filed to
fix, one day after #520 merged.
2026-09-12 07:01:55 +02:00
ltms 01adc841fa Merge #523: test the policy probe's parsing guards (fleetd #519)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 1m32s
Adjudicated and verified by me, not taken from the PR body.

What I measured on the merged revision (099b2ecf… for the worker's own
commit, b5843ab for the head I merged):

* 5 test functions defined, 5 invoked; suite exit 0, "PASS: probe member
  credentials guards"
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, with a green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count: missing policy parser (jq) returned 0 field(s)
  - arity threshold 5 -> 0 -> FAIL: short parser output status

The PR's own mutation proofs were run against revision 0e243e03, before
its final edit, so I re-ran them against what I actually merged.

Two things I fixed at the gate rather than sending back:

* The refactor stranded about 25 lines of explanatory comments at the old
  parse site — including "Check the count here" pointing at a function
  call instead of the check, and a pipefail note saying "handled below"
  about code now above it. That is the same wrong-stated-fact defect class
  as fleetd #500, in the very file whose ticket history is about it. Moved
  each block above the code it explains.
* Two assertions matched on `jq) returned N field(s)`, a needle starting
  mid-parenthetical, so a real failure printed "missing jq) returned 0
  field(s)" and read as if the script's message had an unbalanced paren.
  Widened to `policy parser (jq) returned N field(s)`, which also pins
  that the refusal names the parser it used.

Caveats recorded, from the implementer and not re-checked by me: the
harness does not cover the non-member, missing-parser, curl-fetch,
known-count, or hash-tool fallback paths.

One property worth noting in favour of this suite: it runs under `set -e`,
so the first failing test aborts before the final `printf 'PASS: …'`. That
PASS line is reachable only from the fully successful path.
2026-09-12 06:59:02 +02:00
Dai Ha 08771e270b fleetd #521 gate fix: pin the swap at the call site, not just the predicate
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m11s
The extraction in the previous commit did what #521 asked for — a
should_swap() predicate with a test for each value — and I measured that
it does not close the defect. With the main flow reading
`if should_swap "$DO_BUILD"; then`, changing that line to `if false; then`
left the whole suite at exit 0 with zero FAIL lines. The swap still never
ran, and a redeploy would still report success while starting on no jar.

That is my ticket's fault, not the implementer's: "extract the decision so
the suite can call it" pins the decision and never the wiring. Extraction
moved the untested decision up one level instead of removing it.

Fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally — there is no guard left in the
main flow to get wrong. should_swap() stays, because it is the decision
and is worth naming and testing on its own. Two new tests call
swap_if_built() with a recording stub in place of the real mv, so they
fail if the guard is removed, inverted, or stops being consulted.

Also fixed, found while verifying this:

* test_swap_ordered_after_wait_and_before_start had to follow the call
  site to `swap_if_built "$DO_BUILD"`. Left on the old needle it reported
  "swap_staged_jar (line 215) is not after wait_for_daemon_exit (line
  730)" — true of a function definition, and nothing about step order.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail` an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output,
  no FAIL line, nothing naming what was missing. Each grep now ends in
  `|| true` so the assignment succeeds empty and the guard can speak.

Verified by me on this revision:

* suite exit 0, 0 `^FAIL:` lines, 44 tests defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call
    site in redeploy-fleetd.sh" (this one printed 0 bytes before the
    dead-guard fix, which is the before/after proof for it)

Not fixed here, filed separately: drain_gate_refusal has the same shape.
Replacing `die "$(drain_gate_refusal ...)"` with `die "aborted — nothing
changed"` leaves the suite at exit 0 with output byte-identical to a clean
run, which reinstates the exact wrong message #517 was filed to fix.
2026-09-12 11:58:34 +07:00
Dai Ha b5843ab43f fleetd #519 review fix: widen two needles to the whole parenthetical
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m5s
Both arity assertions matched on `jq) returned N field(s)` — a needle
that starts in the middle of the script's `(parser name)` parenthetical.
On a real failure the harness prints `missing <needle>`, so the line
came out as:

  FAIL: empty parser output count: missing jq) returned 0 field(s)

which reads as if the script's own message had an unbalanced paren. It
does not; the needle was just sliced. Matching on
`policy parser (jq) returned N field(s)` makes the failure readable and
also pins that the refusal names the parser it used, which the narrower
needle did not.

make_jq() PATH-prefixes a fake jq, so `_PARSER_NAME` is deterministically
"jq" in both tests; the wider needle cannot flake on a host without jq.

Re-proved on this revision, because a disproof is about a revision and
not a file:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 on the test under /bin/bash 3.2.57 and bash 5.3.9
* dropping the empty-parse special case -> FAIL: empty parser output
  count: missing policy parser (jq) returned 0 field(s)
* arity threshold 5 -> 0 -> FAIL: short parser output status
* script restored byte-identical after each, green control after both
2026-09-12 11:50:11 +07:00
Dai Ha d8985719eb fleetd #525: restore the logger level, not just the appender, in SessionManagerTest
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
The finally block in onTurnFailedIsLoggedAtWarnWithThePriorState (and four other
tests in this file) called setLevel(Level.WARN) on the shared SessionManager
logger (one test used MemberRegistry.class) but only detached the appender in
finally, never restoring the level. Since logback Logger instances are cached
per class and shared across the whole JVM, the pinned level leaked into every
test that ran after it.

Add CapturedLog, a small AutoCloseable that captures a logger's level and
appender together and restores both on close via try-with-resources, so this
shape cannot be half-fixed again. Convert all 7 addAppender/setLevel call
sites in this file to it, including the 2 sites #522 already fixed locally
(kept per-test pinning as belt-and-braces).

Add a proving test (sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn,
@Order(2), running right after the fixed test at @Order(1)) that fails before
this fix and passes after it. Verified with a mutation: dropping the level
restore in CapturedLog.close() turns the proving test red with
'expected: <DEBUG> but was: <WARN>'; restoring is byte-identical to the
pre-mutation file (sha256 matched) and the suite goes green again.

mvn -f fleetd/pom.xml clean install: Tests run: 1699, Failures: 0, Errors: 0,
Skipped: 0. BUILD SUCCESS.
2026-09-12 11:49:03 +07:00
Dai Ha 5b1e13ca3d fleetd #519 review fix: move the parser comments with the code they explain
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m2s
PR #523 extracted parse_policy_fields() but left about 25 lines of
explanatory comments at the old parse site in main(). That is the same
defect class as fleetd #500 itself — a stated fact that no longer
matches the code next to it — in the very file whose ticket history is
about it.

Three blocks moved, no code touched:

* "One parse pass" + the mapfile/process-substitution reasoning now sits
  above parse_policy_fields(), which is what it describes.
* The arity-check block now sits inside the function, directly above
  `if (( ${#_FIELDS[@]} < 5 ))`. At the old site it said "the slice just
  below this" and "every line below this expects", both pointing at a
  function call rather than the check. Reworded to name main() and its
  slice explicitly.
* The pipefail note said the parser failure was "handled below"; the
  handling is now above it, in the function.

The call site keeps a three-line pointer saying where the reasoning went.

Checked myself, on this revision:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count
  - arity threshold 5 -> 0 -> FAIL: short parser output status
2026-09-12 11:47:16 +07:00
Dai Ha c89a375e5d fleetd #521: extract should_swap so the swap guard can't be silently disabled
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m29s
Mutating the swap step's guard (if [ "$DO_BUILD" = 1 ] -> if false) left the
whole test suite green: test_swap_ordered_after_wait_and_before_start only
checks source positions, which an in-place if-condition edit never moves.
Extracts the decision into should_swap(do_build), following the same shape as
#510's wait_for_daemon_exit and #517's drain_gate_refusal, with a direct test
for each value.
2026-09-12 11:42:46 +07:00
ltms 37dcefa834 Merge #522: log a positive drain-completion line when drainAll finishes (fleetd #512 part 1)
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m36s
Verified by the lead, not taken from the worker's report.

Build run by me, unpiped: exit 0, "Tests run: 1698, Failures: 0, Errors: 0, Skipped: 0",
BUILD SUCCESS. That is +2 on main's 1696, matching the two new tests.

Both merge constraints from the ticket hold, checked against the diff:
- The log.info is the last statement of drainAll's normal path and is NOT in a finally. The only
  finally in the file is at :372, unrelated. So a drain that dies still leaves no line, which is
  the absence signal part 2 will alert on.
- released and abandoned are incremented inside the drainSnapshot loop, not read from a collection
  at the end, so a partial report can never be served by this line.

Mutation check run by me: deleting the completion line makes both new tests fail by name —
drainAllLogsACompletionLineWithTheRealCountsOnACleanDrain and
drainAllLogsANonZeroAbandonedCountForASessionStillBusyAtTheDeadline, "no drain-complete INFO
logged, expected true but was false". Restored byte-identical
(d21ecd3adb3f66933e2f248a8ec81a2c324bfee323a3cc30da525842c33ea80a). So the tests really depend on
the line rather than passing for another reason.

Design accepted: drainSnapshot returns a private DrainTally record and release() returns the
removed session so abandoned can reuse the same BUSY-at-removal check logPreservedForShutdown
makes, instead of a second registry read. Both are private with one call site.

Two follow-ups filed rather than fixed here, see the ticket comments.
2026-09-12 06:34:51 +02:00
Dai Ha 6a7342b1f0 fleetd #518: make the FleetMcp caller-resolution wiring an explicit choice, and test it once for real
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 2m12s
FleetMcp's contextExtractor picked its principal-resolution path off `callers == null`, so
"authorization off" also silently swapped in a second, untested identity heuristic
(legacyPrincipal). Nothing drove that closure through a real MCP request, so the whole wiring
was an unexercised claim.

- callers (CallerResolver) is now required, never null.
- A new AuthorizationMode enum (ENFORCED/UNENFORCED) is a required constructor parameter with
  no default, replacing the null-means-legacy idiom for whether denyFor enforces at all.
- legacyPrincipal is deleted: there is exactly one resolution path now
  (callers.resolve(...)), so the mutation that swapped it for an unconditional legacy call no
  longer compiles ("cannot find symbol: method legacyPrincipal").
- FleetMcpContextExtractorTest boots the real transport on a real Jetty server and drives it
  with a real MCP client, proving fleet_whoami's resolved role comes from CallerResolver's
  token check.
- Adapted FleetMcpAuthzTest/FleetMcpHandoverTest call sites; theLegacyConstructorLeavesTheGateOpen
  keeps its meaning under the new AuthorizationMode.UNENFORCED value.
2026-09-12 11:34:12 +07:00
Dai Ha a5ad7c6561 fleetd #519: test policy probe guards
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m51s
2026-09-12 11:30:35 +07:00
ltms c71ac231e5 Merge #520: pin the drain-gate abort branch and jar_id's absent case (fleetd #517)
CI / contract (push) Successful in 1m28s
CI / build (push) Successful in 1m50s
Verified by the lead, not taken from the worker's report:

- 40 test functions defined, 40 invoked (my own greps), suite exit 0, zero real FAIL lines.
- Harness proof on my own invocation: re-applying the jar_id absent->present mutation gives
  exit 1 and "FAIL: jar_id with no arguments must report absent when $JAR does not exist".
  So a green run from this suite is readable.
- Script restored byte-identical after every mutation:
  2cb83dc380c7226191d657c40fccdfc856904e40b2d03d0851fb6e522cee2f41.
- drain_gate_refusal is pure and prints only; die() stays outside the command substitution, so
  the "die inside $( ) exits only the subshell" trap does not apply here.
- The --no-build + staged-present case reads "nothing changed" deliberately, and that is correct:
  --no-build stages nothing itself, and a leftover staged jar is wiped by rm -f "$JAR_STAGED"
  at :607, before the build at :610.

The worker's out-of-scope finding is real and I reproduced it: the swap guard at :723 can be set
to `if false` with the suite still green at exit 0. Filed separately.
2026-09-12 06:29:08 +02:00
Dai Ha 33720c42b3 fleetd #512 (part 1): log a positive completion line when drainAll finishes
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 1m38s
drainAll used to log nothing on a clean drain — both existing log calls
(drainSnapshot's per-session failure, drainAll's straggler-sweep warning)
sit on abnormal paths, so "drained fine" and "died on the first session"
looked identical: no log line either way.

Add one log.info at the end of drainAll: "drain complete: released=N
abandoned=M (still BUSY at the shutdown deadline)". It fires on the
normal path, including the all-zero case, and folds both drainSnapshot
passes (main snapshot + straggler sweep) into one line.

drainSnapshot now returns a private DrainTally(released, abandoned)
record instead of void, and the private release(paneId, cause) overload
now returns the removed MemberSession (previously void) so drainSnapshot
can read its state at the moment of removal — the same check
logPreservedForShutdown already makes. Both signature changes are
private with a single call site, so the blast radius stays small.
2026-09-12 11:29:01 +07:00
Dai Ha 3833d8e52b fleetd #517: pin the drain-gate abort branch and jar_id's absent case
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m31s
Two mutation-testing survivors in scripts/redeploy-fleetd.sh: a source-text
test pins what a message SAYS but never whether the branch that prints it is
REACHED.

- Extract the drain-gate abort decision into drain_gate_refusal(do_build,
  staged_path), a pure function the suite can call directly for all four
  build/staged combinations. The existing source-text grep test is kept
  alongside it (it catches a re-wording; the new tests catch a dead branch).
- Extend jar_id's test to cover the missing-file path (both the no-argument
  default and an explicit path), which the #511 test never exercised.
2026-09-12 11:23:37 +07:00
ltms b37def9238 Merge #516: the probe refuses with three distinct messages, each naming its own cause (fleetd #500)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m46s
2026-09-12 06:09:50 +02:00
ltms 8f02576df6 Merge #515: pin the two-client completeness fold, and legacyPrincipal earns no authority (fleetd #509)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m54s
2026-09-12 06:06:12 +02:00
ltms 525bc1c5f4 Merge #514: the drain-gate abort message names a recovery that works, and jar_id()'s default is pinned (fleetd #511)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m41s
2026-09-12 06:00:23 +02:00
Dai Ha d59ece6dec fleetd #500: stop a wrong-interpreter or failed-parse reading a policy as empty
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m8s
probe-member-credentials.sh used mapfile < <(producer) to parse the fetched policy. That
hides a producer failure three ways: mapfile is bash 4+ and missing on macOS's /bin/bash
3.2, a process substitution's exit status is never propagated to mapfile, and the
downstream reads (":-" defaults and a slice) never fire set -u on a short or unset array.
All three converge on the same "0 known names" refusal, which blames the policy for a
failure that is actually the interpreter or the parser.

Three distinct guards, each closing one cause with its own message:
- a BASH_VERSINFO gate at the top refuses outright on bash < 4 (exit 3)
- the parser's output is captured via command substitution instead of mapfile < <(...),
  so a non-zero jq/python3 exit is caught at the call while the fact still exists (exit 4)
- an arity check before the field slice refuses a parse that exits 0 but returns fewer
  than 5 fields (exit 5)

The existing "0 known names" guard is now honest: by the time it fires, the three causes
above are already ruled out, so it really does mean the policy has 0 known names.
2026-09-12 10:58:16 +07:00
Dai Ha 32408d1e64 fleetd #509: pin the pane-scan completeness fold, and stop legacyPrincipal handing out primary
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Successful in 1m40s
Unit 1 — PaneLocator.terminalForPid's completeness fold across herdr
clients (PaneLocator.java:117) had no test that varied the number of
clients, so a mutation that keeps only the last client's Lookup.complete()
instead of ANDing every client's outcome survived: 14 of 15 existing tests
agree with the mutant on a single client. Added a two-client test where
the lead client errors on the pane that would have owned the pid (an
incomplete, negative scan) and the member client cleanly finds no panes
(a complete, negative scan) — the real fold ANDs these to false, a
last-wins fold reads it as true. Proved against MUTANTC
(complete = outcome.complete();): the new test fails with
"expected: <false> but was: <true>", the file was restored byte-identical
(sha256 unchanged), and the control run is green.

Unit 2 — FleetMcp.legacyPrincipal's else-branch returned Principal.primary
for ANY caller the connection did not resolve to a worker pane, with none
of CallerResolver.java:254's isLoopback/scanComplete guards. Measured that
no production caller passes null callers (Fleetd.java:696 always
constructs a real CallerResolver) but FleetMcpAuthzTest.mcp(false)
legitimately does, for its "legacy constructor leaves the gate open" test
— so the null-callers path is not dead code to delete (option a), it is a
documented legacy mode (option b). Changed the else-branch to
Principal.anonymous() and widened legacyPrincipal to package-private (like
denyFor) so a new test pins the behavior directly, since it only ever ran
inside a contextExtractor closure no existing test triggers.
2026-09-12 10:58:09 +07:00
Dai Ha 6e23bf8309 fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
The drain-gate abort message told the operator a rerun "with or without
--no-build" would finish the restart. That is wrong: by the time this
message can fire, stage_built_jar has already moved the jar off $JAR, so
--no-build hits require_no_build_jar's own refusal. Reworded to say the
rerun must NOT use --no-build, and why: the built jar is no longer at the
live path that --no-build requires.

Also added a test pinning jar_id()'s no-argument default (reports $JAR,
the live path) and its explicit-argument behavior (reports that path
instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm
-f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn
clean install deletes target/ on the next line regardless).
2026-09-12 10:56:32 +07:00
ltms aa4c0b84c3 Merge #510: never build into the path a running daemon holds (fleetd #493)
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m4s
2026-09-12 05:44:16 +02:00
ltms 40c593cd09 Merge #508: a herdr error during the pane scan is refused, not promoted to primary (fleetd #505)
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m38s
Closes the worker->primary escalation that fleetd #317 left open on the other side of the same
ternary. #317 stopped an UNRESOLVABLE pid being promoted. This stops a RESOLVED pid whose pane scan
failed being promoted: the scan now reports completeness, and an incomplete scan resolves to
anonymous.

Shape chosen: PaneLocator.terminalForPid returns Lookup(terminal, complete); paneOwnsAnyOf returns a
private Ownership enum OWNS/DOES_NOT_OWN/UNKNOWN, so a HerdrException is UNKNOWN rather than a clean
negative. ConnectionIdentity.Caller gains scanComplete as a third field -- resolved() was NOT
widened, correctly: it is documented as testing the lsof sentinel and #505 is a different axis. A
definite match still short-circuits, so a genuinely vanished non-owning pane does not become a
refusal.

Verified by me, not taken from the report.

Trial-merged onto main (136312f) and built the merge:
  mvn clean install -> Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS

Harness proof, by re-running the worker's own mutation rather than trusting it (UNKNOWN ->
DOES_NOT_OWN): Tests run: 63, Failures: 3 -- exactly the three tests meant to pin it, matching their
report line for line:
  CallerResolverTest.aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary
  ConnectionIdentityTest.scanIsIncompleteWhenHerdrErrorsOnThePaneThatOwnsThePid
  PaneLocatorTest.anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative
So the cells can go red, and their numbers are honest. Restored, shasum 81b797a6... byte-identical,
control 63/0.

MY OWN MUTATION, on the half they did not touch, FOUND A GAP. I mutated the two-client completeness
fold in terminalForPid -- `complete = complete && outcome.complete()` -> `complete =
outcome.complete()`, which forgets an earlier client's failure. Two greps proved it applied. Result:
Tests run: 63, Failures: 0. NOT pinned.

It is not an equivalent mutant: with CB-185's two herdr daemons, a first client that scans
incompletely followed by a second that scans cleanly gives `false` originally and `true` mutated, so
an incomplete scan would be reported complete and promoted. It is the #393 shape instead -- no test
assembles that combination, because the single-daemon tests collapse lead and member to one object.
That axis has cost us before: four routing defects passed the suite when one client served both roles
with memberHerdrSocket unset.

Merging anyway, because the production behaviour is correct and the ticket's defect is pinned three
ways -- what is missing is a test for a dimension the PR description claims in prose. Filed as a
follow-up with the numbers above.

Second item for the follow-up, latent not live: FleetMcp.legacyPrincipal still does
`c.terminal() != null ? worker : primary` and drops scanComplete, so it would re-open exactly this
hole. Measured unreachable in production -- Fleetd.java:696 always passes a real CallerResolver, and
the contextExtractor only calls legacyPrincipal when `callers == null`. A defect on paper is not a
reachable defect, so it is not a blocker; it is a trap for whoever next touches that constructor.

That second item is also the fleet01 lead's structural point, and it is the sharper framing: #415's
antidote (make the decision total) is ORTHOGONAL to this defect. Authz.permits is already a
default-less switch over Action and cannot help, because it says nothing about whether the caller was
resolved correctly, and Principal carries no record of how it was resolved. #415 guards a decision
nobody wrote; this was a decision written correctly and fed a bad input. The durable fix direction is
that an unresolved scan should be unrepresentable as a principal rather than checked for at each
consumer.
2026-09-12 05:35:41 +02:00
Dai Ha 979adf82eb fleetd #493: never build into the path a running daemon holds
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m37s
redeploy-fleetd.sh's build step wrote straight into fleetd/target/fleetd.jar
via `mvn clean install` while the OLD daemon was still running from that
exact path. A JVM loads classes lazily, so a class the daemon had not
touched yet could be read from a jar already replaced or removed -- the
failure landed on the shutdown drain (NoClassDefFoundError, exit 143,
looks clean).

Stage the freshly built jar at target/fleetd-new.jar (stage_built_jar),
confirm the OLD pid has actually exited (wait_for_daemon_exit, extracted
from the existing wait loop), and only then swap it into the live path
(swap_staged_jar) -- strictly after the wait, strictly before start. A
failed swap dies without starting. --no-build and --check keep their
existing, truthful behavior (require_no_build_jar; jar_id still reads
the live path by default). A leftover staged jar from an interrupted
run is wiped before the next build. The build still runs before
anything is stopped, so a failed build still never takes the fleet down.

Also fixed: the drain-gate abort message ("aborted -- nothing changed")
now names the staged jar when one exists, since staging already moves
the freshly built jar off the live path before that prompt runs.

Adds unit tests for stage_built_jar, swap_staged_jar, require_no_build_jar,
wait_for_daemon_exit, and a source-order test proving swap sits after the
wait and before start (sourcing stops before the main flow ever runs, so
the ordering itself can only be checked by reading the script's own call
sites).
2026-09-12 10:35:19 +07:00
Dai Ha 36870836aa fleetd #505: a herdr error during the pane scan must not read as a clean negative
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m29s
A transient herdr error on pane.process_info during PaneLocator's pid→pane scan used
to be swallowed into a plain "does not own it", so a real worker whose owning pane
errored mid-scan resolved with a null terminal but a resolved (real) pid — exactly
what CallerResolver's loopback-trust fallback reads as the primary. That is a
worker→primary privilege escalation through the door fleetd #317 did not close: #317
guards a failed lsof lookup (c.resolved()), not a failed herdr pane scan.

Fix: add a third state to the scan instead of widening Caller.resolved() (which stays
centralised next to the lsof sentinel it tests, per #505's explicit instruction not
to reopen that decision). PaneLocator.terminalForPid now returns a Lookup(terminal,
complete) record: a HerdrException on one pane marks that pane's ownership UNKNOWN,
not DOES_NOT_OWN, and the scan is complete only if every pane was either matched or
confirmed not to own the pid. A definite match found elsewhere in the same scan
still short-circuits as complete — a pane that genuinely vanished mid-scan without
being the caller's own does not turn into a refusal.

ConnectionIdentity.Caller carries the new scanComplete flag alongside the unchanged
resolved(). CallerResolver's loopback-trust fallback now requires both resolved()
and scanComplete() before promoting to Principal.primary(); an incomplete scan
resolves anonymous, which fails toward the recoverable error (a refused primary
retries loudly; a promoted worker would not).

Logs a warning naming the pane and which herdr client (of how many) failed, so the
incomplete-scan path is diagnosable rather than silent (fleetd #317's own lesson).
2026-09-12 10:27:42 +07:00
ltms 136312fb11 Merge #503: Injector's readiness-grace warn prints measured elapsed time, never arithmetic on constants (fleetd #501)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m27s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (4f28da6) and built the merge, because a clean auto-merge is not a compiling
merge:
  mvn clean install -> Tests run: 1689, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  InjectorTest -> Tests run: 34, Failures: 0
  Gitea CI on ac351ee: success

The worker mutated the two log arguments. I mutated the half it did not: the STAMP SITE. Changing
the first-sample guard so the clock is re-stamped on every non-ready poll gave 1 failure,
InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants, on a clock-read
counter assertion: "the readiness-not-ready branch should read the clock exactly twice — once to
stamp the first non-ready sample, once at grace expiry". Two greps proved the mutant applied;
restored, shasum byte-identical, control 34/0 green.

Behaviour preserved: the restructure from `else if (p != null && ++t.notReadySincePoll >= N)` to a
nested `else if (p != null) { ... }` keeps the short-circuit, so the counter still increments only
when p != null. All three reset sites now clear notReadySinceMillis alongside notReadySincePoll.

Two notes for the record, neither a blocker.

The poll-count half is an equivalent mutant and the worker said so instead of reporting a kill it
did not get. That is the right call and the code comment states the limit honestly.

The clock is injected through a package-private constructor overload while the public constructors
default it to System::currentTimeMillis. My brief prescribed that shape. With exactly one production
construction site (Fleetd.java:531) a required parameter — fleetd #415's antidote — would have been
just as cheap and would match LeadRollover and the new Fleetd.awaitHerdr. The silent-survivor risk
#415 names is not present here, because the field is final and both public constructors delegate, so
a future constructor cannot compile without supplying it. Recording the choice so the next person
does not read it as an oversight.
2026-09-12 05:07:35 +02:00
ltms 4f28da62a3 Merge #499: redeploy-fleetd.sh gains an "unclear" supervisor state that cannot reach kill, and the detail survives the subshell (fleetd #492, #497)
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m32s
Two commits, b17f37a and 599419f. Verified by the lead, not taken from the worker's report.

b17f37a adds the fifth answer "unclear" for installed-but-not-loaded and for a probe that could not
answer at all, and makes require_drivable_supervisor die on it.

599419f fixes a defect I found in b17f37a: SUPERVISOR_UNCLEAR_DETAIL was set inside
detect_supervisor, which is called as $(detect_supervisor). A subshell only returns stdout, so the
detail never reached the caller, and a ${VAR:-generic} fallback then printed a message naming no
supervisor and no reason. Measured before the fix: plain call -> detail length 142; via $( ) ->
length 0.

Lead verification of the fix, re-running my own original measurement across every branch:

  launchd installed, not loaded   kind=[unclear]   detail_len=142
  systemd installed, not active   kind=[unclear]   detail_len=146
  launchd loaded (normal)         kind=[launchd]   detail_len=0
  systemd loaded (normal)         kind=[systemd]   detail_len=0
  neither -> none                 kind=[none]      detail_len=0
  both loaded -> ambiguous        kind=[ambiguous] detail_len=0

The separator is always emitted, so the ${RAW#*SEP} unpack cannot silently fall back to the whole
string on the four branches that carry no detail.

Harness proof, run by me: making detect_supervisor print the kind alone (two greps proved the mutant
applied) turned the suite red, exit 1, "FAIL: detail does not name the systemd unit it found
installed-but-not-loaded". Restored, shasum byte-identical, control exit 0, bash -n clean on both
files. 23 test functions defined, 23 invoked.

The three case "$SUPERVISOR_KIND" switches now have final arms, chosen by what each caller does with
the value: the reporting switch warns and continues (a diagnostic that aborts goes silent exactly
when the state is novel), while stop and start die.

Not fixed here, filed separately: the "was already not running" branches at :603 and :610 swallow a
failed launchctl unload / systemctl stop with || true and then print "ok" unconditionally.
2026-09-12 05:04:04 +02:00
ltms 708f1795ad Merge #502: awaitHerdr reports three outcomes with measured elapsed time, not one boolean (fleetd #498)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m49s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (8f59019) locally and built the merge, because a clean auto-merge is not a
compiling merge:
  mvn clean install -> Tests run: 1687, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  FleetdAwaitHerdrTest -> Tests run: 6, Failures: 0
  Gitea CI on 274afaf: success

The worker mutated the three return values. I mutated the half it did not: the reap DECISION at the
call site. Changing logHerdrWaitOutcomeAndShouldReap to return true for INTERRUPTED gave 1 failure,
FleetdAwaitHerdrTest.interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed:135 ("an interrupted
wait must not tell main to reap"). Two greps proved the mutant applied; restored, shasum
byte-identical, control 6/0 green.

Behaviour preserved: herdrUp is still true only for ANSWERED, and both its readers (the orphan reap
at :276 and the lead auto-launch gate at :375) see exactly what they saw before.

One correction to the worker's report, which changes nothing in the code. It wrote that a real
interrupt-detection regression "would also hang the daemon's startup thread forever". It would not:
in production `nanos` is System::nanoTime, so the deadline check still fires. The hang it hit was a
test-fixture property — a frozen injected clock with no iteration bound. That is fleetd #486's
shape, now seen in a second class, and it is recorded there.
2026-09-12 05:01:02 +02:00
Dai Ha 599419f9e6 fleetd #492 follow-up: carry the unclear detail across detect_supervisor's subshell boundary
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m53s
b17f37a set SUPERVISOR_UNCLEAR_DETAIL as a global inside detect_supervisor, but the real call
site invokes it as $(detect_supervisor) — a subshell — so that global died with the subshell and
the die() message's ${VAR:-fallback} silently masked the loss with generic text.

- detect_supervisor now packs kind and detail onto its one stdout line (joined by the ASCII unit
  separator byte, $SUPERVISOR_DETAIL_SEP), the only channel that survives $( ). The real call site
  unpacks both with in-shell parameter expansion — no extra subshell.
- Dropped the ${SUPERVISOR_UNCLEAR_DETAIL:-...} fallback at the die() message: under set -u, a
  missing detail now fails loudly instead of silently defaulting (same defect class as #497).
- Added a constraints comment block above detect_supervisor for future callers: stdout-only,
  no ${VAR:-default} papering over a lost value, and every case on the return value needs an
  explicit *) arm.
- Added *) arms to the three `case "$SUPERVISOR_KIND"` switches (report/stop/start): report warns
  and continues (display-only), stop/start die naming the value (they act on it).
- Rewrote the "unclear" test to go through the real call-site shape ($(detect_supervisor) then
  the same split), not a hand-constructed value, and tightened its final assertion to check for
  the actual detail text rather than $SYSTEMD_UNIT alone (the die() boilerplate names the unit
  either way, so that check could pass on a lost value).
2026-09-12 09:58:45 +07:00
Dai Ha ac351ee1de fleetd #501: readiness-grace expiry logs measured elapsed time and the loop's own poll counter, never the configured budget
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m2s
The line at Injector.java:383-387 printed two numbers that read as measurements
but were both compile-time constants: READINESS_GRACE_POLLS for the poll count
(the loop's own Target.notReadySincePoll counter was in scope at the same call
site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed
time — arithmetic on two constants, never a measurement, and wrong in the
direction that says everything ran on schedule.

Fix, copying the LongSupplier-clock shape LeadRollover already uses:
- print t.notReadySincePoll instead of the constant for the poll count (the two
  agree by construction on this branch, so no test can tell them apart — the
  comment says so honestly).
- add Target.notReadySinceMillis, stamped at the first non-ready sample and
  reset at all three sites notReadySincePoll already resets (:304, :355, :389
  pre-fix line numbers), to compute a real elapsed time at expiry.
- inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis)
  through new package-private constructor overloads so a test can supply a
  clock whose advance does not track POLL_INTERVAL_MILLIS.

Tests use a ListAppender to assert on the log message contents, per
LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub
clock returning two literal, non-derived values so it can fail if the fix
regresses to the constant-arithmetic line — proved by mutation: reverting the
elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that
one test red (1 failure); reverting the poll-count print to the constant is an
equivalent mutant (0 failures), because the counter and the constant are
identical at that exact call site by construction.

fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0.
2026-09-12 09:57:57 +07:00
Dai Ha 274afafde6 fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m51s
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare
  boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are
  two distinct, named states instead of the same false (fleetd #497's shape).
- awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required
  parameters, with no defaulted overload (fleetd #415), so a test can drive it.
- The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself
  cannot be driven from a unit test; it logs a distinct message per outcome, always printing the
  measured elapsed time next to the configured budget, never the budget alone.
- Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt
  flag) and the call site (the three distinct log messages), using ListAppender.
2026-09-12 09:54:59 +07:00
ltms 8f59019305 Merge #496: lead-rollover logs measured elapsed time and measured counts, never the configured budget (fleetd #494)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 2m1s
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report.

3fb3311 — the /clear-settle and success paths return measured elapsed time and a real nudge count
c87cc25 — print the measured nudge count; fix the same defect on the sibling turn-settle timeout line
e966cba — the poll count on the same line was still a constant; the test could not tell the difference

Lead verification at e966cba:
  mvn clean install -> Tests run: 1681, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  LeadRolloverTest  -> Tests run: 33, Failures: 0, Errors: 0, Skipped: 0
  Gitea CI on e966cba: success

Mutation M2 (PICKUP_GRACE_POLLS 8 -> 5), run by the lead: 2 failures,
clearGraceReleaseLogIsWarnWithMeasuredElapsed and successLogPrintsMeasuredElapsedForTheWholeRoll.
The emitted line under the mutant read "after 5 consecutive IDLE/DONE polls (4 of those were
nudged)", so both numbers now follow the loop. Restored, shasum byte-identical, control 33/0 green.

Known and deliberate limit, stated in the code comment: on the grace-release branch both counters
equal the constants by construction (one write site each), so no test can prove the difference on
that line. The change buys one source of truth, not provable coverage. The place nudges genuinely
varies with the run is the /clear-timeout warn, which is covered by a discriminating test.
2026-09-12 04:44:10 +02:00
Dai Ha e966cbadf9 fleetd #494 follow-up (2nd pass): the grace-release line still had one constant, and its test could not tell the difference
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m55s
Two more fixes on the same line, LeadRollover.java:600-624:

1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
   loop's own idlePollsAwaitingPickup counter instead of the constant. Same
   defect shape as the nudges fix from c87cc25, one argument over.

2. LeadRolloverTest's clearGraceReleaseLogIsWarnWithMeasuredElapsed asserted
   its nudge-count expectation as `LeadRollover.PICKUP_GRACE_POLLS - 1` — the
   same expression the production code used to build the log line from, so it
   could not discriminate a reverted fix. Rewritten to plain literals
   ("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"), proven to
   trip when PICKUP_GRACE_POLLS's value changes (2 failures under a
   PICKUP_GRACE_POLLS=5 mutation: this test and successLogPrintsMeasured...).

Also corrected the comment above the log line: idlePollsAwaitingPickup and
nudges each have exactly one write site on this loop's release branch, so they
cannot differ from PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 at this call
site — confirmed by re-deriving the loop's control flow, and by reverting just
the nudges argument (M1) and observing 33/0 stayed green even after the test
rewrite. That is an equivalent mutant on this line, not a gap the test rewrite
could close; the honest value of printing the counters is one source of truth
for the loop, not a provable-by-test difference here. The place nudges truly
varies with the run — and is covered by a test that can tell it apart from a
constant — is the /clear-timeout warn's clearResult.nudges() in runRollover.

mvn -Dtest=LeadRolloverTest test: 33/0. mvn clean install: 1681/0, BUILD SUCCESS.

Same-shape sweep (found, not fixed, per instructions):
Injector.java:383-387 — the readiness-grace warn prints the constant
READINESS_GRACE_POLLS where the measured per-target counter
t.notReadySincePoll is in scope (single increment site, just reached the
threshold at this call site — same equivalent-mutant situation as this fix).
2026-09-12 09:39:39 +07:00
Dai Ha b17f37a683 fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m32s
Two situations were silently landing in the "none" answer, which
require_drivable_supervisor accepts and the script then falls back to a raw
kill + nohup — exactly the wrong move when a supervisor actually IS present:

- installed-but-not-loaded, on either supervisor. `systemctl --user is-active`
  answers "no" for activating/deactivating/failed and while an auto-restart is
  pending too, and every one of those is a host that IS under systemd (or
  launchd) and about to act again. `*_installed` already knew this; it was
  only ever consulted for a warning line, never by the decision itself.
- a systemd probe that could not answer at all (e.g. systemctl cannot reach
  the user bus over a non-lingering ssh session) looked identical to a clean
  negative, because both probes redirected stderr straight to /dev/null.

detect_supervisor now returns a fifth answer, "unclear", for both cases.
systemd_loaded/systemd_installed capture systemctl's exit status and stderr
separately and set their own *_ERRORED flag only on a real tool failure
(non-zero exit WITH stderr), never on a clean negative. "none" now means only:
neither supervisor installed, neither loaded, neither probe errored.
require_drivable_supervisor die()s on "unclear" exactly like it already does
on "ambiguous", naming the specific supervisor and reason via the new
SUPERVISOR_UNCLEAR_DETAIL global.

Tests: 4 new cases (systemd/launchd installed-but-not-loaded, a real
systemd_loaded run through a systemctl stub that errors on stderr, and the
die() refusal for "unclear" naming the unit). All 3 new guards were verified
by mutation: each was removed from the real script, the suite caught it (a
new FAIL line naming the exact broken assertion), then the file was restored
byte-identically and the suite went green again.
2026-09-12 09:23:21 +07:00
Dai Ha c87cc25aa6 fleetd #494 follow-up: print the measured nudge count, and fix the sibling turn-settle timeout line
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m30s
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter
  instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the
  constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8
  nudges where 7 went out) and a code read did not catch it — only a mutation test did.
  Printing the counter cannot drift from the loop's real behaviour.
- waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value
  printed as if measured' defect as the three lines fixed in the original #494 commit, but was
  out of scope because the brief named specific lines instead of the shape. Fixed the same way:
  it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the
  timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds()
  as the measured wait.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured
  nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is
  updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a
  new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line.
- Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via
  two greps that the mutant applied and the original constant was gone, ran the grace-release
  test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8
  and reran the full LeadRolloverTest suite as a control (33/33 green).
2026-09-12 09:22:19 +07:00
Dai Ha 3fb331145a fleetd #494: log measured elapsed time, never the configured budget, on a lead-rollover failure/success
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m35s
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
  instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
  prints the measured elapsed time next to the target pane — this is the exact path that
  reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
  whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
  instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
  self-advancing fake clock so the measured elapsed/nudge values are deterministic and
  provably distinct from the configured budget.
2026-09-12 09:08:04 +07:00
Dai Ha dcd505286f fleetd #492: teach redeploy-fleetd.sh systemd --user as a third supervisor
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
launchd, systemd --user, and unsupervised are three different answers, not two.
Refuse (die) rather than fall through to kill+nohup when a supervisor is
detected that this script cannot drive (e.g. both signals fire at once), and
add a post-restart check that fails the run if more than one fleetd process
is alive. detect_supervisor()/require_drivable_supervisor()/
count_daemon_pids()/assert_single_daemon() are pure, overridable functions so
scripts/test-redeploy-fleetd.sh can exercise them without a real launchd or
systemd.
2026-09-12 09:00:59 +07:00
Dai Ha 60fa958da2 docs: a relative handoverPath lands in the LEAD's repo, not fleetd's (#491)
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore
protects the handover file and the distinction is invisible. On fleet01 the
lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory,
and that repo has no .handover rule — measured 2026-09-12.

Tell the lead to check its own workspace's .gitignore before writing, and to
report it rather than committing the file or editing someone's .gitignore.

Refs fleetd #491, #487, #480.
2026-09-12 08:46:17 +07:00
Dai Ha 9950361bc9 docs: #489 is fixed and deployed, but criterion 6 is still unmet
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m33s
The skill said the roll 'does nothing until #489 is merged and redeployed'.
Both have happened, so that sentence now reads as a green light. It is not
one: no roll has bootstrapped a fresh session end to end yet. Say that
plainly, and tell the lead to warn the operator before confirming.

Refs fleetd #489, #480.
2026-09-12 07:55:41 +07:00
ltms 9f74b6619a Merge #490: nudge the /clear submit keystroke before bootstrapText (fleetd #489)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m30s
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".

Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
  including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on f687046: success

A separate mutation of the FIRST gate hung the suite instead of failing it.
That is fleetd #486, not a regression here; the reproduction is recorded there.
2026-09-12 02:52:52 +02:00
Dai Ha f687046450 fleetd #489 follow-up: fix stale class javadoc, off-by-one nudge count, weak test
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m33s
Three review corrections on top of the previous commit:

1. The class javadoc's four-step continuation list (lines 47-58) was stale.
   Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
   own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
   described the old plain re-check ("the original, pre-correction wait...
   still here") - fixed to describe what waitForClearPickupAndSettle
   actually does: nudge while unpicked-up, then wait for a real WORKING ->
   IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.

2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
   polls, not the number of nudges - the 8th poll releases instead of
   nudging again, so 8 polls produce 7 nudges. The log.info in the release
   branch and two javadoc spots said "8 nudges"; fixed all three to state
   the poll count and the nudge count separately and correctly. Behavior
   and the constant are unchanged.

3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
   and sendKeysCallCount, both of which a return-true stub also satisfies.
   Added an assertion on the already-tracked postClearGetCalls counter
   (>= 2), which only a real post-/clear poll loop can produce - this is
   what makes the test fail against a return-true mutant.
2026-09-12 07:45:36 +07:00
Dai Ha c6058652be fleetd #489: nudge the /clear submit keystroke before bootstrapText
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m36s
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled
for IDLE/DONE, which /clear itself never leaves since it starts no real turn,
so it always returned true on the first poll. Combined with a direct
agents.send bypassing Injector (deliberate, to avoid wedging the pane), the
submit Enter that accompanies /clear could race the paste and leave it
unsubmitted — bootstrapText then landed concatenated onto the same input
line, exactly as measured live on 2026-09-12.

Replace that second wait with waitForClearPickupAndSettle, which copies the
pickup-nudge pattern Injector already ships for its own post-turn /clear
housekeeping (fleetd #306): nudge agents.submit while the pane hasn't
reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than
wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is
observed. BLOCKED stays excluded from both the nudge and the boundary check,
same as the (unchanged) first wait — a paused live turn is not settled, and
nudging Enter into an open prompt could wrongly answer it.

Adds four tests to LeadRolloverTest covering the paste-race regression
(nudge ordered between /clear and bootstrapText), a confirmed pickup, a
deadline expiry with no boundary ever reached, and a throwing submit().
2026-09-12 07:29:26 +07:00
Dai Ha 2a95da2ff4 docs: the handover skill must warn that the automatic roll is broken (#489)
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 1m53s
Section 11 said the bootstrap prompt landing was 'not yet proven
end-to-end'. It is now measured failing: the first real roll joined
/clear and bootstrapText into one line. Nothing was cleared, so the
failure is safe, but a lead that reads the old wording would reach for
fleet_handover expecting it to work.

Refs fleetd #489, #480.
2026-09-12 07:21:38 +07:00
Dai Ha 7f9137fcb6 fleetd #480: ignore the handover file, and note the absolute-path guarantee in the handover skill
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m52s
The lead rollover handover file now lives inside the workspace, at the relative
path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one
moment's live state, so it must never enter git history.

The handover skill also now says the handoverPath fleetd hands back is always
absolute, even when the configured value is relative — a lead that resolves it
itself can pick a different file from the one the daemon checks.
2026-09-12 05:46:34 +07:00
Dai Ha 3bf3968bc7 Merge #487: resolve a relative leadRollover.handoverPath against the calling lead's workspace (fleetd #480 follow-up) 2026-09-12 05:43:12 +07:00
Dai Ha 261aa056f9 fleetd #480 follow-up correction 2: guarantee resolveHandoverPath is always absolute
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m32s
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.

Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.

Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.

Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.

Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
2026-09-12 05:41:40 +07:00
Dai Ha 042b8c99dd fleetd #480 follow-up correction: cover Fleetd.leadRollover(...)'s own wiring behaviourally
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m51s
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...)
factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live
one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative
handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live
lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at
construction time (a lead discovered by the tab scan after leadRollover(...) was built still
resolves correctly).

Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always
"no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full
1669-test suite green before this commit. With the new test added, the same one-line mutation now
fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during
verification and is not part of this commit (git diff on Fleetd.java is empty).

FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test
could catch this wiring dropping out, which was true only before this commit and only covered the
factory's own body, not its call site. Restated what each test class actually covers: the source-
text pin covers the call site's argument list; the new behavioural test covers the lambda's body.
2026-09-12 05:32:25 +07:00
Dai Ha 4bfab6b718 fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m4s
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the
calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that
lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores
only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed
back to the lead in the fleet_handover open response, and the default bootstrapText sentence
all see the same absolute location instead of a value resolved against whatever directory the
daemon process happened to start in.

FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it
would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor
(resolvedHandoverPath) method builds the default sentence from the resolved path instead.

Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to
lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on
every call, never off a startup snapshot.
2026-09-12 05:19:49 +07:00
Dai Ha 008a457557 docs: fleet_handover is shipped — update the charter table and the handover skill
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m38s
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to
catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the
canonical block is an incomplete change.

CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover.
Both copies edited identically; the byte-identical sync check prints True.
The row carries the trap rather than just the call: open FIRST, then write the
file, then confirm — because confirm refuses the file as stale unless its
modified time is later than the open request, so the obvious order fails.

.claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover
tool: it does not exist." That was true when it was written this morning and is
now false, which is the worst state for an instruction file to be in. Replaced
with the two real paths (by hand, or with the tool when leadRollover: is
configured), plus a new section 11 giving the three-step order and the five
things that surprise a caller — chiefly that `accepted` does not mean the pane
has been cleared, and that operatorConfirmed is a report of what a human said,
not a confidence level.

Section 11 also states plainly that the bootstrap prompt landing in a freshly
cleared pane is not yet proven end-to-end, and that the recovery is the manual
path — which is why the file is written before confirm, never after.
2026-09-11 07:29:59 +07:00
ltms 9494a6b99a Merge #485: fleetd #480 Unit C — the fleet_handover MCP tool
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m58s
fleet_handover{action: "open"|"confirm"|"cancel"} drives LeadRollover, which #483 and
#484 landed with nothing calling it. Primary-only via a new Authz.Action.HANDOVER, on
the same case line as SPAWN/STOP/DRAIN.

The tool has NO terminal, session or leadTerminal parameter of any kind — the pane is
always callerTerminal(exchange), resolved from the connection. A lead can therefore only
ever roll itself, never another lead. That is charter invariant 3, and it is the second
of the two corrections recorded in LeadRollover's class javadoc.

Registered unconditionally, so the tool surface does not vary with config: with
leadRollover: absent, every action returns a clean NOT_CONFIGURED refusal instead of
failing, and open()'s IllegalStateException (config removed by a hot reload after
construction) is caught and turned into the same refusal. A config-dependent tool set
would have collided with #474's charter tool-surface gate and McpContractDocTest.

Correction round applied before merge, and it is the reason this took two passes.
The unit first shipped with a defaulted 15-argument FleetMcp constructor delegating to
the new 16-argument one with leadRollover = null. I mutated the wiring rather than
reasoning about it: deleting just the leadRollover argument from Fleetd.main's FleetMcp
call compiled with 0 errors and passed all 1659 tests, BUILD SUCCESS — while the live
daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
Neither FleetMcpHandoverTest (it builds its own FleetMcp) nor FleetdLeadRolloverWiringTest
(it pins that LeadRollover is constructed, not that it is passed on) could see it.

The worker then found the defect was wider than I had named: all five shorter
constructors (11/12/13/14/15-arg) formed one defaulting chain into the 16-arg one, each
silently supplying another feature's "off" value — leadChannel, outage, leadSeats, peers,
and finally leadRollover. All five are deleted. FleetMcp now has exactly one public
constructor, so every one of those features is compile-enforced at its call site, not
just this one.

Verified by the lead before merge, on PR head merged with current main (0176378):
- CI run 1728 green on eb05576.
- The identical mutation re-run by me: control against the original -> 1; mutant present
  (MUT-DROP-ARG2) -> 1; original gone -> 0; `mvn -o -q compile` now FAILS —
  "constructor FleetMcp ... cannot be applied to given types; reason: actual and formal
  argument lists differ in length" at Fleetd.java:[698,24]. The antidote holds: required
  parameter = compile error, defaulted overload = silent survivor a green suite vouches for.
- Restored clean (0 changed files), then mvn clean install on the merged tree:
  BUILD SUCCESS 1, BUILD FAILURE 0, Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.
- Three call sites of `new FleetMcp(` measured, all accounted for: Fleetd.java,
  FleetMcpAuthzTest.java, FleetMcpHandoverTest.java.
2026-09-11 02:27:56 +02:00
Dai Ha eb0557621e fleetd #480 correction round: collapse FleetMcp to one required constructor
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 1m32s
FleetMcp had a defaulted 15-argument constructor that delegated to the new
16-argument one with an implicit null for leadRollover. Dropping the
leadRollover argument from Fleetd.main's FleetMcp(...) call fell back to that
shorter overload, compiled fine, and left all 1659 tests green — the live
daemon would then answer NOT_CONFIGURED to fleet_handover forever with
nothing going red.

Delete every overload that could reach the 16-arg constructor with a
silently-defaulted leadRollover (11/12/13/14/15-arg forms all chained to it),
leaving the 16-arg constructor as FleetMcp's sole public constructor. Update
FleetMcpAuthzTest's call site to pass every parameter explicitly (leadChannel
null, OutageSource.none(), LeadSeatSource.none(), List.of(), leadRollover
null) — Fleetd.java and FleetMcpHandoverTest already called the full form.
Proved with mvn -o -q compile: removing the leadRollover argument from
Fleetd.main now fails to compile instead of silently defaulting.

No behaviour changes — NOT_CONFIGURED refusals are unchanged.
2026-09-11 07:25:08 +07:00
ltms a0eed6f01b Merge #484: fleetd #480 Unit E — a BLOCKED pane is not a settled pane
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m35s
LeadRollover's two settle waits used AgentStatus#injectable(), which is
IDLE || BLOCKED || DONE. That is the right rule for Injector ("may I deliver a message
without stepping on a live turn") and the wrong one here ("has the turn actually
ended"), because BLOCKED is a live turn that is merely paused — a pane sitting on an
approval prompt.

Path in: the lead calls confirm(); its turn carries on and hits anything needing
approval; herdr reports blocked; within 250ms the deferred continuation reads that as
settled; /clear is typed into an open prompt. That destroys the lead's live context
mid-turn, which is exactly what the turnSettleSeconds gate added in #483 exists to
prevent. The window is turnSettleSeconds, default 20s.

Both waits now require a real turn boundary — IDLE or DONE. AgentStatus#injectable() is
untouched: it is correct for Injector, LeadHeartbeatLoop, LeadCoordLoop, ReplyPushLoop
and HerdrPeerLauncher's two readiness checks, all of which are delivery gates.

Verified by the lead before merge:
- CI run 1726 green on head e2a91e8.
- Mutation on the half the worker did NOT mutate: dropped the DONE arm, leaving
  `if (status == AgentStatus.IDLE)`. Mutant proven applied with two unrelated proofs
  using different search strings (MUT-DROP-DONE present -> 1; 'IDLE || status' gone
  -> 0). Same single test, same 100s cap: unmutated passes in 0.172s; mutated is still
  running at 100s and does not pass. So doneStatusStillCompletesTheFullRoll really does
  pin the DONE arm, and the fix did not over-tighten to IDLE-only.

Why the earlier #483 battery could not have caught this: it proved the guard FIRES, not
that the guard's predicate was tight enough. LeadRolloverTest drove the fake status with
only "idle" and "working"; neither value discriminates this defect, only "blocked" does.

Correction round applied before merge: both warning lines said "never went idle" and
"did not become injectable". Retiring that word is the whole point of the change, so a
future reader debugging from those lines would have been handed back the wrong mental
model. Both now say "did not reach a turn boundary (IDLE or DONE)".

Filed separately as #486, deliberately not fixed here: with the DONE arm removed the
test hangs rather than fails, because waitUntilAtTurnBoundary is bounded only by an
injected clock that the tests hold fixed. Not a production bug — nowMillis is
System::currentTimeMillis there — but it turns a future regression into a CI hang
instead of a red test.
2026-09-11 02:21:06 +02:00
Dai Ha e2a91e883e fleetd #480 Unit E correction: retire "injectable" wording from the log lines
CI / build (pull_request) Successful in 1m28s
CI / contract (pull_request) Successful in 1m27s
Both waitUntilAtTurnBoundary guard messages still said "never went idle" /
"did not become injectable" — the old mental model the rename was meant to
retire. Made both say what the code now actually waits for: a turn boundary
(IDLE or DONE).
2026-09-11 07:18:51 +07:00
Dai Ha 62646957ea fleetd #480 Unit C: wire fleet_handover MCP tool onto LeadRollover
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m56s
Adds the fleet_handover tool (open/confirm/cancel) as a thin adapter over
LeadRollover, registered unconditionally so the charter tool-surface gate
sees a stable set regardless of whether leadRollover: is configured. With a
null LeadRollover every action degrades to a clean NOT_CONFIGURED refusal
instead of throwing. Gated on a new Authz.Action.HANDOVER (primary-only,
same as SPAWN/STOP/DRAIN). The caller's own connection-resolved terminal is
the only lead identity ever used — the tool's input schema carries no
terminal/session/leadTerminal parameter, so a lead can only ever roll
itself. Fleetd.main now passes its existing leadRollover local into FleetMcp
via a new trailing constructor parameter.
2026-09-11 07:07:24 +07:00
Dai Ha 802c0ab701 fleetd #480 Unit E: BLOCKED is not a settled turn boundary
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m55s
LeadRollover's waitUntilInjectable used AgentStatus#injectable(), which
accepts BLOCKED. A BLOCKED pane is paused mid-turn on a prompt, not
settled — reusing injectable() let /clear (or the bootstrap text after
it) fire into an open approval prompt within the 20s settle window,
destroying the lead's live context.

Renamed the helper to waitUntilAtTurnBoundary and restricted both waits
to IDLE or DONE only, with a comment explaining why this class does not
reuse injectable() (it answers "may I deliver", not "has the turn
ended"). Added tests for BLOCKED-forever on both waits (zero sends /
exactly one send) and for DONE still completing the full roll.
2026-09-11 07:01:28 +07:00
ltms 4bab23e241 Merge #483: fleetd #480 Unit A — lead rollover core (config + executor)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
Two corrections were applied to the original unit before this merge, and the PR
description above still describes the pre-correction shape:

1. confirm() no longer rolls inline. It validates every gate, then hands a one-shot
   continuation to continuationRunner and returns. confirm() is called BY the lead FROM
   its own turn, so the pane is WORKING and cannot report injectable until confirm()
   returns; the old code sent /clear first and then timed out waiting, which destroyed
   the lead's context and started no fresh session. The continuation waits for the
   calling turn to settle FIRST (turnSettleSeconds, a new knob), and if that wait times
   out it sends no /clear at all.
2. The pane to roll comes from the caller's terminal, not PrimaryRegistry. A single-slot
   lookup let lead X's confirm() clear lead Y's pane. open() records the caller's
   terminal; confirm() refuses with NOT_YOUR_ROLLOVER on a mismatch.

Verified by the lead before merge:
- CI run 1722 green on head a942622.
- Mutation battery on the two safety-critical branches, in a scratch worktree at a942622.
  Controls measured against the ORIGINAL first; each mutant proven applied with two
  unrelated proofs using different search strings.
  M1 `if (!turnSettled)` -> `if (false)`: turnThatNeverSettlesSendsNoClearAtAll FAILS
     (LeadRolloverTest:247, expected 0 sends but was 1).
  M2 ownership check -> `if (false)`: aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover
     FAILS (LeadRolloverTest:266). Unmutated harness-proof run: exit 0 with both guard
     WARN lines logged, so the branches really are exercised.

Nothing calls LeadRollover yet. The fleet_handover MCP tool is Unit C.
2026-09-11 01:51:20 +02:00
Dai Ha a94262271b fleetd #480 correction round: defer the roll, and gate it on caller identity
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m3s
Two defects found after the fact, both from the original brief, both fixed here.

1. confirm() is called FROM the calling lead's own turn, so its pane is still
   WORKING and can never report injectable inside that same call. The old
   confirm() sent /clear before polling for that — the poll always timed out,
   but only after /clear had already fired and queued, destroying the lead's
   context with no fresh session ever started and a refusal return that lied
   about what had happened.

   Fix: confirm() now only validates and, if every gate passes, hands a
   one-shot continuation to a new continuationRunner (a real virtual thread in
   production, Runnable::run in tests) and returns RollDecision.approved()
   immediately - "scheduled", not "rolled". The continuation itself does the
   actual work, once the calling turn has ended: wait for the SAME pane to
   report injectable again (new turnSettleSeconds config key, default 20) -
   if this never happens, /clear is NEVER sent, at all - then /clear, then
   wait again (clearSettleSeconds, as before), then bootstrapText. The "no
   timer/scheduler, only confirm() can roll" invariant is restated precisely
   in LeadRollover's class javadoc: it is about initiative, not synchronicity
   - a single-shot continuation of an already-approved confirm() call still
   satisfies it; a recurring background loop would not.

2. confirm() resolved the pane to clear via PrimaryRegistry.primaryTerminal(),
   a single-slot lookup that is correct for a background loop with no caller
   but wrong here: on a daemon with more than one labelled lead tab, lead X's
   confirm() could clear lead Y's pane, violating the charter's "identity
   comes from the connection, never an argument" invariant.

   Fix: open() and confirm() now take the caller's terminal id as a parameter
   (resolved by the MCP layer from the connection - the later MCP-tool unit
   must pass it in, never accept it as a request field). confirm() refuses
   with a new NOT_YOUR_ROLLOVER reason unless it matches the terminal open()
   recorded. LeadRollover no longer depends on PrimaryRegistry at all.

Also: renamed RollResult to RollDecision (rolled -> accepted) to reflect the
new meaning - approved and scheduled, not necessarily cleared yet. Added
turnSettleSeconds to the leadRollover: config block (documented in
fleetd.example.yaml alongside the existing keys) and updated Fleetd.java's
leadRollover(...) factory to drop the primaryRegistry parameter, with
FleetdLeadRolloverWiringTest's source-text pin updated to match.

New tests: turnThatNeverSettlesSendsNoClearAtAll (the branch that matters
most - a turn that never ends means /clear is never sent) and
aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover (NOT_YOUR_ROLLOVER),
plus a settle-after-clear timeout test and an open() input-validation test.
LeadRolloverTest: 11 -> 14 tests.
2026-09-11 06:44:58 +07:00
Dai Ha 5c12865c25 fleetd #480 Unit A: lead rollover core (config block + executor)
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m58s
Adds the opt-in leadRollover: config block and LeadRollover, the executor a
later unit's MCP tool will call. A lead writes a handover file, then open()
records a token and confirm() verifies it (exists, non-empty, fresh) and an
operator confirmation before clearing the lead's own pane via /clear (sent
directly through AgentControl, bypassing Injector, same as
ClaudeCodeLauncher#clearContext) and bootstrapping a fresh session. Nothing
but an explicit confirm() call can ever roll a pane - no timer, no heartbeat,
no background thread anywhere in this class.

Wired into Fleetd.java exactly like LeadHeartbeatLoop: constructed only when
leadRollover: is present at startup, and nothing calls it yet - the MCP tool
is a separate, later unit.

Classified leadRollover: as HOT in ConfigRef (joins placement/
memberCredentials/memberLoginShell/models): the executor holds
Supplier<FleetConfig.LeadRollover> and reads every field fresh per call,
unlike LeadHeartbeatLoop's frozen final fields. The one caveat: the object's
construction is still gated on presence in the startup config snapshot, so a
freshly-added block needs a restart before anything exists to call.

Tests: LeadRolloverTest (14 cases covering the 6 hard requirements - no
object without the config block, only confirm() can roll, missing/empty/
stale handover file each refuse by name, requireOperatorConfirm gating, and
the injected wall-clock supplier) and FleetdLeadRolloverWiringTest (source-
text pin on Fleetd.main's construction call, mirroring
FleetdCompletionResolverWiringTest). Also updated the existing
FleetConfigValidateAllTest, FleetConfigWithDefaultsPreservesEveryComponentTest,
ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest to
account for the new record component.
2026-09-11 06:33:01 +07:00
Dai Ha 3f8c956325 docs: the handover skill must not promise a tool that does not exist
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m30s
The skill shipped in #481 described fleetd #480's automation in the present
tense: "fleetd ticket #480 lets a lead session hand off to a fresh one". The
tool is not built. A session loading the skill would look for a fleet_handover
tool it cannot call, and this repo already knows what a confidently wrong
instruction costs — a wrong comment stops the reader investigating with a false
conclusion, which is worse than no comment.

Now says plainly: today the handoff is manual, #480 will automate the same
cycle, and do not call fleet_handover because it does not exist. The nine
procedure rules are unchanged — only who performs the swap changes when #480
ships, not what the file must contain.

My wording to fix, not the worker's: the brief handed them the present tense.
2026-09-11 06:17:11 +07:00
ltms 2051ceeb26 Merge #481: the handover skill (fleetd #480 Unit B)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m27s
Adds .claude/skills/handover/SKILL.md — the outgoing lead's procedure for writing
the file a fresh lead session inherits.

Reviewed by the lead against the eight requirements in the brief: every number
carries its command, measured-vs-reported is marked, open decisions name their
owner, live hazards, what is not owed, a first section of three re-measurement
commands, a time-and-commit stamp, and a plain statement that the file goes
stale. All eight are present, plus a leave-out section and a template.

Verified by the lead, not taken from the worker's report:
- CLAUDE.md diff is the addendum's skill list only, in the primary-side group.
- The canonical block is byte-identical with the wiki template (sync check True).
- CI run 1718 green on 325d0771, the same commit reviewed.

Follow-up owed by the lead: the skill describes #480's automation in the present
tense, but the tool does not exist yet. The lead will soften that wording so the
skill does not promise a fleet_handover tool a session cannot call.
2026-09-11 01:16:40 +02:00
Dai Ha 325d0771a4 fleetd #480 Unit B: add handover skill for lead session handoff
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m54s
Adds .claude/skills/handover/SKILL.md, the procedure an outgoing lead
follows to write the handover file a fresh lead session inherits when
fleetd clears the pane. Registers the new skill in CLAUDE.md's
primary-side skills list; no other change to CLAUDE.md.
2026-09-11 06:13:49 +07:00
Dai Ha 7b97aae85b docs: move the redeploy procedure out of CLAUDE.md into a skill
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m34s
CLAUDE.md loads into every session. The daemon-redeploy procedure is
needed only when someone redeploys, so it paid for context it did not
use: 58 lines, about 918 est. tokens, every session.

The procedure now lives in .claude/skills/redeploy-fleetd/SKILL.md,
which loads only when invoked. The moved text is byte-identical to
what was removed, plus one new paragraph documenting --no-build
(scripts/redeploy-fleetd.sh:34,62,266) — the script and the wiki
already had that flag, CLAUDE.md never did.

CLAUDE.md keeps an 11-line pointer, because two rules must stay
resident: a merge is not a deployment, and workers must never
redeploy. A session learns it needs the procedure before it needs
the skill, then the next line names the skill.

Also updates the addendum's primary-side skill list, as this file's
own rule for .claude/skills/** changes requires.

CLAUDE.md: 34,442 -> 31,591 chars.
Canonical block untouched — the wiki sync check still prints True.
2026-09-11 05:48:06 +07:00
Dai Ha 49a404ddf3 Merge #474 follow-up: pin main's ConfigRef wiring against the surviving mutation
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m37s
My battery on the #474 merge found one survivor: reverting Fleetd.java:154
from the three-argument ConfigRef constructor to the plain two-argument
one turns the live reload gate off and leaves all 1633 tests green. Both
new #474 tests build their own ConfigRef with the method reference, so
neither reads what main chose.

This adds FleetdConfigRefWiringTest, following the three source-text
precedents already in the tree (FleetdBackendQuarantineWiringTest,
FleetdLeadSeatWiringTest, FleetdCompletionResolverWiringTest) rather
than the weaker sibling pattern that builds the wiring itself. No
production change.

The worker branched fresh off 4466ee0 rather than continuing its old
branch off the stranded 435e022 base. That was its own call and it was
the right one: one merge base, a clean 80-line diff, and no cherry-pick
needed this time.

Its vacuity guard is worth keeping in mind for the next source-text
test: it asserts the file it read contains 'public final class Fleetd'
before asserting anything about the mutation, with a message saying the
assertFalse below would pass vacuously on a broken read. An unrelated
anchor is the right choice there, because a guard that shares the
mutation's text cannot tell a bad read from a real change.
2026-09-10 20:50:13 +07:00
Dai Ha 72d6a6878b fleetd #474 follow-up: pin main's config wiring against M2
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m32s
Fleetd.main's own choice of the three-argument ConfigRef constructor
(with Fleetd::assertChartersNameOnlyRegisteredTools as extraValidation)
was unpinned. Reverting Fleetd.java:154 to the plain two-argument
constructor compiled with 0 errors and left the whole suite green,
because ConfigRefTest and FleetdConfigRefCharterToolSurfaceWiringTest
each build their own ConfigRef directly rather than through main.

Adds FleetdConfigRefWiringTest, a source-text check on Fleetd.java
following the FleetdBackendQuarantineWiringTest/FleetdLeadSeatWiringTest/
FleetdCompletionResolverWiringTest precedent: asserts the exact
three-argument construction is present, asserts the plain two-argument
form is absent, and guards against a vacuous pass on a broken/empty
source read by first asserting an unrelated anchor is present.
2026-09-10 20:48:05 +07:00
Dai Ha 4466ee0ef2 fleetd #474: ConfigRef.reload() runs the charter tool-surface gate too
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m44s
A charter naming an MCP tool the server does not register refused Fleetd.main
at startup but slipped through ConfigRef.reload(), because reload() only ran
FleetConfig.validateAll(), which never looks at what a charter's text names.

CharterToolSurface stays in the mcp package (config must not depend on it), so
ConfigRef now accepts the check as a Consumer<FleetConfig> extraValidation,
run inside reload()'s same try/catch as validateAll(). Fleetd.main wires a new
package-private adapter, Fleetd.assertChartersNameOnlyRegisteredTools, into
both the startup call site and ConfigRef's constructor, so the two call sites
can never check different things.

Tests: ConfigRefTest (reload refuses/accepts, via a locally-built equivalent
consumer since Fleetd's method is package-private to dev.ltms.fleet) and the
new FleetdConfigRefCharterToolSurfaceWiringTest (same proof through the exact
Fleetd::assertChartersNameOnlyRegisteredTools reference production uses).
Verified deleting the new extraValidation.accept(fresh) call site fails both
new "refuses" tests by name.

(cherry picked from commit 97e4c1d658)

Lead review. Cherry-picked, not merged: the worker branched from 435e022,
which is not an ancestor of main (I had reset and rewritten that commit's
message while the worker was already on it). Merging its branch would have
added a second merge commit for work already on main. Same content, no
duplicated history. Rule learned: once a worker is spawned, its base
commit is published.

My own mutation battery, five cells, each a full `mvn -B clean test` on
this commit's own tree:

  CONTROL 1, unmutated       1633 tests, 0 failures, BUILD SUCCESS
  M1 delete accept(fresh)    KILLED  2 by name
  M3 3-arg ctor stores no-op KILLED  2 by name
  M4 set(fresh) before valid KILLED  6
  M5 accept outside try      KILLED  2 errors
  M2 main uses the 2-arg ctor SURVIVED

M2 is a real gap and it is a test gap, not a defect: the production code
here is correct. Reverting Fleetd.java:154 to `new ConfigRef(configPath,
cfg)` turns the live reload gate off and leaves all 1633 tests green,
because both new tests build their own ConfigRef with the method
reference rather than reading what main wires.

That is the fourth Fleetd.main call site to survive a battery (#446 M5
and M7, #466 M1, now this). Extracting a check into a well-tested helper
moves the untested surface UP, into the line that chooses to call it. The
repo already has the answer in three source-text wiring tests
(FleetdBackendQuarantineWiringTest, FleetdLeadSeatWiringTest,
FleetdCompletionResolverWiringTest); the worker followed the other
sibling pattern, which builds the wiring itself and so cannot pin main.
Delegated as a follow-up, with the surviving mutation as its acceptance
criterion.

Also mine to correct: my M3 cell's annotation said the no-op store "must
be 2" and measured 1. The 2-arg constructor delegates rather than
assigning, so my pattern only ever matched the mutant. The cell still
stands on its other proof (requireNonNull left = 0). Third battery in a
row with a wrong CONTROL annotation of my own.
2026-09-10 20:40:49 +07:00
Dai Ha 25ba7f16bb Merge #473: fleet_profiles and fleet_list report which attempt a quarantine is on (fleetd #466 item 2)
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m24s
The escalating cooldown landed in 789b6a8, but the only thing an operator
could see was quarantinedForSeconds. A long number does not say whether
this is the first exhaustion or the fifth, and after escalation those look
the same from outside: 3600 seconds could be a big base cooldown or a
credential on its fourth strike.

Both reports now carry quarantineAttempt beside quarantinedForSeconds.

The part that matters is HOW, not that the field exists.
BackendQuarantine#status(credentialId) does ONE quarantines.get() and
returns both values from the same QuarantineState. It is not two accessors
that a caller pairs up. That is the pattern
CompositePeerLauncher.modelGateState() already documents: a report that
reads a different source from the behaviour it describes will eventually
disagree with it, and the disagreement is invisible because both halves
look right on their own.

This is reporting only. Escalation, the ceiling and the quiet-gap reset are
untouched. The flat two-argument constructor still reports a real growing
attempt count even though its cooldown stays flat - which is the honest
answer, since the streak is real whether or not the cooldown uses it.

The worker's report was lost: its fleet_reply never arrived and its inbox
was empty. Nothing was lost with it, because the brief required pushing
first and putting the report in the PR body. That is the second time that
practice has saved a turn.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:21:14 +07:00
Dai Ha 5ba69c9cf2 fleetd #393: correct a comment that claimed two tests pin a call they cannot
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m51s
The comment above the charter writer said flipping that call back to
putArray was "proven load-bearing" and pointed at two named tests. I
mutated exactly that, on the merge commit, and it SURVIVED at 1618 green -
so neither named test covers it, and neither can.

The reason is structural, not a missing test: that writer runs first
against an empty array, so putArray has nothing to replace and the two
idioms are equivalent there. No test can distinguish them. The worker
measured the same thing independently and said so in their report; the
comment was left over from the pre-fix state, where the mutation being
described was a DIFFERENT one (a later writer destroying the charter).

A comment that names tests which do not cover the line is worse than no
comment. The next person mutates the line, sees green, and concludes the
tests are broken.

The corrected version states what each cell actually measures: this line
is unpinnable and why, the later writers ARE pinned and by which test
names, and deleting this line entirely fails the charter-reaches-the-member
assertion - the hole that predates #393.
2026-09-10 20:17:21 +07:00
Dai Ha 17052bb515 Merge #471: seeded skills reach an opencode member, and instructions[] stops depending on write order (fleetd #393)
memberSkills: copied skill folders into every provisioned worktree's
.claude/skills/ and stopped there. Claude Code reads that directory
natively; opencode never does. So the feature was INERT for opencode
members rather than broken: the copy succeeded, the files were correct,
and nothing ever read them. No test failed because there was nothing to
fail - the feature worked at the only layer it implemented.

An opencode member's only channel for static guidance is the instructions[]
array in its generated config. OpenCodeLauncher now adds each seeded
skill's SKILL.md there. A folder with no SKILL.md is never delivered and
the log names it.

The second half is an ordering hazard fleet01 found by reading, and that I
then measured. Three writers append to instructions[]: the role charter,
the seeded skills, and the IDE rules. The charter used putArray
(CREATE-OR-REPLACE) while the other two used withArray (get-or-create).
That was safe only because the charter ran first against an empty array -
an undeclared constraint that nothing tested. Measured on the earlier
merge: making the skills writer use putArray left 1603 tests green while
silently deleting the charter entry, so an opencode member would launch
with no role contract at all. Worse than the bug being fixed, and
invisible.

All three writers now use withArray, and three tests pin the array's
CONTENTS (never its size - a size assertion passes when putArray swaps two
entries for two others) across the combinations that matter:
charter-only, charter+ide, charter+skills+ide.

WHY NOTHING CAUGHT IT, MEASURED RATHER THAN ASSUMED. fleet01 first said
the missing axis was the COMBINATION of writers, then revised that to a
stronger claim: that nothing asserted the charter reaches an opencode
member at all. I checked the second claim on this merge and it is FALSE.
Deleting the charter writer outright fails 6 tests, and 3 of those existed
before #393: writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet,
roleCharterWithoutMcpOrCustomProviderStillWritesAConfig, and
aPinnedEndpointAndTheFleetMcpCoexistInOneConfig. The charter reaching a
member was already pinned.

So their FIRST diagnosis was the right one. The surviving mutation did not
delete the charter write; it made a LATER writer replace the whole array.
Every pre-existing test had exactly one writer active, and with one writer
putArray and withArray are indistinguishable. The gap was never "is the
charter delivered" - it was "are two writers ever active at once", which is
the combination axis. Recording this because the stronger claim is the more
quotable one and it would have sent the next reader looking for a hole that
is not there.

One honest residue: the charter writer's own idiom cannot be pinned.
Flipping it back to putArray leaves the suite green, and always will,
because it runs first against an empty array where the two idioms are
equivalent. The edit removes an undeclared constraint for the next person
to add a writer; it is not a change any test can detect. The comment in
the source claimed two named tests cover it - that was wrong, and the
commit after this one corrects it rather than leaving a false claim beside
the code.

What this does NOT do: opencode has no equivalent of Claude Code's Skill
tool, so the content is static system-prompt text present from spawn, not
something a member can invoke by name. This closes the DELIVERY gap and
cannot close the ACTIVATION gap. That residue is opencode's design.

Numbers and my own mutation battery are on the ticket, measured on this
merge commit.
2026-09-10 20:17:21 +07:00
Dai Ha e95ed99bf7 fleetd #466 scope item 2: report the quarantine repeat count, not only the seconds
CI / build (pull_request) Successful in 1m31s
CI / contract (pull_request) Successful in 1m33s
Add BackendQuarantine#status(credentialId) -> Optional<Status>, a single
QuarantineState read that answers both remainingSeconds and repeatCount
together -- the same "one accessor" pattern CompositePeerLauncher.
modelGateState() already uses, so the two facts can never disagree.

fleet_profiles/REST GET /profiles and fleet_list's capacity rows
(FleetMcp.profilesView/capacityView) now call status() instead of
remainingSeconds() and add a "quarantineAttempt" field beside
"quarantinedForSeconds": 1 for a first occurrence, 2 for the second
in a row, and so on.

No change to the escalation, ceiling, or reset logic itself -- this
unit is reporting only.
2026-09-10 20:09:13 +07:00
Dai Ha 1477e4358a Merge #472: one canonical tool-name set, and charters are checked against it (fleetd #469)
CI / contract (push) Successful in 57s
CI / build (push) Successful in 2m0s
A role charter is free text in config that tells a member which tools to
call, and nothing checked that those tools exist. A charter naming
bridge_send - a name CB-634 removed - started the daemon cleanly, and the
member found out at run time by calling something that was not there.

#464 shipped a test for this, but it wrote its own charter into a @TempDir,
so nothing anyone put in the real config could fail it. That was a defect
in my acceptance criteria, not in that work.

Now: FleetTool is one enum of the 11 registered wire names, and every
reader goes through it.

- FleetMcp's schema builders pass FleetTool.X.wireName() instead of a
  literal.
- FleetMcp's constructor asserts at startup that what it registers with the
  SDK equals FleetTool.wireNames() exactly, in both directions.
- toolAction(String, Map) resolves arbitrary wire input against
  FleetTool.byWireName() and keeps its run-time throw, which is necessary -
  network input has no closed compile-time form. It then hands off to
  authzAction(FleetTool, Map), a switch over the enum with NO default, so a
  new tool is a compile error at that layer.
- CharterToolSurface lives in mcp, not config, and Fleetd.main calls it
  right after validateAll(). Config must not depend on the MCP server:
  config loads before the server exists.

My brief undercounted the problem and the worker corrected it. I said there
were two tool-name inventories plus a test fixture. There were FIVE: the
registrations, the authz switch, and three separate source-text scrapes of
FleetMcp.java in CharterToolSurfaceTest, FleetMcpAuthzTest and
McpContractDocTest - none of which the ticket mentioned. Fixing those three
was required, not scope creep: once the literals moved into FleetTool their
regexes matched zero names, so one would have failed on its vacuity guard
and the other two would have gone quietly vacuous. Inventories after: one.

Verified independently on origin/main before accepting the wider diff:
three test files did read FleetMcp.java as source text, with a control file
at zero to prove the search discriminated.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:59:03 +07:00
Dai Ha 9e4e423ad6 fleetd #393 follow-up: remove the instructions[] writer-ordering hazard
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 2m11s
OpenCodeLauncher.writeConfig has three writers into the instructions[]
array (charter, seeded skills, IDE rules). The charter writer used
putArray (create-or-REPLACE) instead of withArray (get-or-create), which
"worked" only because it happened to run first against a still-empty
array — an undeclared ordering dependency nothing tested. Found by the
fleet01 lead and verified on this branch's merge: flipping the skills
writer to putArray left the full 1603-test suite green while silently
deleting the charter entry, which would launch an opencode member with
no role contract at all.

Fix: charter's putArray -> withArray (one-word change, behavior-identical
today). Add three tests asserting instructions[] CONTENT as an exact
ordered list (not size) across writer combinations: charter only,
charter + IDE rules, and charter + IDE rules + seeded skills. Mutation
testing (see PR body) shows the skills and IDE-rules writers are each
independently detectable by name; the charter writer's own mutation is
not detectable by any test, because it structurally always runs first
against an empty array, so putArray and withArray are equivalent there.
2026-09-10 19:58:25 +07:00
Dai Ha 6c2d6e93cb fleetd #469: one canonical FleetTool set backs registration, authz and charter checks
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m37s
FleetConfig.validateCharters() only checked that a charter key is a role
wire name and its text is non-blank; #464's CharterToolSurfaceTest compared
charter text against the registered tool surface, but wrote its own charter
into a @TempDir fixture, so nothing anyone wrote into the live fleetd.yaml
could ever fail it.

Add FleetTool, an enum in dev.ltms.fleet.mcp holding the one canonical set
of registered tool wire names. FleetMcp's tool schemas now derive their
names from it, its constructor asserts at startup that what it actually
registers with the SDK equals FleetTool.wireNames() exactly, and its authz
dispatch (toolAction/authzAction) resolves the wire string against FleetTool
before switching on the enum itself with no default -- adding a tool without
pinning its Authz.Action is now a compile error, not just a test gap.

Add CharterToolSurface (mcp package, not config -- config loads before the
MCP server exists) and call it from Fleetd.main right after
cfg.validateAll(), so a charter naming a tool the server does not register
refuses the daemon's startup, naming both the charter key and the unknown
tool. FleetdStartupValidationTest proves this through Fleetd.main itself
against a live-shaped config fixture (bridge_send, CB-634's own removed
name).

CharterToolSurfaceTest, FleetMcpAuthzTest and McpContractDocTest each kept
an independent regex scrape of FleetMcp.java's source for the registered
side of their own comparison -- three more copies of the same list nothing
tied together. All three now read FleetTool.wireNames() instead.

Proved canonical by removal: deleting FleetTool.ACK while ackTool() still
referenced it broke mvn compile in two places (FleetMcp.java:940,:1898);
registering a schema under a literal not backed by FleetTool
("fleet_ack_v2") failed FleetMcp's new startup assertion in every test that
constructs it (12 errors, IllegalStateException at FleetMcp.<init>). Both
reverted before this commit.
2026-09-10 19:55:02 +07:00
Dai Ha 789b6a8716 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
CI / contract (push) Successful in 1m56s
CI / build (push) Successful in 2m4s
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things on the record because they are judgement calls, not facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this codebase
  reports a spawn success back to this class, so "it started working again"
  cannot be observed here. A base cooldown of quiet is the best available
  evidence, and the class doc says that plainly instead of implying the
  stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - the same
formula with multiplier 1.0 and a ceiling equal to the base, which collapses
to the original flat "now + cooldown". Every existing call site keeps its
shape.

The second commit exists because my own mutation on the first merge found
the wiring unpinned: putting Fleetd.main back on the flat constructor left
all 1608 tests green, so the factory was pinned and the decision to use it
was not. FleetdBackendQuarantineWiringTest closes that, following the five
existing *WiringTest files rather than inventing an idiom. It checks source
TEXT, and its class doc says so: it does not prove the call executes, and it
cannot tell "wrong factory" apart from "renamed the anchor" - both fail the
same assertion. That limit is real and recorded rather than papered over.

Build number and my own mutation results are on the ticket and the PR,
measured on this merge commit rather than on the branch.
2026-09-10 19:51:00 +07:00
Dai Ha 01462c9695 fleetd #466 follow-up: pin main's choice of the escalating quarantine factory
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m25s
BackendQuarantineTest proves BackendQuarantine.withEscalation itself
escalates, is ceilinged, and resets. Nothing proved Fleetd.main still
calls it: reverting main to the flat two-argument constructor compiled
with 0 errors and left the full 1608-test suite green, because every
other test builds its own BackendQuarantine directly.

Adds FleetdBackendQuarantineWiringTest, a source-text assertion on
Fleetd.java (same idiom as FleetdLeadSeatWiringTest and
FleetdCompletionResolverWiringTest) pinning that main's declaration is
built from BackendQuarantine.withEscalation(...), not `new
BackendQuarantine(...)`. Measured: reverting to the flat constructor
fails it (expected true, was false); renaming the anchored local
variable while keeping behaviour identical also fails it loudly, not
silently. The class doc states plainly this checks source text only,
not that the call executes or that escalation reaches a real backend.
2026-09-10 19:47:59 +07:00
Dai Ha 8b4ff78546 Merge #470: the exhaustion quarantine escalates instead of retrying flat (fleetd #466)
A flat 30-minute cooldown suits a backend that is out of capacity for the
hour. It does not suit a weekly subscription limit: that keeps reporting
exhausted until the window resets, so the daemon retried it roughly 336
times across a week and learned nothing each time.

BackendQuarantine now tracks, per credential, how many consecutive
exhaustion reports it has seen with no quiet gap between them, and doubles
the cooldown each time, capped at 12x the base (about 6 hours at the 1800s
default). That is about a dozen attempts a week instead of ~336.

Three things I want on the record because they are judgement calls, not
facts:

- The reset is a TIME PROXY, not a success signal. Nothing in this
  codebase reports a spawn success back to this class, so "it started
  working again" cannot be observed here. A base cooldown of quiet is the
  best available evidence. The class doc says this plainly rather than
  implying the stronger thing.
- No automatic probing. That was the operator's design constraint and the
  implementation respects it: the daemon warns and waits, it never pokes a
  limited backend to see whether the limit lifted.
- The multiplier (2.0) and ceiling (12x) are constants, not config surface,
  so no new ConfigRef hot/cold/deferred classification question arises.
  quarantineCooldownSeconds stays Deferred and is now the BASE of the
  backoff; fleetd.example.yaml and FleetConfig's javadoc say so.

Escalation fires on the exhaustion signal alone. Cooling-off
(BackendOutagePolicy, a flat 60s after repeated non-exhaustion errors) is a
separate mechanism and is deliberately NOT escalated: doing so would turn a
transient 5xx storm into a multi-hour outage.

The old two-argument constructor is behaviourally unchanged - it is the
same formula with multiplier 1.0 and a ceiling equal to the base, which
collapses to the original flat "now + cooldown". Every existing call site
keeps its shape.

Build number and my own mutation results are reported in the PR and on the
ticket, measured on this merge commit rather than on the branch.
2026-09-10 19:40:15 +07:00
Dai Ha d4a2cd720c fleetd #393: deliver memberSkills to opencode members, and stop overclaiming seeding success
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m46s
GitWorktrees.seedSkills copies memberSkills:-seeded skill folders into every
provisioned worktree's .claude/skills/ and logged "skill seeding: N of M" as
if that were success — but .claude/skills/ is a Claude Code CLI convention.
opencode has no such discovery, so a kind: opencode member never actually
read a seeded skill even though the log said N of M succeeded.

Two changes, both required:

1. Deliver it. OpenCodeLauncher.skillInstructionFiles scans
   <cwd>/.claude/skills/*/SKILL.md at spawn time (the one point the launcher
   knows both the kind and the cwd) and appends each to the generated
   opencode.json's instructions[] array, the same channel already used for
   the member charter and IDE rules. A skill folder with no SKILL.md is
   named and skipped rather than silently dropped.

2. Stop claiming it where the claim can't be verified. GitWorktrees.seedSkills'
   log now says explicitly that consumption depends on the member's kind and
   points at the launcher's own log; OpenCodeLauncher logs its own kind-aware
   "skill delivery: M of N ..." line once the kind is actually known, naming
   any folder it could not turn into an instructions[] entry.

fleetd.example.yaml's memberSkills: doc previously claimed "Claude Code
members only; an opencode member reads a different path (.opencode/agent)
this key does not touch" — false as of this fix, corrected to name both
kinds and how each consumes it.

Tests: OpenCodeLauncherTest gains two cases driving the real
GitWorktrees#add seeding path (not a hand-built fixture) into an
opencode-kind spawn — one asserting a seeded skill's SKILL.md lands in
instructions[] plus the honest log line, one covering a skill folder
without SKILL.md (delivered skills still flow, the malformed one is named
in the log and excluded from instructions[]). ClaudeCodeLauncher is
untouched — its native .claude/skills/ discovery already worked and is out
of scope.

mvn -B clean test: Tests run: 1603, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
2026-09-10 19:37:00 +07:00
Dai Ha 5a467e1f8b fleetd #466: escalate BackendQuarantine's cooldown on repeated exhaustion
CI / build (pull_request) Successful in 1m34s
CI / contract (pull_request) Successful in 1m36s
A flat 30-minute quarantine retries a weekly subscription limit about
336 times before the window resets. BackendQuarantine now doubles the
cooldown on each consecutive exhaustion of the same credential (no
more than one base cooldown after the previous quarantine's deadline),
capped at 12x the base cooldown (~6h at the 1800s default), and resets
back to the base cooldown once a base-cooldown's worth of quiet has
passed with no further exhaustion.

The flat two-argument constructor is unchanged (equivalent to
multiplier 1.0 / ceiling == base), so all ~20 existing call sites keep
their current shape and behaviour. Production wiring (Fleetd.main)
switches to the new BackendQuarantine.withEscalation factory.

This only touches the exhaustion path (BackendQuarantine's one
production caller is Fleetd.exhaustionSink, fired on BACKEND_EXHAUSTED
alone) and never the separate, unescalated cooling-off mechanism
(BackendOutagePolicy, fixed 60s) that guards against a transient
backend-error storm.
2026-09-10 19:33:09 +07:00
Dai Ha 1fb6176783 Merge #457: exhaustedPattern goes hot, and the warning names the fix (fleetd #446)
CI / contract (push) Successful in 1m31s
CI / build (push) Successful in 1m33s
Three rounds. Round 1 made exhaustedPattern a hot config key, made the
quarantine warning name the fix instead of only the fact, and added model/reason
to fleet_profiles' quarantined rows. Round 2 extracted the warning text into
usageLimitFixWarning/usageLimitFixWarningNoModel and pinned both. Round 3
extracted the sink itself into a static exhaustionSink(...) factory and pinned
what it actually logs, using a ListAppender on this class's own logger.

Where the mutation numbers below come from, stated exactly. The battery ran on
merge commit 3d2d521, tree 1dcca23: origin/main at 235644c plus this branch.
This merge commit's tree is 953ce11, which is that tree plus one file --
CharterToolSurfaceTest, from the #464 merge (49df792) that landed on main while
the battery was running. So the battery did not run on this exact tree. That one
added file was verified green on its own merge. The build on THIS tree, tree
953ce11, is: Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 0 compile-error blocks. That is the battery tree's 1600 plus that one
test, which is the arithmetic the two trees predict. Saying this rather than
implying one tree.

On the battery tree: 1600 tests green, 0 compile-error blocks. FleetMcp.java
auto-merged there against #463's change to the same file, so that build was also
the gate on the auto-merge -- a clean auto-merge is not a compiling merge.

- M5, the round-2 survivor reproduced verbatim: the call site stops using either
  pinned method and logs a literal instead. Round 2 left this green across 1592
  tests. Now KILLED by
  FleetdExhaustionSinkWarningTest.profileWithModelGetsTheActionableFix and
  .profileWithNoModelGetsTheFallbackNotAFix.
- M6, the ternary's two branches swapped, so a profile WITH a model gets the
  no-model text and vice versa. Both extracted methods and both their unit tests
  untouched. KILLED by the same two tests, independently.
- M7, THE HALF THIS ROUND DID NOT PIN, and it SURVIVED. main()'s call to the
  factory replaced with an inert lambda: the factory and its test stay perfect,
  the daemon quarantines nothing and logs nothing. 1600 green.

M7 is the same defect shape as round 2, moved one level up, and it is worth
naming plainly. Extracting a thing in order to pin it CREATES the seam the test
then lands on. Round 2 extracted the warning TEXT, pinned the two methods, and
left the call site that selects between them unpinned -- M5. Round 3 extracted
the SINK, pinned the factory's behaviour, and left main()'s wiring of it
unpinned -- M7. The question to ask on any such fix is which of three things a
test now reaches: the value, the call site, or the selection between values.
Extraction only ever answers the first.

M7 is not a reason to hold this merge. Proving main() wires this factory means
driving daemon startup, which nothing here does -- that is fleetd #460. A
cheaper option exists and is recorded there: a source-reading assertion of the
kind #439 used, which would kill M7 without starting the daemon.

Round 3's own caveat, disclosed by the worker and worth keeping: its first Cell A
run was corrupted by a second concurrent mvn against the same module directory.
It killed that run, checked for stray ForkedBooter processes, and re-ran. The
numbers above are mine, from this battery, not that run.
2026-09-10 19:16:24 +07:00
Dai Ha 49df79203c Merge #464: a test that charter text names only registered tools (fleetd #464)
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m56s
CharterToolSurfaceTest extracts every fleet_* / bridge_* token from configured
launch charters and every tool("fleet_...") FleetMcp registers, then asserts the
first set is a subset of the second.

Verified on the merge commit. Its three acceptance criteria are met:

- Catches the ticket's own example. Fixture charter naming bridge_send, the tool
  CB-634 renamed away: KILLED.
- Fails loudly with no charter text. Fixture stripped of every tool name: KILLED
  by its named.isEmpty() guard, not a silent pass.
- Fails loudly with no registered tools. The tool("...") scrape broken so it
  matches nothing: KILLED by its registered.isEmpty() guard.
- And one cell of my own: the server stops registering fleet_reply, which the
  fixture names. KILLED. This is what proves the 'registered' half reads real
  production source and is not a second fixture.

The scrape finds 11 registered tools: ack, ask, list, poll, profiles, reply,
send, spawn, status, stop, whoami. An independent count of every "fleet_x"
literal in FleetMcp.java is also 11.

WHAT THIS DOES NOT CLOSE, and it is the ticket's actual gap. The charter half is
a @TempDir fixture the test writes itself, so no charter text anyone writes can
make this test fail. Measured: the test mentions fleetd.yaml 0 times, and the
commit changes 0 production files -- FleetConfig.validateCharters() still never
reads charter text (0 lines of its body mention a tool name). So this pins the
comparison logic and acts as a rename tripwire for the two tools the fixture
names. It does not check the live config. That needs a production-side check and
is filed as a follow-up.

That residue is my ticket's fault, not the worker's: the three criteria I wrote
are exactly the three it met.

A note on my own battery, because it nearly published four false kills. The
first run showed rc=1 on all four mutation cells and I would have read that as
four kills. It was zsh: unquoted parameters are NOT word-split, so
'mvn -B $scope test' passed '-Dtest=X -DfailIfNoTests=false' as ONE argument and
surefire ran zero tests. The tell was a missing 'Tests run:' line. The rerun
proves the harness first -- selector alone must report 'Tests run: 1' -- and
every cell now prints its surefire summary count so a void cell cannot pass for
a kill.
2026-09-10 19:10:56 +07:00
Dai Ha 235644c0f0 Merge #467: listFleet's callerIsPrimary default fails closed (fleetd #463)
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m38s
A wrapper overload that is called with no callerIsPrimary argument used to
default it to true, so a forgotten argument silently handed out the lead's
coordination state. It now defaults to false: a missing identity fails closed.

Verified on the merge commit, not the branch:

- The funnel is real. FleetMcp has 7 listFleet declarations and 7 real calls
  (an 8th 'listFleet(' match is a javadoc {@link}). Exactly one call writes a
  literal for the new boolean, and it writes false; exactly one writes the real
  predicate, coordinatorVisibleTo(principal(exchange)) in the MCP handler. No
  call writes true.
- Control battery on the merge: 1584 tests green unmutated.
- M1, the fix reverted at the one line that writes the default (false -> true):
  KILLED by listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey.
- M2, the half this round weakened. The worker rewrote
  listIsByteForByteUnchangedForThePrimaryCaller and dropped its byte-for-byte
  equality assertion, which was correct because that comparison ran against the
  implicit-default overload -- now the path #463 closes. So: does anything still
  notice if the primary's coordinator row silently loses a field? Dropped
  heldCount: KILLED by two tests, that same test and
  listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero.
- M3, a regression check on #439's own gate, 'if (callerIsPrimary)' -> 'if (true)':
  KILLED by three tests.

What is no longer pinned, stated plainly: the primary's output is now checked
field by field, not as a whole string. A field that no test names could
disappear without failing anything. Every field the tests do name is pinned,
proven by M2. The whole-output answer belongs to fleetd #460.

One control in my own battery was wrong and is worth recording: I labelled
', true);' in FleetMcp.java 'must be 0' and it is 2 -- row.put("configured",
true) and m.put("self", true), neither a listFleet delegation. The pattern was
too wide. The count above comes from a walk over each declaration and call
instead.
2026-09-10 19:04:51 +07:00
Dai Ha 7772b41993 fleetd #446 round 3: extract exhaustionSink and pin its caller (Cell A/B)
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 2m13s
Round 2 pinned usageLimitFixWarning/usageLimitFixWarningNoModel's TEXT via
FleetdUsageLimitFixWarningTest, but a mutation battery against the merged PR
proved two gaps in the caller that builds main()'s real quarantine
ExhaustionSink: nothing proved the sink's log.warn actually invokes either
method (Cell A), and nothing proved it picks the right one for a profile
with vs without a configured model: (Cell B).

Extract the inline lambda into a new static Fleetd.exhaustionSink(...)
factory (same refactor-for-testability class the lead approved in round 2
for the two static warning methods), behaviourally unchanged from the
lambda it replaces. FleetdExhaustionSinkWarningTest drives this factory's
return value directly and asserts on the real text a ListAppender attached
to Fleetd's own logger captures, covering both cells from one mechanism:

- Cell A (ternary result replaced by a literal string): confirmed red,
  1594 run / 1 failure, restored, confirmed green (1594/0).
- Cell B (ternary's two branches swapped): confirmed red, 1594 run /
  2 failures (both test methods independently caught it), restored,
  confirmed green (1594/0).
2026-09-10 19:00:07 +07:00
lead eccd0548ce Merge #465: state canonical invariant 5 as a purpose, not a banned tool (fleetd #458)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 2m1s
Verified by me on a local merge of 29e7a06 onto f5e02fe:

- One file, one hunk. 5 changed lines inside the canonical block, all of them
  invariant 5; a line-by-line diff of the block against origin/main shows
  nothing else moved.
- The addendum is untouched: both 'Herdr socket tests (measured 2026-09-10)'
  and 'Only the lead can run that check' appear once on each side.
- Block size: origin/main 17467 chars / 17626 bytes, the merge 17557 chars /
  17718 bytes. The worker's numbers were bytes and were right; I am naming both
  units because len(str) in python is characters and this block is full of em
  dashes, which is a mistake I have made before.
- mvn -B clean test: Tests run: 1583, Failures: 0, Errors: 0, Skipped: 0.
  BUILD SUCCESS, rc=0.

The three things I reserved from the worker, because a provisioned worktree
cannot do them:

- wiki/7-Use-Cases.md updated to match, in the wiki repo's own commit.
- The sync script run in this clone: in sync: True.
- Other projects carrying the block: NONE. 29 CLAUDE.md files under the
  operator's project roots, 18 of them carry the block and 17 of those are this
  repo's own worktrees, which inherit it from git. Only
  LTMS/claude-bridge/CLAUDE.md is a real second copy, and it is this file.

On the worker's criterion-4 question: the addendum note stays. The restated
invariant removes the unsatisfiability, which was the sharp problem, but the
note still names the exact four files the carve-out covers and records why
(fleetd #449, where a fake and the real herdr disagreed for weeks under a green
suite). A rule that is merely satisfiable is not the same as a rule a worker can
apply without re-deriving it.
2026-09-10 18:57:24 +07:00
Dai Ha c4e23eebad fleetd #464: guard charter tool names
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m37s
2026-09-10 18:55:16 +07:00
Dai Ha 7df503dfc2 fleetd #463: default listFleet's callerIsPrimary to false, fail closed
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 2m9s
The compat overload at FleetMcp.java:1373 defaulted callerIsPrimary to a
literal true, so a caller that forgot the argument silently got the
coordinator row (this daemon's coord-id, mailbox state, held-mail previews,
peer reachability) -- lead-to-lead state fleetd #439 just gated. Flip the
default to false: a forgotten argument now yields a missing row instead of
a leaked one.

Six FleetMcpTest methods relied on the implicit true to see the coordinator
row at all; they now pass true explicitly through the canonical overload.
listIsByteForByteUnchangedForThePrimaryCaller's own premise (comparing the
implicit-default path against an explicit-true path) was the shape of the
bug, so it now only exercises the explicit-true path.

Added listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey
to pin the new default: a compat overload called with no callerIsPrimary
argument, against a fully-configured lead channel, must produce a result
with the coordinator key absent -- not empty, not redacted, absent.
2026-09-10 18:55:13 +07:00
Dai Ha f5e02fedd6 plans: commit the fleet01 move plan, with today's state measured on the host
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m58s
This plan was written 2026-09-05 and has been sitting untracked in the
working tree since, so nobody but this machine could read it and nothing
recorded that it existed. plans/ is a tracked directory here.

Committed with a status section measured today over a read-only ssh survey,
because a five-day-old plan committed as-is would read as current:

- Phases 1 to 3 are done. Both systemd user units exist, are active and
  enabled, linger is on, one java process (so the old restart.sh
  double-daemon problem is gone), the live fleetd.yaml is on the host, and
  the jar was rebuilt 2026-09-10 02:10 UTC.
- The checkout has drifted again: main at 4887731, 77 commits behind
  origin/main, 0 ahead. So fleet01's daemon runs code from before this
  week's merges.
- Phases 6 and 7 are not done. The Mac still runs the daemon this fleet
  uses, and fleet01 has 2 weeks of uptime, so no reboot proof exists.
- Section 9 is out of date: claude is installed on fleet01 and a fleet01
  lead is live on the coordination channel, so the headless-login blocker
  it names is solved.

Also records fleet01's four profiles with their weights, because
placement: weighted plus gx at 100 and xf at 80 means an unqualified spawn
there almost never lands on a claude-code profile.

And one search that found the opposite of what I expected, written down so
the next session does not repeat it: opencode IS installed on fleet01, at
~/.opencode/bin/opencode, but only on the INTERACTIVE PATH. zsh -ic finds
it, zsh -lc does not. A herdr pane is a non-login interactive zsh and
fleetd types the launch command into the pane, so gx and xf are not broken
by this. The daemon's own login-shell ExecStart cannot see it, which is the
mirror of the trap that ExecStart exists to fix: credentials live on the
login side, ~/.opencode/bin on the interactive side.
2026-09-10 18:53:11 +07:00
Dai Ha 29e7a06c49 fleetd #458: restate invariant 5 by purpose, not mechanism
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Failing after 1m56s
Invariant 5 banned 'driving the terminal multiplexer directly', naming
herdr CLI and socket as the banned tool. That bans a mechanism. What it
protects is the control plane: nobody may move a fleet session, pane or
peer by a route that skips the bridge's policy checks.

In this repo herdr is itself the subject under test, so four contract
tests must open its socket on purpose (see the addendum's 'Herdr socket
tests' note). Under the old wording, a worker assigned to that code
reads invariant 5 and finds its only path to finish the task banned.

Restate the invariant by purpose: never move fleet state except through
the bridge. herdr's CLI and socket stay as the named example of the
banned route, not the definition of it.
2026-09-10 18:50:41 +07:00
Dai Ha 92a96fcbd8 Merge #462: gate the coordinator row on a named predicate the handler must consult (fleetd #439)
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m51s
Verified by me on a local merge of c1ca627 onto 1348287:

- mvn -B clean test: see the totals below. Ran in fleetd/, redirected to a file,
  exit code captured on its own line.
- Mutation battery on merge 8402923 (same two parents, earlier base 5f1b260),
  4 cells, each with a proof gate on the occurrence count:
    * CONTROL, unmutated: 1583 tests, 0 failures, rc=0.
    * M2r - the handler stops asking who called: coordinatorVisibleTo(principal(exchange))
      replaced by a literal true. KILLED by
      FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo.
      This is the mutant that survived round 1, so the gap the follow-up was for is closed.
    * M3 - is the new source-reading detector vacuous? Renamed its anchor
      (listHandler -> listHandlerX, 2 sites, behaviour identical). The detector FAILED,
      rc=1, as a source-reading test must: it cannot silently pass on an empty scrape.
    * M4 - the predicate itself always says yes: return caller.isPrimary() replaced by
      return true. KILLED by FleetMcpAuthzTest.onlyThePrimaryMaySeeTheCoordinatorRow.
  Tree verified clean before the battery and restored after each cell.
- ANON is pinned too, not only worker and architect: FleetMcpAuthzTest asserts
  coordinatorVisibleTo(ANON) is false.
- listFleet( appears in exactly one main file, mcp/FleetMcp.java, and no REST class builds
  the coordinator row, so the detector's single-file scope covers every live call site
  today. Control for that sweep: 109 main .java files matched a string they all contain.

Not in this PR, and my call, not the worker's: the six compat overloads of listFleet still
default callerIsPrimary = true, which fails open. Safe today because the one production
call site passes the computed value. Filed separately.
2026-09-10 18:40:48 +07:00
Dai Ha 13482872bb contract test: keep the loaded run's passes, discard only its failure
CI / contract (push) Successful in 1m30s
CI / build (push) Successful in 2m6s
The javadoc threw away the whole loaded run as "not clean evidence". The
fixed cell order makes that too strong in one direction. A cell that runs
last on a climbing load has a free explanation for FAILING. It has no free
explanation for PASSING: surviving a worse condition than a fair order
would have given it is evidence in the safe direction. So the old version's
0 of 3 is still discarded, and the three passes are kept with the load each
one ran at.

Also names the cell that tests the swallow explanation head-on, which the
old text left as "nobody has managed that yet". SHELL_READY_TIMEOUT_MS at 0
types input at once — the worst case for "typed before the prompt" — and it
passed 3 of 3 at load 18.42 to 23.65. The wider read window cannot explain
that away, because a swallowed keystroke is lost, not late: the command
never runs, so no amount of polling makes its output appear.

And a warning not to carry the raw load average to another host. Load
average counts differently per core and per operating system, so only load
per core compares. I broke that rule myself when comparing this run with
another host's numbers.

Both points came from the fleet01 lead reviewing 2af13ab.
2026-09-10 18:37:41 +07:00
Dai Ha c1ca6273fc fleetd #439: pin the caller at the fleet_list call site, not just the gate
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m33s
Review of PR #462 found M2: the coordinatorVisibleTo gate (then an inline
principal(exchange).isPrimary() check) could survive a mutation that
replaced the argument with a literal true at the one production call
site, because every existing test drove listFleet directly and supplied
the boolean itself -- nothing exercised the handler's own call.

- Name the decision: FleetMcp.coordinatorVisibleTo(Principal), a small
  package-private predicate next to denyFor/recordPrimarySingleton. The
  fleet_list handler now calls coordinatorVisibleTo(principal(exchange))
  instead of inlining .isPrimary().
- Pin the predicate's role table in FleetMcpAuthzTest
  (onlyThePrimaryMaySeeTheCoordinatorRow), covering primary/worker/
  architect and, newly, anonymous.
- Add a source-reading detector at the boundary
  (theFleetListHandlerActuallyConsultsCoordinatorVisibleTo), same idiom as
  toolsTheServerRegisters/everyRegisteredToolHasItsHandlerActionPinned: it
  reads FleetMcp.java, isolates the listHandler block, asserts (as a
  control) that the block actually contains a listFleet( call, then
  asserts the call's trailing boolean argument is exactly
  coordinatorVisibleTo(principal(exchange)) -- not a literal true/false.

Both mutations from the review were reproduced and killed by these tests,
then reverted; see the PR body for the full break-and-restore transcript.
2026-09-10 18:28:16 +07:00
Dai Ha 5f1b260c81 addendum: the canonical sync check is the lead's, not a member's
CI / contract (push) Successful in 1m4s
CI / build (push) Successful in 1m45s
I put "the sync script must print in sync: True" in a worker's acceptance
criteria for #455. The worker could not run it and said so, honestly, instead
of inventing a pass. My brief was the defect.

A member's provisioned worktree has wiki/ uninitialized, so the script dies
with FileNotFoundError. Measured in three worker worktrees: git submodule
status printed a leading '-' and wiki/ held 0 entries. The primary's own clone
printed a leading '+' and the file was there.

The note is dated, gives the re-measure command, says what each outcome means,
and says to delete it once it stops reproducing - as this file requires of any
measurement in an addendum.

Canonical block untouched: the sync script itself reports in sync: True.
2026-09-10 18:25:32 +07:00
Dai Ha 30d6872779 fleetd #446 follow-up: pin the WARNING text and the fleet_profiles model/reason fields
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m4s
A mutation battery run against merged PR #457 (d02dd1b) proved criteria 2 and 3
shipped without a test that could catch them breaking: deleting either
row.put("model", model) or row.put("reason", reason) in FleetMcp.profilesView
left all 1586 tests green (rc=0), and renaming the "usage-limit fix:" log tag
to something meaningless did too. Criterion 1's own mutation (LiveExhaustedPatterns
snapshotting instead of reading live) was correctly killed by the existing
FleetdExhaustionDetectionArmedWiringTest/LiveExhaustedPatternsTest — only 2 and 3
were unguarded.

- Extracted the exhaustionSink WARNING text out of two inline SLF4J {}-placeholder
  log.warn calls into two static methods, Fleetd.usageLimitFixWarning(profile, model)
  and Fleetd.usageLimitFixWarningNoModel(profile, quarantineCooldownSeconds), the
  same extracted-static-method + dedicated-test idiom as
  exhaustedPatternCoverageLine/errorPatternCoverageLine. log.warn is now called with
  each method's return value as a single already-formatted argument, so the string a
  test asserts on is byte-identical to what fleetd.out receives. No behaviour change:
  same text, same two branches, same call site.
- New FleetdUsageLimitFixWarningTest pins the leading "usage-limit fix:" grep tag and
  three specific facts (profile name, model name, "enabled: false" under
  models.allow) rather than the whole sentence, plus the no-model fallback's profile
  name and cooldown-seconds substitution and its explicit absence of "enabled: false".
- New FleetProfilesQuarantineModelReasonFieldsTest exercises FleetMcp.QuarantineSource
  with a modelFor/reasonFor that actually return values (every existing test used
  QuarantineSource.none() or a 3-arg form defaulting both to null), asserting the
  quarantined row's model/reason are present when supplied and absent (not null, not
  blank) when modelFor/reasonFor return null or a blank string.
- Each of the three: implemented, broken by hand (row.put deleted / tag renamed),
  confirmed the new test goes red, restored, confirmed green again. See PR body and
  this ticket's fleet_reply for the verbatim failure output of all three.
2026-09-10 18:22:51 +07:00
ltms 9d1306d442 Merge #461: project addendum for the herdr socket carve-out (fleetd #455)
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m56s
Doc-only, one file, 13 added lines, inside §Project addendum. Verified by me:

- Canonical block byte-identical: 17468 chars on main and on the branch.
- The sync script in CLAUDE.md prints "in sync: True" both before and after.
- The note's own re-measure command works. Run verbatim, it returns exactly the 4 files
  the note names, and no others:
    fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java
    fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java
  Control that the search reaches the tree: 120 test java files, 63 of them mention herdr.
  That control matters here — a zero-match re-measure command would tell a future session to
  delete a live restriction.
- All four required elements are present: the carve-out, what stays banned (the control plane),
  who it applies to, and the perishable half (dated 2026-09-10, the command, what each outcome
  means, and delete-when-stale).
- Cross-references #458 for the canonical restatement, which is deliberately not in this change.

Pre-send check applied to the note itself: a worker assigned to AgentControlContractTest can now
name one legal action that finishes its task — let the test open the herdr socket in a throwaway
workspace it tears down.
2026-09-10 13:18:17 +02:00
Dai Ha e54e3d87ea fleetd #439: omit fleet_list's coordinator key for non-primary callers
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
The coordinator row is lead-to-lead coordination state (coord-ids, mailbox
facts, held-message previews). fleet_list returned it to every caller,
including a worker or an architect, because coordinatorView() had no way to
know who was asking.

Gate at the call site inside listFleet: a new overload takes
callerIsPrimary and only assembles/attaches the coordinator row when it is
true, so the key is absent (not empty) for a worker or an architect. The
MCP handler now passes principal(exchange).isPrimary(); every other
listFleet overload keeps passing true, so callers with no caller identity
(existing unit tests, the no-op wrappers) are unaffected -- confirmed by a
byte-for-byte comparison test against the pre-fix overload.

Authz's READ case is untouched: it stays shared by fleet_status,
fleet_profiles and fleet_whoami, and the gate here is purely inside
fleet_list's own result assembly.
2026-09-10 18:16:48 +07:00
Dai Ha bf027f10b9 #455: document herdr test socket exception
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m50s
2026-09-10 18:13:50 +07:00
Dai Ha 3c5873dfe2 fleetd #453: point the new javadoc at the right javadoc
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m52s
#456's first paragraph said "HerdrPeerLauncher's own {@link
#spawn(SpawnRequest, PlacementDecision)} javadoc". That link resolves to this
interface's own abstract declaration, not to HerdrPeerLauncher's override, so a
reader who follows it lands on the wrong text. The second paragraph already
used the plain {@code HerdrPeerLauncher.spawn(...)} form; both now match.

Also says who is actually forced to read the paragraph, because #456's
reasoning rests on it and the two cases differ. A class that implements this
interface directly must write a body for spawn(SpawnRequest,
PlacementDecision) - it is abstract here - so it reads this javadoc. A
subclass of HerdrPeerLauncher does not: HerdrPeerLauncher already implements
that method (member/HerdrPeerLauncher.java:611) and the subclass inherits the
body. For a subclass the paragraph is advice, not a gate.

javadoc -Ddoclint=reference: 5 "reference not found", the same 5 in the same 5
untouched files as origin/main at 2af13ab, and none in PeerLauncher.java. Those
5 are ticket #459. mvn compile rc=0.
2026-09-10 18:05:07 +07:00
ltms e29227d5f4 Merge #456: document the override obligation on PeerLauncher.defaultProfileFor/place (fleetd #453)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m34s
Doc-only, one file, 17 added lines. Verified by me on a local merge of 0788d84
onto 2af13ab (merge commit 5b46538):

- mvn -B clean test: Tests run: 1578, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS, rc=0.
- javadoc reference lint (mvn javadoc:javadoc -Ddoclint=reference): the merge has 5
  "reference not found" errors; origin/main at 2af13ab has the same 5, in the same 5
  untouched files. So this PR adds no broken link. Both runs rc=1 for that pre-existing
  reason, which is now ticket #459.
- The quoted phrase is real: "are unoverridden here and just wrap {@link #defaultProfile()}"
  is at member/HerdrPeerLauncher.java:604. place() is not overloaded, so {@link #place}
  is unambiguous.

One inaccuracy I am fixing in a follow-up commit rather than sending the PR back: the first
new paragraph writes "HerdrPeerLauncher's own {@link #spawn(SpawnRequest, PlacementDecision)}
javadoc", but that link resolves to PeerLauncher's own abstract declaration, not to
HerdrPeerLauncher's override. The second paragraph already gets this right with plain
{@code HerdrPeerLauncher.spawn(...)}.
2026-09-10 13:04:09 +02:00
Dai Ha 45aca9eb3e fleetd #446: make exhaustedPattern hot, name the fix in the warning, report it in fleet_profiles
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m34s
The model gate can be turned off at runtime (models.allow[].enabled: false, hot
since fleetd #422), but the usage-limit detector it's meant to react to was
compiled once at Fleetd.main startup into a frozen Map<String,Pattern> — arming
or disarming exhaustedPattern needed a daemon restart. Backwards for a feature
meant to react live.

- New LiveExhaustedPatterns: reads exhaustedPattern off the live config supplier
  per lookup (matching CompositePeerLauncher#models0's live-supplier pattern),
  caching compiled Pattern objects by PROFILE NAME (not pattern text — pattern
  text would grow unboundedly as an operator tunes a regex across reloads;
  profile names are bounded by the small, restart-gated set of configured
  profiles). patternFor() backs CompletionResolver's classification; armed()
  backs fleet_profiles' exhaustionDetectionArmed — both read the same object,
  the fleetd #404 single-accessor rule CompositePeerLauncher.modelGateState()
  established for the model gate.
- Fleetd.java: on BACKEND_EXHAUSTED, log a WARNING naming the profile's model
  and the exact fix (enabled: false under models.allow, hot, no restart; remove
  it again once the window resets) — or, when the profile has no model:
  configured, say quarantine is the only thing keeping spawns off it.
- FleetMcp.QuarantineSource gains modelFor/reasonFor; fleet_profiles'
  quarantined rows gain model/reason fields so a lead can see why without
  reading the daemon log. capacityView (fleet_list) intentionally untouched —
  scoped to fleet_profiles only.
- ConfigRef/FleetConfig docs + fleetd.example.yaml updated: exhaustedPattern
  moves from Deferred to Hot. errorPattern stays deferred on purpose (out of
  scope for this ticket).
- Tests: LiveExhaustedPatternsTest (new, unit-level hotness/caching proof),
  FleetdExhaustionDetectionArmedWiringTest (rewritten — same QuarantineSource
  object, read before and after a reload, asserts the answer flips with no
  restart), ConfigRefTest/ConfigRefProfileCoverageTest updated for the new
  Hot/Deferred classification.
2026-09-10 18:02:00 +07:00
Dai Ha 2af13ab1ff fleetd #449: name the decisive cell and the load the claim was measured at
CI / contract (push) Successful in 54s
CI / build (push) Successful in 2m8s
The previous comment named a cause with no cell behind it. fleet01's review
made the point: a confident wrong mechanism gets copied, and a confident
under-determined one gets copied the same way.

So the comment now names the cell that settles it. Hold the old 1000ms write
sleep and change only the read - 800ms fixed sleep becomes a 5s poll - and the
test goes 0 of 3 to 3 of 3. The read deadline was the whole story.

It also says where: a 12-core macOS host near idle (load 2.6 to 5.9). The
loaded run agreed, but its load climbed from 7 to 50 while the cells ran and
the old version ran last, so it is not clean evidence and the comment says so.

Comment only. No test or production code changed.
2026-09-10 17:59:52 +07:00
Dai Ha 0788d84be8 fleetd #453: document the override obligation on PeerLauncher.defaultProfileFor/place
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m34s
Decision: leave both as default methods (option 1), not abstract. Neither
default is a live defect today — HerdrPeerLauncher is the sole single-profile
implementer and the degenerate answer (ignore role, always defaultProfile())
is correct for it. Making them abstract would force ~10 boilerplate one-line
overrides across 5 unrelated PeerLauncher test doubles (NeverSpawnsLauncher,
RaceLauncher, NoResumeLauncher, ClearContextSpyLauncher, LazyIdLauncher) that
never call either method, for a risk that is speculative (no multi-profile
HerdrPeerLauncher subclass exists or is planned).

Strengthens both javadocs with an explicit MUST-override warning and cross-
references HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)'s existing
#450 javadoc, which already names both methods as "unoverridden here" and
ties that to being a single-profile adapter -- the concrete place a future
multi-profile launcher author would read, since #450 made that method
abstract and any subclass must write its body.
2026-09-10 17:49:37 +07:00
Dai Ha 20c1094cbf fleetd #449: say what the timing fix actually proved, not what it assumed
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m33s
The polling fix that landed in #452 is right, but its javadoc named a
mechanism nobody measured: that input typed before the shell's prompt was
swallowed by the shell's own startup.

I mutated the settle poll away — SHELL_READY_TIMEOUT_MS = 0, so input is
typed at once with no wait — and the test passed 3 of 3. So waitForText is
the load-bearing half, and the proven cause is the old 800ms READ deadline,
not the 1000ms write delay.

The direction is the point: typing at 0ms works where typing at 1000ms
failed. If early input were swallowed, 0ms would be worse than 1000ms. It is
better, so the swallow explanation is unsupported.

waitUntilSettled stays as cheap insurance, now labelled as insurance rather
than as the fix. Comment-only; AgentControlContractTest still green.
2026-09-10 17:36:22 +07:00
ltms 9011c59b9f Merge #452: run the contract tag in CI, fix the stale protocol 14 assertion (fleetd #449)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m33s
Verified on a locally built merge onto c11ad71 (the PR was branched from 822327e, before #448):
1578 unit tests green, then `-Dgroups=contract` gives 30 tests / 0 failures / 0 skipped here.

CI's own run of the tag on d4f93a7: 29 tests, 0 failures, 6 skipped — all 6 herdr tests skip on
the runner, and LeadMailboxTest's 15 tests run there for the first time.

Two mutations killed: restoring `assertEquals(14` fails the test, and pointing the expected env
value at a wrong URL fails it with the real pane text in the message.
2026-09-10 12:35:10 +02:00
ltms c11ad71ed0 Merge #448: fleet_ack errors instead of claiming success on a miss (fleetd #437)
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m33s
Verified by the lead on head 5289eb5, and then again on the MERGE.

Battery on the branch tip:

  FULL BUILD  Tests run: 1577, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     contract test, real broker: Tests run: 9, Failures: 0, Skipped: 0

  M3a  AmqpReplyInbox.ack's `h == null` branch reports true
       -> KILLED  AmqpReplyInboxContractTest.ackReportsHitVsMissAgainstARealBroker
  M3b  AmqpReplyInbox.ack's `perTarget == null` branch reports true
       -> KILLED  same method

M3a is the gap I found in round 1: reporting true for something never held
survived both the default suite and -Pcontract against a real broker. It is now
closed in the adapter this daemon actually runs.

Both branches die, and they die through DIFFERENT assertions, which the worker
worked out and I confirmed by reading the code. held.get(target) is populated by
the deliver callback, not by own(), so after own() with nothing ever delivered
the map entry is still null — the "never held" assertion therefore exercises
`perTarget == null`, and only the double-ack assertion reaches `h == null`. Both
assertions are load-bearing; neither is redundant.

This PR was pushed before #451 landed, so the battery above tested the branch,
not the result. Merging main (bdcf285) in gave 0 conflicts and #448 adds no
PeerLauncher implementer, but a clean auto-merge is not a compiling merge, so I
built the merged tree 0829542:

  FULL BUILD               Tests run: 1578, Failures: 0  BUILD SUCCESS, 0 compile errors
  contract, real broker    Tests run: 9,    Failures: 0
  CompositePeerLauncherTest (#447's guarantee)  Tests run: 76, Failures: 0

The worker also declined to point the contract test at the shared local LavinMQ
broker, because this adapter never deletes queues and a run would leave orphaned
durable queues on the instance backing the live fleet. It started a disposable
rabbitmq:3.13-management container on a throwaway port instead, then removed it.
That was its own judgment and it was correct.

Held peer mail is deliberately NOT ackable: fleet_ack against a coord-id errors
and names fleet_poll{coordId}. See #437 for why refusing is the right answer
today, and why that is a policy choice rather than a structural one.
2026-09-10 12:20:49 +02:00
ltms bdcf285265 Merge #451: make PeerLauncher.spawn(req, decision) abstract (fleetd #450)
CI / contract (push) Successful in 1m17s
CI / build (push) Successful in 1m33s
Verified by the lead on head cfebc57 (base 822327e is current main).

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0

  M1  delete HerdrPeerLauncher's new override
      -> BUILD FAILURE, 1 COMPILATION ERROR block, naming exactly
         ClaudeCodeLauncher.java:[49,14] and OpenCodeLauncher.java:[59,14]

  M2  CONTROL for M1: put the interface method back to a `default` AND delete
      the override
      -> BUILD SUCCESS, 0 compile errors
      This is the row that makes M1 mean something. Without it, M1 only shows
      that the build broke; with it, the break is attributable to the method
      being abstract rather than to anything else the edit disturbed.

  M3  CompositePeerLauncher's override reverts to the re-entering form
      -> KILLED, Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace
      So #447's guarantee survives this refactor of the interface it rests on.

Tree restored clean after each mutation (git status --porcelain empty).

The worker corrected my ticket, and it was right. My #450 body listed five
src/main implementers of PeerLauncher and quoted `grep -rln 'implements
PeerLauncher'` as the source; that command returns two files. I had run a wider
pattern that also matched a comment in ConfigRef and the `extends
HerdrPeerLauncher` line in two subclasses, then quoted the narrow command beside
the wide command's output. Ground truth: ConfigRef implements
Supplier<FleetConfig>; ClaudeCodeLauncher and OpenCodeLauncher extend
HerdrPeerLauncher. So one override in that parent serves both, which is what the
worker built. The ticket body is corrected.

Out-of-scope note carried forward from the worker: defaultProfileFor(MemberRole)
and place(MemberRole) are two more default methods with the same shape. Noted,
not fixed here.
2026-09-10 12:16:47 +02:00
Dai Ha d4f93a7b13 fleetd #449: fix stale herdr protocol 14 javadocs/assertion, diagnose and fix the timing-raced AgentControlContractTest, select contract tests by tag in CI
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m9s
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
  protocol 19 (CB-521) left the client javadoc and the contract test's own
  assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
  renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
  assertion is real by temporarily changing the expected value to 20 (fails),
  then restoring 19 (passes).

- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
  failing, not skipping, on a host with a live herdr socket. Diagnosed with a
  temporary instrumented run (not committed) that polled the pane every
  200ms before and after sending input: the seed shell reliably takes ~2.5s
  to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
  raced that startup. Input typed too early was swallowed by the shell's own
  startup, leaving the typed line followed by the "Restored session" banner
  and no command output — indistinguishable at a glance from the env map
  never reaching the shell. Once the shell was actually ready, the injected
  env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
  fixed sleeps with bounded polling on the actual conditions (pane text
  settling, then the expected output appearing). Ran the fixed test 3x
  standalone, all green.

- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
  name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
  @Tag("contract") test from CI including the herdr ones above -- which is
  how the stale protocol 14 assertion went unnoticed. Changed to
  -Dgroups=contract, which selects the whole tagged group and picks up
  future contract tests automatically.
2026-09-10 17:14:09 +07:00
Dai Ha cfebc575ea fleetd #450: make PeerLauncher.spawn(SpawnRequest, PlacementDecision) abstract
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m5s
The default re-entered the single-argument spawn(SpawnRequest), which re-runs
checks that can refuse the profile place() just chose (#444's window). Only
CompositePeerLauncher overrode it; a future placement-doing launcher could
have inherited the wrong body silently.

Give every current implementer an explicit override, chosen by what it does:
- HerdrPeerLauncher (base of ClaudeCodeLauncher/OpenCodeLauncher, neither of
  which overrides spawn(req) or place()) does no placement filtering of its
  own, so it gets the re-entering form.
- CompositePeerLauncher's routing-form override is untouched.
- 5 test-fake PeerLauncher implementers (SessionManagerTest, FleetdBackendErrorSinkTest)
  get overrides matching their existing spawn(SpawnRequest) shape: delegating
  wrappers delegate, unreachable stubs throw, the single-profile fake re-enters.

ConfigRef does not implement PeerLauncher at all (confirmed in this tree at
822327e) despite the ticket listing it as a src/main implementer.
2026-09-10 17:10:44 +07:00
ltms 822327eed5 Merge #447: pin the place()-to-spawn() window PlacementDecision closes (fleetd #444)
CI / contract (push) Successful in 1m33s
CI / build (push) Successful in 1m36s
Verified by the lead on the exact tree that lands (head e4c703a, base 82fae94 is
an ancestor, so this is the tree I measured):

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     CompositePeerLauncherTest  Tests run: 76, Failures: 0  -> GREEN

  M1  the 2-arg spawn re-enters the 1-arg spawn (the inherited default this
      ticket forbids for a multi-profile launcher)
      -> KILLED  Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace

  M2  drop the stamping: route to the decided profile but do not carry it
      (SpawnRequest routedReq = req)
      -> KILLED  Failures: 1
      same test method

Both mutations proved applied by printing the mutated method, and the tree was
restored clean after each (git status --porcelain empty).

Round 1 of this PR had a fixture weakness I found by mutation: StubLauncher's own
fallback default was "sol", the same profile place() decides, so an unstamped
request landed on spawnCount("sol") by coincidence and M2 survived. e4c703a gives
the adapter "b" as its fallback instead. One fixture now kills both mutations.

src/main is javadoc-only in this PR: 12 added lines, 0 added code lines, measured
by filtering the main-side diff.
2026-09-10 12:00:59 +02:00
Dai Ha 5289eb509f fleetd #437: pin the ack hit/miss contract in the AMQP contract test
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 1m34s
AmqpReplyInboxContractTest is the one contract-group class CI actually
runs, and it never asserted on ack()'s return value at all — so the
exact defect this ticket fixes (reporting success for an ack that
removed nothing) was unpinned in the adapter fleetd runs live.

Add ackReportsHitVsMissAgainstARealBroker: a msgId never held for an
owned target returns false without throwing, a real held reply returns
true and is removed, and acking the same msgId again returns false.
Ran against both broker modes the class supports: Testcontainers
(AMQP_URI unset) and an external broker via AMQP_URI (the CI shape,
using a disposable container — not the shared local LavinMQ instance).
2026-09-10 16:58:55 +07:00
Dai Ha e4c703a51a fleetd #444: separate the adapter's fallback default from the decided profile
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m56s
Review found the fixture's StubLauncher fell back to 'sol' too — the
same profile place() decides — so an UNSTAMPED request could land on
spawnCount('sol') by coincidence, and the assertion's claim that the
request 'actually carried sol' was unproven. Dropping the stamping
(SpawnRequest routedReq = req) while keeping the routing survived the
test unchanged.

Fix: give the adapter 'b' as its own fallback default instead, so an
unstamped request counts against 'b', not 'sol'. Verified both
mutations against the single test in isolation:
  - drop-stamping (routedReq = req): RED, expected <sol> but was <b>
  - re-entering (return spawn(req.withProfile(decision.profile())))
    i.e. M1 from the first round: still RED, PlacementException
    naming the now-quarantined 'sol'
Restored both; full suite green at 1575 tests.

No changes to src/main — PeerLauncher's javadoc from the first round
is unchanged.
2026-09-10 16:50:39 +07:00
Dai Ha 703a05db41 fleetd #437: fleet_ack errors instead of claiming success on a miss
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Successful in 1m37s
ReplyInbox.ack now returns boolean (true = removed, false = nothing to
remove) instead of void, so FleetMcp.ack can finally tell a hit from a
miss. FleetMcp.ack returns an error when the boolean is false, naming
fleet_poll{coordId} for held peer mail, which has no route through this
call. MessageService.ackReply propagates the boolean; drainReplies keeps
ignoring it (its own javadoc already documents that loss window as
deliberate). Updated the tool schema's target description to match.

Rewrote FleetMcpTest's ack tests to publish a real message before
asserting success, and added tests for a never-queued id and a coord-id
target, both now erroring. Added boolean assertions to
InMemoryReplyInboxTest's existing ack cases.
2026-09-10 16:44:08 +07:00
Dai Ha 3f036b2a62 fleetd #444: pin the place()-to-spawn() window PlacementDecision closes
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m53s
Add a test that resolves place(role) while nothing is quarantined,
then quarantines the resolved profile's credential BEFORE spawning
against the held PlacementDecision. CompositePeerLauncher.spawn(req,
decision) must still honor the decision and land on the quarantined
profile, since it never re-runs the explicit-profile enforce* checks.

Verified the test kills the regression: with the override's body
replaced by the re-entering spawn(req.withProfile(...)) form, this
exact test goes RED with a PlacementException naming the now-
quarantined profile; restored, the full suite is green (1575 tests).

Also documents on PeerLauncher's default spawn(req, decision) that a
launcher routing across more than one profile MUST override it,
naming the four enforce* checks the default's re-entry re-applies.
2026-09-10 16:41:52 +07:00
ltms 82fae94c55 Merge #445: pin every startup report call in Fleetd.main (fleetd #442)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m59s
Test written by a worker whose backend died before it could report; evidence
re-run by the lead against the merged tree.

Verified: merge of current main clean (0 conflicts); full build 1574 tests,
0 failures, BUILD SUCCESS, 0 compile errors; control green; deleting each of
reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and
reportExhaustedPatternGap from main() is KILLED by
mainReportsEveryStartupGapBeforeValidationAborts.
2026-09-10 11:30:31 +02:00
Dai Ha b1f34c2e6b fleetd #442: drop the unused java.util.List import
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m8s
The new test never names List — only ListAppender, which has its own
import. An unused import is an IDE warning, and this repo treats
warnings as gates. No behaviour change: FleetdStartupReportTest still
runs 1 test, 0 failures, BUILD SUCCESS, 0 compile errors.
2026-09-10 16:29:56 +07:00
ltms 3f807d9f1b Merge #443: derive coordinator.heldDurable from queue durability + ack mode (fleetd #440)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m35s
Found by the fleet01 lead reviewing #438 after I had merged it. Verified
independently before merging.

The implementation choice is the load-bearing part: LeadMailbox.own() now
assigns queueDeclare's durable flag and basicConsume's autoAck flag to named
locals, passes those SAME locals into the two real AMQP calls (:203/:204), and
derives heldDurable from them (:207). So the reported fact cannot drift from a
duplicate constant - a mutation to either call's argument moves the behaviour
and the report together. LeadChannel.heldDurable() is abstract, so a future
implementer gets a compile error rather than a silent default.

My own battery, merged tree, control green, tree restored clean:
- M1 revert to the literal true -> KILLED by
  FleetMcpTest.listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable
- M3 heldCount forced to 0 (a half the worker did not touch) -> KILLED by
  FleetMcpTest.listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero
- M2 break the derivation itself -> SURVIVED under plain clean install, exactly
  as the worker reported. LeadMailbox needs a real broker, so the only test that
  reaches the real queueDeclare/basicConsume is @Tag("contract"), excluded from
  the default build. The worker ran that arm with -Pcontract and got 3 reds
  including its own new test. Pre-existing structural limit of this class, not
  introduced here, and the worker flagged it rather than hiding it.

Full build, my own run: Tests run: 1573, Failures: 0, Errors: 0, Skipped: 0 -
BUILD SUCCESS, 0 compile errors.

Not blocking, noted for a possible follow-up: one boolean over two independent
facts cannot say WHICH fact was lost. The merged field is still strictly better
than the literal it replaces, because it can now go false at all.
2026-09-10 11:21:41 +02:00
Dai Ha e70263062c fleetd #442: pin startup report calls
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m36s
2026-09-10 14:41:10 +07:00
ltms 1a1e586b62 Merge #433: carry the PlacementDecision instead of re-resolving the profile (fleetd #425)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m47s
Round 4 pins the fix. Verified independently before merging.

My own battery in the worker's tree (control green, tree restored clean):
- M1 revert SessionManager:677 to launcher.spawn(spawnReq) -> KILLED by
  SessionManagerTest.acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedThe
  WorktreeForUnderARotatingPolicy. This is the ticket's own deliverable and it
  survived 186 tests in round 3.
- M3 drop the withProfile stamping in the 2-arg spawn -> KILLED by 3 tests.
- M2 make the 2-arg spawn re-enter the refusing branch -> SURVIVED, but it is a
  near-equivalent mutant, not a gap in this work. Post-#435 all four enforce*
  conditions are ones the routing branch already filtered on, so the two paths
  differ only if placement state moves between place() and spawn(). Filed
  separately.

Full build, my own run this turn, whole log redirected and grepped:
Tests run: 1572, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile
errors. Branch already contains current main.

Read the src/main diff. The new test uses roundRobin() (stateful) and asserts
agreement between the overlay profile and the spawned profile, rather than a
hardcoded name, which is the right shape - select() is stateful, so two calls
disagree by design.
2026-09-10 09:35:02 +02:00
Dai Ha c16d118f09 fleetd #440: derive coordinator.heldDurable from queue durability + ack mode
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m36s
FleetMcp.coordinatorView wrote heldDurable as a literal true, so a change
that broke either the durable queue declare or the manual-ack consume in
LeadMailbox would leave the field, and the full suite, green.

- LeadChannel gets a new heldDurable() method: the conclusion of a durable
  queue declare AND a manual-ack consumer, derived by the implementation
  from what it actually did, never asserted.
- LeadMailbox.own() captures the exact booleans it passes to
  queueDeclare/basicConsume and stores their conjunction; heldDurable()
  returns it.
- FleetMcp.coordinatorView now reads channel.heldDurable() instead of a
  literal; updated the javadoc to say where the fact comes from.
- FakeLeadChannel gets a heldDurable field (default true) + withHeldDurable
  setter so FleetMcpTest can prove the field goes false.
- FleetMcpTest: new test asserts heldDurable:false when the channel says so.
- LeadMailboxTest (contract, real broker): new test asserts heldDurable()
  true against a real LeadMailbox. Verified by hand that flipping own()'s
  autoAck local to true turns this test (and two pre-existing redelivery
  tests) red, and restoring it turns them green again.
2026-09-10 14:31:44 +07:00
Dai Ha 4b10d02207 fleetd #425 rework round 4: mutation-pinning test for the dropped PlacementDecision
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m3s
SessionManager.acquireWithWorktree's unqualified branch must carry the
PlacementDecision it already resolved via launcher.place() into
launcher.spawn(spawnReq, decision) rather than re-deriving it through a
blank-profile launcher.spawn(spawnReq). Every existing test in this file uses
PlacementPolicies.fixed(), which answers select() the same way on every call,
so dropping the decision (handle = launcher.spawn(spawnReq);) was invisible:
186 tests stayed green under that mutation.

acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy
uses PlacementPolicies.roundRobin() instead — deterministic AND stateful, so
two select() calls on the same policy instance disagree (index 0 then index 1
across a two-profile pool). It asserts AGREEMENT between the profile the
worktree's parity overlay was provisioned for and the profile the member
actually spawned on, never a hardcoded expected profile name.

Verified as a real mutation, not a no-op: applying the exact mutation
(handle = launcher.spawn(spawnReq);) turns it red — expected [b.mcp.json]
but was [a.mcp.json] — and reverting turns it green again. Full build:
1572 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-09-10 14:24:14 +07:00
Dai Ha c5fbfdbf4a Merge main into #425 rework branch (brings #438 held-peer-mail read) 2026-09-10 14:11:37 +07:00
ltms 12cff28abb Merge #438: let a lead read its own held peer mail, primary-only (fleetd #421)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 2m6s
fleet_poll{coordId} peeks this daemon's own held lead-to-lead mail and
returns full bodies without acking. New Authz.Action COORD_READ, primary
only — not the architect, which holds READ today. pollAction is now
argument-derived over both target and coordId.

heldView and HELD_PREVIEW_MAX_CHARS are untouched: fleet_list stays a
cheap always-safe scan, and the full read is a separately authorized call.

Verified by me, not taken from the report: merged tree builds 1562 green
(main 1555 + 7 new methods), 0 compile errors, no merge conflicts.
Mutation battery on lines the worker did NOT mutate, control 111 green:
  - COORD_READ widened to the architect  KILLED (AuthzTest + FleetMcpAuthzTest)
  - self-coord-id guard removed          KILLED (FleetMcpTest)
  - full body swapped for the preview    KILLED (FleetMcpTest)

The third mutation targets the ticket's own deliverable, and it is pinned.
Authz.permits has no default, so a new action is a compile error rather
than a silently unhandled case.

Follow-up filed as #439: fleet_list's coordinator row is still READ-gated,
so a worker sees peer coord-ids and 80-char previews of lead-to-lead
bodies. Pre-existing; the implementer flagged it and left it alone.
2026-09-10 09:08:17 +02:00
Dai Ha 77a6a7142e Merge main into #421 branch (brings #434 model-gate observability and #436 fixed-placement cap)
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m48s
2026-09-10 14:06:34 +07:00
Dai Ha 9f3671b801 fleetd #425 rework round 3: rewrite prose after #435 made fixed honour maxLoad
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m56s
fleetd #435 (merged to main) made FixedPlacementPolicy evaluate maxLoad during automatic
selection, the same way weighted/round-robin already did. Six comments across
CompositePeerLauncher.java, PeerLauncher.java, PlacementDecision.java, and SessionManager.java
justified round 2's place()/spawn(req, decision) mechanism by saying fixed "deliberately never
evaluates maxLoad" — that claim is now false, and needed restating, not just deleting.

The honest case after #435: the two-path shape (routing branch falls through an excluded
candidate; explicit-profile branch refuses on it) is still real and still deliberate — an
operator who names a profile should get a refusal, not a silent substitution. What round 1 got
wrong, and what round 2 still needs to prevent, is turning a fall-through into a refusal by
accident: resolving a name via place() and then feeding it back to spawn(SpawnRequest) as an
explicit profile. Before #435 that accident was reachable through maxLoad specifically, because
fixed never evaluated it; #435 closed that specific gap, so a PlacementDecision can no longer be
at-cap in the first place. What survives as the justification for spawn(req, decision): it never
re-evaluates a condition place() already decided, and it closes the window between that decision
and the spawn in which the underlying state could otherwise move — not a failure #435 already
prevents.

Re-measured the sibling paths this round exists to keep in agreement (one profile at maxLoad: 1,
liveCount pinned at 1, PlacementPolicies.fixed(), unqualified spawn): both the with-worktree and
without-worktree paths now throw the identical PlacementException — "worker profile 'a' is at
maxLoad (1 live >= 1 cap), and no available candidate remains" — closed upstream by #435, at
place()/select(), before either path ever reaches a spawn call. The observable asymmetry this PR
was filed to fix is gone; what remains is the structural argument above.

No behavior change: place()/PlacementDecision/spawn(req, decision) are untouched, and
FixedPlacementPolicy/PlacementPolicyUtil are taken wholesale from main's merge.
2026-09-10 14:06:10 +07:00
Dai Ha 84034b34d1 Merge remote-tracking branch 'origin/main' into worker/425-rework-placement-resolve-c58ba1-9 2026-09-10 13:56:54 +07:00
Dai Ha 1e9b2c9b7e fleetd #421: let a lead peek its own held peer mail, primary-only
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m36s
fleet_list truncated held lead-to-lead messages to an 80-char preview with
no way to read the full body, and fleet_poll{target} drained the wrong
inbox (a worker's reply queue, not the coordinator mailbox) -- it silently
returned []. fleet_ack would have destroyed the message unread.

Add a non-destructive read: fleet_poll{coordId} peeks (never acks) this
daemon's own held mail via LeadChannel.peek(). The coordId must equal the
caller's own selfCoordId -- passing a peer's id is refused with a reason,
instead of repeating the original silent-[] confusion.

This is authorization-sensitive: mapping it to the existing READ action
would let any worker read every peer lead's mail in full. READ's openness
rests on "the roster carries no secrets" (Authz.java), which does not hold
for lead-to-lead coordination bodies. Added Authz.Action.COORD_READ,
primary-only (not even the architect, which holds READ today), and made
pollAction's signature depend on both target and coordId so every call
site states explicitly what it passes.

Also fixes fleet_list's "pending: 0" trap: mailbox.pending only counts
broker-ready messages, so a healthy held mailbox reads as empty. Added
heldCount/heldDurable beside held[] so the durability fact isn't implied
only by reading the code.

Mutation-tested: pollAction's COORD_READ->READ mapping, the peek->ack
substitution, the 80-char preview cap widened to 81, and the Authz case
widened to include caller.isWorker() -- each breaks exactly its matching
test and nothing else. The first attempt at the preview-cap test used a
homogeneous "x"*200 body, which a widened cap slipped through unnoticed
(contains() found a shifted match); replaced with a sentinel character at
index 80 to actually pin the boundary.

Updates CLAUDE.md's intent->tool table for fleet_poll's new coordId
semantics, per this repo's own "prompt is part of the product" rule.
wiki/ is a submodule and not committable from a worker's worktree --
wiki-bound content is in the PR body instead.
2026-09-10 13:54:50 +07:00
ltms 5d422f85fa Merge #436: make fixed placement honour maxLoad (fleetd #435)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m37s
fixed was the one automatic policy that ignored maxLoad, and it is the
default for an absent placement: key. It now gates on the same shared
atCap predicate weighted and round-robin use, and falls through to the
next candidate rather than refusing.

Verified here: 1555 tests green (main was 1548, +7 new methods), 0
compile errors. Mutation battery, control 113 green:
  - atCap boundary >= -> >        KILLED (15 tests, across all 3 policies)
  - drop cap check in fixed walk  KILLED (3 tests)
  - weightExcluded -> false       KILLED (2 pre-existing tests)

Neither live host is affected today: both run placement: weighted (Mac
fleetd.yaml:210, fleet01 fleetd.yaml:172), and weighted already skipped
at-cap candidates via PlacementPolicyUtil.available. This aligns fixed
with the other policies and with maxLoad's documented contract.
2026-09-10 08:54:18 +02:00
Dai Ha ed2027b202 fleetd #435: make FixedPlacementPolicy honor maxLoad
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m54s
FixedPlacementPolicy (the default placement policy) never consulted
maxLoad, so an at-cap default was chosen anyway on every unqualified
spawn -- the cap was advisory, not enforced, for the one policy every
config uses by default. weighted/round-robin already gated on it via
PlacementPolicyUtil.available().

Extract the "at cap" predicate into PlacementPolicyUtil.atCap(ctx, c)
so all three policies share one definition, and consult it at both of
FixedPlacementPolicy's filter sites (the default fast path and the
candidate walk), mirroring the existing weightExcluded pattern. An
at-cap default now falls through to the next candidate instead of
refusing the spawn -- only when every candidate is unusable does the
policy still throw, naming the cap in the message. Update the class
javadoc (five exceptions -> six) and the reason-priority comments to
match CompositePeerLauncher's explicit-spawn order (quarantine,
cooling off, max load, model-off).
2026-09-10 13:46:06 +07:00
Dai Ha 6b0a99b2b7 fleetd #425 rework round 2: stop routedProfileFor's caller re-entering the throwing branch
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m34s
Round 1 closed quarantine/cool-off/model-off routing for acquireWithWorktree by resolving the
profile through routedProfileFor(role) and handing that name back to launcher.spawn(SpawnRequest)
as an EXPLICIT profile. That re-resolution has a cost the lead measured directly: naming a
profile explicitly makes CompositePeerLauncher.spawn take its THROWING branch (enforceMaxLoad
included), while the routing branch a blank spawn takes never calls enforceMaxLoad at all, and
FixedPlacementPolicy (the default) deliberately never evaluates maxLoad during automatic
selection. So an at-cap pool-first profile that placement itself would have picked for a plain
unqualified spawn could die at enforceMaxLoad one call later, purely because the worktree path's
route to the spawn passed through an explicit profile name — a new failure a worktree-less
unqualified spawn never hits.

This closes the two-path shape instead of moving it: PeerLauncher gains place(role), returning an
opaque PlacementDecision, and spawn(req, decision), which honors that decision through the SAME
routing branch a blank spawn uses — no enforce* check is newly applied. SessionManager.
acquireWithWorktree now keeps the PlacementDecision from place() and hands it to
spawn(req, decision) for an unqualified request, instead of re-resolving through an explicit
profile name. An explicitly-named profile is unaffected: it still goes through spawn(req) and its
throwing branch, exactly as before.

Also corrects the acquireWithWorktree comment's false claim that round 1 "loses nothing else" —
maxLoad was lost too, as a new hard failure, not a retry. The comment now names it explicitly.

Kept the four round-1 tests (still pass — routedProfileFor now just delegates to place()). Added
one class asserting the invariant itself: an unqualified spawn on a maxLoad-capped profile must
land the same outcome with and without a worktree, asserting on the pair rather than a hardcoded
direction, so it stays correct however fleetd #435 (not this ticket) resolves whether maxLoad
should gate an unqualified spawn at all.
2026-09-10 13:42:06 +07:00
ltms 6b7caba248 Merge #434: make the model gate's own state observable
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m49s
Verified in the worker's tree at 7fd914d: 1544 tests green (base 1535 + 9),
0 compile errors, unpiped mvn clean install. merge-tree against f8b0d42 reports
no conflicts and the two sides share no files.

Read all four production diffs. The design is right: one ModelGateState record
carrying both "armed" and "off" from a single models0() read, so the startup
log line, fleet_profiles' modelGateArmed and the spawn gate cannot disagree —
the fleetd #404 lesson applied properly. disabledModels() now delegates to it
rather than being a second independent read.

Mutated the two subtlest lines, with a control in the same script and the
changed line echoed back with its number:

- modelGateState(): sentinel identity check replaced by the naive
  "m.offIds().isEmpty()" -> 2 failures. FleetProfilesModelGateStateTest
  .modelsBlockWithNothingOffReportsGateArmedAndZeroOff:86 and
  CompositePeerLauncherTest.modelGateStateIsHotReloadedThroughARealConfigRef
  :1519, both "expected: <true> but was: <false>". So the one state this
  ticket exists to expose — a block present with nothing off — is pinned.
- notConfigured() returning a non-empty off set, breaking the invariant the
  record's javadoc states but does not enforce -> 3 failures, including
  disabledModelsIsEmptyWithNoModelsConfigured:1581. So the invariant is
  observable even though the constructor does not check it.

Control run unmutated: 1544 green.

Follow-up on me, not a merge blocker: fleet_profiles gains an operator-visible
field, so this needs a wiki/11-Features.md entry. Workers cannot commit the
wiki submodule, so I am adding it.
2026-09-10 08:32:54 +02:00
Dai Ha f8b0d42a5c fleetd #431 follow-up: profileForSlot's javadoc named a caller that does not exist
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m53s
The javadoc said "what the spawn lifecycle reads". Nothing in src/main calls
profileForSlot at all, in either the ".profileForSlot(" or the
"::profileForSlot" form. My own #431 ticket text repeated that sentence as a
fact and ranked the three accessors by it, and the #432 worker copied it into
the test file's comment and one assertion message. Corrected in all three
places; the ticket correction is posted on #431.

What the spawn lifecycle actually reads for an architect's profile is the
SlotReservation that reserve() returns — SessionManager.java:225,
"reservation == null ? profile : reservation.profile()".

The corrected ranking, measured rather than read off the javadoc:
- nameForSlot is wired, at CallerResolver.java:137 (method reference, which is
  why a ".nameForSlot(" grep missed it)
- isSlot is reached through bind, called at MemberRegistry.java:235 and :376
- profileForSlot has no caller at all

Prose only. No behaviour change.
2026-09-10 13:25:19 +07:00
ltms d1e7d71eee Merge #432: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
Verified in my own tree at a196d34: 1539 tests green, 0 compile errors, 0 files
changed under src/main (test-only, as reported).

Mutated two halves the worker's own proof did not cover, with a control in the
same script and the changed line echoed back each time:

- roleForSlot returning ARCHITECT for any configured slot (flatten is
  role-blind) -> 2 failures, incl. CallerResolverTest
  .aBoundNonArchitectSlotStillResolvesAsAWorker:454. So the role-blind flatten
  cannot grant ARCHITECT through a non-architect pool; that was already pinned.
- nameForSlot parsing the key suffix instead of reading config -> 1 failure,
  MemberRegistryLiveTest.nameForSlotReflectsANameChangedByReload:255. The new
  test has teeth beyond the freeze the worker ran.

Control run unmutated: green.

The ticket's severity ranking was wrong and I corrected it on #431. profileForSlot
has no caller in src/main in either the "." or "::" form, so its javadoc ("what
the spawn lifecycle reads") names a caller that does not exist; nameForSlot is
wired at CallerResolver.java:137; isSlot is reached through bind, at :235 and
:376. A follow-up commit fixes the test prose that repeated my claim.
2026-09-10 08:24:26 +02:00
Dai Ha 7fd914df1a fleetd #422 follow-up: make the model gate's own state observable
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m53s
PeerLauncher.disabledModels() reported an empty set both when no
models: block exists and when a block exists with nothing off, so
fleet_profiles/GET /profiles and the startup log could not tell an
inert gate from an armed one reporting zero. Add
PeerLauncher.ModelGateState (configured + off), a modelGateState()
default method disabledModels() now delegates to, and a
CompositePeerLauncher override that reads models0() once and
distinguishes the null-supplier case (no models: block) from a real,
config-supplied block via identity against the NO_MODELS_CONFIGURED
sentinel — reusing the exact accessor the spawn gate itself reads, per
the fleetd #404 lesson.

Wires the state into a new startup log line (Fleetd.modelGateCoverageLine)
and a new modelGateArmed field in FleetMcp.profilesView, reported
unconditionally alongside the existing modelsOff set.
2026-09-10 13:18:52 +07:00
Dai Ha b066eb1903 fleetd #425 rework: resolve acquireWithWorktree through real placement, not a blind pool-first read
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m13s
e1d7dde (PR #430) kept two good fixes and one regressed one. Kept: (1)
CompositePeerLauncher.defaultProfile() delegating to
defaultProfileFor(MemberRole.DEV) so fleet_profiles' "default" tracks a live
reload, and (2) PeerLauncher.defaultProfileFor(MemberRole). Redone:
acquireWithWorktree's profile pre-resolution.

The regression: acquireWithWorktree pre-resolved via
launcher.defaultProfileFor(memberRole), which just returns the role pool's
FIRST entry, blind to quarantine/cool-off/model-off. That name was then
passed to launcher.spawn as an EXPLICIT profile, which takes
CompositePeerLauncher.spawn's THROWING branch (enforceNotQuarantined /
enforceMaxLoad / enforceModelEnabled) instead of the ROUTING branch a blank
profile gets. So a quarantined or model-off pool-first profile turned a
routine unqualified spawn into a hard PlacementException -- undermining
fleetd #429's "the fleet keeps working when a model is turned off"
guarantee for every worktree spawn.

Fix: add PeerLauncher.routedProfileFor(MemberRole), the profile an
unqualified spawn of that role would actually be routed to right now --
same candidate list, same quarantined/coolingOff/modelOff filtering, same
PlacementPolicy spawn() itself consults. CompositePeerLauncher implements it
by extracting spawn()'s context-building into a shared private
placementContextFor(role, unreachable), so spawn() and routedProfileFor()
can never disagree about which conditions apply to which candidate.
acquireWithWorktree now calls routedProfileFor once and reuses that name for
repoRoot, parityOverlay, and the spawn -- the fleetd #425 defect (the three
disagreeing) stays fixed, now on the routed path instead of the blind one.

An explicit profile named by the caller is untouched -- it still hits the
throwing branch, which is correct for an operator override.

Trade-off carried over from e1d7dde, now precisely scoped: an unqualified
worktree spawn still loses CompositePeerLauncher's cross-candidate retry on
a live PeerUnreachableException (a transport failure at spawn time, which
placement cannot see in advance) -- but NOT the quarantine/cool-off/
model-off routing, which routedProfileFor already resolved before spawn
ever runs. Accepted: a worktree provisioned for the wrong backend is worse
than a spawn that fails cleanly and can be retried by the caller.

Tests: CompositePeerLauncherTest gains
routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy and
...ModelOff..., both under PlacementPolicies.fixed() (the default policy,
not weighted() -- the previous round's tests all used weighted() and never
exercised FixedPlacementPolicy's own inline filter, which is exactly what
regressed). SessionManagerTest gains
acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile, proving
repoRoot/parityOverlay/spawn agree on the ROUTED profile, not just the
pool-reordered-by-reload one the existing #425 tests already covered.
2026-09-10 13:10:30 +07:00
Dai Ha a196d34455 fleetd #431: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m7s
#424 made MemberRegistry.slots() re-read fleet: on every call, but only
roleForSlot was tested against a real reload. profileForSlot, isSlot and
nameForSlot all have the same live-read line and none was pinned — proved by
freezing each to a construction-time snapshot and watching the full suite
stay green.

Adds 4 tests to MemberRegistryLiveTest, each driving a real ConfigRef.reload()
against a @TempDir config file (never two frozen registries compared in
memory, which would test the constructor instead of the reload):
- profileForSlotReflectsAProfileChangedByReload
- isSlotStopsReportingASlotRemovedByReload / isSlotStartsReportingASlotAddedByReload
- nameForSlotReflectsANameChangedByReload

No production change. profileForSlot has no call site anywhere in src/main
yet, so there is no spawn-lifecycle seam to drive the test through beyond the
accessor itself.
2026-09-10 13:06:56 +07:00
Dai Ha 051d320ea0 fleetd #425: fleet_profiles' default and worktree provisioning must read live placement
fleet_profiles' "default" was CompositePeerLauncher.defaultProfile, a value
frozen at construction from cfg.effectiveDefaultProfile(). An unqualified
fleet_spawn instead resolves the dev pool live via defaultProfileFor(DEV) on
every call, so reordering fleet.developers and reloading changed where a
spawn landed without ever changing what fleet_profiles reported.

- CompositePeerLauncher.defaultProfile() now delegates to
  defaultProfileFor(MemberRole.DEV) -- the same live, reload-aware pool read
  placement already uses -- falling back to the frozen field only when no
  profiles are configured at all.
- PeerLauncher gains a default defaultProfileFor(MemberRole) method so a
  generic PeerLauncher reference can ask for a role's live default; the
  default implementation delegates to defaultProfile() for launchers with no
  pool concept of their own.
- SessionManager.acquireWithWorktree resolved a profile via
  launcher.defaultProfile() (DEV-only) to provision repoRoot/parityOverlay,
  then spawned with the original (possibly blank) profile, which re-resolves
  independently through placement -- for any non-DEV role, or across a config
  reload between the two reads, the two resolutions could disagree and
  provision a worktree for a profile the member never runs on. Fixed by
  resolving once, through defaultProfileFor(the caller's actual role), and
  reusing that same resolved name for repoRoot, parityOverlay, and the spawn
  itself. Trade-off: this path now spawns with an explicit profile rather
  than a blank one, so it loses CompositePeerLauncher's cross-candidate retry
  on PeerUnreachableException -- accepted because a worktree provisioned for
  the wrong backend is worse than a spawn that fails cleanly and can be
  retried.

Tests: CompositePeerLauncherTest (live dev-pool reorder + empty-pool
fallback), FleetProfilesLiveDefaultTest (drives FleetMcp.profilesView
directly), SessionManagerTest (worktree overlay follows a reorder, and a
non-DEV role's worktree spawn uses that role's pool, not DEV's).
2026-09-10 12:58:20 +07:00
Dai Ha 766772763f Merge #428: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m32s
fleetd #424. MemberRegistry.slots() now re-reads fleet: through a supplier,
so removing an architect slot demotes the bound pane on its very next request.
The boundEntries cache and entryFor() fallback from the first round are gone:
the slot OCCUPANCY (terminalToSlot) survives a reload, the ARCHITECT role does
not. That split was my own ticket wording's fault -- I asked for a test that a
bound architect "survives the rebuild", which conflated the binding with the
privilege.

Conflict resolved by hand in ConfigRef.java: #422 (models:) and #424
(architects) both rewrote the same Hot bullet. Kept both.

Also corrected two claims #424's own second commit left stale -- 7f672f0
reversed the behaviour but never touched ConfigRef, whose whole job is to tell
the operator what a reload does:
  - the Hot bullet said MemberRegistry's rule "keeps a live session's identity
    even after its slot is removed from config"
  - the reload-report comment said "only a NEW bind is refused"
Both now say what the code does: removal revokes ARCHITECT on the next
request, and only the slot occupancy survives.

Verified by the lead: 1535 tests, 0 failures, 0 compile errors, BUILD SUCCESS
on the merged tree.

Mutation of three halves the worker's own proof did not cover -- profileForSlot,
nameForSlot and isSlot each pointed at a frozen snapshot taken at construction
(live readers 5 -> 4, each mutation naming its method and line). All three
PASSED at 1535. The ticket's own fix is well pinned; these three sibling live
reads are not. Follow-up filed.
2026-09-10 12:54:35 +07:00
ltms eab8185d7b Merge #429: enforce the models.allow on/off gate at spawn, hot — including under fixed placement
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m33s
fleetd #422 + follow-up. Verified by the lead: 1524 tests green, 0 compile errors.

Mutation proof of three halves the worker's own proof did not cover:
- modelOffProfiles() -> Set.of() (the feed into ctx.modelOff): kills 3, incl. fixedPlacementSkipsAnOffModelProfileToo
- enforceModelEnabled() -> no-op (the explicit-spawn gate): kills 3, incl. the real-ConfigRef hot-reload test
- disabledModels() -> Set.of() (the reporting accessor): kills 1, alone
2026-09-10 07:44:35 +02:00
Dai Ha 7f672f0fb8 fleetd #424: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 2m3s
Correction to the #424 fix in PR #428: the ticket asked to revoke a removed
architect slot, but the previous change (boundEntries) kept BOTH the binding
and the ARCHITECT privilege alive for an already-bound session after its slot
left config. That left the ticket's actual headline defect half-open.

The corrected rule: config governs both what may be bound next AND what a
bound slot still grants. Removing a slot now demotes its bound session to
worker on the very next request (roleForSlot/nameForSlot read slots() with no
cache, so CallerResolver.resolve falls through to Principal.worker(...)). The
terminalToSlot binding itself is untouched by a reload, on purpose: dropping
it would double-book the slot key and break unbind's compare-safe contract.

- Delete boundEntries and entryFor; profileForSlot/roleForSlot/nameForSlot/
  isSlot all read slots() directly, live, with no cache.
- Rewrite the class doc's binding rule for the corrected semantic.
- Replace the old "survives removal" test with anArchitectAlreadyBoundToASlotIsDemotedByReload,
  asserted through a real CallerResolver.resolve (not the roleForSlot seam),
  plus two tests for what must NOT change: the binding still occupies the
  slot after removal (a second terminal cannot claim it, even once the slot
  returns to config), and unbind still succeeds for the original terminal.

Verified snapshot()/CallerResolver.members() need no change: snapshot() only
ever reported raw terminalToSlot occupancy, and CallerResolver.members() has
no production caller.
2026-09-10 12:29:35 +07:00
Dai Ha e2801b9bbc fleetd #422 follow-up: gate FixedPlacementPolicy on modelOff too
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName
returns it for an absent/blank name) and it built its own inline candidate
filter instead of calling PlacementPolicyUtil.available(). That filter checked
quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an
unqualified fleet_spawn on any fleet without an explicit placement: policy
could still land on a profile whose model the operator turned off.

- Add ctx.modelOff() to both filter sites: the default-profile fast path and
  the fallback walk over ctx.candidates().
- Add a modelOff refusal reason to the default-profile reasons list, worded as
  an operator decision ("turned off in models.allow"), matching
  enforceModelEnabled. Quarantine and cooling off still take priority when a
  profile is also model-off, matching CompositePeerLauncher's explicit-spawn
  check order.
- Update the two stale "excluded from automatic selection" messages to name
  model-off, consistent with PlacementPolicyUtil.emptyException.
- Update the class javadoc: four exceptions -> five, with a new bullet for
  model-off (fleetd #422).

Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path),
fixedFallbackWalkSkipsModelOffCandidate (fallback walk),
fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and
fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest
gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of
the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under
PlacementPolicies.fixed(). The two existing weighted()-based tests are
untouched.
2026-09-10 12:25:31 +07:00
Dai Ha ea02c7b248 fleetd #422: enforce the model allow-list on/off state at spawn, hot
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 1m30s
Ships the two halves left out of the earlier allow-list ticket in one PR,
since apart they are inert: a gate with no flag always allows, and a flag
nothing reads does nothing.

- FleetConfig.Models.ModelEntry gains `enabled` (default on; absent/true =
  on, false = off). Turning a model off never removes it from `allow:` —
  validateModels() checks membership only, so an off model stays valid
  config and a still-configured profile naming it does not refuse reload.
  Models.offIds() is the one live accessor both the gate and the status
  report read.
- CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent
  spawn-refusal reason (operator intent) — never layered onto
  BackendQuarantine/BackendOutagePolicy, which are backend-reported outage.
  Wired into the explicit-profile branch. modelOffProfiles() feeds the same
  off-model exclusion into PlacementContext for unqualified spawns via
  PlacementPolicyUtil (a new modelOff set, counted into its own bucket in
  emptyException so "all off" is named as the cause, not generic).
  Both read models0(), a live Supplier<FleetConfig.Models>, so a reload
  reaches the very next spawn — no restart.
- PeerLauncher.disabledModels() (default empty) lets fleet_profiles/
  GET /profiles report off models by reading the exact same accessor the
  gate reads (the fleetd #404 lesson: a status field must read the source
  the behaviour reads).
- ConfigRef: `models` reclassified from deferred to hot-excluded — nothing
  about it is baked into a startup-built object anymore; membership is
  re-validated in full on every reload via validateAll(), and the on/off
  half is read live everywhere. Tally: 5 cold, 13 deferred, 3 split, 4
  hot-excluded (25 total). ConfigRefTopLevelCoverageTest and
  ConfigRefTopLevelReportingCoverageTest updated with no new exclusion
  added just to force green.

Tests: FleetConfigTest (old-style fixture stays on; an off model is still
valid config; one model disables every profile naming it),
CompositePeerLauncherTest (explicit refusal wording distinct from
quarantine/cool-off; unqualified spawn skips an off candidate and names
model-off when every candidate is off; a real ConfigRef.reload() proves
the hot path; disabledModels() matches the gate).
2026-09-10 12:08:18 +07:00
Dai Ha ce74e164c6 fleetd #424: make architect-slot identity checks read fleet.architects live
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m32s
MemberRegistry used to flatten fleet.architects into an unmodifiable map at
construction, so removing (revoking) an architect slot from config never
took effect: reserve()/requireSlotFor() kept granting spawns against the
frozen snapshot forever, while ConfigRef told the operator "already
applied" for the wrong consumer.

- MemberRegistry gains a live constructor (MemberRegistry.live(Supplier))
  that re-flattens fleet.architects/developers/reviewers on every
  slots()/slotsFor() call, so reserve() and requireSlotFor() (which both
  read through slotsFor) govern the NEXT spawn with no restart. The frozen
  single-arg constructor is kept for tests and fixed/code-built configs.
- Binding rule: config governs what may be bound next; it never
  retroactively unbinds a live session. A slot removed from config while a
  terminal is bound to it keeps that binding. To keep the bound terminal's
  IDENTITY too (CallerResolver.resolve reads roleForSlot/nameForSlot on
  every request), every successful bind now caches the slot's Entry into a
  new boundEntries map; roleForSlot/nameForSlot/profileForSlot/isSlot fall
  back to it when the slot is no longer live, and unbind clears it in the
  same critical section it clears the binding.
- Fleetd.java now wires MemberRegistry.live(() -> config.get().fleet())
  instead of the frozen constructor.
- ConfigRef: corrected the fleet.leaders split-key message and the class
  doc's Hot bullet — architects is now hot for two independent consumers
  (CompositePeerLauncher for placement, MemberRegistry for identity), not
  only the one the message used to name. An architects-only edit still
  reports nothing beyond "config reloaded", which is now honest since the
  key really is fully hot for both consumers.

Added MemberRegistryLiveTest: real ConfigRef.reload() against a @TempDir
file, both directions (slot removed / slot added) for requireSlotFor and
reserve tested separately, plus a bound-architect-survives-removal test
that checks the binding AND the identity (roleForSlot/nameForSlot).

Mutation-tested: reverting requireSlotFor to a frozen snapshot fails
requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload and its mirror;
reverting reserve the same way fails the two reserve tests; dropping the
boundEntries fallback fails the survives-removal test's roleForSlot
assertion. All three restored before commit.

mvn clean install: Tests run: 1510, Failures: 0, Errors: 0, Skipped: 0 —
BUILD SUCCESS.
2026-09-10 12:05:18 +07:00
ltms e60f892efd Merge pull request 'fleetd #415: split coverage() feature-state wording by pattern fallback semantics' (#423) from worker/415-coverage-wording-2cbf9c-5 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
2026-09-10 06:57:33 +02:00
Dai Ha ce05886831 fleetd #415: pin which UnsetMeaning Fleetd pairs with which pattern key
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m54s
Review found a gap: the earlier tests all called CompletionResolver.coverage()
directly, supplying the UnsetMeaning themselves — proving the enum's wording,
never that Fleetd's two call sites pair the right meaning with the right key.
Swapping the two UnsetMeaning arguments at those call sites (recreating #415's
defect with exhaustedPattern and errorPattern exchanged) compiled with 0 errors
and left all 1506 tests green.

Extract the two coverage-line call sites out of main() into package-private
static factories (Fleetd.exhaustedPatternCoverageLine /
errorPatternCoverageLine), the same pattern already used for capacitySource
and worktreeBranchLookup. Add FleetdPatternCoverageLineTest, which calls both
factories directly and asserts the actual wording each produces for the same
empty-coverage input, including that the two differ.

Also recorded the swap-mutation measurement (0 errors, 1506 green) in
UnsetMeaning's javadoc so a future reader does not delete the new test as
redundant with CompletionResolverTest.
2026-09-10 11:54:02 +07:00
Dai Ha be123d0ac7 fleetd #415: split coverage() feature-state wording by pattern-key fallback semantics
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m4s
CompletionResolver.coverage() measured pattern coverage (how many profiles set
a key) but its 'off' wording read as feature state. That is false for
errorPattern: an unset errorPattern still runs the classification against the
built-in BACKEND_ERROR pattern (CompletionResolver.java:84), so the empty case
is not off.

Add CompletionResolver.UnsetMeaning (OFF / BUILT_IN_DEFAULT), a required
parameter every coverage() call must supply — no defaulted overload, so a
future third pattern key cannot compile without stating what unset means for
it. Fleetd.java now passes UnsetMeaning.OFF for exhaustedPattern (no fallback
exists) and UnsetMeaning.BUILT_IN_DEFAULT for errorPattern.

Tests: updated the three existing empty/full/partial cases to pass the new
parameter, corrected the one test that pinned the old (wrong) errorPattern
wording, and added a test that asserts the same empty-coverage input produces
different wording for the two keys.
2026-09-10 11:40:32 +07:00
ltms 2d09c8b027 Merge pull request 'fleetd #418: barrier the throw-path push-loop test on state decide() reads' (#419) from worker/418-588283-3 into main
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m39s
2026-09-10 06:29:23 +02:00
ltms ed54f0224e Merge pull request 'fleetd #416: fleet_list must enumerate the STARTUP profile set' (#420) from worker/416-3ad1da-1 into main
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m34s
2026-09-10 06:27:36 +02:00
Dai Ha 8d5bc3ee89 fleetd #416: fleet_list must enumerate the STARTUP profile set
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m52s
`profiles` is a DEFERRED key: HerdrPeerLauncher takes Map.copyOf(profiles) once at
construction, so a profile added only to the hot-reloaded map can never be spawned.
CapacitySource was built with `() -> config.get().profiles().keySet()` — the live map —
so fleet_list reported a hot-added profile as free while fleet_spawn on that same
profile failed with "unknown worker profile". fleet_list's contract for `free` says it
runs "the same check the spawn gate itself runs"; it did not.

The set now comes from the startup snapshot, the same shape as the coordinator.peers
wiring three lines below, whose comment already stated the rule. maxLoad stays live on
purpose — ConfigRef documents it as hot, like credentialId and weight — so a maxLoad
edit still takes effect without a restart.

Found by the fleet01 lead, who proved it with a live ghost profile rather than an
argument. Implemented by a sonnet member; its reply was lost to an empty scrape, so the
work was salvaged uncommitted from its worktree and both proof steps were run by the
lead instead:

  full build                          Tests run: 1505, Failures: 0, 0 compile errors
  mutation A: live keySet restored    liveOnlyProfileIsNotListed FAILS (alone)
  mutation B: permanently empty set   startupProfileIsListed FAILS (alone)

Mutation B is the point of the second direction: per #404, a test that only ever checks
the absent case cannot tell a correct lookup from one that returns nothing at all.
2026-09-10 11:25:56 +07:00
Dai Ha ab0cc71aa4 fleetd #418: barrier the throw-path push-loop test on state decide() reads
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m50s
anAskThatLeavesByThrowingStillClosesItsQuestion barriered on Phase.ASKING,
which markAsyncQuestion sets in ask()'s FIRST step. The assertion right
after it depends on ask()'s THIRD step (pushLoop.onQuestionOpened), which
is what actually populates ReplyPushLoop's pendingQuestions map. Under
load the asker thread can be descheduled between those two steps, so the
barrier released before decide() had anything to see, and it correctly
returned STOP instead of the expected INJECT.

Add ReplyPushLoop#pendingQuestionTurnIdsForTest, a package-private test
seam (modeled on MessageService#isCompletionStampedForTest) exposing the
private pendingQuestionTurnIdsFor. The test now waits for its own turnId
to appear there before asserting on decide() — not for decide() itself to
return INJECT, which would make the barrier assert nothing.

Checked every other awaitTicketPhaseOn(..., Phase.ASKING) in the file
(two, in the CB-582 nudge tests): both are followed by a real awaitNudge()
that waits for an actual agent.prompt push-loop call before any assertion
depends on push-loop state, so they are not exposed to this race.

No production behaviour changed.
2026-09-10 11:24:41 +07:00
ltms 4cd9046353 Merge pull request 'fleetd #409: deterministic test for the #399 completion-stamp ordering race' (#414) from worker/deterministic-stamp-race-409-3cb7b6-10 into main
CI / contract (push) Successful in 59s
CI / build (push) Successful in 1m36s
2026-09-10 04:48:44 +02:00
Dai Ha 4e98a74047 fleetd #409: deterministic test for the #399 completion-stamp ordering race
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 2m5s
Widens the completed-hook's real race window (normally instructions-wide, needing
~2x-core host load to hit by chance per #399) by injecting a bounded sleep into the
test clock's completion-stamp read. This makes the ordering invariant — a test must
wait for isCompletionStampedForTest, not just DONE, before advancing the clock past
the TTL — fail deterministically on the first run when the barrier is removed, and
pass deterministically with it present. No production code changed.
2026-09-10 09:39:08 +07:00
ltms a502ba53e0 Merge #404: the armed field reads the startup pattern map, both directions pinned
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m0s
The armed lookup now uses the same compiled map detection uses. Second test added during review after a mutation proved the true direction was unpinned.
2026-09-10 04:20:54 +02:00
Dai Ha 446cc11d1d fleetd #404: pin the armed field's TRUE direction too
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m51s
Mutation-tested during review: replacing the armed lambda with
`profile -> false` left the whole suite green at 1475 tests, because the
only existing test passes an EMPTY startup map. That mutation would make
#395's visibility feature silently dead.

With this test the same mutation fails, and it is the only test that
fails, so nothing else covers this direction.
2026-09-10 09:20:43 +07:00
ltms ddd81fe174 Merge #391: fleet_reply refuses a lead, and the role argument is now mandatory
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m1s
The permissive 3-arg overload is deleted, so a handler that forgets the role no longer compiles.
2026-09-10 04:18:36 +02:00
ltms ae74cc081f Merge #399: wait on a real completion stamp, not on DONE
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m29s
Adds a volatile completedNanos stamp and a bounded test barrier. Mutation-tested: removing the production stamp fails 3 tests.
2026-09-10 04:14:29 +02:00
Dai Ha cf8da1d5fa fleetd #391: require reply caller role
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m45s
2026-09-10 09:13:16 +07:00
Dai Ha b8aedeafcb fleetd #404: test armed startup map behavior
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m30s
2026-09-10 09:13:07 +07:00
Dai Ha 3d61af6f6f Merge #400: the scrub receipt now measures the blank, not the attempt
CI / contract (push) Successful in 1m8s
CI / build (push) Successful in 1m34s
The blanking loop classified a name as blanked from eval's exit status.
zsh coerces a bare NAME= assignment on an integer special parameter
(SECONDS, RANDOM, SHLVL, HISTSIZE, COLUMNS, LINES, USERNAME) to a number
instead of failing, so eval returned 0 with the value untouched. Measured
7 false receipts in 10 names. This is a security receipt, so a count that
overstates the scrub is worse than no count.

The loop now runs the eval unconditionally and decides from the observed
value, read back with the (P) indirection flag. One check covers all
three shapes a name can take: a real blank, a fatal error eval merely
contained, and this silent no-op. Exit status plays no part.

Lead review: mutation removing the '!' unblankable report line is CAUGHT
(2 failures in EnvAllowListScrubTest, both asserting the name is reported
rather than silently dropped). The eval-site identifier guard is
untouched.

Unplanted evidence the fix works: the #394 test
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt began failing under
the fix, because zsh auto-exports SHLVL and the old exit-status bug had
been miscounting it as blanked all along. Its exact-count assertion had
only ever passed because of the bug beside it; it is now a presence
check, since which names a zsh version auto-exports is not this test's
to pin.

Still not pinned, tracked in #394's follow-up: the eval-site identifier
guard has no test behind it.
2026-09-10 09:05:30 +07:00
Dai Ha ed99c209ac Merge #398: a central allow-list of usable models, with the startup validators pinned
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m25s
models: allow: is a single place that names every model the fleet may
use. Absent or empty keeps today's behaviour, so this ships inert until
configured. Once set, a profile naming a model outside the list refuses
to start, and refuses a reload, rather than reaching a backend adapter
as a free-form string.

The allow-list cannot be checked against a provider catalogue: for an
opencode profile fleetd SYNTHESIZES the provider from provider/model
plus baseUrl (OpenCodeLauncher:562-583), so a valid fleetd model id
appears in no published catalogue. An operator-owned list is therefore
the only workable gate.

Includes the #398 follow-up: FleetConfig.validateAll() reflectively
sweeps this class's validateXxx() methods, and both real call sites
(Fleetd.main and ConfigRef.reload) call that one method. Before this,
deleting a validateXxx() call from either caller left the whole suite
green. FleetdStartupValidationTest now drives the real Fleetd.main.

Lead review: mutation on the reload call site is caught (ConfigRefTest,
2 failures). Mutation replacing the reflective sweep with a hardcoded
list is NOT caught (1491 green) — so the sweep is a convenience and the
denominator test is the real guarantee; two false statements in the test
javadoc were corrected to say so (af4c88d).

Recovered work: the worker's agent died mid-turn with the follow-up
uncommitted and the startup call left disabled as
'// MUTATION-TEST-TEMP: cfg.validateAll();'. I restored it before
committing.

NOT covered: the five log-only reportXxx(cfg) calls in main are still
unpinned — filed as #407.
2026-09-10 09:02:18 +07:00
Dai Ha af4c88d54b t398: correct two false statements in FleetConfigValidateAllTest
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m58s
Measured at review: reverting validateAll() to a hardcoded list of
today's six calls leaves the suite green (1491 tests, 0 failures). The
class javadoc claimed that mutation fails a test. It does not — claim 1
pins the generic helper on an unrelated class, claim 2 pins today's six,
and a hardcoded list satisfies both.

The interaction was the real hazard. The denominator assertion IS a
tripwire (declaring a seventh validator fails it), but its failure
message said the sweep reaches new validators 'by construction' and told
the author to just update the expected set. If the sweep were ever
replaced by a name list, the one assertion that fires would hand back a
false all-clear at the moment it fired.

Javadoc now states the measurement, and names the denominator test as
the actual guarantee. The assertion message now says to confirm
validateAll() still delegates to invokeAllValidators(this) BEFORE
updating the expected set.
2026-09-10 09:01:49 +07:00
Dai Ha 44c735f6f5 fleetd #404: report armed detection from startup map
CI / build (pull_request) Successful in 1m25s
CI / contract (pull_request) Successful in 1m20s
2026-09-10 08:56:39 +07:00
Dai Ha b540a1744b fleetd #398 follow-up: pin the startup validators with a reflective validateAll()
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Successful in 1m28s
Mutation testing found that deleting a cfg.validateXxx() call from
Fleetd.main left the full suite green: every test called a validator
directly and none exercised main as the caller.

FleetConfig.validateAll() sweeps this class's own public no-arg void
validateXxx() methods by reflection and invokes each in alphabetical
order, so a newly written validator is wired into both callers
(Fleetd.main and ConfigRef.reload) with no second step to forget.
FleetdStartupValidationTest calls the real Fleetd.main with six configs,
each failing exactly one validator.

Recovered by the lead: the worker's agent died mid-turn with this work
uncommitted, and had left the startup call commented out as
'// MUTATION-TEST-TEMP: cfg.validateAll();' from its own mutation run.
I restored the call before committing. Build after restoring:
Tests run: 1491, Failures: 0, BUILD SUCCESS.

NOT covered, and not claimed to be: the five log-only reporters in
main (reportRequiredSecrets, reportGitHostShape, reportMemberTrustModel,
reportMemberCredentialsGap, and reportExhaustedPatternGap on current
main) are not validateXxx() methods, so the sweep does not reach them
and their call sites stay unpinned.
2026-09-10 08:56:20 +07:00
Dai Ha 0f51d53098 fleetd #399: fix TTL test race by waiting on a real completion stamp, not DONE
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m2s
poll() can report Phase.DONE for a ticket before the whenComplete hook that
stamps Task.completedNanos has run — CompletableFuture.complete() publishes
its result and only then runs dependents. The TTL tests advanced an injected
clock right after observing DONE, so on a host where the hook runs late it
stamps the ADVANCED time and the eviction never happens (fails on Linux,
passes on macOS).

Add a package-private test seam, MessageService.isCompletionStampedForTest,
that reports whether completedNanos is stamped. Both TTL tests now wait on
that (a real volatile read/write happens-before edge) before advancing the
clock, instead of on Phase.DONE. The prune condition in pruneTerminalTickets
is untouched.
2026-09-10 08:54:38 +07:00
Dai Ha c670792ffe #400: classify the blanking loop's result on the observed value, not eval's exit status
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m1s
eval "export NAME=" can return success even when zsh coerces the bare
assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
HISTSIZE, COLUMNS, LINES, USERNAME) instead of failing, leaving the
value unchanged. The old exit-status check then reported the name as
blanked when it was not -- a false receipt.

Classify on the observed effect instead: attempt the export, then read
the name's value back with the (P) indirection flag and decide from
whether it is now empty. One check now covers all three shapes a name
can take here -- a genuine blank, a fatal read-only error eval merely
contains, and this silent no-op -- with the exit status playing no
part in the decision.

Adds a test driving all three shapes through the real scrubScript in
one run (a normal name, LINENO for the fatal case, SECONDS for the
silent no-op), with the parent environment explicitly carrying those
names since a cleared ProcessBuilder parent does not expose them on
its own. Also corrects the previously-merged
unblankableNameInTheMiddleDoesNotAbortNamesAfterIt test, whose "exactly
one failed name" assertion turned out to only pass by accident: zsh
itself auto-exports SHLVL on every shell start, and the old exit-status
bug was silently miscounting it as blanked. The fixed classification
now correctly reports it unblankable too, so the test asserts presence
rather than an exact count.
2026-09-10 08:46:27 +07:00
Dai Ha 3982ace544 fleetd #391: refuse lead fleet replies
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m31s
2026-09-10 08:43:54 +07:00
Dai Ha 7180b1aad0 Merge #395: warn when exhaustedPattern usage-limit detection is off
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m27s
A profile with no exhaustedPattern has usage-limit detection silently
disabled. Startup now reports every unarmed profile (a louder warning for
subscription profiles, which are the ones a limit actually stops), and
fleet_profiles / GET /profiles carry exhaustionDetectionArmed per profile.

Reviewed by the lead: all three requested mutations fail a test, and the
back-compat QuarantineSource ctor defaults to 'not armed' when the source
is unknown. KNOWN GAP, not fixed here: deleting the
reportExhaustedPatternGap(cfg) call at Fleetd.java:140 leaves the suite
green (Tests run: 1472, Failures: 0). That is the same unpinned-startup-
call shape as PR #398's six validators, and fleetd #398's ticket owns it.
2026-09-10 08:43:09 +07:00
Dai Ha c26f695402 fleetd #395: warn when exhaustedPattern detection is silently off
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m25s
exhaustedPattern is opt-in per profile: unset means a usage-limit
refusal on that profile is never classified BACKEND_EXHAUSTED and
never quarantines its credential, with nothing telling the operator.
Add a startup WARN naming every unarmed profile (a louder, separate
WARN for a subscription: true profile, since that is the operator's
own metered plan). Surface the same fact per profile in fleet_profiles
as exhaustionDetectionArmed, so an operator can tell "healthy" from
"can never be caught" without reading fleetd.yaml.
2026-09-10 08:34:33 +07:00
Dai Ha 48877315ca Merge #394: contain the fatal export that aborted the credential scrub
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m31s
The CB-633 allow-list scrub has been dying mid-loop on every fleet01 member
pane and saying nothing. `export UID=` in zsh is not a failed command — it is
a fatal parameter error that terminates the whole sourced file. The blanking
loop is wrapped in `{ ... } 2>/dev/null`, so the message was swallowed and the
report block after the loop never ran.

Root cause found by the fleet01 lead, with xtrace on a live pane's own ZDOTDIR:

    +scrub.zsh:28> _cb633_n=UID
    +scrub.zsh:28> export 'UID='
    +zsh:1> rc=126        <- file aborted

The severity is the INVERSION, and this is their finding, quoted:

  "env lists inherited names first and the names a startup file exports last.
   So the loop blanks the harmless inherited half and dies immediately before
   the operator's own exports — exactly the credentials the policy exists to
   remove. The selection is inverted, not merely partial."

Measured there: UID is name 42 of 57, and a ~/.zshrc decoy at 58 survived on
8 of 8 spawns. "Partial scrub" reads as "we got most of it"; it got precisely
the wrong half.

Fixed with `eval "export ${n}=" 2>/dev/null` rather than a skip-list of the
known-fatal names (UID EUID GID EGID PPID LINENO). A skip-list has to be
complete forever and this is a security control; eval needs no list. Measured:
plain export dies at UID and every later name keeps its value, while the eval
form completes the loop and blanks all of them. PR #396 proposed the skip-list
and is closed in favour of this; its claim that the abort happens "however the
assignment is wrapped" holds for a direct `if ! export` but not for eval,
which reparses in a nested context.

The report now carries `failed N` and `!`-prefixed unblankable names, and
HerdrPeerLauncher WARNs when any name could not be blanked. The old "no report"
WARN no longer claims the daemon knows what the member saw.

Why the suite stayed green: EnvAllowListScrubTest starts zsh from
pb.environment().clear(), and under a cleared parent UID is not an exported
name at all, so the abort could not reproduce in that harness.

Reviewed by mutation, which found a second gap now also closed: the eval is
only safe because names are filtered to ^[A-Za-z_][A-Za-z0-9_]*$. Replacing
that pattern with .* left the class green, so the line the security property
rests on was unpinned. The guard is now re-asserted at the eval site and pinned
by a test. The reachable vector is a VALUE with an embedded newline, not a
hostile name — measured: zsh strips non-identifier env names outright, while
MULTI=$'keep\njunk.fragment' forges 'junk.fragment' as a candidate name out of
its own value.

Closes #394. Refs #396, #388.
2026-09-10 08:30:15 +07:00
Dai Ha 5c08054533 #394 follow-up: re-assert the identifier guard at the eval call site
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m48s
EnvAllowListScrub's blanking loop splices each name into a string
handed to eval ("export ${n}="). That is only safe because every name
reaching _cb633_blank already passed an identifier check in the
enumeration loop -- 20 lines away, in a different loop. Before eval
was introduced a non-conforming name reaching plain `export "$n="`
was inert either way (the quoting neutralized it); eval removed that
safety net, so the enumeration loop's guard became the ONLY thing
standing between a non-identifier string and code execution in the
member's pane, with nothing at the eval site itself defending that
property.

Re-assert the same [A-Za-z_][A-Za-z0-9_]* check immediately before
the eval call, independent of the enumeration loop's own guard (left
untouched, not moved). A name that fails it is counted unblankable
rather than silently dropped, so a bypass of the upstream guard would
leave real evidence in the report.

New test exploits the "junk from multi-line values" gap the
enumeration loop's own comment already documents: a value with an
embedded newline makes `command env`'s text output split into a
spurious extra "name" line that was never a real variable. Runs the
real generated scrubScript() end-to-end under zsh and asserts the
non-conforming fragment is neither blanked nor counted unblankable.
The fragment used is merely non-conforming (contains a dot) --
never command-shaped.

Mutation-verified both guards. Weakening the enumeration guard alone
DOES break the new test (the fragment then reaches the new eval-site
guard and gets counted unblankable, failing the "not unblankable"
assertion). Removing the new eval-site guard alone, with the
enumeration guard intact, does NOT break it: _cb633_blank has exactly
one producer (the enumeration loop), so nothing can reach the eval
site without already having passed the identical check there. That is
expected given the single-source architecture, and it is exactly why
the eval-site guard is defense-in-depth against a future change that
adds a second path into _cb633_blank or decouples the two loops --
not a currently independently-observable divergence.
2026-09-10 08:24:46 +07:00
Dai Ha b6db9c31f5 charter: a peer lead is answered with fleet_send, not fleet_reply
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m0s
`fleet_reply` has no route to a peer lead. `AmqpReplyInbox` publishes to
`agent.<target>.inbox`, mandatory, and a lead's own terminal has no such
queue, so the publish is refused. `MessageService.reply()` has no peer
branch at all — `grep -c 'coord\|LeadMailbox'` on it returns 0. The charter
told every lead to use a tool that cannot work, and both leads here hit it.

Three edits to the canonical block, byte-identical with the wiki template
(pushed as 803726a; the in-sync check in this file reports True):

- the intent->tool row now says `fleet_send{coordId}`, or `{sessionId}` for
  a peer on the same host, and says plainly that `fleet_reply` is refused
- the prose says WHY: `fleet_reply` resolves a member's blocked `fleet_send`,
  while a peer's coord-id message is durable and non-blocking, so there is
  nothing for it to resolve
- lead<->lead item 3 gains the data-point rule: N observations are N data
  points only if they differ in the axis you are trusting

Wording for all three drafted by the fleet01 lead, who verified the missing
queue namespace independently in its own tree. The data-point rule has now
caught three separate errors in a day, in both directions: one cause blamed
for N failures, and N agreeing measurements that shared a single instrument.

The refusal message itself is still wrong — it says "queue not declared or
owned", which sends the reader to the broker instead of to this file. That
half stays open on #391.

Tracked as fleetd #391.
2026-09-10 08:22:35 +07:00
Dai Ha e7b33fe3a0 fleetd: central allow-list of usable models (models: + validateModels())
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m34s
Add an optional top-level `models:` block (Models{allow: List<ModelEntry>})
naming the models any profiles: entry may use. Absent/empty allow: keeps
today's behaviour exactly (no check, no warning). When configured,
FleetConfig.validateModels() fails config load (and reload, via ConfigRef)
naming both the model and the profile, if any profile's model: is outside
the list. The check is one-way: editing profiles: alone can never widen
what is permitted, only models.allow: can.

Wired into Fleetd.main() alongside the other validateXxx() calls, and into
ConfigRef.reload()/DEFERRED_KEYS so a bad edit can't slip in through a
reload either. Each ModelEntry is its own record (not a bare string) so a
later unit can add per-model on/off or load-limit state without changing
the YAML shape. One flat string namespace covers both a bare Claude id and
an opencode provider-prefixed id.
2026-09-10 08:09:27 +07:00
Dai Ha e3e403e5c8 #394: contain a fatal export error instead of letting it abort the scrub
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m52s
EnvAllowListScrub's blanking loop used a plain `export "$n="` on every
name not on the allow-list. For a zsh read-only/special parameter (e.g.
UID) that is a FATAL parameter error, and since the loop runs inside the
sourced startup file, the error aborts the whole file: every name still
to come is never blanked, and scrub-report.txt is never written at all
-- silently, because 2>/dev/null on the group swallows it.

Route each blanking attempt through `eval` instead, which contains the
error to that one iteration. The loop always finishes; a name it could
not blank is now counted separately ("failed" on the report's first
line) and listed !-prefixed rather than disappearing. No skip-list of
known-bad names is added -- every enumerated name is still attempted,
so a name nobody has thought of is still tried and, if it fails, still
counted.

HerdrPeerLauncher: log a WARN when a pane's report carries a nonzero
failed count, and reword the "no report at all" WARN so it no longer
claims the daemon knows the member "saw the full host environment" --
a partial vs. a missing scrub are different situations and only the
first is now distinguishable from the report alone.
2026-09-10 08:08:41 +07:00
Dai Ha 799014e99d Merge #388: scrub a pane shell that is neither login nor interactive
CI / contract (push) Successful in 1m8s
CI / build (push) Successful in 1m38s
2026-09-10 07:21:48 +07:00
Dai Ha 69e09b10fa #388: scrub a pane shell that is neither login nor interactive
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m32s
EnvAllowListScrub generated four zsh startup files but only .zshrc and
.zlogin sourced the scrub — .zshenv (the one file zsh always reads) did
not. A pane shell that is neither login nor interactive reads only
.zshenv and stops, so it was never scrubbed at all (measured on fleet01,
issue #388).

Adding an unguarded scrub to .zshenv (the ticket's own suggested fix) is
wrong: .zshenv is read by every zsh, including a short-lived `zsh -c`
a member's own tooling forks for a single command. Those children are
also neither login nor interactive, so they would scrub the environment
their parent deliberately set for them (GIT_DIR, VIRTUAL_ENV, ...), and
the rewritten scrub-report.txt would describe the last child to exit
instead of the pane.

Fix (per comment 15387, measured): keep .zshrc/.zlogin unconditional,
and add to .zshenv a pass guarded on the exact condition that defines
the gap (neither login nor interactive), plus a per-pane sentinel
(_CB633_SCRUBBED) so it runs once per pane, not once per process. The
sentinel is exported only after the scrub runs, and is folded into the
scrub's own allow-list so a later pass in the same pane cannot blank it
back to empty.

Also corrects the class javadoc's wrong premise (a bare argv[0] proves
NOT login, not "therefore interactive") and its now-stale two-file
walkthrough.

Tests: two new real-zsh tests in EnvAllowListScrubTest run actual
non-login/non-interactive zsh processes (never string-match the
generated files) to prove: a neither-shell pane is scrubbed; a child
that pane forks keeps variables the pane deliberately set for it; the
child does not re-scrub; and scrub-report.txt still describes the pane
after the child exits. Both fail without the production fix (verified
by reverting it and re-running: AssertionFailedError on the sentinel
and on the decoy secret surviving).
2026-09-10 07:21:08 +07:00
Dai Ha b9d09e044e t386: pin the per-member drift baseline the fix's own tests left open
CI / contract (push) Successful in 37s
CI / build (push) Successful in 1m57s
The two tests merged with #386 both start with the member already BUSY, so a
single global drift baseline passes them. This one sleeps the host while nothing
is busy and only then starts a turn, which fails without the per-member map.
2026-09-10 06:59:42 +07:00
Dai Ha fd8650cda4 Merge #386: correct the stall check for a monotonic clock frozen by host sleep 2026-09-10 06:56:37 +07:00
Dai Ha 11050e24ed t384: fix javadoc indentation on the merged shareWithGroup lines
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m29s
2026-09-10 06:54:51 +07:00
Dai Ha 769f282408 #386: give the stall detector a real-time clock, log the divergence
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m26s
FleetHealthMonitor.tick's stalled check compared two monotonic-clock
readings (System.nanoTime(), which macOS freezes across a host sleep),
so a member BUSY for 101 real minutes was never flagged.

The monitor now also takes a wall-clock LongSupplier (realtimeClock),
used only inside the stall check. Each tick measures how far the two
clocks moved apart since the previous tick and folds any positive
divergence into a running total; when a single tick's divergence
exceeds one tick interval (the signature of a sleep, since a tick
cannot run while the process itself is suspended) it logs one WARN
naming how long the detector could not see. The correction is applied
per member, keyed to when that member's current lastActivityAtNanos
was first observed BUSY - not since monitor start - so a sleep that
happened before a member went busy is never charged to it.

Every other use of the monitor's clock (readiness grace, snapshot
timestamp) is unchanged. Backend quarantine/cool-off, the lead tab
scan, the completion resolver, the session reaper and the message
service TTLs are untouched, per the ticket's decision.

Existing FleetHealthMonitor/FleetHealth tests pass unmodified (none of
them ticks a BUSY session more than once, so the drift path never
engages for them). Two new tests: a frozen monotonic clock past the
real-time threshold produces STALL_SUSPECTED, and a single sleep gap
logs the divergence exactly once, not once per tick.
2026-09-10 06:53:30 +07:00
Dai Ha 6f71f40047 Merge #384: pre-create the scrub receipt and give it group write 2026-09-10 06:49:46 +07:00
Dai Ha fb36c5238f Merge #382: SpawnRequest.withProfile() replaces the six-accessor rebuild 2026-09-10 06:49:46 +07:00
Dai Ha cc9cdc938b #384: write shared scrub receipts 2026-09-10 06:45:30 +07:00
Dai Ha 5fede82468 #382: give SpawnRequest a withProfile wither, guard it against the arity trap
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m49s
CompositePeerLauncher:372 rebuilt a routed SpawnRequest from a literal
new SpawnRequest(...) call listing six of the original request's own
accessors. That call is only correct because it happens to match the
canonical 6-arg constructor today; add a 7th component plus the
established back-compat constructor at the old (now-shorter) arity and
this call would silently rebind to it, dropping the new field on every
profile-routed spawn with no compile error — the same defect shape
already guarded on FleetConfig.withDefaults() (#357) and MemberSession
(#358).

Add SpawnRequest.withProfile(String), modeled on
MemberSession.withState/withActivity, and use it at the call site
instead. Add a guard test that resolves the true canonical constructor
by exact component types (never by argument count), gives every
component a distinctive value, and asserts every component but
profileName survives withProfile() unchanged.

Proved the guard against the real mechanism: temporarily dropped the
last (role) argument from withProfile()'s constructor call so it bound
to the 5-arg back-compat constructor — it still compiled, and the new
test failed, catching the silently-defaulted role. Restored the fix
and reconfirmed green.
2026-09-10 06:44:15 +07:00
Dai Ha 1515025804 charter: a profile differs in liveness, not just model and cost
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m31s
The old line said profiles differ in model and cost. That is true and it is
not the reason the default hurts. Measured on two hosts: a default sitting on
an exhausted or withdrawn credential either fails the spawn loudly or, worse,
produces a member that starts fine and then returns nothing.

Wording proposed by the fleet01 lead; merged with the existing 'not in tier'
clause, which is still right. Applied byte-identically to the wiki template.
2026-09-10 06:28:51 +07:00
Dai Ha 7754f53662 Merge t385 follow-up: pin the mid-turn redelivery ack 2026-09-10 06:28:08 +07:00
Dai Ha a507f7b31b t385: pin that a redelivery is acked even while the lead is mid-turn
Verification found the fix's dedup check could be moved behind the
injectable-status gate and every test still passed. That placement matters:
this lead is mid-turn most of the time, so gating the ack on an idle pane
leaves the redelivered message held, and the next recovery delivers it again.

The new test fails on that mutant and passes on the fix.
2026-09-10 06:28:03 +07:00
Dai Ha 71c322f104 Merge #385: a redelivered lead message must not write the pane twice 2026-09-10 06:24:40 +07:00
Dai Ha fde2c15627 t385: a redelivered lead message must not write the pane twice
An AMQP recovery clears the held delivery tags, the broker redelivers with
fresh ones, and the coordination loop wrote the same peer message into the
lead pane again. Measured on the live daemon: one msgId reached the pane 12
times in 9 hours, across 19 recovery events.

LeadCoordLoop now remembers the msgIds it has written to a pane (bounded at
1024) and acks a redelivery without a second write. LeadMailbox.ack no longer
returns quietly for an unknown msgId: it throws, so an ack that never reached
the broker is reported instead of hidden. A repeat ack that this connection
already completed stays quiet, tracked in a bounded set.
2026-09-10 06:24:35 +07:00
Dai Ha 2830735644 t377: raise the logger level in the test, or it asserts on an empty list
CI / contract (push) Successful in 46s
CI / build (push) Failing after 1m32s
The 9 tests the member wrote failed 7 of 9 on first build. The production
code was correct; the test harness was not. logback-test.xml sets
dev.ltms.fleet to WARN, so the INFO shape lines were dropped by the level
check before any appender saw them.

The member copied attach()/detach() from MemberTrustModelReportTest but not
the setLevel(INFO) those siblings do at each call site. Doing it inside
attach()/detach() covers all nine at once, and restores the original level
(null, meaning inherit) rather than a concrete one.

Proven by mutation: making the code log the value fails 4 tests, including
theValueNeverAppearsInLogOutput — 'the full GITEA_HOST value must never
reach the log'. Full build 1459 tests green.
2026-09-09 07:52:43 +07:00
Dai Ha 4e3ac91a22 Merge worker/t377-7f587b-7: startup reports the git host value's shape, never its value 2026-09-09 07:49:09 +07:00
Dai Ha 29d3f0b41f t377: report the git host value's SHAPE at startup, never its value
A member receives the git host as GITEA_HOST and often gets a full URL
(scheme, trailing slash) where it expects a bare host, so it builds
https://https://... and the request never leaves the machine.

Logs set/unset, length, startsWithScheme and trailingSlash next to the
existing secret report. The value itself is never logged, and the value
is still passed to members unchanged — rewriting it here would change
what works on one host and breaks on another.

Written by the gx member on branch worker/t377-7f587b-7; it could not
build, commit or push because the host command classifier refused every
shell command (see #381). Build and verification are mine.
2026-09-09 07:49:02 +07:00
Dai Ha cbc732444f fleetd #365: the push-nudge metric outcome is 'sent', not 'delivered'
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m31s
The label was renamed in the #365 merge. This table still named the old one.
'sent' counts the herdr paste-and-submit call returning, never a confirmation
that the pane read it.
2026-09-09 07:40:57 +07:00
Dai Ha 6f828b8c38 Merge worker/t365-3920c5-3: t365: fleet_reply/REST reply distinguish resolved vs queued; rename nudge metric outcome
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m48s
2026-09-09 07:35:00 +07:00
Dai Ha 145a8c8862 Merge worker/t373-336973-2: t373: pin the production XDG-excludes seam GitWorktreesTest.seedingGitWorktrees builds 2026-09-09 07:35:00 +07:00
Dai Ha 7057291739 Merge worker/t358-6e989b-1: t358: guard MemberSession's 5 rebuild sites and Profile.withProfile() against the back-compat-arity trap 2026-09-09 07:35:00 +07:00
Dai Ha 2302b3bc11 #376: the too-fast failure stops asserting a cause it cannot know
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m31s
A turn that settles inside MIN_TURN_NANOS is still reported FAILED. Only the claim
about WHY is withdrawn, and the pane is still carried.

Measured on fleet01 on 2026-09-08 UTC: an opencode member on mimo-v2.5-free answered
a real question in 1575ms, below the 2000ms floor. The daemon reported the turn as
failed with 'most likely a backend error before any work started'. The answer was
right there in the scrape. With no errorPattern configured — the live state on both
hosts, which both log as 'backend-error classification: off' — that sentence is a
guess, and a reader who believes it stops looking at the pane.

WHAT I REJECTED, because the next person will try it. A worker implemented the
ticket's first suggested direction: inspect the pane inside the floor and resolve a
COMPLETION when the text looks like a real reply. Its test for 'looks like a real
reply' was non-blank plus a '.', '!' or '?' anywhere in the text. That is unsafe
twice over. lastAssistantBlock falls back to the WHOLE pane when it finds no U+23FA
marker, so on a crash the candidate reply is the entire screen; and a crash pane
almost always contains a full stop, in a file path, a version or a hostname. I ran
that implementation against the new guard test and it resolved

    Error: connection reset while loading src/main/java/Foo.java v1.2.3

as a COMPLETION — expected: <FAILED> but was: <COMPLETION>. A loud wrong answer
became a silent one, which is the trade the ticket brief forbade.

The obvious repair does not work either. Requiring the U+23FA marker as positive
evidence would be safe, but that marker is Claude Code chrome and an opencode pane
never carries it — and an opencode member is what raised this ticket. There is no
reliable cross-backend marker for 'this is a real reply', so this path must not try
to judge one. That reasoning is now in the failTooFast javadoc.

Two guard tests. aPlausibleLookingReplyInsideTheFloorStillFails pins the safety
property against exactly the rejected approach; it passes today and fails against
that implementation, which is how it was verified rather than assumed.
theTooFastFailureDoesNotAssertACauseItCannotKnow pins the wording.

One existing test changed. aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink
asserted the phrase 'too fast to be real work', which carried the withdrawn claim. It
now asserts what it was really guarding: the floor alone fails the turn, the reason
stays generic, the pane is carried, and the typed sink is never notified.

MIN_TURN_NANOS is unchanged at 2000ms. mvn clean install: 1441 tests, 0 failures.
2026-09-09 07:28:04 +07:00
Dai Ha 3a004dc1b3 t373: pin the production XDG-excludes seam GitWorktreesTest.seedingGitWorktrees builds
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m32s
fleetd #362 review finding 2 protects GitWorktrees#previouslyEffectiveExcludesFileContent's
Java-side XDG_CONFIG_HOME/HOME read (it never goes through a git subprocess, so no
GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM isolation reaches it) with a gitEnv constructor seam. A
mutation run during the #372/#369 merge found that seam unpinned: stripping hermeticGitEnv(tmp)
from seedingGitWorktrees left every test green, poisoned XDG_CONFIG_HOME or not.

Adds seedingGitWorktreesResolvesTheExcludesFileFallbackInsideItsThrowawayDirectory, which asserts
the property directly (a GitWorktrees built for seeding resolves the fallback inside its own
throwaway directory) using a self-contained marker instead of relying on an externally poisoned
env var. Refactors hermeticGitEnv/seedingGitWorktrees into two-argument overloads (one taking an
explicit XDG_CONFIG_HOME / gitEnv) so the new test can pre-populate the marker before construction
while still going through the same production construction every other seeding test uses; no
behavior change for the 4 existing call sites.
2026-09-09 07:24:22 +07:00
Dai Ha a9a3c12232 t365: fleet_reply/REST reply distinguish resolved vs queued; rename nudge metric outcome
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m51s
fleetd #365. fleet_reply always returned the literal "delivered" and
POST /sessions/{id}/reply always returned {"delivered": true}, whether
the reply resolved a live waiting send/ticket or was merely queued in
the inbox for a later drain (CB-307) — both are successes, but not the
same fact.

MessageService.reply() now returns a ReplyOutcome (RESOLVED_SEND,
RESOLVED_ASYNC_TICKET, or QUEUED) instead of an always-true boolean.
FleetMcp.reply and FleetApp.replyMessage both read it: the MCP tool
result names which happened, and the REST body's "delivered" field is
now accurate, with an added "outcome" field.

Also renames the heartbeat/push-loop nudge metric's "delivered" outcome
to "sent" (LeadHeartbeatLoop, ReplyPushLoop, FleetMetrics): it only
records that the herdr agent.prompt paste-and-submit call succeeded,
never that the lead's pane actually read it — there is no read-receipt
concept at that layer, so "delivered" overclaimed there too.

Tests: MessageServiceTest/FleetMcpTest/FleetAppTest strengthened to
assert the specific outcome per case (a resolved send, a resolved async
ticket, and a queued reply); ReplyPushLoopTest updated for the outcome
rename.
2026-09-09 07:23:41 +07:00
Dai Ha 94ec77a1bc t358: guard MemberSession's 5 rebuild sites and Profile.withProfile() against the back-compat-arity trap
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m53s
Follow-up to #357 (FleetConfig.withDefaults()). Reflectively enumerate each
record's own components, resolve the canonical constructor by exact
component types, build a real non-null value per component, run each
rebuild site, and assert every component survives (except the one it is
documented to change). Exclusion lists pinned at 0 for both.

Re-counted the Profile back-compat ladder directly against the source:
8 constructors (arities 25, 24, 22, 20, 18, 15, 14, 12) against a
canonical arity of 26 — the ticket's own number was explicitly untrusted.
2026-09-09 07:19:12 +07:00
Dai Ha 49f285cfda charter: a measured fact in an addendum must carry its own deletion trigger
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m36s
The operator chose this and its size; the reasoning below is the fleet01 lead's.

Placed in the canonical block's boundary paragraph rather than the orchestration
body. That paragraph already talks about the addendum layer instead of protocol,
every project that mounts the bridge inherits it, and it sits about 3800 characters
before the primary's step list, so it does not dilute the steps a lead reads while
working. Perishability is structurally an addendum problem: the block is
byte-identical across projects by construction, so a dated local measurement in the
block body would already be a layering violation.

What happened. The fleet01 lead's kb addendum held a dated merge-refusal section
that carried an instruction to delete itself once it stopped reproducing. On
2026-09-08 UTC the operator granted merge rights on akb/kb, the lead re-ran the
probe, got 409 'head out of date' where the identical request had returned 405
'User not allowed to merge PR' on 2026-09-06, and deleted the section as instructed.

Why four parts and not one. The lead's finding is that the banner did not work
because it was emphatic. It worked because the falsification condition was
executable: it carried the exact probe, the reason for the all-zeroes
head_commit_id, and what each response code meant. The lead did not have to
reconstruct the experiment or decide what would count as refutation, and just ran
it. A banner saying 'this may be out of date, verify before relying on it' costs
the same space and does nothing, because deciding what would falsify a claim is the
expensive step and a reader in the middle of another task will not pay it. So: the
date, the command, what each outcome means, and the instruction to delete. The
fourth without the second is decoration.

The closing clause is the justification for the machinery. Most stale notes are
merely wrong. This one went stale in the dangerous direction: it would have told a
future lead it could not merge at the exact moment merging became its job, silently
and with confidence. A note that goes harmlessly stale does not need this.

Note what is NOT centralized here. The banner text itself cannot be. What fired for
the lead was a specific instruction sitting on top of the specific stale fact, which
it could not read past on its way to acting. A rule elsewhere saying 'date your
measurements' would not have fired, because nobody reads that rule at the moment
they re-measure. This sentence sets the convention; the trigger still has to live
next to the fact it governs.

Propagated to the wiki template in the same turn, wiki 8c4f152 on main; the sync
check in this file's addendum reports 'in sync: True'.
2026-09-09 04:24:17 +07:00
Dai Ha 5a12ae7930 charter: test a refusal, and do not count a transport failure as one
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m38s
Step 8 gained a refusal paragraph in 2f71a30, which said what a lead does once the
forge refuses a merge. It did not say how a lead establishes that it was refused.
Both halves of this amendment come from the fleet01 lead, measured on akb/kb on
2026-09-08 UTC, and both are ways to be wrong about a permission you never tested.

Do not read a refusal off a permissions field. After the operator granted merge
rights, the lead re-ran its probe: POST .../pulls/53/merge with an all-zeroes
head_commit_id, chosen so the request cannot succeed on its merits and a rejection
can only mean the refusal. It returned 409 'head out of date' where the identical
request returned 405 'User not allowed to merge PR' on 2026-09-06. A 409 is payload
validation and sits after the permission gate, so the grant took. The lead reports
the repository permissions object did not change across that flip -- still
admin:false, push:true, pull:true. I did not read that object myself; my forge token
is a different identity and would return a different one, so this stays the lead's
measurement and not mine. Merge rights on a protected branch live in branch
protection, so a permissions field can be wrong in both directions.

Do not count a transport failure as a refusal. The lead's first attempt returned
HTTP 000, because GITEA_HOST already carries a scheme and a trailing slash and the
URL came out as https://https://git.ltms.dev//api/... Under a 'not 200' test that is
indistinguishable from being refused. A probe exists to separate a refusal from
everything else, so an error that never reached the gate has to be a third answer
that concludes nothing.

Propagated to the wiki template in the same turn, wiki 8c2ef96 on main; the sync
check in this file's addendum reports 'in sync: True'.
2026-09-09 04:09:08 +07:00
Dai Ha 127e6832a9 Merge #374: fleetd holds off idle sleep while any member is live
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 2m3s
Lands PR #355 (fleetd #354's sibling), rebased onto current main by a worker
after 39 commits of drift left it unmergeable.

The problem, measured on the original branch: a fleetd host idle-slept after as
little as one minute (pmset -g custom reported 'sleep 1' on battery). Overnight
the daemon's AMQP link dropped 13 times, and every drop minute had a sleep or
wake event in pmset -g log in the same minute or the one before. The AMQP churn
is the visible symptom; the real cost is a member mid-turn freezing with the
host, and a long turn with nobody typing is exactly the case that goes idle.

IdleSleepGuard holds an OS-level assertion for as long as at least one member is
live. It is driven by SessionManager's existing onAcquire/onRelease hooks rather
than a second member count kept in parallel, so it reads the same registry
fleet_list's numbers come from, and only a real 0->1 or 1->0 crossing touches the
OS. It fails safe: a mechanism that cannot acquire means nothing is ever held,
and it never throws, never blocks a spawn, a release, or shutdown.

Conflict resolution was the whole job, and all three were in config plumbing:
ConfigRef, FleetConfig and ConfigRefTopLevelReportingCoverageTest. The power
package is byte-identical to the original branch commit.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1439, Failures: 0, Errors: 0, BUILD SUCCESS
  (1425 on main + 14 new: 4 caffeinate, 5 guard, 1 wiring, 4 config)

The denominator recount, which the worker flagged as its own weakest number
because this file's count has drifted three times before (#330/#333/#337). I
counted it mechanically rather than reading it: FleetConfig has 24 canonical
record components; COLD_KEYS 5, DEFERRED_KEYS 13, SPLIT_KEYS 3, plus the 3 the
javadoc names as hot-excluded (placement, memberCredentials, memberLoginShell).
5+13+3+3 = 24. The javadoc's '24 components: 5 cold, 13 deferred, 3 split, 3
hot-excluded' is correct. The worker's prose called idleSleepGuard the 25th
constructor argument; it is the 24th. The code is right, the report was off by
one.

Mutation run on merge, on the half the worker verified by READING rather than by
proving -- it said it had checked that withDefaults()'s final call binds the true
canonical constructor. I dropped the trailing idleSleepGuard argument so the call
silently binds the 23-arg back-compat overload. It compiles, which is the whole
hazard. Caught: 1 failure, 3 errors, BUILD FAILURE, and
FleetConfigWithDefaultsPreservesEveryComponentTest names the dropped component
and prints its own denominator -- '24 components, 24 checked, 0 excluded, 23
survived'. That test was added on main after this exact defect happened live when
idleSleepGuard was added on a sibling branch; the worker had to add the missing
entry to it, and doing so is what makes the guard cover this component at all.
2026-09-07 20:31:41 +07:00
Dai Ha 6442a583ae t355: give the withDefaults() coverage test a value for idleSleepGuard
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m6s
FleetConfigWithDefaultsPreservesEveryComponentTest was added on main after
the sleep-guard commit was cherry-picked, specifically to catch this exact
rebase hazard (a withDefaults() call silently rebinding to a stale-arity
back-compat constructor). Its baseValues() map didn't know about the new
idleSleepGuard component yet, so the coverage test itself failed the
name-drift check. Add a real, non-null value for it, consistent with how
the sibling ConfigRefTopLevelReportingCoverageTest already covers it.
2026-09-07 20:25:33 +07:00
Dai Ha 24b96d29ae fleetd must not let the host idle-sleep while members are live
Adds a small IdleSleepGuard (dev.ltms.fleet.power) that holds a macOS
caffeinate -i child while at least one fleet member is live, and
releases it once none are. It hangs off SessionManager's existing
onAcquire/onRelease hooks and SessionManager#size() rather than
tracking members a second way. New idleSleepGuard: config block,
on by default, following the FleetConfig.Health/ConfigReload pattern.
2026-09-07 20:22:11 +07:00
Dai Ha 4721771052 charter: a peer reads your project addendum
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m47s
Fourth item under Lead <-> lead. The addendum is instruction surface -- every
future session on that host obeys it, and a wrong one is obeyed as faithfully as
a right one -- but nothing in the block said to have anyone check it.

Evidence is one addendum, written on fleet01 this week, and it carried two
defects that its author did not see and a non-author did. First: it described a
forge permission wall as if it were policy, so a token regrade would have made
it tell a session NOT to merge at the moment merging became its job. Second,
found only because the first was raised: it DID carry the 'this is a dated
measurement' caveat, at the bottom, after the prohibition -- so a session that
read the instruction and stopped had already taken it as policy. The caveat sat
downstream of the thing it qualified.

Neither is a writing slip. Both are the author being unable to see their own
qualifier placement, which is what a second reader is for.

Deliberately conditional: not every operator runs two leads, so the rule ends
with what a lone lead can still do -- ask which sentence goes false first, and
whether a reader reaches the caveat before acting.

Weaker evidence than the step 8 amendment (n=1 addendum, 2 defects, versus a
measured 405). Recorded as such so it can be dropped if it does not earn the
context it costs.
2026-09-07 05:03:58 +07:00
Dai Ha 2f71a30bd7 charter: say what step 8 means when the forge refuses the merge
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m35s
Step 8 said 'then merge' and nothing about a refusal. fleet01 is the first host
to hit that: on akb/kb its lead gets 405 'User not allowed to merge PR' from the
API, and main is protected, so merging locally and pushing is refused too. Both
routes shut. The charter was telling a lead to do something the forge would not
let it do, and that gap was latent in every copy of the block.

The amendment keeps the rule that the merge decision is never delegated, while
admitting the mechanical merge may not be the lead's to make.

The second sentence is the one that earns its keep, and it came from the fleet01
lead rather than from me: never call a PR 'ready to merge' without having read
the diff. A refusal is exactly when that shortcut is tempting, because no action
is left that forces the lead to look. Without it, a refused merge quietly turns
step 8 from 'read it yourself, then merge' into 'forward the reviewer's verdict'
— the proxy-delegation the same step forbids two lines earlier.

Propagated to the wiki template in the same turn; the sync check in this file's
addendum reports 'in sync: True'.
2026-09-07 05:01:31 +07:00
Dai Ha 22cdebbdbe fleetd #369: state hermeticGitEnv's real scope, measured
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m32s
The comment claimed no test in this class can reach the real machine's home
directory. Measured on the merge: 53 of the 58 new GitWorktrees(...)
constructions here pass no env override, and stripping the override from
seedingGitWorktrees leaves the class green under the poison command that the
same comment cites as proof. Say what it covers and what it does not.
2026-09-06 20:32:21 +07:00
Dai Ha 154971c2b8 Merge #372: GitWorktreesTest no longer reads the operator's real git config
fleetd #369. Every git subprocess the test class starts now goes through one
gitProcessBuilder factory that applies the hermetic environment. Before this,
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus set nothing at all -- so the tests inherited the JVM's
whole real environment, including the operator's default excludes file. That
file applies with no core.excludesFile configured at all, and /dev/null for the
global config does not stop it.

Round 1 pinned this with a call-site count: exactly 2 literal
new ProcessBuilder( occurrences. That catches a NEW helper built the old way,
but it is a proxy, not the property. I measured the gap -- deleting
pb.environment().putAll(hermeticEnv()) from inside the factory left every call
site unchanged, the count stayed 2, and 1413 tests stayed green. Round 2 added
gitProcessBuilderCarriesTheFullHermeticEnvironment, which asserts on what the
factory actually hands to ProcessBuilder#start(). Both checks are kept: they
catch different regressions.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1414, Failures: 0, Errors: 0, BUILD SUCCESS
  poison control on the PRE-FIX file:
    XDG_CONFIG_HOME=<dir with a '*' git/ignore> mvn test -Dtest=GitWorktreesTest
    -> tests=59 failures=56, so the poison genuinely reaches these tests

Mutation run on merge, on a half neither the worker nor the reviewer touched --
made seedingGitWorktrees pass null instead of hermeticGitEnv(tmp), which strips
the hermetic environment from the PRODUCTION GitWorktrees instances rather than
from the test's own subprocesses: tests=61 failures=0 unpoisoned AND poisoned.
That half is unpinned. It is fleetd #362's protection, not this ticket's, and it
guards a different path -- the Java-side XDG read in
previouslyEffectiveExcludesFileContent, which no assertion observes. Out of
scope here; filed as a follow-up rather than held against this PR.

One javadoc sentence corrected in the merge: hermeticGitEnv claimed "no test in
this class can reach the real machine's home directory". 53 of the 58
new GitWorktrees(...) constructions in this file pass no env override at all, so
the claim is true of the 5 seeding sites and of every test-started subprocess,
not of the class.
2026-09-06 20:31:01 +07:00
Dai Ha fef287c346 Merge #371: a stale lead binding no longer swallows a worker's reply nudge
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m55s
fleetd #368. PrimaryRegistry.forgetDelegation fires only when a WORKER is
released, never when the delegating LEAD goes away. The map is keyed by the
worker, so a closed, crashed or relaunched lead left its bindings behind. A
stale non-null entry then beat nudgeTargetFor's single-primary fallback every
time -- and that fallback's own javadoc argues it is correct precisely in the
case the stale entry was hiding.

ReplyPushLoop now probes the recorded lead with agents.status before trusting
it, at all five entry points, and a lead that is really gone is forgotten so
resolution reaches the fallback.

Round 1 caught any RuntimeException and treated it as death, and death here
calls forgetDelegation -- destructive and permanent on ONE reading. That is the
#359 mistake repeated two days later: a socket blip or a decode error on a live
lead would silently unbind it forever. I measured the breadth was unpinned
(narrowing it left 1414 green), and round 2 narrowed it to the one affirmative
signal AgentControl.agentCall itself uses, agent_not_found. Every other failure
is treated as live, because guessing wrong costs one extra retry next tick while
guessing wrong the other way is unrecoverable.

The fallback it lands on is live, not stale: PrimaryRegistry.record overwrites
the single slot on every orchestration-side MCP call, and neither host pins
primary.terminal, so a relaunched lead re-registers on its first tool call.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1415, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge, on a half the worker did not touch -- dropped the
forgetDelegation call while still returning the fallback, so behaviour on the
first nudge is identical and only the self-healing is lost: 2 failures,
BUILD FAILURE. The cleanup is pinned, not just the fallback.
2026-09-06 20:13:07 +07:00
Dai Ha a37acd5ee3 fleetd #369 review round 2: pin the factory's behaviour, not its call count
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m26s
everyGitSubprocessGoesThroughTheHermeticFactory counts ProcessBuilder("git",
...) call sites, so it catches a new helper built the old way, but a
reviewer proved it does not catch gitProcessBuilder itself being gutted:
removing pb.environment().putAll(hermeticEnv()) from inside the factory
leaves every call site unchanged, the count stays 2, and the whole
unpoisoned suite stays green.

Add gitProcessBuilderCarriesTheFullHermeticEnvironment, which inspects what
the factory actually hands to ProcessBuilder#start(): every hermetic key
present with the isolating value, and XDG_CONFIG_HOME pointed inside the
class's own throwaway directory rather than left unset or pointing at the
operator's real one. This fails the moment the hermetic environment stops
being applied, on any machine, with no poison needed. Keep the call-site
count check too — the two catch different regressions.
2026-09-06 20:09:32 +07:00
Dai Ha d6ef0c8013 fleetd #368 review: only agent_not_found may forget a lead binding
CI / build (pull_request) Successful in 1m22s
CI / contract (pull_request) Successful in 1m22s
isLive treated any RuntimeException from the liveness probe as "the lead is
gone", which forgetDelegation then acted on destructively and permanently.
That made a transient herdr hiccup (socket blip, decode error) on a perfectly
live lead indistinguishable from the lead actually being dead — the same
one-bad-reading mistake #359 shipped a guard against for lead-tab liveness.

Narrow isLive to match AgentControl.agentCall's own rule: only an affirmative
HerdrException("agent_not_found") counts as gone. Every other failure is
treated as still live and the binding is left alone.

Adds aTransientLivenessFailureMustNotForgetABindingToAStillLiveLead, which
fails with the bare RuntimeException catch and passes with the narrowed one.
2026-09-06 20:08:18 +07:00
Dai Ha dfb70871b4 fleetd #360 follow-up: the install block must say loginctl enable-linger
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m54s
PR #370 shipped both units but its install block stopped at
`systemctl --user enable --now`. Without lingering a user manager starts at
your first login and stops at your last logout, so the units do not come back
after a reboot -- which is the whole reason this ticket moved fleet01 off the
setsid scripts.

It is easy to miss because leaving it out looks like success: `systemctl --user
enable` reports "enabled" and both units run while you stay logged in. The
issue named this and the PR did not carry it over.

fleet01 itself is fine -- measured `Linger=yes`, both units `enabled`. This is
about the next host that follows these instructions.

Comment only; SystemdUnitSafetyTest still 8 green (a commented line is not an
active directive).
2026-09-06 20:08:05 +07:00
Dai Ha 4ee7b16929 Merge #370: ship systemd units that do not silently disable the daemon
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m1s
fleetd #360. deploy/fleetd.service shipped four mount-namespacing directives
(ProtectSystem, ProtectHome, ProtectKernelTunables, ProtectControlGroups),
PrivateTmp=true, and an ExecStart that ran java directly. Each one starts green
and breaks the daemon in a way nothing logs: lsof goes blind so every caller is
resolved ANONYMOUS and refused; the member ZDOTDIR scrub becomes a no-op; every
credential is empty. deploy/herdr.service did not exist at all, though
fleetd.service's After=/Wants= already named it.

Both units are now the ones running on fleet01, comments included -- the
bisected lsof counts and the reasons live in the files, because the next person
to 'harden' this needs the reason, not the rule.

A unit file has no compile step, so SystemdUnitSafetyTest reads both units plus
deploy/herdr-inner.sh and fails on an active forbidden directive, on
PrivateTmp=true, on an ExecStart that skips the login shell, on a herdr-inner.sh
that does not exec a login shell, and on one that does not set a non-zero pty
size. Each message names the consequence. A vacuity guard pins that all three
files exist and that the DO-NOT-add comment block still mentions every forbidden
directive, so 'comment survives, directive does not' is actually exercised.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1420, Failures: 0, Errors: 0, BUILD SUCCESS

Round 1 shipped herdr-inner.sh untested; I mutated it (dropped both the login
shell and the stty sizing) and got 6 green. Round 2 added those two checks.

Mutation run on merge, on a half the worker never touched -- ProtectHome=read-only
added to herdr.service, the unit it only ever mutated fleetd.service for:
Tests run: 8, Failures: 1, BUILD FAILURE. The parameterisation really covers
both files.
2026-09-06 20:04:48 +07:00
Dai Ha 8557289dc0 fleetd #360 review round 2: extend SystemdUnitSafetyTest to herdr-inner.sh
CI / build (pull_request) Successful in 1m27s
CI / contract (pull_request) Successful in 1m30s
herdr.service's ExecStart only names deploy/herdr-inner.sh, so that script was the only
place herdr's login-shell and pty-size properties lived, and nothing was reading it -- the
same silent-at-startup shape the ticket was about, one file further down the chain. A
mutation dropping both the login shell and the stty sizing left SystemdUnitSafetyTest green.

Add herdrInnerScriptUsesALoginShell and herdrInnerScriptSetsANonZeroPtySize, extend the
vacuity guard to cover herdr-inner.sh too, and tighten execStartUsesALoginShell's check from
a bare contains("-lc") substring match to the same login-shell-invocation regex the new
checks use.
2026-09-06 19:54:36 +07:00
Dai Ha dd2efd8541 fleetd #369: make GitWorktreesTest hermetic against the machine's real git config
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Failing after 1m28s
gitOutput set GIT_CONFIG_GLOBAL/SYSTEM/TERMINAL_PROMPT but not XDG_CONFIG_HOME,
and status/fullStatus (plus every other raw git subprocess in this class) set no
isolation at all — inheriting the JVM's real environment, including the
operator's real ~/.gitconfig and default excludes file. Measured: with a
poisoned XDG_CONFIG_HOME, 56 of 59 tests failed.

Centralize every git subprocess this test starts through one factory,
gitProcessBuilder, which always applies the existing hermeticGitEnv isolation
(extended with XDG_CONFIG_HOME, the same fix #366 already applied to the
production-instance seam). Add a self-check test that counts direct
ProcessBuilder("git", ...) constructions in this file's own source and fails
if a future helper bypasses the factory, so the omission that caused this
ticket is caught by name instead of rediscovered on a poisoned machine.
2026-09-06 19:54:35 +07:00
Dai Ha 5af786d135 fleetd #368: a stale lead delegation binding must not shadow the primary fallback
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m55s
PrimaryRegistry.forgetDelegation only fires when a WORKER is released, never when
the delegating LEAD terminal itself disappears (closed, crashed, or relaunched).
A stale, non-null leadByTarget entry always beat nudgeTargetFor's single-primary
fallback, so a dead lead silently swallowed every reply nudge for its workers.

Fix: ReplyPushLoop now verifies (via the same agents.status check decide() already
uses every tick) that a recorded delegating lead is actually live before trusting
it. A dead lead is treated as if never recorded — self-healing the binding
(mirroring AgentControl.paneByTerminal's self-heal on agent_not_found) and falling
through to PrimaryRegistry's existing fallback.
2026-09-06 19:54:26 +07:00
Dai Ha 380eb63277 fleetd #360: fix systemd units that silently disabled caller identity and the credential scrub
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 2m12s
deploy/fleetd.service started clean on fleet01 but broke the daemon in three ways nothing
logs: ProtectSystem/ProtectHome/ProtectKernelTunables/ProtectControlGroups each put the unit
in its own mount namespace, which blinds fleetd's lsof-based caller lookup and falls every
caller back to ANONYMOUS; PrivateTmp=true silently no-ops the credential scrub the member
pane depends on; and running java directly from ExecStart skips the login shell that sources
the daemon's secrets, so it boots with empty credentials.

Replace the unit with the version verified working on fleet01 for a day, and add the
deploy/herdr.service companion unit it was already depending on via After=/Wants= but which
did not exist in the repo. Add deploy/herdr-inner.sh as the login-shell template
herdr.service's ExecStart wraps in a pty.

Add SystemdUnitSafetyTest (fleetd/src/test/java/dev/ltms/fleet/deploy) to read both unit
files from disk and fail if a forbidden mount-namespacing directive is active, PrivateTmp is
true, or fleetd.service's ExecStart does not go through a login shell -- the only guard
possible for a unit file with no compile step.
2026-09-06 19:45:23 +07:00
Dai Ha 6c61355f8f Merge #367: dead lead tabs stop accumulating, without destroying a live one
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m7s
fleetd #359. LeadTabScanner used to join labelled tabs straight to terminals
with no liveness check, and its javadoc excused that ("a stale name costs
nothing here"). It cost plenty: LeadCoordLoop reads that map to pick which
pane a peer message goes into, so a dead tab was a candidate it could pick.
LeadLauncher, meanwhile, had no cleanup path at all -- every reconcile that
found 0 live created another tab and left the old one.

Both now cross-check agent.list, and neither trusts a single reading of it.
That matters because the daemon's own evidence on fleet01 was agent.list
reporting 0 live while ps showed one real claude. A first cut of this fix
closed tabs on that single reading, which would have closed the operator's
live lead instead of leaving a spare tab. So: a dead tab is flagged, not
closed, and only closed when a later reconcile still finds it dead; and the
scanner grants one grace scan to a terminal it already knew was live.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1412, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge (PendingCloseMarker.strip made identity, so a flagged
tab stops matching its configured label): 4 failures, BUILD FAILURE.

Known and accepted: ensureLeads() runs at startup, so the second reading
arrives at the next restart. A tab that dies mid-session stays flagged and
open until then. Deliberate -- the bug is about repeated restarts, and one
leftover tab is cheaper than closing a live session on unverified evidence.
2026-09-05 16:00:04 +07:00
Dai Ha 96c406b968 #359 review: require two independent readings before destroying a lead's tab
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m33s
Finding 1 (LeadLauncher): closing a labelled tab on a single agent.list
miss could destroy a live lead's session — the ticket's own evidence
showed that exact signal missing a genuinely running agent. A dead
reading now only flags the tab (PendingCloseMarker); it is closed only
if a later, independently-connected reconcile still finds it dead
while flagged. A tab found live again has its flag cleared instead.

Finding 2 (LeadTabScanner): the new agent.list cross-check in scan()
was not covered by get()'s "keep the cache on a failed scan" contract,
which only fires on a thrown HerdrException. A successful-but-short
agent.list could silently drop a lead CallerResolver had already
resolved, demoting it to Role.WORKER. A terminal already reported live
now gets one grace scan before being dropped; a terminal never
reported live gets none, so the original #359 exclusion is unaffected.

Both mechanisms were mutation-tested: reverting either change turns
exactly its own new tests red and nothing else.
2026-09-05 15:54:52 +07:00
Dai Ha a5d81c3f70 Merge remote-tracking branch 'origin/main' into worker/359-dead-lead-tabs-f1253b-4 2026-09-05 15:28:49 +07:00
Dai Ha 92c0f164f1 Merge #366: seed the bridge's worker skills into every provisioned worktree
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m46s
fleetd #362 item 3. A member spawned against a repo that does not ship its
own .claude/skills/ could not load implementer, reviewer or hunter at all.
Every brief starts with "Load the <name> skill", and outside this repo that
line was silently a no-op. memberSkills: <dir> now copies those folders into
each provisioned worktree, skipping any name the target repo already ships.

Two review rounds, both about the same hazard: core.excludesFile is
single-valued, so pointing it at fleetd's own file would SHADOW the
operator's. It now composes instead of replacing, and the XDG default
excludes file is carried forward too.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1402, Failures: 0, Errors: 0, BUILD SUCCESS

Two mutations run on merge:
  drop the XDG fallback          -> 1 failure, BUILD FAILURE
  remove the composition itself  -> 2 failures, BUILD FAILURE
Both directions are pinned.
2026-09-05 13:32:16 +07:00
Dai Ha 9e813ec179 fleetd #362 review fix 2: route the XDG excludesFile fallback through gitEnv too
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m24s
Finding 1 (lead): the XDG fallback branch of previouslyEffectiveExcludesFileContent was
unpinned — deleting it left the suite green (Tests run: 1386, Failures: 0). Added
seedSkillsComposesWithTheXdgDefaultExcludesFileWhenNoneIsConfigured to pin it: isolates
XDG_CONFIG_HOME via the gitEnv seam at a temp dir carrying a synthetic git/ignore, points
GIT_CONFIG_GLOBAL at an empty file so core.excludesFile is genuinely unset (forcing the
fallback branch), seeds a skill, and asserts a file matching the XDG-default pattern still
reads as clean. Reverting the fix (mutating the fallback to resolve to "") turns this test
red with a real pasted failure (see PR body): "expected: <> but was: <?? xdg-fallback-marker>".

Finding 2 (lead, the one that actually needed a code fix): the fallback read XDG_CONFIG_HOME
and HOME straight from the JVM's own environment, not through the gitEnv seam every git
subprocess in this class already honours — so no test could isolate it, and on a machine
carrying a real ~/.config/git/ignore (this dev machine does), every seeding test silently
composed with that real file. Added resolveEnv/resolveHome, which check gitEnv first and
fall back to the JVM's real environment only when the seam doesn't supply a value (production
behaviour, where gitEnv is always Map.of(), is unchanged). Added a hermeticGitEnv() test
helper and routed every seeding test in GitWorktreesTest through it, so no test in the class
can reach the real machine's home directory for this fallback.

Also documents two non-defects the lead asked for one javadoc line each on: the composed
excludesFile is a snapshot taken at seed time, not a live reference to the operator's file;
and excludeSeededSkillsFromGitStatus assumes a fresh worktree (not idempotent, but the
double-seed path does not exist today, so no guard was added for it).
2026-09-05 13:23:39 +07:00
Dai Ha 395b3b5c46 #359: LeadTabScanner drops dead lead tabs; LeadLauncher closes them on relaunch
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m49s
LeadTabScanner.scan() joined labelled tabs to panes with no liveness check at
all, so a tab left behind by a crashed/relaunched lead read as a live lead
forever -- exactly the hazard its own javadoc predicted but excused. It now
cross-checks agent.list, the same signal LeadLauncher already trusted, and
drops any labelled tab with no agent running in it. That alone removes the
duplicate candidates LeadCoordLoop.resolveLocalLead() could pick from,
including the dead one its own WARN's advice (name a lead after
coordinator.selfId) could land a message in.

LeadLauncher never actually stopped the accumulation: relaunching on "0 live"
always created a brand new tab and left the old dead-labelled one right where
it was, so any restart that found 0 live for any reason (a real crash, or a
herdr read that missed a still-running agent) added one more dead tab,
forever. ensureLeads() now closes every dead-labelled tab for a lead as part
of the same reconcile that decides to relaunch, so at most one tab survives
per configured lead once a restart's reconcile has run.
2026-09-05 13:20:37 +07:00
Dai Ha 6e9e464d62 Merge #364: fleet_list reports lead-coordination state instead of guessing at it
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m56s
fleetd #361. fleet_list's coordination block now carries this daemon's own
mailbox state and one row per operator-declared peer, each as a tri-state
status (exists / absent / unknown) rather than a boolean. pending and
consumers appear only when status is "exists", so an unresolved probe can
never render as a measured zero.

Verified on this merge, not taken from the worker's report:
  mvn clean install -> Tests run: 1394, Failures: 0, Errors: 0, BUILD SUCCESS

Mutation run on merge (isMissingQueue always returns true, which restores
the exact defect the ticket fixes): 4 failures, BUILD FAILURE. The
discriminator's false branch is pinned.
2026-09-05 13:19:46 +07:00
Dai Ha 105c065615 Merge origin/main into worker/362-worktree-skills-c03e51-3 (picks up #363) 2026-09-05 13:18:32 +07:00
Dai Ha 01492059d4 fleetd #361 review round 2: pin isMissingQueue's false branch
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m34s
The reviewer's mutation (isMissingQueue always returns true) restored
the exact overstatement fleetd #361 exists to fix -- every declare
failure reading as a confirmed absence -- and still left mvn clean
install green (1389/1389), because no test drove a non-404 shape
through inspect(). The false branch was the whole discriminator
between MailboxState.absent() and MailboxState.unknown(), unpinned.

Widened LeadMailbox.isMissingQueue from private to package-private and
added LeadMailboxIsMissingQueueTest: five hermetic tests (no broker)
covering the true case and all three false shapes isMissingQueue's own
javadoc lists -- a different reply code, a ShutdownSignalException
whose reason isn't a Channel.Close, and an IOException with no such
cause at all (plus an IOException wrapping an unrelated exception
type). Re-ran the reviewer's exact mutation locally: 4 of 5 new tests
went red with the expected assertion messages; reverted, and mvn clean
install is green again at 1394/1394 (1389 + 5 new).

Also added a one-line javadoc note on LeadMailbox.inspect being honest
about which of its two RuntimeException catches is proven by a test
(the createChannel() one, end-to-end against a real broker) and which
stays purely defensive (the declare-site one, for a connection-drops-
mid-call race no test drives on purpose).
2026-09-05 13:12:24 +07:00
Dai Ha f84824ee29 fleetd #362 review fix: compose skill-seeding excludes with the operator's own excludesFile
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m54s
core.excludesFile is single-valued, so pointing it at fleetd's own seeded-skill exclude file
with --replace-all at worktree scope was SHADOWING whatever excludesFile the worktree already
resolved (an operator's global config, most commonly) instead of adding to it. This repo's own
.gitignore does not ignore target/ — only an operator's global excludesFile does — so every
worker's `mvn clean install` would make target/ show up as untracked, and CB-576's deliberately
untracked-inclusive hasUncommitted would then read every such worktree as dirty forever, so it
is never cleaned up.

excludeSeededSkillsFromGitStatus now reads whatever core.excludesFile resolves to BEFORE writing
anything (falling back to git's own $XDG_CONFIG_HOME/git/ignore default when the key is unset
entirely, per gitignore(5)), and writes that content into fleetd's own exclude file ahead of the
seeded skill patterns, so every operator-configured pattern keeps applying inside the seeded
worktree. Proven with a new test, seedSkillsComposesWithAnAlreadyEffectiveGlobalExcludesFile,
which isolates a synthetic "operator's global config" via a new gitEnv test seam on GitWorktrees
(GIT_CONFIG_GLOBAL pointed at a throwaway temp file, never the real machine's config) and drives
the real add() path end to end.

Also documents (FleetConfig javadoc + fleetd.example.yaml) that memberSkills copies every
non-hidden subdirectory of its source wholesale, with no per-file allowlist.
2026-09-05 13:10:27 +07:00
Dai Ha c4d40fbc2b fleetd #361 review: fix false-negative absent, uncancelled probes, and a throw contract gap
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m45s
Three findings from review of #364, fixed on the same branch:

1. LeadChannel.MailboxState.absent() was returned both for a genuinely
   absent mailbox AND for "the probe could not determine anything"
   (timeout, unreachable broker, other declare failure) -- exactly the
   overstatement #361 exists to fix, one level down. MailboxState now
   carries a Presence enum (EXISTS/ABSENT/UNKNOWN) with exists()/known()
   accessors; LeadMailbox.inspect classifies a real AMQP 404 (measured
   against a live broker, not assumed: an IOException wrapping a
   ShutdownSignalException whose Channel.Close reply code is 404) as
   ABSENT and everything else as UNKNOWN. fleet_list's mailbox/peer rows
   now render a "status" of exists/absent/unknown and only include
   pending/consumers when status is "exists", so an unresolved self- or
   peer-probe can never render as a measured zero.

2. FleetMcp.probe's get(timeoutMs) left a timed-out inspect() task
   running forever on its own virtual thread, holding the AMQP channel
   it had already opened -- against a hung (not down) broker this would
   orphan one channel per fleet_list call until the connection's
   channel-max was exhausted, breaking publish() too. probe() now holds
   the Future and calls cancel(true) on timeout/failure so the orphaned
   task is interrupted instead of abandoned, and now returns
   MailboxState.unknown() (never absent()) on timeout/exception.

3. LeadMailbox.inspect only caught IOException, but createChannel() on
   an already-closed connection throws AlreadyClosedException, an
   unchecked RuntimeException (measured against a live broker) -- so it
   could escape the "never throws" contract. Both places in inspect now
   also catch RuntimeException and report unknown().

Tests: MailboxState.exists()/absent()/unknown() call sites updated
across FleetMcpTest/FleetMcpLeadCoordTest; new hermetic tests cover the
tri-state fleet_list rendering (self-probe unknown, a peer that's
absent vs. one that's unknown) and probe cancellation (a LeadChannel
fake that blocks until interrupted, proving probe() doesn't just give
up on it); new @Tag("contract") LeadMailboxTest cases pin the real
exception shapes for both the 404 and the already-closed-connection
paths and prove inspect() reports unknown (never throws) when the
connection is already closed.
2026-09-05 13:01:48 +07:00
Dai Ha 7c684e40d3 fleetd #362 (item 3): seed .claude/skills/ into provisioned worktrees
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m46s
Add memberSkills: <dir> to FleetConfig. GitWorktrees#add copies each
skill folder from that directory into <worktree>/.claude/skills/ so a
member spawned against ANY repo — not only one that already ships its
own skills — can load a bridge skill (e.g. implementer). A skill the
target repo already carries is never overwritten.

Every seeded path is hidden from `git status` in that worktree ONLY,
via a --worktree-scoped core.excludesFile pointing at a file under the
worktree's own private git dir (outside the working tree, so it can
never be committed) — not the shared .git/info/exclude, which a linked
worktree resolves to the repo's common git dir and would otherwise leak
visibility changes into the primary checkout and every sibling
worktree. Proven with a real `git status --porcelain` in
GitWorktreesTest, not by reasoning.

Seeding is best-effort like the existing overlayParity/isolateToolSurface
steps: a missing/unreadable source or a copy/exclude failure is logged
and skipped, never fails the spawn. memberSkills is triaged as a
DEFERRED config key in ConfigRef (baked once into GitWorktrees at
startup, like worktreeGroup), with its own changedDeferredKeys branch
and coverage-test entries.
2026-09-05 12:54:01 +07:00
ltms 6938f52155 Merge #363: make the plugin visible, and fix the drift that made it unusable
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m28s
2026-09-05 07:50:19 +02:00
Dai Ha 0df34f3220 fleetd #361: close the lead-coordination visibility gap
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m34s
Lead-to-lead AMQP coordination had a send half with tools and a receive
half without. This closes three blind spots:

- LeadChannel gains inspect(coordId) -> MailboxState(exists, pending,
  consumers), implemented in LeadMailbox with a throwaway probe channel
  (never the long-lived publish/consume channels) so a passive-declare
  404 on a missing queue can never take down publish() on the same
  instance.
- FleetConfig.Coordinator gains peers: List<String> (defaults to empty,
  blank entries dropped) so a daemon can declare which peer coord-ids
  it expects to reach.
- fleet_list reports coordination state via a new CoordinationSource
  (own coord-id, own mailbox state, held messages as msgId/from/preview
  only, and one row per configured peer with reachability/pending/
  consumers), following the existing OutageSource/QuarantineSource
  "Source record with none()" idiom instead of growing listFleet's
  overload chain by another positional parameter. Every peer probe is
  bounded by a 1.5s timeout on a virtual-thread pool and degrades to
  absent rather than ever slowing or failing fleet_list.
- fleet_send{coordId}'s success text now says "durably confirmed by the
  broker" instead of "delivered", and warns (while still reporting
  success) when the target mailbox has zero consumers attached.

Tests: hermetic unit tests for exists/absent/zero-consumer/old-config-
no-peers-key/new fleet_list shape using FakeLeadChannel, plus a
@Tag("contract") LeadMailboxTest.inspectingAMissingMailboxNeverBreaks
PublishOnTheSameInstance proving the invariant against a real broker.
2026-09-05 12:44:29 +07:00
Dai Ha 457458437f #362: make the plugin visible, and fix the drift that made it unusable
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m31s
CB-527 shipped a Claude Code plugin and a marketplace in this repo. Nothing in
CLAUDE.md or docs/ ever named it, so a later session planned the same feature
from scratch. The wiki Features entry existed and was correct, but wiki/ is a
submodule whose pointer is never advanced, so no session reads it.

Visibility:
- CLAUDE.md addendum now names plugin/ and both structural limits, so every
  session sees it. This is the change that stops the rebuild happening again.
- wiki/11-Features.md records the rename and why the entry alone was not enough.

Drift (each measured against the code, not assumed):
- mount name fleetd -> fleet, matching PeerLauncher.MCP_MOUNT_NAME. The old name
  gave a lead with both a project .mcp.json and the plugin two mounts of one
  daemon and a duplicated fleet_* tool set.
- url is now ${FLEETD_MCP_URL} instead of a hardcoded address, so one plugin can
  serve hosts running the daemon on different ports. Plain ${VAR}, the form
  kb-alms proves works here; ${VAR:-default} is untested and not used.
- plugin claude-bridge -> fleet, marketplace claude-bridge -> fleetd, version
  0.2.0. Breaking for a 0.1.0 install: mcp__fleetd__* becomes mcp__fleet__*.
- README install path ltms/claude-bridge -> the fleet/fleetd remote.
- the setup skill's §5 told operators to pin primary.terminal:. CB-579 replaced
  that with fleet.leaders.*.tab. Replaced, with the duplicate-tab warning (#359).

Scope: the plugin is lead-side only, and cannot be otherwise. The launcher adds
--agent only when <worktree>/.claude/agents/<role>.md exists in the member's own
tree (ClaudeCodeLauncher.java:371,391), and a member's CLAUDE_CONFIG_DIR points
at its profile's config dir (ClaudeCodeLauncher.java:285), so a member never
reads the operator's plugin store. On this Mac all four Claude profiles set
configDir, and the four ccs instances hold four separate copies of the plugin
store -- same md5, different inodes. Seeding member skills through the worktree
is #362 scope item 3, implemented separately.

Note for anyone verifying a plugin: `claude plugin validate` does NOT read
.mcp.json. Replacing it with `{ this is not json at all` still passes, exit 0.

Refs #362, #359
2026-09-05 12:42:20 +07:00
Dai Ha 3759c41f99 Merge #354: the redeploy health gate classifies AMQP errors instead of counting them
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m36s
The gate counted ERROR lines since RESTART_MARK. On a laptop that idle-sleeps after one
minute on battery that meant 6 ERROR lines for an AMQP link that recovered every time,
and a gate that cries wolf is a gate nobody reads.

It now reports three states: no errors; only errors proven to have recovered (quiet, and
the gate passes); anything else (the old warning, unchanged). Attribution is per
connection, using the names #356 put into the log -- a lead-mailbox recovery can no
longer clear an unrecovered reply-inbox reset. A candidate carrying neither name is
unattributable and stays LOUD.

Two earlier rounds were rejected. Round 1 was inert: it matched nothing in the real log,
because the layout abbreviates the logger and 'Connection reset' sits in the stack trace,
not on the ERROR line -- my brief had pointed the worker at fleetd.out, which is untracked
and so absent from its worktree. Round 2 was correct and honest but could not attribute
anything, which is what motivated #356.

Verified on merge beyond the worker's own mutations:
 - ran the classifier against the REAL log, which is still in the pre-#356 format: 6 total,
   0 recovered, 6 unexplained. Old-format lines carry no connection name, so they stay loud
   -- the safe direction, on genuine data rather than a fixture.
 - adversarial fixture the worker did not write: a lead recovery BEFORE any failure banks
   no credit; 2 inbox resets with 1 recovery leaves 1 unexplained; a non-AMQP ERROR stays
   loud. total=3 recovered=1 unexplained=2, as intended.
 - RESTART_MARK still anchors the scanned region.

Caveat carried from the PR: the patterns are source-derived. The daemon has not been
redeployed, so they are not yet confirmed against a live log.
2026-09-05 06:09:29 +07:00
Dai Ha 09159f2857 Classify named AMQP recovery errors
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m16s
2026-09-05 06:05:52 +07:00
Dai Ha 29cd1194c2 Merge remote-tracking branch 'origin/main' into worker/errscan-bed2ca-2 2026-09-05 06:02:29 +07:00
Dai Ha 815e8f8b23 Merge #356: name the AMQP connection in its own log lines
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m38s
Both connections were already named at newConnection() -- 'fleetd-reply-inbox' and
'fleetd-lead-mailbox' -- and neither name ever reached the log: 0 occurrences in
fleetd.out, and both connections logged under the same thread name
'[AMQP Connection 10.10.20.13:5672]'. So when one of the two died and never came
back, the log could not say which.

AmqpConnectionFailureLogger extends DefaultExceptionHandler and overrides only the
protected log(String, Throwable) sink that every handle* method calls virtually, so
the identity is added without changing any handler action.

My brief caused a defect here and the correction is the interesting part. I told the
worker the client 'currently uses ForgivingExceptionHandler', read off the log line
c.r.c.i.ForgivingExceptionHandler -- which names where the LOGGER FIELD is declared,
not the instance's class. javap on the jar shows ConnectionFactory's constructor does
'new DefaultExceptionHandler', and DefaultExceptionHandler extends StrictExceptionHandler
extends ForgivingExceptionHandler. The first version therefore extended the base and
silently dropped strict channel-closing on four listener/consumer paths. Now pinned by
a type assertion on both factories plus a behavioural test that handleConsumerException
still closes the channel once.

Verified on merge with a mutation the worker did not run: it mutated the parent class,
so I mutated the copied private-static isSocketClosedOrConnectionReset in the DANGEROUS
direction (always true => every failure logs at WARN and vanishes from the redeploy
gate's ERROR count). Caught: 'inbox failure line ==> expected: <ERROR> but was: <WARN>'.

Merged main in first; the auto-merge compiled. 1379 green, unpiped.
2026-09-05 06:01:30 +07:00
Dai Ha 1e60ac0745 merge main for verification 2026-09-05 05:59:22 +07:00
Dai Ha 650a4c146b Merge #357: a FleetConfig component dropped by withDefaults() now fails the build
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m44s
Test-only. FleetConfig.java itself is unchanged.

The hazard is the back-compat constructor ladder (21/20/18/17/16/15/14 alongside the
22-arg canonical). Add a component and leave withDefaults()'s call at the old arity and
it binds to a back-compat constructor: it compiles, the suite passes, and the new key is
silently defaulted away on every load().

Verified on merge with a mutation the worker did not run: I made withDefaults() issue a
21-arg call, reproducing the real binding rather than an explicit null. It compiled, and
the guard failed by name -- 'memberLoginShell: ... a component silently dropped by
withDefaults(), the shape of the defect this test exists to catch'.

Exclusion list is empty and its size is pinned, so a future exemption must touch an
assertion rather than grow quietly.
2026-09-05 05:55:22 +07:00
Dai Ha 23f299e105 Preserve strict AMQP exception handling
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m46s
2026-09-05 05:53:44 +07:00
Dai Ha dbf6fef0e9 config: guard FleetConfig.withDefaults() against silently dropping a component
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m53s
Adding a component to FleetConfig follows an established pattern: the
record grows by one arg, and a back-compat constructor is added at the
OLD arity so existing callers keep compiling. That back-compat
constructor also silently captures withDefaults()'s own literal-arity
'return new FleetConfig(...)' call the next time this happens, since
that call is now a legal overload match too. It compiles, every other
test passes, and the new component is defaulted away on every load().
This is not hypothetical - it happened live while building the (now
parked) idle-sleep-guard PR, caught only because that branch's own new
tests asserted on the new field.

Add a reflective test that builds a FleetConfig through the true
canonical constructor (resolved by record-component types, not arg
count - the same pattern ConfigRefTopLevelReportingCoverageTest already
uses in this file) with a real, non-null value in every component, runs
the real withDefaults(), and asserts every value survives unchanged.
Never hardcodes the arity - it enumerates
FleetConfig.class.getRecordComponents() - so it keeps working as the
record grows. No back-compat constructor is touched or removed.
2026-09-05 05:51:59 +07:00
Dai Ha d292522d00 Name AMQP connection failure logs
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m29s
2026-09-05 05:45:37 +07:00
Dai Ha 0241e0d3a8 Keep unattributed AMQP errors loud
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m30s
2026-09-05 05:36:53 +07:00
Dai Ha e4973eb8a4 Classify recovered AMQP redeploy errors
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m51s
2026-09-05 05:29:53 +07:00
Dai Ha b6b88c5f1c #334: pin the fresh-owner gate on ask()'s timeout teardown
CI / build (push) Successful in 2m7s
CI / contract (push) Successful in 34m37s
2026-09-04 17:12:49 +07:00
Dai Ha 86dddfe240 Merge #334: ask()'s timeout closes the turn before it forgets the task mapping 2026-09-04 17:08:40 +07:00
Dai Ha 0d5944af63 fleetd #334: close ask()'s turn before forgetting its Task, closing the last stranding window
CI / build (pull_request) Successful in 2m14s
CI / contract (pull_request) Successful in 19m14s
ask()'s TimeoutException catch used to run clearAsyncQuestion(turnId, true) -- forgetting the
Task's asyncTasksByTurn mapping -- before rendezvous.closeAsk(turnId) ran in the shared finally.
Between those two calls the ask was still "answerable" (askSession(turnId) non-null) but the Task
mapping was already gone, so a racing answer() call found task == null, skipped
finishAsyncTask, and stranded the async ticket at PENDING even though answer() itself reported a
result. #329 fixed one step of this same race; this closes the remaining one.

The fix reorders the fresh owner's teardown: closeAsk runs first, then markAskTimedOut and
clearAsyncQuestion. A racing answer() call now either sees the ask still open (and the Task
mapping guaranteed intact) or sees it already closed (STALE_TURN, before it ever reaches
asyncTasksByTurn). It also gates the whole block by ticket.fresh(), matching the invariant the
finally block already states ("only the fresh owner tears down the shared turn") -- a duplicate
coalesced ask() timing out no longer forgets bookkeeping the fresh owner still needs.

Adds aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket, which pins the exact window
with a new test-only hook (askTimeoutRaceHookForTest) and proves both invariants: a late answer()
racing the timeout sees STALE_TURN, and the async ticket still resolves DONE from the worker's
real reply. Reverting the reorder (verified locally, not committed) makes this test fail with
"expected STALE_TURN but was TIMED_OUT_WORKING".
2026-09-04 17:06:15 +07:00
Dai Ha 4a5030a5c6 #348: drop a chrome skip that cannot fire, and pin the live pattern shape
CI / contract (push) Successful in 2m10s
CI / build (push) Successful in 2m12s
2026-09-04 16:57:36 +07:00
Dai Ha c1c8794c48 Merge #348: a member's prose about a usage limit no longer quarantines a credential 2026-09-04 16:53:19 +07:00
Dai Ha 65a78932c1 Merge #335: a per-task cleanup throw in abandon() no longer strands the tasks behind it
CI / contract (push) Successful in 1m24s
CI / build (push) Successful in 2m8s
2026-09-04 16:44:04 +07:00
Dai Ha f429ca1a50 Avoid exhaustion cooldown for member prose 2026-09-04 16:43:16 +07:00
Dai Ha 73aab3f83e Merge #342: teardown resolves the pane's real tab instead of trusting the delegate's placement config
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m47s
2026-09-04 16:38:31 +07:00
Dai Ha 887aca0183 fleetd#335: abandon()'s per-task cleanup and sendAsync's terminal hook must not swallow throws
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 2m21s
Site 1 (abandon()'s matching loop, reachable): the recovery/put-back branch calls
inbox.publish, which AmqpReplyInbox implements as a real broker round trip that
throws IllegalStateException on an unroutable/unconfirmed/interrupted publish.
An uncaught throw there aborted the loop, stranding every task after it in
`matching` PENDING forever. Fixed by recording each task's own future.complete()
result before any cleanup runs, then wrapping the cleanup in try/catch so one
task's failure cannot stop its siblings from getting their outcome. Reaching the
throwing branch by real timing needs a race the file's own #137 follow-up already
found unreachable through the public API, so the reproducing test uses a
test-only hook (same technique as the existing fleetd #324/#329 hooks) to inject
the throw at that exact point.

Site 2 (sendAsync's task.future.whenComplete, reachable): the returned stage is
discarded, so an uncaught throw from pushLoop.onTicketTerminal vanished with no
log line. Reproduced for real: Fleetd's shutdown hook runs messages.close()
(stops the async executor from taking new work, but does not cancel a send
already in flight) before pushLoop.close() (shuts its scheduler down
immediately) — a ticket completing in that window makes onTicketTerminal's own
scheduler.schedule(...) throw a genuine RejectedExecutionException. Fixed with a
try/catch(Throwable) plus log.error inside the whenComplete action.

Site 3 (the two `finally { asyncTasksByWaiter.remove(reply); rendezvous.close(...);
}` blocks in send() and answer()): read Rendezvous.close/closeAsk and the
ConcurrentHashMap operations behind them — both are plain map ops on a non-null
key with no user-overridable code, so neither can throw. Left unchanged; not a
defect.

Mutation-proven: reverting either fix reproduces the failure it exists to catch
— removing site 1's try/catch aborts abandon() with the injected exception
(MessageServiceTest#aPerTaskCleanupFailureDoesNotStrandTheRemainingMatchingTasks
errors); removing site 2's try/catch leaves the RejectedExecutionException
unlogged (MessageServiceTest#aTicketTerminalPushFailureDoesNotVanishSilently
fails its log assertion). Full suite: mvn clean install, Tests run: 1365,
Failures: 0, Errors: 0, BUILD SUCCESS.
2026-09-04 16:38:14 +07:00
Dai Ha 3fd23ecafa fleetd #342: base tab-cleanup teardown on the pane's real placement, not a delegate's static config
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Failing after 1m38s
HerdrPeerLauncher.stop() used to gate spaces.locatePane() on usesTabPlacement(),
which reads the delegate's OWN configured profiles. When CompositePeerLauncher's
single-daemon stop() shortcut hands a pane to a delegate that never spawned it
(spawnedBy empty after a daemon restart, herdrDaemonCount()==1), that delegate's
placement config says nothing true about how the pane was actually placed, and a
dedicated tab could be skipped and leaked.

Resolve the tab unconditionally instead — WorkspaceControl#locatePane already
tolerates a missing pane by returning null — and let the existing single-occupant
check (tabPaneCount()==1) be the only thing that decides whether to close it, same
as it already protects a shared tab regardless of declared placement.

Adds a mixed-placement CompositePeerLauncherTest (every existing stop-fallback test
configured both adapters as tab placement, so the mis-routing never showed) and
updates FleetAppTest#stopWorkerInPanePlacementClosesOnlyThePane, whose old
assertion (no pane.get on pane placement) documented exactly the skip this fix
removes.
2026-09-04 16:36:11 +07:00
Dai Ha f379847942 Merge #345: the timeout path's use of Cancellation.DELIVERED is now pinned
CI / contract (push) Successful in 47s
CI / build (push) Successful in 2m8s
2026-09-04 16:32:45 +07:00
Dai Ha ea12107497 fleetd #345: test timeout cancellation race
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 2m3s
2026-09-04 16:30:46 +07:00
Dai Ha 591df91de1 Merge #337: the deferred key set proves its own reporting coverage too
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m57s
2026-09-04 16:16:20 +07:00
Dai Ha 6a814176f0 #339: the start-of-line check must skip terminal chrome
CI / build (push) Failing after 1m30s
CI / contract (push) Successful in 1m54s
#339 stopped a member's own prose about an error from recording a credential
outage, by requiring the pattern at the start of its matched line. A bare
lookingAt also rejected a genuine error line rendered as

    | 503 Service Unavailable: upstream credential rejected

The send still failed, but the outage was never recorded. That is the false
negative #339's own invariant 3 named as worse than the false positive it set
out to fix: an unrecorded outage leaves the fleet spawning into a dead
credential.

Measured with a throwaway probe on the raw-scrape path, whose own comment says
to expect leading chrome there: kind=FAILED, sinkNotified=0.

startsWithBackendError now skips a leading run of non-letter, non-digit
characters before the check. That keeps #339's intent: prose still does not
match, because there the pattern sits after words rather than after chrome.
The worker's own prose test still passes.

Mutation: restoring the bare lookingAt fails the new test.
2026-09-04 16:13:19 +07:00
Dai Ha d11d1d157c Merge #339: a backend-error text match must be at the start of its line before it records a credential outage 2026-09-04 16:08:46 +07:00
Dai Ha d057d56156 #341: pin the noise control, and the reverse policy order
The fix reported every distinct unprotected name, but nothing held it there.

Mutation: replacing .filter(unprotectedGapNamesWarned::add) with a filter that
adds and always returns true - so every name is logged on every spawn - left
all 1358 tests green. The Set behaved; nothing proved this class used it as a
guard rather than as a record.

Two tests added:

- theSameUnprotectedNameIsWarnedAboutOnlyOnceAcrossSpawns pins invariant 1, the
  noise control. It now fails on that mutation, showing both duplicate WARNs.
- anAllowListWarnDoesNotSuppressALaterDenyByDefaultWarnForADifferentName covers
  the reverse policy order. The defect was found going deny-by-default then
  allow-list; a guard fixed in one direction is not fixed in the other.
2026-09-04 16:07:50 +07:00
Dai Ha d703ce1313 #337: extend ConfigRefTopLevelReportingCoverageTest to DEFERRED_KEYS
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m27s
ConfigRefTopLevelReportingCoverageTest (added by #333) proved every COLD_KEYS
and SPLIT_KEYS member has a real comparison behind it, but left
DEFERRED_TOP_LEVEL_KEYS unexercised. Re-measured by mutation (drop each
key's branch from changedDeferredKeys, run the suite, restore): 6 of the 11
deferred keys had no behavioural test naming them — guard, leadHeartbeat,
worktreeRoot, spawnReadyTimeoutMs, spawnReadyPollMs, quarantineCooldownSeconds
— which corrects the issue's own guessed list in two ways: lifecycle is
actually covered (ConfigRefTest.aDeferredChangeIsAppliedAndReported), and
worktreeRoot was missing from the issue's list entirely.

Promoted the test-side DEFERRED_TOP_LEVEL_KEYS copy into ConfigRef.DEFERRED_KEYS
(package-private, alongside COLD_KEYS/SPLIT_KEYS) so the reflective test reads
the same set changedDeferredKeys is compared against, and made
changedDeferredKeys package-private so the test can call it directly. Every
DEFERRED_KEYS component turned out to be a scalar or a simple record, so no
exclusion set was needed.

Mutation proof: dropping guard's branch from changedDeferredKeys leaves the
whole suite green except the new
everyDeferredKeyIsActuallyReportedByChangedDeferredKeys test, which fails
naming guard exactly.
2026-09-04 16:06:09 +07:00
Dai Ha b32a30fd47 Merge #341: warn once per distinct unprotected credential name, not once per launcher 2026-09-04 16:02:51 +07:00
Dai Ha 464dbc0930 fleetd#341: a per-name guard so a later spawn's different unprotected name still warns
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m29s
unprotectedGapLogged was one AtomicBoolean guarding two WARN branches in
logCredentialGap that name different env var names (the allow-list
keptByDerivedList branch, and warnGapUnprotected's deny-by-default /
non-zsh-fallback branch). memberCredentials is a live, re-read-per-spawn
supplier, so between two spawns a policy reload can change which names are
in the gap: spawn 1 warns about name A and trips the shared flag, and
spawn 2's gap containing a different name B never gets its WARN.

Replace the AtomicBoolean with unprotectedGapNamesWarned, a
ConcurrentHashMap-backed Set<String> guard keyed per name (same shape as
OpenCodeLauncher.modelCheckSkippedWarned), so each distinct credential-shaped
name is warned about exactly once, ever, regardless of which branch or
which spawn first reports it. allowListGapLogged (the separate INFO guard,
#192) is untouched. Neither WARN's wording changed.
2026-09-04 15:59:43 +07:00
Dai Ha a8cadd9150 Merge #338: a timed-out queued send cancels its message instead of leaving it to be delivered later
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 2m12s
2026-09-04 15:57:39 +07:00
Dai Ha 57b8c0b56d fleetd #339: guard backend error sink
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m50s
2026-09-04 15:57:18 +07:00
Dai Ha 147f50c19e #338: cancel timed-out queued deliveries
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Failing after 2m11s
2026-09-04 15:56:10 +07:00
Dai Ha eee4d576a2 #333: the Cold doc bullet listed four keys, COLD_KEYS has five
CI / build (push) Successful in 1m52s
CI / contract (push) Successful in 2m8s
memberHerdrSocket was missing from the prose. The #333 worker spotted it and
correctly left it alone as outside its scope.

Fixed by pointing the bullet at COLD_KEYS instead of re-listing its contents,
so the prose and the set cannot drift apart a second time.
2026-09-04 15:43:39 +07:00
Dai Ha b4f9d7f53a Merge #333: fleet: is a split key, and split membership now proves a reporting branch exists 2026-09-04 15:39:02 +07:00
Dai Ha 3aca53b967 fleetd#333: fleet.leaders is split too, and split membership now proves reporting exists
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Failing after 1m59s
F1: fleet: was sitting in ConfigRefTopLevelCoverageTest's HOT_EXCLUDED_TOP_LEVEL_KEYS
escape hatch, even though fleet.leaders is read only at startup (LeadTabScanner's
identity map, LeadLauncher.ensureLeads) while the rest of fleet: (role pools,
charters, tabLabel) is live. A reload changing only fleet.leaders reported a bare
"config reloaded" -- the operator edits a lead's tab: label, sees the reload
succeed, and the pane keeps resolving as a worker. Moved fleet into
ConfigRef.SPLIT_KEYS; changedSplitKeys now compares fleet.leaders specifically
(not the whole Fleet record, which would over-claim "restart" for a tabLabel-only
change) and names both halves in the message.

F2: membership in SPLIT_KEYS/COLD_KEYS never proved a matching branch existed in
changedSplitKeys/changedColdKeys -- measured by dropping the coordinator branch
while leaving "coordinator" in SPLIT_KEYS: both ConfigRefTopLevelCoverageTest and
the in-method "kept in step" assert stayed green. Added
ConfigRefTopLevelReportingCoverageTest, the ConfigRefProfileCoverageTest mechanism
one level up: reflection-built FleetConfig pairs that differ in exactly one
top-level component, calling the real (now package-private) changedColdKeys/
changedSplitKeys to prove each COLD_KEYS/SPLIT_KEYS member is actually reported.
Scoped to split+cold, not deferred -- see the new test's javadoc for why and what
that leaves open.

Both findings carry a behavioural test in ConfigRefTest plus a mutation proof
(revert -> real failure -> restore) recorded in the PR description.
2026-09-04 15:36:40 +07:00
Dai Ha 4aa1fae296 #329: a null task in answer() is not only "never an async ticket"
CI / contract (push) Successful in 59s
CI / build (push) Successful in 2m26s
The comment that landed with #329 said a null task means the turn was never
an async ticket. That is wrong, and it makes the guard read as complete.

A genuine async ticket also reaches answer() with task == null. ask() runs
clearAsyncQuestion(turnId, true) in its catch block, which drops the
asyncTasksByTurn entry, while rendezvous.closeAsk(turnId) runs later, in its
finally. Between the two the ask is still answerable and the map entry is
already gone, so answer()'s lookup returns null and the ticket is stranded.

Measured with a throwaway probe firing only that first half: answer() reported
REPLIED while the ticket stayed PENDING with a null reply. The probe used
forgetTurnForTest, so it omits markAskTimedOut; that cannot change the outcome,
because askTimedOut is read only by askAnsweredAsyncTasks, which reply() never
reaches while answer()'s own waiter is live.

Open as fleetd #334. The comment now says so.
2026-09-04 15:33:57 +07:00
Dai Ha 0c865032f9 Merge #329: log an exception thrown after a ticket resolves, complete the ticket from the task answer() already holds, and read orphan.turnId once 2026-09-04 15:26:45 +07:00
Dai Ha ea41bbf6b9 fleetd#329: fix silent async-ticket bugs in MessageService (F1/F2/F3)
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m28s
F2 (sendAsync executor catch): log when completeExceptionally returns
false, so an exception thrown after finishAsyncTask already completed
the ticket's future is no longer silently lost.

F1 (answer()'s stranded async ticket): reuse the Task reference answer()
already looked up before rendezvous.answerAsk(), instead of a second
asyncTasksByTurn lookup by turnId in finishAsyncTask. The second lookup
raced ask()'s unlocked timeout cleanup, which could forget turnId first
and leave the ticket stuck PENDING even though answer() itself returned
REPLIED. The #282 chained-ask guard is unaffected: it is still keyed on
result.outcome() == QUESTION, not on this lookup. Removed the now-unused
finishAsyncTask(String, Reply) overload.

F3 (reply()'s orphan recovery path): read orphan.turnId once instead of
twice, closing the same double-read shape fleetd #324 fixed in
finishAsyncTask.

Each fix has its own test plus a test-only race hook (mirroring #324's
finishAsyncTaskRaceHook) to force the exact interleaving deterministically.
Mutation-tested each fix by reverting it, confirming the real failure
(swallowed exception / PENDING ticket / NullPointerException), then
restoring it.

mvn clean install: Tests run: 1345, Failures: 0, Errors: 0, Skipped: 0,
BUILD SUCCESS.
2026-09-04 15:19:38 +07:00
Dai Ha 7b918c51ff Merge #330: a fourth reload class for split keys, and a top-level coverage checker
CI / build (push) Successful in 1m46s
CI / contract (push) Successful in 1m48s
2026-09-04 15:19:14 +07:00
Dai Ha 554395b104 fleetd#330: split reload class for health/coordinator + top-level coverage
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m17s
Unit 1: ConfigRef gets a fourth reload class, `split`, for keys read both
off the startup snapshot and live off config.get() at different sites
(health:, coordinator:). A split change is accepted (Outcome.applied()
stays true) and reported by name, naming which half is live and which
needs a restart, via a new Outcome.split() field kept separate from
deferred() since the two carry different guarantees for any caller that
branches on them, not just prose in summary(). Class doc updated: four
classes now, denominator note no longer calls health/coordinator
undecided.

Unit 2: ConfigRefTopLevelCoverageTest enumerates FleetConfig's 22
top-level record components and requires each to sit in exactly one of
COLD_KEYS, a pinned "compared in changedDeferredKeys" set, SPLIT_KEYS, or
a pinned hot-exclusion escape hatch — printing its own denominator and
pinning the escape hatch's exact contents the way #323 asked for.

Deviates from the issue's starting values by one key: `profiles` moves
from the suggested Hot bucket into the deferred bucket, because
changedDeferredKeys demonstrably compares it (add/remove and launch
settings), and citing "read live off the config supplier" for the whole
key would be false — most Profile fields are not read live, only
weight/maxLoad/credentialId are (and those are already covered by
ConfigRefProfileCoverageTest). Cold=5, split=2, deferred=11, hot=4,
total=22 — verified against the record and against ConfigRef's code, not
copied from the issue.
2026-09-04 15:16:45 +07:00
Dai Ha 823976c1b5 #326: state primary's real consequence, and write down the denominator
CI / contract (push) Successful in 56s
CI / build (push) Successful in 2m5s
The merged javadoc said a changed primary.terminal leaves a lead 'unresolved as
primary until a restart'. That over-claims. CB-532 made the pin deprecated:
identity comes from leaders:/leadScan:, and Fleetd.java:511 warns about the pin
at startup. A changed pin still needs a restart, but for the fallback nudge
destination, the deprecated identity path, and pushReminders/pushBackoffMs -
not for a lead that uses leaders:.

Also record what I measured. FleetConfig has 22 top-level components; four are
named nowhere in ConfigRef. memberCredentials and memberLoginShell are hot and
correctly absent (both read live off config.get() at spawn). health and
coordinator are undecided, not hot. 'Absent' looks the same for both kinds, and
twice now the forgotten kind hid among the correct kind.
2026-09-04 14:58:07 +07:00
Dai Ha c8388a7f92 Merge #326: primary and configReload are deferred keys, so a reload says a restart is needed 2026-09-04 14:53:50 +07:00
Dai Ha 02e6aef98c Merge #324: read task.turnId once in finishAsyncTask, so a concurrent clear cannot make the removal key null
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m29s
2026-09-04 14:44:42 +07:00
Dai Ha 6d493bc7bb fleetd#326: classify primary and configReload as deferred top-level keys
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 1m44s
ConfigRef.changedDeferredKeys only classified seven top-level FleetConfig
keys (#323 fixed the profile side). Two more keys are read only off the
startup snapshot and were missing:

- primary: Fleetd.java:506/519/520 feed PrimaryRegistry and ReplyPushLoop
  at construction; neither is rebuilt on reload.
- configReload: Fleetd.java:679-680 decide once at startup whether to
  build a ConfigWatcher at all, and with what interval; the watcher that
  would apply a later change is itself built once, so it is deferred
  (not cold — no already-open resource goes inconsistent, a running
  watcher just keeps its original settings).

health and coordinator are deliberately left unclassified: both are read
both off the startup snapshot AND live off the config supplier at a
second call site, so no single bucket is correct for either — see the
PR body for the options writeup and the coordinator.uriEnv exposure
question the issue asked to be answered.

Each fix is proven with a failing-first test in ConfigRefTest and a
revert-quote-restore mutation check (see PR body for the transcripts).
2026-09-04 14:42:24 +07:00
Dai Ha e5cb51a90e #324: read task.turnId once in finishAsyncTask to stop an NPE from ask()'s unlocked forgetting
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 2m1s
answer() holds sessionLocks while finishAsyncTask reads the volatile Task.turnId twice — once to
check it is non-null, once as the ConcurrentHashMap.remove key. ask()'s own timeout path mutates
the same field with no lock, via clearAsyncQuestion(turnId, true). volatile makes each read fresh
but not the pair atomic, so the field can go null between the two reads and remove(null, task)
throws NullPointerException on the lead's own answer() call, even though the reply already
completed on the line above.

Capture task.turnId into a local once and use that for both the check and the removal.

Added a package-private test seam (finishAsyncTaskRaceHook + forgetTurnForTest) so a test can force
the exact interleaving deterministically, by running the identical clearAsyncQuestion(turnId, true)
cleanup ask() uses, at the point between finishAsyncTask's former two reads. Both are inert (null)
in production.
2026-09-04 14:38:52 +07:00
Dai Ha e545c08082 #323: pin the exclusion set — the coverage mechanism's own escape hatch, found by mutation
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 2m10s
2026-09-04 14:32:08 +07:00
Dai Ha efa0deb9b2 Merge #323: the reload classifier proves its own coverage instead of claiming it 2026-09-04 14:29:10 +07:00
Dai Ha ca47e90c01 fleetd#323: close the reload-classifier drift with a reflection coverage test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 2m6s
ConfigRef.sameLaunchSettings' javadoc claimed it compares every component
the launcher reads at spawn. It missed ideProjectDir, ideOpenCommand and
autoCompactWindow, and changedDeferredKeys separately missed worktreeGroup
(baked into the same GitWorktrees as worktreeRoot, Fleetd.java:251). A
reload that changed only one of those keys reported "config reloaded" with
nothing deferred, and the running daemon kept the old value.

Fix the four instances, and add ConfigRefProfileCoverageTest: it enumerates
every FleetConfig.Profile record component by reflection, mutates each one
not in the new ConfigRef.LAUNCH_SETTINGS_EXCLUDED set on a base profile,
and asserts sameLaunchSettings actually notices — so a fifth missed field
fails the build by name instead of drifting silently. It also prints its
own denominator (26 components, 23 compared, 3 excluded) per the ticket's
requirement that a checker must be able to state what it checked.

Also add the `profile` field itself to the comparison (it was neither
compared nor excluded before this fix — the coverage test surfaced it).

Rewrote the sameLaunchSettings javadoc to describe what the coverage test
actually guarantees instead of repeating the unchecked claim.
2026-09-04 14:26:09 +07:00
Dai Ha fa1f49675b Merge #318: a delivery landing after release is refused, not parked in a map nobody reads
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 2m8s
2026-09-04 14:21:29 +07:00
Dai Ha 8426c3528f #316: pin the fail-toward-preserve rule on the late re-check, found by mutation
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m42s
2026-09-04 14:20:36 +07:00
Dai Ha 65f98ba910 Merge #316: the dirty check that authorises the worktree removal is taken after the worker stops 2026-09-04 14:17:13 +07:00
Dai Ha d05205d1eb Add a hunter skill: a sweep and a diff review are different jobs with different output contracts
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m49s
2026-09-04 14:13:08 +07:00
Dai Ha 667254df47 #316: re-check worktree dirtiness after the pane stops, before removing it
CI / contract (pull_request) Successful in 39s
CI / build (pull_request) Successful in 1m46s
SessionManager.releaseRemoved read hasUncommitted() once, while the worker
could still write, then used that stale boolean after launcher.stop() to
authorise `git worktree remove --force`. The same stale read also gated
trySnapshot, so a worker that wrote between the read and the stop lost its
work with neither a preserve nor a snapshot.

Add a second, best-effort hasUncommitted read immediately before the
removal, taken only on the path that is actually about to delete something
(never on a release that already decided to preserve, and never for
SHUTDOWN, which preserves unconditionally). If the tree is now dirty,
preserve it and attempt a fresh snapshot, since the original snapshot never
ran when the pre-stop read said clean. A failing re-check also preserves,
matching the existing CB-581 fail-safe rule.
2026-09-04 14:12:49 +07:00
Dai Ha 2926cd1784 #318: release() no longer strands a delivery that lands while it is running
CI / build (pull_request) Failing after 1m59s
CI / contract (pull_request) Successful in 2m14s
AmqpReplyInbox.release() used held.remove(target) then iterated the old
map. A delivery landing on the consumer work-pool thread after the
remove (basicCancel does not flush one already handed to that pool) hit
deliverCallback's computeIfAbsent, found the key gone, and created a
brand-new map release() never looks at again — delivered-but-unacked
forever, never requeued, never redelivered (#298 only closed the
"already in held when release runs" case).

Fix: release() swaps in a RELEASED tombstone via held.compute(...)
instead of held.remove(...). ConcurrentHashMap serializes compute/
computeIfAbsent calls for the same key against each other, so whichever
of release() and a concurrent deliverCallback runs first is fully
visible to the other — no gap. deliverCallback checks for the
tombstone and nacks-with-requeue instead of recreating a map; peek/ack
treat it as empty; own() clears a stale tombstone so a target is never
poisoned if its id is ever reused (the issue's own text says id reuse
doesn't happen, but the tombstone would otherwise sit in `held` forever
either way).

New test AmqpReplyInboxReleaseRaceTest forces the actual interleaving
with a latch (blocks release() inside its nack loop, which is only
reachable after the tombstone swap, then fires a concurrent delivery)
rather than a sequential call — a sequential test would not have caught
this, since #298's own contract test forces settlement before release()
runs. Mutation-tested: reverting the fix makes this test fail with
"expected: <2> but was: <1>" (m1 never nacked); restored after
confirming that failure.
2026-09-04 14:10:21 +07:00
Dai Ha c801851c66 Correct the isLoopback javadoc: after #305 narrowing this range refuses a caller, it does not promote one
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m47s
2026-09-04 14:10:10 +07:00
Dai Ha de70aa38f1 Merge #317: an unresolved caller is refused, never promoted to primary
CI / contract (push) Successful in 1m1s
CI / build (push) Successful in 1m56s
2026-09-04 14:03:35 +07:00
Dai Ha 53a533afb4 #317: refuse an unresolved caller instead of promoting it to primary
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
ConnectionIdentity.resolve() called pids.pidForLocalPort(remotePort),
which returns -1 both on a real failure and (silently, no log line)
when lsof just finds no matching process. terminalForPid(-1) then
matches no pane, so CallerResolver's loopback-trust fallback could not
tell that caller apart from a genuine primary and handed it
Principal.primary(...) — granting SPAWN, STOP, SEND and DRAIN to a
worker whose PID lookup failed. This is the escalation PaneLocator's
own javadoc already names; CB-161's ancestry walk only helps once a
candidate pid exists, and a failed lookup has none.

Fix: ConnectionIdentity.Caller gets a resolved() predicate (pid > 0),
centralised next to the -1 sentinel it tests for the same reason
isLoopback() is centralised (fleetd #305: two independent copies of
one rule already drifted once). CallerResolver's loopback-trust
fallback now requires c.resolved() before granting PRIMARY; an
unresolved caller gets Principal.anonymous() — the same already-tested
"authenticated as nothing" outcome used everywhere else in that
method, so the refusal is a clean, named, unsurprising result rather
than something that looks like a bug.

Also logs the previously-silent "lsof ran clean, found no match" case
in LsofPeerPidLookup at DEBUG, since that (not a slow lsof — the
waitFor result was already discarded) is the likelier real trigger.

loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary is untouched
and still green: a real pid that owns no pane (the actual primary) is
still resolved() and still PRIMARY. Token mode is unaffected — it
never consults c.pid() at all.

Mutation-tested: reverting only the CallerResolver.java guard
reproduces the escalation exactly (aFailedPeerPidLookupIsRefusedNotPromotedToPrimary
fails with "expected: <ANONYMOUS> but was: <PRIMARY>").
2026-09-04 13:58:14 +07:00
Dai Ha 77ad88631b Merge #315: the fixed placement policy honours the retry loop's unreachable set
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 1m55s
2026-09-04 13:57:24 +07:00
Dai Ha d88017807b #315: fix self-contradicting javadoc left by the previous commit
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 2m31s
FixedPlacementPolicy's class javadoc still opened with "This ignores caps
and reachability" after the previous commit added reachability as the
fourth carve-out that is explicitly NOT ignored — caught by a shape-check
survey run against this same file as part of #315's own request ("look in
placement/ ... for the same shape: a caller/comment that documents an
expectation ... where an implementation does not meet it"). Reworded the
opening sentence: fixed still ignores caps (maxLoad) by design, but
reachability is now a narrower, per-call retry exclusion, not an ignored
concern.
2026-09-04 13:55:28 +07:00
Dai Ha 2159a5a94a #315: FixedPlacementPolicy now honors the retry loop's unreachable set
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 2m33s
CompositePeerLauncher.spawn retries a failed candidate on the next one and
rebuilds PlacementContext "so the policy excludes this profile" (its own
comment), but FixedPlacementPolicy.select never read ctx.unreachable(). Under
the default `fixed` placement policy (used when `placement` is unset or set
to `fixed`), every retry re-picked the same dead default and a second,
healthy, configured profile was never tried. This also covers the wiring-bug
branch (a candidate profile with no owning adapter), which hit the exact same
symptom for the same reason.

Not live on this fleet: fleetd.yaml sets placement: weighted, which already
consults ctx.unreachable() via PlacementPolicyUtil.available(). This is live
only for a deployment that leaves placement unset or sets it to fixed.

Fix is in FixedPlacementPolicy: consult ctx.unreachable() in the same two
places it already consults quarantined/coolingOff (the default check and the
fallback walk over candidates()), and add a fourth reason to the "no
candidate remains" exception. Considered fixing this in
CompositePeerLauncher's retry loop instead (break when select() returns an
already-unreachable profile), but that only fails faster on the same dead
profile — it cannot make the loop advance to a different candidate, because
only the policy decides which candidate is next. The defect is that one
policy implementation does not honor the loop's stated contract, so the fix
belongs in that policy, matching how weighted/round-robin already behave.

Also fixed: the "no reachable worker profile" exception message said
"trying N candidate(s)" where N was unreachable.size(), a count of DISTINCT
profiles (a HashSet dedupes a profile added twice), under wording that reads
as a count of attempts. Reworded to "N distinct candidate(s)" so the count
matches what is measured and the profile list that follows it.

Tests: two new failover tests next to the three existing ones in
CompositePeerLauncherTest (which all use PlacementPolicies.weighted(), which
is why this had no coverage) — one pinned to PlacementPolicies.fixed() for
the unreachable-default case, one for the wiring-bug (no adapter) case.
Mutation-proofed: reverted FixedPlacementPolicy.java, both new tests failed
with the exact bug ("no reachable worker profile available after trying 1
distinct candidate(s): a" / "...c"), then restored the fix.
2026-09-04 13:52:29 +07:00
Dai Ha b9c2cf69f4 Merge #307: a worker's real reply after an ask timeout completes its ticket instead of stranding
CI / build (push) Successful in 2m5s
CI / contract (push) Successful in 2m19s
2026-09-04 13:31:43 +07:00
Dai Ha 8beae50fe7 Merge #308: refuse spawns once the shutdown drain has started, and sweep stragglers
CI / contract (push) Successful in 1m49s
CI / build (push) Successful in 3m0s
2026-09-04 13:28:09 +07:00
Dai Ha b2a58cb966 Merge #309: clean up partial worktree state when git worktree add fails 2026-09-04 13:28:04 +07:00
Dai Ha b8b25cf74c #307: an ask() timeout no longer strands the worker's real reply
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m30s
MessageService.reply()'s async-recovery path (askAnsweredAsyncTasks)
required a live Task.turnId, but ask()'s own TimeoutException handler
calls clearAsyncQuestion(turnId, true) — deliberately forgetting turnId
so hasAsyncQuestion() stops reporting the target BUSY. That made a
worker's eventual real fleet_reply, after an unanswered fleet_ask, fall
through to the inbox: fleet_poll{ticket} stayed PENDING forever and was
later force-failed with the false reason "session released before it
replied".

Fix: a new Task.askTimedOut marker is set (markAskTimedOut) right
before the turnId is forgotten, and askAnsweredAsyncTasks accepts it in
place of a live turnId. The marker never touches asyncTasksByTurn, so
the BUSY-release behaviour (invariant 1) is untouched. The existing
ambiguity guard (candidates.size() > 1 -> inbox, never guess) still
applies unchanged, but is now genuinely reachable rather than pure
defence in depth, since an ask timeout frees its target for a fresh,
independent delegation — the affected javadocs are updated to say so.

Tests: MessageServiceTest.aReplyAfterAnAskTimeoutStillCompletesTheAsyncTicket
(positive, mutation-proven) and
.twoAskTimedOutTicketsOnOneTargetFallBackToTheInboxRatherThanGuess
(negative/ambiguity). FleetMcpTest's
unansweredAsyncAskReturnsTheTicketToPending was renamed and its final
assertion updated — it had pinned the old (buggy) inbox-stranding
behaviour as expected.
2026-09-04 13:23:01 +07:00
Dai Ha f159ca7d27 #310: log when a reap is skipped because the record changed
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m46s
The compare-and-release declines silently. This race is unobservable by
construction, so a reaper that quietly stops reaping is the hardest kind of
behaviour to diagnose later. One debug line names the pane and the likely
cause.
2026-09-04 13:22:55 +07:00
Dai Ha 83f2aea60f #308: refuse a spawn once the shutdown drain has started, and sweep stragglers
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 1m36s
drainAll iterated a one-shot registry snapshot with nothing to refuse a new
fleet_spawn while the drain was still running (mcp.close() only runs 8 calls
after sessions.close() in the shutdown hook). A session registered in that
window was never visited by the drain loop: its pane kept running and its
worktree was never preserved, with the in-memory registry gone at exit.

Fix, both mechanisms as the issue asked for (neither alone is complete):

- SessionManager.acquire now checks a `draining` flag, flipped true at the
  very start of drainAll before the registry snapshot is even taken, and
  throws the new ShuttingDownException (invariant 3: fail loudly, say why).
  FleetMcp.spawn and FleetApp.spawnMember surface it as a clean error/503
  rather than an uncaught RuntimeException.
- The flag alone cannot close the whole race: a caller already past the
  check can still be mid-launcher.spawn() (a real herdr round trip) when
  drainAll snapshots the registry. drainAll now re-reads the registry once
  its main pass finishes and drains whatever straggler landed there too,
  bounded by the SAME whole-drain deadline (invariant 1: timeoutNanos stays
  a budget for the whole drain, never extended for a straggler).
- ReleaseCause.SHUTDOWN still preserves worktrees for both the initial pass
  and the sweep (invariant 2, unchanged release() path).

Tests (SessionManagerTest): a guard test proving acquire() throws once
drainAll has started, and a race test using a launcher double that blocks
the second spawn() and the first stop() call to force, deterministically,
the exact interleaving where a spawn passes the guard before drainAll flips
it and only registers after the initial snapshot — proving the post-loop
sweep catches it.

Shape check (SessionManager.java only, not fixed): reapIdle has the same
shape — a decision made from a roster() snapshot, then acted on via
release(s.paneId()) with no re-check of the session's current state.
2026-09-04 13:21:53 +07:00
Dai Ha 002329adb5 #309: clean partial worktrees after add failure
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 2m2s
2026-09-04 13:18:57 +07:00
Dai Ha a49671ceb9 #310: prevent idle reap from stopping delivered workers
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 2m1s
2026-09-04 13:17:51 +07:00
Dai Ha 76672ff016 #306: the post-turn phase gets the same unknown-stall escape as a turn
CI / build (push) Successful in 1m43s
CI / contract (push) Successful in 1m50s
Four latches gate delivery in Injector, and only awaitingCompletion had a way
out of a sustained unknown streak. CB-109 added that escape because a worker
stuck in a state herdr cannot classify never produces a working->idle
boundary. The same is true during post-turn housekeeping, but the escape was
never extended there.

awaitingPostTurnPickup and postTurnObserved are both released only on an
injectable sample, so a worker that goes unknown and stays there wedges: the
target is polled forever, every later message to it is blocked by the delivery
gate, and no onTurnFailed fires, so the session sits at DONE and looks healthy.
The counter did not even increment, since ++unknownSinceTurn sits inside the
awaitingCompletion short-circuit.

postTurnPending needs no escape; it is cleared on the line after the listener
call that sets it.

The escape does not set turnFailed. The delegated turn already completed and
its waiter already resolved — what is outstanding is the /clear. Failing the
turn would drive SessionManager.onFailed on a session that genuinely finished.

Not reachable in the live configuration: the path needs lifecycle.clearAfterTurn,
which fleetd.yaml does not set. It becomes reachable as soon as anyone turns
that supported knob on.

Fixes #306
2026-09-04 13:06:29 +07:00
Dai Ha 9379f92c23 #305: one definition of loopback, so a worker cannot become the primary
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m40s
ConnectionIdentity and CallerResolver each kept their own isLoopback. They
drifted: the identity resolver accepted only 127.0.0.1, the authorization
check accepted all of 127.0.0.0/8.

A caller from 127.0.0.2 therefore had its identity resolution skipped, so it
carried no terminal, and CallerResolver reads a missing terminal as "not a
worker" — which under loopback-trust, the default mode, is the primary. A
worker got spawn, stop, send and drain. The skip also happens before the PID
ancestry walk, so that defence is bypassed too.

Being strict in ConnectionIdentity was not the safe direction. That predicate
decides whether identity is resolved at all, and resolution is what demotes a
worker, so every address it excluded was one where a worker became the lead.

Measured, not assumed: on Linux the whole 127.0.0.0/8 is bound to lo, and
binding a source of 127.0.0.2 on the fleet host succeeds (curl rc=7, the
connect refused rather than the bind). On macOS the source bind fails (rc=45),
so this workstation was never exposed.

The shared predicate also accepts the IPv4-mapped IPv6 form, which neither
copy handled. That one failed in the safe direction: a primary on
::ffff:127.0.0.1 was refused as anonymous.

No transport-level test binds a real 127.0.0.2 source — it cannot run on
macOS. The reasoning is recorded on the issue.

Fixes #305
2026-09-04 12:58:21 +07:00
Dai Ha 21ff63b11d #304: the member routes report a herdr failure instead of a bare 500
CI / build (push) Successful in 1m34s
CI / contract (push) Successful in 1m53s
POST /members and DELETE /members/{paneId} were the two routes in FleetApp
with no catch (HerdrException). FleetApp has no Javalin exception mapper, so
the exception escaped as the default 500 with the body "Server Error" — no
herdr code, no herdr message. fleet_spawn and fleet_stop catch the same
exception and report a named error, so this was the same one-door-guarded
shape as #297.

Both now go through the existing herdrError mapper: 404 when herdr says the
target is gone, 502 otherwise. That is an answer the caller can act on.

stopMember matters more than spawnMember. SessionManager.release deregisters
the session, notifies the release listener and preserves a dirty worktree
before it calls launcher.stop, so a throw from that stop arrives after the
teardown the caller asked for has already happened. A bare 500 told the caller
to retry and carried nothing to explain what went wrong.

The two existing tests that asserted 500 now assert 502 and check the error
body. Neither was about the status code: one guards that a failed teardown is
not reported as a successful 204, the other that a failed spawn still closes
its tab. Both properties are unchanged.

Fixes #304
2026-09-04 12:48:21 +07:00
Dai Ha 21844b54d7 #302: fleet_reply refuses blank content too, matching fleet_send
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 3m34s
MessageService.reply now throws on blank content. FleetMcp.reply guarded only
against null, and its handler is a bare BiFunction with no try/catch, so a
whitespace-only fleet_reply left the handler as an uncaught
IllegalArgumentException instead of the clean tool error null already got.
fleet_send has always used isBlank here; reply now matches it.

The worker found that null/isBlank difference and reported it as a correction
to my ticket, which had quoted the guard wrongly. It was right: I grepped the
error string and assumed the condition matched its sibling.
2026-09-04 12:36:23 +07:00
Dai Ha 4769481515 Merge #302: a reply with no content is refused, not silently delivered
REST read content with .path("content").asText(""), so a body missing the key
became an empty string that resolved the lead's waiter. The turn completed and
the lead saw a member that finished and reported nothing, indistinguishable
from one that genuinely said nothing. The guard went into MessageService.reply,
which both doors call, rather than being written a second time in FleetApp.
2026-09-04 12:33:47 +07:00
Dai Ha bbbb4c1eb3 fleetd #302: require content in MessageService.reply so a REST reply with no content cannot silently resolve a waiter
CI / build (pull_request) Successful in 1m44s
CI / contract (pull_request) Successful in 3m17s
2026-09-04 12:31:09 +07:00
Dai Ha 38c248e617 Merge #298: a released reply is requeued, not dropped
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m46s
release() cancelled the consumer and dropped its local record of deliveries the
broker still held as outstanding. Cancelling a consumer does not requeue them:
they stay unacked on the still-open channel until it or the connection closes.
So a held-but-undrained reply became permanently unreachable — a worker's
report lost with no error and no log line. It now nacks with requeue, after
cancelling, so a later own() can still receive it.
2026-09-04 12:21:30 +07:00
Dai Ha 9debc0de27 #297: render both profiles doors from one body builder, not two copies
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 1m31s
The merged change shared the QuarantineSource and OutageSource instances
between fleet_profiles and GET /profiles, so the two doors read identical
facts. It then rendered those facts through a character-for-character copy of
the loop, in a different file. Shared inputs do not make duplicated computation
safe: a later edit to the row shape lands on one door and not the other, and
the two disagree about a live outage. That is what #284 was.

The ticket caused this. It said 'read from the same shared instances' and 'do
not change FleetMcp', and together those made copying the loop the only legal
move. Extracting FleetMcp.profilesView and calling it from both is what the
ticket should have asked for.
2026-09-04 12:19:02 +07:00
Dai Ha f34361b263 Merge #297: REST reports what MCP reports
GET /agents and GET /members now map a herdr transport failure into the same
{error, detail} envelope every other handler in FleetApp uses, instead of
letting it escape to Javalin's default handling. GET /profiles now reports the
quarantined and coolingOff states, read from the same shared sources FleetMcp
reads. REST is the door a lead falls back to when its MCP mount drops, so it
was weakest exactly when it was load-bearing.
2026-09-04 12:16:00 +07:00
Dai Ha be5ba22c75 #298: AmqpReplyInbox.release() requeues held deliveries instead of dropping them
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 2m3s
release(target) used to cancel the target's consumer and clear held's local
record for it. Cancelling a consumer does not requeue the broker's in-flight
deliveries — they stay unacked on the still-open channel until a real
connection drop. So a held-but-undrained reply became permanently
unreachable: never acked, never nacked, never requeued, invisible to peek.

Fix: cancel the consumer first (so it can no longer receive redeliveries),
then nack-with-requeue every held delivery for that target before dropping
the local record. Nacking before the cancel was tried first but a real
broker demonstrated a race: the still-active consumer immediately received
the requeued message back, racing held.remove and leaving peek non-empty.
Cancelling first avoids that. A failed requeue is logged at WARN and does
not abort release(), matching the best-effort teardown style #293 settled
for HerdrPeerLauncher.stop().

Extends AmqpReplyInboxContractTest.releaseCancelsConsumer... to assert the
held delivery is recoverable via a later own(), not just absent from peek.
2026-09-04 12:14:00 +07:00
Dai Ha 85c90d440a #297: map HerdrException on GET /agents and /members; GET /profiles reports quarantine + cool-off
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m33s
Two REST-only visibility gaps, both against the same shared instances FleetMcp
reads (BackendQuarantine/BackendOutagePolicy), never recomputed:

- GET /agents and GET /members let a HerdrException escape uncaught, outside
  the {error, detail} envelope every other failure path in FleetApp uses.
  Both now route through the existing herdrError() helper, matching healthz/
  sessionStatus. GET /members is the endpoint's own comment names as the
  out-of-band path a lead falls back to when its MCP mount drops.
- GET /profiles omitted the two outage states fleet_profiles already reports:
  quarantined (CB-578 stage B) and coolingOff (fleetd #201 Unit 5). FleetApp
  now takes the SAME FleetMcp.QuarantineSource/OutageSource instances Fleetd
  wires into FleetMcp (extracted to local vars in Fleetd.java so both doors
  share one object, not two independently-built copies of the same rule).

FleetMcp itself is unchanged. Item 3 of the ticket (a capacity block on
GET /members) is explicitly out of scope and was not added.
2026-09-04 12:12:35 +07:00
Dai Ha e97502d550 drainAll: record that the timeout is a whole-drain budget, not a per-session grace
CI / build (push) Successful in 1m27s
CI / contract (push) Successful in 1m48s
The javadoc said 'for each session that is BUSY, poll up to timeoutNanos',
which reads as a per-session grace period. The deadline is taken once, before
the loop, so the first BUSY session can spend all of it. That is deliberate and
is the safer of the two designs: the drain is one phase of a shutdown sequence
that must finish inside launchd's exit window, and a per-session grace would
overrun it and get the daemon SIGKILLed part-way through, leaving the sessions
not yet reached with no clean release, no preserved-worktree log and no
snapshot. Found by a read-only hunt that read the code correctly and drew the
opposite conclusion from the wording.
2026-09-04 12:09:44 +07:00
Dai Ha aef14ff46e Merge #296: a failed spawn no longer leaks the pane it opened
Two exits created a pane and left it running. The readiness gate propagated an
unrelated herdr error without teardown, and spawnAsPane never closed the pane it
split when the peer failed to start. Neither could be cleaned up by the caller:
SessionManager.acquire never learns the pane id, because spawn throws before it
returns one. It removed the worktree anyway, so the leak was a live backend with
a deleted cwd, invisible to fleet_list and holding a seat nothing decremented.
2026-09-04 12:08:23 +07:00
Dai Ha cba516bda4 fleetd #296: close panes on failed spawn
CI / contract (pull_request) Successful in 1m38s
CI / build (pull_request) Successful in 2m7s
2026-09-04 12:05:26 +07:00
Dai Ha ba51e0c6cc #293: catch RuntimeException on closeTab, matching its sibling guard
CI / contract (push) Successful in 50s
CI / build (push) Successful in 2m15s
No behaviour change today. HerdrCodec wraps every encode/decode failure
and UnixSocketHerdrClient wraps every IOException, so HerdrException is
all closeTab can currently throw.

But releaseZdotdir five lines below catches RuntimeException, and the
whole point of this fix is that nothing here may mask the cleanups
below. Guarding against the expected exception type and staying bare
against any other is the same asymmetry the ticket exists to remove,
one level down. This stops a later change inside
WorkspaceControl.closeTab reopening it.
2026-09-04 11:47:42 +07:00
Dai Ha 086c59848e Merge #293: a failing tab.close no longer masks the cleanups below it 2026-09-04 11:46:00 +07:00
Dai Ha 0c10079755 #293: wrap the bare tab.close in HerdrPeerLauncher.stop()
CI / contract (pull_request) Successful in 1m41s
CI / build (pull_request) Successful in 1m59s
The pane is already closed by the time spaces.closeTab runs, so a failing
tab.close is cosmetic workspace tidying, not a real teardown failure. Left
bare, it propagated out of stop() and masked releaseZdotdir (ZDOTDIR leak)
and, worse, SessionManager.release()'s worktree removal (no self-heal,
no retry — the registry entry is already gone by then).

Wrap it in a try/catch that logs a WARN naming the tab id, matching the
"must not mask a real teardown failure above" comment already on
releaseZdotdir. isAlreadyGone is untouched — this continues past *any*
tab.close failure, not just *_not_found, since the failure is cosmetic
regardless of its cause.

Adds FakeHerdr#tabCloseFailsWith/tabCloseFailsForTab (the tab.close
counterpart to #290's paneCloseFailsForPane) plus two tests: one proving
releaseZdotdir still runs (the generated ZDOTDIR is deleted) and one
proving SessionManager.release() still removes the worktree, both with a
non-not_found tab.close failure.
2026-09-04 11:41:07 +07:00
Dai Ha ece2091b53 #280: states is now touched by two scheduler tasks, so make it concurrent
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m56s
The delayed re-check reads `states` from its own scheduled task, while
`tick` writes and prunes it. Both run on the single-threaded scheduler
Fleetd passes in today, so they are serialised — but nothing in the
class enforces that, and an unsynchronised HashMap read racing a resize
can spin a CPU forever rather than fail visibly.

`priors` and `orphanStreaks` stay plain maps: `tick` is still their only
toucher. The comment says which is which, so the next person does not
have to re-derive it.
2026-09-04 11:32:56 +07:00
Dai Ha b3f917e6f5 Merge #280: one bounded delayed re-check sweeps a ticket whose ask lapsed after GONE 2026-09-04 11:31:00 +07:00
Dai Ha fa39a5f55e #280: sweep a GONE/NEVER_READY target's lapsed fleet_ask once, not just on transition
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m58s
FleetHealthMonitor's fire-once-per-transition rule (CB-580) means a target
that is genuinely mid-fleet_ask when health first classifies it GONE is
correctly skipped (sweepAsking=false). But nothing re-fires abandon() once
that ask lapses on its own 55-115s later: FleetHealth.decide keeps reporting
GONE every tick, and reportTransition's previous==next guard never lets the
sweep run again. The ticket then sat PENDING forever, the same destination
#275 fixed for an explicit teardown, reached here by a health guess instead.

SessionManager.reapIdle only reaps READY/DONE sessions (SessionManager.java:844),
and a session mid-turn (including mid-ask) stays BUSY the whole time
(onDelivered sets BUSY, nothing clears it until the turn completes) — so
SessionReaper never releases such a session and onRelease's sweepAsking=true
path is never reached.

Fix: schedule one bounded, delayed re-check per terminal transition (not a
per-tick retry — that shape was rejected by CB-580). It fires failTerminalTarget
again after a delay that exceeds the worst-case ask-lapse window, and only if
the target is still classified in the same terminal state at that time, so a
recovered or since-released target is never reached into. sweepAsking stays
false throughout, so a ticket whose ask has not yet lapsed is still never
touched — same invariant abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer pins.

Proven with a mutation: neutering recheckTerminalTarget's body made
delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved fail with
"expected: <FAILED> but was: <PENDING>", all 20 other FleetHealthMonitorTest
cases still green; restored and reran clean (1300 tests, 0 failures).
2026-09-04 11:28:12 +07:00
Dai Ha d5128a1d35 #290: use the imports FakeHerdr already has
CI / build (push) Successful in 1m33s
CI / contract (push) Successful in 1m39s
2026-09-04 11:27:51 +07:00
Dai Ha 61097e5cf0 Merge #290: restore coverage for reapIdle's per-session guard 2026-09-04 11:25:45 +07:00
Dai Ha ef507bcd12 #290: restore reapIdle's per-session guard coverage via a new launcher.stop() trigger
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m56s
#283 fixed release() to catch and log a worktree-removal failure, which closed off
reapIdleCountsAllThreeSessionsWhenOnlyItsWorktreeRemovalFails as a trigger for
reapIdle's own per-session try/catch (CB-581) — that test now proves a different,
still-real thing (a swallowed removal failure doesn't shrink the reaped count),
but the try/catch itself lost its test.

Add FakeHerdr.paneCloseFailsForPane(paneId, code) so a test can make exactly one
session's launcher.stop() fail while its siblings still tear down normally
(paneCloseFailsWith already existed but fails every pane, which cannot isolate
one session in a three-session reap). Add
reapIdleSurvivesOneSessionWhoseLauncherStopFails beside the #283 test, using
launcher.stop() as the trigger the ticket names, and prove it catches removal of
reapIdle's try/catch: deleting the guard makes the test fail with the
HerdrException propagating out of reapIdle uncaught (quoted in the PR body).
2026-09-04 11:23:52 +07:00
Dai Ha 94f50e507a #284/#285: one reclaimable rule, one group-share helper
CI / contract (push) Successful in 1m2s
CI / build (push) Successful in 2m11s
Two corrections on top of the merged worker branches.

#284: I told the worker to report a BACKEND_ERROR/FAILED session as
reclaimable. That half of my own ticket was wrong. Once the live count
stops counting a terminal session, its seat is already in `free`;
counting it in `reclaimable` too reports the same seat twice, and
`free + reclaimable` reads as more capacity than maxLoad allows. Worse,
only the profile-level count was widened, so the same fleet_list
response said `reclaimable: 2` while every member row said
`reclaimable: false`.

Both views now call one shared predicate, FleetMcp.reclaimable, so they
cannot drift apart. A test runs it over every MemberSession.State value,
so a state added later cannot slip through unconsidered.

#285: the new per-file chgrp+chmod helper moved from ClaudeCodeLauncher
into EnvAllowListScrub as shareFileWithGroup, next to the directory-wide
shareWithGroup it was copied from. It now reuses that class's own
setGroupAndPermissions and also catches UnsupportedOperationException,
which the copy missed — on a filesystem without POSIX group ownership
the copy threw a raw runtime exception instead of the sibling's
UncheckedIOException.
2026-09-04 11:13:01 +07:00
Dai Ha 75b15086b0 Merge #285: the workspace-trust seed refuses rather than write fleetd's own home 2026-09-04 11:08:44 +07:00
Dai Ha dab9645906 Merge #284: a BACKEND_ERROR or FAILED member no longer holds a spawn seat 2026-09-04 11:04:09 +07:00
Dai Ha e93b5f6512 #282: record that answer()'s QUESTION guard is defence in depth, not load-bearing
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m44s
Measured after merging: removing the guard alone leaves the new test green,
because ask() calls markAsyncQuestion before resolveQuestion, so the task has
already moved to the new turnId. The PR claimed each half was necessary; only
the pair is. Keeping the guard, with the ordering written down so nobody
deletes it as dead code or trusts it as the only protection.
2026-09-04 10:59:10 +07:00
Dai Ha bfabe13e8f Merge #282 (PR #289): a chained second fleet_ask no longer kills its own async ticket 2026-09-04 10:56:42 +07:00
Dai Ha 7e49c6eca2 #285: seedTrustDialog refuses under memberHerdrSocket instead of seeding fleetd's own home
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 3m35s
seedTrustDialog gated only on isProvisionedWorktree(cwd) and, being static, could not
see memberHerdrSocketConfigured() — unlike its sibling writeCharterFile, which already
refuses the spawn when it cannot place a file where a different-uid member can read it.
Under memberHerdrSocket + configDir unset, seedTrustDialog wrote fleetd's OWN
~/.claude.json while believing it was seeding the member's, reintroducing the fleetd
#149 failure (interactive trust dialog, no fleet_reply, silent readiness timeout) for
this one config combination.

Makes seedTrustDialog an instance method so it can see memberHerdrSocketConfigured()
and memberGroup(), and applies writeCharterFile's "refuse, don't degrade" rule: under
memberHerdrSocket it now requires both configDir and worktreeGroup before touching any
file, naming exactly which is missing, and shares the written .claude.json group-
readable (rw-r-----) via a new shareTrustJsonWithGroup so the member's OS user can
actually open it. The memberHerdrSocket-absent path (today's only live mode) is
unchanged.
2026-09-04 10:55:35 +07:00
Dai Ha 11cbfa79b4 fleetd #284: free failed-session capacity
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m19s
2026-09-04 10:55:21 +07:00
Dai Ha c2c2746922 Merge #283 (PR #288): guard release()'s worktree removal, delete the orphaned branch on spawn failure
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m44s
2026-09-04 10:54:36 +07:00
Dai Ha 6ed70700a0 #282: don't let a chained fleet_ask kill its own async ticket
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Successful in 1m38s
answer() opened a fresh forward waiter but, unlike send(), never
registered it in asyncTasksByWaiter. So when a worker chained a
second fleet_ask inside the same resumed turn (before calling
fleet_reply), markAsyncQuestion had no Task to re-associate, and
answer() then completed the async ticket's future with the second
QUESTION as if it were a terminal reply — fleet_poll reported FAILED
while the worker was still alive and mid-conversation.

Fix: register answer()'s waiter in asyncTasksByWaiter (mirroring
send()) so a chained ask can re-arm the ticket under its new turnId,
and guard answer()'s finishAsyncTask call the same way sendAsync's
own lambda already does (skip on Outcome.QUESTION). Also drop the
stale asyncTasksByTurn entry left behind when markAsyncQuestion
re-arms a task under a new turnId, a leak the fix makes reachable
for the first time.

Reachability confirmed by driving the exact sequence through the
public API (sendAsync -> ask -> answer -> ask again) in a new test;
reverting the production change makes it fail with
"expected: <ASKING> but was: <FAILED>", confirming it catches the
regression.
2026-09-04 10:52:54 +07:00
Dai Ha 0087645da4 Merge #281 (PR #286): pin the Authz.Action every handler chooses
CI / contract (push) Successful in 1m1s
CI / build (push) Failing after 1m45s
2026-09-04 10:50:43 +07:00
Dai Ha 19cacf5b62 #283: guard release()'s bare worktree remove; delete orphaned branch on spawn failure
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 2m42s
Two teardown-cleanup leaks in SessionManager, same shape as #274.

Defect 1: release()'s last step (removing a released session's worktree)
was the one cleanup step in the method left unguarded, even though every
sibling step is wrapped because exec() can throw on a non-zero exit or its
own 30s timeout. By the time it ran, the registry entry, retained handle,
and pane were already gone, so a throw here escaped release() with no
retry path and made a fully-torn-down session look like a failed stop.
Now wrapped in try/catch with a WARN, matching the pattern already used
by every other step in this method.

Defect 2: acquireWithWorktree's catch (covering failures after add()
returns — overlayParity, shareWithGroup, launcher.spawn) removed the
worktree but left the branch it provisioned orphaned. #274 already fixed
the sibling failure inside add() itself (GitWorktrees.cleanupAfterAddFailure
deletes both). Extracted that branch-delete into a new Worktrees.deleteBranch
method, reused by both cleanupAfterAddFailure and this catch, so a routine
spawn failure (quarantined credential, backend refusal) no longer leaks a
worker/<slug>-<nonce> branch.

A normal release() still never deletes a branch — only the failed-provision
path does. releaseRemovesWorktreeButDoesNotDeleteBranch pins this, and
spawnFailureAfterAddDeletesTheOrphanedBranch / unchangedRegression* prove
the two paths stay apart.
2026-09-04 10:50:30 +07:00
Dai Ha 4dd12083ab fleetd #284: free backend-error capacity
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m42s
2026-09-04 10:47:09 +07:00
Dai Ha 30e21adec7 #281: cover registered authorization actions
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 1m33s
2026-09-04 10:46:26 +07:00
Dai Ha 719b79f892 wip: pin handler authorization actions
CI / build (pull_request) Successful in 1m30s
CI / contract (pull_request) Successful in 2m27s
2026-09-04 10:39:58 +07:00
Dai Ha 66e5247b6d Merge #275 (PR #279): sweep an ASKING ticket on a definite teardown
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m39s
A member torn down while parked in fleet_ask left its async ticket pending
for good. resolveQuestion had already closed the forward waiter, so
abandon()'s waiter branch found nothing; the 'question == null' guard then
excluded the task from the matching loop. By the time the worker's own ask
lapsed (~55-115s), the released session was gone from the roster, so
nothing was left to call abandon() on that target again. fleet_poll{ticket}
reported PENDING forever.

The ticket told the worker to drop the 'question == null' guard. That was
wrong, and the worker said so with evidence: an existing test
(abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer) deliberately pins that
an ASKING ticket must SURVIVE abandon(), because the primary may be mid
answer() for that same turn. Widening the shared method would have traded
this bug for a worse one — a health guess killing a live conversation.

So the fix splits the two callers by what they actually know:

  - sessions.onRelease (fleet_stop / idle reaper) knows the pane is being
    stopped right now, so it sweeps: sweepAsking=true.
  - FleetHealthMonitor keeps sweepAsking=false. GONE/NEVER_READY is a
    classification from the live agent list, not a teardown it performed.

I verified the reachability chain myself rather than taking it on trust.
FleetHealth.decide returns GONE before it can ever return
DELEGATION_ORPHANED; FleetHealthMonitor.reportTransition returns early when
previous == next; and terminal() is GONE/NEVER_READY only. So after the one
GONE transition fires and no-ops, nothing re-fires. Every link holds.

Verified: the real merge into current main builds green (1283 tests), the
protective test still passes untouched, and my own mutation — reverting the
sweepAsking widening — fails the new test with 'a released target's open ask
can never resume, so it must fail right here'.
2026-09-04 10:18:59 +07:00
Dai Ha 1e41bd63b4 Merge #274 (PR #277): clean up the worktree when add() fails after creating it
CI / contract (push) Successful in 59s
CI / build (push) Successful in 1m34s
GitWorktrees.add() created the worktree and branch, then ran more steps that
can throw — requireCredentialFreeHttpsOrigin among them, which is an
intended security refusal, not an IO accident. Any throw meant add() never
returned, so SessionManager.acquireWithWorktree never learned the path, its
'if (path != null)' cleanup could not fire, and the worktree and branch
leaked with nothing tracking them. Every OTHER exit from that method was
cleaned up correctly; only the exits inside add() were uncounted.

add() now cleans up what it created before rethrowing, reusing remove() and
additionally deleting the branch — a branch that never finished provisioning
has no session and no PR behind it. Worktree first, since a checked-out
branch cannot be deleted. Cleanup failure is logged and never masks the
original exception.

Verified by me: the real merge into current main builds green (1281 tests),
and I reran the mutation myself without git stash — dropping the cleanup
call fails the new test with 'the worktree directory leaked after a
post-creation step threw'.

The test drives add() itself through the existing afterWorktreeAdded seam,
so the failure happens after the worktree exists rather than downstream in
another caller.
2026-09-04 10:14:05 +07:00
Dai Ha 4887d03d88 #275: abandon() sweeps an ASKING ticket only on a definite teardown
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 2m5s
Confirmed reachable: a target torn down for good (fleet_stop / the idle
reaper) while its async ticket sits in fleet_ask (Phase.ASKING) got
permanently stuck. resolveQuestion already closes the forward waiter, the
question == null guard excluded the task from abandon()'s sweep, and by the
time the worker's own fleet_ask lapses (~55-115s) the released session no
longer appears in FleetHealthMonitor's roster, so nothing ever calls
abandon() again. fleet_poll{ticket} then reports PENDING forever.

Add abandon(target, reason, sweepAsking) — sessions.onRelease (a definite
teardown: the pane is being stopped right now) passes true and now fails the
ASKING ticket and closes its reverse-rendezvous ask. FleetHealthMonitor's
health-classification call keeps the 2-arg overload (sweepAsking=false):
a GONE/NEVER_READY reading is a guess from the live agent list, not a
teardown it performed, and abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer
already covers why an active ask must survive that guess (the primary may
be mid-answer for the same turn). hasOrphanedDelegation is left unchanged
for the same reason — it must not flag a live, active ask as orphaned.

Proven with a test driving the real public sequence (sendAsync -> ask ->
abandon(..., true)), not a hand-built task map; reverted the widening to
confirm it goes red, then restored it.
2026-09-04 10:13:56 +07:00
Dai Ha 7d5434455d implementer: never git stash — the stash stack is shared across worktrees
CI / build (push) Successful in 1m41s
CI / contract (push) Successful in 1m48s
A worker's worktree is isolated; refs/stash is not. It is one stack shared
by the primary's checkout and every worker worktree of this repo.

This bit a real worker today. Two ran in parallel; one called git stash
while the other was mid-edit, and the second worker's in-progress change
was silently overwritten by the first's stashed content. It recovered by
retyping the edit and diffing to confirm, and pushed the other worker's
change back onto the stack untouched — but nothing warned either of them,
and nothing would have.

Measured before writing this: 'git stash list' from a worker worktree and
from the primary's checkout return byte-identical output, and refs/stash
is a single common ref, not a per-worktree one.

The branch already IS the isolation, so the skill now points at committing
a wip commit or writing a patch file instead.
2026-09-04 10:10:08 +07:00
Dai Ha f71ee4926e Merge #273 (PR #278): validate exhaustedPattern at config load
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m44s
A malformed exhaustedPattern passed FleetConfig.load and then crashed the
daemon at startup, in Fleetd.main's unguarded Pattern.compile, with a
message naming neither the profile nor the key. Its sibling errorPattern
had a load-time validator whose own javadoc explains exactly why that is
bad. The validator was correct; its coverage was not.

rejectMalformedErrorPattern becomes rejectMalformedProfilePatterns and now
compiles both keys, reporting failures from either in one message.

Verified by me, not taken on the worker's word: the actual merge of this
branch into main builds green (1280 tests), and I reran the mutation myself
— narrowing the loop back to errorPattern turns exactly the two new tests
red, one of them with 'Expected IllegalStateException to be thrown, but
nothing was thrown', which is the defect stated out loud.
2026-09-04 10:09:15 +07:00
Dai Ha 282a2fc2b8 fleetd #274: clean up the worktree and branch when add() fails after creating them
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m22s
GitWorktrees.add() created the worktree and branch, then ran several more
steps that can throw (requireCredentialFreeHttpsOrigin — an intended
security refusal, not only an IO accident — plus the credential-helper and
tool-surface isolation steps). Any exception there meant add() never
returned, so its caller (SessionManager#acquireWithWorktree) never learned
the path: its local `path` stayed null, the `if (path != null)` cleanup
guard never ran, and the worktree directory and branch leaked on disk
forever with nothing tracking them.

Wrap those steps in try/catch; on failure, clean up via the same
`git worktree remove --force` path remove() already uses, additionally
force-delete the new branch (remove() alone deliberately leaves a
released session's branch behind, but a branch that never finished
provisioning has nothing else pointing at it), log the cleanup outcome,
and rethrow the original exception so it is never masked.

Test drives add() itself via the existing afterWorktreeAdded seam with a
mutation that trips requireCredentialFreeHttpsOrigin after the worktree
exists, then asserts both the worktree directory and the branch are gone.
Reverting the fix (git stash on GitWorktrees.java, test unchanged) turns
it red: "the worktree directory leaked after a post-creation step threw
==> expected: <false> but was: <true>". Restored afterward.

mvn clean install: BUILD SUCCESS, Tests run: 1275, Failures: 0, Errors: 0
2026-09-04 10:05:59 +07:00
Dai Ha f04e934b94 fleetd #273: validate exhaustedPattern regex at load, like errorPattern
CI / build (pull_request) Successful in 1m36s
CI / contract (pull_request) Successful in 2m11s
FleetConfig.rejectMalformedErrorPattern only compiled errorPattern eagerly
at config load. exhaustedPattern was compiled unguarded in Fleetd.main,
so profiles.<name>.exhaustedPattern: "[" passed load() and then crashed
the whole daemon at boot with a raw PatternSyntaxException naming neither
the profile nor the key.

Rename the validator to rejectMalformedProfilePatterns and extend it to
also compile every non-blank exhaustedPattern, reporting
profiles.<name>.exhaustedPattern ("<value>"): <message> in the same style
as errorPattern. Both keys are collected and reported together from a
single load. Fleetd.java's compile site is left as-is per scope — it is
now safe because load already rejects a bad value.

Added tests covering: a bad exhaustedPattern is refused; a bad pattern in
each key is reported together in one message; valid patterns still load;
a blank/absent exhaustedPattern is ignored.
2026-09-04 10:05:36 +07:00
Dai Ha 3fae35c357 #276: say whose environment the "allowed N of M" line counted
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m42s
A gap in my own #269 fix. That ticket stopped four sites claiming things
about the member's environment that fleetd cannot see when memberHerdrSocket
is configured, and gave the WARN in logCredentialGap a guard. The INFO line
called two lines earlier never got one:

    logAllowListCoverage(allowed);      // no guard
    logCredentialGap(creds, allowed);   // guarded since #269

Read plainly, "member credentials: allowed 7 of 39" is a statement about the
member's credentials. Under memberHerdrSocket the pane is routed to a second
herdr whose environment fleetd has no channel to inspect, so those counts
come from fleetd's own process instead. Same overclaim #269 existed to
remove, in the line next door.

The method's javadoc does carry the caveat, by cross-reference to another
field's javadoc. That does not help the operator reading fleetd.out.

The counts stay useful, so this is not a WARN and not a refusal — only the
claim is narrowed. The unguarded path keeps its exact original wording, so
the existing assertion on "member credentials: allowed 1 of 3" still holds.

The new test pins the pair together so a later edit cannot fix one line and
leave the other. Mutation-proved: with the guard removed it fails printing
the old line verbatim.

Also worth recording: no test covered #269's own guard — that WARN wording
shipped unverified, and still has no coverage.
2026-09-04 10:01:23 +07:00
Dai Ha 18aecbfe67 #272: fleet_poll{target} is a drain, so gate it as one
CI / build (push) Successful in 1m46s
CI / contract (push) Successful in 2m29s
fleet_poll is two operations behind one tool name. With `ticket` it observes
an async delegation and changes nothing. With `target` it calls
MessageService.drainReplies, which REMOVES the replies — a second call
returns nothing.

The handler gated both branches with a constant Authz.Action.READ, and did
not pass the target at all. READ is open to every authenticated role, so any
worker could read a peer's sessionId out of fleet_list and destroy the
replies that peer had queued for the primary. The gate failed open, and a
drained reply is not recoverable.

Three things already said the tight gate was intended:

  - fleet_ack, four lines below, gates the same drain as DRAIN, with a
    comment giving the exact reasoning missed here ("Acking removes a reply
    from the inbox, so it is a drain, not a read").
  - the REST path checks DRAIN in FleetApp.drainReplies.
  - wiki/2-Message-Server.md lists fleet_poll as lead-only, and the tool
    schema says "drain that worker's inbox".

Nothing that works today breaks: the documented flow is fleet_poll{target}
then fleet_ack{target,msgId}, and fleet_ack is already primary-only. A
worker could never complete that flow — only destroy its first half.

The required action is a function of the arguments, but the handler chose it
before looking at them. pollAction(target) makes that choice explicit. The
ticket branch stays READ on purpose: an architect may fleet_send, so it owns
tickets and must be able to poll them.

Why the suite missed it: FleetMcpAuthzTest checks every Action against every
Role, including "a worker may not DRAIN", and passed the whole time. The
policy table was right; the action fed to it was wrong, and nothing tested
that mapping. The new tests assert against pollAction itself, so the handler
keeps no private copy of the rule.

Mutation-proved: reverting pollAction to a constant READ turns exactly the
two new defect tests red and leaves the ticket-branch test green.

Introduced in 9daf1ec, where Authz.READ's own javadoc ("...task polling")
describes only the ticket half.
2026-09-04 09:54:52 +07:00
Dai Ha 5d75f72473 Merge #271: warn when the model check cannot run for a no-worktree opencode spawn (#267)
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m40s
2026-09-04 08:38:05 +07:00
Dai Ha e028a0ae54 fleetd #267: warn once per profile when the model check can't run
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Successful in 1m50s
OpenCodeLauncher.SessionAwareHandle.agentSessionId() is the only caller of
checkModelMatch (fleetd #175), and it sits behind the fleetd #249 worktree
gate. A spawn with no worktree:true — the ordinary shape of most opencode
spawns — never reached the check at all, and the gap was totally silent.

The check cannot be decoupled from agentSessionId()'s resolved id: doing so
would re-derive 'whatever is newest in the shared directory' and reintroduce
the false-positive risk fleetd #234 fixed (a sibling's differently-configured
model looking like a mismatch for a profile that never actually ran it). The
#249 gate is correct and stays as-is.

Instead, log once per profile at WARN, naming the profile, the same
treatment discoveryUnavailable already gets a few lines above — a logged
UNKNOWN beats a check that silently never runs.
2026-09-04 08:36:12 +07:00
Dai Ha 2fa673d4c0 Merge #270: ArchUnit package-cycle test with explicit accepted exceptions (#131)
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m57s
2026-09-04 08:22:35 +07:00
Dai Ha 9020d01b40 Merge #268: rename sshAuthSock values to omit/inherit with a read-both shim (#266) 2026-09-03 20:27:55 +07:00
Dai Ha 1006805027 fleetd #131: enforce package boundaries with an ArchUnit cycle test
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m1s
Adds PackageCyclesTest, which fails the build on any new cycle between
the top-level dev.ltms.fleet.* packages. Today's five real cycles are
recorded as narrow, explicit exceptions (ignoreDependency per named
pair, both directions), each commented with the ticket step (or a note
that it needs its own) that removes it. No package moves in this PR.

archunit-junit5 1.5.0 (current stable, newer than an earlier 1.4.1
draft). Main code only (DO_NOT_INCLUDE_TESTS) and importPackages(...)
instead of a working-directory-relative target/classes path.
2026-09-03 20:22:46 +07:00
Dai Ha 27aefbf9a0 Merge #269: stop claiming memberHerdrSocket proves a different OS user (#184 item 5)
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m41s
2026-09-03 20:14:44 +07:00
Dai Ha d42c2bc204 fleetd #266: rename SSH agent environment setting
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 2m1s
2026-09-03 20:12:46 +07:00
Dai Ha 3916adc372 fleetd #184: stop claiming memberHerdrSocket proves a different OS user
CI / contract (pull_request) Successful in 1m18s
CI / build (pull_request) Successful in 1m19s
HerdrPeerLauncher asserted, as established fact, that member panes run under
a different OS user whenever memberHerdrSocket is configured. fleetd has no
channel to see the uid at the other end of a herdr unix socket — an operator
may point memberHerdrSocket at a second herdr under the SAME user for pane
isolation, in which case members do inherit fleetd's environment and the
count this WARN told them to disregard is the real gap.

Reworded the class javadoc on hostEnvNames, the WARN in
warnUnknownMemberEnvironment, the javadoc on memberHerdrSocketConfigured(),
and warnCannotShareScrubDirectory's "unreadable by another uid" claim to say
what is actually true: fleetd cannot confirm what OS user the second herdr
runs as, so the member credential gap is UNKNOWN, not known-clean or
known-dirty. No behaviour change — the fallback paths and the honest
UNKNOWN conclusion stay the same, only the stated reason changes.

Matches the framing already used by Fleetd.reportMemberTrustModel on main.

Added unknownEnvironmentWarnStatesUncertaintyNotAnAssertedDifferentUser to
HerdrPeerLauncherAllowListWiringTest asserting the new WARN wording and that
it no longer claims a different OS user as fact.
2026-09-03 20:10:22 +07:00
Dai Ha fa97f598dd Merge #265: state the member trust model at startup (#184)
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m26s
2026-09-03 16:48:44 +07:00
Dai Ha ea9aa4fd77 fleetd #184: report member trust model
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m59s
2026-09-03 16:48:13 +07:00
Dai Ha a2b8caf6b5 #184: keep the reason SSH_AUTH_SOCK matters, and the measurement
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m23s
The correction removed a false claim (blocking the socket breaks git over
SSH) but took a true one with it: the socket is a live handle to the agent,
so a member holding it can sign with every key the agent holds. Without that,
the entry reads as if the setting does not matter, and an operator has no
reason left not to set it to allow. Fixing an overclaim must not leave an
underclaim.

Also record the measurement and the mistake behind the old claim, so the next
person does not re-argue it from scratch.
2026-09-03 16:42:11 +07:00
Dai Ha b5ddbe5757 Merge #264: sshAuthSock is not a control, and the config now says so (#184) 2026-09-03 16:41:52 +07:00
Dai Ha d223a93039 fleetd #184: correct sshAuthSock guidance
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Failing after 1m37s
2026-09-03 16:38:56 +07:00
Dai Ha 96d8191149 Merge #263: free reports what the spawn gate grants (#257)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m13s
2026-09-03 16:33:20 +07:00
Dai Ha f0e7ac73d6 #103: give the operator the string that is actually in the log
The merged fix told the operator to look in journalctl for
"Fleetd.reportRequiredSecrets". That string never appears there — it is a
method name. The logger is d.l.f.Fleetd and the lines read
"startup secret NAME: set|MISSING", so the advice sent the operator looking
for text that does not exist.

Give the grep instead, and say what the report does not cover: it lists only
names a configured profile references, so a secret nothing references is never
reported.
2026-09-03 16:32:23 +07:00
Dai Ha e2fe861b4d Merge #262: the systemd unit names all three secrets it needs (#103) 2026-09-03 16:32:07 +07:00
Dai Ha 279d6f5fbd fleetd #257: free must stop subtracting leadSeatCount
CI / build (pull_request) Successful in 1m43s
CI / contract (pull_request) Successful in 1m51s
fleet_list's free row subtracted leadSeatCount, but the real spawn gate
(CompositePeerLauncher#enforceMaxLoad) only ever compares live against
maxLoad and never reads leadSeatCount. So free could report 0 while a
fleet_spawn on that exact profile still succeeded, and a lead trusting
free:0 gave up on capacity the gate would still grant.

free now always equals max(0, maxLoad - live); leadSeats stays in the
row as an informational fact, never subtracted. Documented in the
fleet_list tool description and fleetd.example.yaml.
2026-09-03 16:28:32 +07:00
Dai Ha 34480cebef Merge #261: compare-and-swap the workspace-trust seed against an external writer (#247)
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m40s
2026-09-03 16:26:37 +07:00
Dai Ha 9d0bf14c46 fleetd #103: document systemd worker secrets
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m31s
2026-09-03 16:23:29 +07:00
Dai Ha 5a3ab5764c fleetd #247: CAS seedTrustDialog against the operator's own live Claude Code
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m49s
TRUST_JSON_LOCK only serialises seedTrustDialog calls this launcher makes
inside its own JVM. It cannot reach the one writer that actually shares the
target file on a real host: the operator's own live Claude Code, whose
CLAUDE_CONFIG_DIR is routinely the very configDir a profile is given, so the
file fleetd writes on every claude-code spawn is that session's own config.
A plain read-modify-write there is a routine lost update: fleetd reads v1,
the operator's session writes v2, fleetd's ATOMIC_MOVE lands v3 built from
v1 and v2 is gone, atomically.

Add a bounded compare-and-swap: before the move, re-read the target's exact
bytes and compare with what the update was built from; on a mismatch,
rebuild from the fresh bytes and retry (up to 5 attempts). Exhausting the
retries writes nothing and logs a WARN naming the file — the member shows
the trust dialog and fails to reach an injectable state instead, which is
visible and recoverable, unlike silently overwriting the operator's live
config. Also warn every time the seed is about to target the default
~/.claude.json (configDir unset), since that is the unsafe default.

This narrows the lost-update window, it does not close it: a write landing
between the final re-read and the ATOMIC_MOVE itself is still lost.
2026-09-03 16:20:20 +07:00
Dai Ha 3bad9f5785 Merge #260: restore the already-gone-worktree teardown test (#116)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 2m9s
2026-09-03 16:13:15 +07:00
Dai Ha 8308c0b68f fleetd #116: recover the already-gone-worktree teardown regression test
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m23s
Ports the intent of the lost CB-576 commit c393600 (worker/cb576-01a04b-17,
never merged, package dev.ltms.bridged.*) onto main's dev.ltms.fleet.*
tree. Adds FakeWorktrees.markGone (tracks which add()'d worktree paths
still "exist", mirroring GitWorktrees.hasUncommitted's Files.exists guard
for the already-gone case) and a regression test,
releaseStillStopsPaneAndNotifiesWhenWorktreeIsGone, asserting that
SessionManager.release still fires notifyReleased, stops the pane, and
falls through to remove() when the worktree is already gone.
2026-09-03 16:11:27 +07:00
Dai Ha 3b3063eb2b #252: the route scrape counted commented-out registrations as live
CI / build (push) Successful in 1m44s
CI / contract (push) Successful in 2m21s
Follow-up to #259, found by verifying the guard rather than trusting it.

The scrape read FleetApp's raw source, so a registration disabled with `//`
still matched. Measured: commenting out `app.get("/tasks/{ticket}", ...)` left
the test GREEN, while deleting the same line was caught. Only the commented-out
shape was blind, and it is the silent direction — the inventory would keep
claiming a route the server no longer serves.

Drop whole-line comments before scraping. Only lines starting with //, * or /*
are dropped, deliberately not every // on a line: that would also cut a string
literal containing // (a URL) and could silently delete a real registration
sharing the line. The remaining gap is a trailing comment beside real code; no
registration here has that shape, and the vacuity test catches a scrape that
loses registrations wholesale.

Proof: with the fix, the comment-out mutation fails naming
"Removed ... [GET /tasks/{ticket}]"; reverted, FleetApp.java confirmed clean.
mvn clean install: 0 compile errors, 1257 tests, BUILD SUCCESS.
2026-09-03 15:59:35 +07:00
Dai Ha df9086263d Merge #259: guard the REST route inventory against drift (#252) 2026-09-03 15:56:15 +07:00
Dai Ha eb568ff451 fleetd #252: guard test for the REST route inventory
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m27s
FleetApp's route list has never been checked against anything and has
already drifted once (GET /member-credentials shipped hours before the
#252 ticket and was missing from its list). Add
RestRouteInventoryTest, modelled on McpContractDocTest, which scrapes
FleetApp.java's app.<verb>("path") calls with a regex and compares
them against an explicit expected inventory, failing loudly with the
added/removed routes when they diverge.
2026-09-03 15:54:45 +07:00
Dai Ha ac790e4cce #258: stop two test fixtures writing the operator's real ~/.claude.json
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m29s
seedTrustDialog targets ~/.claude.json when a profile sets no configDir.
Two IDE-overlay fixtures built a worktree-shaped @TempDir, which opens the
#149 isProvisionedWorktree gate, and left configDir null — so every test run
added two project entries to the operator's real file. 116 had accumulated,
none of them still existing on disk, and 32 of those came from current code.

The #149 gate only closed the opposite case: a fixture with cwd unset falling
back to user.dir. A fixture that builds a worktree on purpose walks straight
through it.

- ideProfile/ideProfileModule now take configDir first and mandatory, so each
  fixture states where the trust seed goes.
- noFixtureSeededTheDefaultClaudeJson snapshots the temp-dir project keys in
  @BeforeAll and fails in @AfterAll on any key this class added. Differential,
  not absolute: an absolute check would fail on every host still carrying the
  historical entries, and such a check gets deleted rather than fixed.

Not done: a blanket -Duser.home redirect in surefire. EnvAllowListScrubTest
tests the credential scrub against the operator's real login chain and guards
with assumeTrue($HOME/.zshrc exists), so the redirect would silently skip two
security tests.

Proof: with the bug put back on one fixture the guard fails and names the path;
reverted and confirmed identical with diff -q. A full suite run under a fake
home now creates no .claude.json at all. mvn clean install: 0 compile errors,
1255 tests, BUILD SUCCESS.
2026-09-03 14:01:45 +07:00
Dai Ha 6b5f3f472f Merge #111: the credential probe reads the policy instead of copying it
CI / contract (push) Successful in 1m19s
CI / build (push) Successful in 1m30s
scripts/probe-member-credentials.sh carried its own NAMES array of 31
names. The live policy has 34. The probe reported 26 blocked against a
policy that blocks 29, exited 0, and printed a table that looked
complete. A verification tool that under-reports is worse than none,
because its clean output stops anyone looking.

Same defect as #114, fixed the same way: DELETE the second copy rather
than correct it. The NAMES array is gone, not updated.

The daemon now serves GET /member-credentials — names and counts, never
a value; MemberCredentialPolicyView reads no environment at all, so
there is nothing to redact by construction. The probe fetches it and
refuses with a non-zero exit when the daemon is unreachable, the policy
is absent or empty, or knownCount disagrees with the length of known[].
No local fallback: a verification tool must not quietly degrade into a
weaker check.

The startup log line and the endpoint now share that one class, so the
counting exists once. That also protects a subtlety I measured before
briefing this: blocked is NOT known - allowed. Live, known=34 and
allow=7, but only 5 of those 7 appear in known, so blocked=29 and the
naive subtraction gives 27. The view reuses creds.blockedSet(), the
existing derivation, so it keeps 29.

Worker's mutation: MemberCredentialPolicyView.of(...) forced to return
ABSENT turned 4 tests red with 0 compile errors — including
MemberCredentialsGapReportTest, which proves the startup log really
does run through this path. Reverted and confirmed with diff -q.

It also caught a bug in its own first draft: jq's // operator treats
false and 0 as missing, so `.present // empty` turned a genuine
"present": false into "unknown". Fixed by reading the fields directly.

NOT yet verified: acceptance criterion 5, the live 34/29/5 run. The
route does not exist until the daemon is redeployed onto this jar, so
that check comes next and is mine, not the worker's.

Merged clean, then built on the merged tree: 1255 tests, 0 failures,
0 compile errors.
2026-09-03 13:32:16 +07:00
Dai Ha 38dec72152 Merge #155: refuse the spawn when allow-list policy cannot be enforced
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m15s
policy=allow-list is enforced by a ZDOTDIR scrub, and a non-zsh login
shell ignores ZDOTDIR entirely, so no scrub runs. The launcher already
DETECTED this and logged a WARN — then degraded to the weaker overlay
and spawned anyway. The operator asked for the blocking control and
silently got the weaker one, which is the defect the ticket is about.

Detection existed; refusal did not. Under policy=allow-list a non-zsh
shell now throws IllegalArgumentException before any ZDOTDIR or env
work, naming the actual shell and giving three ways out. Under
policy=deny-by-default nothing changes: that overlay is applied to the
pane before any shell runs, so it does not depend on the shell.

Checked against the live config myself, because this refuses spawns and
no worker can see fleetd.yaml:

  memberHerdrSocket : NOT set  -> the shell comes from fleetd's own
                                  $SHELL, not the unset memberLoginShell
  policy            : allow-list
  fleetd's $SHELL   : zsh, proven by behaviour rather than by reading
                      the process env — the daemon log shows the ZDOTDIR
                      scrub generating a directory 147 times, most
                      recently minutes ago, and that only happens when
                      isZshShell() returned true

So the new refusal cannot fire on this host. Had memberHerdrSocket been
set, the unset memberLoginShell would have read as "<unset>", non-zsh,
and refused every spawn — worth knowing before anyone sets that key.

Worker's mutation evidence, re-stated: `if (!zsh)` -> `if (false)` turned
the refusal test RED with 0 compile errors, then reverted clean.

NOT verified: a live non-zsh member spawn. Forcing it means changing the
daemon's own environment, and the value of the test does not justify
that. The unit tests drive the real launcher.spawn entry point.

Merged clean, then built on the merged tree: 1250 tests, 0 failures,
0 compile errors.
2026-09-03 13:29:53 +07:00
Dai Ha 0e8bfb74fc Merge #176 stage 2: group subscription profiles by account, not by name
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m34s
Stage 1 shipped INERT on this host and every test was green. The matcher
compared effectiveCredentialId(), which fell back to the profile's own
NAME when credentialId was unset. This host runs the lead on `opus` and
members on `sonnet`; both are subscription:true with no credentialId, so
it compared "opus" against "sonnet", never matched, and charged 0 seats.

Every stage-1 test put the lead on the SAME profile name as the target,
so the fixture encoded the one shape the live config does not have.

Stage 2 returns a "<subscription>" sentinel when credentialId is unset
and subscription is true. An explicit credentialId still wins, so an
operator with two genuinely separate Claude logins can keep them apart.

Verified by me on the live config shape, not by reasoning:

  opus.effectiveCredentialId()   = <subscription>
  sonnet.effectiveCredentialId() = <subscription>
  seats charged to sonnet = 1     (was 0 before this change)

Only opus and sonnet join the sentinel group on this host; local,
local-direct, gx, xf, sol and terra are unaffected. free is clamped
with Math.max(0, ...), so the subtraction cannot report a negative.

Second, wider consequence, flagged by the worker and confirmed here:
CompositePeerLauncher.credentialIdFor feeds enforceNotQuarantined and
enforceNotCoolingOff, so quarantining one subscription profile now also
refuses spawns on the other. That is correct — one Claude subscription
hitting a usage limit really does take out every profile on it — but it
is a behavioural change beyond fleet_list's numbers.

Checked all 5 logical callers of effectiveCredentialId(); every one
wants "this account", none wants "this exact profile".

Merged clean, then built: 1250 tests, 0 failures, 0 compile errors.
An auto-merge with no conflicts is not a compiling merge, so the build
was run on the merged tree before this landed.
2026-09-03 13:24:19 +07:00
Dai Ha 51f7b0a3ca fleetd #111: probe reads the live memberCredentials policy, no hardcoded name list
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m22s
scripts/probe-member-credentials.sh carried its own hand-maintained NAMES array
(31 names, recorded 2026-08-16), so a name added later to fleetd.yaml's
memberCredentials.known was never checked and the probe still exited 0 with a
clean-looking table. Same drift shape as #114's tool catalogue.

- New dev.ltms.fleet.member.MemberCredentialPolicyView: the single place that
  turns a MemberCredentials policy into names + counts (never a value). Reused
  by Fleetd.reportMemberCredentialsGap (startup log line) and by the new
  GET /member-credentials REST endpoint (FleetApp), so the two can no longer
  drift apart the way the probe and the policy did.
- FleetApp gains one route + handler + a Supplier<MemberCredentialPolicyView>
  constructor param (legacy constructors default to ::absent, so existing call
  sites are unaffected).
- probe-member-credentials.sh now fetches its name list from
  GET /member-credentials instead of carrying one. No local fallback: an
  unreachable daemon, an empty/absent policy, or a knownCount/known[] length
  mismatch all refuse with a non-zero exit rather than silently checking zero
  names. Prints "policy contains N; this run checked N" so the two numbers are
  visibly equal.
2026-09-03 13:21:35 +07:00
Dai Ha 21c539f22e #113: derive the config guard from the record tree, both directions
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m48s
The existing guard walks KNOWN_TOP_LEVEL_KEYS and anchors its regex at
column 0, so it sees only top-level keys. Every nested key was outside
its scope and nothing said so, which is the shape #113 collects: a
checker narrower than it looks, whose green run stops anyone looking.

Two derived guards replace the assumption:

  everyNestedConfigKeyIsDocumentedInTheExample
      walks FleetConfig's record components (17 records, 83 distinct
      key names) and requires each to be documented in the example.

  everyLiveKeyInTheExampleBindsToARecordComponent
      resolves every live key path in the example against the record
      tree, so a documented key that binds to nothing fails here
      instead of being silently ignored in production.

Neither carries a list, so a key added to any nested record is covered
the moment it compiles (criterion 2).

Both mutations run through the real caller, not the helper (criterion 1,
which asks for exactly that):

  removed every mention of paneProbeIntervalSeconds from the example
      -> FAILS, naming health.paneProbeIntervalSeconds
  added a live bind.totallyMadeUpKnob to the example
      -> FAILS, naming bind.totallyMadeUpKnob

0 compile errors in both; both reverted and confirmed with diff -q.
The first attempt at mutation 1 removed only the `key:` line and the
run stayed green — correctly, because the key was still documented in
prose. An incomplete mutation proves nothing, so it was redone.

Denominators (criterion 3): both guards print how many keys they
checked, and the floor for "did the walk descend?" is derived from
KNOWN_TOP_LEVEL_KEYS.size() rather than being a literal.

Scope is stated in the javadoc rather than implied: the guards do not
check a key sits at the right path, do not parse commented prose for
the reverse direction, and do not prove a parsed key is read by
anything. paneProbeIntervalSeconds is parsed and read by nothing, and
these guards pass it -- the example already says so in its own text.

everyOptionalKnobDocumentedInTheExampleBinds keeps its hand-written
list but is re-documented as a value-binding spot check, explicitly
not a coverage guard; coverage now comes from the two derived tests.

broker.uri is documented only in the example's prose convention
(`#  uri  -> ...`), never as a copy-pasteable `uri:` key, because
writing it out invites pasting a password into a file -- the thing
uriEnv exists to avoid. The matcher accepts that convention rather
than pushing the file toward doing it.

Full build: 1236 tests, 0 failures, 0 compile errors.
2026-09-03 13:18:50 +07:00
Dai Ha bd2774b5f1 fleetd #155: refuse a member spawn under memberCredentials.policy=allow-list on a non-zsh shell
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m56s
The ZDOTDIR scrub that enforces policy=allow-list only runs on zsh. The daemon
already detected a non-zsh login shell (isZshShell/warnNonZsh, from #213), but
degraded to the weaker CB-596 overlay and spawned anyway — the exact "control
silently does nothing" defect this ticket is about. Now a non-zsh shell under
policy=allow-list refuses the spawn (IllegalArgumentException, naming the
shell), surfaced by FleetMcp.spawn's existing catch(IllegalArgumentException).
policy=deny-by-default is unaffected in substance (its overlay never depended
on the shell) but now also logs a one-time WARN naming the shell, since the
stronger allow-list control is unavailable there.

The worktreeRoot/worktreeGroup-missing degrade path under memberHerdrSocket
is untouched — that gap is fleetd #213's scope, not this one.
2026-09-03 13:13:23 +07:00
Dai Ha c50f5b2d61 fleetd #176 stage 2: make effectiveCredentialId() subscription-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m44s
Stage 1's lead-seat matcher (leadSeatLookup) was correct but inert on
the live host: the lead runs on profile 'opus', members on 'sonnet',
both subscription:true with no explicit credentialId. Because
effectiveCredentialId() fell back to the profile's own name, opus and
sonnet never matched even though they share one Claude login, so the
matcher charged zero seats.

FleetConfig.Profile.effectiveCredentialId() now falls back to a shared
sentinel (SUBSCRIPTION_CREDENTIAL_ID = "<subscription>") instead of the
profile name when subscription:true and credentialId is unset. An
explicit credentialId still wins, so two separate Claude logins on one
host can still be kept apart.

This is also BackendQuarantine's and BackendOutagePolicy's grouping
key and CompositePeerLauncher's spawn-time enforcement key, so the fix
also links quarantine/cool-off across subscription profiles sharing an
account -- intentional: one usage limit really does take out every
profile on that login, mirroring credentialId: openai-shared already
doing this for off-subscription profiles. Every caller was reviewed;
none wants "this exact profile" over "this account".

Tests added:
- FleetdLeadSeatLookupTest: the live shape itself (lead on a
  DIFFERENT subscription profile than the target, same account,
  neither sets credentialId) -- the case stage 1's suite never covered
- FleetMcpTest: quarantining one subscription profile's shared
  account zeroes free on another sharing it, via the same
  effectiveCredentialId()-driven wiring Fleetd.main uses

Mutation-tested: reverting the subscription branch to the old
fall-back-to-profile-name behavior sends both new tests RED with 0
compile errors; reverting the mutation restores byte-identical
(diff -q) source and green tests.

fleetd.example.yaml's fleetd #176 notes are rewritten for the sentinel
semantics and when to override it with an explicit credentialId.
2026-09-03 13:10:10 +07:00
Dai Ha 01a840cc14 #248 follow-up: drive the real backendErrorSink, not a copy of it
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m29s
BackendOutageFlowTest held a ~30-line hand-copy of the lambda in
Fleetd.main, under a comment promising it mirrored production "EXACTLY".
That promise was the defect. The test proved the copy, so any change to
the real sink left the flow test green.

#248 made Fleetd.backendErrorSink(...) public for exactly this reason.
The test now calls it.

Measured, same mutation in the real sink (an early return after
sessions.onBackendError, dropping the cool-off and the lead nudge):

  old test (hand-copy):  Tests run: 5, Failures: 0  -- blind
  new test (real sink):  Tests run: 5, Failures: 4  -- catches it

0 compile errors in both runs, so both are real results. Production
reverted and confirmed with diff -q.

Full build: 1234 tests, 0 failures, 0 compile errors.
2026-09-03 13:06:11 +07:00
Dai Ha c4deef08be fleetd #249: withhold agentSessionId when the cwd is not a provisioned worktree
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m22s
A member spawned without a worktree inherits the lead's cwd, which holds
many old opencode session rows. sessionIdForDirectory picks the most
recently updated row for that directory, so a brand-new member - which
has not written its own row yet - resolves to somebody else's session.
Measured: a row three days old, from a different profile.

The damage was at the tool surface. fleet_list told the lead that
agentSessionId is the id to pass as resumeSessionId, so acting on it
would resume a stranger's conversation, with foreign context, and
nothing to distinguish that from a correct resume.

Fixed by refusing to answer rather than by making the heuristic smarter.
#234 already established the heuristic cannot be made reliable at that
layer, and its javadoc records why, so the SQL is untouched.

agentSessionId() now returns null for a non-provisioned cwd, and
spawn() refuses a resumeSessionId request for one outright, before
anything starts. isProvisionedWorktree moved to HerdrPeerLauncher so
both adapters share it. fleet_list and fleet_spawn descriptions no
longer describe the id as always safe to resume.

Verified rather than taken on trust:
- the refusal reaches the lead as a readable message, not a stack trace
  - FleetMcp.spawn already catches IllegalArgumentException and returns
  error(e.getMessage()).
- the message tells the lead to pass fleet_spawn{worktree:<slug>}, which
  is valid: worktree is typed string, 'true' or a ticket slug.
- 13 existing tests moved off a placeholder "/work/dir" onto a real
  provisioned-worktree fixture. They cover #175/#234 model-mismatch
  machinery and would otherwise have tripped the new gate incidentally.

Worker's mutation evidence, both reverted and diff-confirmed:
- gate at OpenCodeLauncher:812 -> if(false): RED at
  OpenCodeLauncherTest:448, expected <null> but was <ses_someone_elses>.
- resume refusal at OpenCodeLauncher:683 -> 'false &&': RED at
  OpenCodeLauncherTest:379, expected IllegalArgumentException.
Both 0 compile errors.

Closes #249. PR #253.
2026-09-03 12:55:29 +07:00
Dai Ha c796eac09c fleetd #176: subtract the lead's own subscription seat from free
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m16s
maxLoad counted panes, never subscription seats: a subscription:true
profile's lead is itself a live claude session on that same account,
so free overstated capacity by the lead's own seat (measured free:1
with a real ceiling of 0, and free:3 on an idle fleet with a real
ceiling of 2).

Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/
OutageSource) and Fleetd.leadSeatLookup, which derives the seat count
from fleet.leaders.<name>.profile matched against the target profile
by effectiveCredentialId() - no hardcoded "-1", and no new config key:
profile: already exists for this exact "which account does this lead
share" question. maxLoad itself is left untouched; only free (and a
new, additive-only leadSeats field) changes.

Exhaustion quarantine (cause 2 in the ticket) already forced free to 0
via the same BackendQuarantine capacityView already reads - confirmed
by reading the exhaustionSink wiring, no code change needed there.
2026-09-03 12:55:08 +07:00
Dai Ha e897e5257b fleetd #114: delete the drifted tool catalogue, keep the flows, guard the names
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
docs/MCP-Contract.md was written 2026-07-14, before any MCP code existed,
and never caught up. CLAUDE.md points every session in the fleet at it.

Audited against the code today. The drift was not confined to the tool
table the ticket reported:

  section 3  still described the OLD identity rule - "any connection that
             does not map to a known worker is treated as a primary".
             That was a real privilege bug, fixed since by the ancestry
             walk in #161. The page still taught it.
  section 4  names port 8080 (the mount is 8765) and says the pom does
             not yet carry an MCP dependency.
  section 5  named fleet_read and fleet_cancel, which do not exist, and
             omitted fleet_poll, fleet_ack, fleet_profiles, fleet_whoami
             and fleet_list, which do.
  section 8  says turn_id where the code says turnId, and has no row for
             the exhausted outcome CB-578 added.
  sections
  9, 10, 11  pre-build planning: "new work" columns, open decisions long
             since decided, CB-1xx placeholders.

Every one of those is the same defect: a hand-maintained second copy of
something the code already states. So the copy is deleted rather than
corrected - correcting it just restarts the clock.

What survives is the flows and the status gating, because a flow is a
shape rather than a name, and shapes are what this page was ever good
for. They are rewritten with the names checked against the code, and
extended with what has been learned since: the ~60s cap on a blocking
send, the ~55s ask window, and the three ways the turn-done fallback
loses a report (clipped, echoed brief, slow member).

389 lines -> 188.

The names that remain are guarded. McpContractDocTest fails if the page
names a fleet_* tool FleetMcp does not register, and - because an empty
set is a subset of everything - a second test pins that both sides
actually found names, so the check cannot pass by checking nothing. A
third pins the "this is not the tool reference" sentence, which is the
fix itself: without it someone helpfully re-adds a tool table.

Mutation-tested both ways, 0 compile errors each: adding `fleet_read` to
the doc fails theDocNamesNoToolThatDoesNotExist ("names [fleet_read] ...
Checked 6 name(s)"); removing the disclaimer fails
theDocStillDisclaimsBeingTheToolReference.

All 5 mermaid diagrams render under mermaid-cli.

CLAUDE.md's pointer said "section 6 only" and now names the guard
instead. It is in the project addendum, so the canonical block is
untouched - verified still byte-identical with the wiki template.

REST is split out to #252: 14 routes, documented nowhere, and it IS a
supported operator surface - one of them drains on read.

1232 tests, 0 failures.
2026-09-03 12:52:12 +07:00
Dai Ha 2afa3652bb fleetd #249: withhold agentSessionId for a non-provisioned opencode cwd
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m52s
OpenCodeSessionDiscovery.sessionIdForDirectory keys on the worker's cwd, which
is reliable only when fleetd provisioned a unique git worktree for that
member. Without one (the default no-worktree spawn), the cwd is shared with
other sessions, and "most recently updated row for this directory" can pick a
stranger's session — fleet_list would then hand a lead an agentSessionId that
resumes someone else's conversation.

Move isProvisionedWorktree from ClaudeCodeLauncher to the shared
HerdrPeerLauncher base (both adapters need it now). OpenCodeLauncher.spawn now
refuses a resumeSessionId spawn outright when the target cwd is not a
provisioned worktree (fleetd can never verify or re-report that identity), and
SessionAwareHandle.agentSessionId() withholds the id — returns null rather
than guessing — for any member spawned without one, resumed or not. Corrected
fleet_list/fleet_spawn's tool descriptions, which previously implied
agentSessionId is always a safe resume handle.
2026-09-03 12:51:56 +07:00
Dai Ha 80092ff359 fleetd #247: stop writing a trust key Claude Code strips on every save
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m28s
seedTrustDialog wrote two keys into the shared .claude.json:
hasTrustDialogAccepted and hasCompletedProjectOnboarding. Only the first
one survives.

Measured live on 2026-09-03, minutes after a spawn seeded the file:

  hasTrustDialogAccepted:       28 of 28 project entries
  hasCompletedProjectOnboarding: 0 of 28 project entries

Our entry was written by the running jar and the key was already gone,
so it was written and then removed. It is absent from the 27 entries
Claude Code wrote for itself too, which says Claude Code normalises the
whole file when it saves and drops that key every time.

That reframes #247. I filed it as a race - a save landing between our
read and our ATOMIC_MOVE. It is not a race. The other writer removes
this key as its steady-state behaviour, with no window involved. So the
compare-and-swap retry proposed there would not have helped: it would
re-add a key that gets stripped again on the next save.

The seed's whole job is to stop the workspace-trust dialog blocking a
member (#149). The live probe reached idle with hasTrustDialogAccepted
alone, so the second key was never doing that job. Writing it only added
a contested key to a file two processes share, and made the next reader
think it mattered.

The atomic write and the lock stay. Both are still correct, both are
cheap, and hasTrustDialogAccepted is genuinely shared state.

The new assertion is assertFalse, not a deletion. Removing the old
assertion would leave nothing to stop someone re-adding the key later as
a plausible-looking completeness fix. Mutation-tested: restoring the
production line fails
seedTrustDialogWritesOnlyTheTrustFlagAndNotTheOnboardingKey:2198 with 0
compile errors.

1229 tests, 0 failures.
2026-09-03 12:41:20 +07:00
Dai Ha 9d37f3aa29 fleetd #201: the coverage line must name the key its caller actually means
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m56s
Found by reading a real boot log after the redeploy, not by a test.

coverage() is shared by two call sites — CB-578's exhaustedPattern line
and Unit 5's errorPattern line — but its 'off' branch hard-coded the
word exhaustedPattern. So this daemon printed:

  backend-exhausted classification (CB-578 stage A): partial
      (configured: [sol, terra]; not configured: [...])
  backend-error classification (fleetd #201 Unit 5): off
      (no profile has an exhaustedPattern configured; profiles: [...])

Two lines, one directly under the other, disagreeing about whether any
profile has an exhaustedPattern. Both were individually defensible and
together they were nonsense. Worse, the message sends an operator to
set the wrong key: the thing that is missing is errorPattern.

coverage now takes the key name. I changed the signature rather than
adding an overload, so the compiler found all three existing callers
instead of leaving them silently on the old path.

Every earlier coverage test passed the exhaustion case only, which is
why none of them could see this. The new test pins the errorPattern
case. Reverting the fix turns it red with 0 compile errors.

1229 tests, 0 failures, BUILD SUCCESS.

This is the second defect in two hours found only by reading the live
startup log — see #115, where the noise of a false warning had been
hiding a correct line saying a whole feature was off.
2026-09-03 12:29:27 +07:00
Dai Ha eaf89abaf6 fleetd #248: make Fleetd's CompletionResolver wiring provable
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m40s
Before this, dropping either #241's worktree lookup or Unit 5's
backend-error pair at Fleetd.main's new CompletionResolver(...) call
left all 1216 tests green with 0 compile errors. Every existing test
built its own CompletionResolver, so they proved the class and never
the wiring. BackendOutageFlowTest was the sharpest case: it copies
main's sink lambda line-for-line, so it proves the copy and cannot
notice the original being deleted.

The three inline arguments are now package-private static factories on
Fleetd, following the deliverableTo pattern the file already had, each
with its own behaviour test. backendErrorSink is public so a
cross-package test can drive the real production object rather than a
hand-mirrored copy.

The test that was actually missing is a source-text assertion. That is
the honest fallback for a composition root with no seam, and it is
labelled [SOURCE TEXT] in every test name and message so it cannot be
misread as a behaviour check. It is not vacuous: two tests pin that the
variables are assigned from the factories, and two pin that those
variables reach the call site, so renaming a variable while assigning
an inert value does not slip through.

Known cost, accepted: the assertions match exact source substrings, so
reformatting that statement will break them. That is the price of
covering a main method, and a spurious failure here is loud and
obvious, which is the right direction to fail.

Verified by the lead, both mutations re-run against the merged code —
see the merge check.

PR #251
2026-09-03 12:22:32 +07:00
Dai Ha d895f02bc1 fleetd #248: prove main() wires CompletionResolver's arguments, not just the class
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 1m40s
Fleetd.main built three of CompletionResolver's 8 constructor arguments inline
(a worktree/branch lookup lambda, and the backend-error pattern lookup + sink
locals). Dropping any of them at the call site compiled clean and left every
existing test green, because every existing test constructs its own
CompletionResolver and only ever proves the class, never main's wiring.

Extract each into a static factory on Fleetd (worktreeBranchLookup,
backendErrorPatternLookup, backendErrorSink — the same static-factory pattern
Fleetd.deliverableTo already uses), test each factory's own behaviour, and add
a source-text assertion (FleetdCompletionResolverWiringTest) proving main's
CompletionResolver call still passes all three. backendErrorSink is public so
BackendOutageFlowTest can exercise the real production sink directly instead
of the hand-mirrored copy its own class doc used to describe.

No production behaviour changes — mechanical extraction only.
2026-09-03 12:20:28 +07:00
Dai Ha 43206cac2f fleetd #148 point 2: drop .envrc from the default parity overlay
CI / contract (push) Successful in 1m9s
CI / build (push) Successful in 1m18s
The default is now [.env], not [.env, .envrc]. .env is data, so copying
it into a worker worktree can only move values. .envrc is executable
shell that direnv runs on every cd, so copying it moves behaviour. Those
are different risks and should not share a default.

The knob is unchanged. An operator who wants .envrc copied writes
parityOverlay: ['.env', '.envrc'] and owns that choice; a new test pins
that escape hatch, because without it this would be a removal rather
than a re-default.

Decision recorded on the ticket, with the evidence it asked for first:
this checkout has no .env and no .envrc, and direnv is not on PATH, so
there was no live exposure. Point 1 (extend the credential scrub to
direnv) is declined and the reason is on the ticket — the scrub is a
one-shot .zlogin and a direnv hook runs on every cd, so no amount of
work on the scrub can cover it. Not copying the executable file is the
smaller change and removes the need.

The worker also fixed WorktreeSessionManagerTest, which hardcoded the
same default at another layer and broke the build. Outside its named
scope, correctly flagged rather than done silently.

Verified by the lead: 1216 tests, 0 failures, 0 compile errors.

PR #250
2026-09-03 12:03:21 +07:00
Dai Ha 8bba3a8184 fleetd #148 (point 2): drop .envrc from the default parityOverlay
CI / contract (pull_request) Successful in 1m20s
CI / build (pull_request) Successful in 1m32s
.env is data; .envrc is executable shell that direnv runs on every cd, so
copying it into a worker moves behaviour, not just values. The default
parityOverlay is now [.env] only. The knob is unchanged: an operator who
wants .envrc copied can still write parityOverlay: [.env, .envrc]
explicitly.

Updates FleetConfig's default and javadoc, fleetd.example.yaml's two
mentions of the default, and the FleetConfigTest coverage: renamed the
default test, added parityOverlayExplicitEnvrcOptInStillWorks to prove
the .envrc opt-in escape hatch still works, and fixed
WorktreeSessionManagerTest#worktreeAcquireRunsParityOverlayWithProfileDefaults
which also hardcoded the old default.
2026-09-03 12:00:13 +07:00
Dai Ha 5cf3ca9a89 fleetd #241: never hand the lead back its own brief as the member's report
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m27s
The completion fallback scrapes a member's pane when a turn ends with no
fleet_reply. If the pane still shows the brief the lead injected, the
scrape returned that brief, and the lead read its own words as the
member's answer. A silent member looked like a member that had reported.

echoesInjectedBrief now recognises that case and refuses it. Round 1 used
plain containment in both directions, which destroyed real reports: a
genuine report that quotes the brief contains it. Round 2 keeps the safe
direction unbounded (the brief contains the scrape) and bounds the other
one at MAX_ECHO_EXCESS_CHARS, so a scrape only counts as an echo when it
adds almost nothing to the brief.

Merge note — the Fleetd.java conflict:

This call site was changed by both #201/#227 Unit 5 (backendErrorPatterns
+ backendErrorSink) and by this ticket (the worktree/branch lookup). I
resolved it onto the full 8-argument constructor so neither feature is
dropped; nowNanos has to be passed explicitly to reach that overload.

Verified by the lead: 1215 tests, 0 failures, 0 compile errors.

I also measured whether the resolution itself is protected, and it is
NOT. Both mutations at this call site stay green:
  - drop the worktree lookup (pass _ -> null): 1215 tests, 0 failures
  - drop Unit 5's patterns/sink (legacy()/none()): 1215 tests, 0 failures
Nothing in the suite covers Fleetd's composition root, so either feature
could be silently unwired here and the build would still be clean. The
tests prove the seams, not the caller. Filed separately rather than
fixed in a merge commit.

PR #245, branch worker/cb241-fallback-echo-1175e9-11
2026-09-03 11:54:06 +07:00
Dai Ha ac474981e4 fleetd #201/#227 Unit 5: wire the backend-error cool-off into config, placement and the MCP surface
A profile's credential that throws two distinct backend errors inside 60
seconds now cools off for 60 seconds. Automatic placement skips it,
an explicit fleet_spawn naming it is refused before the adapter is
called, and fleet_list/fleet_profiles report it as coolingOffForSeconds
next to the separate CB-578 quarantinedForSeconds.

Verified by the lead: see the merge check below. The worker ran 7
mutations, all killed with 0 compile errors; M7 was NOT killed on the
first pass (the assertion only checked .contains("quarantined"), which
is true of both the correct message and the mutated fallback), and the
worker strengthened it to assertEquals on the exact literal and kept
that change. That is the right call and it is reported honestly.

Two deviations, both justified in the PR:
- FixedPlacementPolicy needed the same coolingOff filter because it
  filters candidates inline instead of using PlacementPolicyUtil.
- BackendOutageFlowTest sits in dev.ltms.fleet.inject because
  CompletionResolver.InFlight is package-private there.

The startup coverage log line is a code-reading claim, not a captured
line from a live daemon. The worker said so rather than overclaiming.

PR #246, branch worker/cb201-unit5-wiring-6c12e6-8
2026-09-03 11:48:42 +07:00
Dai Ha ba04b2359b fleetd #201/#227 unit 5: wire errorPattern, cool-off spawn gate, and fleet views
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m52s
Wires the already-merged units into production:
- Per-profile errorPattern config (beside exhaustedPattern), compiled once at
  startup; falls back to the legacy (?i)\bAPI Error\s*: pattern when unset.
  Startup logs configured-vs-legacy coverage, same as exhaustedPattern.
- One production BackendErrorSink in Fleetd.java: mark backend_error on the
  session, resolve the profile's credential fail-loud (never
  Optional.ifPresent), record it in BackendOutagePolicy, and push a lead
  nudge on a new incident.
- CompositePeerLauncher's explicit and automatic spawn paths both refuse a
  cooling-off credential; exhaustion quarantine wins when both are active.
  PlacementContext gets a separate coolingOff set so refusal text says
  "cooling off", never "exhausted".
- fleet_list/fleet_profiles report coolingOffForSeconds as an independent
  fact from quarantinedForSeconds; both can appear together.
- fleetd.example.yaml documents errorPattern and the 2/60/60 cool-off policy;
  CLAUDE.md tells leads how to read the two independent outage states.

Also: FixedPlacementPolicy.java, not in the original file list, needed the
same coolingOff filtering as PlacementPolicyUtil (it does its own inline
candidate filtering rather than delegating).

1190 tests, 0 failures (up from the 1163 baseline); BUILD SUCCESS.
2026-09-03 11:45:11 +07:00
Dai Ha 2e5b63f6f6 fleetd #149: seed the workspace-trust entry before a claude-code spawn
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m29s
Claude Code asks 'is this a project you trust?' the first time it starts in a directory it
has not seen. It is interactive with no timeout, and every member spawned with worktree:true
lands in a brand-new directory. The member never reaches its first turn and never replies,
while herdr reports blocked/interactive_ready — which reads as healthy.

ClaudeCodeLauncher now seeds projects.<cwd>.hasTrustDialogAccepted in the profile's
.claude.json before the process starts. This is not a new grant: the operator already
trusted the repo by configuring the profile against it, and a worktree is a checkout of it.

The write is gated on isProvisionedWorktree(cwd) — a .git that is a regular gitdir-pointer
file, never a real checkout. That gate exists because an earlier revision of this change,
run under mutation testing, wrote to the operator's real ~/.claude.json and truncated it
from 72KB to 919 bytes. Tests using a null configDir fall back to the real user.home, so
an ungated seed reaches real files.

The write is atomic (sibling temp file + ATOMIC_MOVE, never truncate-in-place) and the
whole read-modify-write is under a lock, because .claude.json is large, live, and rewritten
by Claude Code itself while fleetd runs. Two parallel spawns are normal here.

Verified by the lead: 1173 tests, 0 failures. A truncating write turns the torn-read test
red; removing the lock turns concurrentSeedsForDifferentCwdsBothSurvive red. Both with 0
compile errors. copyPosixPermissionsIfPresent is NOT covered by a test — its mutation stays
green — but createTempFile is 0600 on POSIX by default, so the not-world-readable property
holds without it; the line only preserves a non-default mode.
2026-09-03 11:44:03 +07:00
Dai Ha 3437d6313d fleetd #241: bound the echo match so a real report is never swallowed
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Successful in 1m25s
Round 1 used plain bidirectional containment. The direction that catches the real bug --
the pane holds the brief plus a status bar, so the scrape contains the brief -- also fires
when a member restates the whole brief and then writes a genuine report under it. That
threw the report away and told the lead nothing was produced, which is worse than the bug
being fixed: it destroys a delivery instead of merely obscuring one.

The safe direction (the scrape is a fragment of the brief) stays unbounded, because a
fragment of the brief is by definition not a report. The dangerous direction now requires
the scrape to add at most MAX_ECHO_EXCESS_CHARS beyond the brief, which is the amount of
TUI chrome a real echo carries.

Work by the cb241 worker, committed by the lead: its backend stopped answering after the
fix was written, so two turns ended with no commit and no reply. Verified by the lead:
1169 tests, 0 failures; removing the bound turns pinsTheMaximumTuiChromeExcess and
completionFallbackKeepsARealReportThatRestatesTheWholeBrief red with 0 compile errors.
2026-09-03 11:39:11 +07:00
Dai Ha 743377d6cd fleetd #149 review round 2: make the trust-dialog seed atomic and lock-protected
CI / build (pull_request) Successful in 1m21s
CI / contract (pull_request) Successful in 1m25s
Files.writeString truncates the target in place before writing, so there
was a window where .claude.json could be observed empty or half-written
- exactly the shape of the incident this ticket already hit once, but
reachable in production too: a crash/kill mid-write, or two concurrent
claude-code spawns (normal here - several run in parallel routinely)
racing a naive read-modify-write and silently discarding one spawn's
entry.

Two independent fixes, each with its own dedicated test proving it (not
the other):

- ClaudeCodeLauncher.writeAtomically: serialise to a sibling temp file in
  the same directory, then Files.move with ATOMIC_MOVE + REPLACE_EXISTING,
  preserving the target's existing POSIX permissions (.claude.json ships
  0600). A reader now only ever observes the fully-old or fully-new file,
  never a torn one. Package-visible so a test can drive it directly.
- TRUST_JSON_LOCK: a process-wide lock around seedTrustDialog's whole
  read-modify-write, so two concurrent spawns for different cwds both
  keep their entry instead of the second write discarding the first.
  Sufficient because every spawn on this daemon runs in one JVM; it does
  NOT protect against a second daemon process or the operator's own live
  Claude Code writing at the same instant - writeAtomically covers that
  case instead.

Both fail soft, same as before: any I/O failure here must never block a
spawn.

Four new tests: a large (30-project) existing file survives without
collapsing (asserted on the restored key set, not just that the result
parses); two concurrent spawns for different cwds both keep their entry
(CountDownLatch-synchronised, not a sleep); existing 0600 permissions
survive the write; and a direct test of writeAtomically with a busy-poll
reader thread proving a concurrent reader never observes a torn file.

See PR body for the full mutation-testing table, including an
honest note on which of these tests the atomicity mutation actually
caught (not the one implied by the numbering in review) and why.
2026-09-03 11:38:10 +07:00
Dai Ha 5952d559c7 fleetd #134 point 3: tell the lead which files a worktree neutralizes
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m26s
The daemon log and the worktree git config are both new in 205ad82, but neither helps a
lead who is writing a brief. This is the line that does: never brief a worker to edit
.mcp.json, opencode.json or .autoenv in its worktree, because the edit cannot be
committed and nothing will say so.

Goes in the project addendum, not the canonical block, so the byte-sync with
wiki/7-Use-Cases.md is unaffected — re-checked and still true.
2026-09-03 11:32:03 +07:00
Dai Ha 205ad823b0 fleetd #134: make tool-surface neutralization visible to the daemon and the worker
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m34s
isolateToolSurface replaces .mcp.json, opencode.json and .autoenv with stubs in every
provisioned worktree and marks them --skip-worktree. That neutralisation is correct and
is unchanged here — the committed files would mount the primary's credentials.

The problem was that it was invisible. A worker told to edit opencode.json read a 3-byte
stub and reported, truthfully and wrongly, that the mount key did not exist. A missing
file would have prompted a question; a plausible stub did not.

Two changes, both visibility only. The daemon now logs one info summary per provisioning
with the denominator, the files neutralised, and the consequence. And the list is recorded
in worktree-scoped git config (fleet.neutralizedConfig / fleet.neutralizedConfigNote) so a
worker can discover it from inside its own worktree with
'git config --worktree --get-all fleet.neutralizedConfig'.

Worktree-scoped config was chosen over a file in the working tree because it lives in
.git/worktrees/<nonce>/config.worktree and so can never appear in git status, and because
configureEnvironmentCredentialHelper already uses the same mechanism in the same add() call.

Verified by the lead: baseline 1172 tests, 0 failures. Reverting the summary to log.debug
goes red (2 tests), and recording into --local rather than --worktree — which would leak
the record into the shared repo config — goes red too. Both with 0 compile errors.
2026-09-03 11:31:24 +07:00
Dai Ha d654ccb818 fleetd #134: make tool-surface neutralization visible to the daemon and the worker
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m35s
isolateToolSurface replaced .mcp.json/opencode.json/.autoenv with neutral stubs and
marked them --skip-worktree, but said nothing anywhere. A real worker read a 3-byte
{} stub for opencode.json, where the repo's real file is 30+ lines, and truthfully
(but wrongly) reported a mount key did not exist.

Two readers, two fixes:
- the daemon operator gets one info log per provisioning, naming the denominator,
  what was neutralized, and why anything was not (same shape as overlayParity's
  fix in #148 point 3).
- the worker gets the same fact recorded in worktree-scoped git config
  (fleet.neutralizedConfig / fleet.neutralizedConfigNote), discoverable with
  `git config --worktree --get-all fleet.neutralizedConfig` from inside its own
  worktree, without asking the lead. Not a working-tree file: this repo already
  uses worktree-scoped config for the credential helper and the SSH->HTTPS
  rewrite, and it lives under .git/worktrees/<nonce>/ so it can never appear in
  `git status` for the worker to trip on or commit.

The neutralization itself (stub content, --skip-worktree marking) is unchanged.
2026-09-03 11:24:41 +07:00
Dai Ha a89dcc9b7e fleetd #149: seed the workspace-trust entry before a claude-code spawn
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m43s
A claude-code member spawned into a fresh worktree hits an interactive,
un-timed workspace-trust prompt on its first start in a directory it has
never seen. It never reaches its first turn and never mounts the bridge.

Fix: ClaudeCodeLauncher.seedTrustDialog writes
projects.<cwd>.hasTrustDialogAccepted / hasCompletedProjectOnboarding into
the profile's configDir/.claude.json (or ~/.claude.json when configDir is
unset) BEFORE the herdr spawn call, additively (existing keys/projects are
preserved). Gated to isProvisionedWorktree(cwd) - a .git that is a regular
gitdir-pointer file, never a real checkout's .git directory - the same
signal writeIdeOverlay already used, now shared between both.

That gate is a fix for a real incident hit while building this: an
earlier ungated version ran against this file's own pre-existing tests
(configDir=null, no cwd -> falls back to the real user.dir and
~/.claude.json) and corrupted the operator's actual ~/.claude.json down
to a single entry during a mutation-testing run. See PR body for the
full incident report.

FakeHerdr gained onAgentStart(Runnable) so a test can assert the seed
is on disk at the exact instant herdr's agent.start call is reached -
i.e. strictly before the peer process itself would start.
2026-09-03 11:23:06 +07:00
Dai Ha 0d7b4fb026 fleetd #134 + #148 point 3: make the parity overlay say what it did
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m35s
Both defects lived in one method. overlayParity logged every step at debug, so at the
default level the copy was silent and nobody could tell which overlay files a member
actually got. It also marked a copied tracked file --skip-worktree and said nothing, so a
worker editing that file later found git ignoring the change with no error anywhere.

The summary now reports the denominator, not a bare count: 'copied 1 of 2 candidates:
.env (.envrc absent)'. A bare 'copied 1' is the same under-reporting shape as #113.
Neutralised files are named with the consequence in the message itself.

No marker file is written into the worktree: acceptance criterion 1 requires the worktree
to hold exactly the configured overlay set, so a marker would violate the fix it documents.

The copy and mark logic is unchanged — only logging is new.

Verified by the lead: baseline 1168 tests, 0 failures. Reverting either log.info to
log.debug goes red (2 reds and 1 red, 0 compile errors each), which is the regression that
matters since the whole fix is the log level.
2026-09-03 11:14:21 +07:00
Dai Ha 321d8dcbb5 fleetd #241: suppress echoed fallback briefs
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m20s
2026-09-03 11:10:48 +07:00
Dai Ha ef8c97871e fleetd #134/#148 point 3: make overlayParity's copy and skip-worktree visible
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m53s
overlayParity logged everything at debug, so at the default level nobody could
tell which overlay files a spawn actually received (#148 pt 3), and a tracked
file marked --skip-worktree gave no warning that it can no longer be edited
from that worktree (#134).

Report the outcome at info: a per-spawn summary naming the denominator (every
configured candidate), what was copied, and why anything was not — plus a
separate line naming every file marked --skip-worktree, stating plainly that
it cannot be committed from this worktree. No worktree-local marker file: the
worktree must hold exactly the configured overlay set and nothing else, so an
extra file would violate that invariant.
2026-09-03 11:09:50 +07:00
Dai Ha 26bafe824b fleetd #201/#227 unit 1: classify a backend error at the scrape
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m54s
CompletionResolver already reads the pane on every turn, so the classifier lives there
rather than in a new watcher. A matched pattern resolves the waiter as a failure and
fires BackendErrorSink, always inside the resolveFailure win-gate so exactly one thread
reports one incident.

The too-fast path now takes a fresh scrape instead of reusing the pre-turn text, so a
backend that dies immediately is still classified. BackendErrorSink is a single-method
functional interface by design — see #234 for what a default overload does to a lambda.
2026-09-03 10:53:51 +07:00
Dai Ha 838a701109 fleetd #234: carry the profile hint to the exhaustion sink
An exhausted opencode member could not always be mapped back to a credential, because
the launcher knew the profile and the sink did not. The sink now takes a profile hint.

The interface is inverted on purpose: the three-argument method is the single abstract
method and the two-argument one is the default. A lambda can only implement the abstract
method, so every lambda is now forced to carry the profile. The first round of this fix
added the third argument as a default overload, and the production forwarder in Fleetd
was a two-argument lambda — so the fix compiled, passed its tests, and never ran.

Verified by the lead: deleting the forwardingTo factory's override, and rewriting the
Fleetd call site as a plain lambda, both fail to compile now rather than passing silently.
2026-09-03 10:52:01 +07:00
Dai Ha e5eb3534c7 fleetd #201/#227 unit 3: lead outage nudge
Backend incidents become a fourth source inside ReplyPushLoop, not a new scheduler — a
second injector would race the one control that already owns lead-pane delivery. One
notice per incident per affected lead (not per member), one-shot, waiting while the lead
pane is not injectable, and combined into the same nudge as any pending failed ticket.

A classified target that cannot be mapped to a credential gets its own truthful notice
and its own pending/delivered records. It previously reused the incident message, which
told the lead a credential was 'cooling for 0 remaining seconds' when nothing was cooling,
and smuggled the free-text reason into the profiles field.

Verified by the lead: 1129 tests green; routing the unmapped path back through
onBackendIncident turns unmappedBackendTargetUsesTheKnownLeadSchedule red with 0 compile
errors, and that test now asserts the whole rendered message rather than two substrings.
2026-09-03 10:49:25 +07:00
Dai Ha 959c83534f fleetd #201/#227 unit 2: credential outage policy
Adds BackendOutagePolicy — a credential-keyed state machine on an injected monotonic
clock. Two classified backend errors from two DISTINCT targets on one credential inside
60 seconds mint one incident and start a 60-second cool-off. Errors during cool-off
neither extend it nor mint another; expiry clears evidence, so two fresh errors rearm.

Correlated on credentialId, never on profile name or error text. Deliberately not
BackendQuarantine: that restarts a 1800-second cooldown per exhaustion, and its name
would make every refusal say 'backend exhausted', which is a different condition.

Threshold counts distinct targets rather than raw events (lead decision): the classifier
is a heuristic and a valid member report can quote an 'API Error:' line, so one member
repeating that line must not remove a healthy credential's capacity. A real outage hits
every member on the credential, so true detection is unaffected.

Verified by the lead: 1135 tests green; reverting evidenceCount() to reasons.size()
turns two BackendOutagePolicyTest cases red with 0 compile errors.
2026-09-03 10:49:10 +07:00
Dai Ha 31b028e860 fleetd #234 round 4: invert ExhaustionSink's abstract method so the bug class is unrepresentable
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m23s
Round 3's factory fixed the two known call sites but the underlying shape
was still there: a lambda written against ExhaustionSink binds to whichever
overload is abstract, and the 2-arg form held that position, so ANY lambda
-- a call-site forwarder, a hand-built test double, a future caller who has
never heard of fleetd #234 -- could still silently take the hint-dropping
default. Two rounds shipped exactly that mistake in two different places.

Fix: made the 3-arg onExhausted(target, reason, profile) the interface's
single abstract method; the 2-arg form is now a default that delegates with
a null profile. A lambda declared against ExhaustionSink today is forced by
the compiler to take three parameters -- there is no overload left for it to
bind to that can drop the hint. This is enforced by the type system, not by
a test that has to remember to check for it.

Knock-on changes:
- ExhaustionSink.none() -- a 3-arg lambda, still a genuine no-op, now safe
  by construction rather than by care.
- ExhaustionSink.forwardingTo(...) -- collapses to a one-line 3-arg lambda;
  kept as a named factory (round 3's lesson: a test must call the real
  object, not rebuild its shape).
- Fleetd.java's real sink and the two OpenCodeLauncherTest sinks that used
  to be anonymous classes overriding both overloads are now plain lambdas
  too -- the 2-arg override each carried was pure boilerplate once the
  interface provides it as a default.
- CompletionResolver.java itself: UNCHANGED, zero diff (confirmed via
  `git diff --stat` before staging) -- its two call sites still call the
  2-arg onExhausted(target, reason), which is now the default and behaves
  identically. CompletionResolverTest (41 tests, 0 failures) proves this;
  its five ExhaustionSink lambdas needed a mechanical third parameter added
  to keep compiling against the new abstract method, no assertion changed.

Mutation proof, re-run against the new shape: forwardingTo's body edited to
call the 2-arg default instead of passing the hint through (the equivalent
of round 3's "delete the 3-arg override" now that there is only one method
to break) -- both new tests go red with the same assertions as round 3:

  ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
  OpenCodeLauncherTest...ForwardingHop:  expected: <true> but was: <false>
  Tests run: 68, Failures: 2

Restored, re-ran: green (Tests run: 109, Failures: 0, including
CompletionResolverTest).

Compiler proof (not committed -- a scratch file outside the worktree,
compiled with the real ExhaustionSink.java on the classpath, then deleted):

  ExhaustionSink forwarder = (target, reason) -> System.out.println(target + reason);

  error: incompatible types: incompatible parameter types in lambda expression

A 2-arg lambda against this interface no longer compiles at all.

mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:46:27 +07:00
Dai Ha 7840e9adf6 CB-201: make unmapped target notice truthful
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m15s
2026-09-03 10:44:18 +07:00
Dai Ha c3672f5472 fleetd #201/#227 unit 4: durable BACKEND_ERROR member outcome
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m47s
Adds MemberSession.State.BACKEND_ERROR with a nullable failureReason, surfaced in
rosterView, and SessionManager.onBackendError(target, reason). The CAS loop accepts
both sides of the completion race (BUSY and DONE); BACKEND_ERROR is terminal.

completeTurn now returns early when its CAS loses, so a stale DONE copy can no longer
release the pane or reset its context behind a member that just went BACKEND_ERROR.

Verified by the lead: 1130 tests green; mutating the completeTurn early return back to
the old fall-through turns losingCompletionDoesNotReleaseOrClearABackendErrorMember red
with 0 compile errors.
2026-09-03 10:43:58 +07:00
Dai Ha cf54aed451 CB-201 unit 2 review fix: threshold counts distinct targets, not raw events
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m46s
Two errors from the same target inside the window must never trip the
outage threshold on their own (a valid member report can legitimately
quote an "API Error:" line twice) — only two DIFFERENT targets on the
same credential do. Change evidenceCount() to targets.size() instead
of reasons.size(); reasons() still keeps every event, including
same-target repeats, so it can be longer than evidenceCount(). A real
outage still hits every target on the credential, so this loses no
true-positive coverage while cutting a real false-positive path.
2026-09-03 10:42:37 +07:00
Dai Ha bbf68f3e3c CB-201: cover losing completion CAS
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m37s
2026-09-03 10:35:36 +07:00
Dai Ha 776743cbe2 fleetd#201 Unit 1: typed backend-error classification in CompletionResolver
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m45s
Replace the hardcoded API-Error check with a target-keyed BackendErrorPatternLookup
plus a BackendErrorSink, mirroring the existing ExhaustedPatternLookup/ExhaustionSink
pair. Classifies in all three paths (normal block, #211 raw-scrape fallback, and the
fleetd#164 MIN_TURN_NANOS floor). The sink fires only after Rendezvous.resolveFailure
wins for the exact waiter. A target with no configured pattern still falls back to the
narrow (?i)\bAPI Error\s*: compatibility pattern. Existing constructors keep compiling
via BackendErrorPatternLookup.legacy() / BackendErrorSink.none() defaults.

Public send result is unchanged (still a failed send) — the typed sink event is the
internal seam Unit 5 will consume.
2026-09-03 10:33:59 +07:00
Dai Ha 826e0aeb2a CB-201 unit 2: credential-keyed backend outage policy
CI / build (pull_request) Successful in 1m11s
CI / contract (pull_request) Successful in 1m24s
Add BackendOutagePolicy: two classified backend errors on the same
credentialId within a 60s window mint one Incident and start a 60s
cool-off for that credential, one atomic ConcurrentHashMap.compute()
per credentialId so a concurrent second and third event can never
both cross the threshold. Errors during cool-off are ignored outright
(no extension, no incident); once cool-off elapses the next error
clears old evidence, requiring two fresh errors to rearm. This is a
new class, deliberately not BackendQuarantine (wrong store, wrong
1800s duration, misleading "exhausted" semantics for a 60s transient
fault). Knows nothing about panes, profiles, sessions, launchers, or
leads — takes events in, returns incidents out.
2026-09-03 10:30:40 +07:00
Dai Ha c935b181dd fleetd #234 round 3: make ExhaustionSink's forwarder a shared factory, not a rebuilt-per-caller shape
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m36s
Round 2's tests never reached Fleetd.java at all: both new tests declared
their OWN local copy of the forwarding shape instead of calling production's.
Mutating Fleetd.java's real forwarder back into the broken lambda left those
copies untouched, so the whole suite stayed green while production had
regressed to exactly the bug being fixed -- proven live by the reviewer.

Fix: extracted the forwarding shape into one named factory,
ExhaustionSink.forwardingTo(Supplier<ExhaustionSink> target), with the
"why a lambda here is wrong" explanation moved onto it (the one place the
shape is now written). Fleetd.java's forwarder collapses to one line:

    ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);

Both new tests now call this same factory instead of rebuilding an anonymous
class inline, so they exercise the identical object production builds:
- ExhaustionSinkForwardingHazardTest: calls ExhaustionSink.forwardingTo
  directly and asserts the hint reaches the real sink through it.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
  same factory call, inside the full Fleetd-shaped construction order
  (forwarder built first, real sink pointed at via the AtomicReference
  afterward), driven through the real SessionManager.acquire() path.

Mutation proof, this time on production code only: deleted the factory's
3-arg override (falls back to the interface default, dropping the hint) --
both new tests go red with no test file touched:

  ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
  OpenCodeLauncherTest...ForwardingHop:  expected: <true> but was: <false>
  Tests run: 68, Failures: 2

Restored, re-ran: green (Tests run: 68, Failures: 0). Confirmed Fleetd.java
carries no lambda ExhaustionSink anywhere (grep). ExhaustionSink.none() stays
a lambda on purpose -- both its overloads are true no-ops regardless of
arity, so there is no hint to drop.

mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:30:37 +07:00
Dai Ha ee932fd85b CB-201: nudge leads about backend outages
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Successful in 1m55s
2026-09-03 10:28:51 +07:00
Dai Ha fe2e5ede34 CB-201: retain backend failure outcome
CI / build (pull_request) Successful in 1m11s
CI / contract (pull_request) Successful in 1m23s
2026-09-03 10:26:54 +07:00
Dai Ha c325054242 fleetd #234 round 2: fix the ExhaustionSink forwarding hop Fleetd.java actually uses
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m18s
The round-1 fix was dead on the real production path. Fleetd.java:177 builds
a forwarding sink (needed because the adapters are constructed before
`sessions` exists, breaking a genuine cycle) as a LAMBDA:

    ExhaustionSink forwardingExhaustionSink =
            (target, reason) -> exhaustionSinkRef.get().onExhausted(target, reason);

A lambda can only implement the interface's one abstract method (the 2-arg
overload), so it silently inherited the 3-arg overload's default body, which
drops the profile hint and calls back into the 2-arg method. OpenCodeLauncher
is constructed with this forwarder, so the hint it supplies (its own
already-known profile name) was thrown away before it ever reached the real
sink built later in Fleetd.main -- reproducing the exact silent no-op round 1
was sent to fix. The 1127 tests from round 1 all injected a sink directly
into OpenCodeLauncher and never went through this forwarding hop, so none of
them could see it.

Fix: forwardingExhaustionSink is now an anonymous class overriding both
overloads, each delegating to whatever exhaustionSinkRef currently holds.

Audited every other ExhaustionSink value in main/: the only other one is
ExhaustionSink.none() (a lambda), which is safe regardless of arity since
both its 2-arg body and the inherited 3-arg default are true no-ops.

New tests:
- ExhaustionSinkForwardingHazardTest: isolates the hazard at the interface
  level (a lambda forwarder drops the hint; an anonymous-class forwarder
  does not), independent of Fleetd.java's specific wiring.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
  replicates Fleetd.java's actual construction order (forwarder built and
  handed to the launcher first, real sink built and pointed at via the
  AtomicReference afterward) and drives the quarantine through it via the
  real SessionManager.acquire() path.

Both proven by mutation: temporarily rewriting each fixed forwarder back
into the pre-fix lambda makes its test fail with a real assertion message
(both matched exactly: "expected: <gx> but was: <null>" for the interface
proof, "expected: <true> but was: <false>" for the composed-wiring test);
restoring makes it pass again. No reverts were committed.

mvn clean install: Tests run: 1130, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:22:31 +07:00
Dai Ha 7662e2d0c8 fleetd #201/#227: refine backend outage work
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m49s
2026-09-03 10:21:44 +07:00
Dai Ha 4877992a70 fleetd #234: key the opencode model check on the resolved session id, and make the spawn-time quarantine actually happen
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m29s
Defect 1: OpenCodeSessionDiscovery.actualModelForDirectory queried
WHERE directory = ?, the same heuristic sessionIdForDirectory uses. Since a
default fleet_spawn (no worktree:) shares the lead's cwd with every other
worker and every past session ever run there, the model read-back could
silently compare against a DIFFERENT session's row. Renamed to
actualModelForSessionId(sessionId), keyed on the primary key id instead, and
made OpenCodeLauncher's SessionAwareHandle cache the resolved id once
non-null (AtomicReference) so a later sibling row in the same directory can
never flip which session's evidence is read. sessionIdForDirectory (#209) is
left directory-based on purpose, with a comment explaining why the heuristic
is unavoidable at that layer.

Defect 2: the ERROR log claimed "quarantining this profile's credential" but
Fleetd's ExhaustionSink lambda resolved target -> roster -> profile ->
credential, while OpenCodeLauncher's model-mismatch check fires from
agentSessionId() during SessionManager.acquire(), before the session is
registered in the roster -- the lookup found nothing and silently no-opped.
Added a default 3-arg ExhaustionSink.onExhausted(target, reason, profile)
overload (defaults to the 2-arg method, so CompletionResolver's two call
sites are unchanged); OpenCodeLauncher now passes its own already-known
profile name; Fleetd's sink became an anonymous class that tries the roster
first, falls back to the hint, and logs loudly at ERROR naming target/reason
when neither resolves, instead of silently no-oping.

Both fixes proven by mutation: reverting each independently makes its new
test fail with a real assertion message, restoring makes it pass again.

mvn clean install: Tests run: 1127, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:09:26 +07:00
ltms 1c051c4e47 fleetd #115: a profile that needs no token no longer warns, and no longer crashes
CI / contract (push) Successful in 1m2s
CI / build (push) Successful in 1m50s
tokenEnv no longer defaults to FLEETD_WORKER_TOKEN. A profile that names no
tokenEnv is stating it needs none, which is different from one naming a variable
that turns out to be unset. requiredSecretEnvVars now warns only for an explicitly
declared tokenEnv, so the permanent false alarm about a token no profile needs is
gone and the real warnings beside it stay trustworthy.

Decision recorded in a code comment: an explicit tokenEnv is checked for every
kind, opencode included. opencode can use its own provider credentials, but an
explicit tokenEnv declares a required host secret for its configured provider.

Second commit fixes a crash the first one introduced. Making tokenEnv nullable
changed what every reader of that value can receive, and ClaudeCodeLauncher:261
passed it straight to env.apply — System::getenv in production, which throws on a
null name. The live profile local-direct (kind: claude-code, baseUrl set, no
tokenEnv) would have crashed on spawn. It is weight: 0 today, so the failure would
have surfaced whenever someone re-enabled it. Now uses the superclass helper
resolveEnv, which already tolerates a null name and which OpenCodeLauncher was
already using.

Reader audit, all four: ClaudeCodeLauncher fixed; OpenCodeLauncher already safe via
resolveEnv; ConfigRef uses Objects.equals; MemberEnvAllowList drops null and blank
names in addIfPresent.

Verified by the lead before merge: reverting the resolveEnv fix makes the new test
fail with a NullPointerException from the null variable name, and restoring it
passes. Independent build: 1125 tests, 0 failures, 0 errors, 0 skipped.

Not verified: the ticket's acceptance criterion 4, a real boot on this host showing
no FLEETD_WORKER_TOKEN line while the WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN lines
are unchanged. A worker cannot restart the daemon it talks through. The lead checks
that at the next redeploy.
2026-09-03 05:02:27 +02:00
Dai Ha adb7a67880 CB-612: support tokenless Claude Code profiles
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m24s
2026-09-03 09:59:21 +07:00
Dai Ha 723fe494e9 CB-612: suppress unneeded token warnings
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Successful in 1m17s
2026-09-03 09:50:49 +07:00
ltms 0f08b93659 fleetd #176: spawn gate fails fast on a dead backend, and refines UNKNOWN safely
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m47s
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.

Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.

Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".

Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.

Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
2026-09-03 04:20:23 +02:00
Dai Ha 432c1d92d1 fleetd #176 review: cover the namePrefix guard, the adapter-refinement corroboration that had no test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m14s
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.

Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.

mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:17:10 +07:00
Dai Ha 60b7e67b42 fleetd #176: fail fast when the spawn gate's backend exits mid-wait, and safely refine its UNKNOWN status
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m26s
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.

Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.

Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
 - spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
 - spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
 - refinedIdleIsNotAcceptedWhenAgentTypeIsNull
 - refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
 - refinementNeverFiresWhenRawStatusIsAlreadyInjectable

mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:10:49 +07:00
ltms 3fbd43fe3f fleetd #175: read back the model opencode actually resolved, and quarantine a silent substitution
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m29s
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
2026-09-02 13:01:33 +02:00
Dai Ha 32ebf065ac fleetd #175 review: a missing providerID in opencode's model JSON is UNKNOWN, not a mismatch
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m13s
parseModel already tolerates a model JSON with an id but no providerID (a real shape
opencode can write). checkModelMatch's providerMatches check did not: a provider-
prefixed profile whose id matched but whose evidence had no providerID was reported
as a mismatch and quarantined on incomplete data, which acceptance rule 4 forbids.

Compare the provider only when BOTH the profile requested one AND the evidence has
one. A genuine id mismatch is still caught either way — narrows the check, does not
disable it.
2026-09-02 17:59:04 +07:00
Dai Ha 1178b3f684 fleetd #175: check opencode's actual model against the profile, quarantine on a real mismatch
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m49s
opencode does not fail on an unknown -m <model> flag — it silently falls back to a
default model, which can be a paid credential. Extends the existing late-resolve
path (#209's SessionManager -> handle.agentSessionId() re-poll) so that once the
opencode session row exists, OpenCodeSessionDiscovery also reads its `model` JSON
column and OpenCodeLauncher's SessionAwareHandle compares it against the profile's
configured model.

Comparison rule: split the profile's model on the first '/' into provider+id. Compare
id always; compare provider only when the profile specified one. A bare model name
with no '/' matches on id alone. Absent/unparseable evidence is UNKNOWN, never a
mismatch, so a working profile is never quarantined on missing data. A real mismatch
logs an ERROR naming both models and the profile, then quarantines through the
existing ExhaustionSink path (wired via an AtomicReference forwarding sink in
Fleetd.java to break the sessions/workers/adapters construction cycle).
2026-09-02 17:53:59 +07:00
ltms 2a434ced2f fleetd #222: keep the member's charter file out of fleetd's own java.io.tmpdir
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m28s
Verified by the lead before merge: read the full production diff, confirmed the memberHerdrSocket-absent branch is the literal unmodified Files.createTempFile call in its own branch, and that the refusal names the missing key. Measured behaviour recorded: claude 2.1.258 exits 1 immediately on an unreadable --append-system-prompt-file, so the pre-fix bug was the loud readiness-gate failure, not a silent charter-less member. The /tmp full-suite failure the worker reported was the two other #225 copies, fixed by #230 which is already on main; main is built and checked after this merge.
2026-09-02 02:56:36 +02:00
ltms cc919aa2b6 fleetd #226: reserve an architect slot before launch, so no member runs on a charter it will not hold
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m31s
Verified by the lead before merge. Read the full production diff: locking is consistent (every MemberRegistry method uses synchronized(terminalToSlot)), and the refusal happens before the launcher starts a process. Proved the restored fallback test is real by removing the DEV fallback from SessionManager and re-running it: it failed with "expected: <DEV> but was: <ARCHITECT>", then reverted. Independent build: BUILD SUCCESS, 1092 tests, 0 failures, 0 skipped.
2026-09-02 02:54:52 +02:00
Dai Ha 6417b0edd9 fleetd#222: fix currentUserGroup() to resolve the real primary group (fleetd#225)
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m43s
The helper I copied from OpenCodeLauncherTest read the CWD's owning
group instead of the process's real primary group, so it silently
picked up whatever group owns the directory Maven was started from
(staff in a home checkout, wheel under /private/tmp on macOS) rather
than a group the operator is actually in. Replaced with the id -gn
based resolution that landed on #230 for the other two copies of this
helper, same shape and skip wording.
2026-09-02 07:54:11 +07:00
ltms 39c7ce76f3 fleetd #224 + #225: make worktreeRoot group-traversable, and stop two tests depending on the checkout location
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m41s
Verified by the lead before merge: read the full production diff, confirmed `group` is normalised to null at GitWorktrees:148 so the `group == null` guard is complete, and confirmed the refusal runs before `git worktree add` so a failure leaves no half-made worktree. Independent build in the worker's worktree: BUILD SUCCESS, 1093 tests, 0 failures, 0 skipped.
2026-09-02 02:52:12 +02:00
Dai Ha b40f477210 #226 retain architect fallback coverage
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m33s
2026-09-02 07:50:34 +07:00
Dai Ha 5c56cb347f fleetd #224 / #225: share worktreeRoot with the group, and fix the group-detection test helper
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m42s
#224: GitWorktrees#add created worktreeRoot with the daemon's umask and never shared it with
worktreeGroup, even though shareWithGroup shares every child underneath it (each worktree, and
the repo's common git dir). Under memberHerdrSocket: the member pane runs as a different OS
user, which needs execute on every ancestor directory to reach anything underneath, no matter
how carefully each child is shared — so a member could not read the opencode.json #219 places
under this root, could not reach its own worktree, and could not read #213's ZDOTDIR scrub when
placed here either.

Fix: add() now calls a new shareRootWithGroup(root) right after creating the root, chgrp+chmod
g+x on the root itself (non-recursive — each child is still shared individually by its own call
site). No-op when worktreeGroup is unset, so behaviour is byte-identical in today's only live
mode. On failure (group missing, or operator not a member of it) the spawn is refused with a
WorktreeException naming the root, its current mode, and the group — mirroring shareWithGroup's
existing refusal shape — before `git worktree add` ever runs, so no partial worktree is left
behind.

Also adds the assertion the #221 reviewer flagged as missing: a test driving
EnvAllowListScrub#shareWithGroup directly against a directory holding several flat files
(opencode.json, member-charter.md, ide-rules.md, plus an unrelated one) and asserting every one
of them gets group-readable/never-group-writable permissions, not just the two files someone
happened to think of.

#225: OpenCodeLauncherTest/HerdrPeerLauncherAllowListWiringTest's currentUserGroup() read the
group that owns the current working directory, not the process's own primary group, despite its
comment claiming the latter. Those coincide only by accident: a home checkout is typically owned
by a group the operator belongs to (staff), while a checkout under /private/tmp on macOS is
group wheel, which the operator is usually not a member of — so the same test fails for real
depending on where the repo happens to be checked out, and the existing assumeTrue only guarded
against "no POSIX groups at all", never "a resolvable but wrong group". Fixed by resolving the
process's REAL primary group via `id -gn` instead, with assumeTrue (skip, not fail) only when
that itself cannot be resolved on the host. The permission assertions these tests exist for are
unchanged.

Verified `mvn clean install` green from both a home checkout and a /private/tmp copy (mirroring
the exact repro in #225): 1093 tests, 0 failures, 0 errors in both locations.
2026-09-02 07:47:13 +07:00
Dai Ha e694deace3 #226 reserve architect slots before launch
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Failing after 1m33s
2026-09-02 07:44:41 +07:00
Dai Ha 748367b7d6 fleetd#222: put the claude-code charter file where the member can read it
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m16s
ClaudeCodeLauncher#writeCharterFile used Files.createTempFile with no
directory argument, which resolves against fleetd's own java.io.tmpdir
(macOS: the per-user $TMPDIR, mode 0700). Under memberHerdrSocket: the
member pane runs as a different OS user and cannot read that directory,
and since #220 the charter file is the ONLY delivery path for
--append-system-prompt-file. Following #219's refusal decision (a
charter is the member's turn contract, not a degradable control): with
memberHerdrSocket configured, the charter now goes into a fresh
per-spawn directory under worktreeRoot, shared read-only via
EnvAllowListScrub.shareWithGroup (reusing #213/#219's mechanism); a
missing worktreeRoot/worktreeGroup refuses the spawn by name instead of
writing an unreadable file. With memberHerdrSocket absent the path is
unchanged.
2026-09-02 07:44:34 +07:00
ltms dcf5fb3be3 Merge pull request 'CB-619 / fleetd #123: refuse an architect spawn with no matching slot' (#223) from worker/cb-123-role-demotion-c600f7-2 into main
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m55s
2026-09-01 10:41:47 +02:00
ltms f9d2ee2a2b Merge pull request 'fleetd#219: fix OpenCodeLauncher config/discovery roots under memberHerdrSocket' (#221) from worker/cb-219-opencode-roots-1f677e-1 into main
CI / contract (push) Successful in 45s
CI / build (push) Successful in 2m16s
2026-09-01 10:33:04 +02:00
Dai Ha 866c7f2e9a CB-619 / fleetd #123: refuse an architect spawn with no matching slot
CI / build (pull_request) Successful in 1m16s
CI / contract (pull_request) Successful in 1m18s
An explicit-profile spawn bypasses role-pool placement (CompositePeerLauncher
only constrains an UNQUALIFIED spawn to fleet.<role>), so it was the one path
that could ask for role=architect on a profile no architect slot carries.
MemberRegistry silently held the session as a plain worker while GET /members
still reported the requested "architect" and only fleet_whoami (which reads
live bindings, not the request) told the truth.

- MemberLifecycle.requireSlotFor(role, profile): refuses the acquire before
  anything spawns when no configured architect slot carries the profile,
  naming the role, the profile, and the pools that do carry it. No-op for
  dev/reviewer, which are placement candidates only, never a live identity
  binding — refusing a profile mismatch there would break the documented
  fleet_spawn{profile:"opus"} (role defaults to dev) flow.
- MemberLifecycle.acquired(...) now returns the role the session actually
  holds, so a residual race (a slot exists but every instance is already
  bound to a different terminal) still falls back to dev honestly instead of
  lying — this case logs at WARN (was INFO), naming profile and terminal.
- SessionManager now records the role acquired() returns on MemberSession,
  never the requested role, so GET /members and fleet_list can no longer
  report a role the member does not hold; no changes needed to memberView/
  rosterView, which just read session.role().

An architect's identity IS the slot it is bound to — binding a role with no
slot to bind means inventing an identity out of nothing, which is the quiet
failure the whole role system exists to prevent.

Tests: SessionManagerTest and FleetMcpTest each drive a real spawn through
FleetMcp.spawn -> SessionManager.acquire -> the real ClaudeCodeLauncher (via
FakeHerdr), then assert on GET /members and fleet_whoami for that same
session — not on MemberRegistry.bind directly (fleetd issue #113's mistake).
2026-09-01 15:29:39 +07:00
Dai Ha d9168de43e fleetd#219: OpenCodeLauncher config/discovery roots must not assume fleetd's own filesystem
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m48s
Site 1 (config root): under memberHerdrSocket, writeConfig() now places the
ephemeral opencode.json directory under worktreeRoot and shares it read-only
with worktreeGroup, reusing EnvAllowListScrub#shareWithGroup (widened to
package-private and generalized) — the same mechanism #213 built for the
ZDOTDIR scrub, rather than a second copy. Unlike the ZDOTDIR scrub's
degrade-to-overlay fallback, a missing worktreeRoot/worktreeGroup here
REFUSES the spawn (IllegalStateException from buildLaunch): this file is the
member's only way to learn where the bridge MCP is, so writing it somewhere
unreadable would just produce an undeliverable member with no signal
pointing at the cause. memberHerdrSocket absent stays byte-identical.

Site 2 (discovery root): under memberHerdrSocket, agentSessionId() now
declares session discovery unavailable and logs one WARN per launcher
instance instead of silently scanning fleetd's own $HOME (opencode.db lives
under the MEMBER's home under this config key). Decision + reasoning for why
this is a declare-unavailable rather than a new config key is in
defaultDiscoveryRoot()'s javadoc.

Widened HerdrPeerLauncher#memberHerdrSocketConfigured/memberScrubParentDir/
memberGroup to package-private so OpenCodeLauncher reuses the exact same
config resolution rather than re-deriving it.

Same-shape finding (not fixed, out of scope): ClaudeCodeLauncher#writeCharterFile
(line ~465) writes the role-charter temp file via Files.createTempFile with no
directory argument, i.e. under java.io.tmpdir — the same site-1 shape, unfixed
for the Claude Code adapter.
2026-09-01 14:38:48 +07:00
Dai Ha c3fa1136d4 #220: keep the launch command inside the pane's 1024-byte line
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m37s
herdr does not exec a member's launch command — it TYPES it into the pane,
and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Past that the
tail is dropped and NOTHING reports it: herdr answers "agent started", the
backend exits on the mangled argument it was handed, the pane closes, and the
only symptom is the readiness gate timing out 20 seconds later with no reason.

That is what broke every claude-code spawn after #214. The reply charter rode
inline on --append-system-prompt, so the command was already 978 bytes; adding
--session-id <uuid> made it 1028, and the 4 bytes cut off the end turned
--autocompact 250000 into --autocompact 25, which claude rejects. Measured on
the live pane, the cut is at byte 1024 exactly.

- ClaudeCodeLauncher: the charter ALWAYS travels as --append-system-prompt-file.
  The file path already existed for the two-charter case; the inline form only
  ever saved a temp file, and it cost ~800 bytes of the line budget. This takes
  the prose off the command line for good.
- HerdrPeerLauncher.checkPaneCommandFits: refuse a command that cannot fit,
  naming the byte count and the longest argument, instead of spawning something
  that cannot work. The estimate is deliberately conservative — fleetd cannot
  see herdr's quoting, and an under-estimate would let the silent truncation
  back in.
- HerdrPeerLauncher.waitUntilInjectableOrThrow: log the pane tail and the last
  herdr status BEFORE stop() closes the pane. Without it the gate reports only
  that it timed out, which is true of every cause. This is what found the bug,
  and it stays.

The guard also catches a case that was already over the limit: a profile with
ideMcpUrl set assembles 1084 bytes. It is now impossible to ship that silently.

3 tests, all watched failing first: with the inline charter restored the guard
fires in the new fit test, in the pre-existing autocompact test and in the IDE
mount test. Full suite 1081 tests green. Proven live: sonnet spawns again, the
member obeys the file-delivered charter and ends its turn with fleet_reply, and
fleet_list reports the #214 agentSessionId.
2026-09-01 14:11:28 +07:00
Dai Ha cabcd87b66 #211 follow-up: the raw-scrape fallback must carry the pane, not just the matched line
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m50s
The normal backend-error path appends the pane tail to the failure reason on
purpose (fleetd#164): the BACKEND_ERROR pattern is a heuristic, and a member
that reported *about* an error while forgetting fleet_reply matches it too, so
dropping the rest of the pane destroys the report.

The new raw-scrape fallback did not do that. It matters more there, not less:
the fallback only runs when the trimmed assistant block was empty, so the raw
scrape is the ONLY copy of whatever the member managed to say. A lead read the
matched line and nothing else.

Clipped to the same cap the normal path uses, since a raw screen has no
boundary trimming to bound its size.

The assertion was watched failing without the fix:
  AssertionFailedError: fleetd#164: the failure must carry the pane, not only
  the matched line ... expected: <true> but was: <false>

mvn clean install: Tests run: 1079, Failures: 0, Errors: 0, Skipped: 0
2026-09-01 13:26:05 +07:00
Dai Ha bf0e09b1a2 #214: mint a session id on every claude-code spawn, so every member is resumable 2026-09-01 13:23:32 +07:00
Dai Ha e1eb50ce65 #213: fix ZDOTDIR credential scrub gate + directory under memberHerdrSocket 2026-09-01 13:23:32 +07:00
Dai Ha 049e7d9d54 fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback 2026-09-01 13:23:32 +07:00
Dai Ha ff3b49cd1e #214: mint a session id on every claude-code spawn, so every member is resumable
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m52s
2026-09-01 11:27:10 +07:00
Dai Ha 0373b6c41b #213: fix ZDOTDIR credential scrub gate + directory under memberHerdrSocket
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m17s
The memberCredentials.policy: allow-list ZDOTDIR scrub decided zsh-vs-not
using fleetd's own process $SHELL and wrote the generated scrub dir into
fleetd's own java.io.tmpdir. Under memberHerdrSocket: (member panes run as
a different OS user than fleetd's own process) this silently protects
nothing: the wrong shell decides the gate, and the directory can be
unreachable to the member.

- New FleetConfig.memberLoginShell: the member OS user's login shell,
  only ever read when memberHerdrSocket: is configured; fleetd's own
  $SHELL keeps deciding everything when memberHerdrSocket: is absent
  (byte-identical to before).
- HerdrPeerLauncher.applyEnvironmentAllowListPolicy: memberHerdrSocket +
  memberLoginShell not configured/non-zsh falls back to the CB-596
  sentinel overlay (warn loudly, never refuse to spawn). memberHerdrSocket
  + zsh memberLoginShell generates the ZDOTDIR under the configured
  worktreeRoot instead of java.io.tmpdir, shared with the existing
  worktreeGroup (reused, not a new key).
- EnvAllowListScrub: new generate(parentDir, allowedNames, group) overload
  shares the generated directory via pure-Java POSIX group ownership
  (rwxr-x--- dir, rw-r----- files) — no external process spawn.
- fleetd.example.yaml documents memberLoginShell: and worktreeGroup:'s
  reuse for the scrub directory (the live fleetd.yaml is gitignored).

4 new tests in HerdrPeerLauncherAllowListWiringTest cover the acceptance
criteria; 3 of the 4 were watched failing against the pre-fix code.
2026-09-01 11:07:11 +07:00
Dai Ha 445a45f6e1 fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m17s
CompletionResolver.resolve() returned an empty-scrape failure before the
BACKEND_EXHAUSTED / BACKEND_ERROR classification ever ran, whenever
lastAssistantBlock() found no usable text — most commonly a pane with no ⏺
marker at all, whose boundary scan starts at the top of the raw screen and
breaks immediately on the first line of TUI chrome. Since BACKEND_EXHAUSTED
is the only caller of exhaustionSink, this meant an exhausted backend was
recorded as "produced nothing" instead of being quarantined.

Fix: run the same two classifications against the raw (untrimmed) scrape as
a fallback, only inside the empty-scrape failure branch. A pane that already
yields a usable assistant block never reaches this branch, so the existing
narrow match is unchanged. lastAssistantBlock stays the sole source of the
reply text; only classification ever consults the raw scrape.
2026-09-01 10:50:10 +07:00
Dai Ha 966c58a3b8 #137: don't hand one stranded reply to every open async ticket (#215)
CI / contract (push) Successful in 1m31s
CI / build (push) Successful in 1m41s
abandon() drained the target's stranded reply once and then reused that same
Reply for every open task it walked past. Two open tickets on one target
therefore both came back REPLIED with the same text — one of them a reply the
worker never gave for that delegation.

One worker answer can settle at most one delegation. It now goes to the oldest
open task (lowest createdNanos) and every other open task keeps the ordinary
WORKER_FAILED path. If the chosen task turns out to be already resolved by
another path, the drained reply is published back to the inbox instead of being
dropped.

reply()'s matching side gets the same rule: more than one candidate means the
reply goes to the inbox rather than to a guess.

Reachability, checked rather than assumed:
- matching.size() >= 2 alone is reachable and was a real bug before this change.
- More than one candidate in reply() is not reachable today — hasAsyncQuestion
  matches any task with a stamped turnId, and answer()'s
  clearAsyncQuestion(turnId, false) leaves that stamp until the resumed turn
  resolves. Kept as defensive code, documented, no test seam added.
- matching.size() >= 2 together with a live strand is not reachable either:
  send() and answer() are the only two lock holders and both open a Rendezvous
  waiter inside the lock, so "lock held" and "waiter open" are one fact, and an
  acceptance always clears the strand first.

I checked the last point by building it in a scratch worktree: it can be forced
by closing an accepted send's waiter directly through Rendezvous, and it does go
red against the pre-fix loop — but that breaks the lock-and-waiter invariant
from outside the class, so no such test is added. A note in abandon()'s javadoc
says so, to save the next reader the same round trip.

Worker's REPORT-cb137.md left out of main.

mvn clean install: Tests run: 1071, Failures: 0, Errors: 0, Skipped: 0
2026-08-31 23:23:07 +07:00
Dai Ha de026b8f8a #137 follow-up: document defect-1 and defect-2 reachability, drop unconstructible test
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m28s
Defect 1 (reply()'s askAnsweredAsyncTasks returning >1 candidate): confirmed
unreachable today. Documented why in three places — hasAsyncQuestion matches
any task with a stamped turnId (not just an open question), and answer()'s
clearAsyncQuestion(turnId, false) leaves that stamp in place until the
resumed turn's own future resolves — so a second task can never reach the
same eligible state while a first one holds it. Kept the defensive
inbox-fallback branch as defence in depth against that guarantee weakening,
per review instruction; no test seam added.

Defect 2 (abandon()'s broader `matching` filter applying a stranded reply to
more than one task): matching.size() >= 2 alone IS reachable (already
covered by abandonFailsEveryPendingAsyncTicketForTheReleasedTarget) and was
a real pre-fix bug (97f6c33's parent reused one drained reply for every
matching task). But hadStrandedReply == true together with matching.size()
>= 2, at the instant abandon() runs, is not constructible through the public
API: send() and answer() are the only two sites that ever hold a target's
session lock, and both open a Rendezvous waiter for that target as the first
thing they do while holding it — so "lock held" and "waiter open" are the
same fact throughout this class, and reply()'s fast path always resolves an
open waiter directly instead of stranding. A strand can only be created
while no task is accepted, and the moment the lock is next taken, that
acceptance clears the strand again before abandon() can observe both facts
together. Documented this in abandon()'s javadoc and removed the earlier
attempt at a deterministic test for the conjunction, whose apparent failure
was an invalid premise (the "accepted" task's own acceptance silently
cleared the strand it was meant to race against), not the fix being absent.
2026-08-31 23:17:50 +07:00
Dai Ha 97f6c33a45 #137: don't guess when a target has more than one open async task
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m17s
The CB-205-recovery fix in #205 assumed a target has at most one open
async task, with no guard. Fix three consequences:

- reply(): askAnsweredAsyncTask -> askAnsweredAsyncTasks (List). Exactly
  one candidate completes it (unchanged). Zero falls to the inbox
  (unchanged). More than one now ALSO falls to the inbox instead of
  picking an arbitrary ConcurrentHashMap iteration order, and logs a
  WARN naming the target and every candidate ticket.

- abandon(): a stranded reply now settles at most one matching task —
  the oldest by Task#createdNanos (a new field, the tiebreaker). Every
  other matching task keeps WORKER_FAILED, same as today.

- abandon(): if the chosen recovery task's complete() loses a race
  (another path resolved it first), the drained reply is republished
  to the inbox instead of being silently dropped.

Single-task behavior is unchanged; only the ambiguous case changes.
2026-08-31 22:46:23 +07:00
Dai Ha 735c837604 #209: resolve agentSessionId lazily via the retained PeerHandle
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m18s
SessionManager called handle.agentSessionId() once at spawn and froze it in the
immutable MemberSession. For opencode that value is always null - the session
row does not exist yet when the pane is created - so fleet_list never reported
an agentSessionId and fleet_spawn{resumeSessionId} was unusable for that
backend. The handle's javadoc said 'the caller re-calls later'; no caller did,
and SessionManager did not even retain the handle.

Retain the PeerHandle per pane and re-resolve while the stored id is still
null, sticky once found, CAS-swapped into the registry.

roster() stays non-resolving - it is the roster supplier for LeadHeartbeatLoop
and FleetHealthMonitor, and resolving there would open opencode's 841MB SQLite
database on every tick, for every unresolved member, forever. rosterResolved()
carries the resolve and is used only by fleet_list and the REST roster, the two
surfaces that report the id. get(paneId) and release() resolve too, both
caller-driven.

A test pins the split: the plain roster() must never call agentSessionId()
again.
2026-08-31 22:21:16 +07:00
Dai Ha 457dc0330d #185 stage 2: the credential gap detector must not report on the wrong environment
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m59s
Under memberHerdrSocket: member panes run as a different OS user, so
hostEnvNames (fleetd's own environment) no longer describes what a member
inherits. logCredentialGap now reports "unknown, not clean" in that mode
instead of its usual UNBLOCKED/blanked conclusions, once per launcher, naming
the config key and scoping its count to fleetd's own process.

Byte-identical when memberHerdrSocket: is absent, pinned by a test.
2026-08-31 22:17:40 +07:00
Dai Ha a1a9015217 #199: /members returns rows under "members" 2026-08-31 22:17:35 +07:00
Dai Ha 5a811a3695 #199: GET /members returns its rows under "members", with "workers" kept as an alias
The endpoint became /members in the CB-634 rename but the body key stayed
"workers", so a caller that read "members" saw an empty fleet and reported
no members at all.

Emit the canonical "members" key. Keep "workers" as a deprecated alias so an
existing REST consumer keeps working - the out-of-band path a lead falls back
to when its MCP mount drops reads this endpoint.

The test now pins both keys and asserts they carry the same rows, so the alias
cannot silently drift. Watched failing without the fix:
"GET /members must return its rows under \"members\" ==> expected: not <null>"
2026-08-31 22:17:22 +07:00
Dai Ha 1fdaa74eb3 fleetd #209 follow-up: keep roster() off the resolve path
CI / build (pull_request) Successful in 1m20s
CI / contract (pull_request) Successful in 1m21s
roster() is the supplier for LeadHeartbeatLoop and FleetHealthMonitor
(both timer-driven) and for placement/exhaustion checks and the
metrics scrape — none of which read agentSessionId. Resolving there
meant every tick could open a lazy-resolving adapter's (opencode's)
on-disk session database once per member whose id was still unknown,
with no bound: a member whose id never appears would pay that cost for
the life of the process.

roster() goes back to its pre-#209 behavior (no resolve, no I/O). A
new rosterResolved() carries the resolve logic, and is used only by
the two surfaces that actually report agentSessionId to a caller:
fleet_list (FleetMcp.listFleet) and the REST roster
(FleetApp.listMembers). fleet_whoami's roster().stream() at
FleetMcp.java:768 does not surface the field, so it stays on the plain
roster(). get(paneId) (fleet_status) and the release() resolve are
caller-driven, not timers, and are unchanged.

Retargeted the roster-facing tests from #209 at rosterResolved(), and
added plainRosterDoesNotResolveAgentSessionId, which pins the split by
asserting the handle's agentSessionId() is not called again by
roster().
2026-08-31 22:15:16 +07:00
Dai Ha 7d4a4339c2 Merge branch 'worker/cb-137-ask-ticket-e7760c-2' of https://git.ltms.dev/fleet/fleetd into worker/cb137-ambiguous-task-4df3d8-4 2026-08-31 22:13:37 +07:00
Dai Ha 847e8bd3fa #185 stage 2: stop the credential-gap detector reporting on the wrong environment
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m39s
When memberHerdrSocket: is configured, member panes run under a different OS
user than fleetd's own process, so hostEnvNames (fleetd's own environment)
no longer describes what a member pane inherits. logCredentialGap now checks
for that config key and, when set, logs a single WARN saying the gap is
UNKNOWN (not clean) and names the key, instead of printing the "inherits
them UNBLOCKED" / "the scrub blanks them" conclusions as fact. Behaviour is
byte-identical when memberHerdrSocket is absent (the default and only mode
this host runs).
2026-08-31 22:12:56 +07:00
Dai Ha f9fb387427 fleetd #209: re-poll agentSessionId against the retained PeerHandle
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m40s
SessionManager used to call PeerHandle.agentSessionId() exactly once at
spawn and freeze the answer into the immutable MemberSession. For
opencode that call always came back null, because opencode has not
written its on-disk session row yet when the pane is created, and no
caller ever re-asked the handle — it went out of scope at the end of
the spawn method. fleet_list therefore never reported agentSessionId
for an opencode member, and resumeSessionId was unusable for it.

Retain each spawn's PeerHandle in SessionManager, keyed by paneId, and
re-resolve a still-null agentSessionId against it from roster(), get(),
and release() (so a released member's detail also carries a
late-resolved id). Resolution is bounded: only sessions with a still-
null id do any work, a resolved id is never looked up again, and a
throwing handle degrades to "unresolved" rather than breaking the
caller. MemberSession gains a withAgentSessionId wither in the same
style as withState/withActivity.
2026-08-31 22:04:20 +07:00
Dai Ha a55079afbd #185: opt-in worktreeGroup, so a member running as another OS user can write its worktree
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m39s
Stage 3 of #185. A provisioned worktree and the repo's git store are made
group-writable when worktreeGroup names an OS group; absent, nothing changes.
The share pass runs after overlayParity, not inside add(), because
overlayParity copies more files in after add() returns.

This isolates credentials, not the repository: a member in the group can
still write the operator's git objects and refs.
2026-08-31 21:47:03 +07:00
Dai Ha e18ad4723b Merge remote-tracking branch 'refs/remotes/origin/cb206' 2026-08-31 21:47:03 +07:00
Dai Ha 6de8ac8972 #206: read opencode session ids from opencode.db, and pin the read-only open
opencode migrated its session store from a JSON file tree to SQLite in
January. OpenCodeSessionDiscovery still scanned the frozen tree, so it
returned null for every member: agentSessionId was never known and
resumeSessionId silently did nothing for every opencode profile, through
57 member spawns, with nothing logging that the search found nothing.
2026-08-31 21:46:32 +07:00
Dai Ha a052975420 #206: pin the read-only open with a test that actually fails without it
CI / build (pull_request) Successful in 1m18s
CI / contract (pull_request) Successful in 1m44s
setReadOnly(true) is the whole thing keeping fleetd out of the operator's
live 841MB opencode.db, and no test failed when it was removed.

The obvious test does not work. Making the database file unwritable and
checking the read still succeeds passes either way, because SQLite silently
downgrades a read-write open of an unwritable file to read-only. I wrote that
test, watched it pass with the flag removed, and threw it away.

What works: extract a package-private openReadOnly(), then ask that connection
to INSERT and require the refusal. Watched red with the flag removed, green
with it restored.

Also switches the test's INSERT helper to a PreparedStatement -- hand-escaped
SQL in a test is a pattern that gets copied into main code.
2026-08-31 21:46:07 +07:00
Dai Ha 6d82ca95a4 #185: skip absent git paths, and resolve the git dir instead of assuming .git
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m38s
Two defects in the stage-3 share pass, both of which would have failed EVERY
provisioning spawn once worktreeGroup was set, not only the two-user case.

.git/logs was handed to chgrp unguarded while packed-refs was guarded. It does
not exist with core.logAllRefUpdates=false, or before the first ref update, and
chgrp on a missing path exits non-zero -- surfacing as a WorktreeException that
blames a group which is in fact fine. Every path is now skipped when absent.

repoRoot + "/.git" was hardcoded. That is a FILE, not a directory, when the
checkout is itself a linked worktree -- the very thing this class creates for
every member. It now asks git: rev-parse --git-common-dir, resolved against
repoRoot because git answers relatively for an ordinary checkout.

Both new tests were watched failing with the fix removed before being kept.
2026-08-31 21:34:49 +07:00
Dai Ha 8067ee4ec4 fleetd #185: opt-in worktreeGroup config for group-shared worktrees
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Successful in 1m45s
Adds worktreeGroup (top-level FleetConfig key), Worktrees.shareWithGroup
(GitWorktrees impl: git config core.sharedRepository group + one-time
chgrp/chmod g+rwX/setgid fix-up over the worktree, .git/objects, refs,
logs, worktrees, and packed-refs when present), and wires SessionManager
to call it AFTER overlayParity so overlay files are covered too. Off by
default (byte-identical behaviour when unset). Documents the
credentials-not-repository caveat in the javadoc and example config.
2026-08-31 15:53:01 +07:00
Dai Ha 3743789e8d fleetd #206: read opencode session ids from opencode.db (SQLite), not the frozen JSON tree
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m44s
opencode migrated its session store to SQLite in January 2026; the JSON tree under
storage/session/<projectID>/ses_*.json stopped being written, so OpenCodeSessionDiscovery
returned null for every member forever, and fleet_spawn{resumeSessionId} was unreachable.

- Add org.xerial:sqlite-jdbc 3.53.4.0, opened read-only (SQLiteConfig.setReadOnly), so it
  never disturbs a live opencode process writing the WAL-mode database.
- Rewrite sessionIdForDirectory to run a parameterized SELECT ... WHERE directory = ?
  ORDER BY time_updated DESC LIMIT 1 against the session table. Still never throws: a
  missing database, a locked/corrupt one, or no matching row all return null.
- Log the silence that let this go unnoticed: WARN once per instance when opencode.db
  itself is missing (the layout moved again), DEBUG when it exists but no row matches
  yet (the normal interim answer right after a spawn).
- Replace the JSON-fixture tests with a synthetic-SQLite-db fixture; delete the tests
  that only proved the old JSON scan worked.
2026-08-31 15:52:06 +07:00
Dai Ha 388aba7632 #137: an answered turn's reply completes its own ticket, instead of a false failure
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m17s
A wait:false delegation whose worker used fleet_ask ended with fleet_poll{ticket}
reporting 'the worker session was released before it replied' -- naming a
worktree, a branch and a snapshot commit, so it read as lost work. The worker had
in fact replied in full.

The ticket guessed the ask rendezvous detached the turn. The real cause is
narrower: answer() (behind fleet_send{turnId}) waits only for the lead's own
bounded MCP call window. A resumed turn doing real work -- edits, a build, a
push, a PR -- routinely outlives it. On timeout answer()'s finally closed the
waiter, so the worker's later fleet_reply found none and fell to the session
inbox, leaving the ticket's future unresolved until fleet_stop forced it FAILED.

reply() now looks for the async task parked on this exact answered turn and
completes it with the real reply. That is safe against completing the wrong
ticket: answer() calls clearAsyncQuestion(turnId, false), so the task keeps its
turnId and stays in asyncTasksByTurn, and hasAsyncQuestion therefore still makes
send() return BUSY for a second async send to that target. At most one candidate
task can exist per target.

abandon() keeps an independent check: if a reply was stranded, a released session
reports REPLIED with that text rather than a failure -- so the recovery hint that
implies lost work never prints once a reply exists.

Verified before merging: build green unpiped, and both new tests drive the full
delegation path (send async, ask, answer, reply, poll the ticket) rather than
handing a Reply to a sink, which is the trap this ticket called out.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 14:24:54 +07:00
Dai Ha ad587eafa3 #172: keep the broker URI, password and all, out of every member pane
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m15s
broker.uriEnv names an environment variable holding amqp://user:password@host,
and it was reaching every member. Its name is not credential-shaped -- no TOKEN,
KEY or SECRET in it -- so every name-pattern heuristic missed it, and it sat on
neither credential list.

fleetd already knows the name: the operator wrote it in broker.uriEnv. So derive
the exclusion from the config rather than hoping an operator also remembers to
deny it. coordinator.uriEnv has the same shape and is excluded too; on this host
both resolve to the same variable.

Excluded even when the operator lists the name under memberCredentials.allow:,
following the SSH_AUTH_SOCK precedent. There is no override, because a member has
no legitimate use for the broker password.

Reviewer finding, recorded rather than overstated: this is only a hard guarantee
under policy: allow-list, where the ZDOTDIR scrub runs after the pane's shell has
sourced the operator's chain. Under deny-list the name is removed from the
pre-shell env only, and a login shell re-exports it. That is deny-list's existing
weakness rather than a regression here, but the javadoc now says so plainly
instead of implying a guarantee that path cannot give.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 14:21:48 +07:00
Dai Ha 08ce9aef11 #161: resolve a pane by process ancestry, closing a worker->primary escalation
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m52s
PaneLocator's javadoc always claimed it found 'the agent pane whose process
tree contains' a pid. It did not: paneOwnsPid matched only the pane's shell_pid
and its foreground_processes. A process a member spawned -- python3, curl, any
helper opening its own connection to 127.0.0.1:8765 -- matched no pane, so
CallerResolver fell through to loopback-trust and resolved it as the PRIMARY.
A member escalated to lead by shelling out.

terminalForPid now builds the caller's ancestor set once (bounded at 32
generations, with a cycle guard) and matches any ancestor against a pane's pids.
The set is reused across both herdr clients on the CB-185 two-daemon path.

This only ever ADDS matches, which is the safe direction: the failure mode of
the fix is a member correctly restricted, while the failure mode of the bug is a
member acting as the lead. The no-match case still returns null, so the lead --
which maps to a pane named by leaders: -- still resolves as primary.

Ancestry is walked through a new ParentResolver seam so the tests drive it from
a fake pid->parent map rather than spawning real processes.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 14:15:39 +07:00
Dai Ha ee5f8b932b #137: complete an async ticket's own reply after answer() times out
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m44s
fleet_send{turnId} (MessageService.answer) blocks the primary only for its
own bounded MCP-call window (25s default, 120s max) — far shorter than a
worker's resumed turn can genuinely take. When that window expires, answer()
closes its rendezvous waiter, so the worker's eventual fleet_reply has no
live waiter to resolve and falls back to the session inbox. The async
ticket's future was never completed by that path, so fleet_poll{ticket}
stayed PENDING until fleet_stop's abandon() forced it FAILED with a
misleading "the worker session was released before it replied" reason,
even though the reply had genuinely arrived.

- MessageService.reply(): before falling to the inbox, look for the async
  task this exact turn belongs to (already answered — question cleared,
  turnId still stamped — but not yet resolved) and complete it directly
  with the real reply, so fleet_poll{ticket} returns it.
- MessageService.abandon(): defense in depth, independent of the above —
  never write a false WORKER_FAILED once a reply reached the inbox for
  this target; recover and use its real content instead.
- Two new tests drive the full delegation path (async send -> ask ->
  answer with a short timeout -> reply -> poll/abandon), not a reply sink
  directly; both fail with the fix disabled and pass with it restored.
2026-08-31 14:12:23 +07:00
Dai Ha 4ac688b6d9 #164: classify a backend-error scrape as WORKER_FAILED, carrying the whole pane
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m22s
main already shipped the core of #164 in 3bfa828: the MIN_TURN_NANOS floor and
the hard fail on an empty or unreadable scrape. This adds the one case that was
still resolving as a success -- a scrape that reads cleanly but whose content is
the backend's own rejection (e.g. "API Error: 400 invalid request body").

The BACKEND_ERROR pattern is deliberately narrow. A growing list of ad-hoc error
strings rots as backends change their wording, and broader backend-error
surfacing is #164 point 3.

Because the pattern is a heuristic, it also matches a member that forgot
fleet_reply while reporting *about* a backend error. So the failure reason
carries the whole pane tail, not just the matched line: a genuine backend error
reads as before, and a false positive keeps its report instead of losing it.

Checked before merging: 3bfa828 is an ancestor of main; the branch was current
with main; mvn clean install green unpiped (1037 tests, 0 [ERROR] lines); and
each of the 5 new tests fails with the fix commented out.

Co-authored-by: fleetd worker <worker@ltms.dev>
2026-08-31 10:56:40 +07:00
ltms a814d1ef00 #197: measure the async ticket TTL from completion, not from creation
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m17s
Lead-authored and lead-verified: full mvn clean install green at 1032 tests, and the new test proved by restoring the old comparison and watching it fail.
2026-08-31 05:10:46 +02:00
Dai Ha ea98856130 #197: measure the async ticket TTL from completion, not from creation
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m32s
pruneTerminalTickets compared the cutoff against createdNanos, so the real
window to collect a reply was "TTL minus however long the task ran". A
delegation that ran longer than the 10-minute TTL was already past the cutoff
the moment it finished, so the next prune destroyed its reply.

That is the normal case here, not an edge case. Real delegated work runs well
past ten minutes. Three workers in one session did, and two of their complete
reports were lost. The reply lives only in Task.future, so pruning it discards
the worker's whole report, and fleet_poll{target} returns [] rather than
holding it — there is no fallback.

Task now stamps completedNanos from a whenComplete hook registered in its
constructor, so every completion path stamps it (a reply, the completion
fallback, a timeout, a failure, an abandon on teardown) without each one having
to remember to. The stamp is a boxed Long, not a long with a sentinel:
System.nanoTime may return any value, so no number can mean "not stamped yet".
A task that is done but not yet stamped is left for the next sweep.

createdNanos had no other reader and is removed.

The TTL still bounds tasks — an uncollected finished ticket is still evicted
once the TTL passes since it finished. Both halves are pinned by a test, and
the first one was proved by restoring the old comparison and watching it fail.

Tests run: 1032, Failures: 0, Errors: 0, Skipped: 0
2026-08-31 10:10:16 +07:00
ltms a49e96835a CB-189: cover every remote, both URLs, and any non-SSH scheme in the credential check
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m11s
Lead-verified: merged onto current main (which already carries #194 and #196), full `mvn clean install` green at 1030 tests. Confirmed no plain `exec` call that reads a remote URL remains.

Review round 2 closed the half-fix: execRedacted was added but applied only to the new calls, leaving the three pre-existing URL readers (lines 143, 171, 334) still copying stdout into exception messages — the exact hole CB-189 named.

Kept the check before the origin strip, on the worker's reasoning: the credential really is in the config at that moment, so WARN-found followed by INFO-fixed is the full audit trail, whereas moving it after would silence the origin case entirely.
2026-08-31 04:32:14 +02:00
ltms 63c19dcba7 CB-185: fix two blockers to switching on memberHerdrSocket
CI / build (push) Successful in 1m31s
CI / contract (push) Successful in 1m19s
Lead-verified: merged with #194 onto an integration branch off main, full `mvn clean install` green at 1025 tests.

Review round 2 fixed the stale ambiguous-pane message and, more importantly, a real abort: probeOwner called list() unguarded, so one unreachable daemon made panes on a different healthy daemon un-stoppable too — the same bug blocker 1 exists to fix, through a new door. Worker proved it by removing the guard and quoting the failure.
2026-08-31 04:30:57 +02:00
ltms 2823349c8e CB-192: fix false credential-gap WARN under allow-list+zsh, split its log guard
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m11s
Lead-verified: merged with #196 onto an integration branch off main, full `mvn clean install` green at 1025 tests. Deny-by-default WARN confirmed byte-identical to main.

Review round 2 fixed a defect I found in round 1: the new INFO asserted that gap names were not on the derived allow-list without ever checking, so a name derived from a profile's tokenEnv/gitTokenEnv/env: would be reported as safe while the member actually inherited it. Now split with MemberEnvAllowList.keeps — the same predicate the generated scrub evaluates.
2026-08-31 04:30:48 +02:00
Dai Ha 6fc301d62c CB-189 review fix: redact the three pre-existing URL-reading exec calls
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m42s
Review found that execRedacted was applied only to the new remote-enumeration code
and left three pre-existing calls reading remote.origin.url through the plain,
unredacted exec: removeUserInfoFromHttpsOrigin, requireCredentialFreeHttpsOrigin, and
configureHttpsUrlRewriteForSshOrigin. A non-zero exit or timeout on any of those could
still have copied the credentialed URL into a WorktreeException message. Switches all
three to execRedacted; the set-url write in removeUserInfoFromHttpsOrigin is left on
plain exec with a comment explaining why (it writes the already-stripped URL, not a
read).

Widens the shared exec(Map, boolean, String...) overload to package-private, the same
test-seam pattern already used by the afterWorktreeAdded constructor parameter, and
adds a test that drives it directly with a synthetic failing command whose stdout
carries a marker (passed via env, not argv, so the always-printed command line can't
carry it) and asserts the marker never reaches the exception message.
2026-08-31 09:29:45 +07:00
Dai Ha bc99d64786 CB-185: fix ambiguous-pane message and unreachable-daemon abort in probeOwner (review)
CI / build (pull_request) Successful in 1m6s
CI / contract (pull_request) Successful in 1m5s
Lead review of PR #196 found two issues in CompositePeerLauncher.probeOwner:

1. The "more than one daemon claims this pane" throw kept the old
   pre-fix message ("no owning herdr daemon was recorded"), which was
   only true of the code it replaced. Reworded to say what actually
   happened: N configured herdr daemons report this pane, so it is
   genuinely ambiguous. Updated the one test pinning the old string.

2. probeOwner let list() propagate straight out of the probe loop, so
   one unreachable daemon aborted the whole probe and made a pane on a
   DIFFERENT, healthy daemon un-stoppable too — resurrecting the exact
   bug blocker 1 fixes. Now catches HerdrException per daemon, logs the
   exception class only, and treats that daemon as not knowing the pane
   so probing continues. New test proves this: verified it fails with
   the try/catch removed (HerdrException propagates and the stop that
   should succeed via the healthy daemon throws instead), then restored.

Full mvn clean install: 1019 tests, 0 failures, 0 errors.
2026-08-31 09:28:29 +07:00
Dai Ha 615af4ed0a CB-192 review fix: split the allow-list gap by what the scrub actually keeps
CI / build (pull_request) Successful in 1m5s
CI / contract (pull_request) Successful in 1m20s
Lead review on PR #194 found that the allow-list INFO wording claimed the
whole gap ("credential-shaped names on neither known: nor allow:") is blanked
by the scrub, without checking that against effectiveAllowed. effectiveAllowed
is a SUPERSET of known+allow — MemberEnvAllowList.derive also unions in every
profile's gitTokenEnv/gitHostEnv/tokenEnv/env: keys, and derivedAllowedNames
further unions in the spawn's own env keys — so a gap name can still be kept
by the derived list (e.g. a profile's tokenEnv names it) and reach the member
unblocked while the INFO said "no member pane keeps them". That inversion is
exactly what #192 exists to remove.

logCredentialGap now splits the gap with MemberEnvAllowList.keeps (the same
predicate the generated scrub itself evaluates, so this cannot drift from
what the scrub does): names it keeps get a WARN, guarded by the same
unprotectedGapLogged flag as the deny-by-default case (same severity — a name
reaching a member unprotected is equally serious either way); names it
blanks keep the existing INFO, guarded by allowListGapLogged. The
deny-by-default WARN text and the non-zsh fallback are untouched.

Added allowListWarnsWhenTheDerivedAllowListKeepsAnUncoveredName and
allowListSplitsAMixedGapBetweenTheWarnAndTheInfo to ClaudeCodeLauncherTest.
2026-08-31 09:28:10 +07:00
Dai Ha 045d229728 CB-185: fix two blockers to switching on memberHerdrSocket (#185)
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m39s
1. CompositePeerLauncher.stop() was permanently un-stoppable for any
   member that survived a daemon restart, because spawnedBy is in-memory
   only. On a cache miss with more than one configured herdr daemon, probe
   each distinct daemon's agent.list() for the pane instead of refusing
   outright: exactly one owner routes and caches; zero owners is treated
   as already-stopped (a no-op, matching the tolerance HerdrPeerLauncher
   already gives an already-gone pane); more than one owner is the
   genuine per-daemon-pane-id ambiguity and still throws.

2. FleetApp#healthz always reported the LEAD daemon's herdr version/
   protocol even when a second (member) daemon was configured, so a
   member-daemon protocol mismatch was invisible behind a green
   /healthz while every spawn silently failed. Added a separate "member"
   key alongside the unchanged "herdr" key, and a "protocolMismatch"
   flag when the two differ. Verified scripts/redeploy-fleetd.sh and
   scripts/rename-checkout.sh only check the HTTP status code and print
   the body verbatim — neither parses a specific field — so adding a key
   is safe.

Both fixes are covered by tests written to fail without the fix
(verified by reverting each fix and watching the new tests fail, then
restoring). Full `mvn clean install`: 1018 tests, 0 failures, 0 errors.
2026-08-31 09:20:14 +07:00
Dai Ha d1fd5700f5 CB-189: cover every remote, both URLs, and any non-SSH scheme in the credential check
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m7s
GitWorktrees only ever inspected origin's HTTPS fetch URL for embedded credentials. A
credential on any other remote, on a pushurl, or on a plain http:// URL passed through
unreported. Adds an additive, reporting-only check that enumerates every remote and both
its fetch and push URLs, flagging non-empty user-info on any non-SSH-family scheme.

The existing origin/https strip-and-refuse behaviour is untouched. The new check is
wrapped so it can never abort a provision, and on failure logs only the exception's
class, never its message, since the enumerating `git remote` call is not redacted.
Also adds execRedacted, an exec variant that never copies captured stdout into a
WorktreeException message, for commands whose stdout may itself be a credentialed URL.
2026-08-31 09:17:54 +07:00
Dai Ha 2d55b0b9a5 #168: correct the audit's MISSING claim — the feature is documented, its index row was not
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m42s
The memberHerdrSocket section exists at 11-Features.md:2174; what was absent was
its row in the index table. My omission when I added the section. Wiki fixed at
b24965c.
2026-08-31 09:17:41 +07:00
Dai Ha d89ae94a2e CB-192: fix false credential-gap WARN under allow-list+zsh, split its log guard
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 1m7s
logCredentialGap(creds) always emitted the WARN wording ("every member pane
inherits them UNBLOCKED"), even under memberCredentials.policy: allow-list on
a zsh login shell, where the generated ZDOTDIR scrub genuinely blanks the
name. The line reported the control working as though it were a hole.

Pass an effectiveAllowed set instead: null keeps the WARN (deny-by-default,
and the allow-list non-zsh fallback, where nothing is ever scrubbed); the
derived allow-list set (only reachable after applyEnvironmentAllowListPolicy's
own zsh gate) selects a new INFO wording that says the scrub will blank the
name instead of claiming it is inherited unblocked.

Also split the single credentialGapLogged AtomicBoolean into two guards
(unprotectedGapLogged / allowListGapLogged) — one per report kind. Since
memberCredentials is a live, re-read-per-spawn supplier, a shared flag let a
harmless allow-list INFO on one spawn permanently suppress a later spawn's
real deny-by-default WARN after a policy reload.

Fixes gitea #192.
2026-08-31 09:17:20 +07:00
ltms e6193c4098 Merge pull request '#168: audit current wiki snapshot' (#193) from worker/cb-168-wiki-audit-3ef3d1-3 into main
CI / build (push) Successful in 1m6s
CI / contract (push) Successful in 1m5s
2026-08-31 04:17:07 +02:00
Dai Ha b66f0677ed #168: audit current wiki snapshot
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m37s
2026-08-31 09:14:22 +07:00
ltms 23ada1981e CB-185: route members to a separate herdr daemon (#186)
CI / build (push) Successful in 1m6s
CI / contract (push) Successful in 10m30s
2026-08-29 01:24:31 +02:00
ltms a22480c117 CB-185: route PaneLocator, StatusRefiner and FleetApp to the right herdr daemon (#188)
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m7s
2026-08-29 01:24:25 +02:00
Dai Ha 24f404f989 CB-185: fix three connection-identity/status/health gaps a second herdr daemon exposes
memberHerdrSocket splits lead operations from member operations onto two herdr
daemons. Three seams still assumed one shared daemon and broke silently when the
two clients differ (all three collapse to today's behaviour when they are the
same object):

1. ConnectionIdentity's PaneLocator was pinned to the member daemon only, so a
   lead's own MCP connection (which lives on the LEAD daemon) resolved to
   terminal == null, breaking fleet_reply/fleet_ask/fleet_whoami for a lead.
   PaneLocator now searches the lead client first, then the member client.

2. StatusPoller's StatusRefiner was pinned to the member daemon, so refining an
   UNKNOWN status for a lead target read the wrong daemon's pane content and
   never left UNKNOWN, wedging status-gated delivery to that lead forever.
   StatusRefiner gained a refine(target, raw, control) overload and the poller
   now refines through the same AgentControl the raw status was sampled from.

3. FleetApp was constructed with the raw lead-only herdr client, so /healthz
   stayed green while the member daemon was down (every spawn then fails
   invisibly) and GET /sessions silently dropped every member workspace.
   FleetApp now takes both clients: healthz requires both to answer, sessions
   merges workspaces from both.

Each fix has a test proven to fail without it (verified by reverting the
production change and re-running): FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest assert the actual Fleetd.java wiring (the same
technique as FleetdHerdrControlConstructionTest); StatusPollerRoutingTest and
the new PaneLocatorTest/FleetAppTwoDaemonTest cases exercise the real
production classes end to end rather than a hand-built object graph.
2026-08-29 06:20:32 +07:00
ltms a237fbff9d #185: refuse an unowned paneId when more than one herdr daemon could own it (#187)
CI / build (push) Successful in 1m4s
CI / contract (push) Successful in 1m17s
2026-08-29 01:10:21 +02:00
Dai Ha 31d5516991 #185 review: keep the stop owner on failure, key list() by daemon
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m30s
Three fixes on top of the pane-id PR, from my own read and the reviewer's:

- stop() removed the spawnedBy record BEFORE the delegate accepted the stop. A
  delegate that threw left the pane alive with its owner forgotten, so the retry
  fell into the ambiguous branch and refused the id for good. Remove after.
- list() deduplicated on the raw pane id. Pane ids are per-daemon counters, so
  two daemons can each hold w1:p1 on different panes, and one of the two real
  agents was silently dropped from fleet_list and every view built on it. The
  key is now (owning daemon, pane id). Delegates sharing one daemon still
  collapse, which is what the dedupe was for.
- The class javadoc still stated the single-herdr-connection premise as fact,
  next to the bullet this PR had just corrected for stop(). Fixed there too.

Also drops a redundantly qualified java.util.Collections.

Tests: 993 run, 0 failures, BUILD SUCCESS.
2026-08-29 06:09:43 +07:00
Ha Trong Dai 17af61e8dd CB-185: route message status by target
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m40s
2026-08-28 09:45:39 +07:00
Ha Trong Dai 6af87b6ad6 CB-185: share routed herdr controls
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m40s
2026-08-28 09:42:45 +07:00
Ha Trong Dai 5ba05d0bdb #185: count herdr owners for stop fallback
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m8s
2026-08-28 09:42:24 +07:00
Ha Trong Dai fc655e78c2 CB-185: route members to separate herdr
CI / contract (pull_request) Successful in 37s
CI / build (pull_request) Successful in 1m28s
2026-08-28 09:37:44 +07:00
Ha Trong Dai 25726a5ae7 #185: reject ambiguous unowned pane ids 2026-08-28 09:36:39 +07:00
Dai Ha 11c3ff67b6 fleets-status: fleet01 IS ssh-reachable; correct the 'denied' claim
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m12s
The skill said SSH to fleet01 is denied, so every report wrote 'not
reachable' for that fleet's daemon facts. That is true only for the user
dai.ha. The host alias fleet01 maps to user ltms and key auth works.

Checked 2026-08-28 while measuring #185: ssh fleet01 connects, and ltms
has passwordless sudo there. So fleet01's PID, uptime, jar and /healthz
can be reported over SSH even though its REST port is unreachable.
2026-08-28 09:06:54 +07:00
Dai Ha d867c87100 #184: correct the false ssh-agent premise in the URL-rewrite javadoc
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m37s
The javadoc said a member cannot authenticate at all once memberCredentials
blocks SSH_AUTH_SOCK, "there is no private key file on this host, only an
ssh-agent socket". That is wrong, and it was written after looking only in
~/.ssh, which holds nothing but Include lines.

Measured: ssh -G git.ltms.dev resolves an IdentityFile under the shared-env
directory. That file exists, is readable by this user, and has no passphrase.
A live member with SSH_AUTH_SOCK blanked pushed to the forge over SSH.

The rewrite itself is unchanged and still worth having. Only its stated reason
was wrong: it routes a member through its own scoped token instead of the
operator's ssh identity, which is what makes a member's pushes attributable
and revocable. It is not what stands between a member and the forge.
2026-08-28 06:35:46 +07:00
Dai Ha b0c4cedfab #157: redact remote-URL user-info before it reaches a log
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m25s
The four log lines added with the worktree HTTPS rewrite echoed the origin
URL verbatim, and one of them echoed the ssh:// authority, which carries
user-info. An ssh authority is normally just git@, so in practice this
changes nothing -- but a remote URL is not obviously a credential channel,
and that is precisely why one has leaked here three times (#157, #182).

Redact at the log call, not after it surprises someone.
2026-08-28 06:24:08 +07:00
ltms 85417d5215 #157: rewrite an SSH origin to HTTPS inside the provisioned worktree only
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m37s
Git never consults a credential.helper for an SSH transport, so #177's helper was inert on this repo — whose origin is ssh://. Once allow-list policy blocks SSH_AUTH_SOCK, a member on an SSH origin cannot authenticate at all: there is no private key file on this host, only an agent socket.

A worktree-scoped `url.<https>.insteadOf <ssh>` gives the member HTTPS for fetch and push while the primary checkout keeps SSH untouched. Host and port are parsed from the origin, never hardcoded — a test with a synthetic host proves it. The scp-like shorthand is left alone deliberately, since its host:path split is defined by ssh_config aliases rather than URI syntax.

Verified by the lead in an independent worktree: Tests run: 986, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:22:57 +02:00
Dai Ha 4accc746bd #157: rewrite SSH origin to HTTPS in the worktree so the credential helper is reachable
CI / build (pull_request) Successful in 1m10s
CI / contract (pull_request) Successful in 1m16s
2026-08-28 06:20:10 +07:00
ltms 21c4c8cbef #157: keep the forge token out of git config, via an environment credential helper
CI / build (push) Successful in 1m4s
CI / contract (push) Successful in 47s
The member credential scrub removes environment variables. It cannot remove a token written into git config inside the repo the member works in, so `git remote -v` handed a member a credential it was deliberately not given.

Provisioning now strips HTTPS user info from the origin before `git worktree add`, refuses the worktree if user info survives, and configures a per-worktree credential helper that reads WORKER_GITEA_TOKEN at call time. Nothing is persisted.

The helper emits BOTH username and password, and resets the inherited helper list first. An earlier revision emitted only `username=`, which made git fall through to the next helper — on a Mac that is osxkeychain, so a member would have authenticated with the operator's stored credential while every test passed and `git remote -v` looked clean. See #182.

`worktreeCredentialHelperCompletesWithoutUsingAnInheritedHelper` plants a synthetic operator helper in an isolated global config and proves the worktree helper wins. The worker confirmed it fails when the reset is removed. All credential tests pin GIT_CONFIG_GLOBAL and GIT_CONFIG_SYSTEM so they can neither read nor write real credentials.

Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:10:38 +02:00
ltms 430f5b0dae #164: never resolve a send with an empty scrape or a sub-floor turn
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m27s
A turn that dies on a backend error produces the same working -> idle transition as a real one, just faster and with nothing on screen. The resolver accepted that as a completed turn and handed the caller HTTP 200 with an empty reply, so a lost turn and a successful empty answer were indistinguishable.

Now: an empty or unreadable scrape fails, naming the member; and a BUSY -> DONE inside MIN_TURN_NANOS (2s) fails as a crash signature.

One existing test encoded the bug — it asserted a failed scrape resolved as a success carrying "" — and has been inverted rather than worked around.

Verified by the lead in an independent worktree: Tests run: 973, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:10:27 +02:00
ltms 42731833d0 CB-633: union memberCredentials.allow into the member env allow-list
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m38s
`policy: allow-list` silently ignored every name an operator wrote under `allow:` unless a profile happened to carry it too, so turning the policy on would have blanked credentials working members depend on. Derivation now unions the operator's list.

`SSH_AUTH_SOCK` stays governed only by `sshAuthSock`, even when listed under `allow:` — it is a live handle to the operator's ssh-agent, not a value.

Adds one INFO line per allow-list spawn, `member credentials: allowed N of M`, emitted only after the shell gate so it can never report coverage on a path where the scrub does not run.

Verified by the lead in an independent worktree: Tests run: 976, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:10:18 +02:00
ltms 65acf066ad #154: pin the AMQP reply inbox prefetch bound
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m39s
No behaviour change. #154 supposed that ownership drains a whole queue into the in-memory `held` map, making `x-max-length` and per-message TTL decorative. Measurement says otherwise: `basicConsume` is manual-ack, `deliverCallback` acks only duplicates, and `basicQos` is set on the one shared channel before any consumer starts — so total `held` is bounded by the prefetch window across all targets.

Adds a fake-broker test that drives the real `own()` path and fails if receipt ever starts acking, plus a javadoc line naming the prefetch window at the point of first mention.

Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 01:06:43 +02:00
Dai Ha ee8f570fd7 #157: isolate worktree credential helpers
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 2m46s
2026-08-28 06:06:12 +07:00
Dai Ha fa3f910d44 #154: pin AMQP reply inbox prefetch
CI / build (pull_request) Successful in 2m16s
CI / contract (pull_request) Successful in 2m18s
2026-08-28 06:03:54 +07:00
Dai Ha 46ac6e4e38 #157: use environment git credential helper
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 1m37s
2026-08-28 06:02:09 +07:00
Dai Ha 3bfa82839b fleetd#164: an empty or suspiciously fast scrape must fail, never resolve as a success
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 1m20s
CompletionResolver.resolve() used to hand the caller a successful "" reply whenever a
turn's scrape came back empty (whether the read failed, or genuinely produced nothing),
making a lost turn indistinguishable from a real empty answer. It also had no way to
tell a crashed backend's near-instant BUSY -> DONE transition apart from a genuine
completion.

Add MIN_TURN_NANOS (2s), a named floor below which a completed turn is treated as a
crash signature and failed rather than resolved as a reply. Fail on any empty scrape
(read failure or a clean-but-empty read) instead of resolving with "". Both failures
name the member and carry whatever is on the pane for context.

Thread an injectable LongSupplier clock through CompletionResolver (matching the
SessionManager/MessageService nowNanos pattern) so the floor is testable without a
real sleep.
2026-08-28 06:01:00 +07:00
Dai Ha 65ccf2e4ad CB-633 follow-up: only log allowed N of M when the scrub actually runs
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m44s
The coverage line was logged before the zsh gate, so a non-zsh
spawn (where nothing is scrubbed — overlayBlockedCredentials is the
fallback instead) printed 'allowed N of M' as if the derived
allow-list scrub had run. Move the log after the gate so it only
fires on the path that actually generates the ZDOTDIR scrub; the
non-zsh fallback keeps logCredentialGap's WARN as its only signal.

Added a test proving no 'allowed N of M' line is emitted on the
non-zsh fallback, through the real HerdrPeerLauncher#spawn path.
2026-08-28 06:00:38 +07:00
ltms 7a3b27f76f #150: report lead readiness from the delivery gate
CI / contract (push) Successful in 1m6s
CI / build (push) Successful in 1m7s
Share one deliverability predicate between Injector and FleetApp, so the status endpoint reports the same answer the injector acts on instead of re-deriving it from one of that predicate's two inputs.

Verified by the lead in an independent worktree: Tests run: 971, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-08-28 00:59:44 +02:00
Dai Ha 82e7be564c CB-633 follow-up: union memberCredentials.allow into the derived env allow-list
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m16s
MemberEnvAllowList.derive only ever looked at profile fields, so
memberCredentials.allow: was silently ignored under
policy: allow-list — turning the policy on would have blanked
credentials working members already depended on.

- derive(profiles, configuredAllow) unions memberCredentials.allow
  into the derived set, with SSH_AUTH_SOCK explicitly excluded from
  that union (it stays governed only by sshAuthSock: allow).
- HerdrPeerLauncher threads MemberCredentials.allowSet() into the
  derivation instead of calling the profiles-only overload.
- Added a per-spawn INFO log 'member credentials: allowed N of M'
  (N/M from the daemon's own env, the existing hostEnvNames proxy),
  never logging a blocked name or a value.
2026-08-28 05:56:27 +07:00
Dai Ha c5e24197bf #150: report lead readiness from delivery gate
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m37s
2026-08-28 05:54:40 +07:00
Dai Ha bcb402b688 #157: convert forge worktree origins to SSH
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m11s
2026-08-28 05:54:15 +07:00
Dai Ha 7c4170ff6d CB-643: join the message-layer evidence to the health monitor
CI / build (push) Successful in 1m8s
CI / contract (push) Successful in 1m9s
CB-640 published the three message-layer facts and CB-641 wired the herdr
and time ones. This joins them, so every HealthSnapshot field now carries
real evidence and the NOT_YET_OBSERVED placeholder is gone. That constant
is what made 8 of the 9 fault states unreachable, GONE and NEVER_READY
included, which is why CB-580's failTarget never fired.

hasOrphanedDelegation is a true snapshot, but it can read true for one
tick during an ordinary race: an async ticket exists before its virtual
thread reaches rendezvous.open, so for that instant nothing is accepted or
queued behind it. decide maps the field straight to DELEGATION_ORPHANED
with no smoothing, so one racy read would log a fault that clears on the
next tick. The monitor now requires two consecutive observations. That
costs one interval on a real orphan and removes the false positive.

Two tests drive real ticks against a genuinely orphaned ticket (an
unanswered fleet_ask that lapsed back to PENDING), not the seam: one tick
reports nothing, two report once, and a single clean tick in between
resets the streak.

Also correct two config comments. paneProbeIntervalSeconds is parsed and
read by nothing, so its "minimum 60" note promised a floor that does not
exist.

970 tests green.
2026-08-27 22:15:39 +07:00
Dai Ha a6095743f0 Merge CB-641: wire herdr and time evidence into the fleet health monitor 2026-08-27 22:06:50 +07:00
Dai Ha 66e776d178 CB-641: wire herdr health evidence
CI / contract (pull_request) Successful in 1m9s
CI / build (pull_request) Successful in 1m11s
2026-08-27 22:03:20 +07:00
Dai Ha 26f64cba45 Merge CB-640: MessageService evidence accessors for the health monitor
CI / build (push) Successful in 1m4s
CI / contract (push) Successful in 1m28s
2026-08-27 22:02:36 +07:00
Dai Ha 51047848f1 CB-642: make the fleets-status redaction global — a non-global sed leaks a second URI on the same line
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m29s
2026-08-27 22:01:15 +07:00
Dai Ha d8c0b657e8 Merge CB-642: /fleets-status skill — multi-fleet status over the shared LavinMQ 2026-08-27 22:00:49 +07:00
Dai Ha 312c0584ce CB-640: add MessageService message-layer health evidence accessors
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m11s
hasQueuedDelivery/hasStrandedReply/hasOrphanedDelegation surface three of the
message-layer facts FleetHealthMonitor needs but currently hardcodes to
NOT_YET_OBSERVED. Additive only — no existing public method's signature or
behavior changes.
2026-08-27 21:58:21 +07:00
Dai Ha da2625acfb CB-642: add shared fleets status skill
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m36s
2026-08-27 21:55:27 +07:00
Dai Ha 97d9cebc59 CB-634: skip the zsh scrub-report test when /bin/zsh is absent
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m27s
scrubWritesAnAllowedNofMReport spawned /bin/zsh with no guard, so it
errored on the Gitea CI runner (Linux ARM64 container, no zsh) while
the two sibling zsh tests already skipped there via assumeTrue. Add the
same guard so CI skips instead of failing; the Mac build still runs it.
2026-08-25 04:18:45 +02:00
Dai Ha 3ce76a5d69 CB-634: finish the bridge_* -> fleet_* tool rename in docs and comments
CI / contract (push) Successful in 1m5s
CI / build (push) Failing after 1m35s
The alias removal left bridge_* tool names in prose. Fix them:
- README no longer claims the old bridge_* names still answer (they were removed).
- pom + LeadTabScanner comments name fleet_* tools.
- FleetMcp comment no longer mentions the removed deprecated twin.
- docs/MCP-Contract.md and e2e swept bridge_* -> fleet_*; e2e ask files renamed.
The historical mcp__bridge__* mount-name note in CLAUDE.md is kept on purpose.
949 tests pass.
2026-08-25 04:06:43 +02:00
Dai Ha 2e138a199b CB-634: one shared "fleet" workspace + rename bridged -> fleetd cutover
Two changes ship together here.

1. One shared herdr workspace. The lead and every worker now live in one
   workspace called "fleet", so the operator sees one "session" with many
   windows, not two. Before, the lead sat in a "leads" workspace and workers
   in "bridged-workers", which read as two sessions. The lead is still told
   apart from workers by its exact tab label ("lead: <name>"), so putting them
   in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
   for split layouts; Fleetd now passes an empty exclude set.

2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
   launchd/systemd units, module dir, and MCP mount).
   - Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
   - Log line, comments, docs, and CLAUDE.md updated to say fleetd.
   - Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
     bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
   - Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
     bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
   - Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
     still read as a fallback, and still gitignored.
   - MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
     server name is "fleet". The mount name in the local .mcp.json becomes
     "fleet" (gitignored, not in this commit).
   - Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
     BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.

Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.

Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.

The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).

949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
2026-08-25 04:01:08 +02:00
Dai Ha 450a5ed9c5 Merge CB-634 + CB-636 + CB-637: members-IDE overlay, per-profile auto-compact window, cross-host lead coordination
CI / contract (push) Successful in 1m16s
CI / build (push) Failing after 1m30s
- CB-634: deliver IDE guidance as an on-disk CLAUDE.local.md / opencode overlay,
  pinned to the module dir, best-effort auto-open (opt-in per profile via ideMcpUrl).
- CB-636: per-profile autoCompactWindow -> --autocompact for claude-code,
  provider.<p>.models.<m>.limit.context for opencode. Range-validated [100000,1000000].
- CB-637: cross-host lead-to-lead over a shared AMQP coordination vhost
  (coordinator: block, fleet_send{coordId}, LeadMailbox + LeadCoordLoop).

952 unit tests + 5 LeadMailbox contract tests green. Live E2E verified on the Mac daemon.
2026-08-24 19:11:44 +02:00
Dai Ha edabccd885 CB-637: relabel lead-comms (was colliding CB-635) + document coordId route in the CLAUDE.md tool table
The lead-to-lead wiring shipped with CB-635 in its comments, but CB-635 is
already the broker.uriEnv / unreachable-broker work. Relabel the mailbox +
fleet_send{coordId} + receive loop to CB-637 so a ticket number names one
feature. Add the cross-host peer-lead row to the primary intent->tool table
(kept byte-identical with the wiki template).
2026-08-24 17:48:51 +02:00
Dai Ha 1dbe3a03fc Merge lead-comms-wiring: wire LeadMailbox into the daemon (fleet_send{coordId} + receive loop) 2026-08-24 17:42:11 +02:00
Dai Ha 6058b8472b lead comms: wire LeadMailbox into the daemon so lead-to-lead messages flow
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Failing after 1m22s
The sibling ticket landed the mechanism (LeadMailbox, LeadMessage, the
coordinator: config block) but nothing opened it, nothing sent through it, and
nothing read it. This is the wiring.

- LeadChannel: a small interface LeadMailbox now implements (publish/peek/ack
  plus a selfCoordId() accessor). It exists so FleetMcp and the receive loop can
  be tested with a fake instead of a live broker. LeadMailbox's AMQP logic is
  untouched — the diff is the implements clause, four @Override marks and the
  accessor.

- Fleetd.openLeadMailbox: opens this daemon's mailbox after the reply inbox is
  selected, with the same env-injected seam selectReplyInbox uses. Every "off"
  path returns null and the daemon still starts: no coordinator block (silent),
  a uriEnv that does not resolve (INFO), a configured broker with no selfId
  (WARN — a mailbox is named after the coord-id that owns it), or a broker that
  refuses at boot (WARN, credentials stripped). Closed in the ordered shutdown
  hook, after the loop that reads it has stopped.

- fleet_send{coordId}: publishes a LeadMessage(from=selfCoordId, to=coordId) to
  the peer's mailbox and returns the broker-confirmed receipt. coordId is
  mutually exclusive with sessionId/turnId and is rejected by name rather than
  resolved by precedence. An unroutable/nacked/timed-out publish comes back as a
  tool error naming the coordId, never a crash. The worker send/reply path is
  not touched.

- LeadCoordLoop: the receive half. Each tick peeks the mailbox, resolves the
  local lead pane, and — only at a turn boundary — injects "[lead <from>] <text>"
  and acks. Anything not delivered stays unacked and is retried, so a message is
  never dropped; one message per tick, so every delivery is gated on a status
  read that already saw the previous one.

- fleet_list reports {selfId, configured} when coordination is on, so an
  operator can find the coord-id a peer must use to reach them. Omitted
  entirely when it is off.

Tests: 20 new hermetic tests (no broker) across routing, delivery and startup
selection. mvn clean install: Tests run: 944, Failures: 0, Errors: 0, Skipped: 0
— BUILD SUCCESS.
2026-08-24 17:39:29 +02:00
Dai Ha f29968c334 Merge autocompact-window: per-profile autoCompactWindow for claude-code + opencode members 2026-08-24 17:32:53 +02:00
Dai Ha 7b98cca967 LeadMailbox: durable leader-to-leader mailbox over a shared coordination vhost
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m35s
Adds the broker-side mechanism for lead-to-lead messages across daemons/hosts
(unit 1 of 2): a LeadMessage envelope carrying from/to coord-ids, an
AMQP-backed LeadMailbox modeled closely on AmqpReplyInbox (consume-and-hold,
deferred manual ack, confirm-mode publish, recovery handling), and a new
optional coordinator: config block (separate vhost from broker:, leader
traffic only). Config parsing + accessors only — FleetMcp/Fleetd/Injector/
MessageService and the send path are untouched; wiring is a separate ticket.
2026-08-24 17:23:42 +02:00
Dai Ha 2757bc7185 autoCompactWindow: per-profile bounded auto-compaction, both backends
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Failing after 1m37s
Add opt-in Integer autoCompactWindow to FleetConfig.Profile (last field,
null/unset = today's behaviour). Validated at config load to [100000,
1000000] — the band Claude Code's own --autocompact flag accepts.

Claude Code: appends --autocompact <window> to argv (mirrors --model),
so it survives the ccs <profile> wrapper.

opencode: has no absolute compact-at-N knob (only compaction.auto/prune/
reserved/tail_turns/preserve_recent_tokens), so the window is applied as
the resolved model's own limit.context (+ a required limit.output:16384
default) in the generated opencode.json, merged via get-or-create nodes
so it does not clobber a custom-provider block. Only applies when model:
resolves to "provider/model"; otherwise logs a WARN naming the profile
rather than silently doing nothing.

Docs added to fleetd.example.yaml explaining the cross-backend semantics
difference (compacts AT the window vs. WITHIN it).
2026-08-24 16:54:40 +02:00
Dai Ha 7655f1b51a CB-634: pin the IDE overlay to the module dir + best-effort auto-open
The overlay pinned project_path to the worktree root. For a repo whose Maven
module is a subdir (this repo's pom is in `bridged/`, not at the root), opening
the root imports no module and every ide_* call resolves nothing. Pin and open
the module dir instead.

Two new opt-in per-Profile keys, both read only when ideMcpUrl is set:
- ideProjectDir: repo-relative module dir the IDE opens and the overlay pins;
  blank keeps the old worktree-root behaviour.
- ideOpenCommand: host command that opens that dir in the IDE at spawn, with
  {dir} substituted and run through /bin/sh -c so env (e.g. DISPLAY) can be set
  inline. Best-effort and non-fatal — a failure never fails the spawn. Blank
  keeps the manual-open behaviour. No close half yet (deferred).

Shared helpers PeerLauncher.ideProjectPath / openInIde back both launchers.
The two Profile fields ride a back-compat constructor, so every existing call
site and YAML compiles and behaves unchanged.

Tests: overlay content pins the module dir when ideProjectDir is set;
ideProjectPath resolution; openInIde no-op on a blank command. 918 tests green.
2026-08-24 06:47:37 +02:00
Dai Ha a5efb7c676 CB-634: write the overlay exclude to the common git dir, not the per-worktree gitdir
git reads info/exclude from the common dir for a linked worktree (only
info/sparse-checkout is per-worktree), so the entry written into
<common>/worktrees/<name>/info/exclude was never honoured and CLAUDE.local.md
showed as untracked -- at risk of being swept into a worker's PR. Derive the
common dir (<common>/worktrees/<name> -> <common>) and write there. Found by
dogfooding a real spawn on fleet01; the test now uses the real worktree layout
and asserts the entry lands in the common dir, not the per-worktree gitdir.
2026-08-23 20:06:36 +02:00
Dai Ha 5a8cf4cb4d Merge CB-634 overlay redesign: deliver IDE guidance as on-disk CLAUDE.local.md / opencode instructions overlay 2026-08-23 16:45:50 +02:00
Dai Ha a3fc7e4df8 CB-634: deliver IDE guidance as an on-disk overlay, not the system-prompt charter
Move the IDE guidance text to PeerLauncher.ideOverlayText (shared by both
launchers). ClaudeCodeLauncher drops it from the reply-charter file and writes
CLAUDE.local.md into a provisioned worktree instead, gated on a .git FILE
(safety: never writes into the primary's real .git-DIRECTORY checkout) and
registers it in info/exclude. OpenCodeLauncher mounts the intellij server and
adds the rules file to the instructions array.
2026-08-23 16:37:08 +02:00
Dai Ha d811b30df3 CB-634 (draft): mount the IDE Index MCP into a member, opt-in per profile
Adds `ideMcpUrl` to FleetConfig.Profile (default off). When set, the
Claude Code launcher mounts the IDE Index MCP as a second inline
--mcp-config server named `intellij`, and appends an IDE charter that
pins every ide_* call to the member's own worktree (spec.cwd()). The
charter order is role -> ide -> reply, one --append-system-prompt-file,
reply last (CB-618). The mount gate now fires on ideMcpUrl alone, not
only mcpUrl. ConfigRef treats an ideMcpUrl change as deferred, like the
other launch flags.

Never touches .mcp.json or CLAUDE.md — the mount and the rule arrive as
launch flags, so a project's own config is untouched.

Not yet done (see fleetd #162): the bridged-owned IDE lifecycle
(open on provision, close before worktree removal), and the opencode
adapter (separate ticket). fleetd.example.yaml documents ideMcpUrl and
fixes the stale parityOverlay default.

911 tests green.
2026-08-23 16:00:36 +02:00
Dai Ha 3f4ac2b24e CB-635: --check reports whether broker.uriEnv resolves in a login shell
CI / contract (push) Successful in 46s
CI / build (push) Failing after 1m30s
An empty uriEnv no longer stops the daemon (#152), so the failure is quiet: bridged
starts, falls back to the in-memory reply inbox, and held reports stop surviving a
restart. --check is the only thing that says so before the fact. The var name is read
out of bridged.yaml so a renamed key cannot make the check lie.
2026-08-23 14:03:08 +02:00
Dai Ha 4644359128 Merge CB-635: broker.uriEnv keeps the AMQP password out of the config, and an unreachable broker no longer stops the daemon (#151, #152)
CI / contract (push) Successful in 1m11s
CI / build (push) Failing after 1m42s
2026-08-23 13:58:21 +02:00
Dai Ha 95a8dbcea9 broker: uriEnv config + non-fatal unreachable broker on boot
CI / build (pull_request) Failing after 56s
CI / contract (pull_request) Successful in 1m6s
#151: Broker gains uriEnv beside uri, taking the AMQP URI from an env var so
the password stays out of fleetd.yaml (same pattern as auth.tokenEnv). uriEnv
wins when set; isConfigured() treats a uriEnv naming an unset/blank variable
as unconfigured. A configured uriEnv is added to the startup required-secrets
report. Never logs the resolved URI (it carries the password).

#152: AmqpReplyInbox.open throwing at boot no longer stops the daemon. The
selection at the call site catches the failure and falls back to the in-memory
inbox for the process lifetime, warning loudly that durable cross-restart
delivery is off and logging the failed URI with credentials stripped.
2026-08-23 13:49:26 +02:00
ltms 83b50753fe Merge CB-633: constrain a member's environment with a derived allow-list
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m23s
Round 2 fixes the defect that mattered: the scrub lived in .zlogin only, and a
herdr pane is a login shell on macOS but a plain interactive one on Linux. It
would have protected nothing on the vhost it was built for, in silence.

Verified by the lead: 901 tests, 0 failures, mvn clean install green. Both new
tests mutation-checked -- unwiring the control fails one, reverting the scrub to
.zlogin alone fails the other while the login-shell test still passes.

NOT yet deployed: the running daemon still holds the old jar.
2026-08-23 08:22:06 +02:00
Dai Ha 6f968d59a4 CB-633 round 2: the scrub ran on macOS and did nothing on Linux
CI / build (pull_request) Failing after 1m8s
CI / contract (pull_request) Successful in 1m10s
Six review findings, from the peer lead `vms` and a reviewer worker. The first
one is a real defect that would have shipped as a dead control.

1. The scrub only ran in a login shell. It lived in the generated `.zlogin`,
   and zsh reads `.zlogin` only for a login shell. herdr does not open the same
   kind of shell everywhere: measured on herdr 0.8.0, a macOS pane runs `-zsh`
   (login) while a Linux pane runs a plain `/usr/bin/zsh`. So on the vhost this
   was being built for, `.zlogin` never ran and every member kept the whole
   secret store, in silence.

   The scrub body now lives in a generated `scrub.zsh` that BOTH `.zshrc` and
   `.zlogin` source, each after sourcing its own `$HOME` counterpart. Linux
   runs the first, macOS runs both, and the second pass is not merely harmless
   -- it re-scrubs anything the operator's `~/.zlogin` exported after `~/.zshrc`
   had finished. Re-running is idempotent.

2. `INFRASTRUCTURE_PASSTHROUGH` listed names that are not infrastructure:
   ANTHROPIC_AUTH_TOKEN, GITEA_TOKEN, GITEA_HOST, ANTHROPIC_BASE_URL,
   ANTHROPIC_MODEL, CLAUDE_CONFIG_DIR, OPENCODE_CONFIG, BRIDGED_MEMBER. I read
   each injection point and confirmed every one of them reaches `launch.env()`
   only when actually injected, so `allowed.addAll(launch.env().keySet())`
   already covers the legitimate case. As static entries they were pure leak
   surface: a host that happened to export ANTHROPIC_AUTH_TOKEN would have had
   it passed straight through.

3. A missing `scrub-report.txt` at teardown was logged at debug. The report is
   the only evidence the scrub ran at all. Its absence has an innocent reading
   and a serious one, and we cannot tell them apart from the daemon -- so it is
   now a WARN that says exactly that. Logging it at debug is how a control that
   quietly stopped working stays unnoticed.

4. Nothing tested that the control was wired in. Deleting the single
   `applyEnvironmentAllowListPolicy(cfg, launch)` line left all 896 tests green
   while turning the feature completely off -- the CB-586/CB-611 shape again.
   `HerdrPeerLauncherAllowListWiringTest` starts a real spawn and asserts on the
   env that reached herdr. Mutation-checked: unwiring that line fails it.

5. `EnvAllowListScrubTest` now also runs `zsh -i` with no `-l`, which is the
   Linux pane shape, so finding 1 is tested from a Mac. Mutation-checked:
   putting the scrub back in `.zlogin` alone fails that test alone, while the
   login-shell test still passes -- which is exactly the blind spot that let
   the bug through.

6. Two ZDOTDIR leaks closed. A failed spawn has no pane id, so its directory
   was never keyed for teardown; it is now removed on the way out. And
   `deleteOnExit` covers a clean shutdown and nothing else, so `generate` now
   reaps sibling directories older than 24h left by a killed daemon.

Also: `policy:` is lowercased with Locale.ROOT, and the `.zlogin`-only claim is
corrected in fleetd.example.yaml, FleetConfig and HerdrPeerLauncher.

901 tests, 0 failures, `mvn clean install` green.
2026-08-23 08:17:52 +02:00
Dai Ha d432df8e5c CB-633: memberCredentials policy=allow-list — derived ZDOTDIR env scrub
CI / build (pull_request) Failing after 1m4s
CI / contract (pull_request) Successful in 1m19s
Move member environment control out of the pane-creation env overlay
(defeated by any file the login shell sources) into a per-spawn ZDOTDIR
directory whose .zlogin runs LAST, after the operator's whole chain, and
blanks every exported variable not on an allow-list DERIVED from what the
launcher itself injects (profiles' tokenEnv/gitTokenEnv/gitHostEnv/env
keys + an infrastructure set) — never hand-typed.

- memberCredentials.policy: allow-list (deny-by-default/deny-list stay
  default and unchanged); known:/allow: become reporting only under it.
- memberCredentials.sshAuthSock knob, blocked by default; allowing it is
  an explicit decision (operator ssh-agent handle).
- Non-zsh login shell: loud WARN, protection off, fallback to the old
  enumerated-name overlay.
- Scrub writes an 'allowed N of M' denominator report, read at teardown;
  credential-shaped blanked names go to WARN (names only, never values).
- Equality test against a real login zsh from a clean parent: surviving
  non-empty exports EQUAL baseline ∩ derived allow-list.
2026-08-23 07:43:38 +02:00
Dai Ha 4b822731e6 CB-632: rename the member's MCP mount bridge -> fleet
CI / contract (push) Successful in 40s
CI / build (push) Successful in 1m22s
Every launcher writes the bridge's MCP server into the config it hands its peer,
and it named that server "bridge". So a member addressed its tools as
mcp__bridge__fleet_send while the tools themselves are already fleet_*. The
mount is named "fleet" now, and a member's tools are mcp__fleet__*.

The name was a bare literal in three files: ClaudeCodeLauncher and LeadLauncher
build a --mcp-config JSON string, OpenCodeLauncher writes an opencode.json node.
Three hand-written copies of one name is how a rename lands in two of them, so
the name is now one constant, PeerLauncher.MCP_MOUNT_NAME.

The mount name is local to the peer — it is the label its own client puts on the
server, and nothing in the daemon reads it back. Renaming it changes no wire
call.

Tests. Each launcher's test now asserts the mount is named fleet AND that
nothing writes "bridge"; the second half is the part that would have caught a
half-done rename. LeadLauncherTest never checked the name at all, only the URL,
so it gained the assertion rather than had one updated.

CLAUDE.md's role-detection ladder quoted mcp__bridge__* as the marker of a
spawned member. It names mcp__fleet__* now, and says that a member spawned
before this change still reports the old prefix. The portable block stays
byte-identical with the wiki template (wiki 569a917).

Build: cd bridged && mvn clean install, then read target/surefire-reports/*.xml
directly — 884 tests, 0 failures, 0 errors.
2026-08-23 07:04:44 +02:00
Dai Ha e38eac1a33 CB-632: ignore fleetd.yaml too, and fix the docs that named the old example
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m31s
Unit 3 renamed bridged.example.yaml to fleetd.example.yaml and taught Fleetd to
read fleetd.yaml first. Two things it left behind.

bridged/.gitignore still ignored only bridged.yaml. An operator who follows the
new comment and copies the example to fleetd.yaml gets an untracked live config
holding tokens, and git offers to commit it. Both names are ignored now, because
both names work until the cutover.

Three design docs still pointed readers at bridged.example.yaml, a file that no
longer exists under that name.
2026-08-23 07:00:30 +02:00
Dai Ha 8101290933 CB-632: rename the exported metrics bridged_* -> fleet_*
Part of #145 (CB-632). Doing this NOW, ahead of the rest of the path
renames, for one reason: the vms lead is about to wire a monitoring
dashboard to these series. Renaming a metric after a dashboard points at
it breaks continuity and silently leaves a dead panel. Renaming it before
costs nothing, so it goes first rather than at the cutover.

Nine series renamed, all declared in FleetMetrics.

Two real defects found while doing it:

  - FleetApp had "bridged_auth_failures_total" written as a LITERAL
    instead of using FleetMetrics.AUTH_FAILURES -- a second hand-written
    copy of a name, which is how these drift. It now uses the constant,
    so there is one source for that name again.
  - Nothing guarded the prefix. One test does assert a wire name
    (FleetAppAuthTest checks the real /metrics body for fleet_sessions),
    which is good, but it covers one series out of nine. The literal
    above was covered by nothing at all.

So this adds MetricNamesTest, which reads the constants reflectively
rather than listing them -- a test that lists the nine names is itself a
second hand-written copy, and would pass while a tenth went unchecked.
It asserts its own denominator too: "no name starts with bridged_" is
true of an empty set, so a sweep that found nothing would pass loudly.
Asserting the count of 9 makes a broken sweep fail instead.

Also renamed three herdr contract-test workspace labels, __bridged_* ->
__fleet_*. Those are throwaway workspaces created by ensureWorkspace, not
metrics, but they are the same word.

Verified: mvn clean install green, 52 classes, 881 tests, 0 failures.
Test count is up by 3 -- the new prefix, denominator and uniqueness
checks. No bridged_ string remains anywhere outside wiki/.
2026-08-23 06:59:37 +02:00
Dai Ha 6e7fc12f89 CB-632 unit 3: point text references at fleetd.example.yaml
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m36s
2026-08-23 06:51:48 +02:00
Dai Ha 50df14a50f CB-632 unit 3: prefer fleetd.yaml, fall back to bridged.yaml; log the chosen file 2026-08-23 06:51:35 +02:00
Dai Ha ccd882f2fc CB-632 unit 3: rename the example config file to fleetd.example.yaml 2026-08-23 06:48:16 +02:00
Dai Ha ecc590f344 CB-632 unit 5: rename the daemon and classes in the docs prose
Part of #145 (CB-632). Documentation only, plus one internal literal.

Unit 1 renamed the package and classes, which left every doc describing
classes that no longer exist. This fixes the prose across README.md,
docs/ and bridged/docs/ -- 18 files.

Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and
"bridged" where it names the daemon as a product rather than a path.

Also renamed two literals, because a doc that disagrees with the code is
worse than one that is out of date:

  - bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey
    OpenCodeLauncher sends when a profile resolves no token, to a local
    endpoint that does not check it. No test asserts the old string.
  - the vnd.ltms.bridged.* media type in the M4 design doc. It appears
    in no Java file, so nothing implements it yet.

Deliberately NOT renamed, because each is still literally true today and
changes only at the cutover:

  - paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar,
    .bridged-worktrees, deploy/dev.ltms.bridged.plist,
    scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh
  - bridged_* metric names -- renaming these after the monitoring is
    wired would break dashboard continuity, so they move before it is
  - bridge_* MCP tool names, which answer alongside fleet_* on purpose
  - BRIDGED_* environment variables, read by a file outside this repo

Method note: perl, not sed. BSD sed has no \b and no lookaround, and a
word-boundary expression there fails silently. The prose replace uses
(?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an
identifier, then every remaining hit was read by hand.

Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
2026-08-23 06:46:34 +02:00
Dai Ha 3b7cdf9365 CB-632: the addendum named a class unit 1 renamed
mcp/BridgeMcp is now mcp/FleetMcp. This line is in the project addendum,
not the canonical block, so the wiki template is untouched -- the sync
check still returns True.

Part of #145.
2026-08-23 06:39:54 +02:00
Dai Ha 9b50dd69d8 CB-632 unit 1: rename the Java package and classes to fleet
Part of #145 (CB-632), under epic #125.

The product is called fleet and the daemon is called fleetd, but the code
still said bridge everywhere. This renames the Java half:

  package dev.ltms.bridged -> dev.ltms.fleet
  Bridged        -> Fleetd          (the main class)
  BridgedConfig  -> FleetConfig
  BridgeMcp      -> FleetMcp
  BridgedApp     -> FleetApp
  BridgedMetrics -> FleetMetrics

The package root is dev.ltms.fleet, not dev.ltms.fleetd. The trailing d
means daemon, which names a process, not a namespace.

What this commit deliberately does NOT change:

  - The module directory stays bridged/, and <finalName> stays bridged.
    The installed launchd plist names bridged/target/bridged.jar and its
    KeepAlive is armed, so renaming the jar on its own strands a restart.
    Both change at the cutover, together with the plist, in one step.
  - The bridge_* MCP tool aliases. CB-622 shipped both names on purpose.
    One test names a local variable viaBridge because it holds the result
    of the deprecated call; the rename collided with it and the compiler
    caught it. That variable is back.
  - BRIDGED_* env var names, and bridged.yaml. Both are operator
    contracts and need a read-both shim, which is a later unit.

Two things a plain search-and-replace would have missed:

  - logback.xml and logback-test.xml name the package twice, once as a
    turboFilter class= attribute. The compiler never checks those.
  - BSD sed does not support \b. The word-boundary expression matched
    nothing and said nothing, while the other ten in the same command
    worked. Checked the leftovers instead of trusting the exit code.

Verified: mvn clean install green, 51 test classes, 878 tests, 0 failures
-- the same count as before the rename.
2026-08-23 06:27:43 +02:00
ltms 08f1a79800 Merge #143: scripts/rename-checkout.sh (CB-624)
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m22s
Adds the rename script for CB-624 (#128), with a read-only --check mode.

Two workers ran this ticket in parallel and were kept apart on purpose: one wrote the script, one inventoried the same rename read-only without seeing it. The inventory found three surfaces the script had missed, all verified on the live machine before being acted on: ~/.claude.json's projects key (trust state, enabledMcpjsonServers, allowedTools, lastSessionId), ~/.config/herdr/session.json, and the JetBrains recentProjects.xml / trusted-paths.xml pair in 2026.1 and 2026.2.

All three are report-only by decision, not by oversight, and the reason is written into the script header: ~/.claude.json is global config written by live sessions, herdr/session.json is live process state, and IntelliJ rewrites its own files when the project is reopened.

While adding them the author found a trap worth recording: JetBrains stores $USER_HOME$/LTMS/claude-bridge, not the absolute path, so a plain absolute-path grep reports 0 hits on files full of them. The script now counts both forms and says which form matched.

Every apply-mode branch is unexecuted. Running it moves the directory this system runs from and stops the daemon the workers talk through, so the script is well-reasoned, not proven. Its first real run is its test. Verified before merge: CI run 215 green on both jobs, bash -n passes. shellcheck is not installed on this machine, so it never ran.
2026-08-23 06:09:18 +02:00
Dai Ha 51bdec22e7 CB-624: report three more rename surfaces as report-only
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Successful in 1m14s
--check gains section 6 counting ~/.claude.json (the projects entry keyed
by the old absolute path), ~/.config/herdr/session.json, and JetBrains
recentProjects.xml / trusted-paths.xml across every IntelliJIdea* version.
Apply mode's closing summary lists the same three with manual follow-ups.

None of the three is rewritten automatically, on purpose: ~/.claude.json
is global live Claude Code config, herdr session.json is live process
state, and JetBrains rewrites its own files when the project is reopened
at the new path. Header comment states the reason for each so it does not
read as an oversight.

JetBrains stores these paths as its $USER_HOME$ macro rather than a
literal absolute path, so the counter matches both forms and says which
form the hits used.
2026-08-23 06:07:16 +02:00
Dai Ha d43f670285 CB-624: add scripts/rename-checkout.sh to rename the checkout safely
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m32s
Renames ~/LTMS/claude-bridge -> ~/LTMS/fleetd as one auditable command,
modeled on redeploy-bridged.sh. --check reports every surface holding the
old absolute path (checkout files, launchd plist, Claude Code project
state, worktree .git pointers, running daemon). Apply mode stops the
daemon first (launchctl-aware), moves the checkout and the Claude Code
project-state slug dir derived from both paths, repairs worktree gitdir
pointers, rewrites bridged.yaml and the installed plist if they exist,
restarts, and verifies /healthz plus a fresh 'bridged listening' line
anchored to a pre-stop marker.
2026-08-23 05:55:04 +02:00
ltms c135583402 Merge #139: drop the unused rabbitmq host port mapping
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m19s
The contract job reaches the broker by network alias (AMQP_URI=amqp://guest:guest@rabbitmq:5672), so the host mapping 5672:5672 was never used. It only bound a port on the runner host, which made two concurrent runs collide and the service container fail to start.

Seen on 2026-08-22: run 209 (PR) contract=success and run 210 (the merge of that same code) contract=failure with every step, including checkout, marked cancelled. Both started at 22:08. Re-running 210 alone on the identical commit passed.

PR #139's own run 212 is green on both jobs with the mapping gone, which proves the alias path still works.

Also the worker-PR proof for CB-623 (#127): opened by the agent account against fleet/fleetd after the org transfer.
2026-08-23 05:41:26 +02:00
Dai Ha 8d2893b67e CB-ci: drop host port mapping for rabbitmq service in contract job
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m33s
2026-08-23 05:35:58 +02:00
Dai Ha 555715ced9 CB-623: point every path at fleet/fleetd after the org transfer
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m29s
The repo moved lms/claude-bridge -> fleet/claude-bridge -> fleet/fleetd.
Gitea redirects hold, so most of this is not urgent, but one line was a
real break: the implementer skill posts a worker's PR to a hardcoded
repo path, so every worker PR would have gone to the old address.

  .claude/skills/implementer/SKILL.md  the worker PR endpoint (functional)
  .gitmodules                          wiki submodule URL
  CLAUDE.md + wiki/7-Use-Cases.md      the canonical block, kept byte-identical
  README.md                            clone command and wiki link
  deploy/bridged.service               Documentation=
  plugin/.claude-plugin/plugin.json    homepage + repository
  docs/*.md                            issue and wiki links

The wiki is not a separate repo. /repos/lms/claude-bridge.wiki returns 404
and lms owned no .wiki entity, so the wiki moved with the repo; both the old
and the new wiki SSH URLs resolve to the same sha. Ticket step 4 assumed a
second transfer that does not exist.
2026-08-23 05:22:04 +02:00
ltms 446a9d395c Merge CB-622: canonical block to fleet_*, charter quote fixed
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m22s
2026-08-22 22:08:44 +02:00
Dai Ha 8311f2db7f CB-622: update the canonical block to fleet_*, and fix the broken charter quote
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m28s
The eleven tool names become fleet_* throughout the portable block. The
fallback ladder keeps mcp__bridge__* unchanged and that is deliberate: both
launchers still mount the server as "bridge", so a spawned member really
does see that prefix. Renaming the mount is a separate surface CB-622 did
not touch.

Separately, a pre-existing defect: the ladder quoted the reply charter as
"You are an off-subscription worker in the claude-bridge fleet", while
REPLY_CHARTER says "You are a spawned member in the claude-bridge fleet".
A member matching that quote found nothing, so the ladder's first rung
could never fire. Now quoted verbatim.

wiki/7-Use-Cases.md is advanced to the matching template commit; the sync
check prints in sync: True.
2026-08-22 22:08:12 +02:00
ltms f4570ff274 Merge CB-622 lead follow-up
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m1s
2026-08-22 21:59:46 +02:00
Dai Ha 3aea4e1ec6 CB-622 lead follow-up: the three changes no worker could make
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m30s
1. opencode.json — mount key bridged -> fleetd. Unit C could not do this:
   the file is neutralized by the worktree overlay, so every worker sees a
   stub and correctly reported it held no mount key. Tracked as CB-628.

2. e2e/bridge_ask_transcript.md keeps its name. It is a dated record of a
   run on 2026-07-16 that really did call bridge_ask, and its first line
   says so. The harness now writes fleet_ask_transcript.md for new runs;
   its header already says 'Live fleet_ask', so the two now agree.

3. docs/MCP-Contract.md line 20 — the historical banner describes section
   6, and section 6 now says fleet_ask.
2026-08-22 21:59:23 +02:00
ltms 516366b796 Merge CB-622 Unit B: rename bridge_* to fleet_* in the Markdown docs
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 40s
Lead-verified: 27 .md files, 165 occurrences, no wiki/, no CLAUDE.md. The docs/MCP-Contract.md boundary held exactly — every hunk falls in 223-286, inside section 6 (210-293). The transcript rename is reverted by the lead in a follow-up: that file is a dated record of a run that really did call bridge_ask.
2026-08-22 21:58:16 +02:00
ltms d82d0157cb Merge CB-622 Unit C: rename bridge_* to fleet_* in scripts, e2e and config
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m5s
Lead-verified: 9 files, +68/-68, nothing under wiki/, root .mcp.json untouched. The worker's opencode.json finding was correct about its worktree and led to CB-628 (#134); the tracked file is fixed separately by the lead.
2026-08-22 21:58:09 +02:00
ltms 1deb0c90d4 Merge CB-622 Unit A: dual-register the MCP tools as fleet_* and bridge_*
CI / contract (push) Successful in 37s
CI / build (push) Successful in 1m30s
Lead-verified: both names reach the same handler instance; warn-once uses a per-name Set, not a numeric sentinel; Authz keys on an action enum so the deprecated names keep their authorization. 878 tests green in a local integration with #132, #133 and the lead's opencode.json fix. Live MCP verification is still outstanding and is the lead's step after redeploy.
2026-08-22 21:58:03 +02:00
Dai Ha 41a3114d03 CB-622 Unit A: register fleet_* MCP tools, keep bridge_* working
CI / build (pull_request) Successful in 1m2s
CI / contract (pull_request) Successful in 1m8s
fleet_* is now the documented tool name for all eleven MCP tools; each
bridge_* twin is registered against the exact same handler (no logic
duplication) and its description leads with a DEPRECATED notice. A
bridge_* call logs one WARN naming the old and new name, once per name
for the life of the process (a Set, not a numeric sentinel).

REPLY_CHARTER in HerdrPeerLauncher now tells a spawned member to call
fleet_reply — the one rule that must survive with no repo checkout.

All other bridge_* string literals across mcp/, Javadoc, and tests were
renamed to fleet_* for consistency with the new documented name.
2026-08-22 21:54:32 +02:00
Dai Ha 76e4b577a0 CB-622: rename bridge_* tool names to fleet_* in Markdown docs
CI / build (pull_request) Successful in 1m6s
CI / contract (pull_request) Successful in 1m6s
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/
send/spawn/status/stop/whoami) to their fleet_* names across the Markdown
documentation. fleet_* is written as the normal name; one deprecation note
in README.md says bridge_* still works for one release.

docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293):
sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are
deliberately left with old names so dead text does not look maintained.

Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md
to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and
.claude/skills/port-to-opencode/SKILL.md are owned by other units and are
untouched.
2026-08-22 21:52:24 +02:00
Dai Ha 03473a286b CB-622: rename bridge_* tools to fleet_* in scripts, e2e tests and config; mount name bridged -> fleetd
CI / build (pull_request) Successful in 58s
CI / contract (pull_request) Successful in 1m9s
2026-08-22 21:49:15 +02:00
Dai Ha 5cabd09705 wiki: CB-618 correction to the role-agent entry
CI / contract (push) Successful in 48s
CI / build (push) Failing after 1m0s
2026-08-22 12:37:12 +02:00
Dai Ha cc47672b7c CB-618: one charter file, and agent files need a name:
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m27s
Two defects in CB-617 that only a live spawn could find. Both were shipped
green: every test passed because every test read the argv we built, and none
ran the binary that has to accept it.

1. Claude Code refuses to start when both --append-system-prompt and
   --append-system-prompt-file are on the command line:

     Error: Cannot use both --append-system-prompt and
     --append-system-prompt-file. Please use only one.

   CB-617 put the role charter on the file flag and left the reply charter on
   the inline flag, so every Claude-profile spawn with a role charter died at
   launch. The pane exited on its own and bridged reported it as
   spawn_timeout ("did not reach injectable state within 20000ms"), which
   hides the real cause.

   Both charters now go in the one file, role charter first and reply charter
   last — last is where the reply rule must sit, because it is the rule that
   must survive. A member with only a reply charter keeps the proven inline
   flag, which is also the only form that reaches a member with no repo
   checkout.

2. A Claude Code agent definition needs `name:` in its frontmatter. Ours had
   only `description:`, so the files were skipped and --agent architect failed
   with "not found. Available agents: claude, Explore, ...". Added to all
   three. The OpenCode files take their name from the filename and are
   unchanged.

Checked on this host, in this repo, with the real binary:

  claude --model claude-sonnet-5 --agent architect \
    --append-system-prompt-file /tmp/combined.md -p '...'
  -> ROLEOK, REPLYOK, yes

so the agent definition, the role charter and the reply charter all compose.

The updated test now asserts the constraint that actually binds: with a role
charter present, --append-system-prompt must be absent, and the file must open
with the role charter and end with the reply charter.

874 tests pass.
2026-08-22 12:32:44 +02:00
Dai Ha 37edd9134b Merge CB-617 Unit B (#122): role agent definitions for both backends
CI / contract (push) Successful in 56s
CI / build (push) Successful in 1m16s
2026-08-22 12:23:08 +02:00
Dai Ha 80f167b1f7 Merge CB-617 Unit A (#121): role charter travels by file, not argv 2026-08-22 12:23:08 +02:00
Dai Ha bd6547fca3 CB-617: add role agent definitions
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m31s
2026-08-22 12:21:16 +02:00
Dai Ha 7e97f5bff5 CB-617 Unit A: role charter via file, not argv
CI / build (pull_request) Successful in 1m16s
CI / contract (pull_request) Successful in 1m28s
herdr refused to shell-encode a multi-line inline --append-system-prompt
argument (invalid_agent_argument), which broke any Claude Code profile with
a configured multi-line fleet.charters.<role>. Split delivery: the role
charter now writes to a temp file and mounts via
--append-system-prompt-file; the one-line REPLY_CHARTER keeps its inline
--append-system-prompt delivery, since it must reach a member with no repo
checkout. Also pass --agent <role> when a role's agent-definition file
exists under the worker's cwd (.claude/agents/<role>.md for claude-code,
.opencode/agent/<role>.md for opencode); absent either file, nothing extra
is added and the member still spawns.

The base's LaunchSpec now carries roleCharter/replyCharter/cwd alongside
the existing composed charter field, so CharterReceipt keeps fingerprinting
the same composed text it always did -- unchanged, since the digest covers
the full logical charter content regardless of how it is delivered.
2026-08-22 12:04:38 +02:00
Dai Ha 7d41ccccee docs: replace the decommissioned ollama.ltms.dev host
CI / build (push) Failing after 1m4s
CI / contract (push) Successful in 1m31s
A page audit found the dead host still named in four places outside the wiki.
ollama.ltms.dev no longer exists; the one front door is llm.ltms.dev, and the
path differs by backend - /anthropic for claude-code, /v1 for opencode.
Anyone following README.md today points a member at nothing.

README.md also now links the new wiki chapter 13, the operator guide, since
the README's own next step used to be the never-written Setup page.

Not touched: SubscriptionGuard.java:14 still names ollama.ltms.dev, but only
in a javadoc line describing the Stage-1 example allowlist. It is a comment
about history, not a live default, so it stays until that class is next
edited for its own reasons.
2026-08-17 16:19:19 +02:00
Dai Ha bf616e192a CB-610: document subscription:, the knob that bills the operator's Claude plan
CI / contract (push) Successful in 1m28s
CI / build (push) Successful in 1m29s
profiles.<name>.subscription is read in five places (BridgedConfig record +
isSubscription + validateSubscriptionProfiles, ClaudeCodeLauncher, and the
startup secret check that deliberately skips it) and was in bridged.example.yaml
nowhere. bridged.yaml is gitignored, so the example is the only place an operator
can learn a key exists - which meant a fresh host had no way to discover the one
switch that moves cost onto the operator's own subscription. Two profiles here
have it set.

Documents what changes when it is true, that it is mutually exclusive with
baseUrl and refused at load, that the startup secret check skips such profiles
so a clean secrets report says nothing about them, and that maxLoad is the only
throttle against the operator's plan - there is no metering or budget refusal.

94 config tests pass; the example still loads.
2026-08-17 16:08:20 +02:00
Dai Ha f8bd5d0c51 CB-609: bridge_ack takes target, not ticket; mark MCP-Contract.md as historical
CI / contract (push) Successful in 1m6s
CI / build (push) Successful in 1m17s
The instruction surface named a parameter the tool refuses. bridge_ack requires
target and msgId (BridgeMcp.java:1001) and errors with 'target and msgId are
required' otherwise, but step 5 of the primary procedure said bridge_ack{ticket,
msgId}. The table two sections below it was already correct, so the file
contradicted itself and the wrong half was in the numbered steps a lead follows.
The same line is fixed in the byte-identical wiki template.

docs/MCP-Contract.md still opened with 'Greenfield - no MCP code exists yet' and
CLAUDE.md pointed every session at it. Audited against BridgeMcp.java: it names
two tools that do not exist, omits four that do, gets nearly every parameter name
wrong, and uses /workers paths the daemon does not serve. Its status banner now
says so and names the live schema as the authority; CLAUDE.md's pointer is
narrowed to section 6, which is the part that did survive. Rewrite is CB-609.
2026-08-17 14:26:08 +02:00
Dai Ha ac40de1d30 Merge CB-596: member credential policy is config-driven deny-by-default
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m15s
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow with a
memberCredentials: block: policy + allow + known, validated at config load
in the shape CB-606 established.

What this does and does not do, because the distinction matters:

IT DOES enumerate, report and validate. No credential name is hardcoded in
Java any more. An unrecognised policy: refuses at load. A credential-shaped
env var on neither known: nor allow: is named in a WARN. An absent block
warns loudly at startup, because absence now blocks nothing at all and there
is no hardcoded fallback left.

IT DOES NOT, on its own, block anything the login shell re-exports. The env
overlay is applied at pane creation and the pane's login shell runs after it.
An exec-time fix was attempted and has no seam: herdr protocol 19's
agent.start takes a fixed kind plus trailing CLI args, with no env map and no
argv[0] control. The enforcement therefore lives in the operator's secrets.sh,
which sets BRIDGED_MEMBER-guarded sentinels for the 29 blocked names.

Verified by me, not taken on report: my own unpiped mvn clean install gives
870 tests, BUILD SUCCESS; loading the live gitignored bridged.yaml with the
new code yields 34 known / 5 allow / 29 blocked with an empty intersection;
and that blocked set diffs clean against the secrets.sh block, so the two
lists agree today.
2026-08-17 13:25:39 +02:00
Dai Ha 7930a31b94 CB-596 round 2: the exec-time argv-prefix fix has no seam — stop and report
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m6s
herdr protocol 19's agent.start takes a fixed `kind` (herdr resolves the
executable) plus trailing CLI args for that binary; only tab.create/pane.split
accept an env map, and that IS the round-1 pane-creation overlay already
shipped. There is no argv/env control point that runs after the pane's login
shell and before the agent process starts, so the proposed `env NAME=value ...`
argv prefix cannot be implemented against this API. Documented the finding and
corrected bridged.example.yaml's round-1 comments, which had overclaimed that
the overlay survives the login shell.

Kept everything else: added a startup WARN (Bridged.reportMemberCredentialsGap)
when memberCredentials: is absent or its known: list is empty, so CB-592's
protection loss is never silent, mirroring CB-594's reportRequiredSecrets.
2026-08-17 09:13:57 +02:00
Dai Ha 850fb12807 CB-596: member credential blocking is config-driven deny-by-default, not one hardcoded name
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m26s
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow in HerdrPeerLauncher
with BridgedConfig.MemberCredentials (memberCredentials: policy/allow/known).
Every known name not also allowed is overlaid with a sentinel; an unrecognized
policy value refuses at load; a credential-shaped host env var on neither list
is logged as a gap (name only, never a value). bridged.yaml is gitignored, so
the 31 measured names + 4-name allowlist ship as a commented block in
bridged.example.yaml for the operator to apply live.
2026-08-17 09:02:21 +02:00
Dai Ha b14b66ab03 CB-586: the retention sweep never ran — Long.MIN_VALUE overflowed the gate
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m36s
SessionReaper.lastWipSweepNanos started at Long.MIN_VALUE as a "never swept
yet" sentinel. That sentinel cannot be compared by subtraction. nanoTime() is
positive on this platform, so `now - Long.MIN_VALUE` wraps to a large negative
number, the gate `delta < WIP_SWEEP_INTERVAL_NANOS` reads it as "swept moments
ago", and the method returns before the assignment that would have fixed the
field. The sweep never ran once, for the life of the process, and nothing in
the log said so.

Measured:
  System.nanoTime()      = 31305820625625   (positive)
  now - Long.MIN_VALUE   = -9223340731034150183
  interval (6h in nanos) = 21600000000000
  gate 'delta < interval' -> true  => returns early, every iteration, forever

Fix: a separate `sweptOnce` boolean holds "never yet", so the subtraction only
runs once both operands come from nanoTime. The first pass always sweeps — a
restart is a fine moment for it, the 24h age floor keeps it safe, and the
feature becomes observable right after a redeploy instead of six hours later.

The existing tests all passed because they call SessionManager.sweepWipRefs
directly, which walks around the gate. The new test asserts through the reaper
loop instead: it spawns a worktree session, starts the reaper, and requires a
real prune call at the seam with the 24h floor intact. Removing the fix makes
it fail with "it never reached the seam".

860 tests, mvn clean install, BUILD SUCCESS.
2026-08-16 20:13:37 +02:00
Dai Ha 15ff6bcde5 Merge CB-586: prune refs/wip snapshots whose content is already on main
A refs/wip snapshot ref is deleted only when both hold: its commit's tree is
already reachable from main, and it is older than 24h. Reachability is the safety
floor — a snapshot exists because the work was committed nowhere else, so an
unreachable one is the last copy and is never swept. Every deletion logs the ref
and the sha.

/members gains wipRefs{count,costBytes} so the growth is visible.

Verified against real git, not only the fakes: a recoverable+old ref is deleted,
a recoverable+young one survives the age floor, and the last copy survives. A repo
with no main deletes nothing, and a repo with no snapshots is a clean no-op.

Closes #67
2026-08-16 20:09:16 +02:00
Dai Ha 7822772905 CB-582: bridge_status now reports an open question — say so, and keep the warning
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m7s
The prompt is part of the product: CB-582 changed what bridge_status returns, so
the primary's step 5 was no longer the whole truth.

The second half matters more than the first. A nudge makes a lead more likely to
notice an ask; it does not widen the ~55s window, which is bounded by the
WORKER's own MCP client timeout, not by anything the daemon chooses. Without that
sentence a lead reads 'the ask now nudges me' as 'asking works now' and briefs a
worker to ask — which is the failure CB-582 was filed about.
2026-08-16 19:05:06 +02:00
Dai Ha 0efe1567c0 CB-586: prune refs/wip/* older than 24h whose tree is reachable from main
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m21s
Add the CB-586 retention rule to GitWorktrees and drive it from the reaper:
a snapshot is deleted only when its tree content is already reachable from
main AND the ref is older than 24h. Reachability keeps the last copy of a
worker's work; the age floor stops a fresh snapshot being swept while a
lead is still looking at it. Every deletion logs the ref name and commit
sha so it is recoverable from the reflog. The /members response gains a
wipRefs{count,costBytes} census the operator can read without shelling
into the repo.
2026-08-16 19:02:33 +02:00
Dai Ha 837fed7690 CB-596: the credential probe, as one auditable command
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m39s
Issue #82 step 1 is a measurement, and the classifier refuses an ad-hoc pipeline
that enumerates credential names inside a member — correctly. This is the seam:
one file the operator reads once and then runs, instead of approving a shell
pipeline they have to take on trust.

It never prints a credential value or any part of one. #82's criterion 1 asked
for a 6-character prefix; this prints a truncated SHA-256 instead. A prefix of a
short secret is most of the secret and would end up pasted into a ticket, while
the hash answers every question the prefix was for — is it set, is it the same
value as over there, is it the CB-592 sentinel.

Refuses to run unless BRIDGED_MEMBER=1, since the finding is what a MEMBER holds;
--allow-outside-member takes the comparison reading and labels it as such.

I have not run the reading path. That is the operator's call, which is the whole
point of the ticket.
2026-08-16 19:01:42 +02:00
Dai Ha aa4ee64a34 Merge CB-606: refuse an unrecognized auth.mode or placement at config load
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m40s
Three fields had CB-604's shape — lower-cased, compared against one string,
never checked against the valid set. The auth.mode one was the worst: a typo of
'token' silently behaved as loopback-trust, and validateAuthExposure() only fires
on a non-loopback bind, so a loopback bind hid it end to end. The daemon started
clean and authenticated nobody while the operator believed token mode was on.

All three now refuse at config load, naming the value and the accepted set. The
top-level placement policy was already validated but only lazily at first spawn;
it now calls PlacementPolicies.fromName eagerly at load, so a bad name cannot
start a daemon that merely looks healthy.

Verified by probing the real BridgedConfig.load with eleven values, including the
critical auth.mode typo on a loopback bind, and by loading the live gitignored
bridged.yaml — which the worker cannot see and so could not check.

Closes #106
2026-08-16 18:58:49 +02:00
Dai Ha bb750cdba3 CB-606: refuse an unrecognized auth.mode, per-profile placement, or top-level placement policy at config load
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m25s
An auth.mode typo (e.g. "toekn") used to silently fall back to loopback-trust with no signal
anywhere — validateAuthExposure() only checks the pairing on a non-loopback bind, so on the
common loopback bind the daemon started cleanly and authenticated nobody. Per-profile placement
had the same shape, falling back to legacy pane placement. The top-level placement policy name
was already validated by PlacementPolicies.fromName, but only lazily at first spawn through
CompositePeerLauncher's Supplier; it is now checked eagerly at load, calling fromName itself as
the single source of truth.
2026-08-16 18:54:35 +02:00
Dai Ha d56c77b368 CB-582: close the question when ask() leaves by throwing
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m38s
Found reviewing CB-582 before the merge, not by the implementer.

ask() clears its question on three paths — no-waiter, timed out, and (from
answer()) answered. It can also leave by throwing: an interrupt while blocked on
the answer, or an ExecutionException from the answer future. Those run only the
finally block, which tore down the rendezvous turn but not the push loop's copy.

The result was a question that stayed pending for good: named in every nudge
until it hit its own cap, then left in pendingQuestions with no remover at all.

Teardown now happens where the rendezvous teardown already happens, so the two
cannot drift apart again. Closing a turnId that was never pending is a no-op, so
the normal paths are unaffected.

The new test fails on the pre-fix code with expected: <STOP> but was: <INJECT>.
2026-08-16 18:47:18 +02:00
Dai Ha 83e2ff06cf Merge CB-582: nudge the lead when a worker pauses on bridge_ask
An async (wait:false) delegation opens a ~55s reverse-rendezvous window when its
worker calls bridge_ask. A lead polling on its normal minutes-long cadence never
sees that window, so the worker times out and proceeds without an answer.

The question is now a third source in the CB-588 per-lead push schedule, with its
own per-item nudge count (CB-598's shape), and it is surfaced by bridge_status and
by REST /sessions/{id}/status and /tasks/{ticket} (which previously dropped turnId
on an ASKING phase, so a REST caller could see the question but not answer it).

Closes #61
2026-08-16 18:45:14 +02:00
Dai Ha f5deaafd06 Merge CB-604: refuse an unknown profile kind at config load
CI / build (push) Successful in 1m2s
CI / contract (push) Successful in 1m4s
kind: was lower-cased and compared against one string, so a typo like
'opencod' was accepted and routed to the claude-code adapter. With argv:
unset the launch command became the misspelled string itself, and the
daemon tried to run a program named after the typo. Nothing said a word
until the spawn failed.

It now throws at load, naming the profile, the bad value and the
accepted set - matching rejectNegativeMaxLoad and the duplicate-adapter
check, which already treat routing mistakes as fatal.

Verified here by probing the real BridgedConfig.load with five values:
opencod refused with the full message; opencode, OpenCode, claude-code
and an absent kind all accepted with the right adapter. 832 tests,
BUILD SUCCESS, unpiped.

Closes #102
2026-08-16 18:39:50 +02:00
Dai Ha fe46311266 CB-604: reject an unrecognized profile kind at config load
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m10s
2026-08-16 18:37:21 +02:00
Dai Ha 08968bb1b7 CB-582: make a pending bridge_ask question visible on the lead's poll cadence
CI / build (pull_request) Successful in 1m23s
CI / contract (pull_request) Successful in 1m23s
bridge_ask blocks the worker's turn for ~55s by default (BridgedApp.java,
BridgeMcp.java) — a value deliberately kept just under the worker's own MCP
client's ~60s call cap so the daemon can return a clean timeout before the
client severs the call, not a value that can usefully be widened. A lead
following the charter's wait:false + poll cadence is minutes away, so the
window closes long before a poll would ever see the question — and until now
bridge_poll on such a ticket just read as ordinary "pending" progress.

bridge_poll(ticket) already surfaced Phase.ASKING with the question and
turnId (CB-205); this ships the two pieces that were still missing:

- The lead's own pane is now nudged the instant a question opens, reusing
  the CB-588 ReplyPushLoop push mechanism (a third source alongside queued
  replies and terminal tickets) rather than a new path. The nudge is capped
  by the loop's existing maxReminders budget, and stops the moment the
  question is answered or lapses.
- bridge_status(sessionId) and REST GET /sessions/{id}/status now also show
  an open question and how to answer it, via a new
  MessageService.pendingAsk() lookup — covering the case where a lead checks
  status directly rather than the ticket.
- The REST /tasks/{ticket} endpoint was silently missing turnId on an ASKING
  phase (only the MCP layer's formatted text carried it) — fixed as part of
  making the state genuinely visible over both surfaces.

An unanswered question still behaves as today: the worker proceeds and its
reply says the ask went unanswered — not a hard failure.
2026-08-16 18:35:45 +02:00
Dai Ha 27bbd11f06 Merge CB-584: carry the agent session id on a failed ticket's detail
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m23s
Most of issue #65 was already shipped in 5d5b3bd - MemberSession
records agentSessionId, PeerHandle.agentSessionId() has no default, the
roster exposes it, and bridge_spawn accepts sessionName and
resumeSessionId. That commit left one item for follow-up: carrying the
id on ReleaseDetail.

This does that item. CB-578 stage C already lets a lead re-dispatch onto
the same worktree after a failure; without the session id that is a cold
start. With it, the work and the thread both survive.

Also updates the bridge_spawn and bridge_list rows in the CLAUDE.md
intent table, which the earlier commit missed.

Verified here: trial merge onto main builds 830 tests BUILD SUCCESS,
unpiped. Confirmed against the code that items 1-4 really were already
on main, so the scope-down is correct rather than work skipped.

Closes #65
2026-08-16 18:29:24 +02:00
Dai Ha 48d7841fbf CB-589: document that weighted placement is not cheapest-first
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m43s
The weight ratio does not express a preference order. weighted spreads
spawns across every profile with a free slot, so paid spawns happen
while the free box is idle - and a profile at maxLoad freezes its score,
so it can lose the next pick after a slot frees.

The real fix is a cost-first policy (CB-589). This documents the
workaround and its trap next to the key, because bridged.yaml is
gitignored: a fresh host starts without the workaround and quietly pays,
with nothing to tell the operator why.

Comments only. BridgedConfigTest: 85 tests, BUILD SUCCESS.
2026-08-16 18:27:19 +02:00
Dai Ha cc1df11f69 CB-584: carry agentSessionId on a failed ticket's ReleaseDetail
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Successful in 1m42s
Closes the one piece issue #65 deliberately left out of 5d5b3bd: a
released member's agentSessionId now rides alongside worktree, branch
and snapshotRef on ReleaseDetail, and Bridged's onRelease handler
names it in the abandon reason, so a lead can resume the member's
conversation instead of only re-dispatching a fresh one onto the same
files.

Also updates the CLAUDE.md bridge_spawn/bridge_list table row, which
5d5b3bd shipped the sessionName/resumeSessionId/agentSessionId surface
for but never updated.
2026-08-16 18:25:24 +02:00
Dai Ha fdfd4ac491 Merge CB-600: make installing the launchd agent safe
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
Three gaps that only bite once the agent is loaded, plus one wrong
comment.

The script computed its log path from its own location while the plist
hard-codes one. Run from a different checkout, every post-restart check
would read the wrong file and report a clean restart while the daemon
crash-looped. It now compares the two and fails, not warns.

A failed 'launchctl load' after a successful 'unload -w' left the agent
stopped AND persistently disabled - worse than before the redeploy. It
now retries once, then dies naming the exact recovery command.

The plist now says plainly that ThrottleInterval paces restarts but does
not bound them, and what actually stops the loop.

Verified here: ran the script with --check from the merged tree and it
behaves exactly as before, so the unsupervised path - my only restart
route - is intact. Exercised the log-path check against match, mismatch
and missing-plist fixtures using a truncated copy with no mutating code
in it: ok/1/1. 829 tests BUILD SUCCESS.

Closes #91
2026-08-16 18:17:11 +02:00
Dai Ha d5dd5639ae Merge CB-602: guard against a config key that never reaches the example
bridged.yaml is gitignored, so bridged.example.yaml is the only
committed description of the config schema. Two tests already covered
example -> code; nothing covered code -> example, so a brand-new key
could ship undocumented and no test would notice.

A new test compares BridgedConfig.KNOWN_TOP_LEVEL_KEYS against the
example scanned as TEXT, so a key documented only as a comment counts as
documented. That is what makes the guard correct rather than annoying:
most of the example is commented on purpose.

Verified here: added an undocumented key and watched the test fail with
an actionable message naming it; then documented that key as a comment
only and watched it pass. Probe reverted, tree clean.

Closes #96
2026-08-16 18:17:11 +02:00
Dai Ha 7a120b3256 Merge CB-598: per-item reminder counts so backoff-window work is never orphaned
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m39s
The reminder count was one counter per lead per source, carried forward
across ticks. A counter carried forward has no memory of which item it
counted, so work arriving during the backoff window inherited an
already-capped count and was never named in a nudge.

tick() now recomputes each source's count fresh from the minimum count
among the items actually pending, tracked per item. A fresh item keeps
its source eligible; an older capped item still rides along in the text
without spending more budget. decide() is unchanged.

Verified here: read the diff; the bumped set is exactly the set named in
the nudge, and the empty early-return skips the bump. Trial merge onto
main builds 826 tests BUILD SUCCESS, unpiped.

Closes #87
2026-08-16 18:13:29 +02:00
Dai Ha 863d477966 CB-603: make FakeHerdr.calls thread-safe
Background loops call the fake from their own scheduler threads while a
test polls called() from the test thread. The list was a plain
ArrayList, so a nudge landing mid-stream threw
ConcurrentModificationException out of called().

It surfaced while I was verifying CB-598, which nudges more often, but
the race is on main today and is unrelated to that change.

824 tests, BUILD SUCCESS.
2026-08-16 18:12:23 +02:00
Dai Ha cec48832be CB-600: make it safe to install the launchd agent
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m26s
- redeploy-bridged.sh now refuses (not warns) a supervised restart when
  its computed log path disagrees with the loaded plist's StandardOutPath
  — otherwise every post-restart check reads the wrong file and can
  report a clean restart while the daemon crash-loops. The check is a
  pure, testable function; the script gained a source-for-test guard so
  it can be exercised without installing the agent or touching launchd.
- a failed 'launchctl load' after a successful 'unload' now retries once
  and, on ultimate failure, tells the operator the agent is stopped AND
  disabled plus the exact recovery command, instead of leaving that
  silently worse than the pre-redeploy state.
- the plist documents honestly that the crash loop launchd retries is
  unbounded (ThrottleInterval only paces it), and what actually stops it.
- fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup
  throw is ~370 lines below its call site, not a few lines above it, and
  only fires in auth.mode: token.
2026-08-16 18:08:58 +02:00
Dai Ha a36b7ccd7c Merge CB-601: make the recovery-race test deterministic
CI / build (push) Successful in 1m5s
CI / contract (push) Successful in 1m7s
The test asserted one of two interleavings that are both correct, and
steered toward it with a 5 ms Thread.sleep. Under load the other
interleaving happened and main went red on a correct implementation.

The head start is now a latch counted down from inside the sweep's
guarded loop, so the ordering is guaranteed, not likely. Test file only;
the production guard is unchanged.

Verified here: 822 tests BUILD SUCCESS; 10/10 passes while a full clean
install ran in parallel; 5/5 failures with the guard removed, so the test
still catches the bug it exists for.

Closes #95
2026-08-16 18:08:18 +02:00
Dai Ha 32bf324a1e CB-602: guard against a config key that never reaches the example
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m24s
BridgedConfig.KNOWN_TOP_LEVEL_KEYS is now package-private so a test can assert
every key the parser accepts appears in bridged.example.yaml — live or
commented-out, since the file is gitignored and the example is the only
committed description of the config schema. The existing tests only checked
the example->code direction; this adds code->example.
2026-08-16 18:06:36 +02:00
Dai Ha 16de9df000 CB-601: make the recovery-race test's head start deterministic, not a sleep
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m9s
2026-08-16 18:02:57 +02:00
ltms 65c9deb4d1 Merge CB-597: correct bridged.example.yaml, including two knobs that do nothing
CI / contract (push) Successful in 42s
CI / build (push) Successful in 2m2s
My premise for this ticket was wrong and the worker corrected it. I reported five whole sections missing from the example; nothing was missing. My comparison script only counted uncommented lines, so every section documented as a commented-out example looked absent. Earlier tickets had each updated the example alongside their feature.

What it found instead is more useful than what I asked for — real inaccuracies, found by tracing each field through the parser and its consumers:

- `health.workingSuspectAfterSeconds` and `paneProbeIntervalSeconds` documented enforced minimums that do not exist. I checked: both names appear **only** in the `Health` record declaration and are read by nothing. Only `intervalSeconds` is clamped, and it is silently raised to 15 rather than rejected.
- `notifications.mode: webhook` only flips what `bridge_list` reports as `healthCoverage`. It sends no webhook — "webhook" appears in one `configured()` boolean and there is no delivery code in the repo.
- `lifecycle.clearAfterTurn` was undocumented, and is a no-op for any peer kind other than claude-code.
- The reload doc claimed the whole `fleet:` block is hot; `fleet.leaders` is built once at startup and is not rebuilt, so a change is silently accepted and does nothing until a restart.
- The `fleet.leaders` demotion consequence is now stated next to the block itself: an unmatched pane is silently an ordinary worker and every orchestration call it makes is refused, with no startup error.

Documenting a knob as dead is worth more than documenting it as working. Someone tuning `workingSuspectAfterSeconds` would otherwise have concluded their monitor was broken.

Comments only — no parsing or production code touched. Verified by the lead: parses cleanly under the project's own snakeyaml 1.30, top-level live keys `[bind, herdrSocket, profiles, placement, fleet, guard]`, the rest correctly commented examples. Both dead-knob claims verified by grep against `src/main` rather than taken on the worker's word.
2026-08-16 17:59:41 +02:00
ltms 28ae27b8e1 Merge CB-599: a capacity refusal now tells the caller why
CI / build (push) Failing after 1m19s
CI / contract (push) Successful in 1m25s
`PlacementException extends IllegalStateException`, and neither spawn path caught that type, so it escaped to Javalin's default handler as a bare `500 Server Error` with a text/plain body — while every other failure on the same endpoint returned structured JSON. The reason existed and was good, but only in the daemon log.

I hit this live while orchestrating: asked for a member on a full profile, got a blank 500, guessed another profile, got a blank 500 again, and spent two round trips learning things the daemon already knew.

Both surfaces now catch it. REST returns 503 with `{"error":"no_capacity","detail":...}`; MCP returns the same reason in the `isError` shape it already uses for every other spawn failure. 503 is right because the request was valid and will likely succeed later — the caller did nothing wrong, so 400 would have been a lie.

The other throw sites all funnel through the same type, so quarantine cooldowns, weight-0 exclusion, and the all-at-cap / all-quarantined / all-unreachable messages now reach callers too. That last group matters most: those three distinguish "wait a moment" from "your backends are gone", and all three used to arrive as the identical blank 500.

Tests assert the caller can read the *reason*, not merely that the status changed — one per surface.

Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 824, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

No exception message was reworded. This change delivers messages that were already written.
2026-08-16 17:54:44 +02:00
Dai Ha 81a0cf4710 CB-598: track reminder counts per pending item, not per lead per source
CI / build (pull_request) Successful in 1m12s
CI / contract (pull_request) Successful in 1m11s
Work that arrived during the ~15s push_backoff_ms window between two
ticks landed in the pending map before the next tick's start-of-tick
snapshot, so a shared per-lead-per-source counter (carried forward via
scheduleNext(lead, count+1, ...)) already treated it as exhausted
backlog even though no nudge had ever named it. ReplyPushLoop.tick now
recomputes each source's reminder count fresh every tick as the
minimum nudge count among that source's currently pending items, so a
freshly-arrived item (count 0) keeps its source eligible regardless of
how depleted an older, still-undrained sibling's count is. decide()
itself is unchanged.
2026-08-16 17:49:21 +02:00
Dai Ha 8a837a2830 CB-599: surface a capacity refusal's reason instead of a bare 500
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m35s
PlacementException extends IllegalStateException, which neither BridgedApp
nor BridgeMcp's spawn catch blocks handled, so a maxLoad/quarantine/
all-exhausted refusal fell through to a blank 500 on REST and lost its
message on MCP. Catch it on both surfaces, ahead of the unrelated
IllegalArgumentException(unknown_profile) mapping, and return its message
structured: REST as {"error":"no_capacity","detail":...} with status 503,
MCP as an isError result prefixed "no capacity: ...".
2026-08-16 17:43:07 +02:00
Dai Ha 0a2b3a4a56 CB-597: fix inaccuracies the example config already had, none actually missing
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Failing after 1m21s
Audited bridged.example.yaml against BridgedConfig's KNOWN_TOP_LEVEL_KEYS and
found every top-level key already documented (broker, health, lifecycle,
configReload, quarantineCooldownSeconds, fleet.leaders/architects/reviewers
included) — CB-573/CB-566/CB-559/CB-579/CB-527/528 each updated the example
alongside their feature. What was actually wrong:

- health.workingSuspectAfterSeconds/paneProbeIntervalSeconds claimed enforced
  minimums (300/60) that don't exist in code — only intervalSeconds is
  clamped (floor 15); the other two are parsed but never read anywhere.
- notifications.mode: webhook was undocumented as only flipping the
  healthCoverage label bridge_list reports — no webhook is ever sent.
- lifecycle.clearAfterTurn was missing entirely.
- the HOT bullet under configReload claimed the whole fleet: block reloads
  live, but ConfigRef's own javadoc carves out fleet.leaders as needing a
  restart with no deferred-list warning — added that exception.
- fleet.leaders' demotion consequence (unmatched tab -> silent WORKER
  demotion, no startup error) is now stated inline next to the block, not
  just implied by the multi-lead rationale higher up.
2026-08-16 17:37:28 +02:00
ltms 613ece92dc Merge CB-594: make supervision and a working fleet possible at the same time
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m45s
The launchd unit was a CHANGEME template that had never been installed, and it could not have worked if it were: launchd does not source a login shell, so the daemon would have started with no forge or gateway token, and the failure would only appear much later as workers unable to open a PR.

Four parts: a wrapper that execs one login shell in place so the job inherits the secret store; a startup report naming which required secret env vars resolved and which are MISSING, by name only, never a value; the plist filled in with this host's real verified paths; and `redeploy-bridged.sh` detecting the agent and switching stop/start to `launchctl unload -w` / `load -w`, falling back to the existing kill + nohup when it is not installed.

Verified by the lead. My own unpiped build of the branch: 813 tests, BUILD SUCCESS. Trial-merged onto current main (which had moved twice) and rebuilt: Tests run: 822, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. All three host paths in the plist exist; no CHANGEME remains; the secret report leaks no values.

An independent reviewer, briefed only from the diff, verified the parts that matter today and said merge. It confirmed live that the unsupervised path is unchanged, that `launchctl list` exits 113 for an absent agent so the detection reads correctly, that the wrapper round-trips arguments containing spaces and quotes and stays a single exec chain, and that neither branch of the secret report can print a value.

Its two open findings only bite once the agent is actually loaded, which has not happened and is the operator's call. Filed as #91 — that must land before the agent is ever installed. The important one is that the script computes its log path from its own location while the plist hard-codes an absolute one; if they ever disagree, the post-restart ERROR check reads the wrong file and reports "ok" while the daemon crash-loops.

The agent is deliberately NOT installed by this merge. Nothing here changes how the daemon runs today.
2026-08-16 17:32:57 +02:00
ltms 8d4206c2b5 Merge CB-590: one nudge schedule per lead, with a reminder budget per source
CI / contract (push) Successful in 1m7s
CI / build (push) Failing after 1m15s
Carries PR #84 (its head `78ca24d` is an ancestor of this one), so this single merge delivers both rounds.

CB-590 collapses the CB-307 reply schedule and the CB-588 ticket schedule into one per lead, which closes the double-injection race. The review round found that this also collapsed the two reminder caps into one shared budget — so a busy reply stream could exhaust the cap and a ticket arriving afterwards would never be nudged at all. That was a regression, not a pre-existing wart: before this PR the two sources had independent counters.

The follow-up keeps the single schedule and gives each source its own budget. `decide()` returns INJECT while either source has pending work under its own cap, and STOP only when neither does. `stopOrRestart` and its snapshot-diff race logic are untouched.

Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 814, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. `ReplyPushLoopTest`: 41 tests.

Known and deliberately out of scope: an item arriving during the 15s backoff is already inside the tick's "before" snapshot, so at-cap work can be abandoned rather than nudged once. That shape predates this PR in the ticket-only loop and is filed as #87 (CB-598).
2026-08-16 17:28:57 +02:00
ltms 03286a589b Merge CB-527/CB-528: bound AMQP prefetch, confirm publishes, and close the recovery race
CI / contract (push) Successful in 1m4s
CI / build (push) Failing after 1m18s
Carries PR #83 (its head is an ancestor of this one), so this single merge delivers both rounds.

CB-527 bounds prefetch. CB-528 makes a publish wait for its broker confirm, so a failed publish is never reported as success, and correlates a `mandatory` Return back to the right publish.

The review round on #83 found one real race: `failPendingPublishesOnRecovery` swept the pending maps without holding `publishChannelLock`, so a publish issued on the already-recovered channel could be failed by the sweep — the exact inversion of what CB-528 exists to prevent, and undetectable by dedup because `MessageService.reply` mints a fresh msgId per call. Fixed by taking the same lock. `close()` now fails in-flight publishes promptly instead of letting them time out after 10s.

Verified by the lead, not taken on the worker's word:
- `mvn -f bridged/pom.xml clean install` unpiped: Tests run: 809, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
- `mvn test -Pcontract -Dtest=AmqpReplyInboxContractTest` against a real broker: Tests run: 8 — BUILD SUCCESS

That second run mattered: #83's own new tests are contract-tagged, so the default suite and CI prove nothing about them.
2026-08-16 17:27:21 +02:00
Dai Ha 88b9503c3b CB-590 follow-up: give each nudge source its own reminder budget
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m46s
decide(lead, reminderCount) shared one counter across the reply and
ticket sources after PR #84 collapsed both onto a single per-lead
schedule. A reply stream that used up the whole budget could then
make decide() STOP even for a ticket that had never been nudged and
had coalesced onto the same still-active schedule — stranding it with
no live schedule left, since stopOrRestart's racedIn check does not
save work that was already present in the "before" snapshot.

decide() now tracks a per-source count (replyReminderCount,
ticketReminderCount) and returns INJECT while either source is still
under its own cap, STOP only when both are exhausted. Still exactly
one schedule per lead; stopOrRestart's snapshot-diff logic is
untouched.
2026-08-16 17:26:07 +02:00
Dai Ha c553d795d8 CB-528: close the recovery race in AmqpReplyInbox
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m20s
failPendingPublishesOnRecovery walked and cleared pendingBySeq/pendingByMsgId
without holding publishChannelLock, so a publish() that registered while the
sweep was still iterating could be failed even though it published on the
already-recovered channel — a successful publish reported as failed, and
since MessageService.reply mints a fresh msgId per retry, dedup can't catch
the resulting duplicate. Guard the sweep with publishChannelLock: publish()
only holds it for the seq/map-put/basicPublish, so the sweep can only ever
wait for an in-flight basicPublish to return, never a broker round trip.

Also make close() fail in-flight publishes immediately with a clear message
instead of leaving them to idle out the 10s confirm timeout, and record the
(currently unreachable) msgId-uniqueness assumption pendingByMsgId relies on.

AmqpReplyInboxRecoveryRaceTest drives the sweep and a real publish() against
each other directly (no live broker reconnect) using Proxy-backed fake AMQP
channels and a large in-flight backlog to make the race window observable;
confirmed it fails without the guard (reply "fresh" wrongly failed as
"connection recovered mid-publish") and passes with it.
2026-08-16 17:24:43 +02:00
Dai Ha 3ba6d6784c CB-594: make supervision and a working fleet possible at the same time
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m40s
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets
WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd
never sources secrets.sh itself). bridged now logs at startup which required
token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a
hand-written list) resolved or are MISSING, by name only. Fills in the real
paths in deploy/dev.ltms.bridged.plist for this host and points it at the
wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and
uses launchctl unload/load instead of a raw kill+nohup, because a bare
SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false
reads as a crash and would race the script's own restart; --check reports
installed/loaded state and stays read-only.
2026-08-16 17:13:17 +02:00
Dai Ha 78ca24dc3f CB-590: collapse the CB-307 and CB-588 nudge schedules into one per lead
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m33s
Both reply-queued and ticket-terminal nudges could independently decide
to inject into the same lead pane in the same window, since they ran as
two separate schedules keyed differently (worker target vs. lead) that
never checked each other. Replace both with a single per-lead schedule
(activeLeads) that drains pending reply targets and pending tickets
together, sends at most one combined nudge per tick, and shares one
reminder cap across both sources — so two injections into the same pane
can no longer overlap, and work queued while the lead is busy is never
lost, only deferred.
2026-08-16 16:56:04 +02:00
Dai Ha 4fa6553db5 CB-527/CB-528: bound AMQP prefetch and confirm publishes before claiming durable
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m15s
CB-527: basicQos(prefetch) on the consume channel before basicConsume, configurable
via broker.prefetch (default 32), so an undrained inbox backlog stays on the broker
instead of growing the JVM heap without limit.

CB-528: publish moves to its own confirm-mode channel with mandatory=true and a
return listener, so an unroutable or unconfirmed reply now throws instead of
vanishing silently. The confirm callback checks the per-message returned flag
(set by the return listener, which the broker always fires before the matching
confirm) so an acked-but-returned publish is still reported as a failure. The ack
path stays on its own channel/lock and never waits on a publish confirm.
2026-08-16 16:55:40 +02:00
Dai Ha 2124e043ce CB-593: correct the member MCP claim — measured, not assumed
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m19s
CLAUDE.md told every member 'You mount only the bridge MCP' and told the lead
'a worker mounts only the bridge MCP and cannot run your other tooling'. Both
were false for Claude Code members.

Measured by spawning one member per backend and asking each what it actually has:

  opencode (gx)      11 bridge tools only                     claim TRUE
  claude-code (local) 11 bridge + 45 gitea + 2 context7       claim FALSE

Source is ~/.claude.json user-scope mcpServers; --mcp-config adds to that scope
rather than replacing it, so the worktree parity overlay (which correctly
neutralises .mcp.json and opencode.json) cannot see or stop it.

The forge tools are mounted but not usable: CB-592's blocked sentinel means
get_me and list_issues both fail with 'invalid username, password or token'.
That is defence in depth working in a path it was not designed for, so the text
now says a mounted tool is not a working tool rather than pretending the tools
are absent.

Block propagated byte-identically to wiki 7-Use-Cases.md (wiki 1d95e3f).
2026-08-16 09:46:42 +02:00
Dai Ha e4f3620acb CB-591: record the final 86400s route timeout and the request:0s trap
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m19s
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after
the silent-truncation risk was discussed. They tried request: 0s first: it
removes the total-duration timer, but on an AIGatewayRoute the idle timeout
is derived from the request timeout, so 0s also removed any bound on a
stalled connection.

At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a
runaway request now ends as a clean FAILED ticket we raised instead of a
silently truncated 200. While the gateway sat at 1800s the two numbers were
equal and did not nest.
2026-08-15 21:29:49 +02:00
Dai Ha 032a59a34d CB-591: fleet moved onto the gateway — both ceilings fixed and re-verified
CI / contract (push) Successful in 40s
CI / build (push) Successful in 1m19s
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.

systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:

    listener buffer    32 KiB -> 32 Mi   ours: 1.2 MB body -> 200 (was 413)
    LLM route timeout  60s    -> 1800s   ours: 101s stream -> 200,
                                               message_stop present, 4000/4000

Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.

Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.

§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix #42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.

The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.

Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.

bridged.yaml carries the same notes inline (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 20:47:41 +02:00
Dai Ha e689090024 CB-591: correct the root cause — Envoy's buffer limit, not Caddy
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m35s
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.

Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:

    listener default/llm/http    per_connection_buffer_limit_bytes: 32768

Nobody chose 32 KiB; it was inherited.

Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.

Consequences recorded in the doc:

  * DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
    profiles to it would be designing around a bug.
  * Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
    deployed, pending their operator's approval. We do not re-test until they
    confirm — a half-changed system gives a number neither side can trust.
  * In standalone `aigw run` a SecurityPolicy is accepted and then silently
    ignored, so "the config was accepted" proves nothing there. They will verify
    by re-reading the live config_dump and sending a large request. Same
    silent-default shape this repo keeps hitting, one layer down.

bridged.yaml carries the same correction (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 19:47:01 +02:00
Dai Ha 1cc34888fd CB-591: record the live result — blocked by a 32 KiB body limit at the gateway
CI / contract (push) Successful in 44s
CI / build (push) Successful in 55s
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.

llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:

    /v1        32695 bytes -> 200        /anthropic  32095 bytes -> 200
    /v1        32795 bytes -> 413        /anthropic  32855 bytes -> 413

That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.

The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.

opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.

Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.

Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.

Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.

Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.

Refs: gitea #76
2026-08-15 19:27:45 +02:00
Dai Ha 0331ecd5d3 CB-592: add the BRIDGED_MEMBER marker — the sentinel alone cannot hold
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m39s
Live check on a member pane showed the CB-592 shadow did NOT take effect:
GITEA_ACCESS_TOKEN inside the pane was still the real admin token.

Measured cause. The overlay itself works — GITEA_TOKEN is injected the same
way, is exported by no shell file, and does reach the pane. The sentinel loses
one step later. A herdr pane runs a LOGIN shell, ~/.zprofile line 41 sources
${SHARED_ENV}/tools/secrets.sh, and that file does a plain unconditional
`export GITEA_ACCESS_TOKEN=...`. A login shell overwrites a value already in
the environment, so the real token is put back before the member starts.
Confirmed directly:

    GITEA_ACCESS_TOKEN=cb592-sentinel zsh -lc ...
    -> RESULT: sentinel was OVERWRITTEN by the login shell

This defeats any launcher-side overlay for any name secrets.sh exports. No
change in this repo can win it alone.

So this adds the half that does survive: BRIDGED_MEMBER=1, a name secrets.sh
never exports. It is a no-op until the operator guards the export:

    [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...

Setting it now costs nothing and makes that one line the whole remaining fix.
The sentinel stays: it is correct for any peer kind whose pane does not start
a login shell, and it keeps the intent explicit where every adapter passes.

Also corrects the javadoc and the test javadoc, which both claimed a
protection that was measured not to hold.

The other reported failure was my own bad test, not a regression. The probe
called /api/v1/user, which a minimal write:repository token cannot read. Same
token on the repo endpoint answers 200, so CB-302 is intact:

    GITEA_ACCESS_TOKEN: /user=200  /repos/lms/claude-bridge=200
    WORKER_GITEA_TOKEN: /user=403  /repos/lms/claude-bridge=200

Tests 805 -> 807. Both new tests proved to discriminate by reverting the
marker: everySpawnMarksThePaneAsAMember and
aProfileEnvEntryCannotClearTheMemberMarker both fail without it.

Refs: gitea #77
2026-08-15 18:37:52 +02:00
Dai Ha 831a918c30 Merge CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's environment
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m22s
A live probe showed every spawned member carried the admin GITEA_ACCESS_TOKEN: 108
environment variables in a member's pane against 99 in the primary's. The operator's
rule is that only the leader and architects may use it; everyone else uses
WORKER_GITEA_TOKEN. We were not enforcing that at all.

The cause is invisible from inside the launcher. baseEnv builds a fresh map holding only
PATH and the profile's env:, so a member looks like it gets a small explicit environment.
That map is an OVERLAY: WorkspaceControl.createTab/splitPane send only the keys it
contains, and herdr spawns the pane from its own login-shell environment, so every key we
never mention passes straight through — admin token included.

The fix puts a non-blank sentinel over the key in baseEnv, applied AFTER the profile's
env: so no profile, present or future, can restore the real token by naming it in config.
One place, every adapter, including peer kinds not yet written — deliberately not a
per-profile bridged.yaml entry, which is the silent-default shape this repo has shipped
nine times.

A non-blank sentinel rather than the empty string, on purpose: whether an empty overlay
value overrides an inherited variable or is skipped as blank cannot be settled from this
repo, because herdr's merge happens in an external process. baseEnv's own PATH seeding
(CB-511) already relies on a non-blank value replacing an inherited one, so this reuses
the shape that is demonstrated to work rather than the one that is merely plausible.

CB-302's repo-scoped GITEA_TOKEN grant is untouched — a worker can still open its own PR.
The subscription boundary was checked and is unaffected: the primary's pane carries no
ANTHROPIC_* at all, so nothing is inherited there.

Closes gitea #77. Live verification follows separately: the daemon must be redeployed
before this reaches any pane.
2026-08-15 18:27:05 +02:00
Dai Ha 3db5277ae8 CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's herdr overlay
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m37s
herdr spawns a pane from its own login-shell process env and layers our map on
top, so any key baseEnv never mentions passes straight through — including the
admin forge token. baseEnv now puts a non-blank sentinel for
GITEA_ACCESS_TOKEN, applied after the profile's own env: so no profile can
restore it. One place, every adapter, every profile including future ones.
CB-302's GITEA_TOKEN grant (applyGitToken) is untouched.
2026-08-15 18:25:16 +02:00
Dai Ha 6939e0cbbc CB-592: the tracked opencode.json must name the worker forge token, not the admin one
Operator's rule, 2026-08-15: only the leader and architects may use GITEA_ACCESS_TOKEN;
everyone else uses WORKER_GITEA_TOKEN.

opencode.json is TRACKED, so it ships in every worker worktree, and it mounted the gitea
MCP with {env:GITEA_ACCESS_TOKEN}. A live probe confirmed that variable actually resolves
inside a member: herdr spawns each pane from its own login-shell environment and layers
the launcher's map on top, so a member sees 108 variables rather than the small explicit
set baseEnv appears to build. That gave an opencode member admin forge TOOLS — enough to
merge its own PR, which both CLAUDE.md and the member contract forbid.

This is the narrow half of the fix: it removes the tooling. The admin token is still
present as a string in every member's environment, which is the real defect and is
tracked as CB-592 (gitea #77) — that fix belongs in the launcher, in one place, not
per-profile in bridged.yaml where a sixth profile would silently reopen it.

.mcp.json keeps GITEA_ACCESS_TOKEN and is correct to: it is skip-worktree, the primary's
own local copy, and the primary is the lead. That is the pattern this change follows —
the shared tracked file grants least privilege, and anything needing more overrides
locally.
2026-08-15 18:20:55 +02:00
Dai Ha f0095bf8b2 CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m38s
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.

The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.

Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
2026-08-15 17:06:31 +02:00
Dai Ha 5206679efd Merge CB-588: nudge the lead when an async ticket goes terminal
An async delegation ticket (bridge_send wait:false) resolves on MessageService.reply's
rendezvous fast path, which returns before onReplyQueued. So CB-307's push loop only ever
heard about the durable-inbox case, and the mode CLAUDE.md tells leads to prefer never
nudged anyone. Closes gitea #72.

Adds a second, independent reminder schedule keyed by the lead terminal, so several
tickets finishing together coalesce into one nudge. The CB-307 path is untouched.

Three defects were found in review and fixed before merge:
 * a pendingTickets entry outlived the ticket it named. poll() returns null once
   pruneTerminalTickets drops a ticket, so ticketCollected was never reached and the
   entry leaked for the daemon's life, riding along on every later nudge and sending
   the lead after a ticket bridge_poll can no longer find.
 * a lost nudge: a ticket landing between decideTickets returning STOP and
   activeLeads.remove coalesced onto a schedule that was about to die. That is the
   exact failure this ticket exists to remove, reintroduced in a narrow window.
 * the success direction was unpinned in tests, and the comment listing the paths that
   complete the future was short by several.

The obvious fix for the second one was wrong: restarting on any pending ticket defeats
the reminder cap, because a never-collected ticket at cap is expected to still be there.
The fix diffs against a snapshot taken before the decision, so only a ticket that truly
arrived during the window restarts the schedule.

Verified on my own unpiped build: 802 tests, 0 failures, BUILD SUCCESS.
Two reviewers on the diff; the loop-gating finding they raised is split out as CB-590.
2026-08-15 17:06:12 +02:00
Dai Ha ac044e7573 CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m31s
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
  slot on STOP, restart only if a ticket landed that the pre-decision snapshot
  did not already account for. A naive "restart on any pending ticket" version
  was tried first and reverted: it defeated the reminder cap by restarting
  forever on a stale, never-collected ticket (broke
  successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
  Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
  from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
  (now package-private) rather than forcing the underlying thread race:
  aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
  aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
  Both were verified to fail against deliberately-reverted versions of the fix
  before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
  (must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
  (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
  on throw, and answer() -> finishAsyncTask(turnId, result)).
2026-08-15 17:02:05 +02:00
Dai Ha 6ebad2a91f CB-588 follow-up: reclaim a pruned ticket's pendingTickets entry too
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m0s
tasks is the sole authority on whether a ticket exists, but
pruneTerminalTickets dropped entries from it without telling
ReplyPushLoop. ticketCollected(ticket) was only ever called from
MessageService.poll's terminal branch, which pruneTerminalTickets
short-circuits past once a ticket is gone (poll returns null at the
top). A ticket the lead never polled — or one the reminder cap already
gave up on — was pruned from tasks but never collected in
ReplyPushLoop.pendingTickets, so it rode along on every later nudge to
the same lead forever, naming a ticket bridge_poll could no longer
find, and the map itself never shrank.

pruneTerminalTickets now calls pushLoop.ticketCollected for every
ticket it actually removes (guarded on pushLoop != null), so a pending
nudge entry lives exactly as long as its ticket is pollable. Reused
tasks as the only removal trigger rather than adding a second live
query back into MessageService — no new source of truth.

Added an injectable clock (LongSupplier nowNanos, defaulting to
System::nanoTime) to MessageService, mirroring the SessionManager/
SessionReaper nowNanos seam, so a test can cross the 10-minute
TICKET_TTL_NANOS deterministically instead of sleeping for real.
Confirmed the new regression test fails against the prior
pruneTerminalTickets (a stale ticket rides along on a later coalesced
nudge) before restoring the fix.
2026-08-15 16:41:23 +02:00
Dai Ha 3d10ed385c CB-588: nudge the lead when an async ticket reaches a terminal phase
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m46s
An async bridge_send(wait:false) registers a rendezvous waiter, so its
reply always takes MessageService.reply's fast path and returns before
ReplyPushLoop.onReplyQueued is ever called — the exact mode the charter
tells leads to prefer never nudged.

Add ReplyPushLoop.onTicketTerminal(ticket, target, failed), a second
entry point reusing the loop's status gating, bounded/backoff reminders
and metrics, keyed by the nudge-receiving lead so several tickets
finishing together coalesce into one nudge naming the count. MessageService
wires it via task.future.whenComplete in sendAsync (covers reply,
completion fallback, wedge, and abandon() alike) and calls the new
ticketCollected(ticket) from poll() once a terminal view is handed back,
so an already-collected ticket is never nudged again. The CB-307 inbox
path (onReplyQueued/decide/NUDGE_FORMAT) is untouched.
2026-08-15 16:10:22 +02:00
Dai Ha 6d0c94dbdb Correct enforceMaxLoad's comment after CB-585
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m16s
The comment said non-positive means unlimited at load. That stopped being
true when CB-585 made an explicit maxLoad: 0 survive as a real cap of zero
and made a negative value refuse config load. The code below it was already
right — only the comment described the old normalisation. Flagged by the
CB-585 worker, which correctly stayed out of a file not on its list.
2026-08-15 16:02:35 +02:00
Dai Ha e01563a550 Merge cb584-session-resume-b62247-4
CI / build (push) Successful in 1m20s
CI / contract (push) Successful in 1m44s
2026-08-15 15:58:29 +02:00
Dai Ha 30e3225a3c Merge cb585-maxload-zero-4585c4-3 2026-08-15 15:58:29 +02:00
Dai Ha ef186a1516 Merge cb587-snapshot-index-flags-24e236-2 2026-08-15 15:58:29 +02:00
Dai Ha 5d5b3bdc76 CB-584: persist agentSessionId at acquire and expose it via bridge_spawn/bridge_list
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 2m18s
Wires up the two links that made session resume unreachable: MemberSession
now records agentSessionId from PeerHandle at acquire (both the plain and
worktree paths), and it survives onto the roster (rosterView, so both
bridge_list and GET /members show it). bridge_spawn accepts sessionName and
resumeSessionId, threading them into the SpawnRequest fields that already
existed but were never reachable from the MCP surface.

A resumeSessionId now requires an explicit profile (a resumed conversation
is tied to the specific backend that started it, so an unqualified spawn
routed by placement has no safe candidate to check) and is refused, naming
Capability.SESSION_RESUME, when that profile's adapter does not declare it
— PeerLauncher gains capabilitiesFor(profileName) so a mixed fleet is
checked per-adapter rather than against the fleet-wide capability union.

PeerHandle.agentSessionId() loses its default, the same fix CB-571 (c00a86b)
made for charterReceipt() one method above it — both existing adapters
already overrode it, so this only closes the landmine for a future one.

Out of scope, left for follow-up: carrying agentSessionId on a failed
ticket's ReleaseDetail (issue #65 criterion 5) — CB-578 stage C just
landed in that file and this ticket deliberately stayed out of it.
2026-08-15 15:48:06 +02:00
Dai Ha 2f8c98dac9 CB-585: maxLoad: 0 caps a profile at zero, negative refused at load
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m18s
2026-08-15 15:38:50 +02:00
Dai Ha 2db7189067 CB-587: seed the snapshot's temp index from the worktree's real index
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m24s
--skip-worktree is an index flag. GitWorktrees.snapshot staged into a fresh
empty temp index, which carried none of the real index's skip-worktree bits,
so add -A staged local on-disk content for files git status correctly hides
(e.g. .mcp.json). Copy the real index (resolved via git rev-parse --git-path
index, correct for linked worktrees) into the temp index before staging, so
add -A skips exactly what git status skips.
2026-08-15 15:36:06 +02:00
Dai Ha 94476ac109 CB-529: drainReplies javadoc said the ack is local — false for AMQP
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m5s
The doc claimed an in-flight failure re-surfaces the messages on a later
drain, because the ack is local. That was true when InMemoryReplyInbox was
the only inbox, and became false without anyone noticing when the AMQP
adapter landed: there the ack is a broker-side basicAck, so a crash while
writing the response loses the reply outright. Re-polling cannot recover it,
since the broker has already forgotten it.

Behaviour is unchanged and the window stays accepted — acking on the next
poll instead would double-deliver on every normal drain. The point is that
the comment sat exactly where the next implementer would read it and said
the opposite of what happens.

Closes #12.
2026-08-15 15:34:59 +02:00
Dai Ha 5fe02b7c98 Fix redeploy script reporting failure on a successful restart
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
Found by running it. The script launched the daemon with a relative jar path
(cwd is bridged/) but detected it with an absolute one, so pgrep never matched.
The daemon restarted correctly and booted clean, and the script still failed
with 'no process appeared' — the worst shape of bug for a deploy tool, because
it invites a second restart on a daemon that is already healthy.

Detection now matches both path forms, and the launch uses the absolute path so
ps names which checkout is running.
2026-08-15 15:18:44 +02:00
Dai Ha a1052f4fd1 Add scripts/redeploy-bridged.sh so the lead can deploy in one command
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m23s
The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.

It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.

The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
2026-08-15 15:14:21 +02:00
Dai Ha 8c9904a7c4 Record that the classifier still blocks the daemon kill
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
A CLAUDE.md rule grants intent, not tool permission. Tried the redeploy right
after writing the section and the classifier refused `kill`, so the section
would have been misleading as written. States the honest split until a Bash
permission rule exists: the lead builds and verifies, the operator runs the
stop and start with `!`.
2026-08-15 13:57:18 +02:00
Dai Ha 23af5dfdff The lead may redeploy the daemon — write down the procedure and its traps
A merge is not a deployment: the running bridged holds the jar it started
with, so merged code does nothing until the daemon is rebuilt and restarted.
Calling that work shipped is a false report. This makes the redeploy the
lead's job rather than something handed back to the operator, and records the
five things that have gone wrong doing it here — chiefly that starting the
daemon from a non-login shell empties WORKER_GITEA_TOKEN, which nothing logs
and which only surfaces later as workers that cannot open a PR.

Goes in the project addendum, not the canonical block; the block is unchanged
and still byte-identical with the wiki template.
2026-08-15 13:54:38 +02:00
Dai Ha 3a10f6ad17 Note that stage C's snapshot makes the parityOverlay gitignore rule load-bearing
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m19s
CB-578 stage C commits a preserved dirty worktree to refs/wip/<branch> with
git add -A. The parity overlay copies the primary's environment files into
every worktree, so a non-gitignored overlay path now reaches a durable git
object instead of only sitting on disk. add -A respects .gitignore, which is
what stops it — so the rule documented for CB-581 is now what keeps a secret
out of a commit, not just out of a directory.
2026-08-15 13:09:01 +02:00
Dai Ha a3842c873d Merge cb578c-92885c-1 2026-08-15 13:07:50 +02:00
Dai Ha c29c3f063d Merge cb554-e8e224-3 2026-08-15 13:07:50 +02:00
Dai Ha 758d62a396 Merge cb583-b82d11-2 2026-08-15 13:07:50 +02:00
Dai Ha 6725642274 CB-578 stage C: snapshot a dirty worktree into refs/wip before it can be lost
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m21s
Adds Worktrees.snapshot(worktreePath, branch, message): stages into a
temporary GIT_INDEX_FILE (never the worker's real index/HEAD), writes
the tree, commit-trees it onto the worktree's current HEAD, and points
refs/wip/<branch> at the result. add -A (never -f) respects .gitignore.

SessionManager.release() now snapshots any dirty worktree before the
preserve-or-remove decision, regardless of release cause (COMPLETED or
SHUTDOWN) — preserving on disk alone is one `worktree remove --force`
away from gone. A failing snapshot never escalates: the worktree is
still preserved, the pane still stops, and the release listener is
still notified.

onRelease's listener now receives a ReleaseDetail (terminal, worktree
path, branch, snapshot ref) instead of a bare terminal id, so Bridged's
WORKER_FAILED wiring can put the same three facts into a failed
ticket's detail — a lead can re-dispatch onto the same tree instead of
starting from the base commit.
2026-08-15 12:56:12 +02:00
Dai Ha a1f4dc365a CB-554: weight <= 0 excludes a profile from automatic placement, not 1.0
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 52s
2026-08-15 12:54:02 +02:00
Dai Ha 165b62ee20 CB-583: make bridge_list's capacity view quarantine-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m32s
A quarantined profile's capacity row now forces free:0 and names the
quarantine (credentialId, quarantinedForSeconds), reusing the same
QuarantineSource bridge_profiles already reads instead of a second
lookup. An ordinary fleet's capacity rows are unchanged (no new keys).
2026-08-15 12:51:36 +02:00
Dai Ha 72d3481de3 Merge CB-578 stage B: quarantine the exhausted credential, not the profile
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Closes issue #50. Verified by the lead: own unpiped build of the branch merged
onto main — 740 tests, BUILD SUCCESS, exit 0.

Stage A only classified a usage-limit refusal. Nothing acted on it, so the
fleet would spawn another member onto the same exhausted account and fail the
same way. Now a BACKEND_EXHAUSTED classification quarantines the CREDENTIAL
for a cooldown: an explicit spawn onto it is refused with the profile, the
credential and the seconds remaining, placement skips it under every policy,
and it lifts itself on an injected clock.

Keyed by credential, not profile name, because sol and terra are two models on
one OpenAI account. Quarantining only the profile that reported the refusal
would leave its sibling live, and the next spawn walks onto the same dead
account. A profile that sets no credentialId quarantines alone under its own
name, so an existing config behaves exactly as before.

The implementer also found a real bug outside the brief: ConfigRef's
sameLaunchSettings never compared exhaustedPattern, so a reload changing only
that key reported 'config reloaded' while the value is actually deferred —
the precise failure that file's own doc calls the worst outcome a reload can
produce. Fixed with a regression test.

Two things I changed at review. The example.yaml conflict with e09cac6 was
resolved toward the branch, which was a strict superset. And
BackendQuarantine.none() held a clock frozen at 0, so a quarantine() call on
it would have locked a credential out for the life of the daemon — two
production CompositePeerLauncher constructors default to none(). It is now a
real no-op with a test.

Left open deliberately: the implementer's bridge_ask about the credentialId
name timed out after 55s with no answer, because I was polling minutes apart.
It proceeded on its own judgment and chose well. The 55s ask window against an
async lead is a bridge problem, not a worker problem.
2026-08-15 10:33:59 +02:00
Dai Ha 29ccb747ad CB-578 stage B: make BackendQuarantine.none() actually inert
Added at merge review. none() held a clock frozen at 0 with a 1ns cooldown, so
a quarantine() call on it recorded a deadline that could never pass — the
credential would be locked out for the life of the daemon. Two production
CompositePeerLauncher constructors default to none(), so that failure would
have been silent and permanent.

The implementer documented the limitation honestly rather than hiding it, but
a stand-in named none() should not need the caveat. quarantine() is now a
no-op on that instance, with a test asserting it. An inert value must omit the
fact, never invent one.
2026-08-15 10:33:44 +02:00
Dai Ha 4749e27453 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	bridged/bridged.example.yaml
2026-08-15 10:31:41 +02:00
Dai Ha e501d39988 CB-578 stage B: quarantine the exhausted credential, not the profile
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m21s
A BACKEND_EXHAUSTED classification (stage A) now puts that profile's
credential into a BackendQuarantine for a configurable cooldown. A spawn
onto a quarantined profile is refused naming the credential and roughly
when it lifts; weighted/round-robin/fixed placement skip a quarantined
candidate; the quarantine lifts itself on the injected clock; and it is
visible on bridge_profiles.

Keyed by credential, not by profile name, via the new Profile.credentialId
(profiles sharing one credential quarantine together — e.g. two models on
one account) and effectiveCredentialId() (unset ⇒ quarantines alone,
today's behaviour unchanged). Fixed a related gap along the way: a reload
changing exhaustedPattern was silently reported "applied" even though it's
deferred — sameLaunchSettings() now catches it too.
2026-08-15 10:29:54 +02:00
Dai Ha 6b6cf25862 CB-581: neutralise the parityOverlay false-preserve risk, and record the rule
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m32s
Issue #57 criterion 5 asked for an explicit decision on this rather than
silence. It is closer to live than the issue assumed.

The DEFAULT parityOverlay is List.of(".env", ".envrc"), so it applies to every
profile, and neither path was gitignored. Since CB-576 a release preserves any
worktree that git status --porcelain calls dirty, and that deliberately counts
untracked files — the work lost in CB-576 was a file nobody had added. So one
.env at the repo root would make every COMPLETED release preserve its
worktree, and worktrees would accumulate with no error to notice.

Inert today only because neither file exists here. Both are now gitignored,
which they deserve on their own as environment files. The javadoc carries the
rule for the next overlay path: it must be gitignored, or tracked and
skip-worktree'd.
2026-08-15 10:17:26 +02:00
Dai Ha 180de840eb Merge CB-581: a throw in release() no longer orphans the pane or aborts reapIdle
CI / contract (push) Successful in 43s
CI / build (push) Successful in 52s
Verified by the lead: own build of the branch merged onto main — 717 tests,
BUILD SUCCESS, exit 0.

release() ran an unprotected sequence. hasUncommitted shells out to git status
and throws on a non-zero exit, which skipped both notifyReleased and
launcher.stop. The session was already out of the registry, so nothing retried
it: a live pane kept burning a fleet slot while absent from the roster, and any
send blocked on it was never resolved. reapIdle called release bare inside a
loop, so one such session aborted the whole pass and skipped every session
after it.

Now the dirty-check block is guarded and fails toward preserving the worktree —
'we could not tell' must not be treated as 'it is clean', because deleting on a
guess destroys work with no other copy. notifyReleased runs in a finally and
launcher.stop runs unconditionally, so the pane always stops. reapIdle catches
per session, matching the shape drainAll already used.

Checked while reviewing: notifyReleased cannot throw out of the finally — it
already guards each listener and only logs. The pane stop is genuinely
unconditional.
2026-08-15 10:14:41 +02:00
Dai Ha 8fd2d7e5e7 CB-581: fail-safe release() so a throw never orphans the pane or aborts reapIdle
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m1s
hasUncommitted shells out to git and can throw; release() now catches that
inside a try/finally so notifyReleased and launcher.stop always run, and
defaults to preserving the worktree on a throw (can't tell dirty vs clean,
so don't risk deleting unrecoverable work). reapIdle wraps each per-session
release in try/catch, matching drainAll, so one bad session no longer
skips the rest of the reaping pass.
2026-08-15 10:09:32 +02:00
Dai Ha 129dd4a838 M4: record what Unit 2 has actually landed, and correct criterion 15
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m16s
Unit 2 was written as one block, but CB-571/576/580 and unit 2a have since
built parts of it. A symbol survey of main source plus the merge history puts
it at 1 of 21 criteria done, with three partials.

Criterion 15 says a normal COMPLETED release removes the worktree. Since
CB-576 that is false on purpose — a dirty worktree is preserved, because
deleting it destroys work that cannot be recovered. The criterion is wrong,
not the code.
2026-08-15 10:08:38 +02:00
Dai Ha e09cac6f1f CB-578 stage A: document exhaustedPattern as a deferred reload key
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m31s
The patterns are compiled once in Bridged.main from the startup config
snapshot, so adding one to a profile does nothing until a restart. Neither
the per-key docs nor the reload-class table said so.
2026-08-15 10:06:41 +02:00
Dai Ha c00a86b32c Merge CB-571: charter receipt on the roster, with no silent null for OpenCode
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m19s
Verified by the lead: own build of this branch merged onto main — 710 tests,
BUILD SUCCESS, exit 0.

Supersedes PR #54. A reviewer found that OpenCodeLauncher.SessionAwareHandle
wrapped the base WorkerHandle but never overrode charterReceipt(), so it
inherited the interface default of null while the real receipt sat on its
delegate — sol and terra would have shown no charterSource/charterSha256 on
the roster while Claude Code members showed both.

The fix is the root one, not the one-line override: PeerHandle.charterReceipt()
is no longer a default, so the compiler forces every implementation to answer.
This repo had shipped that same class of defect — a defaulted dependency that
compiles, passes tests, and quietly turns a feature off — eight times before
this one.
2026-08-15 10:02:36 +02:00
Dai Ha d3ae0350a2 Merge remote-tracking branch 'origin/main' into fix-charter 2026-08-15 10:01:10 +02:00
Dai Ha 2c2196a1f1 Merge CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m24s
Verified by the lead: own build of the branch merged onto main — 704 tests,
BUILD SUCCESS, exit 0. No vendor wording in any Java source (grep clean); a
profile with no exhaustedPattern keeps today's completion-fallback path exactly.

Accepted the implementer's deviation from the brief. The brief asked for a
terminal HealthState BACKEND_EXHAUSTED. FleetHealth.decide() is a pure
classifier over HealthSnapshot, which carries only booleans, AgentStatus and
MemberSession.State — never pane text. FleetHealth's own javadoc already says
ERROR_ON_SCREEN 'is not decided yet because it needs a bounded pane detection
read'. A second declared-but-unproduced value would repeat a gap the class
already documents as a problem, so the signal was put where the evidence
actually lives: the completion scrape.
2026-08-15 10:00:34 +02:00
Dai Ha 35ade14630 CB-571: make PeerHandle.charterReceipt() abstract, fix OpenCode adapter's silent null
CI / build (pull_request) Successful in 50s
CI / contract (pull_request) Successful in 1m21s
SessionAwareHandle wrapped the base's WorkerHandle but never overrode
charterReceipt(), so it silently inherited the interface default (null)
while the real receipt sat on its delegate. sol/terra never got a
charterSource/charterSha256 roster row.

Deletes the default so every PeerHandle must answer explicitly; the
compiler now catches this class of gap instead of a roster field
quietly going missing.
2026-08-15 09:55:52 +02:00
Dai Ha 541df87272 CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (pull_request) Failing after 59s
CI / contract (pull_request) Successful in 1m10s
A backend that refuses on a subscription usage limit leaves the pane healthy but
the turn ends with no bridge_reply; the completion fallback used to scrape and
hand that refusal back as if it were a real answer. CompletionResolver now
matches the scrape against a per-profile exhaustedPattern (config, never a
vendor string) and resolves the send as Rendezvous.Kind/Outcome.BACKEND_EXHAUSTED
with a reason carrying the matched line, kept distinct from GONE/WORKER_FAILED.
A profile with no pattern configured is unaffected. Coverage is logged at
startup via CompletionResolver.coverage(...), naming which profiles have a
pattern and which don't, following FleetHealthMonitor.coverage's pattern.
2026-08-15 09:55:06 +02:00
Dai Ha 337b6ccd6e Merge remote-tracking branch 'origin/main' into fix-charter
# Conflicts:
#	bridged/src/test/java/dev/ltms/bridged/session/SessionManagerTest.java
2026-08-15 09:52:39 +02:00
ltms 42f46dfe9a Merge CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m4s
Verified by the lead: merged onto main (9088d2b) in a scratch worktree, mvn -f bridged/pom.xml
clean install unpiped — MVN_EXIT=0, Tests run: 696, Failures: 0, BUILD SUCCESS. main alone measures
692, so this adds 4 net tests. Merges cleanly; Bridged.java auto-merged against CB-580.

Reviewed by the lead reading the full production diff and the three test files. The member's own
report was lost to the idle reaper before collection, so there was no author write-up.

Closes the live bug: LeadTabScanner.scan() no longer merges the config pin over the scan result, and
the cache no longer seeds from it, so a lead disappears once its tab is gone. The ghost this fixes
had begun throwing agent_not_found from ReplyPushLoop.decide on every tick, against a dead terminal
that still owned two live members.

Beyond the brief, and correct: `tab` is required for every leader, not only non-creatable ones,
because LeadLauncher also uses it to label a tab it creates. terminal: is rejected by a raw-YAML
check rather than by record shape — the only way to beat @JsonIgnoreProperties(ignoreUnknown = true).
The silent-default trap is avoided: the back-compat constructor still takes `tab` positionally.

Behaviour change worth knowing: the scanner's initial cache is now empty instead of the config pins,
so a herdr failure on the very first scan yields no leads until a scan succeeds. That is unavoidable
once the pins are gone, it fails loudly rather than silently, and it is covered by
aFailedFirstScanReturnsEmptyRatherThanThrowing.

Operator action required: bridged.yaml must replace fleet.leaders.<name>.terminal with tab. Already
done for this deployment.
2026-08-15 09:29:00 +02:00
ltms 9088d2b2c5 Merge CB-580: fail a ticket when its member reaches a terminal health state
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Verified by the lead on 0af902e: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 689, Failures: 0, Errors: 0, BUILD SUCCESS.

Reviewed by the lead reading the diff. The member's own report was lost to the idle reaper before
it was collected, so there was no author write-up to review against.

All three defects that sank 3b2f395 are absent:
  * failTarget is required — one constructor, Objects.requireNonNull, no defaulting overload
    anywhere in the repo, and Bridged.java:365 updated to pass messages::abandon.
  * The failure fires inside reportTransition after its `if (previous == next) return;` guard, so an
    unchanged tick cannot reach it.
  * No AtomicReference; MessageService is passed directly, so there is no empty window.

abandon(String target, String reason) confirmed as CB-568's target-wide operation that resolves
waiters as a failure rather than letting them time out.

Known limitation, accepted and covered by its own test: states.put records the new state before the
bounded retries run, so if all three attempts throw, the tickets stay pending and no later tick
retries. It is logged at WARN, and the retries carry no backoff.
2026-08-15 08:57:03 +02:00
ltms 500bfa2c33 Merge CB-576: release preserves a dirty worktree instead of deleting it
CI / build (push) Successful in 51s
CI / contract (push) Successful in 1m1s
Verified by the lead on b525b0f: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 686, Failures: 0, Errors: 0, BUILD SUCCESS.

Review accepted the required-interface-method shape (no defaulting overload) and the decision to
count untracked files as dirty — the work lost in the incident was a file that was never added.

One blocking defect was found and fixed in b525b0f: hasUncommitted called git with no existence
check, so a missing worktree threw WorktreeException from inside release() after registry.remove()
but before notifyReleased() and launcher.stop(), orphaning the pane and stranding a blocked
bridge_send caller. It now mirrors remove()'s already-gone tolerance.

Waived on merge, tracked as follow-up: a SessionManager-level test that teardown completes when the
worktree is gone, and the stronger fix behind it — reapIdle calls release() with no try/catch while
drainAll wraps it, so any exception in that window aborts the whole reaping pass.
2026-08-15 08:54:33 +02:00
Dai Ha 9ca9c43dfa CB-571: retag charter-receipt references from the taken CB-575
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 54s
CB-575 already names the merged MCP-cancellation-filter change, so the
charter-receipt comments used the wrong number. Retag to CB-571, the number
this work was authored against.
2026-08-15 08:48:57 +02:00
Dai Ha b525b0f08f CB-576: hasUncommitted tolerates an already-gone worktree
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m37s
2026-08-15 08:47:12 +02:00
Dai Ha 1966c69994 CB-575: charter receipt on spawn, in the roster and in the logs
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m34s
Record a CharterReceipt (role, source, sha-256 digest, byte count) for every
launch, store it on the MemberSession, expose it in bridge_list and GET
/members, and log it at spawn as digest+role only. Redact the charter argv
argument in the legacy pane-placement spawn log so the charter text never
reaches the daemon log. The charter prose itself is never recorded.
2026-08-15 08:10:56 +02:00
Dai Ha 976eff8ad1 CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m18s
Leader.terminal -> Leader.tab (exact tab label, case-insensitive match).
LeadTabScanner matches an exact tab->name map instead of stripping a
shared tabPrefix, and no longer merges configured leads into every
scan result -- a stale pin can no longer outlive its tab.
LeadLauncher.tabLabel() returns the configured tab directly; the
terminalId pinned-terminal fallback in liveLeads() is gone.
Config load now rejects a leftover fleet.leaders.*.terminal key
instead of silently ignoring it. primary.terminal is untouched.
2026-08-15 08:05:21 +02:00
Dai Ha 0af902ec43 CB-580: fail a ticket when its member reaches a terminal health state
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m5s
FleetHealthMonitor now requires a failTarget BiConsumer<String,String>
collaborator (no defaulting overload) and calls it exactly once when a
member transitions into GONE or NEVER_READY, via CB-568's idempotent
target-wide abandon() operation. The reason string names the real
terminal state. failTarget invocation retries up to
MAX_FAIL_TARGET_ATTEMPTS (3) within the same transition if it throws,
and never refires on a later tick where the state is unchanged.

Bridged.java wires messages::abandon as the production failTarget.
2026-08-15 07:54:28 +02:00
Dai Ha 9118ce2537 CB-576: release preserves a dirty worktree instead of deleting it
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 1m35s
2026-08-15 07:34:06 +02:00
Dai Ha 2f48e08f1f Merge CB-577 follow-up: drop the target-keyed async index
CI / contract (push) Successful in 43s
CI / build (push) Successful in 55s
asyncTasksByWaiter correlates an async question by the exact rendezvous
waiter, so the target-keyed set it replaced can no longer decide
anything. Keeping it meant two indexes of the same fact, one of them
ambiguous whenever a target has two accepted tickets.

The race the old test modelled by reflection is gone with it: identity
keys make 'some other task reached this target' unrepresentable, so
there is no longer a wrong task for the question to land on.

Also carries the criterion-1 doc correction, which is identical to
aac29d6 on a different parent.
2026-08-15 06:40:31 +02:00
Dai Ha fec284e7cb Merge M4 unit 2a: accepted-turn identity carried to delivery
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m17s
TurnToken, owned by MessageService, binds a target to the exact
rendezvous waiter for one accepted send. Injector.Pending carries it and
the delivery callback hands it to CompletionResolver, so the baseline is
bound to the send it belongs to by construction rather than by a lookup
that could pick a different one.

The callback signature is required, not a defaulted overload: a delivery
with no token is exactly the unbound baseline this unit forbids, so a
default would let a caller silently produce it.

The token deliberately omits the session turn number. MessageService
owns acceptance but never learns of delivery, and
CompletionResolver.onDelivered runs before SessionManager.onDelivered,
so the number does not exist yet at the only point the token could
capture it. docs/M4-Fleet-Health.md criterion 1 records this and the two
rejected alternatives.

Still open for the next slice: the missing/post-restart baseline test and
the no-replay test.
2026-08-15 06:36:46 +02:00
Dai Ha 8e2e4c5e73 M4 unit 2a: migrate test call sites to the required turn token
The delivery callback now requires a TurnToken, so 54 test call sites
had to pass one. They use an explicit TestTurnTokens.inert(target)
rather than a defaulted overload, because a delivery with no token is
the unbound baseline this unit forbids.

The first version of inert() returned a fresh CompletableFuture as the
waiter, which turned captureBaselineSkipsTheReadWhenNoSendIsWaiting red:
the resolver saw a non-null waiter, concluded a turn was in flight, and
scraped a pane no send was blocked on. An inert value must omit the
fact, not invent it, so the waiter is now null and the production skip
fires as designed.
2026-08-15 06:36:35 +02:00
Dai Ha 0edc6615fc M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:32:02 +02:00
Dai Ha b745e159de CB-573: carry accepted turn tokens on delivery 2026-08-15 06:30:41 +02:00
Dai Ha aac29d604c M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m32s
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:29:27 +02:00
Dai Ha c884802b13 CB-577: remove obsolete async target tracking
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-15 06:28:53 +02:00
Dai Ha 5f5573a24e Merge CB-577: correlate an async question by its exact waiter
CI / contract (push) Failing after 0s
CI / build (push) Successful in 1m13s
markAsyncQuestion picked the first not-done task out of an unordered
set, so between resolveQuestion waking the first async send and the
question being recorded, a queued second send could join the set and
take the question. A lead answering with bridge_send{turnId} would then
resume a turn it did not mean to.

Each async task is now indexed by its exact rendezvous waiter, which has
identity semantics, so no other task can hold the same key. The question
is recorded before resolveQuestion, with a rollback when no waiter is
there, which closes the window rather than narrowing it.

An unanswered async question stays PENDING — the worker resumes after
its ask times out, so the delegation is not failed — and its stale
target tracking is now cleared instead of leaking.
2026-08-15 06:27:16 +02:00
Dai Ha 74b0087ebb CB-577: model async question ownership race
CI / build (pull_request) Successful in 1m29s
CI / contract (pull_request) Failing after 0s
2026-08-15 06:24:05 +02:00
Dai Ha 927e0151d4 CB-577: test async question waiter ownership 2026-08-15 06:23:22 +02:00
Dai Ha 5275922d1d CB-577: handle questions without async waiters 2026-08-15 06:22:22 +02:00
Dai Ha e186c7945a CB-577: correlate async questions to turns 2026-08-15 06:21:53 +02:00
Dai Ha 33a6e77f0e Merge CB-573: dormant fleet health monitor (M4 unit 1)
CI / build (push) Successful in 49s
CI / contract (push) Successful in 1m14s
An opt-in whole-fleet observer, separate from the 250ms delivery poller.
One AgentControl.list and one roster snapshot per tick, joined and fed to
the FleetHealth classifier, because a fault is a disagreement between the
two views at the same instant. Absent a health: block nothing is built
and no herdr call is made.

Adds bridge_list healthCoverage: off, detection-only, or full. Detection
is deliberately separate from notification, so a single-lead setup with
no webhook still gets detection and is told its coverage is partial
rather than being refused.

Two review fixes worth naming. tick() rescheduled itself as its last
statement with no try/catch, and a ScheduledExecutorService does not
re-run a task that threw — so the first agents.list failure would have
stopped health permanently and silently, which is exactly when the
control link is down. It now catches Throwable and reschedules in a
finally. And the snapshot fields this unit cannot supply are the named
constant NOT_YET_OBSERVED rather than bare false literals, because false
means no fault to this classifier.
2026-08-15 06:19:15 +02:00
Dai Ha 4e47489d53 Merge CB-568c: fail every pending async ticket on teardown
A released target left its second async ticket pending for the full
30-minute async timeout. abandon resolved only the rendezvous waiter,
and async tickets live in a separate map that could not even represent
two tasks on one target.

abandon now sweeps every non-question async ticket for the target, and
asyncTasksByTarget holds a set. The sweep is a plain loop: the first
attempt used Stream.anyMatch, which short-circuits on the first true, so
it completed one ticket and left the rest pending — the exact bug it was
fixing. Its test passed only because the rendezvous path failed the
first ticket anyway; the test now uses three tickets so a single
completion cannot satisfy it.

An ASKING ticket is an active turn, not a pending send, so the sweep
skips it and CB-574 is unaffected.
2026-08-15 06:18:28 +02:00
Dai Ha c76b2f149e CB-573: keep health monitoring after failures
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m15s
2026-08-15 06:17:28 +02:00
Dai Ha 826fffe05b CB-573: add dormant fleet health monitor 2026-08-15 06:15:24 +02:00
Dai Ha 75f57cdba7 CB-568c: fail every queued async ticket
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m34s
2026-08-15 06:15:21 +02:00
Dai Ha f556af5d4e CB-568: fail queued async tickets on teardown 2026-08-15 06:14:28 +02:00
Dai Ha d5f33f0c6e Merge CB-568: a dropped send reports the real cause
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m0s
Injector.drop knew the precise cause (herdr agent_not_found) but the
sender was told only 'worker unreachable or stuck', so a lead could not
tell a dead pane from a stalled model.

TurnListener.onTurnFailed gains a reason, defaulting to the old one-arg
form. CompletionResolver prefers that reason, then the pane scrape, then
the old fixed text.

drop now fires onTurnFailed unconditionally. That is the substantive
fix: the sender blocks on the rendezvous waiter, not on the delivered
future, so failing delivered() alone never woke it and a queued send sat
until its timeout.
2026-08-15 06:08:41 +02:00
Dai Ha 16e17b32ad CB-568: preserve dropped turn causes
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m26s
2026-08-15 06:06:07 +02:00
Dai Ha 988494e18b Merge M4 fleet-health design (docs only)
The full 5-unit design behind M4: evidence model, classification
precedence, the automatic-vs-lead action boundary, worktree safety on
release, typed inbox and lead routing, capacity, and human escalation.

Two decisions worth keeping visible. Detection is split from
notification, so health works in a single-lead setup with no webhook and
reports partial coverage instead of refusing to run. And capacity stays
a view: the bridge reports free slots but never spawns, reassigns, or
stops a member to improve utilisation, because only the lead holds the
work list.

Section 13 records eleven things nobody checked, including live LavinMQ,
OpenCode pane fixtures, and multi-lead routing.
2026-08-15 06:05:26 +02:00
Dai Ha c4549a5e20 Merge CB-575: filter the routine MCP cancellation WARN
The MCP SDK 2.0.0 registers no handler for notifications/cancelled, so every
client abort logged a WARN. M4 fleet health treats WARN as action-needed, so
that noise had a cost. A Logback TurboFilter denies only that one event:
right logger, WARN level, the SDK's exact format string, and a
JSONRPCNotification whose method is notifications/cancelled. Everything else
is NEUTRAL. If a later SDK handles cancellation the filter stops matching.
2026-08-15 06:00:40 +02:00
Dai Ha 46fa4f38d5 CB-575: filter MCP cancellation warnings
CI / build (pull_request) Successful in 49s
CI / contract (pull_request) Successful in 1m1s
2026-08-15 05:57:45 +02:00
Dai Ha 20e0e68ad7 M4: align design status with CB-573
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m28s
2026-08-15 05:56:00 +02:00
Dai Ha 17468a234a M4: document fleet health design 2026-08-15 05:55:00 +02:00
Dai Ha 8a53d5bfc6 Merge CB-574: an async delegation can now receive a worker's question
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m40s
A worker on a wait:false delegation called bridge_ask and the lead never
saw the question. Outcome.QUESTION is deliberately non-terminal, but
taskView tested r.completed() and fell into the failure branch, so the
ticket was marked FAILED and both the question text and its turnId were
discarded. The worker blocked for 55s, gave up, and had to abandon its
task. CLAUDE.md tells leads to prefer wait:false and to answer an ask with
bridge_send{turnId, content}; those two could not both be followed.

bridge_poll now returns a non-terminal ASKING phase carrying the question
and its turnId, and the ticket stays live so the worker's real reply still
lands on it. An unanswered ask returns the ticket to PENDING, because only
the question wait ended - the delegated turn continues. The 55s/115s ask
caps are unchanged: they exist because the worker's own MCP call would time
out, so widening them would only move the failure.

Two defects found reviewing the first revision, both from replacing
supplyAsync with a manually completed future:

- an exception inside the send left the future uncompleted, so the ticket
  stayed PENDING for the life of the daemon. Now caught and completed
  exceptionally.
- correlation was keyed by target, one entry per worker, registered before
  the session lock. With two tickets outstanding on one target the second
  overwrote the first, so a late reply could resolve the wrong ticket.
  Correlation is now per turn, the target entry exists only while that send
  owns the lock, and a reply with no live waiter still goes to the durable
  inbox as before.
2026-08-15 05:50:45 +02:00
Dai Ha 695da7418e Merge CB-573 (part 1): health classification model and the bridge_list capacity view
CI / contract (push) Successful in 45s
CI / build (push) Successful in 55s
Fleet capacity was invisible. A finished member held a terra slot until a
spawn was refused with 'at maxLoad: 2 live >= 2 cap', and nothing had told
the lead the slot was still held. bridge_list now reports, per configured
profile, maxLoad / live / free / reclaimable, and per member idleForSeconds
and reclaimable.

live comes from the same liveCountRef function placement consumes, so the
advertised free slots cannot drift from what bridge_spawn will accept. The
profile list is the union of configured and roster profiles: an empty
configured profile still appears with its full capacity, and a member whose
profile was removed from config stays visible rather than vanishing.

reclaimable is advisory. The bridge never spawns, stops or retasks a member
to improve utilisation: it has capacity facts but no work list, and choosing
work needs authority it does not have.

Also lands the pure health classifier, its precedence chain, the MUTE counter
and the pane budget. The classifier never reports IDLE while an accepted
delivery is open — IDLE is a claim that nothing is outstanding, and the
capacity view reads exactly that field.

Capacity dependencies are one required CapacitySource rather than defaulted
constructor arguments. A defaulted liveCount would report free slots that do
not exist, which is the dangerous direction; CapacitySource.none() omits the
block instead of inventing zeros.
2026-08-15 05:49:51 +02:00
Dai Ha abd26c796b CB-574: retain async task correlation
CI / build (pull_request) Failing after 1m14s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:49:27 +02:00
Dai Ha 24559d81ac CB-573: require explicit capacity source
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 05:49:04 +02:00
Dai Ha 6c1c2c3994 CB-573: report empty configured profile capacity
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m0s
2026-08-15 05:45:39 +02:00
Dai Ha 01fab15713 CB-573: add fleet capacity view
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m25s
2026-08-15 05:43:46 +02:00
Dai Ha 0cd00e71c3 CB-574: surface async worker questions
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 05:43:43 +02:00
Dai Ha bf0ff2adbf CB-573: keep active delegations out of idle
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m27s
2026-08-15 05:40:37 +02:00
Dai Ha ed4bbc1c56 CB-573: add pure health classification model
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:37:30 +02:00
Dai Ha c456402cc5 Merge CB-572: reject a configured profile name as a send target
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
A lead sent to sessionId "sol" — a profile name, not a terminal id. The
bridge accepted it, handed out a ticket, failed 60s later inside the
injector, and still reported the ticket as pending 20 minutes on. The
sender never learned anything and a whole delegation was lost.

bridge_send now rejects a target that exactly matches a configured
profile name, on both the blocking and the wait:false path, before any
ticket is issued. The error names the value and points at bridge_list.

The check is deliberately narrow. A target absent from the member roster
may still be a peer lead's terminal or a herdr-owned pane, so only a
value the bridge can prove is a profile is refused. profiles is a
required parameter on both send methods — the earlier revision kept
overloads that defaulted it to an empty set, which is the same silent
disable shape as CB-561.
2026-08-15 05:35:12 +02:00
Dai Ha 4f0bf667b1 CB-572: require profiles for send validation
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m2s
2026-08-15 05:33:15 +02:00
Dai Ha 61af9aa574 CB-572: reject profile names as send targets
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m20s
2026-08-15 05:29:37 +02:00
Dai Ha 3b59b34e76 CB-570 follow-up: rewrap an over-long javadoc line
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m17s
2026-08-15 05:12:58 +02:00
Dai Ha 65b38997f7 Merge CB-570: OpenCode receives the composed charter (PR #40)
One file, member-charter.md, not two. Two files would have made the U5
digest non-comparable between the Claude adapter and this one, which is
the whole point of the receipt; and nobody verified how OpenCode merges
multiple instruction files, so array order was an unverified dependency.

The OPENCODE_CONFIG condition widens to include a charter. It used to be
hasMcp() || hasCustomProvider(cfg), so a profile with a role charter but
no MCP and no custom provider would have got no config file and therefore
no charter — the feature silently doing nothing for that profile.

The file stays in the per-spawn temp dir, never the worktree: the
worktree is removed on release, the parity overlay already writes into
it, and CB-525's lesson was that config the bridge copied into a worktree
made a worker operate on the wrong tree. Being outside the repo is also
what stops it being committed, which a .gitignore line does not.
2026-08-15 05:12:26 +02:00
Dai Ha 7510f7649c Merge CB-569: Claude receives the composed charter (PR #39)
argvWithBridge used to gate the charter on cfg.hasMcp(), because the only
charter was the reply rule and telling a peer to call a tool it was not
given is a bug. A role charter is identity, not a tool instruction, so
the two gates are now separate: the MCP mount still depends on mcpUrl,
while --append-system-prompt depends only on the base having composed
something. A profile with a role charter and no MCP now gets its charter.

A null charter adds no flag at all. An empty --append-system-prompt is
not the same as no system prompt.
2026-08-15 05:11:39 +02:00
Dai Ha 619792a81c CB-570: deliver composed OpenCode charter
CI / build (pull_request) Successful in 58s
CI / contract (pull_request) Successful in 1m13s
2026-08-15 05:10:49 +02:00
Dai Ha cea1183f75 CB-569: pass composed charter to Claude
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-15 05:09:30 +02:00
Dai Ha e2af4c5ae4 CB-567 follow-up: realign the changed constructor signatures
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m26s
The supplier swap left three parameter lists indented one column past
their siblings and one javadoc line broken mid-sentence. No behaviour
change.
2026-08-15 05:06:03 +02:00
Dai Ha dfd5f82894 Merge CB-567: read the charter per spawn, compose it once (PR #38)
HerdrPeerLauncher now takes Supplier<BridgedConfig.Fleet> instead of
Supplier<String> tabLabelTemplate, and reads it once per spawn. A field
taken at construction would have made the charter deferred, and deferred
looks exactly like working — which is why the test uses a mutable
supplier and spawns twice, rather than ConfigRef.fixed().

Composition happens once in the base, not in each adapter: two copies
drift while both adapter-local tests keep passing. Role charter first,
reply charter last, because the final instruction is the one that must
not be overridden. The reply charter stays gated on hasMcp() — telling a
peer to call a tool it was not given is a bug — while the role charter is
not, being identity rather than a tool instruction.

Both REPLY_CHARTER copies collapse into one, and it now says 'spawned
member' rather than 'off-subscription worker'. The old text made the
launch prompt contradict bridge_whoami for an architect; both architects
read it in their own prompts and reported it.

The old buildLaunch overloads are removed rather than kept as defaults: a
surviving one is the same shape as a stale snapshot, a route that drops
role and charter while looking healthy. That removal also let the
OpenCode adapter drop its ThreadLocal resume-id hack, since LaunchSpec
now carries the value down the same path.
2026-08-15 05:04:08 +02:00
Dai Ha e81944cef6 CB-567: compose charters per spawn
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m3s
2026-08-15 05:00:48 +02:00
Dai Ha 7bdd39ab9a CB-566 follow-up: say why the five-arg Fleet constructor is kept
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m8s
Reviewing the merge I read the five-arg constructor as dead code and
removed it. That was wrong: the tests call it as BridgedConfig.Fleet,
which my grep for 'new Fleet(' did not match, and the build failed on
eight call sites. It is restored with a javadoc that says why keeping it
is safe here even though an overload that drops a new field is normally
the shape to avoid — nothing reads a charter through a constructor, and
Jackson binds the canonical one, so it cannot swallow an operator's YAML.

Also drop a redundant java.util.Arrays qualifier (the class is already
imported) and rewrap a javadoc line the change had left over-long.
2026-08-15 04:54:09 +02:00
Dai Ha f4b38f040e Merge CB-566: the charter config surface (PR #37)
fleet.charters is a validated Map<String,String>, not a record: Fleet is
@JsonIgnoreProperties(ignoreUnknown = true), so a record field named
architetc would be dropped in silence and the operator would never learn
of the typo. A map lets validateCharters see the bad key and refuse it.

The key sits under fleet: because ConfigRef already treats that block as
hot and changedDeferredKeys does not list it. A new top-level key would
inherit nothing, and forgetting to classify it means a reload prints
'config reloaded' and does nothing.

A blank value is refused while an absent one is fine: an absent key means
the operator configured no charter, a blank one means they tried and
failed. Refusing at both startup and reload is the point — wiring only
one of the two paths is the whole bug.
2026-08-15 04:52:48 +02:00
Dai Ha 799668e129 CB-566: add fleet charter config
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 55s
2026-08-15 04:50:23 +02:00
Dai Ha 890190263e CB-560/562/563: the shipped block tells the truth again
The canonical CLAUDE.md block is the instruction surface this daemon ships to
every agent that mounts it, so a code change that silently invalidates it is an
incomplete change. CB-548 added a third principal kind and the block was never
revisited. Four statements in it were simply false:

  * whoami was documented as returning only primary or worker;
  * "spawn/stop/send/drain are lead-only" — Authz permits SEND to an architect;
  * delivery was documented as idle/blocked — injectable() is IDLE|BLOCKED|DONE,
    and a spawned member must also have mounted the bridge MCP, which is the
    exact condition that made every architect undeliverable for a day;
  * the tool table said bridge_list returns `workers` — the JSON key is
    `members`.

Also corrected: the fallback ladder claimed each one-way signal identifies a
"worker", but an architect gets the same charter, the same mount and the same
env, so those signals identify a spawned member and only bridge_whoami
separates the two. The safe default stays "act as a worker" — it is the most
restricted member role.

The turn contract now covers both member kinds, and says why the completion
fallback is not a substitute for bridge_reply: it returns at most the last 4000
characters, so a long report reaches the lead with its end cut off. That is not
hypothetical — it happened twice today.

Found by a reviewer asked whether the block still matches the code. Verified
against injectable(), Authz, BridgeMcp.listFleet and LeadLauncher before
applying. The wiki template is updated in the same shape and re-checked
byte-identical.
2026-08-15 04:37:03 +02:00
Dai Ha f8522edacd Merge CB-564: give every silent failure path a voice
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m15s
The bridge cannot classify what it never emits. Fleet-health monitoring —
detect a wedged member, decide, escalate — is blocked on that, so this is the
foundation rather than the feature.

A survey of Injector, CompletionResolver, StatusPoller, SessionManager,
SessionReaper and MessageService found six conditions that ended a member's
usefulness while saying nothing useful:

  * Injector.drop()             — worker gone, queue cleared: SILENT
  * CompletionResolver.fail()   — "via turn-stall fallback" at DEBUG, no reason
  * SessionManager.onFailed()   — "session marked failed" at DEBUG, no stage
  * SessionManager.reapIdle()   — indistinguishable from any other release
  * SessionManager.acquire()    — spawn failure rethrown with no log at all
  * MessageService.abandon()    — failed a caller's request at DEBUG

The first two are the exact phrases that misled the CB-560 diagnosis: both name
a symptom and neither names a cause. They are now WARN and carry the reason,
the stage, and the counts.

Conditions already loud were left alone, and so were two by-design timeouts in
MessageService — an async model exists precisely for those, and promoting them
would turn healthy operation into noise.

Behaviour is unchanged: every edit is a log statement.

Merge note: the recycle test deleted by CB-565 conflicted with a test added
here. Resolved by keeping the new onTurnFailed assertion and dropping the
recycle test, which tests a method that no longer exists.
2026-08-15 04:36:09 +02:00
Dai Ha b958747855 Merge CB-565: remove the recycle trap (PR #35)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m27s
SessionManager.recycle called the 4-argument acquire overload, which defaults
the role to DEV and requests no worktree. So it carried profile, cwd and owner
across and silently dropped two things: the member's role, and its worktree.

A recycled architect would have come back a plain worker, never rebound to its
slot, with nothing logged. Worse, a recycled member would have come back with
no worktree at all — and a member's uncommitted work exists in exactly one
place. Role loss is recoverable; that is not.

Nothing called it. The only reference outside its own javadoc was one test, and
the context-cap path calls release, not recycle. So this was a trap waiting for
its first caller, and that caller would not have noticed either loss.

Deleted rather than repaired, on the CB-561 precedent: an API that looks correct
and silently drops a property is worse than no API. The no-reuse invariant it
documented is still true and is now stated directly.

Found by a reviewer asked to hunt the rest of the CB-548 fallout. The worktree
half was found while verifying the report.
2026-08-15 04:31:46 +02:00
Dai Ha 078bde2c02 CB-564: give a voice to silent member-failure paths
Injector.drop, CompletionResolver.fail, SessionManager.onFailed/reapIdle/
acquire spawn failures, and MessageService.abandon used to fail a member
or a caller's request with no log, a bare DEBUG, or a log that named only
the symptom ("session marked failed", "failed send via turn-stall
fallback"). Each now logs at WARN and names the real cause and the
numbers involved. Observability only — no behaviour changed.
2026-08-15 04:31:42 +02:00
Dai Ha 0b10ea987b CB-565: remove unsafe session recycle
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 04:29:50 +02:00
Dai Ha 1553d38182 Merge CB-563: a clipped pane scrape says it was clipped (PR #34)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m12s
When a member ends its turn without bridge_reply, CompletionResolver scrapes
the pane and resolves the waiting send with that text. The scrape is capped at
MAX_SCRAPE_CHARS (4000), and nothing told the caller when the cap had bitten.
A delegating lead could act on a report missing its end and believe it was
complete. That happened to me today: a member's full engineering report arrived
cut at exactly 4000 characters, and the only hint was a DEBUG line reading
"(4000 chars scraped)", which reads like a size and not like a warning.

The returned text now carries a marker when, and only when, it was clipped, and
the clip is logged at WARN with the original length and the cap.

The cap itself is unchanged. The problem was silence, not the number.

The CB-115 misattribution guard still compares the unmarked clipped tail to the
unmarked baseline, and the marker is appended only afterwards. Verified in the
code, not taken on report: clip() strips before truncating and the new length
check uses the same stripped length, so there is no off-by-one either.
2026-08-15 04:26:57 +02:00
Dai Ha d04b075996 CB-563: mark clipped completion scrapes
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 04:25:54 +02:00
Dai Ha 644927636d Merge CB-562: the readiness gate says why it gave up (PR #33)
CI / build (push) Successful in 1m1s
CI / contract (push) Successful in 1m14s
When the injector's readiness grace expired it cleared the queue and logged
nothing. The failure then surfaced elsewhere as a turn-stall, which names the
wrong cause. Diagnosing CB-560 cost two live spawns and a wrong first
hypothesis for exactly this reason: the logs said "session marked failed" and
"failed send via turn-stall fallback", and neither says the message was never
typed into the pane at all.

The expiry now logs the target, the number of messages being failed, the grace
in polls and seconds, and the real cause in plain words.

The grace in seconds is derived, not written down twice: Bridged's own
INJECT_POLL_MILLIS is deleted and Injector.POLL_INTERVAL_MILLIS is the single
source, passed to every StatusPoller. A cadence change can no longer leave a
log line confidently stating the wrong duration.

Behaviour is unchanged. This is the first structured health event in the
daemon, and the foundation the fleet-health work will build on.
2026-08-15 04:21:29 +02:00
Dai Ha 95e45007aa CB-562: single source for the injector poll cadence; tighten count assertion
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m17s
2026-08-15 04:20:29 +02:00
Dai Ha 7a583c4045 CB-562: log why the readiness gate gave up on a target
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:52:25 +02:00
Dai Ha 65bce058bb Merge CB-561: one public way to build a CallerResolver (PR #32)
CI / build (push) Successful in 1m17s
CI / contract (push) Successful in 1m16s
CallerResolver had 5 public constructors and 4 public factories, and only one
of them could ever produce an architect. The rest defaulted memberSlotRoles to
`_ -> null`, so every architect quietly fell through to Principal.worker().
Nothing logged, nothing threw — the role was simply off.

That is the same failure shape as CB-560, so the fix is structural rather than
a warning: withLeadsAndMembers(.., MemberRegistry) is now the only public
construction path. Two overloads with no caller at all are deleted; the rest
are package-private and marked test-only. No path remains that accepts
architect bindings without a slot-role lookup, so no runtime WARN is needed.

Also checked and closed: the suspected slot leak on shutdown drain is not
real. SessionManager.release calls memberLifecycle.released() for every
removed session, and drainAll routes every session through release,
SPAWNING included. Verified in code.
2026-08-14 21:50:04 +02:00
Dai Ha 48b437083b Merge CB-560: mark spawned members present (PR #31)
Merging CB-548 made an architect resolve as Role.ARCHITECT, and BridgeMcp
marked MCP presence only for Role.WORKER. So an architect was never present.
That one line broke two things, because SessionManager.asPresence() is the
same object the injector's readiness gate reads:

  * the session never left SPAWNING;
  * Bridged.deliverableTo was false, so the injector held every delivery,
    waited out READINESS_GRACE_POLLS (~60s) and failed the send.

Confirmed live twice, on claude-code and on opencode. Both logged
"session marked failed" then "failed send via turn-stall fallback", neither
of which names the real cause.

Principal.isSpawnedMember() now means "a spawned member with a pane" —
WORKER or ARCHITECT, never PRIMARY. The Bridged.deliverableTo javadoc, which
still stated the worker-only rule as fact, is corrected: it is the clearest
description of this invariant anywhere, and leaving it stale is how the bug
comes back.
2026-08-14 21:50:04 +02:00
Dai Ha fa0612859b CB-560: document spawned member presence
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:48:58 +02:00
Dai Ha 4ffbcd0b7d CB-561: require member registry for architects
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-14 21:47:53 +02:00
Dai Ha a0cd053fd9 CB-560: mark architect members present
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:45:28 +02:00
Dai Ha f9a5e066b5 Merge PR #30: bind a spawned architect to its slot (CB-548's missing half)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m18s
MemberRegistry.bind() had 18 call sites, every one in its own unit test.
Nothing in production ever bound a terminal to a slot, so CallerResolver's
architect branch was unreachable, every spawned architect resolved as WORKER,
and Authz's architect row was dead code. Since SEND is granted to
isPrimary() || isArchitect(), no architect could message anyone.

The spawn lifecycle now binds on both acquire paths and unbinds on release,
through a MemberLifecycle seam with a no-op default, so all six SessionManager
constructors are unchanged.

The escalation guard is deliberately double-sided. MemberRegistry flattens
every pool, so dev:* and reviewer:* slots sit in the same map the architect
resolver reads. The lifecycle binds only ARCHITECT roles, AND CallerResolver
independently checks roleForSlot(slot) == ARCHITECT before granting. Either
alone would be enough today; together, a regression in one cannot escalate a
worker. A reviewer confirmed a third barrier already existed: Authz gates
SPAWN on isPrimary(), so only a lead can request role=ARCHITECT at all.

Verified by the primary: mvn clean install, 640 tests, 0 failures.

Two findings recorded, neither blocking:

1. Known race, low severity. registry.put() and memberLifecycle.acquired()
   are not atomic. A release landing between them unbinds nothing (no binding
   exists yet), then acquire binds a dead terminal that no later release will
   ever clear - the slot leaks. Only the shutdown drain can reach a SPAWNING
   session (the idle reaper skips it, and an explicit stop needs a paneId the
   spawn has not returned), so the leak dies with the process. Note that
   binding before put does NOT fix it: the drain then iterates a roster the
   session is not in yet. A real fix needs atomicity.

2. The legacy Supplier-based withLeadsAndMembers overload and the Map-form
   constructors now silently disable architect resolution: memberSlotRoles
   defaults to _ -> null, so the architect branch can never be taken there.
   Production uses the registry form, and the tests were migrated, so nothing
   fails - but a future caller gets workers with no error.
2026-08-14 21:25:45 +02:00
Dai Ha e9bc192160 CB-548: bind architect sessions to member slots
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 55s
2026-08-14 21:10:48 +02:00
Dai Ha e0a57988ad Merge PR #29: drop .claude/settings.local.json from the default parityOverlay
CI / build (push) Successful in 1m21s
CI / contract (push) Successful in 1m21s
A member no longer receives the primary's pre-approved tool grants by default.
The grants were inert — GitWorktrees.isolateToolSurface already strips each
worktree's .mcp.json to an empty server map — so this is defence in depth: two
independent guards instead of one. Same reasoning as CB-525, which removed the
sibling .mcp.json and left this file behind.

An explicit parityOverlay: in config is unaffected; only the default changes.

Verified by the primary: mvn clean install, 637 tests, 0 failures.
2026-08-14 21:02:25 +02:00
Dai Ha b414a74c26 CB-559: drop .claude/settings.local.json from the default parity overlay
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m7s
2026-08-14 20:56:16 +02:00
Dai Ha 2c3796d598 Merge: keep every credential in one store
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m13s
2026-08-14 20:38:36 +02:00
Dai Ha 7e0ff9ab06 Keep every credential in one store, not a per-repo copy
The repo carried a gitignored .secrets/ directory with four files. Two of them
(context7-token, gitea-token) were byte-identical copies of variables the login
shell already exported. One (gitea-host) is not a secret. The fourth
(worker-gitea-token) was the only copy anywhere, and nothing exported it, so
bridged read gitTokenEnv from an environment that never had it and every worker
push got an empty token.

All four values now live in the operator's single sourced secrets file, verified by
sha256 before the copies were removed. opencode.json reads them as {env:...}, which
.mcp.json already did. A second copy of a secret is the problem: the copy you forget
is the one that leaks or goes stale.

This makes worktree isolation load-bearing rather than a workaround. opencode.json is
tracked, so it lands in every worktree. It used to fail there, because {file:.secrets/}
pointed at files a worktree never receives and OpenCode refuses to start on a dangling
reference. With {env:...} the reference resolves, and a member would silently inherit
the primary's admin-scoped GITEA_ACCESS_TOKEN. GitWorktrees already neutralizes the
file; only its stated reason changes, and it is now a confidentiality boundary.

The port-to-opencode skill taught {file:.secrets/} as the preferred pattern, so it is
rewritten to teach the central store and to say why we moved. .gitignore keeps the
.secrets/ line as a backstop against habit.

Includes the wiki pointer, which also carries the CB-559 config-reload correction.
2026-08-14 20:38:31 +02:00
Dai Ha 6d1565d8b6 Merge CB-559 correction: profile launch settings are deferred, not hot 2026-08-14 20:37:47 +02:00
Dai Ha e18e002d2f CB-559: a profile's launch settings are deferred, not hot
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.

What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.

changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
2026-08-14 20:37:42 +02:00
Dai Ha c18572ea9d Bump wiki to 8b1eb68 — the CB-557/558/559 Features entries
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m14s
2026-08-14 19:52:02 +02:00
Dai Ha e9d02bc5e8 Merge CB-557/558/559: the fleet block, role pools, lead auto-launch, config reload
Four commits that together reshape how the daemon is told who it may run.

CB-557 (schema) — one `fleet:` block replaces `leaders:`, `members:`, `leadScan:`
and `defaultProfile:`. Role is the containing map key, not a `role:` field, so a
misspelled pool name declares nothing instead of producing a member with no
contract. Tab labels are role-first and counted per role+profile.

CB-557 (routing) — an unqualified spawn picks its profile from its role's pool.
Role and profile stay orthogonal: a reviewer may run on the same profile as the
dev whose diff it reads, so the two cannot be one field. An explicit
`bridge_spawn{profile:…}` stays exempt, because it carries no role and would
otherwise be refused for a profile the operator named.

CB-558 — the daemon launches a declared lead when fewer than `instances` are
running. An auto-launched lead is NOT a member: no reply charter, never
registered with SessionManager (the idle reaper would kill it), and no Anthropic
binding in its env. A lead counts as live only when herdr reports a running
agent, so one crash does not disable auto-launch forever, and an unreachable
herdr starts nothing at all.

CB-559 — `bridged.yaml` can be re-read without a restart, opt-in through
`configReload:`. Keys are hot (`fleet:`, `placement:`, an existing profile's
fields), deferred (lifecycle, guard, adding a profile) or cold (bind,
herdrSocket, broker, auth). A cold change refuses the WHOLE reload rather than
half-applying it, because a daemon matching no file on disk is worse for an
operator than no reload at all.

634 tests, IDE diagnostics clean. Verified live against the running daemon: the
lead launcher recognised the existing lead and started nothing; the reload
applied a hot change, refused a `bind:` change exactly once (not once per tick),
and reverted cleanly.
2026-08-14 19:51:34 +02:00
Dai Ha a2108a8a14 CB-559: re-read bridged.yaml without restarting the daemon
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.

ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.

Keys fall into three classes, and the difference is what already exists when
the reload happens:

  hot       fleet: (pools + tabLabel), placement:, and an existing profile's
            weight / maxLoad / model / tabLabel — live on the next spawn.
  deferred  lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
            and adding/removing a profile — accepted, but the startup wiring
            keeps the old value. The reload logs these by name.
  cold      bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.

A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.

A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.

ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.

MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.

634 tests.
2026-08-14 18:23:32 +02:00
Dai Ha 57f8fa257a CB-558: launch a declared lead at startup when none is live
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.

A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.

Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.

Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.

Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.

Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.

617 tests pass (16 new), IDE-clean.
2026-08-14 16:47:57 +02:00
Dai Ha 61944fc045 CB-557: place an unqualified spawn inside its role's pool
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.

Three parts:

SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.

CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.

Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.

An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.

Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.

601 tests pass.
2026-08-14 16:37:29 +02:00
Dai Ha 4b48d2d921 CB-557: fleet role pools — role is the config key, and the tab label says it
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.

Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.

`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.

Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.

Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.

Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.

Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.

595 tests pass.
2026-08-14 16:32:54 +02:00
Dai Ha 3aa145cbef Merge CB-557: member taxonomy in the API surface
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m1s
2026-08-14 07:08:26 +02:00
Dai Ha 4875127daa CB-557: adopt the member taxonomy in the API surface
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.

MCP:
  bridge_spawn gains role: architect | dev | reviewer (default dev). An
    unknown role is refused with the valid spellings in the message.
  bridge_list returns "members" instead of "workers"; each row carries both
    role (what it is for) and profile (which backend it runs on).
  The spawn result echoes the role back, so a spawn that fell back to dev is
    visible rather than silent.

REST:
  GET/POST /members and DELETE /members/{paneId} replace /workers.
  POST accepts role= as a query param or a body field; an unknown role is 400.

Code:
  dev.ltms.bridged.worker package -> dev.ltms.bridged.member
  WorkerSession   -> MemberSession, plus a MemberRole role component
  WorkerPresence  -> MemberPresence
  SessionManager.acquire gains a role parameter; the existing overloads keep
    working and default to DEV, which is exactly what "worker" used to mean.

ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.

Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.

mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:08:26 +02:00
Dai Ha d14a624421 Merge CB-557: member taxonomy in config
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m32s
2026-08-14 07:02:22 +02:00
Dai Ha 246f50b778 CB-557: adopt the member taxonomy in config
A member is anything a lead spawns. Every member carries two independent
attributes:

  role    — which contract: architect, dev or reviewer. It picks the launch
            charter, the role file, the playbook skill and the authz row.
  profile — which backend: model, CLI adapter, credentials, cost.

They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.

Config changes (breaking — we are in active development, so no aliases):

  workers:       -> profiles:        it was never a list of workers; it is a
                                     catalogue of backends
  defaultWorker: -> defaultProfile:
  architects:    -> members:         each slot now names its role

An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.

Also:
  - BridgedConfig.Worker    -> BridgedConfig.Profile
  - ArchitectRegistry       -> MemberRegistry
  - new peer.MemberRole enum, validated at startup
  - profiles map is normalized once in the compact constructor, so the raw
    map and the derived one can no longer disagree
  - the legacy singular worker: block is dropped
  - workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
    (the record component now owns the plain name)

mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:02:17 +02:00
Dai Ha 15b53c6cfa Merge CB-553: enforce maxLoad on explicit-profile spawns
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m8s
2026-08-14 06:51:21 +02:00
Dai Ha 3a7ef0adbd CB-553: enforce maxLoad on explicit-profile spawns (no cap bypass) 2026-08-14 06:50:28 +02:00
Dai Ha 293a305748 Merge CB-551: idle-lead heartbeat
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m21s
2026-08-13 21:21:33 +02:00
Dai Ha e5038c6d13 CB-551: idle-lead heartbeat — nudge an idle lead back to work on a timer
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-13 21:10:21 +02:00
Dai Ha 13ea79f6fd CB-523 (opencode): generated worker config pins compaction.auto
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m10s
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.

OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.

The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.

Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
2026-08-13 21:03:52 +02:00
Dai Ha f55d3c203b Merge CB-552: supersede rejected orchestrator design; sync wiki Features
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m9s
2026-08-13 20:59:50 +02:00
Dai Ha 37c4d47d3a Merge CB-544: shutdown drain preserves worker worktrees
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m17s
2026-08-13 20:58:06 +02:00
Dai Ha 04e21c9243 CB-544: shutdown drain preserves worker worktrees (no data loss)
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m29s
2026-08-13 20:36:32 +02:00
Dai Ha 83cac07f6e CB-552: update wiki documentation pointer 2026-08-13 20:19:46 +02:00
Dai Ha f472e0f782 CB-552: supersede rejected orchestrator design 2026-08-13 20:19:46 +02:00
ltms 91c9f981c5 Merge CB-548 send ownership and rendezvous hardening
CI / build (push) Successful in 1m16s
CI / contract (push) Successful in 1m17s
2026-08-13 19:44:02 +02:00
Dai Ha dd906526c0 CB-548: guard primary singleton to PRIMARY callers; open waiter before enqueue
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 53s
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
  fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
  as the fallback; the per-target delegation map does not cure the singleton. New
  BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
  (PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
  enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
  callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
  no stale waiter and no queued, orphanable message.

Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
2026-08-13 19:42:48 +02:00
Dai Ha ec3001796a CB-548: record delegator ownership only when a send is accepted
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
2026-08-13 19:42:48 +02:00
ltms 86cf4c285a CB-547b: opencode session identity — resume (-s) + post-hoc discovery (#18)
CI / build (push) Successful in 48s
CI / contract (push) Successful in 1m17s
2026-08-13 19:41:06 +02:00
Dai Ha cd18887b69 CB-547b: opencode session identity — resume (-s) + post-hoc session discovery
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m16s
2026-08-13 19:40:17 +02:00
ltms 4637c68295 CB-543: neutralize worktree-hostile tracked configs (#22)
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m14s
2026-08-13 19:39:54 +02:00
Dai Ha 0114bd1fa7 CB-543: neutralize tracked opencode.json and .autoenv in provisioned worktrees
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m26s
2026-08-13 19:30:19 +02:00
ltms 6dee84ca71 Merge CB-548 architect config and authorization foundation
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m17s
2026-08-13 18:33:11 +02:00
Dai Ha f004a0c654 CB-548: make duplicate-architect-slot detection top-level-only and depth-safe
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m13s
2026-08-13 18:31:29 +02:00
Dai Ha 6123576c68 CB-548: correct architect premise — profile-only slots, registry-owned bindings, dup-key rejection
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 18:05:53 +02:00
412 changed files with 93504 additions and 15772 deletions
+5 -5
View File
@@ -1,15 +1,15 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"name": "fleetd",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the fleetd MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"name": "fleet",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"description": "Mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.2.0",
"author": {
"name": "LTMS"
}
+24
View File
@@ -0,0 +1,24 @@
---
name: architect
description: Refine work into clear, independent units before implementation.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
You never commit production code and never open a pull request.
A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Report only work you actually did and the real output of checks you ran. Do not
claim a result from a tool you could not use. The primary's IDE tools are not yours.
A mounted forge tool may use a blocked credential and fail by design.
The launcher provides the required bridge reply instructions for every member.
+29
View File
@@ -0,0 +1,29 @@
---
name: dev
description: Implement one assigned unit, test it, and open a pull request.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Implement the change and run the full required build in your worktree. Read the
complete output and report its real result. Do not hide failures with a pipe. State
only checks you actually ran. The primary's IDE tools are not yours. A mounted forge
tool may use a blocked credential and fail by design.
Stage only files you changed. Never use `git add -A` or `git add .`. Never commit
`.mcp.json` or `wiki/`. Commit with a clear message, push your branch, and open your
own pull request against `main`. Never merge.
Your handoff must name the pull request or why it was not created, the branch, the
files changed, the build result, and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
+23
View File
@@ -0,0 +1,23 @@
---
name: hunter
description: Sweep one assigned scope for defects and report ranked findings without changes.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You sweep the assigned package or scope for real defects. Read the full assigned scope before you
judge it. Report several ranked findings when the evidence supports them. Change nothing: do not
edit code, commit, push, or open a pull request.
You may run the build or tests to check a finding. Read the complete output and report the real
result. Do not hide failures with a pipe. State only checks you actually ran. The primary's IDE
tools are not yours. A mounted forge tool may use a blocked credential and fail by design.
Do only the assigned scope. Note anything outside it in one line and do not investigate it further.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an unclear requirement
or two defensible fixes. Do not ask about something you can decide by reading more code.
Your handoff must name the files you read, each ranked finding or `NO FINDINGS`, the checks you ran,
and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
+34
View File
@@ -0,0 +1,34 @@
---
name: reviewer
description: Review one assigned scope and report the most important real issue.
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
Read the whole assigned scope before judging it. Review only that scope. If you see
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
Report the single most important real issue in this form:
```
1. <path>:<line>
2. issue: <one sentence: what is wrong and why it matters>
3. fix: <one line: the concrete change>
4. severity: high | medium | low
```
If there is no real issue, report `NO ISSUE` and one line saying why. A clean review
is valid. Do not invent an issue. Use high for a wrong result, data loss, security,
or a hang or crash on a real path. Use medium for an edge-path bug or a correctness
risk under load or concurrency. Use low for clarity, a latent foot-gun, or a smell
with no current failure.
The launcher provides the required bridge reply instructions for every member.
+274
View File
@@ -0,0 +1,274 @@
---
name: fleets-status
description: Report the status of every fleet that shares one LavinMQ instance. Use for local daemon health, broker-wide fleet presence, and cross-host lead coordination checks.
---
# Status of every fleet on the shared LavinMQ instance
**The headline: always report what is missing.** This skill starts with the local fleet, then adds
broker-wide facts when its read-only credential exists. A missing fleet must appear as `unknown` or
`not reachable`, with the reason and the fix. Never leave it out.
The known topology has one LavinMQ instance on `10.10.20.13` (`fleet01`). AMQP uses port `5672`,
and the management API uses port `15672`. The Mac fleet owns vhost `/mac`. The fleet01 fleet owns
vhost `/fleet01`.
## 1. Protect credentials before any probe
**Hard rule — never print `LAVINMQ_URI`.** It is an AMQP URI with its password inline. It only
resolves in a login shell because `${SHARED_ENV}/tools/secrets.sh` supplies it. A non-login shell
can make every broker probe look empty.
- Never run `echo "$LAVINMQ_URI"`.
- Never put `${LAVINMQ_URI:-something}` in output. That form expands to the secret value when set.
- Parse the user, host, and password into shell or Python variables. Use them without printing them.
- Prefer `resolves` or `does not resolve` over any part of the value.
- Every command that can read `LAVINMQ_URI` must send all output through this redaction before it
reaches the report:
```bash
sed -E 's#://[^@]*@#://<redacted>@#g'
```
**The `g` flag is not optional.** Without it `sed` replaces only the first match on each line, so a
line carrying two URIs leaks the second one. `scripts/redeploy-fleetd.sh --check` prints lines like
that. Checked on 2026-08-27: without `g`, `amqp://u1:p1@h1/mac and http://u2:p2@h2:15672/api`
redacts the first pair and prints `u2:p2` in the clear.
Keep `pipefail` on when applying that filter. Otherwise the filter can hide a failed probe. Apply
the same no-print rule to the management password below, even though it is not in an AMQP URI.
## 2. Tier 1 — this fleet (always run)
Start here even when the broker tier is blocked. Work from the local fleetd checkout.
First run the read-only deployment check. It already checks the daemon process, deployed jar versus
the checkout `HEAD`, launchd state, and whether each configured token resolves in a login shell.
Do not copy those checks into new shell code. The script reads `LAVINMQ_URI`, so redact all output:
```bash
set -o pipefail
scripts/redeploy-fleetd.sh --check 2>&1 \
| sed -E 's#://[^@]*@#://<redacted>@#g'
git rev-parse HEAD
```
Treat jar drift as a top-level warning. A merge is not a deployment. State the running jar result
as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into a match.
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
for PID in $PIDS; do
ps -p "$PID" -o pid=,etime=,lstart=,command=
done
fi
```
Read the full health response. Keep the HTTP status because `503` means fleetd is running but herdr
is not reachable. Report both `herdr.version` and `herdr.protocol` when present:
```bash
curl -sS --max-time 5 -w '\nHTTP %{http_code}\n' http://127.0.0.1:8765/healthz
```
Call `fleet_whoami`, then call `fleet_list`. Preserve its sections in the report:
- `leads`, including which row is this lead;
- `members`, including state, role, profile, branch, and worktree when present;
- every per-profile `capacity` row, including `maxLoad`, `live`, `free`, and quarantine facts;
- the exact `healthCoverage` value.
Do not describe an empty `members` list as an empty fleet. It says only that no members are spawned.
Also do not hide a profile with `free: 0`; say whether load or credential quarantine caused it.
Show only WARN, ERROR, and SEVERE lines after the last `fleetd listening` line. This anchor stops an
old incident from looking current:
```bash
python3 - <<'PY'
from pathlib import Path
import re
path = Path("fleetd/fleetd.out")
if not path.exists():
print("cannot check current WARN/ERROR: fleetd/fleetd.out does not exist")
else:
lines = path.read_text(errors="replace").splitlines()
starts = [i for i, line in enumerate(lines) if "fleetd listening" in line]
if not starts:
print("cannot anchor WARN/ERROR: no 'fleetd listening' line exists")
else:
current = lines[starts[-1]:]
alerts = [line for line in current if re.search(r"\b(?:WARN|ERROR|SEVERE)\b", line)]
print(f"current WARN/ERROR/SEVERE count: {len(alerts)}")
for line in alerts[-50:]:
print(line)
PY
```
**What this tier cannot see:** it proves facts only about the Mac daemon at `127.0.0.1:8765`.
It cannot show the fleet01 daemon, broker queue depth, or broker consumers. The fleet01 REST service
at `10.10.20.13:8765` is not reachable from the Mac. Say this in the report rather than omitting
fleet01.
**But fleet01 IS reachable over SSH — checked 2026-08-28.** An older version of this line said SSH
was denied. That is true only for the user `dai.ha`. The host alias `fleet01` maps to user `ltms`,
and `ssh fleet01` works with key auth:
```bash
ssh -o BatchMode=yes -o ConnectTimeout=6 fleet01 'echo $(id -un)@$(hostname)'
```
So fleet01's daemon PID, uptime, jar and `/healthz` **can** be reported — over SSH, not over REST.
Do that rather than writing `not reachable`. `ltms` also has passwordless sudo there.
## 3. Tier 2 — the shared broker (run when management access exists)
**This tier is blocked today.** The AMQP user in `LAVINMQ_URI` can connect on port `5672`, but gets
HTTP `401` from the management API on port `15672`. An AMQP connection does not grant monitoring
access.
The operator must create a separate, read-only LavinMQ management user with the `monitoring` tag.
It needs access to inspect both `/mac` and `/fleet01`. Store its values as
`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` in
`${SHARED_ENV}/tools/secrets.sh`. Do not reuse or print the AMQP URI. Full multi-fleet status stays
blocked until this user exists.
When both variables resolve, run this from a login shell. It calls `GET /api/overview`,
`GET /api/vhosts`, `GET /api/queues`, and `GET /api/connections`. It prints selected status fields,
but never the user, password, Authorization header, or AMQP URI:
```bash
zsh -lc 'python3 - "$@"' -- <<'PY'
import base64
import json
import os
import sys
import urllib.error
import urllib.request
base = "http://10.10.20.13:15672"
user = os.environ.get("LAVINMQ_MANAGEMENT_USER", "")
password = os.environ.get("LAVINMQ_MANAGEMENT_PASSWORD", "")
if not user or not password:
print("broker tier: BLOCKED — management credential does not resolve in a login shell")
sys.exit(0)
token = base64.b64encode(f"{user}:{password}".encode()).decode()
def get(path):
request = urllib.request.Request(
base + path,
headers={"Authorization": "Basic " + token, "Accept": "application/json"},
)
with urllib.request.urlopen(request, timeout=5) as response:
return json.load(response)
try:
overview = get("/api/overview")
vhosts = get("/api/vhosts")
queues = get("/api/queues")
connections = get("/api/connections")
except urllib.error.HTTPError as error:
print(f"broker tier: BLOCKED — management API returned HTTP {error.code}")
sys.exit(0)
except Exception as error:
print(f"broker tier: BLOCKED — management API is not reachable: {type(error).__name__}")
sys.exit(0)
fleet_names = {"/mac": "Mac fleet", "/fleet01": "fleet01 fleet"}
print(json.dumps({
"overview": {
"lavinmq_version": overview.get("lavinmq_version"),
"rabbitmq_version": overview.get("rabbitmq_version"),
"queue_totals": overview.get("queue_totals", {}),
"object_totals": overview.get("object_totals", {}),
},
"fleets": [
{
"fleet": fleet_names.get(vhost.get("name"), "UNKNOWN FLEET"),
"vhost": vhost.get("name"),
"queues": [
{
"name": queue.get("name"),
"messages": queue.get("messages", 0),
"messages_ready": queue.get("messages_ready", 0),
"messages_unacknowledged": queue.get("messages_unacknowledged", 0),
"consumers": queue.get("consumers", 0),
}
for queue in queues if queue.get("vhost") == vhost.get("name")
],
"connections": [
{
"name": connection.get("name"),
"peer_host": connection.get("peer_host"),
"state": connection.get("state"),
}
for connection in connections if connection.get("vhost") == vhost.get("name")
],
}
for vhost in vhosts
],
}, indent=2, sort_keys=True))
PY
```
Map `/mac` to the Mac fleet and `/fleet01` to the fleet01 fleet. Keep any other vhost in the
report as `UNKNOWN FLEET`; do not drop it. For each vhost, total the ready, unacknowledged, and all
messages. Report every queue's consumer count and each live connection.
**A vhost with queues but zero consumers means that fleet's daemon is down while its durable state
survives. Call this out as a top-level warning.** This is the main reason to use the management API
instead of calling each remote daemon.
**What this tier cannot see:** without the new `monitoring` credential it cannot enumerate any
vhost, queue, depth, consumer, or connection. With the credential it still cannot report fleet01's
daemon PID, uptime, jar revision, `/healthz`, herdr version, or member capacity. Those need reachable
fleet01 REST or SSH access, which the Mac does not have today.
## 4. Tier 3 — cross-fleet lead coordination
Use the queue data from Tier 2. Select queues whose names match `lead.<coordId>.inbox`. Report each
queue's vhost, depth, consumer count, and the `coordId` between the prefix and suffix.
- A lead inbox with a consumer shows that a lead mailbox is live on that vhost.
- A durable lead inbox with zero consumers shows saved coordination state, but no live receiver.
- No lead inbox is not proof that coordination is disabled. The daemon may be down before declaring
its queue, or this account may not be allowed to see the vhost.
This Mac fleet currently sets both `broker.uriEnv` and `coordinator.uriEnv` to the same variable,
`LAVINMQ_URI`. Therefore its coordinator connects to `/mac`. Cross-host `fleet_send{coordId}` routes
only when both leads share the same coordinator vhost. If the fleet01 lead uses `/fleet01` for its
coordinator, the leads cannot see each other and the send will not route.
**Open question:** the fleet01 coordinator vhost has not been checked. Surface this question in
every report until a live `lead.<coordId>.inbox` consumer or fleet01's config proves the answer. Do
not claim that fleet01 uses `/fleet01` just because its member queues do.
Also compare these broker facts with the `leads` rows from local `fleet_list`. A missing remote lead
is `not visible from this coordinator`, not `down`, unless the broker consumer facts prove it.
**What this tier cannot see:** without Tier 2 management access it cannot list lead inboxes or their
consumers. Even with that access, a stopped fleet01 daemon leaves only durable queue history. That
history cannot prove which coordinator URI its current config would use after restart.
## 5. Report all fleets
Use one row per known or discovered fleet. Include blocked rows.
| Fleet | Daemon | Deployment | Herdr | Members/capacity | Queues/consumers | Lead coordination | Cannot check |
|---|---|---|---|---|---|---|---|
| Mac (`/mac`) | PID + uptime | jar vs `HEAD` | health + version + protocol | `fleet_list` + `healthCoverage` | facts or blocked reason | inbox facts or open question | exact missing facts |
| fleet01 (`/fleet01`) | reachable/down/unknown | value or `not reachable` | value or `not reachable` | value or `not reachable` | facts or blocked reason | inbox facts plus coordinator-vhost question | exact missing facts and fix |
Add rows for unknown vhosts. End with three short sections: `Current warnings`, `Checks that were
blocked`, and `Operator action`. Until the management user exists, `Operator action` must say:
> Create a read-only LavinMQ management user with the `monitoring` tag and access to `/mac` and
> `/fleet01`. Put its user and password in `${SHARED_ENV}/tools/secrets.sh` as
> `LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD`.
+266
View File
@@ -0,0 +1,266 @@
---
name: handover
description: Procedure for an outgoing lead to write the handover file that a fresh lead session inherits. Load this when your context is filling up and you are about to be replaced, whether you hand off by hand or fleetd does it for you. The file is the new lead's only inheritance — follow it exactly.
---
# Handover — write the file the next lead depends on
A lead session fills up its context and has to be replaced by a fresh one. The outgoing lead
writes a handover file, and the new session reads that file and carries on.
**There are two ways to hand off, and the file is the same either way.**
- **By hand.** You write the file, then tell the operator where it is. The operator starts the new
session and points it at the file. This always works.
- **With `fleet_handover`** (fleetd #480, merged 2026-09-11). You ask fleetd to do the swap: it
checks the file, clears your pane, and tells the fresh session to read it. This needs
`leadRollover:` in `fleetd.yaml`; without it every action answers a clean refusal naming
`NOT_CONFIGURED`, and you fall back to the manual path. Section 11 below is the procedure.
Nothing else in this skill changes between the two. Only who performs the swap changes.
**The new lead's only inheritance is that file.** It does not see your conversation, your plan,
or your screen. If the file is thin or wrong, the new lead re-derives what you already knew, and
that wastes hours. Writing a good handover file is real work. It is not paperwork you rush
through at the end of a session.
This skill is the procedure for writing it. Every rule below earned its place because a past
handover got it wrong.
## 1. Confirm you are the right session to write this
Run `fleet_whoami` first. It must answer `primary`. Only a primary (lead) session writes a
handover file. A worker's job ends with its own pull request, not a fleet-wide handoff.
## 2. Every number needs a command, run in this turn
A number is a claim: a count, a commit hash, a process id, a percentage, a queue depth. Before
you write one, run the command that produces it — now, in this turn, against the live state.
Never take a number from:
- earlier in your own conversation — the state has moved since then,
- a peer lead's report — that is their measurement, not yours,
- your own memory of an earlier session.
Put the command, or its real output, next to the number. That lets the next lead re-run it and
check it still matches. If you cannot measure something yourself, say so instead of guessing:
"the fleet01 lead reports 91 commits behind; I have not checked this myself."
## 3. Say what you measured and what you did not
Mark every claim as one of two things:
- **"I checked this myself, in the code or on this host, at `<time>`."**
- **"I did not check this myself; `<who>` reported it."**
Never present someone else's measurement as your own. This matters most for cross-host claims —
a peer lead's daemon, a worker's report, or something the operator said earlier that you cannot
re-verify from here.
## 4. Record open decisions, and who owns them
List three things:
- what the operator actually asked for, in their own words where you have them,
- what is still unanswered,
- any question you decided yourself instead of asking, with your reason.
Write the decision so it cannot be mistaken for the operator's instruction. Say plainly: "the
operator never answered X; I decided Y, because Z." Without this, the next lead either silently
reopens a closed question or assumes the operator chose something they never did.
## 5. Record live hazards
List anything that will break if the next lead does the obvious thing next. This includes:
- unpushed commits or unmerged branches,
- a build, a spawn, or a redeploy still running,
- code merged to `main` but not yet redeployed to the live daemon,
- any trap that looks safe and is not — say what goes wrong and why, not only that something is
"tricky."
## 6. Record what is explicitly not owed
List work that is finished, and work that another party has said they do not want touched. Name
who said so and when. Without this line, the next lead re-does closed work or reopens a question
a peer already declined to revisit.
## 7. Open the file with three re-measurement commands
The file's own first section must give the next lead three concrete commands to run before
acting on anything else in the file:
1. confirm role — for example `fleet_whoami`,
2. confirm the state of the working tree — for example `git status` and
`git rev-list origin/main..HEAD`,
3. read the live fleet — for example `fleet_list`.
Record what each command answered when you wrote the file, and tell the reader to run it again
rather than trust your answer. The point of this section is that the reader checks live state
before acting on any claim in the rest of the file, including yours.
## 8. Stamp the file with time and commit
Near the top of the file, write:
- the date and time you wrote it,
- the commit the tree was on (`git rev-parse HEAD`),
- whether the tree was clean (`git status`).
Without this, nobody can tell how old the file is, or which code it describes.
## 9. State plainly that the file goes stale fast
Say near the top: **re-measure anything you act on.** The file goes stale the moment anyone
merges a branch, spawns a member, or restarts the daemon. Everything in the file is a snapshot
of one moment, not a live fact.
## 10. What to leave out
Do not include:
- narration of how the session felt, or how hard something was,
- anything the repo already records — code structure, git history, or a rule already written in
`CLAUDE.md`. Point at it instead of repeating it,
- advice that is only true for the session that is ending — a half-open terminal, a local
variable, a train of thought with no state behind it.
A handover file is a record of state and decisions. It is not a diary.
## 11. Using `fleet_handover` (only if `leadRollover:` is configured)
**Run the three steps in this order. The order is not a style choice — the wrong order is
refused.**
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
`handoverPath` you must write to. Nothing has happened to your pane yet.
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
own workspace before it hands it to you. The daemon and your pane can run in different
directories, so a path you resolve yourself can point at a different file from the one the daemon
will check.
**Check that the path is ignored by git before you write to it (#491).** A relative
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
will show up as untracked content in that repo. The file is a snapshot of live state and must
never be committed, so tell the operator rather than committing it or silently editing their
`.gitignore`.
2. **Write the handover file at that path**, following sections 1–10 above.
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
Why that order: `confirm` refuses with `HANDOVER_STALE` unless the file was modified **after** the
`open` request. That check stops a leftover file from an earlier session being accepted as this
one's handover. So writing the file first and then calling `open` — the obvious order — always
fails.
`{action: "cancel", token}` drops a pending request without rolling.
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
get another one.
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
from your connection, so you can only ever roll yourself.
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
are confident. Ask, wait for the answer, then pass what they said.
- **Whether you must ask at all depends on `leadRollover.requireOperatorConfirm`. Check it; do not
assume.** The default is `true` (`FleetConfig.java:1426`), and then `confirm` refuses unless you
also pass `operatorConfirmed: true`. **This host set it to `false` on 2026-09-22**, on the
operator's explicit grant, because they do not want to approve routine context rolls. Where it is
`false`, the three handover-file checks are the whole gate: the file must exist, be fresher than
`maxDocAgeSeconds`, and have been modified after the open request.
Read the live value rather than trusting this line:
```bash
grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml
```
No match means the key is unset, so the default `true` applies and you must ask. The key is
**deferred, not hot** — it is read once at boot, so an edit does nothing until the daemon is
redeployed.
**Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the
daemon no longer requires it.** `LeadHeartbeatLoop.contextNotice()` hardcodes "ask the operator"
and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is
merged and deployed, the nudge matches the config and this warning can be deleted.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
handover file, with the configured `bootstrapText` arriving as its first message. No context was
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
at all. Re-measure both numbers with:
```bash
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
grep -c "lead-rollover:" fleetd/fleetd.out # positive control: must be larger
grep -c "Unknown command" fleetd/fleetd.out # the old failure: expect 0
```
Run the control line too. A broken pattern returns a clean `0` that reads exactly like good news.
If the first number stops growing across rolls, or `Unknown command` returns anything above 0,
the bootstrap has regressed and this paragraph is stale again.
**You still write the file before you confirm, and never the other way round.** That order is not
about the bootstrap being unreliable. It is what the daemon checks: the handover file must have
been modified *after* the open request, or `confirm` refuses it as stale.
- **One warning in the log is normal and is not a failure.** Every one of the three rolls above also
logged `/clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls —
releasing rather than wedging the roll`. The daemon could not see the pane go WORKING after
`/clear`, so it released instead of hanging. The roll then succeeded anyway. That is the safe
branch behaving correctly. Do not report it as a broken roll.
## Writing style
Write in plain English. Use everyday words, one idea per sentence, and active voice. Keep every
class, method, file, flag, and config key exactly as it appears in the code — replacing a
precise term with a vague one makes the sentence wrong, not simpler. Explain an abbreviation the
first time you use it.
If a diagram genuinely helps, put it in the `.md` file as a fenced ` ```mermaid ` block with no
hardcoded colors, so it stays readable on light and dark backgrounds. Quote any label that has
brackets, colons, or slashes.
## Template
```markdown
# Handover — <fleet name> lead session, <date and time>
Written at commit `<output of git rev-parse HEAD>`. Tree was <clean, or dirty: `<git status
summary>`>. Re-measure anything you act on — this file goes stale the moment anyone merges,
spawns, or restarts.
## 0. Do these three things first
1. Confirm your role: `fleet_whoami` — must answer `primary`. (Answered `<result>` at `<time>`.)
2. Confirm tree state: `git status`, `git rev-list origin/main..HEAD`. (`<result>` at `<time>`.)
3. Read the live fleet: `fleet_list`. (`<result>` at `<time>`.)
## 1. What the operator asked for
<the live instructions, in their words where you have them; what is still open; any decision
you made yourself, and why>
## 2. Open decisions, and who owns them
<one line per decision: who owns it, what is unanswered>
## 3. Live hazards
<one entry per hazard: what breaks, and why, if the next lead does the obvious thing>
## 4. What is not owed
<finished work, and work another party has declined; name who said so and when>
```
+102
View File
@@ -0,0 +1,102 @@
---
name: hunter
description: Defect-hunt procedure for a fleetd worker — sweep an assigned package for real bugs and report several ranked findings without fixing anything. Load this when the lead asks you to hunt or audit a scope rather than review one diff. Do NOT load `reviewer` for this; the two want different output.
---
# Hunter worker — procedure
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies.
**This skill is not `reviewer`.** `reviewer` judges one diff and reports the *single* most
important issue in about 90 words. A hunt sweeps a whole package and reports *several* findings
in a long structured form. Loading both gives you two contradictory output contracts, and the
usual result is a worker that writes a good report into its terminal and ends the turn without
sending it. Load exactly one.
## 0. Read this before you read code: how the report gets home
Your terminal reaches nobody. The lead sees **only** the text inside your `fleet_reply` call.
A long report is exactly the case where this goes wrong, so plan for it:
- **Write the report into the `fleet_reply` argument itself.** Do not compose it in your terminal
and then summarise it into the call.
- If the report is long, **send it anyway** — one `fleet_reply` with everything.
- If you end the turn without replying, the bridge scrapes your pane instead. That scrape carries
at most the last 4000 characters, and on a hunt it usually captures the tail of the lead's own
brief rather than your findings. The lead then has nothing and has to ask you again.
## 1. Change nothing
A hunt is read-only. Do not edit a production file, do not "quickly fix" what you find, and do
not run a formatter. You may run the build and tests to *check* a claim, and you should say so
when you did.
## 2. Read the whole scope first
Read every file in the assigned package before you judge any of it. A defect that a caller
elsewhere in the same package makes unreachable is not a defect, and you cannot know that from
one file.
Stay inside the scope. If a defect there depends on a class outside it, read that class to
confirm — but the defect itself must live in the scope you were given.
## 3. The bar — this matters more than the count
**Name the path into the bad state.** Say which caller, in which state, reaches it. A defect on
paper is not a reachable defect. If you cannot name that path, keep the finding but mark it
`unproven` and say exactly what you could not check. Do not drop it, and do not dress it up.
**Say which direction the harm goes.** Data loss, privilege escalation and silent wrong answers
are worth reporting even when the window is narrow. A finding whose worst outcome is a worse log
line is not worth a block.
Two workers once ran the same scope: the one that applied the direction-of-harm filter found ten
real defects, the one that did not found none. Fewer findings the lead can act on beat many the
lead has to triage.
## 4. Shapes that have produced real merged fixes here
Read for these first:
1. **A one-way gate.** A guard added after an incident closes only the direction that incident
came from. Do not only ask what closes the gate — ask **which states still open it**.
2. **A value read once, then used later to authorise something destructive**, after something
else has had a chance to change it.
3. **A failure downgraded to a value that looks like a legitimate result** — `-1`, `null`, an
empty list, `false` — which a caller then trusts.
4. **A lock held for one half of a read-modify-write and not the other**, or two collections
updated under different locks.
5. **A comment or javadoc stating an invariant the code no longer keeps.** Comments are
load-bearing in this repo; a stale one has already caused a bug.
## 5. What you cannot check, and must not claim you did
- `fleetd/fleetd.yaml` is gitignored and **absent from your worktree**. You cannot read it. If a
finding depends on live configuration, name the key and say you could not check it.
- `.mcp.json`, `opencode.json` and `.autoenv` in your worktree are neutralised stubs, not the
repo's real files.
- The `wiki/` submodule pointer is months old. Do not cite it.
Reporting a fact you took from the lead's brief as something you measured yourself is a false
report, even when the fact is correct. Say where each fact came from.
## 6. The report — what goes in `fleet_reply`
One block per finding, most severe first:
```
FINDING N — <one line>
file:line
Path in: <which caller, in which state, reaches this>
Direction: <data loss | escalation | silent wrong answer | outage | ...>
Window/trigger: <when it actually happens>
Confidence: <confirmed by reading | unproven — say what you could not check>
Why nothing else catches it: <the guard or test you checked, and why it misses>
```
End with one line naming every file you read, so the lead knows the denominator.
**Nothing clears the bar?** Reply `NO FINDINGS`, name the files you read, and say what you ruled
out. A clean sweep is a valid result; an invented defect is worse than none.
+19 -6
View File
@@ -1,11 +1,11 @@
---
name: implementer
description: Implementer-role procedure for a bridged worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over bridged.
description: Implementer-role procedure for a fleetd worker — verify your worktree, implement the scope, commit, push, open your own PR, and hand off the PR URL. Load this when the lead delegates you an implementation task over fleetd.
---
# Implementer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
@@ -39,6 +39,19 @@ a worker made all 59 of its edits in the primary's tree and never noticed.
test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-toplevel)"
```
**Never run `git stash` (or `git stash pop`/`apply`/`drop`).** Your worktree is isolated, but the
stash is **not**: `refs/stash` is one stack shared by the primary's checkout and every other
worker's worktree of this repo. Measured on 2026-09-04 — `git stash list` from a worker's worktree
and from the primary's tree returned byte-identical output. So a `git stash` you run can be popped
into someone else's tree, and a `git stash pop` you run can drop **another worker's** uncommitted
edits on top of yours. This has already happened here: two workers were running in parallel and one
of them had its in-progress edit silently overwritten by the other's stash.
The branch is your isolation, so use it instead. To set work aside, commit it on your own branch
(`git commit -m "wip: ..."`) and carry on; to try something and back out, use
`git diff > /tmp/<your-branch>.patch` then `git checkout -- <file>`. Both stay inside your worktree.
If you find a stash entry you did not create, leave it alone and say so in your report.
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
@@ -49,7 +62,7 @@ test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-t
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
cd "$(git rev-parse --show-toplevel)/fleetd" && mvn clean install
echo "exit=$?"
```
@@ -85,7 +98,7 @@ host (`GITEA_HOST`) into your env for exactly this — the token can create a PR
merge**.
```bash
API="${GITEA_HOST%/}/api/v1/repos/lms/claude-bridge/pulls"
API="${GITEA_HOST%/}/api/v1/repos/fleet/fleetd/pulls"
BRANCH="$(git branch --show-current)"
curl -sS -X POST "$API" \
-H "Authorization: token ${GITEA_TOKEN}" \
@@ -103,7 +116,7 @@ fix it if the cause is yours (e.g. branch not pushed yet), and report the failur
inventing a URL. If `GITEA_TOKEN` is unset your profile was not granted PR-create: push the branch
and report its name so the lead opens the PR.
## 6. Hand off — what goes in `bridge_reply`
## 6. Hand off — what goes in `fleet_reply`
The reply is the entire handoff; the lead cannot see your terminal.
@@ -130,7 +143,7 @@ sequenceDiagram
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
G-->>I: html_url
I->>L: bridge_reply(PR url, branch, files, tests)
I->>L: fleet_reply(PR url, branch, files, tests)
Note over L,G: the lead reviews the PR and merges on green — you never merge
```
+20 -12
View File
@@ -53,7 +53,7 @@ Map each entry from `.mcp.json`:
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"fleetd": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
@@ -92,24 +92,32 @@ route by who launches opencode:
| Launcher | Route |
|---|---|
| a human, from a terminal | `{file:…}` — see below |
| a human, from a terminal | one central store, sourced by the login shell |
| a spawner (bridge, CI, IDE) | `{env:…}`, with the spawner injecting the variable |
For the human case prefer **`{file:…}` with a workspace-relative path**, kept in a gitignored
`.secrets/` directory beside `opencode.json`:
**For the human case, keep every credential in one file the login shell sources.** Here that file
is `${SHARED_ENV}/tools/secrets.sh`, sourced from `${SHARED_ENV}/.ltms`, kept at mode 600 and never
committed. `opencode.json` then names variables and holds no values:
```json
"headers": { "Authorization": "Bearer {file:.secrets/api-token}" }
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" }
```
Verified: opencode resolves relative `{file:}` paths against the project root, so this needs no
shell setup at all — no rc export leaking the secret to every process, no direnv dependency.
Both routes end at the same syntax, and that is the point. The file does not change when a human
launches opencode instead of the bridge.
**The catch, and state it out loud:** `.secrets/` is gitignored, so a peer running in a git worktree
does **not** get it — worktrees receive tracked files only, the same rule that makes `opencode.json`
itself worth committing. Spawned peers must therefore be fed through `{env:…}` by whatever launches
them. Check the variable *names* match: a spawner often injects under a different name than your
shell uses, and the config has no fallback.
**`{file:…}` also works, and this project moved away from it.** Opencode resolves a relative
`{file:}` path against the project root, so a gitignored `.secrets/` beside `opencode.json` needs no
shell setup at all. It did not fail; the problem is that it makes a second copy of the token. The
same secret then lives in two places, and the copy you forget is the one that leaks or goes stale.
One store with many references is easier to rotate and to audit.
**One catch survives either choice, so state it out loud:** what a human's shell exports does not
reach a spawned peer, and neither does a gitignored `.secrets/` — a git worktree receives tracked
files only. Spawned peers must be fed through `{env:…}` by whatever launches them. Check the
variable *names* match: a spawner often injects under a different name than your shell uses, and
the config has no fallback. In this repo the bridge goes further: it neutralizes a worktree's
`opencode.json`, so a member cannot inherit the primary's credentials by accident.
## 4. Do not port machine-local MCP servers
+72
View File
@@ -0,0 +1,72 @@
---
name: redeploy-fleetd
description: Rebuild and restart the live fleetd daemon after a merge (lead / primary only). Load this before redeploying — it holds the script, the drain step, the permission grant, and the five checks that have each gone wrong here before. Workers must never do this.
---
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-fleetd.sh --check # report state, change nothing
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `fleet_list`, then `fleet_stop` each member, and collect anything
you still want with `fleet_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-fleetd.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
+9 -4
View File
@@ -1,15 +1,20 @@
---
name: reviewer
description: Reviewer-role procedure for a bridged worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over bridged.
description: Reviewer-role procedure for a fleetd worker — how to work a review scope and the exact shape of the finding to report. Load this when the lead delegates you a code review over fleetd.
---
# Reviewer worker — procedure
The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, honest reporting,
The turn contract (one `fleet_reply`, `fleet_ask` for the lead's decisions, honest reporting,
never merge) is in **`CLAUDE.md` → Bridge communication → Worker** and already applies. This
skill is only the *review procedure*: how to work the scope, and the exact shape of what you
send back.
**Wrong skill for a sweep.** This one reviews *one* diff or scope and reports the *single* most
important issue. If the lead asked you to hunt or audit a whole package for several defects, load
`hunter` instead and ignore this file — the two want different output, and following both is how a
worker ends its turn with a good report that never gets sent.
## 1. Read the whole scope before you judge
The delegation names your scope — a file, a diff, a PR, a function. **Read all of it first.**
@@ -24,13 +29,13 @@ wrong.
covers what it was given.
- Do **not** edit files or run the build. You review; the owner acts.
## 3. Reach for `bridge_ask` only for a genuine fork
## 3. Reach for `fleet_ask` only for a genuine fork
Ambiguous requirement, a missing acceptance criterion, "intended or a bug?", or two defensible
fixes with different consequences — those are the lead's call, and guessing produces a
confident-but-wrong finding. Anything you could settle by reading more code is yours to settle.
## 4. The finding — what goes in `bridge_reply`
## 4. The finding — what goes in `fleet_reply`
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
+37 -12
View File
@@ -14,7 +14,7 @@ jobs:
# does not depend on the wiki repo being reachable.
- uses: actions/checkout@v4
# The runner image ships an older default-jdk; bridged sets maven.compiler.release=25, so
# The runner image ships an older default-jdk; fleetd sets maven.compiler.release=25, so
# provision the JDK explicitly rather than apt-installing whatever "default" means today.
- name: Set up JDK 25
uses: actions/setup-java@v4
@@ -32,13 +32,17 @@ jobs:
mvn -version
- name: Build and test
working-directory: bridged
working-directory: fleetd
# This IS the mock-socket surface CB-503 asks for: the pom's `default-excludes` profile
# already sets excludedGroups=contract, so the @Tag("contract") tests — which need a live
# herdr socket and a RabbitMQ container — are excluded without any flag here. Everything
# that runs does so against the fake UDS herdr and fake ccs/claude stubs.
run: mvn -B clean install
- name: javadoc reference lint
working-directory: fleetd
run: mvn -B -DskipTests javadoc:javadoc -Ddoclint=reference
# Deliberately NOT actions/upload-artifact: this Gitea instance presents as GHES, and
# @actions/artifact v2+ (i.e. upload-artifact@v4) refuses to run there —
# "GHESNotSupportedError ... not currently supported on GHES", which red-Xes an otherwise
@@ -46,7 +50,7 @@ jobs:
# the log instead, where they are actually readable.
- name: Failing test output
if: failure()
working-directory: bridged
working-directory: fleetd
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
@@ -54,12 +58,29 @@ jobs:
done
exit 0
# fleetd #550 — nothing ran scripts/test-redeploy-fleetd.sh in CI before this, on any platform,
# so it had run only on macOS by hand and two Linux-only bugs (this issue's items 1 and 2)
# survived undetected: shasum is a macOS-only tool (it ships with Perl; GNU coreutils, i.e. every
# mainstream Linux distro including this runner's ubuntu-latest, does not have it and ships
# sha256sum instead). The gate here is the step's own exit code, nothing else: a `run:` step in
# Gitea/GitHub Actions already fails the job on a non-zero exit with no extra scripting needed,
# so this deliberately does NOT grep the output for a `FAIL:` count. That is the #550 item-2
# lesson one level up — a suite that dies before it runs a single test prints zero FAIL lines,
# which is exactly what a clean pass also prints, so counting FAIL lines can never be the gate.
shell-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: redeploy-fleetd.sh shell suite
run: bash scripts/test-redeploy-fleetd.sh
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# profiles in fleetd/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
@@ -69,8 +90,6 @@ jobs:
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
@@ -89,16 +108,22 @@ jobs:
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
# tag selects the whole group, so a test added to it later runs here automatically. A prior
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
# noticed until the herdr protocol drifted out from under a test that never ran here
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
# step output rather than assuming which.
- name: Contract tests
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
working-directory: fleetd
run: mvn -B -Pcontract test -Dgroups=contract
- name: Failing test output
if: failure()
working-directory: bridged
working-directory: fleetd
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
+21 -6
View File
@@ -1,12 +1,27 @@
# Workspace-scoped secrets for tools that read no settings cascade of their own
# (opencode resolves these via {file:.secrets/...} in opencode.json).
# No secret belongs in this repo any more: every credential lives in one shell-level store
# (${SHARED_ENV}/tools/secrets.sh), and opencode.json reads it as {env:...}. This line stays as a
# backstop, so a workspace-scoped copy that someone re-creates by habit still cannot be committed.
.secrets/
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
# The default profile parityOverlay copies these primary→worktree, so they appear in EVERY worker
# worktree. Two reasons they must be ignored. They hold environment values, which is reason enough.
# And since CB-576 a release preserves any worktree that `git status --porcelain` calls dirty —
# untracked files included, deliberately. An untracked overlay file would therefore make every
# COMPLETED release preserve its worktree, and worktrees would pile up with no error to notice.
.env
.envrc
# Daemon runtime artefacts. fleetd appends its log wherever it is launched from, so both the
# repo root and fleetd/ collect one; neither belongs in git.
fleetd.out
fleetd/fleetd.out
logs/
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
# has no business in git history.
.handover/
+1 -1
View File
@@ -1,3 +1,3 @@
[submodule "wiki"]
path = wiki
url = ssh://git@git.ltms.dev:2224/lms/claude-bridge.wiki.git
url = ssh://git@git.ltms.dev:2224/fleet/fleetd.wiki.git
+24
View File
@@ -0,0 +1,24 @@
---
description: Refine work into clear, independent units before implementation.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You are an architect in this fleet. You refine work before anyone builds it: scope,
acceptance criteria, risks, and a unit split. You read the repo and write analysis.
You never commit production code and never open a pull request.
A design task is worked by two architects. Design alone first, then exchange and
say plainly where you disagree. Do not concede just to agree.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Report only work you actually did and the real output of checks you ran. Do not
claim a result from a tool you could not use. The primary's IDE tools are not yours.
A mounted forge tool may use a blocked credential and fail by design.
The launcher provides the required bridge reply instructions for every member.
+29
View File
@@ -0,0 +1,29 @@
---
description: Implement one assigned unit, test it, and open a pull request.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You implement the one unit you were given and nothing else. Work in your assigned
git worktree and branch. Never check out, rebase onto, or push to `main`. Confirm
the worktree root and branch before you edit. Use only paths under that root.
Do only the assigned scope. Note anything outside that scope in one line and do not
investigate it further. Use `fleet_ask{question}` only when a decision belongs to
the lead, such as an unclear requirement or two defensible fixes. Do not ask about
something you can decide by reading more code.
Implement the change and run the full required build in your worktree. Read the
complete output and report its real result. Do not hide failures with a pipe. State
only checks you actually ran. The primary's IDE tools are not yours. A mounted forge
tool may use a blocked credential and fail by design.
Stage only files you changed. Never use `git add -A` or `git add .`. Never commit
`.mcp.json` or `wiki/`. Commit with a clear message, push your branch, and open your
own pull request against `main`. Never merge.
Your handoff must name the pull request or why it was not created, the branch, the
files changed, the build result, and any caveat for review.
The launcher provides the required bridge reply instructions for every member.
+34
View File
@@ -0,0 +1,34 @@
---
description: Review one assigned scope and report the most important real issue.
mode: primary
---
<!-- CB-617: The model comes from fleetd.yaml because the launch flag overrides model here on both backends. -->
You review the diff you were given. Report bugs, risks, and missing tests. You do
not change code.
Read the whole assigned scope before judging it. Review only that scope. If you see
something outside it, note it in one line and do not investigate it further. Do not
run the build. The owner makes changes and runs checks.
Use `fleet_ask{question}` only when a decision belongs to the lead, such as an
unclear requirement or two defensible fixes. Do not ask about something you can
decide by reading more code.
Report the single most important real issue in this form:
```
1. <path>:<line>
2. issue: <one sentence: what is wrong and why it matters>
3. fix: <one line: the concrete change>
4. severity: high | medium | low
```
If there is no real issue, report `NO ISSUE` and one line saying why. A clean review
is valid. Do not invent an issue. Use high for a wrong result, data loss, security,
or a hang or crash on a real path. Use medium for an edge-path bug or a correctness
risk under load or concurrency. Use low for clarity, a latent foot-gun, or a smell
with no current failure.
The launcher provides the required bridge reply instructions for every member.
+255 -90
View File
@@ -4,59 +4,72 @@
> **Canonical block.** Everything down to §Layering is the portable bridge charter, copied verbatim
> into every project that mounts the bridge MCP. Keep it byte-identical with the template in the
> wiki ([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable
> wiki ([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable
> CLAUDE.md block*); improvements go to the template first, then out to each project. Anything
> specific to *this* repo lives under §Project addendum below, never inline above it.
>
> **Anything you measure in an addendum is perishable.** Date it, give the command that
> re-measures it and what each outcome means, and tell the reader to delete the section once
> it stops reproducing. The four parts work together: deciding what would falsify a claim is
> the expensive step, and a reader in the middle of another task will not pay it, so a bare
> "verify before relying on this" costs the same space and does nothing. The case this is for
> is a note that goes stale as a live restriction — it will tell a future session it cannot do
> the thing at the moment doing it becomes the job.
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
`fleetd` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `fleet_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are lead-only; reply/ask are
only-as-itself — any peer may answer for its own pane, and for no other. A call outside your
role is refused, not queued.
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
### Primary (lead) — run this on every task, in order
**Delegate by default — that is the job.** With the bridge mounted you are an orchestrator on a
metered subscription, and workers are cheap, parallel, and disposable. The default answer to "who
does this?" is **a worker**, not you. Reach for `bridge_send` before you reach for `Edit`. The steps
does this?" is **a worker**, not you. Reach for `fleet_send` before you reach for `Edit`. The steps
below are the procedure — run them in order, every task, not only the big ones.
0. **Know your role** — `bridge_whoami`, once per session, before anything else.
0. **Know your role** — `fleet_whoami`, once per session, before anything else.
1. **Split.** Write the unit list. Every unit carries: scope · the files or PR in question ·
acceptance criteria · exactly what to report back. A unit with no acceptance criteria is not
ready to delegate — refine it or keep it.
@@ -66,20 +79,37 @@ below are the procedure — run them in order, every task, not only the big ones
delegate. The keep-list is closed: the conversation with the user, decomposition and planning,
the final judgment call, verification, merges, and anything that depends on context only you
hold. Nothing else is yours by default.
3. **Spawn every delegated unit first** — `bridge_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model and cost, not in
tier, so the default is rarely what you want.
4. **Then send them all** — `bridge_send{sessionId, content, wait:false}`. Line 1 of every brief is
3. **Spawn every delegated unit first** — `fleet_spawn{profile, worktree:true, ticket}`, one per
unit, *before* sending any. Pass `profile` explicitly: profiles differ in model, cost and
LIVENESS, not in tier, so the default is rarely what you want. The default is whatever the
daemon reports, and on a host where it sits on an exhausted or withdrawn credential every
unqualified spawn fails — sometimes loudly, sometimes as a member that spawns fine and then
produces nothing. `fleet_profiles` reports the default; check it once per session.
4. **Then send them all** — `fleet_send{sessionId, content, wait:false}`. Line 1 of every brief is
`Load the <name> skill.` naming the worker's playbook; those skills are opt-in and that line is
what makes them reliable. Where the project ships no such skill, spell the procedure out in the
brief instead. The brief is self-contained — the worker sees your message and the repo, nothing
of your context, your plan, or your screen.
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
returns a ticket, and is then never delivered — measured here three times in one session, and the
member was released still executing a brief that had been retracted twice. The receipt is true and
it is a fact about the *mailbox*; what you needed was a fact about the *pane*. **A push delivery
needs the recipient free at send time; a pull channel needs only that they look before acting.** So
put every correction on the **ticket**, which they can read whenever they look, and send as the
notification. That obliges you, not them: **all corrections go to the ticket, and the brief is
write-once.** The member cannot check which source is newer — it just always prefers the ticket —
so the day you revise a brief in place instead of commenting, it obeys your rule and does the wrong
thing. A *first* brief for a unit not yet running is not a correction, and may be the whole spec.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge MCP server it appears to have holds a blocked credential and fails every call, and a piped
command (`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a
fact. Its injected repo-scoped `GITEA_TOKEN` is a different credential and does work, so a worker
reporting that it opened its own PR is reporting something it really can do.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -87,40 +117,66 @@ below are the procedure — run them in order, every task, not only the big ones
read it yourself.
8. **Adjudicate, merge, tear down — yours alone.** Read the diff yourself: fully if it is small,
targeted at the reported findings and the risky paths if it is large. Reviewer findings direct
your attention; they never substitute for it. Then merge, then `bridge_stop{paneId}`.
your attention; they never substitute for it. Then merge, then `fleet_stop{paneId}`.
**If the forge refuses you the merge** — a protected branch, a token without the grant — the
adjudication is still yours. Read the diff, decide, and hand the operator a merge-ready queue
with the refusal quoted. Never report a PR as merged, and never call one "ready to merge"
without having read the diff yourself. A refusal is exactly when that shortcut is tempting,
because no action is left that forces you to look, and taking it turns this step into
forwarding a reviewer's verdict — which is delegating the merge by proxy, two lines above.
**Test a refusal; do not read it off a permissions field.** A protected branch holds its merge
rights separately from the repository permissions, so that field can say yes while the merge is
refused, and still say no after a grant makes it work. Probe instead, with a request that cannot
succeed on its merits, so a rejection can only mean the refusal. Treat a transport failure as a
third answer that proves nothing: a timeout, a DNS error or a bad URL is not a refusal, and
counting it as one makes you sure of something you never measured.
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by
prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
**When a decision blocks you, consult architects — not the operator.** Spawn one or more architect
members, give them the question and the evidence you have, and act on what they agree. They are
authorized to settle it, not only to advise. Architects first form independent positions, then
compare them. If they still disagree after that comparison, they return both positions and their
checked evidence; the lead decides. Go to the operator only for an action the fleet has no
authority to take, such as spending money, granting access, or making a promise to someone else.
**Then write the decision on the ticket.** Taking the operator out of the loop also removes the
signal they used to get, because that signal was the block itself — work stopped, so they found
out. A ticket comment replaces it, and it reaches them whether or not they are at a terminal when
you decide.
| Intent | Tool |
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `workers` · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
| Confirm your own role | `fleet_whoami` |
| See backends available | `fleet_profiles` |
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `workers`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own workers, and its own judgment. An empty
`workers` array means no workers are spawned; it says nothing about peers.
`fleet_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to workers — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a worker's
**A lead never assigns a task to another lead.** Work goes to members — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a member's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
worker and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
member and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
@@ -133,53 +189,152 @@ The traffic between leads is coordination and nothing else:
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
not varying. N *agreeing measurements* are also one data point when they share an instrument —
two hosts, two operators and the same formula is one formula, not two confirmations.
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
Being messaged by a peer does not make you its worker: answer the way you would open —
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
simply complies has thrown away the reason there are two of you.
### Worker — the turn contract
### Member (worker or architect) — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
3. **`bridge_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
3. **`fleet_ask{question}`** when a decision is genuinely the lead's (ambiguous requirement, two
defensible fixes, "bug or intended?"). It blocks and you resume the *same* turn with the answer.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
4. **Re-read the ticket before you act on anything you were told earlier**, and again before you
commit. A message reaches you only while you are free to receive it; the ticket is there whenever
you look. **If a ticket comment contradicts your brief, the ticket comment is newer and it wins.**
5. **End the turn with exactly one `fleet_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `fleet_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
6. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge MCP server you
may find there holds a deliberately blocked credential and fails every call, by design. That is
not your only forge route, and the two must not be confused: the repo-scoped `GITEA_TOKEN` the
daemon injects into your environment does work, and using it to open your own PR is part of the
job. A blocked MCP tool is never a reason to skip that step.
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role agent definition files | role contract and per-job procedure | a member whose launcher binds its role to the matching file in its worktree |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
`CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they* must obey belongs in the
charter, not here.
A rule belongs in **exactly one** layer — the outermost one that must obey it. A member without a
repo checkout still gets the launcher's reply charter, which is why that one rule stays there.
Peers that don't read `CLAUDE.md` (non-Claude adapters) get the charter only, so any rule *they*
must obey belongs in the charter, not here.
## Project addendum — claude-bridge (not part of the canonical block)
- **This repo is the bridge.** The daemon is `bridged`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/BridgeMcp` (tools), `auth/Authz` (the role table),
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
throwaway workspace that the test tears down. This only covers
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
banned. That is the control plane that invariant 5 protects. Re-measure with
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
result means tests still open the socket and this note still applies. An empty result means nobody
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
is not part of this change.
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
distinct backend errors (a non-exhaustion failure such as an HTTP 5xx) within 60 seconds — a
short, fixed 60s cooldown, not configurable per profile. Each check runs independently, so a
profile can show both at once. In the JSON: a cooling profile carries `credentialId` and
`coolingOffForSeconds`; a quarantined profile carries `quarantinedForSeconds`; a profile hit by
both carries all three fields, and either state alone already sets that profile's `free` to `0`.
A `fleet_spawn` naming a cooling-off profile is refused before it ever reaches the backend
adapter, with a message naming the credential and the remaining seconds ("cooling off after
repeated backend errors") — distinct wording from a quarantine refusal, so don't conflate the
two when reading a spawn failure.
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR),
`reviewer` (one diff → one structured finding) and `hunter` (sweep a package → several ranked
findings, change nothing). Name exactly one in every delegation. **`reviewer` and `hunter` are
not interchangeable** — `reviewer` caps the answer at one finding in about 90 words, so naming
it for a multi-finding sweep hands the worker two contradictory output contracts. That has
already cost three workers' turns: each wrote a good report to its terminal and ended the turn
with no `fleet_reply`, and the scrape returned the tail of the brief instead.
Spawn `implementer` with role `dev`, `reviewer` with role `reviewer`, and `hunter` with role
`hunter`.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
`port-to-opencode` (make an OpenCode session a participant in this workspace),
`fleets-status` (report every fleet that shares one LavinMQ instance),
`redeploy-fleetd` (rebuild and restart the live daemon after a merge) and
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
unmentioned by every instruction file, so it drifted and a later session planned it from scratch
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
committed copies would otherwise mount the primary's IDE and forge servers (fleetd #134). The
worktree's copy of each is a stub, **not** the repo's real file, so a worker that reads one and
reports what it found is reporting on the stub. The daemon logs a per-spawn summary, but the
worker cannot see that log. From inside its own worktree a worker — or a lead debugging one —
reads the list with `git config --worktree --get-all fleet.neutralizedConfig`, and the
consequence with `git config --worktree --get fleet.neutralizedConfigNote`. Never brief a worker
to edit one of these files: the edit cannot be committed, and it will not tell you so.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, the turn-done
fallback and status gating — are diagrammed in `docs/MCP-Contract.md`. That page is now flows
only: its pre-build tool catalogue, parameter tables and REST paths were deleted rather than
corrected, because a hand-maintained second copy of the tool surface is what drifted for a month
while this line pointed every session at it (CB-609 / #114). **The live MCP schema is the tool
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `fleetd` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. The lead **may and
should** redeploy rather than hand the job back to the operator. **Workers must never do this** — a
worker has no business restarting the daemon it is talking through, and stopping it kills the
worker's own channel mid-turn.
**Load the `redeploy-fleetd` skill before you redeploy.** It holds `scripts/redeploy-fleetd.sh`
and its flags, the drain step, the operator's permission grant, and the five checks that have each
gone wrong here before. Do not hand-roll the steps from memory.
### The prompt is part of the product — update it with the code (mandatory)
@@ -192,15 +347,15 @@ Before you call any work done, check the row that matches what you touched:
| You changed… | Re-read and update… |
|---|---|
| a `bridge_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| a `fleet_*` tool — added, removed, renamed, or its params/semantics | the primary's intent→tool table; any rule that names that tool |
| `Authz` / the role table | invariant 3, and the primary-only vs worker-only claims |
| `ConnectionIdentity` / how a caller is resolved | the `bridge_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__bridge__*`), and the layering table's top row |
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
@@ -211,7 +366,7 @@ is a Roadmap line. A change that touches none of the three earns no entry, and t
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
([Use Cases](https://git.ltms.dev/fleet/fleetd/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
block*) must stay byte-identical, and other projects carrying the block need the same edit. Verify
rather than trust:
@@ -227,6 +382,16 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
PY
```
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
prints a commit with no `-`, delete this paragraph.
## IDE MCP tools & validation workflow (enforced)
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
@@ -234,10 +399,10 @@ PY
Two IDE MCP servers are connected: **intellij-index** (semantic code intelligence) and
**jetbrains** (file problems, reformat, debugger). IntelliJ has multiple projects open; our
module is **`bridged`**. Always pass these to IDE MCP tools:
module is **`fleetd`**. Always pass these to IDE MCP tools:
- `project_path` = `/Users/dai.ha/LTMS/claude-bridge/bridged`
- IDE paths are relative to `bridged/` (e.g. `src/main/java/dev/ltms/bridged/...`)
- `project_path` = `/Users/dai.ha/LTMS/claude-bridge/fleetd`
- IDE paths are relative to `fleetd/` (e.g. `src/main/java/dev/ltms/fleet/...`)
### After editing any file — mandatory
@@ -251,7 +416,7 @@ module is **`bridged`**. Always pass these to IDE MCP tools:
whole-project gate before declaring work done or committing.
**Whenever dependencies change (or a `pom.xml` edit), validate CVEs with
`jetbrains get_file_problems{filePath: "bridged/pom.xml"}`** — its Mend.io check reflects the
`jetbrains get_file_problems{filePath: "fleetd/pom.xml"}`** — its Mend.io check reflects the
dependencies on disk. (Note: `ide_diagnostics` / intellij-index does NOT re-resolve dependencies
after a pom edit without a full Maven reimport, so it reports stale CVE results — don't trust it
for this.) Treat a CVE warning like any other: bump to a patched version and confirm
+32 -29
View File
@@ -9,32 +9,33 @@ Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (which drives a he
process*, so it inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a
cheaper/local model.
## Leading approach — herdr-centric message server (`bridged`)
## Leading approach — herdr-centric message server (`fleetd`)
A small always-on message server, **`bridged`**, controls
A small always-on message server, **`fleetd`**, controls
[herdr](https://herdr.dev) (an agent multiplexer) over its Unix-socket API and exposes a
clean 2-way messaging API as an **MCP server that both the primary and the workers mount** —
one unified Claude setup and the **sole communication gateway** (REST/SSE stays for non-Claude
clients; any broker is `bridged`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns
clients; any broker is `fleetd`-internal, below the gateway).
herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `fleetd` owns
policy (subscription boundary, session lifecycle, status-gated delivery) and the client
contract. The worker `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev` + a
bearer token; the primary Opus stays env-clean and calls `bridged`'s MCP tools.
contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at the gateway,
`https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and calls
`fleetd`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["bridged — standalone daemon (not a claude process)"]
subgraph BD["fleetd — standalone daemon (not a claude process)"]
SRV["SERVER face<br/>MCP · REST/SSE · policy"]
CLI["CLIENT face<br/>status-gated injector · herdr socket"]
SRV --> CLI
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
M["ollama.ltms.dev<br/>(worker model)"]
M["llm.ltms.dev<br/>(the one gateway)"]
OPUS -->|"MCP bridge_send (blocks)"| SRV
W -.->|"MCP bridge_reply"| SRV
OPUS -->|"MCP fleet_send (blocks)"| SRV
W -.->|"MCP fleet_reply"| SRV
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
HERDR -->|"drives PTY"| W
W -->|"inference"| M
@@ -46,24 +47,26 @@ flowchart LR
```
- **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on
Pro/Max). Only the *secondary* process is off-subscription — and `bridged` itself is a
Pro/Max). Only the *secondary* process is off-subscription — and `fleetd` itself is a
plain daemon (no Anthropic quota), so it may poll/subscribe freely.
- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every
- **One gateway (unified MCP setup):** `fleetd` is the **sole communication path** for every
Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add`
line, same on both) and talk over MCP tools — `bridge_send` / `bridge_reply` /
`bridge_status` (with `bridge_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
- **How the primary consumes a reply:** a single **blocking MCP call** (`bridge_send`);
`bridged` holds it open until the worker calls `bridge_reply` or its turn hits
line, same on both) and talk over MCP tools — `fleet_send` / `fleet_reply` /
`fleet_status` (with `fleet_ask` planned for the blocked-worker path). **No Claude session
ever addresses a broker, a peer, or the network
directly**; any queue is `fleetd`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so
mounting the bridge is subscription-safe by construction.
**Tool naming:** the tools are `fleet_*` (renamed from `bridge_*` in CB-622). The old
`bridge_*` names were removed in CB-634 — only `fleet_*` answers now.
- **How the primary consumes a reply:** a single **blocking MCP call** (`fleet_send`);
`fleetd` holds it open until the worker calls `fleet_reply` or its turn hits
`agent_status=done`, then returns the reply as the tool result. No cross-turn busy-poll, so
no quota burn. SSE is an optional side-channel for humans/dashboards watching status.
- **Worker → primary** rides `bridged`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `bridged` **injects the primary's idle pane** when it's
- **Worker → primary** rides `fleetd`'s **MCP rendezvous** — the reply resolves the primary's
blocking call (or, for detached work, `fleetd` **injects the primary's idle pane** when it's
ready), so *no keystroke-into-primary and no broker are involved, even single-host*. The one
exception: a split-host primary that isn't a herdr pane wakes via its own `Stop`-hook, which
polls **`bridged`** (never a broker). See the wiki for the two topologies.
polls **`fleetd`** (never a broker). See the wiki for the two topologies.
- **Different model per process** sidesteps Claude Code's lack of per-subagent provider
routing — the worker isn't a subagent, it's its own configured process.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a
@@ -72,11 +75,11 @@ flowchart LR
## Docs
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/lms/claude-bridge/wiki)**,
Full design, setup, and operations live in the **[wiki](https://git.ltms.dev/fleet/fleetd/wiki)**,
vendored here as a submodule under [`wiki/`](./wiki):
```bash
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/lms/claude-bridge.git
git clone --recurse-submodules ssh://git@git.ltms.dev:2224/fleet/fleetd.git
# or, after a plain clone:
git submodule update --init
```
@@ -86,7 +89,7 @@ Gitea wiki.
## Status
🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in
🟢 **Implemented & dogfooded** — the herdr-centric **`fleetd`** message server is built and in
real use: an Opus primary delegates tasks to off-subscription workers that reply through the
bridge (code reviews delegated this way have produced committed bug fixes). Selected as the
primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a
@@ -97,9 +100,9 @@ tests run separately via `mvn test -Pcontract`):
- **Core gateway** — herdr socket client (contract-tested vs live 0.7.0); guard-checked worker
spawn with `ANTHROPIC_BASE_URL` injected only into the worker's env; status-gated injector;
blocking `bridge_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `bridge_send` / `bridge_reply` / `bridge_status` (messaging) and `bridge_spawn`
/ `bridge_list` / `bridge_stop` / `bridge_profiles` / `bridge_poll` (fleet). Caller identity is
blocking `fleet_send` with reply rendezvous; MCP server as a thin adapter over the REST core.
- **MCP tools** — `fleet_send` / `fleet_reply` / `fleet_status` (messaging) and `fleet_spawn`
/ `fleet_list` / `fleet_stop` / `fleet_profiles` / `fleet_poll` (fleet). Caller identity is
connection-based (loopback peer PID → herdr pane), so the same mount serves primary and workers.
- **Delivery reliability** — completion fallback (a confirmed `working→idle` turn resolves a
send); async fire-and-poll (beats the caller's MCP call timeout for long tasks); and failure
@@ -107,7 +110,7 @@ tests run separately via `mvn test -Pcontract`):
- **Fleet** — multiple worker profiles, each with an independent base_url guard check; workers
inherit the primary's working directory (never `$HOME`); a readiness gate holds delivery until
a worker's Claude has connected the bridge MCP (no paste lost into its boot window).
- **Blocked-worker path** — `bridge_ask` reverse rendezvous: a worker pauses its delegated turn to
- **Blocked-worker path** — `fleet_ask` reverse rendezvous: a worker pauses its delegated turn to
ask the primary and resumes the *same* turn with the answer (CB-205).
- **Session lifecycle** — session manager with spawn/reuse/recycle, `idle_ttl` reaper, `context_cap`,
and graceful drain on shutdown (CB-301/CB-303); per-worker git worktrees on their own branch with
-14
View File
@@ -1,14 +0,0 @@
# Build output
target/
dependency-reduced-pom.xml
# Local runtime config (copy from bridged.example.yaml)
bridged.yaml
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
logs/
# Editor / OS
*.iml
.idea/
.DS_Store
-286
View File
@@ -1,286 +0,0 @@
# bridged configuration (example). Copy to bridged.yaml and adjust.
#
# bridged is the sole gateway between primary/worker Claude sessions and herdr.
# It is NOT a Claude process and must never carry ANTHROPIC_BASE_URL.
# REST + MCP listen address. Keep it on loopback unless you also switch auth.mode to `token`
# below — bridged REFUSES TO START on a non-loopback bind under loopback-trust (see auth).
bind:
host: 127.0.0.1
port: 8765
# API authentication (CB-501). Governs how a caller that is NOT an on-host worker pane proves it
# is the primary. Worker identity never depends on this: a loopback peer PID that maps to a herdr
# pane is unforgeable and is always honoured, so turning auth on cannot lock the fleet out.
#
# mode: loopback-trust → DEFAULT, and the historical behaviour: any loopback caller that is not
# a worker is the primary, no credential needed. Sound ONLY because the
# OS refuses remote connections to a loopback socket.
# mode: token → such a caller must send `Authorization: Bearer <token>`; without it it
# is anonymous and authorized for nothing. REQUIRED for a non-loopback
# bind — the daemon fails fast otherwise, because "unauthenticated ⇒
# primary" on a reachable port would hand spawn/stop/send to anyone.
# tokenEnv → host env var holding the token (never the literal value). Default
# BRIDGED_API_TOKEN. Read only in token mode; empty ⇒ startup fails.
#
# TLS is deliberately NOT terminated in the daemon (CB-501 D3): run a reverse proxy in front and
# let it own certificate lifecycle, e.g.
# location / { proxy_pass http://127.0.0.1:8765; proxy_set_header Authorization $http_authorization; }
# The broker link gets TLS from its own URI (amqps://…) — see `broker` below.
# auth:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# terminal → the ONLY field identity depends on; get it from that session's bridge_whoami
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead is never spawned — it pre-exists, which is exactly why it must be named rather than created.
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges. If both name the same terminal, the `leaders:` entry wins.
# leaders:
# opus-5.0:
# terminal: term_0123456789abcd
# kind: claude
# gpt-sol-5.6:
# terminal: term_fedcba9876543
# kind: opencode
# model: openai/gpt-5.6-terra
# CB-531: FIND LEADS BY TAB NAME instead of pasting terminal ids. `leaders:` above needs an id that
# only exists once the session is running, so adding a lead is: open a tab, start the agent, ask it
# bridge_whoami, edit this file, restart the daemon. This block replaces all of that with a naming
# convention — label the tab `lead: <name>` when you open it and the pane is recognised on the next
# rescan, with no config edit and no restart. Reopen the tab later and the id changes; the label
# does not.
#
# bridged NEVER writes these labels. It renames worker tabs (see `tabLabel` below) but reads lead
# tabs read-only, so what is in the tab bar is always what you typed. Two things keep the convention
# from being a way to claim leadership: the configured worker spaces are excluded from the scan, so
# nothing bridged places can land in a matching tab; and startup REFUSES a `tabPrefix` that any
# worker `tabLabel` also matches, so the two namespaces cannot overlap by accident.
#
# Opt-in on purpose — this widens who resolves as a lead, so upgrading the daemon must never switch
# it on for you. Absent block = leads come only from `leaders:`/`primary:`, exactly as before.
# leadScan:
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive; default "lead:")
# intervalSeconds: 10 # rescan cadence, and the worst case before a new tab is recognised
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
# (--append-system-prompt) as launch flags; nothing is written to the profile.
# tokenEnv → host env var holding the worker's auth token (value never stored in config);
# omit for a backend that needs no token (e.g. a local ollama).
# cwd → pin this profile's working directory (CB-112). Omit to inherit the primary's
# cwd on an MCP spawn, else the daemon's cwd — never $HOME. See
# docs/Worker-Startup-and-Trust.md.
# configDir → CLAUDE_CONFIG_DIR for the worker, so it inherits that profile's
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
# A worker's environment does NOT come from your shell. bridged hands herdr an
# explicit env map and herdr merges it into ITS OWN process env — so before
# CB-511 a worker inherited whatever PATH the herdr server happened to be
# started with, which on a long-lived herdr can predate your toolchain entirely
# and leave workers unable to run `mvn` or `java` at all.
# bridged now propagates ITS OWN PATH to every worker by default; set `env:`
# only to override that or add more (JAVA_HOME, …). Since the default is the
# daemon's PATH, make sure the daemon is started with a good one — see the PATH
# lines in deploy/dev.ltms.bridged.plist and deploy/bridged.service.
#
# Adapter-owned variables always win over `env:`: ANTHROPIC_BASE_URL and the
# rest of the ANTHROPIC_*/CLAUDE_* wiring are applied after it, so an `env:`
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
# The bridge MCP + reply charter mount via a generated OPENCODE_CONFIG file, and the model is a
# `provider/model` selector. Placement, tabs, cwd, and the readiness gate are shared with Claude.
#
# Dogfood-verified 2026-07-29 against opencode 1.18.5 (spawn → readiness gate → bridge_send →
# structured bridge_reply → teardown). The `opencode/*-free` models run on opencode's own gateway
# and need NO credentials — check `opencode models` for the current free list, since the names
# change. That also makes the worker off-subscription by construction.
# opencode-free:
# kind: opencode
# model: opencode/north-mini-code-free # `provider/model` selector, injected as `-m`
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
#
# CB-508: point an opencode profile at your OWN OpenAI-compatible endpoint (local vLLM, llama.cpp,
# LM Studio, TGI…) instead of opencode's gateway. Setting `baseUrl` on a `kind: opencode` profile
# makes the bridge emit a custom `provider` block into the generated opencode.json — opencode has
# no ANTHROPIC_BASE_URL seam, so this is how the endpoint is pinned.
# baseUrl → a bare host:port gets `/v1` appended (where these servers mount the API); a URL that
# already has a path is used verbatim, so a custom mount point still works.
# model → MUST be "<provider>/<model>". The provider half names the generated block; the model
# half must match an id the server reports at /v1/models. One field drives both the
# declaration and the `-m` flag, so they cannot drift apart. A bare model name with a
# baseUrl set is rejected at spawn rather than silently using the default gateway.
# tokenEnv → optional; its value becomes the provider apiKey. Most local servers ignore the key,
# so a placeholder is used when unset (the AI SDK still requires a non-empty one).
# NOTE: no `guard` entry is needed even with a baseUrl set. The SubscriptionGuard exists to stop a
# worker borrowing the primary's Anthropic subscription, and an opencode process has no Anthropic
# credential path at all.
# opencode-local:
# kind: opencode
# baseUrl: http://127.0.0.1:8000
# model: local-vllm/deepseek-v4-flash
# placement: tab
# workspace: bridged-workers
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
defaultWorker: gx10
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Every profile above must have its host listed here.
guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
# unknown keys are ignored, so a snake_case key would be silently dropped (default kept).
# spawnReadyTimeoutMs: 20000
# spawnReadyPollMs: 300
# Worktree provisioning root (CB-301-ext). Where per-worker git worktrees are checked out so
# each worker owns an isolated branch instead of sharing the primary's tree. Omit to default
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
# Set a broker uri to swap in the AMQP-backed inbox: worker replies with no open send are held
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
# pane — status-gated (only when injectable, never mid-turn) and bounded. Ack = drain: the loop
# stops as soon as the primary's inbox is empty.
# terminal → pin the primary's herdr terminal id. Omit to learn it from the connection on
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
@@ -1,426 +0,0 @@
package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.LeadTabScanner;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.auth.ArchitectRegistry;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
* starts listening. Before anything else it asserts its own environment is clean —
* {@code bridged} is not a Claude process and must never carry a base_url.
*/
public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadScan();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateArchitects();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
: UnixSocketHerdrClient.defaultSocketPath();
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
claudeProfiles.put(name, w);
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.defaultProfile(),
cfg.workerProfiles(),
PlacementPolicies.fromName(cfg.placement()),
profileName -> liveCountRef.get().apply(profileName));
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
} else {
log.warn("herdr did not answer within {}s — starting anyway; /healthz will report "
+ "degraded until it comes up. Orphaned worker panes (if any) were NOT reaped.",
HERDR_WAIT_SECONDS);
}
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
int contextCap = 0;
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
if (cfg.lifecycle() != null
&& cfg.lifecycle().idleTtlSeconds() != null
&& cfg.lifecycle().idleTtlSeconds() > 0) {
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
reaper.start();
} else {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the static registry, discover leads by the tab labels the operator
// writes. Opt-in, so a config with no `leadScan:` block resolves exactly as it did under
// CB-530 — the supplier is then a constant and never touches herdr.
final Supplier<Map<String, String>> leads;
if (cfg.leadScan() != null) {
var scan = cfg.leadScan();
Set<String> workerSpaces = cfg.workerProfiles().values().stream()
.map(BridgedConfig.Worker::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
leads = new LeadTabScanner(herdr, scan.tabPrefix(), workerSpaces, leadTerminals,
TimeUnit.SECONDS.toNanos(scan.intervalSeconds()), System::nanoTime);
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, worker spaces {} "
+ "excluded)",
scan.tabPrefix(), scan.intervalSeconds(), workerSpaces);
} else {
leads = () -> leadTerminals;
}
// CB-548: config-declared architect slots. Slots live in config (name → strong-model
// profile); the terminal → slot binding is the live half, sourced from the slots' declared
// terminals today and swapped for a live binding by the later spawn lifecycle. The registry
// is what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
ArchitectRegistry architects = new ArchitectRegistry(
cfg.architects() == null ? Map.of() : cfg.architects(),
() -> cfg.architectTerminals());
if (!architects.slots().isEmpty()) {
log.info("architect slots: {} configured {}, terminals {}", architects.slots().size(),
architects.slots().keySet(), cfg.architectTerminals().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
completion.onTurnComplete(target);
sessions.onTurnComplete(target);
}
@Override
public boolean hasPostTurnAction(String target) {
return sessions.hasPostTurnAction(target);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completion.resolveBeforePostAction(target);
return sessions.onTurnCompleteWithPostAction(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
}
@Override
public void onTurnFailed(String target) {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
// absent, bridged stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-push-").unstarted(r));
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
primaryRegistry.forgetDelegation(terminal); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
if (cfg.auth().tokenMode()) {
String token = System.getenv(cfg.auth().tokenEnv());
if (token == null || token.isBlank()) {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = CallerResolver.withLeadsAndArchitects(identity, true, token, leads,
architects::terminalBindings);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndArchitects(identity, false, null, leads,
architects::terminalBindings);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
poller.stop();
messages.close();
pushLoop.close();
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
if (replyInbox instanceof AutoCloseable closeable) {
try {
closeable.close();
} catch (Exception e) {
log.debug("reply inbox close: {}", e.toString());
}
}
herdr.close();
}));
Javalin app = new BridgedApp(herdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("bridged listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a worker whose
* agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> worker's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link WorkerPresence} — {@code BridgeMcp} marks presence
* only for a worker, deliberately, since that map doubles as the worker roster's availability
* signal and a lead counted there would show up as an available worker. So without the second
* disjunct a lead is permanently un-deliverable: every lead→lead send sat on the gate for
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(WorkerPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
* @return true if herdr answered, false if it never did
*/
private static boolean awaitHerdr(HerdrClient herdr) {
long deadline = System.nanoTime() + HERDR_WAIT_SECONDS * 1_000_000_000L;
boolean waited = false;
while (true) {
try {
herdr.call("ping");
if (waited) {
log.info("herdr is up");
}
return true;
} catch (HerdrException e) {
if (System.nanoTime() >= deadline) {
return false;
}
if (!waited) {
log.info("waiting up to {}s for the herdr socket…", HERDR_WAIT_SECONDS);
waited = true;
}
try {
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
return false;
}
}
}
}
private Bridged() {
}
}
@@ -1,73 +0,0 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import java.util.Map;
import java.util.function.Supplier;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
* profile it points at, plus the live binding from a live architect's herdr terminal to its slot.
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the <em>future</em> spawn lifecycle will read when it stands the
* slot up. A read-only snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — a {@link Supplier} consulted on every read, so a binding
* injected <em>after</em> startup (an operator pin, or the later lifecycle once it spawns a
* session) takes effect without a restart. {@link CallerResolver} reads this to turn a pane
* into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only exposes the map the
* resolver resolves against and the profile lookup that lifecycle will call. Nothing here
* creates or manages an architect session.
*/
public final class ArchitectRegistry {
private final Map<String, BridgedConfig.Architect> slots;
private final Supplier<Map<String, String>> terminalBindings;
public ArchitectRegistry(Map<String, BridgedConfig.Architect> slots,
Supplier<Map<String, String>> terminalBindings) {
this.slots = slots == null ? Map.of() : Map.copyOf(slots);
this.terminalBindings = terminalBindings == null ? Map::of : terminalBindings;
}
/** The configured slots, keyed by gateway-local unique name. Unmodifiable snapshot. */
public Map<String, BridgedConfig.Architect> slots() {
return slots;
}
/**
* The live {@code terminal_id → slot name} bindings, re-read on every call.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code bridge_whoami}/the roster will read to say which slot a pane hosts.
*/
public Map<String, String> terminalBindings() {
return terminalBindings.get();
}
/** The slot a live terminal is bound to, or {@code null} if it is no architect slot. */
public String slotForTerminal(String terminal) {
return terminal == null ? null : terminalBindings.get().get(terminal);
}
/**
* The strong-model profile a slot runs under — what the future spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
BridgedConfig.Architect a = slots.get(slotName);
return (a == null || a.profile() == null) ? null : a.profile();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
}
}
@@ -1,830 +0,0 @@
package dev.ltms.bridged.config;
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
/**
* {@code bridged} configuration, loaded from a YAML file (see
* {@code bridged.example.yaml}). Unknown keys are ignored so config can grow ahead
* of the code — but an unknown <em>top-level</em> key is logged as a WARN at load (CB-530), because
* silently dropping a whole block is indistinguishable from honouring it.
*
* @param bind REST/MCP listen host:port
* @param herdrSocket path to herdr's Unix socket ({@code null} → client default)
* @param worker single worker profile (legacy; superseded by {@code workers})
* @param workers named worker profiles, keyed by profile name (multi-backend fleet)
* @param defaultWorker which {@code workers} key a no-argument spawn uses ({@code null} → the
* single {@code worker}, or the sole/first profile)
* @param guard subscription-boundary allowlist
* @param worktreeRoot nullable root directory for provisioned worktrees; defaults to a sibling
* of the repo root
* @param lifecycle session lifecycle limits ({@code null} = all disabled)
* @param spawnReadyTimeoutMs max ms to wait for a spawned worker to reach an injectable state
* ({@code null} / 0 disables the poll gate — legacy non-blocking behaviour)
* @param spawnReadyPollMs poll interval while waiting for the worker to become injectable
* @param broker external AMQP broker for durable reply delivery ({@code null} → in-memory,
* soft-state {@code ReplyInbox}; present → the AMQP-backed adapter, CB-307 Stage 2)
* @param primary optional pinned primary terminal config ({@code null} → derived from connection);
* a non-blank {@code terminal} seeds {@code PrimaryRegistry} and prevents
* connection-derived overrides, CB-307
* @param leaders named panes that orchestrate rather than are orchestrated (CB-530), keyed by
* lead name; supersedes the singular {@code primary} pin, which stays honoured.
* See {@link #leaderTerminals()} for how the two merge
* @param architects CB-548 architect slots, keyed by gateway-local unique slot name; each points
* at a strong-model profile, and the identity a live session is matched by is
* its {@code terminal} binding (see {@link #architectTerminals()}). A slot is
* the hook the future spawn lifecycle reads a profile back from — nothing here
* spawns it.
* @param leadScan opt-in discovery of leads by tab label (CB-531); {@code null} ⇒ no scanning,
* and only {@code leaders:}/{@code primary:} name a lead
* @param placement how to choose a worker profile for an unqualified spawn:
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
Bind bind,
String herdrSocket,
Worker worker,
Map<String, Worker> workers,
String defaultWorker,
Guard guard,
String worktreeRoot,
Lifecycle lifecycle,
Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs,
Broker broker,
Primary primary,
Map<String, Leader> leaders,
Map<String, Architect> architects,
LeadScan leadScan,
String placement,
Auth auth) {
@JsonIgnoreProperties(ignoreUnknown = true)
public record Bind(String host, int port) {
public Bind {
if (host == null || host.isBlank()) host = "127.0.0.1";
if (port <= 0) port = 8765;
}
}
/**
* @param profile ccs profile a worker is spawned under (Stage-1: {@code ltms-local})
* @param baseUrl the off-subscription endpoint injected as {@code ANTHROPIC_BASE_URL}
* @param model model alias, injected as {@code ANTHROPIC_MODEL} (may be {@code null})
* @param configDir {@code CLAUDE_CONFIG_DIR} so the worker inherits the profile's
* skills/MCP/hooks (may be {@code null})
* @param tokenEnv name of the host env var holding the worker's auth token; its value
* is injected as {@code ANTHROPIC_AUTH_TOKEN} (never stored in config)
* @param argv launch command; defaults to {@code ["claude"]}
* @param placement where a worker lands: {@code "tab"} (default — its own tab in the
* worker space) or {@code "pane"} (legacy — split the focused tab)
* @param workspace label of the dedicated worker space; found-or-created on first
* spawn (default {@code "bridged-workers"}). A future per-session
* layout is just a distinct label here — the shared space is default.
* @param tabLabel template for a worker tab's label; {@code {profile}}/{@code {model}}
* and {@code {n}} (per-worker number, to keep sibling tabs distinct)
* are substituted (default {@code "worker: {profile} #{n}"})
* @param mcpUrl bridge MCP URL to provision into the worker's {@code configDir} so it
* can call {@code bridge_reply} ({@code null}/blank → no provisioning; the
* worker won't reply, only the fallback/timeout resolves the send)
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param weight relative selection weight for {@code placement: weighted}. Absent or
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
* not sum to 1.0.
* @param maxLoad max live workers allowed on this profile at one time; absent or
* non-positive ⇒ unlimited. Live means any session the registry still owns
* (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.worker.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
* each adapter drives only its own kind. Normalised to lower-case; blank ⇒ the
* default. It selects the adapter, not the transport — placement, tabs, cwd, and
* the readiness gate are kind-independent and stay in the shared base.
* @param subscription {@code true} to run this profile's workers on the operator's Claude
* subscription, on purpose (CB-539). When set, the launcher neither requires
* nor injects {@code ANTHROPIC_BASE_URL} / {@code ANTHROPIC_AUTH_TOKEN}, and the
* {@code SubscriptionGuard} base_url requirement is skipped <em>for this
* profile only</em>. Absent/{@code false} (the default) keeps today's hard
* refusal: a claude-code profile with no base_url may not spawn, because
* spawning one would bill the subscription. Mutually exclusive with
* {@code baseUrl} — setting both is a configuration error (the two state
* opposite intents). For the same reason, an {@code env:} entry naming
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN} is refused at
* config load (CB-542): on the subscription path no guard would vet it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd,
List<String> parityOverlay,
String gitTokenEnv, String gitHostEnv,
String kind,
Map<String, String> env,
Float weight,
Integer maxLoad,
Boolean subscription) {
/** Peer kind spawned by {@link dev.ltms.bridged.worker.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
public Worker {
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
// argv (e.g. `opencode`) and must not inherit the Claude binary — so only default when unset
// AND this is the claude-code kind.
String k = (kind == null || kind.isBlank()) ? KIND_CLAUDE_CODE : kind.toLowerCase();
argv = (argv == null || argv.isEmpty())
? (KIND_CLAUDE_CODE.equals(k) ? List.of("claude") : List.of(k))
: List.copyOf(argv);
kind = k;
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_WORKER_TOKEN" : tokenEnv;
placement = (placement == null || placement.isBlank()) ? "tab" : placement.toLowerCase();
workspace = (workspace == null || workspace.isBlank()) ? "bridged-workers" : workspace;
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? "worker: {profile} #{n}" : tabLabel;
// CB-525: .mcp.json is deliberately NOT here. Replicating the primary's MCP config gave a
// worker the primary's IDE servers, which are bound to the primary's checkout — so its
// navigation returned paths outside its own worktree. GitWorktrees now neutralizes that
// file instead; a worker's tools are whatever its launcher mounts.
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".claude/settings.local.json", ".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
subscription = (subscription != null && subscription) ? Boolean.TRUE : Boolean.FALSE;
}
/**
* Backward-compatible constructor without the CB-302 git-forge fields — the worker is
* granted no PR-create token (push over SSH is unaffected). Keeps pre-CB-302 call sites
* (and any {@code workers:} YAML that omits the git keys) working unchanged.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null, null);
}
/**
* Backward-compatible constructor with the CB-302 git-forge fields but no explicit peer
* {@code kind} — defaults to {@link #KIND_CLAUDE_CODE}. Keeps pre-CB-402 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null, null);
}
/**
* Backward-compatible constructor without the CB-511 {@code env:} passthrough — the worker
* gets the daemon's PATH and nothing else. Keeps pre-CB-511 call sites working.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Worker withProfile(String p) {
return new Worker(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
public boolean isClaudeCode() {
return KIND_CLAUDE_CODE.equals(kind);
}
/** True when this profile is served by the opencode adapter (CB-402). */
public boolean isOpenCode() {
return KIND_OPENCODE.equals(kind);
}
/**
* True when this profile runs on the operator's Claude subscription, on purpose (CB-539).
* Absent/{@code false} (the default) keeps the hard refusal: a claude-code profile with no
* base_url may not spawn, because doing so would bill the subscription.
*/
public boolean isSubscription() {
return Boolean.TRUE.equals(subscription);
}
/**
* Backward-compatible constructor without the CB-539 subscription flag — the worker stays on
* the off-subscription boundary (the default). Keeps pre-CB-539 call sites (and any YAML
* that omits the flag) compiling and behaving identically.
*/
public Worker(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null);
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
public boolean hasGitToken() {
return gitTokenEnv != null && !gitTokenEnv.isBlank();
}
/**
* True when this profile's {@code env:} block names a worker-side Anthropic binding variable.
* Those two keys are the adapter's, never the operator's: on a claude-code worker
* {@code ANTHROPIC_BASE_URL} is the endpoint the {@code SubscriptionGuard} vetted, and
* {@code ANTHROPIC_AUTH_TOKEN} is injected from {@code tokenEnv}. An {@code env:} entry for
* either is a bypass vector — it is what the subscription path would otherwise let survive
* unguarded — so it is rejected at config load (see
* {@link BridgedConfig#validateSubscriptionProfiles()}).
*/
public boolean envCarriesAnthropicBinding() {
return env != null
&& (env.containsKey("ANTHROPIC_BASE_URL") || env.containsKey("ANTHROPIC_AUTH_TOKEN"));
}
/** True when workers should land in their own tab in the worker space. */
public boolean tabPlacement() {
return "tab".equals(placement);
}
/** True when the bridge MCP should be mounted into a spawned worker (via launch flags). */
public boolean hasMcp() {
return mcpUrl != null && !mcpUrl.isBlank();
}
/**
* Render {@link #tabLabel} for the {@code n}-th worker (substitutes
* {@code {profile}}/{@code {model}}/{@code {n}}), so sibling worker tabs are distinct.
*/
public String renderTabLabel(long n) {
return tabLabel
.replace("{profile}", profile == null ? "" : profile)
.replace("{model}", model == null ? "" : model)
.replace("{n}", Long.toString(n));
}
}
/**
* Session lifecycle limits. All knobs are opt-in: {@code null} or {@code 0} disables the
* feature so existing configs keep the previous behaviour.
*
* @param idleTtlSeconds max seconds a {@code READY}/{@code DONE} session may sit idle
* before it is reaped ({@code null} → disabled)
* @param contextCap max delegated turns a session serves before force-release
* ({@code null} → disabled)
* @param drainTimeoutSeconds seconds to wait for {@code BUSY} sessions to finish before
* forced teardown on shutdown (default 5 when unset)
* @param clearAfterTurn whether a reusable worker discards its conversation context after
* every completed delegated turn (default false)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Lifecycle(Integer idleTtlSeconds, Integer contextCap, Integer drainTimeoutSeconds,
boolean clearAfterTurn) {
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
* ⇒ the broker block is treated as absent (in-memory adapter).
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri) {
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return uri != null && !uri.isBlank();
}
}
/**
* Optional pinned primary terminal config (CB-307). When present with a non-blank
* {@code terminal}, the bridge uses this as the primary's herdr identity instead of
* deriving it from the MCP connection. It feeds two consumers: the push loop (where to nudge
* when replies land), and caller resolution — a caller whose connection maps to this pane is
* the primary, where the pane match would otherwise classify it as a worker. Pin it when the
* primary runs <em>inside</em> a herdr pane; it also helps off-host or non-herdr primaries,
* where connection-derived identity is unavailable and only the nudge target matters.
*
* @param terminal the primary's herdr {@code terminal_id} ({@code null}/blank → derive)
* @param pushReminders max reminder nudges before giving up (default 5)
* @param pushBackoffMs delay between reminders in ms (default 15000)
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Primary(String terminal, Integer pushReminders, Integer pushBackoffMs) {
/** @return configured reminder cap, or 5 */
public int remindersOrDefault() {
return pushReminders != null ? pushReminders : 5;
}
/** @return configured backoff in ms, or 15000 */
public long backoffMsOrDefault() {
return pushBackoffMs != null ? pushBackoffMs.longValue() : 15_000L;
}
}
/**
* One entry of the CB-530 {@code leaders:} registry — a pane that orchestrates rather than one
* that is orchestrated.
*
* <p>Why a registry and not a second {@code primary:}: {@code primary.terminal} is singular by
* construction, so a session in any other pane resolves as a worker. That is correct while one
* lead drives a fleet, and wrong the moment two leads (say an Opus lead and an opencode lead)
* work as peers — the second is silently demoted and refused every orchestration call.
*
* <p>{@code kind} and {@code model} are descriptive only at this stage: they document what runs
* in the pane and are reported back by {@code bridge_whoami}. Nothing spawns a lead — a lead
* pre-exists, which is precisely why it must be recognised by configuration rather than created.
*
* @param terminal the lead's herdr {@code terminal_id}; the only field identity depends on
* @param kind which agent runs there ({@code claude}, {@code opencode}, …); descriptive
* @param model the model or selector it runs, for operators reading the roster; descriptive
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Leader(String terminal, String kind, String model) {
}
/**
* One entry of the CB-548 {@code architects:} registry — a gateway-local named slot that points
* at a strong-model profile.
*
* <p>A lead and an architect differ in <em>authority</em>, not in how identity is established:
* both are recognised by configuration rather than spawned. A lead resolves to
* {@link dev.ltms.bridged.auth.Role#PRIMARY} and owns the whole lifecycle (spawn/stop/drain);
* an architect resolves to {@link dev.ltms.bridged.auth.Role#ARCHITECT}, which delegates turns
* ({@code SEND}) and replies/asks as its own pane but cannot stand up or tear down workers —
* lifecycle stays in one pair of hands.
*
* <p>Why a {@code profile} reference: an architect is meant to run a strong model, and the slot
* records which {@code workers:} profile that is — the value the future spawn lifecycle reads.
* It must name a configured profile, enforced by {@link #validateArchitects()} (a stale or
* typo'd reference fails at startup rather than silently spawning the wrong backend later).
*
* @param terminal the architect's herdr {@code terminal_id}; the field identity is matched by,
* via the live terminal→slot binding. Optional at config time — binding may be
* injected live — but a slot with no binding matches nothing yet.
* @param profile the name of the strong-model {@code workers:} profile this slot runs;
* required and validated against {@link #workerProfiles()}
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Architect(String terminal, String profile) {
}
/**
* Discover leads by tab label instead of by pasted {@code terminal_id} (CB-531).
*
* <p>Why: a lead is not spawned, so its {@code terminal_id} exists only once a human has opened
* the tab and started the agent — which makes {@code leaders:} a three-step ritual (start it,
* ask it its id, edit config, restart) repeated per lead. Naming the tab is one step, done at
* the moment the operator is already there. The convention also survives what an id does not:
* close the tab and reopen it and the id changes, while the label is retyped as-is.
*
* <p>Deliberately opt-in ({@code null} ⇒ off). Turning it on widens who resolves as
* {@link dev.ltms.bridged.auth.Role#PRIMARY}, and a config that never asked for it must not
* acquire that by upgrading the daemon.
*
* <p>bridged never writes these labels — see {@link dev.ltms.bridged.herdr.LeadTabScanner} for
* why that one-way direction is what keeps the convention trustworthy.
*
* @param tabPrefix label prefix marking a lead's tab, matched case-insensitively; the
* remainder is the lead's name ({@code "lead: opus-5.0"} → {@code
* opus-5.0}). Default {@code "lead:"}
* @param intervalSeconds how long a scan is cached before herdr is asked again; also the worst
* case before a newly-labelled tab is recognised. Default 10
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadScan(String tabPrefix, Integer intervalSeconds) {
public LeadScan {
tabPrefix = (tabPrefix == null || tabPrefix.isBlank()) ? "lead:" : tabPrefix.strip();
intervalSeconds = (intervalSeconds == null || intervalSeconds <= 0) ? 10 : intervalSeconds;
}
}
/**
* The terminal → lead-name map that {@link dev.ltms.bridged.auth.CallerResolver} resolves
* against, merging the {@code leaders:} registry with the legacy singular {@code primary:} pin.
*
* <p>Precedence: an explicit {@code leaders:} entry wins over the {@code primary:} pin for the
* same terminal. The pin is the older, less expressive spelling of the same fact, so when both
* name a pane the named entry is the one an operator meant. The pin is still honoured on its
* own — a config carrying only {@code primary:} behaves exactly as it did before CB-530.
*
* @return an unmodifiable map, empty when neither block is configured (nothing is pinned, and
* every pane therefore resolves as a worker — the pre-CB-307 behaviour)
*/
public Map<String, String> leaderTerminals() {
Map<String, String> byTerminal = new LinkedHashMap<>();
if (leaders != null) {
leaders.forEach((name, leader) -> {
if (leader != null && leader.terminal() != null && !leader.terminal().isBlank()) {
byTerminal.put(leader.terminal(), name);
}
});
}
if (primary != null && primary.terminal() != null && !primary.terminal().isBlank()) {
byTerminal.putIfAbsent(primary.terminal(), "primary");
}
return Collections.unmodifiableMap(byTerminal);
}
/**
* The terminal → architect-slot-name map that {@link dev.ltms.bridged.auth.CallerResolver}
* resolves against (CB-548), derived from the {@code architects:} registry.
*
* <p>Keyed by terminal because a live session is matched by its pane; the value is the
* gateway-local slot name. Slot names are inherently unique (a map key); a duplicate terminal
* across two slots is last-wins here (the later entry overrides), which {@code leadership} has
* always tolerated rather than refused. This is consumed as the <em>initial</em> live binding —
* the supplier that feeds the resolver may be swapped for a live one by the future lifecycle.
*
* @return an unmodifiable map, empty when no architect slot is configured
*/
public Map<String, String> architectTerminals() {
Map<String, String> byTerminal = new LinkedHashMap<>();
if (architects != null) {
architects.forEach((name, arch) -> {
if (arch != null && arch.terminal() != null && !arch.terminal().isBlank()) {
byTerminal.put(arch.terminal(), name);
}
});
}
return Collections.unmodifiableMap(byTerminal);
}
/**
* API authentication (CB-501). Governs how a caller that is <em>not</em> an on-host worker
* pane proves it is the primary.
*
* <p>Worker identity never depends on this block: a loopback peer PID that maps to a herdr
* pane is unforgeable and is always honoured (see
* {@link dev.ltms.bridged.mcp.ConnectionIdentity}). This only decides what happens for
* <em>everyone else</em>.
*
* @param mode {@code "loopback-trust"} (default) — any loopback caller that is not a known
* worker is the primary, no credential needed; this is the historical
* behaviour, now chosen explicitly rather than implied. {@code "token"} — such
* a caller must present {@code Authorization: Bearer <token>} or it is
* {@code ANONYMOUS} and authorized for nothing.
* @param tokenEnv name of the host env var holding the bearer token; the literal value is
* never stored in config. Defaults to {@code BRIDGED_API_TOKEN}. Only read
* when {@code mode} is {@code token}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Auth(String mode, String tokenEnv) {
/** Historical behaviour: loopback non-worker ⇒ primary, no credential. */
public static final String MODE_LOOPBACK_TRUST = "loopback-trust";
/** A non-worker caller must present a valid bearer token to be the primary. */
public static final String MODE_TOKEN = "token";
public Auth {
mode = (mode == null || mode.isBlank()) ? MODE_LOOPBACK_TRUST : mode.toLowerCase();
tokenEnv = (tokenEnv == null || tokenEnv.isBlank()) ? "BRIDGED_API_TOKEN" : tokenEnv;
}
/** True when a bearer token is required of every non-worker caller. */
public boolean tokenMode() {
return MODE_TOKEN.equals(mode);
}
}
/**
* Subscription boundary. Only these hosts may back a worker's
* {@code ANTHROPIC_BASE_URL}; the primary must carry none.
*
* @param offSubscriptionHosts hostnames allowed for worker base_urls
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Guard(List<String> offSubscriptionHosts) {
public Guard {
offSubscriptionHosts = offSubscriptionHosts == null ? List.of() : List.copyOf(offSubscriptionHosts);
}
public Set<String> hostSet() {
return Set.copyOf(offSubscriptionHosts);
}
}
/**
* The effective worker profiles, keyed by profile name. Prefers the {@code workers} map (each
* value's {@code profile} defaulted to its key); falls back to the legacy singular {@code worker}
* (keyed by its own profile). Empty if neither is configured.
*/
public Map<String, Worker> workerProfiles() {
if (workers != null && !workers.isEmpty()) {
Map<String, Worker> out = new LinkedHashMap<>();
workers.forEach((name, w) -> out.put(name,
(w.profile() == null || w.profile().isBlank()) ? w.withProfile(name) : w));
// Deliberately NOT Map.copyOf: its iteration order is salted per JVM run, which would
// discard the YAML definition order built above. Placement tie-breaks on candidate
// order (see WeightedRoundRobinPolicy), so losing it makes equal-weight placement
// non-reproducible across restarts. Unmodifiable-wrap instead of copy-and-scramble.
return Collections.unmodifiableMap(out);
}
if (worker != null) {
String name = (worker.profile() == null || worker.profile().isBlank()) ? "default" : worker.profile();
return Map.of(name, worker);
}
return Map.of();
}
/**
* The profile a no-argument spawn uses: {@code defaultWorker} if set, else the legacy single
* {@code worker}'s profile, else the sole/first configured profile, else {@code null}.
*/
public String defaultProfile() {
if (defaultWorker != null && !defaultWorker.isBlank()) {
return defaultWorker;
}
if (worker != null && worker.profile() != null && !worker.profile().isBlank()) {
return worker.profile();
}
Map<String, Worker> p = workerProfiles();
return p.isEmpty() ? null : p.keySet().iterator().next();
}
private static final Logger log = LoggerFactory.getLogger(BridgedConfig.class);
private static final ObjectMapper YAML = new ObjectMapper(new YAMLFactory());
/**
* Top-level keys this version understands. Used only to warn about the rest — see
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
*/
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "worker", "workers", "defaultWorker", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "leaders",
"architects", "leadScan", "placement", "auth");
/** Load and validate config from {@code path}. */
public static BridgedConfig load(Path path) {
try {
String yaml = Files.readString(path);
warnUnknownTopLevelKeys(yaml, path);
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
return cfg.withDefaults();
} catch (IOException e) {
throw new UncheckedIOException("cannot read bridged config at " + path, e);
}
}
/**
* Log a WARN naming any top-level key this version does not understand (CB-530).
*
* <p>Why this exists: every record here is {@code @JsonIgnoreProperties(ignoreUnknown = true)},
* which is deliberate — config must be allowed to grow ahead of the code, and a rolled-back
* daemon must still start. The cost is that a whole block can be written, parsed, dropped, and
* never mentioned again. That is exactly how a hand-written {@code leaders:} registry came to
* look configured while being inert: the daemon started, nothing complained, and the only way
* to discover it was reading the config class.
*
* <p>A warning rather than a failure, on purpose. Failing closed would turn "the config names
* something this build has not learned yet" into a daemon that will not boot — which is the
* forward-compatibility this annotation was chosen to preserve. Loud, not fatal.
*/
private static void warnUnknownTopLevelKeys(String yaml, Path path) {
List<String> unknown = unknownTopLevelKeys(yaml);
if (!unknown.isEmpty()) {
log.warn("{}: ignoring unknown top-level config key(s) {} — this build does not "
+ "understand them, so they have NO effect. Check for a typo, or a "
+ "feature not in this version.",
path, unknown);
}
}
/**
* The top-level keys in {@code yaml} that this build does not understand, sorted. Package-private
* so the guardrail is asserted directly rather than through a log appender.
*
* @return empty when everything is known, or when {@code yaml} is not a mapping at all (a
* malformed file is {@code readValue}'s error to report, not this method's)
*/
static List<String> unknownTopLevelKeys(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return List.of();
}
if (raw == null) {
return List.of();
}
return raw.keySet().stream()
.map(String::valueOf)
.filter(k -> !KNOWN_TOP_LEVEL_KEYS.contains(k))
.sorted()
.toList();
}
/** Fill in nested defaults so callers never see nulls for structural fields. */
public BridgedConfig withDefaults() {
Bind b = bind != null ? bind : new Bind(null, 0);
Guard g = guard != null ? guard : new Guard(List.of());
Lifecycle l = lifecycle != null ? lifecycle : new Lifecycle(null, null, null, false);
Integer timeout = (spawnReadyTimeoutMs != null) ? spawnReadyTimeoutMs : 20000;
Integer pollMs = (spawnReadyPollMs != null) ? spawnReadyPollMs : 300;
Auth a = auth != null ? auth : new Auth(null, null);
String placementOrDefault = (placement != null && !placement.isBlank()) ? placement : "fixed";
// broker is left as-is: null (or an empty/blank uri) keeps the in-memory soft-state inbox.
// primary is left as-is: null defaults to connection-derived identity.
// leadScan is left as-is: null is "off", and LeadScan's own compact constructor defaults the
// fields of a block that IS present. Defaulting it here would switch the feature on for
// every config that never mentioned it.
// architects is left as-is: null is "none configured", and Architect's fields have no
// defaults to fill. Defaulting it here would change nothing, so leave the call natural.
return new BridgedConfig(b, herdrSocket, worker, workers, defaultWorker, g, worktreeRoot, l, timeout, pollMs, broker, primary, leaders, architects, leadScan, placementOrDefault, a);
}
/**
* Reject a configuration whose network exposure outruns its authentication (CB-501).
*
* <p>{@code loopback-trust} means "any caller that is not a known worker pane is the primary" —
* safe only because the OS refuses non-local connections to a loopback bind. Widen
* {@code bind.host} without switching to {@code token} mode and that sentence becomes "any
* client that can reach this port is the primary", which is the most privileged role on the
* bus. Rather than document the hazard, make it unrepresentable: fail fast at startup.
*
* @throws IllegalStateException when a non-loopback bind is paired with {@code loopback-trust}
*/
public void validateAuthExposure() {
String host = bind().host();
if (isLoopbackBind(host) || auth().tokenMode()) {
return;
}
throw new IllegalStateException(
"refusing to start: bind.host=" + host + " is not loopback, but auth.mode="
+ auth().mode() + ". A non-loopback bind treats every unauthenticated "
+ "caller as the primary (spawn/stop/send/drain on any session). Set "
+ "auth.mode: token (with auth.tokenEnv) before exposing this port, or "
+ "bind to 127.0.0.1 and put a reverse proxy in front.");
}
/**
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
*
* <p>The scan reads a tab label and concludes "a lead lives here". bridged also <em>writes</em>
* tab labels — every worker gets {@code tabLabel} rendered into its tab. Choose a
* {@code leadScan.tabPrefix} that a worker template matches and the daemon starts labelling its
* own workers as leads, promoting the entire fleet to {@link dev.ltms.bridged.auth.Role#PRIMARY}
* with no message and no diff. The worker-space exclusion in
* {@link dev.ltms.bridged.herdr.LeadTabScanner} already blocks the realistic path, but defence
* that depends on one workspace label holding is not defence enough for a privilege boundary.
*
* <p>Fatal rather than a warning, unlike {@link #warnUnknownTopLevelKeys}: an unknown key means
* a feature does nothing, while this means a feature does the opposite of what it says.
*
* @throws IllegalStateException when any worker profile's {@code tabLabel} starts with the
* configured lead prefix
*/
public void validateLeadScan() {
if (leadScan == null) {
return;
}
String prefix = leadScan.tabPrefix();
List<String> clashing = workerProfiles().entrySet().stream()
.filter(e -> e.getValue().tabLabel() != null
&& e.getValue().tabLabel().strip()
.regionMatches(true, 0, prefix, 0, prefix.length()))
.map(Map.Entry::getKey)
.sorted()
.toList();
if (clashing.isEmpty()) {
return;
}
throw new IllegalStateException(
"refusing to start: leadScan.tabPrefix=\"" + prefix + "\" also matches the tabLabel "
+ "of worker profile(s) " + clashing + ". Every worker spawned under them "
+ "would be read back as a lead and granted spawn/stop/send on the whole "
+ "fleet. Change one of the two so worker tabs and lead tabs cannot be "
+ "confused.");
}
/**
* Reject a subscription profile whose {@code env:} block tries to reseat the Anthropic binding
* (CB-542).
*
* <p>Why this must be fatal rather than sanitised: {@code subscription: true} deliberately
* stops the launcher from writing {@code ANTHROPIC_BASE_URL}/{@code ANTHROPIC_AUTH_TOKEN} and
* skips the {@code SubscriptionGuard} for that profile. But the profile's {@code env:} map is
* layered into the worker environment separately, so an {@code ANTHROPIC_BASE_URL} sitting
* there would survive into the worker having passed no guard at all — {@code subscription: true}
* plus an {@code env:} repoint is a contradiction just like {@code subscription: true} plus a
* {@code baseUrl}. The launcher also hard-strips these two keys from the worker env as a
* belt-and-braces measure; this method is the loud, load-time refusal so the operator is told
* about the mistake instead of having it silently cleaned up.
*
* @throws IllegalStateException when any subscription profile's {@code env:} names
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN},
* naming the profile and the offending key(s)
*/
public void validateSubscriptionProfiles() {
List<String> bad = new java.util.ArrayList<>();
workerProfiles().forEach((name, w) -> {
if (w.isSubscription() && w.envCarriesAnthropicBinding()) {
List<String> keys = w.env().keySet().stream()
.filter(k -> k.equals("ANTHROPIC_BASE_URL") || k.equals("ANTHROPIC_AUTH_TOKEN"))
.sorted()
.toList();
bad.add("worker profile '" + name + "' carries " + keys + " in env: — "
+ "subscription: true forbids ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN there, "
+ "because on the subscription no guard vets them (they would repoint the "
+ "worker past the SubscriptionGuard). Remove them from env: (or drop "
+ "subscription: true).");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/**
* Reject an architect slot whose profile reference does not resolve (CB-548).
*
* <p>An architect's {@code profile} is the strong-model {@code workers:} profile the future
* spawn lifecycle will read to stand the slot up. A reference that names no configured profile
* is a typo or a stale config — and unlike a worker spawn (which fails loudly at its call site
* when it cannot resolve), an architect slot fails only when something later tries to use it.
* This config class <em>does</em> have access to {@link #workerProfiles()}, so the reference
* is validated at startup and the mistake is named then, not discovered months later by a
* spawn that quietly has no backend to use.
*
* <p>Slot-name uniqueness needs no check here: the registry is a {@code Map} keyed by name, so
* duplicates are unrepresentable by construction.
*
* @throws IllegalStateException when any architect slot is missing or names an unknown profile,
* naming the slot and the offending reference
*/
public void validateArchitects() {
if (architects == null) {
return;
}
Map<String, Worker> profiles = workerProfiles();
List<String> bad = new java.util.ArrayList<>();
architects.forEach((name, arch) -> {
if (arch == null || arch.profile() == null || arch.profile().isBlank()) {
bad.add("architect slot '" + name + "' has no profile: — give it the name of a "
+ "workers: profile (the strong-model backend it runs).");
return;
}
if (!profiles.containsKey(arch.profile())) {
bad.add("architect slot '" + name + "' references profile '" + arch.profile()
+ "', which is not a configured workers: profile (have: " + profiles.keySet()
+ ").");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
return true; // Bind's own default is 127.0.0.1
}
String h = host.trim().toLowerCase();
return h.equals("127.0.0.1") || h.equals("::1") || h.equals("localhost")
|| h.startsWith("127.");
}
}
@@ -1,164 +0,0 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Set;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* Discovers which panes host a lead by scanning herdr for tabs the operator labelled by convention
* (CB-531), and hands {@link dev.ltms.bridged.auth.CallerResolver} the resulting
* {@code terminal_id → lead name} map.
*
* <p><strong>Why scan at all.</strong> A lead is never spawned — a human opens a tab and starts an
* agent in it — so the daemon cannot learn a lead's {@code terminal_id} at creation time the way it
* does a worker's. CB-530 solved that by having the operator paste each id into {@code leaders:},
* which works but costs a config edit and a daemon restart per lead, and the id is only obtainable
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>bridged never renames a lead tab. The operator's label is read-only input, so what is in
* the tab bar is always what the human wrote — no round-trip where the daemon's own rename
* becomes the evidence for its next decision.</li>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
* </ol>
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code BridgedConfig.validateLeadScan} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
private final HerdrClient herdr;
private final String tabPrefix;
private final Set<String> excludedWorkspaceLabels;
private final Map<String, String> configuredLeads;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached;
private long scannedAtNanos;
private boolean everScanned;
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabPrefix a tab whose label starts with this (case-insensitively) hosts a
* lead; the rest of the label, trimmed, is the lead's name
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param configuredLeads the static {@code leaders:}/{@code primary:} registry, merged
* over every scan result. Explicit config outranks the
* convention, and survives a scan that cannot run at all
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, String tabPrefix, Set<String> excludedWorkspaceLabels,
Map<String, String> configuredLeads, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabPrefix = tabPrefix == null || tabPrefix.isBlank() ? "lead:" : tabPrefix.strip();
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.configuredLeads = configuredLeads == null ? Map.of() : Map.copyOf(configuredLeads);
this.ttlNanos = ttlNanos;
this.clock = clock;
this.cached = this.configuredLeads;
}
/**
* The current {@code terminal_id → lead name} map, rescanning when the cache has expired.
*
* <p>Synchronized so a burst of concurrent calls produces one scan rather than one each; a scan
* is a handful of RPCs over a Unix socket and is rate-limited to one per TTL.
*/
@Override
public synchronized Map<String, String> get() {
long now = clock.getAsLong();
if (everScanned && now - scannedAtNanos < ttlNanos) {
return cached;
}
// Stamp before scanning, not after: a herdr that is down must cost one attempt per TTL, not
// one per request.
scannedAtNanos = now;
everScanned = true;
try {
Map<String, String> fresh = scan();
if (!fresh.equals(cached)) {
log.info("lead panes: {}", fresh);
}
cached = fresh;
} catch (HerdrException e) {
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
cached.size(), e.getMessage());
}
return cached;
}
/** One full pass: labelled tabs → their panes → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace ws = Workspace.from(w);
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
continue;
}
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
String name = leadNameOf(tab.label());
if (name != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), name);
}
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
}
}
byTerminal.putAll(configuredLeads); // an explicit pin outranks a label
return Collections.unmodifiableMap(byTerminal);
}
/**
* The lead name a tab label declares, or {@code null} if it declares none.
*
* <p>{@code "lead: opus-5.0"} → {@code "opus-5.0"}. A bare {@code "lead:"} names nobody and is
* rejected: an unnamed lead would resolve as {@code PRIMARY} with nothing to attribute it to.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
String l = label.strip();
if (!l.regionMatches(true, 0, tabPrefix, 0, tabPrefix.length())) {
return null;
}
String name = l.substring(tabPrefix.length()).strip();
return name.isEmpty() ? null : name;
}
}
@@ -1,59 +0,0 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.Map;
/**
* Resolves which herdr pane a process belongs to — the herdr half of connection-based MCP
* identity (CB-105). Given the PID that opened an MCP connection, {@link #terminalForPid} finds
* the agent pane whose process tree contains it, so {@code bridged} can tell <em>which worker</em>
* is calling without the worker sending anything spoofable.
*
* <p>herdr owns the PID→pane truth: {@code pane.process_info} reports each pane's {@code shell_pid}
* and foreground process PIDs. This scans agent panes; a spawn-time {@code pid→terminal} cache is
* the obvious optimization once wired into {@code ClaudeCodeLauncher}.
*/
public final class PaneLocator {
private final HerdrClient herdr;
public PaneLocator(HerdrClient herdr) {
this.herdr = herdr;
}
/**
* The {@code terminal_id} of the agent pane whose process tree contains {@code pid}, or
* {@code null} if no agent pane owns it (e.g. the caller is the primary, or off-host).
*/
public String terminalForPid(long pid) {
if (pid <= 0) {
return null;
}
for (JsonNode pane : herdr.call("pane.list", Map.of()).path("panes")) {
String paneId = pane.path("pane_id").asText(null);
if (paneId != null && paneOwnsPid(paneId, pid)) {
return pane.path("terminal_id").asText(null);
}
}
return null;
}
private boolean paneOwnsPid(String paneId, long pid) {
JsonNode info;
try {
info = herdr.call("pane.process_info", Map.of("pane_id", paneId)).path("process_info");
} catch (HerdrException e) {
return false; // pane vanished mid-scan — just skip it
}
if (info.path("shell_pid").asLong(-1) == pid) {
return true;
}
for (JsonNode p : info.path("foreground_processes")) {
if (p.path("pid").asLong(-1) == pid) {
return true;
}
}
return false;
}
}
@@ -1,265 +0,0 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
* {@link Rendezvous} so a blocking {@code bridge_send} resolves even when the worker finishes its
* task without ever calling {@code bridge_reply} — the common case for a real delegated coding task.
*
* <p>On a confirmed {@code working → idle} boundary it scrapes the worker's recent transcript and
* resolves the awaiting send with that tail (a {@link Rendezvous.Kind#COMPLETION} resolution, so the
* caller can tell a scrape from a structured reply). It scrapes only when a send is actually waiting
* — a fleet worker's own turns, or a send that already timed out, cost no herdr traffic. An explicit
* {@code bridge_reply} that raced in first wins; {@link Rendezvous#resolveCompletion} is then a no-op.
*
* <p>It also handles the CB-109 stall signal ({@link #onTurnFailed}): a worker that ran a turn then
* wedged in an {@code unknown} state resolves the send as a failure (with the error screen as
* context) rather than leaving it to time out.
*
* <p>The scrape is cleaned to the last {@code ⏺} assistant block (stripping TUI chrome) and guarded
* against misattribution (CB-115): the pane content is baselined on delivery ({@link #onDelivered}),
* and a completion whose scrape is unchanged from that baseline — the previous turn's wind-down
* sampled as this turn's boundary on a rapid back-to-back send — is suppressed rather than resolving
* the send with a stale answer.
*
* <p><strong>Waiter-specific resolution (CB-116).</strong> On delivery we also capture the exact
* {@link Rendezvous} waiter this turn belongs to, and the completion/failure fallbacks resolve
* <em>that</em> waiter — never "whatever send is waiting now". A completion fallback runs on a virtual
* thread and can land after the worker's {@code bridge_reply} already resolved the turn and the
* <em>next</em> send opened its own waiter on the same session; resolving the current waiter would
* then deliver turn N's stale scrape as turn N+1's answer. Targeting the captured waiter makes a late
* completion a harmless no-op (its waiter is already done) instead of a cross-turn stale reply.
*
* <p>Wired as the {@link Injector}'s {@link TurnListener}; the handlers hand off to a virtual thread
* so the scrape's herdr round-trip never stalls the status poller. The captured waiter is read on the
* poller thread (before any next-turn delivery can overwrite it) and passed into the virtual thread.
*/
public final class CompletionResolver implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(CompletionResolver.class);
/**
* herdr {@code agent.read} source for the completion scrape. {@code recent} returns the tail of
* the transcript (the worker's last output), which is what a delegator wants when the worker
* didn't structure a reply.
*/
static final String SCRAPE_SOURCE = "recent";
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private final AgentControl agents;
private final Rendezvous rendezvous;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
* delivering send opened, plus the assistant block present when it was delivered.
*
* <p>The {@code waiter} is what makes a late fallback safe (CB-116): we resolve it, not "whoever
* is waiting now", so a completion that fires after the next send has opened its own waiter is a
* no-op rather than a cross-turn stale reply. The {@code baseline} is the CB-115 staleness
* reference: a completion scrape equal to it means the worker produced no new output (the previous
* turn's wind-down sampled as this boundary), so it is suppressed. Overwritten on each delivery;
* cleared when the turn resolves. Package-private so tests can capture and replay a specific turn.
*/
record InFlight(CompletableFuture<Rendezvous.Resolution> waiter, String baseline) {
}
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
this.agents = agents;
this.rendezvous = rendezvous;
}
@Override
public void onDelivered(String target) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
}
String baseline;
try {
// Clip to the same cap resolve() applies to the tail (line ~134): the CB-115 misattribution
// guard compares baseline.equals(tail), so both sides must be the same capped representation.
// An unclipped baseline vs a clipped tail would never match for a >MAX_SCRAPE_CHARS block,
// defeating the guard and letting a stale completion resolve the send.
baseline = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
baseline = null; // fail open: no baseline ⇒ no suppression
log.debug("delivery baseline for {} failed: {}", target, e.getMessage());
}
inFlight.put(target, new InFlight(waiter, baseline));
}
/** The turn currently baselined for {@code target}, or {@code null} — a test hook for the captureBaseline path. */
InFlight inFlight(String target) {
return inFlight.get(target);
}
@Override
public void onTurnComplete(String target) {
// Read the in-flight turn on the poller thread — before any next-turn delivery can overwrite
// it — then off-load the scrape (a herdr round-trip we must not block polling on) to a vthread.
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
/**
* Resolve the completed turn before adapter housekeeping can erase its rendered output. This is
* intentionally synchronous and used only when a post-turn context reset is enabled; the normal
* path remains off-loaded so polling is not blocked by a scrape.
*/
public void resolveBeforePostAction(String target) {
resolve(target, inFlight.get(target));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
void resolve(String target, InFlight turn) {
CompletableFuture<Rendezvous.Resolution> waiter = turn == null ? null : turn.waiter();
if (waiter == null || waiter.isDone()) {
// Nobody is blocked on THIS turn (it had no send, or its bridge_reply already won). Skip
// the scrape; resolving the current waiter here would be the CB-116 cross-turn stale reply.
inFlight.remove(target, turn);
return;
}
String tail;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
log.warn("completion scrape for {} failed; resolving with an empty tail: {}",
target, e.getMessage());
tail = "";
scrapeFailed = true;
}
// Misattribution guard (CB-115): if the scrape is byte-identical to the pane content at
// delivery, this turn produced no new output — the boundary belongs to the previous turn's
// wind-down (common on rapid back-to-back sends). Suppress rather than resolve the send with
// a stale answer; the real bridge_reply (or a later genuine completion) resolves it instead.
// A scrape that failed to read is exempt — an empty tail there is "couldn't see", not "no change".
String baseline = turn.baseline();
if (!scrapeFailed && baseline != null && baseline.equals(tail)) {
log.debug("suppressing misattributed completion for {} (no output change since delivery)",
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
inFlight.remove(target, turn);
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
CompletableFuture<Rendezvous.Resolution> waiter =
turn != null ? turn.waiter() : rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.debug("failed send to {} via turn-stall fallback", target);
}
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
return trimmed.length() <= MAX_SCRAPE_CHARS
? trimmed
: trimmed.substring(trimmed.length() - MAX_SCRAPE_CHARS);
}
/**
* Extract the last assistant message from a raw Claude Code pane scrape (CB-115). Claude Code
* prefixes each assistant turn with {@code ⏺}; the delegator wants that answer, not the TUI
* chrome around it. Take everything from the final {@code ⏺} onward and stop at the <em>first</em>
* hard interface boundary below it — the spinner/status line, input box, {@code ❯} prompt (which
* may echo the <em>next</em> turn's text), footer, or tips/warnings. Stopping at the first
* boundary (rather than trimming only trailing chrome) is what keeps a following turn's echoed
* prompt out of this reply. Blank lines are not boundaries, so a multi-paragraph answer survives;
* trailing blanks are trimmed at the end. With no {@code ⏺} marker (an unusual render) the whole
* text is scanned the same way, so we never lose the reply.
*
* <p>Package-private and pure so it is unit-testable without herdr.
*/
static String lastAssistantBlock(String raw) {
if (raw == null || raw.isBlank()) return "";
int marker = raw.lastIndexOf('⏺');
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
StringBuilder out = new StringBuilder();
int kept = 0;
for (String line : block.split("\n", -1)) {
if (isBoundary(line)) break; // first TUI boundary ends the assistant message
if (kept++ > 0) out.append('\n');
out.append(line);
}
return out.toString().strip();
}
/**
* A hard TUI boundary line that marks the end of an assistant message and the start of interface
* chrome (input box, prompt, spinner, footer, tips/warnings). Blank lines are <em>not</em>
* boundaries — an answer may contain them — so they are kept and trimmed only if trailing.
*/
private static boolean isBoundary(String line) {
String t = line.strip();
if (t.isEmpty()) return false;
// A horizontal rule / all box-drawing separators (e.g. "──────").
if (t.chars().allMatch(c -> c == '─' || c == '—' || c == '━' || c == '═' || c == '-')) {
return true;
}
String lower = t.toLowerCase();
return t.startsWith("╭") || t.startsWith("│") || t.startsWith("╰") || t.startsWith("┌")
|| t.startsWith("└") || t.startsWith("❯") || t.startsWith("⏵")
|| t.startsWith("⎿") || t.startsWith("⚠")
// Status/spinner lines Claude Code renders below a settled or in-flight turn,
// e.g. "✻ Baked for 21s", "✶ Forming…".
|| t.startsWith("✻") || t.startsWith("✳") || t.startsWith("✽") || t.startsWith("·")
|| t.startsWith("●") || t.startsWith("◐") || t.startsWith("✢") || t.startsWith("✶")
|| lower.contains("auto mode") || lower.contains("for shortcuts")
|| lower.contains("esc to interrupt") || lower.contains("bypass permissions");
}
}
@@ -1,399 +0,0 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Deque;
import java.util.List;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Consumer;
import java.util.function.Predicate;
import java.util.stream.Collectors;
/**
* The status-gated injector (CB-103): the single writer that delivers a message into a
* worker only when it is safe — {@code idle} or {@code blocked}, never mid-turn.
*
* <p>Per target it holds a FIFO queue and delivers <strong>at most one message per turn</strong>:
* after a send it waits for the worker to pick the message up (go {@code working}) before
* delivering the next, so two rapid deliveries never interleave into one turn. Each
* status-check-then-send for a target is serialized on the target's monitor, closing the
* TOCTOU window between "is it idle?" and "send" — one writer per worker.
*
* <p><strong>Polling caveat.</strong> This is driven by {@link #onStatus} sampling (a
* {@code StatusPoller}), not by reliable status <em>edges</em>. A turn can begin and end
* entirely between two polls, so the {@code working} pickup may never be sampled. To avoid
* wedging a queue forever, an awaited pickup is released after {@link #PICKUP_GRACE_POLLS}
* consecutive injectable samples (the worker has plainly moved on). Only {@code working} — not
* a transient {@code unknown} — counts as a real pickup, so a detection glitch can't prematurely
* release the latch. Perfectly reliable turn boundaries require a herdr {@code events.subscribe}
* stream; that is the intended upgrade and would replace only the sampling, not this queue.
*
* <p><strong>Turn completion (CB-106).</strong> Beyond delivery, the injector reports when a
* delegated turn <em>finishes</em>: after a delivery is picked up (a real {@code working} sample),
* the next injectable sample is a confirmed {@code working → idle} boundary and fires
* {@link TurnListener#onTurnComplete}. Completion is only ever synthesized from a <em>confirmed</em>
* turn — the pickup-grace path (a turn too fast to sample) unwedges the queue but does not fire
* completion, since without a sampled {@code working} there is no trustworthy "the worker just
* finished the task" signal to act on.
*/
public final class Injector {
private static final Logger log = LoggerFactory.getLogger(Injector.class);
/**
* How many consecutive injectable samples (with no {@code working} in between) after a send
* before we assume the turn completed unobserved and release the pickup latch. At the
* default 250ms poll interval this is a ~2s grace — far longer than a worker takes to start
* a turn, so it only fires on a genuinely missed pickup edge.
*/
private static final int PICKUP_GRACE_POLLS = 8;
/**
* How many consecutive {@code unknown} samples while a delegation is outstanding before we
* declare it stalled and fire {@link TurnListener#onTurnFailed} (CB-109). A worker wedged in a
* state herdr can't classify (e.g. an API-error screen) stays {@code unknown} indefinitely and
* would otherwise never resolve; any {@code working}/{@code idle} sample resets the streak, so a
* transient detection glitch cannot trip it. At the 250ms poll interval this is ~30s — far longer
* than any real detection blip, and still vastly better than the async send's timeout.
*/
private static final int TURN_STALL_GRACE_POLLS = 120;
/**
* How many consecutive injectable samples a queued-but-undelivered message may wait on the
* {@link #ready} gate before we give up and fail it (CB-114). The gate holds a message out of a
* worker's boot window (herdr reports {@code idle} while its Claude is still starting), but a
* worker whose Claude crashes during boot — or never connects the bridge MCP — stays "idle and
* not ready" forever: {@link #ready} never accepts it, the message is never delivered, and the
* target would be polled indefinitely with its caller's future never completing. After this
* grace the queued messages are failed and the target released. At the 250ms poll interval this
* is ~60s — deliberately longer than {@link #TURN_STALL_GRACE_POLLS}, since a first boot (spawn
* + model load + MCP connect) legitimately takes longer than an in-turn detection blip.
*/
private static final int READINESS_GRACE_POLLS = 240;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
private final Consumer<String> forget; // CB-114: clear a gone worker's readiness/presence
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
public Injector(AgentControl agents) {
this(agents, TurnListener.NOOP);
}
/** Delivery plus turn-completion signalling (CB-106); every target is treated as available. */
public Injector(AgentControl agents, TurnListener turnListener) {
this(agents, turnListener, _ -> true);
}
/**
* Delivery, completion signalling (CB-106), and a readiness gate (CB-113): a message is delivered
* only when {@code ready} accepts the target — i.e. the worker's Claude has connected the bridge
* MCP. This holds the first delivery out of the worker's boot window, where herdr already reports
* {@code idle} but the TUI would drop an injected paste.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready) {
this(agents, turnListener, ready, _ -> {
});
}
/**
* Delivery, completion signalling (CB-106), a readiness gate (CB-113), and readiness cleanup
* (CB-114): {@code forget} is invoked with a target when its worker is gone — dropped
* (pane crash) or timed out on the readiness gate — so its stale presence/readiness is cleared
* and does not linger past the worker's life.
*/
public Injector(AgentControl agents, TurnListener turnListener, Predicate<String> ready,
Consumer<String> forget) {
this.agents = agents;
this.turnListener = turnListener;
this.ready = ready;
this.forget = forget;
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, CompletableFuture<Void> delivered) {
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
private static final class Target {
final Deque<Pending> queue = new ArrayDeque<>();
boolean awaitingPickup; // sent a message, waiting for the worker to pick it up
int injectableSincePickup; // consecutive injectable samples while awaitingPickup
boolean awaitingCompletion; // a delivered message's turn is not yet known-complete
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
synchronized void add(Pending p) {
queue.add(p);
}
}
/**
* Queue {@code text} for delivery to {@code target} (a herdr {@code terminal_id}). Returns
* immediately with a future that completes when the message is actually sent — the worker
* may be mid-turn, in which case delivery waits for the next injectable window.
*
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
return t;
});
return delivered;
}
/**
* Feed a fresh status observation for {@code target}. Delivers the head of the queue iff the
* worker is injectable and no earlier message is still awaiting pickup. Serialized per target
* so the check and the send cannot race another delivery to the same worker; the delivered
* future is completed <em>after</em> the monitor is released so a caller's continuation never
* runs on the poller thread while it holds the lock.
*/
public void onStatus(String target, AgentStatus status) {
Target t = targets.get(target);
if (t == null) return;
Pending sent = null;
RuntimeException sendError = null;
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
t.postTurnObserved = true;
}
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.unknownSinceTurn = 0;
t.notReadySincePoll = 0;
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
if (t.awaitingPostTurnPickup) {
if (++t.injectableSincePostTurnPickup >= PICKUP_GRACE_POLLS) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
} else {
resubmit = true;
}
} else if (t.postTurnObserved) {
t.postTurnObserved = false;
}
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
// Release the latch rather than wedge — and give up on synthesizing a
// completion for this message, since without a confirmed `working` we cannot
// trust that a task-processing turn actually ran.
t.awaitingPickup = false;
t.injectableSincePickup = 0;
t.awaitingCompletion = false;
t.turnObserved = false;
} else {
// Delivered but still idle → the worker hasn't picked it up; the submit
// keystroke likely raced the paste (esp. right as the TUI became ready).
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
resubmit = true;
}
}
if (!t.awaitingPickup) {
// A confirmed turn (a `working` sample was seen) that has now returned to idle is
// a trustworthy `working → idle` completion boundary.
if (t.awaitingCompletion && t.turnObserved) {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
if (turnListener.hasPostTurnAction(target)) {
t.postTurnPending = true;
startPostTurn = true;
}
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion && !t.postTurnPending
&& !t.awaitingPostTurnPickup && !t.postTurnObserved) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
try {
agents.send(target, p.text());
t.queue.poll();
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (RuntimeException e) {
// Delivery failed at herdr; drop the poisoned message and surface it
// rather than blocking the queue behind it.
t.queue.poll();
sent = p;
sendError = e;
}
} else if (p != null && ++t.notReadySincePoll >= READINESS_GRACE_POLLS) {
// The worker has been idle-but-not-ready for the whole grace: its Claude
// never connected the bridge MCP (crashed during boot, or wedged on a
// startup prompt). The readiness gate would hold this message forever, so
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
t.queue.clear();
t.notReadySincePoll = 0;
}
}
}
} else {
// UNKNOWN (or any other non-injectable, non-working): not a safe window nor a
// reliable pickup signal, so we never deliver or release the pickup latch here. But
// an outstanding delegation whose worker has gone unresponsive — stuck in a state
// herdr can't classify (CB-109) — will never yield a working→idle boundary. After a
// sustained streak, declare it failed so the awaiting send resolves rather than
// riding out the async timeout. (This also frees a delivery that wedged before it
// was ever picked up, which the injectable-only pickup grace could never release.)
if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
t.awaitingPickup = false;
t.awaitingCompletion = false;
t.turnObserved = false;
t.unknownSinceTurn = 0;
turnFailed = true;
}
}
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
targets.remove(target, t);
}
}
// Fire listeners / herdr calls after releasing the monitor so nothing runs on the poller
// thread while it holds the target lock.
if (resubmit) {
try {
agents.submit(target); // nudge a raced Enter so the pending paste submits
} catch (RuntimeException e) {
log.debug("resubmit to {} failed (will retry next poll): {}", target, e.getMessage());
}
}
if (notReady != null) {
// Worker never became available: forget its (never-set) readiness, unblock every queued
// caller, and route the awaiting send through the same failure path as a stalled turn so
// a blocking or async waiter resolves WORKER_FAILED rather than riding out the timeout.
forget.accept(target);
RuntimeException cause = new IllegalStateException(
target + " never became available (no bridge MCP connection within the boot window)");
for (Pending p : notReady) {
p.delivered().completeExceptionally(cause);
}
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
}
if (sent != null) {
if (sendError != null) {
log.warn("inject to {} failed, dropped message: {}", target, sendError.getMessage());
sent.delivered().completeExceptionally(sendError);
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
sent.delivered().complete(null);
}
}
}
/**
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
* awaited turn completion (so the {@code working → idle} boundary is observed).
*/
public Set<String> activeTargets() {
return targets.entrySet().stream()
.filter(e -> {
synchronized (e.getValue()) {
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion
|| t.postTurnPending || t.awaitingPostTurnPickup || t.postTurnObserved;
}
})
.map(java.util.Map.Entry::getKey)
.collect(Collectors.toSet());
}
/**
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
* turn was not yet resolved (CB-110 — the worker vanished mid-turn, e.g. its pane crashed), fire
* {@link TurnListener#onTurnFailed} for it: a delivered message is no longer in the queue, so
* failing queued waiters alone would leave that send's rendezvous hanging until the async
* timeout. Futures and listeners are completed after the monitor is released.
*/
public void drop(String target, Throwable cause) {
Target t = targets.remove(target);
if (t == null) return;
List<Pending> pending;
boolean hadDeliveredTurn;
synchronized (t) {
pending = new ArrayList<>(t.queue);
t.queue.clear();
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
t.postTurnPending = false;
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
}
}
@@ -1,93 +0,0 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.HerdrException;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Set;
/**
* Drives the {@link Injector} by sampling each active worker's {@code agent_status} and
* feeding it in. A single virtual-thread loop polls only targets that have work outstanding,
* so an idle bridge does no herdr traffic.
*
* <p>This is the Stage-1 gate signal. It can later be replaced (or fronted) by a herdr
* {@code events.subscribe} stream without touching the {@link Injector} — the injector only
* consumes {@code onStatus} calls, however they are produced.
*/
public final class StatusPoller {
private static final Logger log = LoggerFactory.getLogger(StatusPoller.class);
private final AgentControl agents;
private final Injector injector;
private final StatusRefiner refiner;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
public StatusPoller(AgentControl agents, Injector injector, long intervalMillis) {
this(agents, injector, new StatusRefiner(agents), intervalMillis);
}
public StatusPoller(AgentControl agents, Injector injector, StatusRefiner refiner,
long intervalMillis) {
this.agents = agents;
this.injector = injector;
this.refiner = refiner;
this.intervalMillis = intervalMillis;
}
/** Start the polling loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("status-poller").start(this::loop);
log.info("status poller started (interval {}ms)", intervalMillis);
}
private void loop() {
while (running) {
Set<String> active = injector.activeTargets();
for (String target : active) {
if (!running) return;
try {
// herdr's agent_status can misreport a settled worker as `unknown`; refine it
// against the pane content before it drives delivery/completion (CB-115).
AgentStatus status = refiner.refine(target, agents.status(target));
injector.onStatus(target, status);
} catch (HerdrException e) {
// The worker's agent is gone — stop trying and unblock its waiters.
if (e.code() != null && e.code().endsWith("_not_found")) {
log.debug("target {} gone; dropping its queue", target);
injector.drop(target, e);
} else {
log.debug("status poll for {} failed (will retry): {}", target, e.getMessage());
}
} catch (RuntimeException e) {
// Never let one target's unexpected error (e.g. an odd agent.get shape) kill
// the single poller thread and stall injection for every worker.
log.warn("unexpected error polling {}; skipping this round", target, e);
}
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the polling loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -1,892 +0,0 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.json.jackson3.JacksonMcpJsonMapperSupplier;
import io.modelcontextprotocol.server.McpServer;
import io.modelcontextprotocol.server.McpSyncServer;
import io.modelcontextprotocol.server.McpSyncServerExchange;
import io.modelcontextprotocol.server.transport.HttpServletStreamableServerTransportProvider;
import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The MCP SERVER face (CB-105): a Streamable-HTTP MCP server whose tools are <em>thin adapters</em>
* over the same {@link MessageService}/{@link Rendezvous} the REST routes use — so the two are
* validated by parity, not by re-implementing behaviour. The primary Opus calls {@code bridge_send}
* / {@code bridge_status}; the worker calls {@code bridge_reply}.
*
* <p>Beyond delegation the primary also manages the fleet here (CB-108): {@code bridge_spawn} /
* {@code bridge_list} / {@code bridge_stop} drive the {@link PeerLauncher} SPI so a worker's whole
* lifecycle is managed through MCP, with each adapter's subscription boundary enforced inside it.
*
* <p>The tool <em>logic</em> lives in package-private static methods returning a
* {@link McpSchema.CallToolResult}, so it is unit-testable without standing up the HTTP transport;
* the SDK owns the wire protocol. Mount {@link #servlet()} at {@code /mcp} on the daemon's Jetty.
*/
public final class BridgeMcp {
private static final long DEFAULT_TIMEOUT_MS = 25_000;
private static final long MAX_TIMEOUT_MS = 120_000;
// bridge_ask blocks the WORKER's own MCP call, which its client caps near 60s — default under
// that so the bridge returns a clean timeout before the client severs the call (CB-205).
private static final long ASK_DEFAULT_TIMEOUT_MS = 55_000;
private static final long ASK_MAX_TIMEOUT_MS = 115_000;
private static final ObjectMapper MAPPER = new ObjectMapper(); // worker-view JSON projections
/** Transport-context key under which the extractor stashes the resolved caller identity. */
static final String CALLER_TERMINAL = "callerTerminal";
/** Transport-context key under which the extractor stashes the caller's PID (for cwd inherit). */
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
/** Transport-context key for the lead's configured name, when the caller is one (CB-530). */
static final String CALLER_NAME = "callerName";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
.mcpEndpoint("/mcp")
// Resolve the caller from the connection (peer PID → herdr pane) in one lookup: the
// worker terminal for bridge_reply (no spoofable arg), and the PID so bridge_spawn can
// inherit the primary's cwd (CB-112). Any contact from a worker marks it available
// (CB-113) — its MCP initialize is the reliable "the agent is up" signal.
.contextExtractor(req -> {
// One resolution per call, shared with the REST surface via CallerResolver so
// the two paths cannot drift on who a caller is.
Principal p = callers != null
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
// CB-532: guard on the ROLE, not on the terminal being null. A named lead now
// carries its pane too, and enrolling a lead in the worker presence map would
// have it counted as an available worker.
if (p.isWorker()) presence.markPresent(p.terminal());
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name(),
CALLER_NAME, orEmpty(p.name())));
})
.build();
this.server = McpServer.sync(transport)
.serverInfo("bridge", "0.1.0")
.capabilities(McpSchema.ServerCapabilities.builder().tools(true).build())
.toolCall(sendTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SEND,
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
// CB-532: remember WHICH lead is waiting on this worker, so its reply nudge goes
// back to that lead rather than to whichever one happened to send first.
primaryRegistry.recordDelegation(str(a, "sessionId"), caller);
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
}
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
.toolCall(replyTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.REPLY, self);
if (denied != null) return denied;
return reply(messages, self, str(req.arguments(), "content"));
})
// bridge_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
.toolCall(askTool(), (exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.ASK, self);
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
})
.toolCall(statusTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
})
.toolCall(pollTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
Map<String, Object> a = req.arguments();
return poll(messages, str(a, "ticket"), str(a, "target"));
})
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
.toolCall(ackTool(), (exchange, req) -> {
Map<String, Object> a = req.arguments();
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.DRAIN, str(a, "target"));
if (denied != null) return denied;
return ack(messages, str(a, "target"), str(a, "msgId"));
})
// Fleet management (CB-108): spawn/list/stop over ClaudeCodeLauncher.
.toolCall(spawnTool(), (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.STOP, paneId);
if (denied != null) return denied;
return stop(sessions, paneId);
})
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return whoami(principal(exchange), sessions);
})
.build();
this.authz = callers;
this.metrics = metrics;
}
/**
* Pre-CB-501 identity: worker if the connection maps to a pane, otherwise the primary. Used
* only by the legacy constructor, where authorization is not enforced anyway.
*/
private static Principal legacyPrincipal(ConnectionIdentity identity, String addr, int port) {
ConnectionIdentity.Caller c = identity.resolve(addr, port);
return c.terminal() != null
? Principal.worker(c.terminal(), c.pid())
: Principal.primary(c.pid());
}
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange), callerName(exchange));
}
/**
* Rebuild a {@link Principal} from the three values the context extractor stashed.
*
* <p>Split out from {@link #principal(McpSyncServerExchange)} so the identity rules are
* reachable without an {@code McpSyncServerExchange} — that is an SDK type this project has no
* mocking library to fabricate, which is why this logic had no test at all until CB-513.
*
* @param role the stashed {@link Role} name, or {@code null} on the legacy path
* @param terminal the worker terminal, or {@code null} for a non-worker
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
return principalFrom(role, terminal, pid, null);
}
/** As {@link #principalFrom(Object, String, long)}, carrying a lead's name (CB-530). */
static Principal principalFrom(Object role, String terminal, long pid, String name) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid, name);
}
/**
* Gate a tool call on the CB-505 table. Returns {@code null} when the call may proceed, or the
* error result to return when it may not.
*/
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
String target) {
return denyFor(principal(exchange), action, target);
}
/**
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}.
*
* <p>This surface exists because the enforcement was previously unreachable from a test: no
* test constructs a {@code BridgeMcp}, so the whole MCP-side gate ran zero times in the suite
* while the REST-side equivalent had ten tests. A security control nothing exercises is a
* claim, not a control.
*
* @return {@code null} when the call may proceed, or the error result to return when it may not
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (authz == null) {
return null; // legacy constructor: authorization not enforced
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
}
String reason = Authz.isUnauthenticated(caller) ? "unauthenticated" : "forbidden";
AuditLog.denied(caller, action, target, reason);
if (metrics != null) {
metrics.inc(BridgedMetrics.AUTH_FAILURES, "reason", reason);
}
return error(reason + ": " + caller.describe() + " may not " + action);
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The lead name resolved from this call's connection, or {@code null} (CB-530). */
private static String callerName(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_NAME);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
try {
return v == null ? -1 : Long.parseLong(v.toString());
} catch (NumberFormatException e) {
return -1;
}
}
private static String orEmpty(String s) {
return s == null ? "" : s;
}
/** The Streamable-HTTP servlet to mount at {@code /mcp} on the daemon's Jetty. */
public HttpServlet servlet() {
return transport;
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
}
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code bridge_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout);
}
/**
* {@code bridge_ask} (CB-205): a worker pauses its delegated turn to ask the primary, blocking
* until the primary answers. The worker is identified by its connection ({@code callerTerminal}),
* never an argument — a {@code null} means the caller is not a known worker.
*/
static McpSchema.CallToolResult ask(MessageService messages, String callerTerminal, String question, Long timeoutMs) {
if (callerTerminal == null) {
return error("bridge_ask is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (isBlank(question)) {
return error("question is required");
}
long timeout = Math.clamp(timeoutMs == null ? ASK_DEFAULT_TIMEOUT_MS : timeoutMs, 1, ASK_MAX_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(callerTerminal, question, timeout);
return switch (r.outcome()) {
case ANSWERED -> text(r.answer());
case NO_WAITER -> error("no primary is awaiting this turn — bridge_ask only works while a "
+ "bridge_send delegation is open to answer it");
case TIMED_OUT -> text("[no answer within " + timeout + "ms — the primary did not respond; "
+ "proceed on your best judgement, then call bridge_reply to end the turn]");
};
}
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called bridge_reply — hand back the scraped
// transcript tail, flagged so the primary knows it isn't a structured reply.
case COMPLETED_UNREPLIED -> text(
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
};
}
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
if (replies.isEmpty()) {
return text("[]");
}
return text(json(replies));
}
if (isBlank(ticket)) {
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
return switch (v.phase()) {
case DONE -> text(v.replySource() != null && v.replySource().equals("transcript")
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
/**
* {@code bridge_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*/
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("bridge_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (content == null) {
return error("content is required");
}
messages.reply(callerTerminal, content);
return text("delivered");
}
/** {@code bridge_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
return error("target and msgId are required");
}
messages.ackReply(target, msgId);
return text("acknowledged " + msgId);
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
return text(messages.status(sessionId).name().toLowerCase());
} catch (HerdrException e) {
return error("herdr error for session " + sessionId + ": " + e.getMessage());
}
}
/**
* {@code bridge_whoami}: the caller's own identity, as the daemon already resolved it.
*
* <p>Every other tool <em>consumes</em> this identity — the authorization gate, the reply
* rendezvous, the cwd inherit — but none reported it, so an agent had to infer its own role
* from side channels the daemon does not control: a charter string in its system prompt, the
* name its MCP mount happens to carry, or {@code ANTHROPIC_BASE_URL} (which Claude-model
* workers do not set). The failure mode of guessing is asymmetric and silent: a primary that
* mistakes itself for a worker is refused by {@link Authz} and learns immediately, while a
* worker that mistakes itself for the primary ends its turn without {@code bridge_reply} and
* the sender simply receives nothing. This tool removes the guess.
*
* <p>For a worker the session registry adds what it knows about that session. A worker the
* registry has no record of — one that outlived a daemon restart — still gets its role and
* {@code sessionId}, which is the load-bearing part.
*/
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (caller.isArchitect()) {
// CB-548: the role reads "architect"; the name is the gateway-local slot the pane is
// bound to, and the pane itself so a peer knows where to reach it.
if (caller.name() != null) {
m.put("architect", caller.name());
}
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
// primary for authorization; the name is additive so no existing reader breaks.
if (caller.name() != null) {
m.put("leader", caller.name());
}
// CB-532: a lead's own pane, so it can tell a peer where to reach it — and so an
// operator can read off which tab hosts which lead without going to herdr.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
m.put("sessionId", caller.terminal());
sessions.roster().stream()
.filter(s -> caller.terminal().equals(s.terminalId()))
.findFirst()
.ifPresent(s -> {
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
if (s.ownerTerminal() != null) {
m.put("owner", s.ownerTerminal());
}
});
return text(json(m));
}
// --- fleet management logic (CB-108 / CB-301) --------------------------------------------
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
return error("herdr error spawning worker: " + e.getMessage());
}
}
/** Build a {@link WorktreeRequest} from {@code bridge_spawn}'s optional {@code worktree}/{@code ticket} args. */
private static WorktreeRequest worktreeRequest(Map<String, Object> a) {
Object w = a.get("worktree");
if (w == null || Boolean.FALSE.equals(w)) {
return null;
}
String ticket = str(a, "ticket");
if (w instanceof String s) {
if (s.isBlank() || "false".equalsIgnoreCase(s)) {
return null;
}
if ("true".equalsIgnoreCase(s)) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(s, null);
}
if (w instanceof Boolean b && b) {
if (isBlank(ticket)) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/**
* {@code bridge_list}: the whole fleet — {@code leads} and {@code workers} — each merged with
* live herdr status. CB-519 decoupled the registry key (a host-unique id) from the herdr pane
* coordinate, so the join is on the terminal id, which both the session and the live agent carry.
*
* <p>CB-535 added the {@code leads} half. Until then this listed the worker roster alone, and a
* lead asking "who else is here?" got an empty array — which reads as <em>no peers</em> but
* actually means <em>no workers spawned</em>. There was no way at all for a lead to learn a
* peer's address; it had to be carried across by a human. Both halves are reported even when a
* half is empty, so an empty {@code workers} can no longer be mistaken for an empty fleet.
*
* <p>Leads are drawn from the resolver rather than from a second registry, so an address listed
* here is one that would actually resolve as a lead — see {@link CallerResolver#leads()}. The
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* @param leads terminal_id → lead name, live from the resolver
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
return text(json(Map.of("leads", leadRows, "workers", out)));
} catch (HerdrException e) {
return error("herdr error listing the fleet: " + e.getMessage());
}
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code bridge_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
String selfTerm) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", terminal);
m.put("name", name);
m.put("status", live == null || live.status() == null
? "unknown" : live.status().name().toLowerCase());
if (terminal.equals(selfTerm)) {
m.put("self", true);
}
return m;
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
return error("paneId is required");
}
try {
sessions.release(paneId);
return text("stopped " + paneId);
} catch (HerdrException e) {
return error("herdr error stopping " + paneId + ": " + e.getMessage());
}
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
private static String json(Object o) {
try {
return MAPPER.writeValueAsString(o);
} catch (Exception e) {
return String.valueOf(o);
}
}
// --- tool schemas --------------------------------------------------------------------------
private static McpSchema.Tool sendTool() {
return tool("bridge_send",
"Delegate a task to a worker session. By default blocks until the worker replies and "
+ "returns its reply (or a 'still working / queued' note on timeout). Pass wait:false "
+ "for a long task to return a ticket immediately, then poll it with bridge_poll. To "
+ "answer a worker's bridge_ask, pass its turnId (with content) instead of sessionId.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id (herdr terminal_id) to delegate to"),
"content", stringProp("The task/message to send to the worker (or your answer, with turnId)"),
"timeoutMs", Map.of("type", "integer", "description", "Max ms to wait for a reply (blocking mode)"),
"wait", Map.of("type", "boolean",
"description", "Block for the reply (default true); false returns a ticket to poll"),
"turnId", stringProp("When answering a worker's bridge_ask, its question turnId — "
+ "routes your answer back into the same turn (omit for a normal delegation)")),
List.of("content")));
}
private static McpSchema.Tool askTool() {
// No target/session arg — the worker's identity is resolved from the connection.
return tool("bridge_ask",
"Pause your current delegated turn to ask the primary a question, blocking until it "
+ "answers — then resume the same turn with the answer. Use this when only the "
+ "primary has a decision or detail you need to continue. You do not address the "
+ "primary; identity is resolved from your connection.",
objectSchema(Map.of(
"question", stringProp("The question to put to the primary"),
"timeoutMs", Map.of("type", "integer",
"description", "Max ms to wait for the primary's answer")),
List.of("question")));
}
private static McpSchema.Tool pollTool() {
return tool("bridge_poll",
"Check an async delegation (a bridge_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by bridge_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
List.of()));
}
private static McpSchema.Tool ackTool() {
return tool("bridge_ack",
"Acknowledge (remove) a specific reply from a worker's inbox. Use when the primary "
+ "has processed a reply and wants to confirm it, leaving other pending replies "
+ "in the inbox for later drain.",
objectSchema(Map.of(
"target", stringProp("Worker session id whose inbox to ack from"),
"msgId", stringProp("The message id to acknowledge")),
List.of("target", "msgId")));
}
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to bridge_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'workers' are the "
+ "sessions delegated to — each with sessionId, paneId, profile, state, "
+ "optional worktree/branch/owner, and live herdr status. An empty 'workers' "
+ "means no workers are spawned; it says nothing about peers.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool stopTool() {
return tool("bridge_stop",
"Tear down a worker session by its paneId (from bridge_spawn or bridge_list).",
objectSchema(Map.of(
"paneId", stringProp("The worker's paneId to stop")),
List.of("paneId")));
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the caller's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked bridge_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
+ "it — never to answer a worker, whose turn it is not.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
}
private static McpSchema.Tool statusTool() {
return tool("bridge_status",
"Get the live lifecycle status (idle/working/blocked/unknown) of a worker session.",
objectSchema(Map.of(
"sessionId", stringProp("The worker session id to query")),
List.of("sessionId")));
}
private static McpSchema.Tool whoamiTool() {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
+ "messaged you, never to answer a worker), 'architect' (you delegate turns "
+ "and reply/ask as your own pane, but cannot spawn/stop/drain), or 'worker' "
+ "(you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus 'leader'/'architect' naming "
+ "which one you are, your own sessionId, and profile/worktree/branch when "
+ "you are a worker. Call this first when following role-conditional "
+ "instructions rather than guessing.",
objectSchema(Map.of(), List.of()));
}
// --- small helpers -------------------------------------------------------------------------
// The SDK 2.0.0 deprecates its own Tool builders without a stable replacement — isolate it here.
@SuppressWarnings("deprecation")
private static McpSchema.Tool tool(String name, String description, Map<String, Object> inputSchema) {
return McpSchema.Tool.builder(name).description(description).inputSchema(inputSchema).build();
}
private static Map<String, Object> objectSchema(Map<String, Object> properties, List<String> required) {
return Map.of("type", "object", "properties", properties, "required", required);
}
private static Map<String, Object> stringProp(String description) {
return Map.of("type", "string", "description", description);
}
private static McpSchema.CallToolResult text(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s == null ? "" : s).build();
}
private static McpSchema.CallToolResult error(String s) {
return McpSchema.CallToolResult.builder().addTextContent(s).isError(true).build();
}
private static String str(Map<String, Object> args, String key) {
Object v = args.get(key);
return v == null ? null : v.toString();
}
private static Long timeoutMs(Map<String, Object> args) {
Object v = args.get("timeoutMs");
return v instanceof Number n ? n.longValue() : null;
}
private static long clamp(long ms) {
return Math.clamp(ms, 1, MAX_TIMEOUT_MS);
}
private static boolean isBlank(String s) {
return s == null || s.isBlank();
}
}
@@ -1,66 +0,0 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.PaneLocator;
/**
* Resolves <em>who is calling</em> an MCP tool from the connection alone — the anti-spoofing
* identity model of the MCP contract. It ties the connection's loopback peer PID (from the OS)
* to a herdr agent pane (from herdr), yielding the caller's worker {@code terminal_id}. A caller
* that maps to no worker pane — the primary, or an off-host client — resolves to {@code null}.
*
* <p>Both sources are authoritative and unforgeable: the OS reports the real connecting PID, and
* herdr owns the PID→pane mapping. A worker cannot claim to be another worker, nor the primary.
* Single-host only (the herd shares the {@code bridged} host); the token path is the split-host
* fallback.
*/
public final class ConnectionIdentity {
private final PaneLocator panes;
private final PeerPidLookup pids;
private final ProcessCwdLookup cwds;
/** Identity only (no cwd resolution — {@link #cwdForPid} returns {@code null}). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids) {
this(panes, pids, _ -> null);
}
/** Identity plus cwd resolution (CB-112 — inherit the primary's directory on spawn). */
public ConnectionIdentity(PaneLocator panes, PeerPidLookup pids, ProcessCwdLookup cwds) {
this.panes = panes;
this.pids = pids;
this.cwds = cwds;
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client) and its {@code pid} (or {@code -1} if not resolvable).
*/
public record Caller(String terminal, long pid) {
}
/** Resolve the caller's terminal and PID from one peer-PID lookup. */
public Caller resolve(String remoteAddr, int remotePort) {
if (!isLoopback(remoteAddr)) {
return new Caller(null, -1); // only same-host callers can be workers
}
long pid = pids.pidForLocalPort(remotePort);
return new Caller(panes.terminalForPid(pid), pid);
}
/**
* The calling worker's {@code terminal_id}, or {@code null} if the caller is not a known
* on-host worker (treat as the primary).
*/
public String callerTerminal(String remoteAddr, int remotePort) {
return resolve(remoteAddr, remotePort).terminal();
}
/** The working directory of {@code pid} (the primary's cwd on an MCP spawn), or {@code null}. */
public String cwdForPid(long pid) {
return pid > 0 ? cwds.cwdForPid(pid) : null;
}
private static boolean isLoopback(String addr) {
return "127.0.0.1".equals(addr) || "::1".equals(addr) || "0:0:0:0:0:0:0:1".equals(addr);
}
}
@@ -1,242 +0,0 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
* publish to an agent owned by another gateway; in that case the publisher must not compete for
* deliveries.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
* (broker delivery latency). Callers that need the reply drained poll (as the primary already does);
* the contract test waits for visibility. This is inherent to broker-backed delivery, not a defect.
*
* <p>The default deploy targets LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1 (URI-only swap),
* so the {@code @Tag("contract")} integration test runs against a RabbitMQ container.
*/
public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final Logger log = LoggerFactory.getLogger(AmqpReplyInbox.class);
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
public static AmqpReplyInbox open(String uri) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this.connection = connection;
try {
this.channel = connection.createChannel();
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@Override
public void handleRecoveryStarted(Recoverable recoverable) {
// no-op: we act once recovery completes
}
});
}
}
@Override
public void own(String target) {
synchronized (channelLock) {
if (consumerTags.containsKey(target)) {
return; // already owning this target
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
}
}
}
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
}
@Override
public List<InboxMessage> peek(String target) {
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
}
synchronized (perTarget) {
return perTarget.values().stream().map(Held::message).toList();
}
}
@Override
public void ack(String target, String msgId) {
var perTarget = held.get(target);
if (perTarget == null) {
return;
}
Held h;
synchronized (perTarget) {
h = perTarget.remove(msgId);
}
if (h == null) {
return; // never held (or already acked) — no-op
}
try {
synchronized (channelLock) {
channel.basicAck(h.deliveryTag(), false);
}
} catch (IOException e) {
// Ack didn't reach the broker: restore the entry so a later ack (or a redelivery after
// reconnect) can retry. Keeps the at-least-once contract — a reply is never silently lost.
synchronized (perTarget) {
perTarget.putIfAbsent(msgId, h);
}
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
long tag = delivery.getEnvelope().getDeliveryTag();
if (msgId == null || msgId.isBlank()) {
msgId = Long.toHexString(tag); // synthesize an id so dedup still has a key
}
String content = new String(delivery.getBody(), StandardCharsets.UTF_8);
var perTarget = held.computeIfAbsent(target, _ -> new LinkedHashMap<>());
boolean duplicate;
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
duplicate = true;
} else {
perTarget.put(msgId, new Held(tag, new InboxMessage(msgId, target, content)));
duplicate = false;
}
}
if (duplicate) {
// Redelivered duplicate: ack the new tag and drop it so the broker stops resending.
synchronized (channelLock) {
channel.basicAck(tag, false);
}
}
};
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
log.debug("AMQP connection close: {}", e.toString());
}
}
}
@@ -1,527 +0,0 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
* the worker returns a <em>structured reply</em> via {@code bridge_reply} (the {@link Rendezvous}),
* then hand that reply back. Delivery is the {@link Injector}'s job (the background poller sends it
* when the worker is injectable); this service never drives the injector or scrapes the terminal —
* completion is the worker's explicit reply, not a guess about {@code agent_status}.
*
* <p>Sends are serialized per session so exactly one reply can be outstanding per worker, which is
* what lets a reply map unambiguously to its send (no cross-talk between concurrent callers).
*
* <p>If the worker never replies within the timeout, the caller gets a typed "still working" /
* "queued" outcome — the message may still be mid-flight. A finished-but-unreplied turn is caught
* by the CB-106 completion fallback (see {@link Rendezvous#resolveCompletion}).
*
* <p><strong>Async fire-and-poll (CB-107).</strong> A caller's MCP client caps a blocking call at
* ~60s, but a real delegated task runs for minutes. {@link #sendAsync} therefore runs the same
* blocking {@link #send} on a background virtual thread and hands back a <em>ticket</em> the caller
* polls with {@link #poll}. The blocking and async paths share one code path (and the same per-target
* serialization), so async inherits the reply + completion resolution behaviour for free.
*/
public final class MessageService {
private static final Logger log = LoggerFactory.getLogger(MessageService.class);
/**
* The window a fire-and-poll send waits for resolution — generous, since no caller is blocked on
* it; a real delegated task resolves (reply or completion) well within this, and only a genuinely
* hung worker rides it out.
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
/** The worker called {@code bridge_reply}; {@code text} holds the structured answer. */
REPLIED,
/**
* The worker's delegated turn finished without a {@code bridge_reply} (CB-106 fallback);
* {@code text} is the scraped transcript tail rather than a structured answer.
*/
COMPLETED_UNREPLIED,
/**
* The worker ran the turn then wedged in an unrecoverable state (CB-109); {@code text} is the
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
TIMED_OUT_WORKING,
/** Timed out before delivery — the message is still queued for the worker. */
TIMED_OUT_QUEUED,
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code bridge_ask} already timed out or was answered.
*/
STALE_TURN
}
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code bridge_reply}
* for {@link Outcome#REPLIED}, a scraped transcript tail for
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
public Reply(Outcome outcome, String text) {
this(outcome, text, null);
}
/** Whether the worker's turn actually finished with an answer (replied or scraped). */
public boolean completed() {
return outcome == Outcome.REPLIED || outcome == Outcome.COMPLETED_UNREPLIED;
}
}
/** How a worker's {@code bridge_ask} (CB-205) resolved. */
public enum AskOutcome {
/** The primary answered; {@link AskResult#answer} carries it. */
ANSWERED,
/** No delegation was open to surface the question to — the worker has no one to ask. */
NO_WAITER,
/** The primary did not answer within the window. */
TIMED_OUT
}
/** The outcome of a worker's {@code bridge_ask}: how it resolved and (if answered) the answer. */
public record AskResult(AskOutcome outcome, String answer) {
}
/** Lifecycle phase of an async delegation ticket. */
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
FAILED
}
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
/**
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
this(agents, injector, rendezvous, inbox, pushLoop, null);
}
/**
* As above, with a metric registry (CB-502). Instrumenting here rather than at the REST and MCP
* edges means both surfaces are counted by one piece of code and cannot drift.
*
* @param metrics nullable — when null, nothing is recorded
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox) {
this(agents, injector, rendezvous, inbox, null);
}
/** Backward-compatible constructor that uses a default {@link InMemoryReplyInbox}. */
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous) {
this(agents, injector, rendezvous, new InMemoryReplyInbox());
}
/** Current lifecycle status of a worker (the {@code GET /sessions/{id}/status} surface). */
public AgentStatus status(String target) {
return agents.status(target);
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
* result is <em>not</em> a failure — the reply is held for later drain.
*
* <p><strong>Do NOT use this for mid-turn questions.</strong> {@code bridge_ask} /
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* @return always {@code true} — the reply either resolved a live send or was queued
*/
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
count(BridgedMetrics.REPLIES, "path", "rendezvous");
return true; // a live send took it — unchanged fast path
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// A rising inbox share is the signal CB-307 exists to make visible: the worker finished but
// nobody was waiting, so delivery now depends on the push loop and a drain.
count(BridgedMetrics.REPLIES, "path", "inbox");
if (pushLoop != null) {
pushLoop.onReplyQueued(session);
}
return true; // held, not lost
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
if (metrics != null) {
metrics.inc(name, labels);
}
}
/** Count a send's terminal outcome and pass the reply through unchanged. */
private Reply recorded(Reply r) {
String label = sendOutcomeLabel(r.outcome());
if (label != null) {
count(BridgedMetrics.SENDS, "outcome", label);
}
return r;
}
/** Map a terminal send outcome to its metric label, or {@code null} for non-terminal ones. */
private static String sendOutcomeLabel(Outcome o) {
return switch (o) {
case REPLIED -> "replied";
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
/**
* Abandon any send still waiting on {@code target} because its session has gone away (CB-516).
*
* <p>Without this, tearing a worker down left its rendezvous waiter open: a blocking
* {@code bridge_send} kept blocking, and an async one kept reporting {@code PENDING} until
* {@link #ASYNC_TIMEOUT_MS} — thirty minutes — even though the worker provably no longer
* existed and the delegation could never complete. Worse, {@code poll} already had the evidence
* (it calls {@code liveStatus} to build its detail string and gets back {@code "unknown"}) and
* reported {@code PENDING} anyway.
*
* <p>Resolving the waiter as a failure — rather than letting it time out — also means the
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* @return true if a live waiter was failed
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
}
return failed;
}
/**
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
* so that a subsequent drain or peek no longer returns it.
*/
public void ackReply(String target, String msgId) {
inbox.ack(target, msgId);
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
public List<ReplyInbox.InboxMessage> drainReplies(String target) {
var messages = inbox.peek(target);
for (var msg : messages) {
inbox.ack(target, msg.msgId());
}
return messages;
}
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
} finally {
rendezvous.close(target, reply);
}
} finally {
lock.unlock();
}
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code bridge_send}, then block this (worker) call until
* the primary answers via {@link #answer} or {@code timeoutMillis} elapses. Identity is the
* worker's own session — it does not address the primary.
*
* <p>Returns {@link AskOutcome#NO_WAITER} when no delegation is open to surface the question to
* (nothing to answer it), {@link AskOutcome#ANSWERED} with the primary's answer, or
* {@link AskOutcome#TIMED_OUT} if the primary stayed silent. The worker resumes its turn either
* way — an answered ask hands back the answer; an unanswered one leaves it to proceed alone.
*/
public AskResult ask(String workerSession, String question, long timeoutMillis) {
Rendezvous.AskTicket ticket = rendezvous.openAsk(workerSession);
// Only the freshly-opening caller surfaces the question; a coalesced duplicate simply blocks on
// the shared answer future that the fresh owner is already responsible for.
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the primary's answer for " + workerSession, e);
} finally {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
}
}
}
/**
* The primary's answer to a worker's {@code bridge_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
* eventual {@code bridge_reply} as it finishes the resumed turn. The worker session is derived from
* {@code turnId}, never a caller argument.
*
* <p>Unlike {@link #send} this does not re-inject through the {@link Injector}: the worker is
* mid-turn (already picked up), so the answer flows back through its own open {@code bridge_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
if (!rendezvous.answerAsk(turnId, content)) {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
return new Reply(Outcome.TIMED_OUT_WORKING, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + workerSession, e);
} finally {
rendezvous.close(workerSession, reply);
}
} finally {
lock.unlock();
}
}
/**
* Fire-and-poll variant of {@link #send}: deliver {@code content} to {@code target} on a
* background virtual thread and return immediately with a ticket to {@link #poll}. This is how a
* long task is delegated without tripping the caller's MCP client call timeout.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
return agents.status(target).name().toLowerCase();
} catch (RuntimeException e) {
return "unknown";
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
}
/** Map a rendezvous {@link Rendezvous.Kind} onto its send {@link Outcome} (shared by send/answer). */
private static Outcome outcomeOf(Rendezvous.Kind kind) {
return switch (kind) {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case QUESTION -> Outcome.QUESTION;
};
}
private static boolean tryLock(ReentrantLock lock, long millis) {
try {
return lock.tryLock(Math.max(0, millis), TimeUnit.MILLISECONDS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting the session send lock", e);
}
}
private static long remainingMillis(long deadlineNanos) {
return (deadlineNanos - System.nanoTime()) / 1_000_000L;
}
}
@@ -1,188 +0,0 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final ScheduledExecutorService scheduler;
private final int maxReminders;
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs) {
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, null);
}
/** As above, with a metric registry (CB-512) so push outcomes are counted. */
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.scheduler = scheduler;
this.maxReminders = maxReminders;
this.backoffMs = backoffMs;
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(BridgedMetrics.PUSH_NUDGES, "outcome", outcome);
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
/** The action the loop should take for a target at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
// CB-532: the destination is per-delegation — the lead that sent this worker its work, not
// "the primary". With two leads orchestrating one fleet the singular question has no right
// answer, and answering it anyway interrupted whichever lead happened to call bridge_send
// first with results it never asked for.
var nudgeTarget = primaryRegistry.nudgeTargetFor(target);
if (nudgeTarget.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, stopping reminder", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
countNudge("exhausted");
return Action.STOP;
}
String leadTerminal = nudgeTarget.get();
AgentStatus status;
try {
status = agents.status(leadTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for lead {}, will retry", leadTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: lead {} is {} (not injectable), waiting", leadTerminal, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
}
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
// Re-read rather than threading it down from decide(): the delegating lead can change
// between the decision and the injection, and the nudge should follow the current one.
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: lead for {} disappeared before the nudge could be sent", target);
return;
}
String leadTerminal = lead.get();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(leadTerminal, nudge);
log.debug("push: nudge {}/{} sent to lead {} for target {}",
reminderCount + 1, maxReminders, leadTerminal, target);
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge lead {} for target {} (reminder {}/{}): {}",
leadTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
}
/** @see #stop() */
public void close() {
stop();
}
}
@@ -1,94 +0,0 @@
package dev.ltms.bridged.peer;
import java.util.List;
import java.util.Set;
/**
* SPI for materializing a connected peer — the only way the bridge core creates or tears down
* a peer process. Every launcher is a first-party, in-tree adapter selected by (future) profile
* config; today's single adapter is the {@code ClaudeCodeLauncher} / Claude Code over herdr.
*
* <p>The core delegates spawn and teardown to this interface without knowing how the peer is set
* up. Environment variables, CLI flags, subscription guards, transport (herdr tab/pane) layout,
* and naming conventions are all adapter-private — the core sees only the returned
* {@link PeerHandle} whose {@code id()} is the registry/routing key.
*
* <p>The interface is a superset of what {@code SessionManager} and {@code Bridged.main} call
* on the concrete launcher today.
*/
public interface PeerLauncher {
/**
* The set of {@link Capability capabilities} this launcher declares. A peer whose profile
* opts into a git-forge token should include {@link Capability#SELF_PR}; the base set for
* the Claude Code herdr adapter is always {@code MID_TURN_ASK, WORKTREE, ORPHAN_REAP}.
*/
Set<Capability> capabilities();
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
*
* @param req the spawn parameters (profile, requested cwd, caller cwd)
* @return a handle whose {@link PeerHandle#id()} is the registry/routing key
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
PeerHandle spawn(SpawnRequest req);
/**
* The configured worker profile names — the set of names {@code spawn(profileName)} accepts.
*/
Set<String> profiles();
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*/
String defaultProfile();
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
*
* @return the resolved absolute path, never null/blank
*/
String effectiveCwd(SpawnRequest req);
/**
* The parity-overlay file list for {@code profileName} (default list when unset). Used by
* worktree provisioning to copy config files into the isolated checkout before spawning.
*/
List<String> parityOverlay(String profileName);
/**
* The set of all agents this launcher currently tracks, transport-specific. Each element
* exposes at minimum a pane-like {@code id()} matching this launcher's {@link PeerHandle}
* scheme, plus transport-level status. Callers merge this set with the session registry to
* build a live roster view.
*/
List<?> list();
/**
* Reap orphaned peers left behind by a prior daemon process. Only peers whose naming scheme
* matches this launcher's and whose nonce differs from the current process are eligible.
* Best-effort: a failure to list or to stop any one peer is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
int reapOrphanWorkers();
/**
* Tear a peer down by its registry/routing key ({@link PeerHandle#id()}). Tolerates an
* already-gone peer. Also cleans up launcher-private resources (e.g. empty dedicated tabs)
* when safe to do so.
*/
void stop(String id);
/**
* Discard the context of the peer identified by {@code id}. Implementations must bypass normal
* bridge delivery/turn accounting. Unsupported peer kinds return {@code false} without sending
* a guessed command.
*
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
}
@@ -1,24 +0,0 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
*
* <p>{@code sessionName} and {@code resumeSessionId} carry the session's durable identity (CB-547a):
* the bridge's LOGICAL name for the session (stable across restarts, meaningful to an operator)
* and the peer's OWN prior session id to resume, respectively. Both are <em>opted in</em> — either
* may be null/blank, in which case the launcher derives a display name and mints a fresh session.
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
/** Back-compat: a spawn with no session identity (fresh session, launcher-derived name). */
public SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
this(profileName, requestedCwd, callerCwd, null, null);
}
}
@@ -1,22 +0,0 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
return new PlacementCandidate(d, null, 1.0f, null);
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
}
throw new PlacementException("no worker profiles configured");
}
}
@@ -1,19 +0,0 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
import java.util.function.Function;
/**
* Everything a {@link PlacementPolicy} needs to make one selection.
*
* @param defaultProfile profile a {@code fixed} policy should return (may be {@code null})
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
}
@@ -1,65 +0,0 @@
package dev.ltms.bridged.placement;
import java.util.ArrayList;
import java.util.List;
/**
* Shared filtering and empty-set reporting used by the built-in placement policies.
*/
final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* Candidates that are not known-unreachable and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (ctx.unreachable().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
}
/**
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
*/
static PlacementException emptyException(PlacementContext ctx) {
int atCap = 0;
int unreachable = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
}
}
int total = ctx.candidates().size();
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
if (unreachable == total) {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
}
}
@@ -1,563 +0,0 @@
package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.auth.AuditLog;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
import io.javalin.http.Context;
import jakarta.servlet.http.HttpServlet;
import org.eclipse.jetty.servlet.ServletHolder;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* The REST surface — {@code bridged}'s contract and its testability seam. Every
* feature is reachable here without Claude or MCP in the loop, so each is an
* acceptance test against plain HTTP. MCP tools (later) are thin adapters over these
* same endpoints and are validated by parity, not by re-implementing behaviour.
*
* <p>Built from injected collaborators so tests supply fakes and run on an ephemeral
* port; {@code main} supplies the real Unix-socket client and worker service.
*/
public final class BridgedApp {
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
/** Blocking window for a worker's bridge_ask (CB-205); the worker's MCP client caps its own call. */
private static final long DEFAULT_ASK_TIMEOUT_MS = 55_000;
private static final long MAX_ASK_TIMEOUT_MS = 115_000;
/** Context attribute under which the resolved caller is stashed by the auth filter. */
private static final String CALLER = "bridged.caller";
private final HerdrClient herdr;
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
private final ObjectMapper mapper = new ObjectMapper();
/**
* Legacy constructor — no identity resolution and no authorization, exactly as the REST surface
* behaved before CB-501. Retained so existing acceptance tests keep exercising handler
* behaviour without each needing an auth fixture.
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
/**
* @param auth resolves each request's {@link Principal}; {@code null} disables authorization
* entirely (legacy). {@code main} always supplies one.
* @param metrics registry to instrument and expose at {@code GET /metrics}; {@code null} omits
* the endpoint
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
this.sessions = sessions;
this.messages = messages;
this.presence = presence;
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
public Javalin build() {
Javalin app = Javalin.create(cfg -> {
cfg.showJavalinBanner = false;
if (mcpServlet != null) {
// The MCP server shares the daemon's port; Jetty routes /mcp to its servlet.
cfg.jetty.modifyServletContextHandler(h ->
h.addServlet(new ServletHolder(mcpServlet), "/mcp"));
}
});
// CB-501: resolve identity once per request, before any handler. /mcp does NOT pass through
// here — it is a raw servlet on Jetty's context handler — so BridgeMcp enforces separately
// against the same CallerResolver. Any check that lives in only one place is not a control.
if (auth != null) {
app.before(ctx -> ctx.attribute(CALLER,
auth.resolve(ctx.req().getRemoteAddr(), ctx.req().getRemotePort(),
ctx.header("Authorization"))));
}
app.get("/healthz", this::healthz);
if (metrics != null) {
app.get("/metrics", this::metrics);
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
app.post("/sessions/{id}/ask", this::askMessage); // bridge_ask (worker → primary, CB-205)
app.get("/sessions/{id}/status", this::sessionStatus); // bridge_status
app.get("/tasks/{ticket}", this::taskStatus); // poll an async (wait:false) send
return app;
}
/**
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
* proceed; otherwise writes the error response and returns {@code false}.
*
* <p>401 vs 403 is a real distinction here: 401 means "you presented no usable identity" (a
* credential problem the caller can fix), 403 means "you are authenticated, but this is not
* yours" (a worker reaching for another worker's session, or for orchestration).
*/
private boolean allow(Context ctx, Authz.Action action, String target) {
if (auth == null) {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return true;
}
if (Authz.isUnauthenticated(caller)) {
AuditLog.denied(caller, action, target, "unauthenticated");
countAuthFailure("unauthenticated");
ctx.status(401).json(Map.of("error", "unauthenticated",
"detail", "present Authorization: Bearer <token>"));
} else {
AuditLog.denied(caller, action, target, "forbidden");
countAuthFailure("forbidden");
ctx.status(403).json(Map.of("error", "forbidden",
"detail", caller.describe() + " may not " + action + " on "
+ (target == null ? "this resource" : target)));
}
return false;
}
private void countAuthFailure(String reason) {
if (metrics != null) {
metrics.inc("bridged_auth_failures_total", "reason", reason);
}
}
/** Prometheus scrape endpoint (CB-502). */
private void metrics(Context ctx) {
if (!allow(ctx, Authz.Action.METRICS, null)) {
return;
}
ctx.status(200).contentType("text/plain; version=0.0.4; charset=utf-8").result(metrics.render());
}
/** Liveness + herdr reachability. 200 when herdr answers ping, 503 otherwise. */
private void healthz(Context ctx) {
try {
JsonNode pong = herdr.call("ping");
ctx.status(200).json(Map.of(
"status", "ok",
"herdr", Map.of(
"version", pong.path("version").asText(""),
"protocol", pong.path("protocol").asInt())));
} catch (HerdrException e) {
ctx.status(503).json(Map.of(
"status", "degraded",
"herdr", "unreachable",
"detail", e.getMessage()));
}
}
/** Sessions view derived from herdr {@code workspace.list} (one workspace → one row). */
private void sessions(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
JsonNode result = herdr.call("workspace.list");
List<Map<String, Object>> out = new ArrayList<>();
for (JsonNode w : result.path("workspaces")) {
out.add(Map.of(
"id", w.path("workspace_id").asText(""),
"label", w.path("label").asText(""),
"focused", w.path("focused").asBoolean(false),
"paneCount", w.path("pane_count").asInt(),
"agentStatus", w.path("agent_status").asText("unknown")));
}
ctx.status(200).json(Map.of("sessions", out));
}
/** Discovery: every agent herdr tracks, keyed by its Claude session UUID. */
private void agents(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of("agents",
workers.list().stream().map(Agent.class::cast).map(BridgedApp::view).toList()));
}
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
// CB-519: the registry key is a host-unique id, not the pane coordinate — join on terminal.
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
/** The configured worker profiles and which one a no-argument spawn uses. */
private void profiles(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
ctx.status(200).json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
}
/**
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
String ticket = ctx.queryParam("ticket");
if (profile == null || profile.isBlank() || cwd == null || cwd.isBlank()
|| worktree == null || worktree.isBlank()) {
try {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
if (ticket == null || ticket.isBlank()) ticket = b.path("ticket").asText(null);
}
} catch (Exception ignored) {
// A malformed/empty body just means "no overrides" → fall through to defaults.
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
try {
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
} catch (PeerUnreachableException e) {
ctx.status(502).json(Map.of("error", "spawn_timeout", "detail", e.getMessage()));
}
}
private static WorktreeRequest worktreeRequest(String worktree, String ticket) {
if (worktree == null || worktree.isBlank() || "false".equalsIgnoreCase(worktree)) {
return null;
}
if ("true".equalsIgnoreCase(worktree)) {
if (ticket == null || ticket.isBlank()) {
throw new IllegalArgumentException("worktree=true requires a ticket slug");
}
return new WorktreeRequest(ticket, null);
}
return new WorktreeRequest(worktree, null);
}
private static String blankToNull(String s) {
return (s == null || s.isBlank()) ? null : s;
}
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
}
sessions.release(paneId);
ctx.status(204);
}
/**
* The blocking delegation call (CB-104): inject {@code content} into the worker via the
* status-gated injector and block until the worker returns a structured {@code bridge_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.SEND, id)) {
return;
}
String content;
String turnId;
long timeout;
boolean wait;
try {
JsonNode body = mapper.readTree(ctx.body());
content = body.path("content").asText("");
turnId = body.path("turnId").asText(null);
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (content.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
// Answering a worker's bridge_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
return;
}
if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return;
}
try {
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/**
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
* bridge_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
*/
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
switch (reply.outcome()) {
case QUESTION -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "question",
"question", reply.text(), "turnId", reply.turnId()));
case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured bridge_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
String source = reply.outcome() == MessageService.Outcome.REPLIED ? "reply" : "transcript";
ctx.status(200).json(Map.of("sessionId", id, "reply", reply.text(), "replySource", source));
}
default -> ctx.status(202).json(Map.of(
"sessionId", id,
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
}
/**
* A worker's mid-turn question ({@code bridge_ask}, CB-205) — surfaces to the primary's open
* blocking send and blocks until it answers. 200 with the answer, 409 if no delegation is open,
* 202 if the primary stayed silent.
*/
private void askMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.ASK, id)) {
return;
}
String question;
long timeout;
try {
JsonNode body = mapper.readTree(ctx.body());
question = body.path("question").asText("");
timeout = body.path("timeoutMs").asLong(DEFAULT_ASK_TIMEOUT_MS);
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
if (question.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "question is required"));
return;
}
timeout = Math.clamp(timeout, 1, MAX_ASK_TIMEOUT_MS);
MessageService.AskResult r = messages.ask(id, question, timeout);
switch (r.outcome()) {
case ANSWERED -> ctx.status(200).json(Map.of("sessionId", id, "answered", true, "answer", r.answer()));
case NO_WAITER -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "no_pending_send",
"detail", "no primary is awaiting this turn to answer a question"));
case TIMED_OUT -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "no_answer",
"detail", "the primary did not answer within " + timeout + "ms"));
}
}
/**
* The worker's structured reply ({@code bridge_reply}) — resolves the blocking send awaiting
* on this session, or queues the reply in the inbox when no send is open (CB-307).
*/
private void replyMessage(Context ctx) {
String id = ctx.pathParam("id");
// The rule that matters: a worker may reply only as itself. Over MCP this was already true
// structurally (identity comes from the connection, never an argument); over REST the path
// id was simply trusted, so this is where the invariant actually gets enforced.
if (!allow(ctx, Authz.Action.REPLY, id)) {
return;
}
String content;
try {
content = mapper.readTree(ctx.body()).path("content").asText("");
} catch (Exception e) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
messages.reply(id, content);
ctx.status(200).json(Map.of("sessionId", id, "delivered", true));
}
/**
* Drain the reply inbox for a worker session — peek + ack any replies that arrived when no send
* was open. At-least-once: draining removes them from the inbox so a subsequent read returns
* nothing; an in-flight failure between the drain and the caller's processing re-surfaces them.
*/
private void drainReplies(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.DRAIN, id)) {
return;
}
var replies = messages.drainReplies(id);
ctx.status(200).json(Map.of("sessionId", id, "replies",
replies.stream().map(m -> Map.of(
"msgId", m.msgId(),
"content", m.content())).toList()));
}
/**
* Live lifecycle status of a worker (MCP `bridge_status` wraps this in CB-105), plus its
* <em>readiness</em> (CB-113): {@code ready} is true once the worker's Claude has connected the
* bridge MCP — the reliable "available to receive a task" signal, unlike bare {@code idle}, which
* is also true during boot.
*/
private void sessionStatus(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, Authz.Action.READ, id)) {
return;
}
try {
ctx.status(200).json(Map.of(
"sessionId", id,
"status", messages.status(id).name().toLowerCase(),
"ready", presence.isPresent(id)));
} catch (HerdrException e) {
herdrError(ctx, e);
}
}
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
private void taskStatus(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
}
Map<String, Object> body = new LinkedHashMap<>();
body.put("ticket", v.ticket());
body.put("phase", v.phase().name().toLowerCase());
if (v.reply() != null) {
body.put("reply", v.reply());
body.put("replySource", v.replySource());
}
if (v.detail() != null) {
body.put("detail", v.detail());
}
ctx.status(200).json(body);
}
/** Map a herdr failure: unknown target → 404, anything else → 502 (herdr is upstream). */
private static void herdrError(Context ctx, HerdrException e) {
if (e.code() != null && e.code().endsWith("_not_found")) {
ctx.status(404).json(Map.of("error", "session_not_found", "detail", e.getMessage()));
} else {
ctx.status(502).json(Map.of("error", "herdr_error", "detail", e.getMessage()));
}
}
/** Stable JSON projection of an agent (null-safe for the start-time shape). */
private static Map<String, Object> view(Agent a) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", a.terminalId());
m.put("paneId", a.paneId());
m.put("workspaceId", a.workspaceId());
m.put("tabId", a.tabId());
m.put("sessionId", a.sessionId());
m.put("agentType", a.agentType());
m.put("status", a.status().name().toLowerCase());
return m;
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(WorkerSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
m.put("profile", s.profile());
m.put("cwd", s.cwd());
m.put("ownerTerminal", s.ownerTerminal());
m.put("state", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
}
if (s.branch() != null) {
m.put("branch", s.branch());
}
return m;
}
}
@@ -1,217 +0,0 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.UncheckedIOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.List;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
/**
* Production {@link Worktrees} implementation that shells {@code git} via {@link ProcessBuilder}.
* Non-zero exits become {@link WorktreeException}. Worktree directories live under a configurable
* root (default: a sibling {@code .bridged-worktrees} of the repo root) so they are never nested
* inside the primary working tree.
*/
public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
/** Project-level MCP config. Present in the repo, so every worktree checks the primary's out. */
private static final String MCP_CONFIG = ".mcp.json";
/** What {@link #isolateToolSurface} writes: a valid, explicitly empty server map. */
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null);
}
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
public GitWorktrees(String configuredRoot) {
this.configuredRoot = configuredRoot;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
String base = (baseRef == null || baseRef.isBlank()) ? "HEAD" : baseRef;
String nonce = nonce();
Path root = resolveRoot(repoRoot);
Path path = root.resolve(nonce);
try {
Files.createDirectories(root);
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
isolateToolSurface(wt);
return wt;
}
/**
* Neutralize the worktree's project MCP config so a worker inherits only the tools its launcher
* mounts (the bridge, via {@code --mcp-config}) — never the primary's.
*
* <p>This is unconditional, and it is not the same job as the parity overlay. The repo's own
* committed {@code .mcp.json} declares the primary's IDE servers, so a fresh checkout mounts them
* whether or not the overlay copies anything; a worker that inherits them navigates and edits
* through tools bound to the <em>primary's</em> IntelliJ project, which silently hands it absolute
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
* that did not contain its changes.
*
* <p>Writing an empty server map (rather than deleting the file) keeps a project-level
* {@code .mcp.json} present and explicit, and the {@code --skip-worktree} bit keeps the
* neutralized copy from ever showing up as a local modification the worker might commit.
*/
private void isolateToolSurface(String worktreePath) {
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
Path mcp = root.resolve(MCP_CONFIG);
try {
Files.writeString(mcp, NEUTRAL_MCP_CONFIG);
} catch (IOException e) {
throw new WorktreeException("cannot neutralize " + MCP_CONFIG + " in the worktree: "
+ e.getMessage(), e);
}
if (isTracked(root, MCP_CONFIG)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", MCP_CONFIG);
}
log.debug("neutralized {} — worker tool surface is launcher-mounted only", MCP_CONFIG);
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing to remove", worktreePath);
return;
}
log.info("removing worktree {}", worktreePath);
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
return;
}
Path srcRoot = Path.of(repoRoot).toAbsolutePath().normalize();
Path dstRoot = Path.of(worktreePath).toAbsolutePath().normalize();
for (String rel : overlay) {
Path src = srcRoot.resolve(rel).normalize();
if (!Files.exists(src)) {
log.debug("parity overlay source missing — skipping {}", rel);
continue;
}
Path dst = dstRoot.resolve(rel).normalize();
try {
Files.createDirectories(dst.getParent());
Files.copy(src, dst, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.COPY_ATTRIBUTES);
log.debug("copied parity overlay {}", rel);
} catch (IOException e) {
throw new WorktreeException("cannot copy overlay " + rel + ": " + e.getMessage(), e);
}
if (isTracked(dstRoot, rel)) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", rel);
log.debug("marked overlay --skip-worktree {}", rel);
}
}
}
@Override
public String repoRoot(String cwd) {
String out = exec("git", "-C", cwd, "rev-parse", "--show-toplevel");
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
return Path.of(configuredRoot).toAbsolutePath().normalize();
}
Path repo = Path.of(repoRoot).toAbsolutePath().normalize();
return repo.resolveSibling(".bridged-worktrees");
}
private String nonce() {
return String.format("%06x", random.nextInt(1 << 24)) + "-" + seq.incrementAndGet();
}
private boolean isTracked(Path worktreeRoot, String rel) {
return exitCode("git", "-C", worktreeRoot.toString(), "ls-files", "--error-unmatch", rel) == 0;
}
/**
* Run a command and return its stdout. Non-zero exit → {@link WorktreeException} with both
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try (BufferedReader r = new BufferedReader(new InputStreamReader(p.getInputStream(), StandardCharsets.UTF_8))) {
out = r.lines().collect(Collectors.joining("\n"));
} catch (IOException e) {
p.destroyForcibly();
throw new UncheckedIOException(e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command) + "\n" + out);
}
code = p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
if (code != 0) {
throw new WorktreeException("exit " + code + " for: " + String.join(" ", command)
+ (out.isBlank() ? "" : "\n" + out));
}
return out;
}
private int exitCode(String... command) {
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
try {
if (!p.waitFor(30, TimeUnit.SECONDS)) {
p.destroyForcibly();
throw new WorktreeException("command timed out: " + String.join(" ", command));
}
return p.exitValue();
} catch (InterruptedException e) {
p.destroyForcibly();
Thread.currentThread().interrupt();
throw new WorktreeException("interrupted waiting for command: " + String.join(" ", command), e);
}
}
}
@@ -1,537 +0,0 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code bridged} process spawned.
* Delegates spawn/teardown to a {@link PeerLauncher} (which performs subscription-guarded env
* setup and process/materialization) and adds lifecycle tracking, ownership, and deterministic
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
public final class SessionManager implements TurnListener {
private static final Logger log = LoggerFactory.getLogger(SessionManager.class);
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
private final boolean clearAfterTurn;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0, false);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0, false);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0, false);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
this(launcher, worktrees, nowNanos, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap, boolean clearAfterTurn) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
this.clearAfterTurn = clearAfterTurn;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
null,
null);
registry.put(handle.id(), session);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
*/
public void onAcquire(Consumer<String> listener) {
if (listener != null) {
acquireListeners.add(listener);
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
if (listener != null) {
releaseListeners.add(listener);
}
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : acquireListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("acquire listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : releaseListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
} catch (RuntimeException e) {
if (path != null) {
try {
worktrees.remove(repoRoot, path);
} catch (RuntimeException cleanup) {
log.warn("failed to clean up worktree {} after spawn error: {}", path, cleanup.getMessage());
}
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
path,
branch);
registry.put(handle.id(), session);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
private String nonce() {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
/**
* The profile to record for a session. A launcher that performed dynamic selection tells us
* the actual profile via {@link PeerHandle#profile()}; otherwise fall back to what the caller
* requested (or the launcher's default for a no-profile spawn).
*/
private String resolveProfile(PeerHandle handle, String requestedProfile) {
String fromHandle = handle.profile();
if (fromHandle != null && !fromHandle.isBlank()) {
return fromHandle;
}
if (requestedProfile != null && !requestedProfile.isBlank()) {
return requestedProfile;
}
return launcher.defaultProfile();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
return List.copyOf(registry.values());
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
m.put("profile", session.profile());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
}
if (session.branch() != null) {
m.put("branch", session.branch());
}
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
}
}
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
completeTurn(target, false);
}
@Override
public boolean hasPostTurnAction(String target) {
if (!clearAfterTurn) return false;
WorkerSession current = findByTerminal(target);
return current != null && current.state() == WorkerSession.State.BUSY
&& (contextCap <= 0 || current.turnCount() < contextCap);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
return completeTurn(target, true);
}
private boolean completeTurn(String target, boolean startContextReset) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return false;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
return false;
}
if (!startContextReset || !clearAfterTurn) return false;
try {
return launcher.clearContext(current.paneId());
} catch (RuntimeException e) {
log.warn("context reset failed for terminal={} pane={}; continuing without reset: {}",
target, current.paneId(), e.getMessage());
return false;
}
}
/** Lifecycle hook: the worker's delegated turn failed. */
@Override
public void onTurnFailed(String target) {
onFailed(target);
}
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
}
}
/**
* Best-effort reap of sessions that have been idle longer than {@code idleTtlNanos}. Only
* {@code READY} and {@code DONE} sessions are eligible — never a {@code SPAWNING} or
* {@code BUSY} worker. Returns the number of sessions released.
*/
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
release(s.paneId());
reaped++;
}
}
return reaped;
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
break;
}
try {
long remaining = deadline - System.nanoTime();
Thread.sleep(Math.min(TimeUnit.NANOSECONDS.toMillis(remaining), 50));
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
break;
}
}
}
release(s.paneId());
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
}
}
/**
* Close this manager by draining all sessions. The timeout comes from configuration when set,
* otherwise a sensible default.
*/
public void close(Integer drainTimeoutSeconds) {
int seconds = (drainTimeoutSeconds != null && drainTimeoutSeconds > 0) ? drainTimeoutSeconds : 5;
drainAll(TimeUnit.SECONDS.toNanos(seconds));
}
/** Number of sessions currently registered. */
public int size() {
return registry.size();
}
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
* a no-op, and {@link dev.ltms.bridged.inject.WorkerPresence#markPresent} honours it — but
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
*/
private WorkerSession findByTerminal(String terminalId) {
if (terminalId == null) return null;
for (WorkerSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
log.debug("session transitioned terminal={} pane={} {} -> {}",
terminalId, current.paneId(), from, to);
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
this.sessions = sessions;
}
@Override
public void markPresent(String terminal) {
if (terminal == null || terminal.isBlank()) {
return; // the primary's contact carries no worker terminal — not a readiness signal
}
super.markPresent(terminal);
sessions.onReady(terminal);
}
}
}
@@ -1,70 +0,0 @@
package dev.ltms.bridged.session;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.concurrent.TimeUnit;
/**
* Periodic virtual-thread reaper that tears down {@code READY}/{@code DONE} sessions which have
* exceeded their idle TTL. Modeled on {@link dev.ltms.bridged.inject.StatusPoller}: a single
* virtual-thread loop, idempotent start/stop, and no {@code ScheduledExecutorService}.
*/
public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
this(sessions, idleTtlSeconds, DEFAULT_INTERVAL_MILLIS);
}
/** Construct a reaper with an explicit polling interval (useful for tests). */
public SessionReaper(SessionManager sessions, long idleTtlSeconds, long intervalMillis) {
this.sessions = sessions;
this.idleTtlNanos = TimeUnit.SECONDS.toNanos(idleTtlSeconds);
this.intervalMillis = intervalMillis;
}
/** Start the reaper loop on a virtual thread. Idempotent. */
public synchronized void start() {
if (running) return;
running = true;
thread = Thread.ofVirtual().name("session-reaper").start(this::loop);
log.info("session reaper started (idle ttl {}s, interval {}ms)",
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos), intervalMillis);
}
private void loop() {
while (running) {
try {
sessions.reapIdle(idleTtlNanos);
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
sleep();
}
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
running = false;
}
}
/** Stop the reaper loop. Idempotent. */
public synchronized void stop() {
running = false;
if (thread != null) thread.interrupt();
}
}
@@ -1,61 +0,0 @@
package dev.ltms.bridged.session;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId the host-unique opaque id (CB-519) — the registry key and the argument to
* teardown. Despite the historical name this is the {@link
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
* launcher-private herdr pane coordinate.
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
String paneId,
String terminalId,
String profile,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -1,18 +0,0 @@
package dev.ltms.bridged.session;
import java.util.List;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
}
@@ -1,318 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>A legacy spawn with no session identity is a fresh, launcher-derived session — delegate to
* the session-aware form with no name and no resume id.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
return buildLaunch(cfg, null, null);
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags. When the request carries session identity (CB-547a) it is
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg, String sessionName, String resumeSessionId) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
// requirement is skipped FOR IT ONLY. Every other profile keeps the hard boundary below.
boolean onSubscription = cfg.isSubscription();
String baseUrl = cfg.baseUrl();
if (onSubscription) {
// NO SILENT CONTRADICTION: subscription:true + a baseUrl state opposite intents; refuse
// loudly rather than pick a winner.
if (baseUrl != null && !baseUrl.isBlank()) {
throw new IllegalStateException("profile '" + cfg.profile()
+ "' sets both subscription: true and a baseUrl ('" + baseUrl + "') — the two "
+ "are contradictory: a subscription profile must not point at an endpoint. "
+ "Drop baseUrl, or drop subscription: true.");
}
// Visible without anyone going looking for it: this worker bills the subscription.
log.warn("spawning profile '{}' on the Claude subscription (subscription: true) — this "
+ "worker WILL bill the operator's subscription", cfg.profile());
} else {
guard.assertWorker(baseUrl); // hard stop before we spawn anything
}
Map<String, String> workerEnv = baseEnv(cfg);
if (onSubscription) {
// CB-542 belt-and-braces: on the subscription path no guard vets these two keys, and the
// profile's env: is layered in by baseEnv — so strip any that rode in there. Config load
// already rejects this (loudly, naming the profile); this makes the boundary hold even
// for a profile built in code that never passed through that validation.
workerEnv.remove("ANTHROPIC_BASE_URL");
workerEnv.remove("ANTHROPIC_AUTH_TOKEN");
} else {
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
}
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
applyGitToken(workerEnv, cfg);
// CB-547a: Claude Code can MINT its own session id, so bridged chooses it — a fresh spawn
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
// handle is known BEFORE the agent has written anything; a resume spawn adopts its prior
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has no MCP — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg));
String agentSessionId = applySessionIdentity(argv, sessionName, resumeSessionId);
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
}
if (resuming) {
argv.add("-r");
argv.add(resumeSessionId);
return resumeSessionId;
}
String minted = UUID.randomUUID().toString();
argv.add("--session-id");
argv.add(minted);
return minted;
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
/**
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
*
* <p>The env var alone is not a reliable pin for this adapter, because the argv is usually a
* launcher rather than {@code claude} itself — {@code ["ccs", "<profile>"]} — and {@code ccs}
* exports its profile's own model family ({@code ANTHROPIC_MODEL}, {@code DEFAULT_OPUS/SONNET/
* HAIKU}, {@code CLAUDE_CODE_SUBAGENT_MODEL}) over whatever it inherited. A worker profile that
* set {@code model:} therefore got silently overruled by its own launcher. Claude Code's
* {@code --model} flag outranks the environment, and {@code ccs <profile> [claude-args...]}
* passes trailing arguments through, so the flag survives the wrapper.
*
* <p>Appended last so it also outranks anything in the operator's own {@code argv}. Profiles
* that deliberately leave {@code model:} unset (letting {@code ccs} own model selection, as
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
*/
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Worker cfg) {
if (cfg.model() == null || cfg.model().isBlank()) {
return argv;
}
List<String> withModel = mutableArgv(argv);
withModel.add("--model");
withModel.add(cfg.model());
return withModel;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.CONTEXT_RESET, Capability.ORPHAN_REAP,
Capability.SESSION_NAME, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
@Override
public boolean clearContext(String id) {
String target = agentTarget(id);
if (target == null) {
return false;
}
// This deliberately bypasses Injector: /clear is housekeeping, not a delegated turn.
agents().send(target, "/clear");
return true;
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,275 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Map<String, BridgedConfig.Worker> profileConfigs;
private final PlacementPolicy placementPolicy;
private final Function<String, Integer> liveCount;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), name -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Worker> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this.profileConfigs = Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs));
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the policy entirely.
HerdrPeerLauncher d = route(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
List<PlacementCandidate> candidates = candidates();
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
PlacementCandidate chosen;
try {
chosen = placementPolicy.select(ctx);
} catch (RuntimeException e) {
// No candidate left (all at cap or all unreachable). The policy already threw a clear
// message; do not wrap it in a generic PeerUnreachableException.
throw e;
}
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/** Build the candidate list from the configured profiles, in definition order. */
private List<PlacementCandidate> candidates() {
List<PlacementCandidate> out = new ArrayList<>();
for (Map.Entry<String, BridgedConfig.Worker> e : profileConfigs.entrySet()) {
BridgedConfig.Worker w = e.getValue();
out.add(new PlacementCandidate(e.getKey(), null, w.weight(), w.maxLoad()));
}
return out;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public boolean clearContext(String id) {
HerdrPeerLauncher delegate = spawnedBy.get(id);
if (delegate == null) {
log.debug("clearContext({}) ignored — no recorded owning adapter", id);
return false;
}
return delegate.clearContext(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -1,670 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
// CB-519: PeerHandle.id() is a host-unique opaque UUID, decoupled from the herdr pane id. The
// routing/registry key is the UUID; the herdr pane id is a launcher-private placement/teardown
// coordinate. This map bridges the two so stop(id) can resolve a host-unique key back to the
// exact pane it must tear down. The pane id is launcher-private (never the routing key) — see
// PeerHandle.id().
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
private final AtomicBoolean resetUnsupportedLogged = new AtomicBoolean();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/**
* Session-aware variant of {@link #buildLaunch(BridgedConfig.Worker)} (CB-547a). Default
* discards the session identity and delegates to the profile-only form, so an adapter that
* carries no durable peer session (opencode, say) inherits byte-identical behaviour and needs
* no change. An adapter that does (Claude Code) overrides this to mint/resume the id and to
* surface it on the returned {@link Launch#agentSessionId()}.
*
* @param cfg the resolved profile to spawn
* @param sessionName the bridge's logical session name, or null/blank for launcher-derived
* @param resumeSessionId the peer's own prior session id to resume, or null/blank for fresh
*/
protected Launch buildLaunch(BridgedConfig.Worker cfg, String sessionName, String resumeSessionId) {
return buildLaunch(cfg);
}
/** Direct transport access for peer-specific, non-turn control operations. */
protected final AgentControl agents() {
return agents;
}
/** Resolve the public peer id to the launcher's private herdr target. */
protected final String agentTarget(String id) {
return paneByAgentId.get(id);
}
@Override
public boolean clearContext(String id) {
if (resetUnsupportedLogged.compareAndSet(false, true)) {
log.warn("context reset is unsupported for peer kind {}; clearAfterTurn is a no-op",
namePrefix);
}
return false;
}
/**
* A peer-specific launch: the herdr {@code env} map and {@code argv}, plus — for an adapter
* that carries durable session identity (CB-547a) — the peer's OWN session id
* ({@link PeerHandle#agentSessionId()}), known before the peer has written anything. Null for
* a launch that carries no identity.
*/
protected record Launch(Map<String, String> env, List<String> argv, String agentSessionId) {
/** A launch without a discoverable agent session id (an adapter that carries none). */
Launch(Map<String, String> env, List<String> argv) {
this(env, argv, null);
}
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/** A started peer plus the launch's agent-session id (the resume handle, or null). */
private record Spawned(Agent agent, String agentSessionId) {
}
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd, null, null).agent();
}
/**
* Spawn a peer with session identity (CB-547a). {@code sessionName} and {@code resumeSessionId}
* are threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Worker,
* String, String)}, and the launch's resolved agent-session id is returned alongside the agent
* so the caller can put it on the {@link PeerHandle}.
*/
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg, sessionName, resumeSessionId);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
return new Spawned(agent, launch.agentSessionId());
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} is a fresh <em>host-unique</em> opaque
* UUID (CB-519), deliberately decoupled from the herdr pane id: the id is the registry/routing
* key and must never collide across daemon processes on the same host, while the herdr pane id
* stays a launcher-private placement/teardown coordinate, remembered here so {@link #stop}
* can resolve the host-unique key back to its pane. When {@code spawnReadyTimeoutMs > 0},
* blocks until the peer's herdr status is injectable or the timeout elapses; on timeout the
* pane is closed (no orphan) and a {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Spawned spawned = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId());
Agent agent = spawned.agent();
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
// CB-519: the handle id is a host-unique UUID; the herdr pane it maps to stays internal.
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile(),
req.sessionName(), spawned.agentSessionId());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
if (tab.rootPaneId() == null) {
// Protocol 19 starts the agent INTO the seed pane — without one there is nowhere
// to start, and a partial tab would be left behind.
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " had no seed pane in the create response — cannot start a peer in it");
}
started = startUniquelyNamed(cfg, argv, tab.rootPaneId());
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now, in the seed pane itself (no shell pane to drop — protocol 19).
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
// we log and still return it so the caller gets its paneId and can tear it down.
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
}
Agent peer = startUniquelyNamed(cfg, argv, paneId).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down: close the pane, and close its tab <em>only</em> when the peer is that tab's
* sole occupant. The single-pane check is what makes this safe regardless of how the peer was
* placed (or a placement-config change across a restart): a pane-placement peer sitting in one
* of the user's shared tabs has siblings, so its tab is never closed — we only ever remove a
* tab we created to hold one peer.
*
* <p>{@code idOrPane} is the {@link PeerHandle#id()} of a peer this launcher spawned (CB-519's
* host-unique opaque UUID), resolved through {@link #paneByAgentId} to the pane it must tear
* down. An argument that is not one of our ids is treated as a raw herdr pane id — the
* {@link #reapOrphanWorkers() orphan-reap} and spawn-gate-timeout paths, plus any caller that
* passes a pane directly, keep working without an owning id.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* and the session identity the launch resolved (CB-547a): the bridge's logical name and the
* peer's own session id, both null when the spawn carried no identity.
*/
private record WorkerHandle(String id, String terminalId, String profile,
String sessionName, String agentSessionId) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
@@ -1,331 +0,0 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
* provider-agnostic terminal coding agent. Its whole reason for existing is to prove the
* {@code PeerLauncher} SPI is genuinely provider-neutral: opencode shares none of Claude Code's
* private launch seams, yet reuses every line of shared transport in the base (tab/pane placement,
* the CB-306 readiness gate, unique naming + CB-117 reap, teardown, listing, cwd).
*
* <p>The divergences from {@link ClaudeCodeLauncher}, all confined to {@link #buildLaunch}:
* <ul>
* <li><strong>No subscription boundary.</strong> opencode carries no {@code ANTHROPIC_BASE_URL}
* and there is no {@link dev.ltms.bridged.guard.SubscriptionGuard} — the guard is a
* Claude-private concern, not part of the SPI. opencode reads the operator's own provider
* credentials from its global {@code auth.json}; the bridge injects none.</li>
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
* <li><strong>{@code opencode} name prefix</strong> so reap matches {@code opencode-*} panes and
* never another adapter's.</li>
* </ul>
*/
public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "opencode";
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
}
/**
* Production constructor with the spawn-ready gate enabled. Polls {@code agents.status()} until
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
* they can inspect the generated {@code opencode.json}/charter under.
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.configRoot = configRoot;
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(argvWithAuto(cfg), cfg));
}
/**
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
* whatever provider opencode resolves by default.
*
* <p>Note this reuses {@code baseUrl}, the same field the Claude adapter injects as
* {@code ANTHROPIC_BASE_URL} — but it does <em>not</em> go through {@code SubscriptionGuard}.
* That asymmetry is deliberate and safe: the guard exists to stop a worker borrowing the
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/**
* The launch argv plus the unconditional {@code --auto} flag, which auto-approves the
* permissions opencode does not explicitly deny. It is unconditional, not a preference: a
* spawned peer has no human at its pane — the bridge spawned it — so one that stops at an
* approval prompt is a wedged agent, indistinguishable from a legitimate mid-turn wait and
* unable to end its turn with {@code bridge_reply}. opencode's own help calls this
* "dangerous!", but the blast radius here is already bounded by design: a worker runs in its
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
* gate.
*/
private List<String> argvWithAuto(BridgedConfig.Worker cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--auto");
return argv;
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(List<String> argv, BridgedConfig.Worker cfg) {
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
}
return argv;
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
charter.toFile().deleteOnExit();
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
}
Path cfgFile = dir.resolve("opencode.json");
// Built with Jackson rather than string concatenation: the provider block is nested and
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
"cannot write opencode config for profile " + cfg.profile(), e);
}
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
*
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
ObjectNode provider = root.putObject("provider").putObject(providerId);
provider.put("npm", "@ai-sdk/openai-compatible");
provider.put("name", providerId + " (bridged)");
ObjectNode options = provider.putObject("options");
options.put("baseURL", openAiBaseUrl(cfg.baseUrl()));
// vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one.
String token = resolveEnv(cfg.tokenEnv());
options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token);
provider.putObject("models").putObject(modelId).put("name", modelId);
}
/**
* Split {@code model:} into its {@code provider} and {@code model} halves. A pinned endpoint
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
throw new IllegalArgumentException(
"profile " + cfg.profile() + " sets baseUrl (a pinned opencode endpoint) so"
+ " model: must be \"<provider>/<model>\", e.g."
+ " \"local-vllm/deepseek-v4-flash\"; got "
+ (model == null ? "null" : '"' + model + '"'));
}
return new String[]{model.substring(0, slash), model.substring(slash + 1)};
}
/**
* The OpenAI-compatible base URL for {@code baseUrl}. A bare {@code host:port} gets {@code /v1}
* appended (where these servers put the API); a URL that already carries a path is taken as-is,
* so an endpoint mounted somewhere unusual is still reachable.
*/
private static String openAiBaseUrl(String baseUrl) {
String trimmed = baseUrl.trim();
while (trimmed.endsWith("/")) {
trimmed = trimmed.substring(0, trimmed.length() - 1);
}
int schemeEnd = trimmed.indexOf("://");
String afterScheme = schemeEnd < 0 ? trimmed : trimmed.substring(schemeEnd + 3);
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
/**
* Whether {@code name} is an opencode bridge worker started by a <em>different</em> process than
* {@code currentNonce}. A thin {@code opencode}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,67 +0,0 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import java.util.HashMap;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-548 — the architect-slot registry: the config snapshot of slot → profile, and the live
* terminal → slot binding the resolver reads. The role the binding produces is asserted in
* {@link CallerResolverTest}; this pins the registry object itself.
*/
class ArchitectRegistryTest {
private static final Map<String, BridgedConfig.Architect> SLOTS = Map.of(
"lead-designer", new BridgedConfig.Architect("term_design", "sonnet"),
"reviewer", new BridgedConfig.Architect(null, "gx10"));
private final ArchitectRegistry registry =
new ArchitectRegistry(SLOTS, () -> Map.of("term_design", "lead-designer"));
@Test
void exposesTheConfiguredSlots() {
assertEquals(SLOTS.keySet(), registry.slots().keySet());
assertTrue(registry.isSlot("reviewer"));
assertFalse(registry.isSlot("nope"));
}
@Test
void theSpawnLifecycleReadsTheProfileBackFromASlot() {
assertEquals("sonnet", registry.profileForSlot("lead-designer"));
assertEquals("gx10", registry.profileForSlot("reviewer"));
assertNull(registry.profileForSlot("unknown"), "an unknown slot has no profile");
}
@Test
void resolvesTheSlotOfALiveTerminal() {
assertEquals("lead-designer", registry.slotForTerminal("term_design"));
assertNull(registry.slotForTerminal("term_unbound"));
assertNull(registry.slotForTerminal(null), "no terminal ⇒ no slot");
}
@Test
void theBindingIsLiveReReadPerCall() {
Map<String, String> live = new HashMap<>();
ArchitectRegistry r = new ArchitectRegistry(SLOTS, () -> live);
assertNull(r.slotForTerminal("term_design"));
live.put("term_design", "lead-designer"); // injected after construction
assertEquals("lead-designer", r.slotForTerminal("term_design"));
}
@Test
void theSlotSnapshotIsFixedByConstruction() {
Map<String, BridgedConfig.Architect> mutable = new HashMap<>(SLOTS);
ArchitectRegistry r = new ArchitectRegistry(mutable, Map::of);
mutable.put("hijack", new BridgedConfig.Architect("t", "gx10"));
assertFalse(r.isSlot("hijack"), "a handed-over map is not offered as live state");
}
}
@@ -1,899 +0,0 @@
package dev.ltms.bridged.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
class BridgedConfigTest {
@Test
void loadsFullConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8080
herdrSocket: ~/.config/herdr/herdr.sock
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000
model: coder
guard:
offSubscriptionHosts:
- gx00.gw
- ollama.ltms.dev
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(8080, cfg.bind().port());
assertEquals("ltms-local", cfg.worker().profile());
assertTrue(cfg.guard().hostSet().contains("gx00.gw"));
assertTrue(cfg.guard().hostSet().contains("ollama.ltms.dev"));
}
@Test
void appliesDefaultsForMissingSections(@TempDir Path dir) throws Exception {
Path f = dir.resolve("minimal.yaml");
Files.writeString(f, "bind:\n host: 0.0.0.0\n port: 9000\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(9000, cfg.bind().port());
assertNotNull(cfg.guard(), "guard must default to empty, never null");
assertTrue(cfg.guard().offSubscriptionHosts().isEmpty());
assertFalse(cfg.lifecycle().clearAfterTurn(), "context clearing is opt-in");
}
@Test
void singleWorkerBecomesAOneEntryProfileMapWithItselfAsDefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("single.yaml");
Files.writeString(f, """
worker:
profile: ltms-local
baseUrl: http://gx00.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("ltms-local"), cfg.workerProfiles().keySet(), "legacy worker → one profile");
assertEquals("ltms-local", cfg.defaultProfile());
}
@Test
void loadsMultipleWorkerProfilesWithADefault(@TempDir Path dir) throws Exception {
Path f = dir.resolve("multi.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
ollama:
baseUrl: http://ollama.ltms.dev
argv: ["ccs", "ollama"]
defaultWorker: gx10
guard:
offSubscriptionHosts: [gx10.gw, ollama.ltms.dev]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("gx10", "ollama"), cfg.workerProfiles().keySet());
// Order, not just membership: placement breaks an exact-weight tie on definition order, so a
// hash-ordered map here would make equal-weight placement differ from one restart to the next.
assertEquals(java.util.List.of("gx10", "ollama"),
java.util.List.copyOf(cfg.workerProfiles().keySet()),
"workerProfiles must preserve YAML definition order");
assertEquals("gx10", cfg.defaultProfile());
assertEquals("ollama", cfg.workerProfiles().get("ollama").profile(), "profile defaults to its map key");
assertEquals("http://gx10.gw:8000", cfg.workerProfiles().get("gx10").baseUrl());
}
@Test
void ignoresUnknownKeys(@TempDir Path dir) throws Exception {
Path f = dir.resolve("future.yaml");
Files.writeString(f, "bind:\n port: 8080\nfutureFeature:\n enabled: true\n");
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
/**
* CB-530. Unknown keys stay ignored — config must be allowed to run ahead of the code — but they
* must be NAMED at load. A whole block that parses, is dropped, and is never mentioned again is
* indistinguishable from one that works: that is exactly how a hand-written `leaders:` registry
* came to look configured while being inert.
*/
@Test
void unknownTopLevelKeysAreNamedSoADroppedBlockCannotLookLikeAWorkingOne() {
assertEquals(List.of("futureFeature", "leedars"),
BridgedConfig.unknownTopLevelKeys(
"bind:\n port: 8080\nleedars:\n a: b\nfutureFeature: true\n"),
"a typo'd key is the common case and must be reported by name");
}
@Test
void everyKeyThisBuildUnderstandsIsAbsentFromTheUnknownList() {
assertTrue(BridgedConfig.unknownTopLevelKeys("""
bind:
port: 8080
herdrSocket: /tmp/s
workers: {}
defaultWorker: a
guard: {}
worktreeRoot: /tmp
lifecycle: {}
spawnReadyTimeoutMs: 1
spawnReadyPollMs: 1
broker: {}
primary: {}
leaders: {}
architects: {}
leadScan: {}
placement: fixed
auth: {}
""").isEmpty(), "the known-key set must not drift from the record components");
}
@Test
void aMalformedOrEmptyDocumentIsNotReportedAsUnknownKeys() {
assertTrue(BridgedConfig.unknownTopLevelKeys("").isEmpty());
assertTrue(BridgedConfig.unknownTopLevelKeys("just a scalar").isEmpty());
}
// ── CB-531: lead discovery by tab label ─────────────────────────────────────────────────────
@Test
void leadScanIsOffUnlessTheBlockIsPresent(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-scan.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertNull(BridgedConfig.load(f).leadScan(),
"turning this on widens who resolves as PRIMARY — upgrading the daemon must not do that");
}
@Test
void leadScanDefaultsItsFieldsWhenTheBlockIsPresentButBare(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bare-scan.yaml");
Files.writeString(f, "bind:\n port: 8080\nleadScan: {}\n");
BridgedConfig.LeadScan scan = BridgedConfig.load(f).leadScan();
assertEquals("lead:", scan.tabPrefix());
assertEquals(10, scan.intervalSeconds());
}
@Test
void leadScanReadsAnExplicitPrefixAndInterval(@TempDir Path dir) throws Exception {
Path f = dir.resolve("scan.yaml");
Files.writeString(f, """
bind:
port: 8080
leadScan:
tabPrefix: "drive:"
intervalSeconds: 30
""");
BridgedConfig.LeadScan scan = BridgedConfig.load(f).leadScan();
assertEquals("drive:", scan.tabPrefix());
assertEquals(30, scan.intervalSeconds());
}
/**
* The hazard the guard exists for: bridged writes worker tab labels and reads lead tab labels.
* Overlap the two and every worker it spawns is read back as a lead.
*/
@Test
void aLeadPrefixThatAWorkerTabLabelAlsoMatchesRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collide.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
tabLabel: "lead: {profile} #{n}"
leadScan:
tabPrefix: "lead:"
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateLeadScan);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
@Test
void theDefaultWorkerTabLabelDoesNotCollideWithTheDefaultLeadPrefix(@TempDir Path dir) throws Exception {
Path f = dir.resolve("ok.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx00.gw:8000
leadScan: {}
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateLeadScan());
}
@Test
void theCollisionGuardIsANoOpWhenScanningIsOff(@TempDir Path dir) throws Exception {
Path f = dir.resolve("off.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
tabLabel: "lead: {profile}"
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateLeadScan(),
"a label that collides with a convention nobody reads is not a problem");
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
@Test
void leadersBlockRegistersEveryPaneByName(@TempDir Path dir) throws Exception {
Path f = dir.resolve("leaders.yaml");
Files.writeString(f, """
bind:
port: 8080
leaders:
opus-5.0:
terminal: term_opus
kind: claude
gpt-sol-5.6:
terminal: term_sol
kind: opencode
model: openai/gpt-5.6-terra
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("opus-5.0", "gpt-sol-5.6"), cfg.leaders().keySet());
assertEquals("opencode", cfg.leaders().get("gpt-sol-5.6").kind());
assertEquals("openai/gpt-5.6-terra", cfg.leaders().get("gpt-sol-5.6").model());
// The whole point: BOTH panes resolve as leads, so neither is demoted to worker.
assertEquals(Map.of("term_opus", "opus-5.0", "term_sol", "gpt-sol-5.6"),
cfg.leaderTerminals());
}
@Test
void aLegacyPrimaryPinAloneStillRegistersAsALeadNamedPrimary(@TempDir Path dir) throws Exception {
Path f = dir.resolve("legacy-pin.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: term_fixed\n");
assertEquals(Map.of("term_fixed", "primary"), BridgedConfig.load(f).leaderTerminals(),
"configs that never migrate must behave exactly as they did before CB-530");
}
@Test
void anExplicitLeadersEntryWinsOverThePinForTheSameTerminal(@TempDir Path dir) throws Exception {
Path f = dir.resolve("both.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_shared
leaders:
opus-5.0:
terminal: term_shared
""");
assertEquals(Map.of("term_shared", "opus-5.0"), BridgedConfig.load(f).leaderTerminals(),
"the pin is the older spelling of the same fact; the named entry is what was meant");
}
@Test
void bothBlocksTogetherRegisterTheUnionOfTheirTerminals(@TempDir Path dir) throws Exception {
Path f = dir.resolve("union.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_pinned
leaders:
gpt-sol-5.6:
terminal: term_sol
""");
assertEquals(Map.of("term_pinned", "primary", "term_sol", "gpt-sol-5.6"),
BridgedConfig.load(f).leaderTerminals());
}
@Test
void neitherBlockLeavesNothingRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("none.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertTrue(BridgedConfig.load(f).leaderTerminals().isEmpty());
}
/** A lead entry with no terminal identifies nothing — it must not register a null key. */
@Test
void aLeadWithoutATerminalIsNotRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-terminal.yaml");
Files.writeString(f, """
bind:
port: 8080
leaders:
sketch:
kind: opencode
real:
terminal: term_real
""");
assertEquals(Map.of("term_real", "real"), BridgedConfig.load(f).leaderTerminals());
}
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
@Test
void architectsBlockBindsSlotsByGatewayLocalName(@TempDir Path dir) throws Exception {
Path f = dir.resolve("architects.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: sonnet
reviewer:
profile: sonnet
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(Set.of("lead-designer", "reviewer"), cfg.architects().keySet(),
"slot names are the keys — gateway-local unique by construction");
assertEquals("sonnet", cfg.architects().get("lead-designer").profile(),
"each slot carries its strong-model profile reference");
assertEquals("term_design", cfg.architects().get("lead-designer").terminal());
// A slot with no terminal binds nothing yet — the live binding may supply it later.
assertTrue(cfg.architects().get("reviewer").terminal() == null
|| cfg.architects().get("reviewer").terminal().isBlank());
}
@Test
void architectTerminalsMapsEachBoundSlotByItsPane(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-terminals.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: sonnet
reviewer:
terminal: term_review
profile: sonnet
unbound:
profile: sonnet
""");
assertEquals(Map.of("term_design", "lead-designer", "term_review", "reviewer"),
BridgedConfig.load(f).architectTerminals(),
"a slot with no terminal registers no binding; the value is the slot name");
}
@Test
void noArchitectsBlockLeavesNothingBound(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-arch.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.architects());
assertTrue(cfg.architectTerminals().isEmpty(),
"no architects: block ⇒ no architect identity, exactly as before CB-548");
}
@Test
void anArchitectSlotMayResolveToTheSoleProfileWithoutPrivileging(@TempDir Path dir) throws Exception {
// Even a single unqualified worker profile can back an architect slot — the reference is
// by name, not by position, so an explicit name is required.
Path f = dir.resolve("arch-single.yaml");
Files.writeString(f, """
bind:
port: 8080
worker:
profile: ltms-local
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: ltms-local
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertDoesNotThrow(cfg::validateArchitects);
assertEquals("ltms-local", cfg.architects().get("lead-designer").profile());
}
@Test
void anArchitectProfileThatIsNotConfiguredRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-bad-profile.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: sonnet
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateArchitects);
assertTrue(e.getMessage().contains("lead-designer"), "the refusal names the slot");
assertTrue(e.getMessage().contains("sonnet"), "the refusal names the offending profile");
}
@Test
void anArchitectSlotMissingAProfileRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-no-profile.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: ""
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateArchitects);
assertTrue(e.getMessage().contains("lead-designer"), "the refusal names the slot");
}
@Test
void aValidArchitectRegistryPassesValidation(@TempDir Path dir) throws Exception {
Path f = dir.resolve("arch-ok.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
sonnet:
baseUrl: http://gx10.gw:8000
gx10:
baseUrl: http://gx10.gw:8000
architects:
lead-designer:
terminal: term_design
profile: sonnet
reviewer:
profile: gx10
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateArchitects());
}
@Test
void absentArchitectsBlockPassesValidation(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-arch.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateArchitects());
}
@Test
void absentBrokerBlockLeavesInboxSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-broker.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.broker(), "no broker: block → null → in-memory inbox is selected");
}
@Test
void brokerBlockWithUriEnablesAmqpAdapter(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker.yaml");
Files.writeString(f, """
bind:
port: 8080
broker:
uri: amqp://guest:guest@127.0.0.1:5672/
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertTrue(cfg.broker().isConfigured(), "a non-blank uri enables the AMQP adapter");
assertEquals("amqp://guest:guest@127.0.0.1:5672/", cfg.broker().uri());
}
@Test
void brokerBlockWithBlankUriStaysSoftState(@TempDir Path dir) throws Exception {
Path f = dir.resolve("broker-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nbroker:\n uri: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.broker());
assertFalse(cfg.broker().isConfigured(), "an empty uri must not enable AMQP");
}
@Test
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-primary.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNull(cfg.primary(), "no primary: block → null → connection-derived identity");
}
@Test
void primaryBlockWithTerminalPinsIdentity(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-pinned.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_fixed
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertEquals("term_fixed", cfg.primary().terminal());
}
@Test
void primaryBlockWithBlankTerminalDefaultsToDerived(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-blank.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: \"\"\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.primary());
assertTrue(cfg.primary().terminal() == null || cfg.primary().terminal().isBlank(),
"a blank terminal in yaml should be treated as absent — null or empty are equivalent");
}
// --- CB-402: peer kind discriminator -------------------------------------------------------
@Test
void workerKindDefaultsToClaudeCodeWhenOmitted(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-absent.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
argv: ["ccs", "gx10"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_CLAUDE_CODE, cfg.workerProfiles().get("gx10").kind(),
"a worker with no kind: is a claude-code worker (backward compatible)");
}
@Test
void opencodeKindIsNormalizedToLowerCase(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-opencode.yaml");
Files.writeString(f, """
workers:
gemini:
kind: OpenCode
model: google/gemini-2.5-pro
argv: ["opencode"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Worker.KIND_OPENCODE, cfg.workerProfiles().get("gemini").kind(),
"kind is normalised to lower-case so YAML casing does not matter");
}
@Test
void kindPredicatesReflectTheResolvedKind(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-predicates.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker claude = cfg.workerProfiles().get("claude");
BridgedConfig.Worker gemini = cfg.workerProfiles().get("gemini");
assertTrue(claude.isClaudeCode(), "the default-kind worker is claude-code");
assertFalse(claude.isOpenCode(), "a claude-code worker is not opencode");
assertTrue(gemini.isOpenCode(), "the kind: opencode worker is opencode");
assertFalse(gemini.isClaudeCode(), "an opencode worker is not claude-code");
}
@Test
void argvDefaultsToTheKindBinaryWhenUnset(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-argv.yaml");
Files.writeString(f, """
workers:
claude:
baseUrl: http://gx10.gw:8000
gemini:
kind: opencode
model: google/gemini-2.5-pro
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(java.util.List.of("claude"), cfg.workerProfiles().get("claude").argv(),
"a claude-code worker with no argv defaults to the claude binary");
assertEquals(java.util.List.of("opencode"), cfg.workerProfiles().get("gemini").argv(),
"an opencode worker with no argv defaults to the opencode binary, never claude");
}
@Test
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-auth-block.yaml");
Files.writeString(f, "bind:\n host: 127.0.0.1\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
assertNotNull(cfg.auth(), "auth must default rather than be null");
assertFalse(cfg.auth().tokenMode());
assertEquals("BRIDGED_API_TOKEN", cfg.auth().tokenEnv(), "documented default env var");
assertDoesNotThrow(cfg::validateAuthExposure, "loopback + loopback-trust is the safe pairing");
}
/**
* CB-501's highest-value check. Under loopback-trust, "not a known worker" means "the primary" —
* sound only while the OS refuses remote connections. Widening the bind without token mode
* would silently promote every reachable client to the most privileged role on the bus.
*/
@Test
void aNonLoopbackBindWithoutTokenModeIsRefusedAtStartup(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed.yaml");
Files.writeString(f, "bind:\n host: 0.0.0.0\n port: 8765\n");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAuthExposure);
assertTrue(e.getMessage().contains("auth.mode: token"),
"the error must say how to fix it, not just that it refused");
}
@Test
void aNonLoopbackBindIsAllowedOnceTokenModeIsOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exposed-with-token.yaml");
Files.writeString(f, """
bind:
host: 0.0.0.0
port: 8765
auth:
mode: token
tokenEnv: MY_TOKEN
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertTrue(cfg.auth().tokenMode());
assertEquals("MY_TOKEN", cfg.auth().tokenEnv());
assertDoesNotThrow(cfg::validateAuthExposure);
}
@Test
void loopbackFormsAreAllRecognised(@TempDir Path dir) throws Exception {
for (String host : new String[]{"127.0.0.1", "localhost", "::1", "127.0.0.53"}) {
Path f = dir.resolve("lb-" + host.replace(':', '_') + ".yaml");
Files.writeString(f, "bind:\n host: \"" + host + "\"\n port: 8765\n");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateAuthExposure(),
host + " is loopback and must not trip the exposure guard");
}
}
/**
* The shipped {@code bridged.example.yaml} must actually parse. Config binds through a plain
* Jackson mapper with {@code ignoreUnknown = true}, so a misspelled key in the example is
* silently dropped and the operator gets a default they did not ask for — exactly how a
* {@code spawn_ready_timeout_ms} typo survived in the example until the CB-5xx wrap-up.
*/
@Test
void shippedExampleConfigParses() {
Path example = Path.of("bridged.example.yaml");
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
BridgedConfig cfg = BridgedConfig.load(example);
assertEquals(8765, cfg.bind().port(), "example binds the documented default port");
assertTrue(cfg.workerProfiles().containsKey("gx10"), "example documents the gx10 profile");
assertEquals("gx10", cfg.defaultProfile(), "example's defaultWorker resolves");
assertTrue(cfg.guard().hostSet().contains("gx01.gw"),
"every example profile's base_url host must be in the example allowlist");
}
/**
* Every optional knob the example documents must bind under the exact spelling used there.
* Keep this list in step with {@code bridged.example.yaml}: a rename that updates the record
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
*/
@Test
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
Path f = dir.resolve("all-knobs.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
spawnReadyTimeoutMs: 25000
spawnReadyPollMs: 400
worktreeRoot: /tmp/bridged-worktrees
workers:
gx10:
kind: claude-code
baseUrl: http://gx01.gw:8000
configDir: /tmp/ccs/gx10
cwd: /tmp/repo
parityOverlay: [".mcp.json", ".env"]
gitTokenEnv: GITEA_TOKEN
gitHostEnv: GITEA_HOST
weight: 0.5
maxLoad: 2
placement: weighted
lifecycle:
idleTtlSeconds: 300
contextCap: 10
drainTimeoutSeconds: 5
clearAfterTurn: true
broker:
uri: amqp://guest:guest@127.0.0.1:5672
primary:
terminal: term_abc123
pushReminders: 5
pushBackoffMs: 15000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(25000, cfg.spawnReadyTimeoutMs(), "spawnReadyTimeoutMs is camelCase, not snake_case");
assertEquals(400, cfg.spawnReadyPollMs(), "spawnReadyPollMs is camelCase, not snake_case");
assertEquals("/tmp/bridged-worktrees", cfg.worktreeRoot());
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals("/tmp/ccs/gx10", w.configDir());
assertEquals("/tmp/repo", w.cwd());
assertEquals(java.util.List.of(".mcp.json", ".env"), w.parityOverlay());
assertTrue(w.hasGitToken(), "gitTokenEnv binds and enables the CB-302 PR grant");
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(0.5f, w.weight(), 0.0001f, "weight binds as a float");
assertEquals(2, w.maxLoad(), "maxLoad binds as an integer");
assertEquals("weighted", cfg.placement(), "placement binds at the top level");
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
assertEquals(10, cfg.lifecycle().contextCap());
assertEquals(5, cfg.lifecycle().drainTimeoutSeconds());
assertTrue(cfg.lifecycle().clearAfterTurn());
assertEquals("amqp://guest:guest@127.0.0.1:5672", cfg.broker().uri());
assertEquals("term_abc123", cfg.primary().terminal());
assertEquals(5, cfg.primary().remindersOrDefault());
assertEquals(15000L, cfg.primary().backoffMsOrDefault());
}
@Test
void placementDefaultsToFixedForExistingConfigs(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-placement.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals("fixed", cfg.placement(), "omitted placement must default to fixed");
}
@Test
void workerWeightAndMaxLoadDefaultSanely(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-weight.yaml");
Files.writeString(f, """
bind:
port: 8080
workers:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Worker w = cfg.workerProfiles().get("gx10");
assertEquals(1.0f, w.weight(), 0.0001f, "absent weight defaults to 1.0");
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
}
@Test
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
Path f = dir.resolve("subscription.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
opted:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertTrue(cfg.workerProfiles().get("sonnet").isSubscription(),
"subscription: true binds as an explicit opt-in");
assertFalse(cfg.workerProfiles().get("opted").isSubscription(),
"a profile without the key stays off-subscription (the default)");
}
// ── CB-542: subscription:true must not smuggle an unguarded endpoint via env: ───────────────
@Test
void aSubscriptionProfileWithAnthropicBaseUrlInEnvIsRejected(@TempDir Path dir) throws Exception {
Path f = dir.resolve("baseUrl.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validateSubscriptionProfiles);
assertTrue(e.getMessage().contains("sonnet"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("ANTHROPIC_BASE_URL"),
"the refusal names the offending key: " + e.getMessage());
}
@Test
void aSubscriptionProfileWithAnthropicAuthTokenInEnvIsRejected(@TempDir Path dir) throws Exception {
Path f = dir.resolve("authToken.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_AUTH_TOKEN: sk-ant-not-on-any-allowlist
""");
BridgedConfig cfg = BridgedConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validateSubscriptionProfiles);
assertTrue(e.getMessage().contains("sonnet"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("ANTHROPIC_AUTH_TOKEN"),
"the refusal names the offending key: " + e.getMessage());
}
@Test
void aSubscriptionProfileWithACleanEnvPassesValidation(@TempDir Path dir) throws Exception {
Path f = dir.resolve("clean.yaml");
Files.writeString(f, """
workers:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
JAVA_HOME: /opt/jdk
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertDoesNotThrow(cfg::validateSubscriptionProfiles,
"a subscription profile may carry env: — just not the Anthropic binding keys");
}
@Test
void aNonSubscriptionProfileMayCarryAnthropicEnvKeys(@TempDir Path dir) throws Exception {
// The override is only dangerous on the subscription path, where no guard could vet it. A
// plain profile's env: is still overwritten by the launcher's guard-checked value (CB-511).
Path f = dir.resolve("nonsub.yaml");
Files.writeString(f, """
workers:
gx10:
baseUrl: http://gx10.gw:8000
env:
ANTHROPIC_BASE_URL: http://something
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertDoesNotThrow(cfg::validateSubscriptionProfiles,
"only subscription:true profiles are checked — off-subscription ones keep the baseUrl guard");
}
}
@@ -1,56 +0,0 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the worker env seam against a REAL herdr. Under protocol 19 (CB-521) the
* env map is injected at PANE CREATION ({@code tab.create}), not {@code agent.start} — and
* {@code agent.start} now only launches supported agent kinds, so this probes the seed pane's
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@Tag("contract")
class AgentControlContractTest {
private boolean noSocket() {
return !Files.exists(UnixSocketHerdrClient.defaultSocketPath());
}
@Test
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_env_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null,
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
try {
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the seed shell; saw: " + visible);
} finally {
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
}
}
}
@@ -1,262 +0,0 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-531. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that an
* operator-labelled tab is what makes a pane a lead, and — just as importantly — what does not.
*/
class LeadTabScannerTest {
private static final ObjectMapper MAPPER = new ObjectMapper();
private static final long TTL = TimeUnit.SECONDS.toNanos(10);
/**
* A herdr whose workspace/tab/pane topology is declared per test. Counts calls so the caching
* contract can be asserted, and can be made to fail on demand.
*/
private static final class TopologyHerdr implements HerdrClient {
/** workspace_id → label. */
final Map<String, String> workspaces = new LinkedHashMap<>();
/** tab_id → [workspace_id, label]. */
final Map<String, String[]> tabs = new LinkedHashMap<>();
/** pane_id → [tab_id, terminal_id]. */
final Map<String, String[]> panes = new LinkedHashMap<>();
int calls;
boolean failing;
TopologyHerdr workspace(String id, String label) {
workspaces.put(id, label);
return this;
}
TopologyHerdr tab(String tabId, String workspaceId, String label) {
tabs.put(tabId, new String[]{workspaceId, label});
return this;
}
TopologyHerdr pane(String paneId, String tabId, String terminalId) {
panes.put(paneId, new String[]{tabId, terminalId});
return this;
}
@Override
public JsonNode call(String method, Object params) {
calls++;
if (failing) {
throw new HerdrException("socket closed");
}
List<String> items = new ArrayList<>();
switch (method) {
case "workspace.list" -> {
workspaces.forEach((id, label) -> items.add(
"{\"workspace_id\":\"%s\",\"label\":\"%s\"}".formatted(id, label)));
return read("{\"workspaces\":[%s]}".formatted(String.join(",", items)));
}
case "tab.list" -> {
String ws = String.valueOf(((Map<?, ?>) params).get("workspace_id"));
tabs.forEach((id, t) -> {
if (ws.equals(t[0])) {
items.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":%s,"
+ "\"pane_count\":1}").formatted(id, t[0],
t[1] == null ? "null" : "\"" + t[1] + "\""));
}
});
return read("{\"tabs\":[%s]}".formatted(String.join(",", items)));
}
case "pane.list" -> {
panes.forEach((id, p) -> items.add(
"{\"pane_id\":\"%s\",\"tab_id\":\"%s\",\"terminal_id\":\"%s\"}"
.formatted(id, p[0], p[1])));
return read("{\"panes\":[%s]}".formatted(String.join(",", items)));
}
default -> throw new AssertionError("unexpected herdr call: " + method);
}
}
private static JsonNode read(String json) {
try {
return MAPPER.readTree(json);
} catch (Exception e) {
throw new AssertionError(e);
}
}
@Override
public void close() {
}
}
/**
* The usual shape: one user space with lead tabs, one worker space bridged owns.
*
* <p>Not closed: {@code close()} is a no-op on this fake, and every test needs the handle after
* the scanner is built (to mutate the topology or read {@code calls}).
*/
@SuppressWarnings("resource")
private TopologyHerdr twoLeads() {
return new TopologyHerdr()
.workspace("w1", "main")
.workspace("w9", "bridged-workers")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "lead: gpt-sol-5.6")
.tab("w1:t3", "w1", "notes")
.tab("w9:t1", "w9", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_gpt")
.pane("w1:p3", "w1:t3", "term_notes")
.pane("w9:p1", "w9:t1", "term_worker");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> configured,
AtomicLong clock) {
return new LeadTabScanner(herdr, "lead:", Set.of("bridged-workers"), configured, TTL,
clock::get);
}
@Test
void everyLabelledTabBecomesALeadNamedByItsLabel() {
Map<String, String> leads = scanner(twoLeads(), Map.of(), new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), leads,
"two leads discovered from labels alone — no terminal_id was ever configured");
}
@Test
void anUnlabelledTabContributesNothing() {
assertFalse(scanner(twoLeads(), Map.of(), new AtomicLong()).get().containsKey("term_notes"));
}
/**
* The guard that matters: bridged labels its own worker tabs, so if a worker space were scanned
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
* worker template never collides.
*/
@Test
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
TopologyHerdr herdr = twoLeads().tab("w9:t2", "w9", "lead: impostor")
.pane("w9:p2", "w9:t2", "term_impostor");
assertFalse(scanner(herdr, Map.of(), new AtomicLong()).get().containsKey("term_impostor"));
}
@Test
void aBarePrefixNamesNobodyAndIsRejected() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", "lead:").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of(), scanner(herdr, Map.of(), new AtomicLong()).get(),
"a lead with no name would resolve as PRIMARY with nothing to attribute it to");
}
@Test
void thePrefixMatchesCaseInsensitivelyAndTheNameIsTrimmed() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", " LEAD: opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of("term_a", "opus-5.0"), scanner(herdr, Map.of(), new AtomicLong()).get());
}
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing bridged placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0", scanner(herdr, Map.of(), new AtomicLong()).get().get("term_opus_split"));
}
@Test
void anExplicitlyConfiguredLeadIsMergedInAndOutranksALabel() {
Map<String, String> configured = Map.of("term_opus", "pinned-name", "term_extra", "from-config");
Map<String, String> leads = scanner(twoLeads(), configured, new AtomicLong()).get();
assertEquals("pinned-name", leads.get("term_opus"), "an explicit pin is the operator's last word");
assertEquals("from-config", leads.get("term_extra"), "a configured lead needs no tab at all");
assertEquals("gpt-sol-5.6", leads.get("term_gpt"));
}
// ── caching ─────────────────────────────────────────────────────────────────────────────────
@Test
void aSecondLookupWithinTheTtlDoesNotTouchHerdr() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
s.get();
int afterFirst = herdr.calls;
clock.addAndGet(TTL - 1);
s.get();
assertEquals(afterFirst, herdr.calls,
"resolve() runs on every request — an un-cached scan would put herdr on that path");
}
@Test
void aTabLabelledAfterStartupIsPickedUpOnceTheTtlExpires() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
assertFalse(s.get().containsKey("term_notes"));
herdr.tab("w1:t3", "w1", "lead: late-arrival"); // the operator renames their tab
clock.addAndGet(TTL);
assertEquals("late-arrival", s.get().get("term_notes"),
"the whole point over `leaders:`: no config edit, no restart");
}
@Test
void aFailedScanKeepsTheLeadsAlreadyKnownRatherThanDemotingThem() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
Map<String, String> before = s.get();
herdr.failing = true;
clock.addAndGet(TTL);
assertEquals(before, s.get(),
"a herdr hiccup must not silently demote a live lead to a worker mid-session");
}
@Test
void aFailedFirstScanStillHonoursTheConfiguredLeads() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
Map<String, String> leads = scanner(herdr, Map.of("term_x", "opus-5.0"), new AtomicLong()).get();
assertEquals(Map.of("term_x", "opus-5.0"), leads,
"config-named leads must not depend on herdr answering at all");
}
@Test
void aDownHerdrIsRetriedOncePerTtlNotOncePerRequest() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
s.get();
int afterFirst = herdr.calls;
s.get();
s.get();
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
}
}
@@ -1,27 +0,0 @@
package dev.ltms.bridged.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Unit tests for PID → pane resolution (the herdr half of connection-based MCP identity). */
class PaneLocatorTest {
private final PaneLocator loc = new PaneLocator(new FakeHerdr());
@Test
void resolvesTerminalForAForegroundPid() {
assertEquals("term_a", loc.terminalForPid(FakeHerdr.WORKER_PID));
}
@Test
void nullForAPidInNoPane() {
assertNull(loc.terminalForPid(999_999));
}
@Test
void nullForNonPositivePid() {
assertNull(loc.terminalForPid(0));
assertNull(loc.terminalForPid(-1));
}
}
@@ -1,302 +0,0 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
class CompletionResolverTest {
@Test
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.resolve("term_a", null); // no in-flight turn captured for this target
assertFalse(herdr.called("agent.read"),
"a turn nobody is blocked on must not cost a transcript scrape");
}
@Test
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
assertFalse(herdr.called("agent.read"),
"a wedge nobody is blocked on must not cost a transcript scrape");
}
@Test
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
}
// --- CB-115 clean scrape: extract the last assistant block ----------------
@Test
void extractsTheLastAssistantBlockStrippingChrome() {
String raw = """
⏺ Reading the file…
⏺ Done. The bug was an off-by-one in the loop bound.
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
⏵⏵ auto mode on · ? for shortcuts
""";
assertEquals("Done. The bug was an off-by-one in the loop bound.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void keepsMultiLineAssistantContent() {
String raw = "⏺ Line one.\nLine two.\n❯ ";
assertEquals("Line one.\nLine two.", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void fallsBackToRawTextWhenThereIsNoMarker() {
String raw = "plain worker output with no glyph";
assertEquals("plain worker output with no glyph", CompletionResolver.lastAssistantBlock(raw));
}
@Test
void blankScrapeYieldsEmpty() {
assertTrue(CompletionResolver.lastAssistantBlock("").isEmpty());
assertTrue(CompletionResolver.lastAssistantBlock(null).isEmpty());
}
@Test
void stripsSpinnerAndRuleChrome() {
String raw = """
⏺ Channel check confirmed — your message got through.
✻ Brewed for 11s
─────────────────────────────────────
""";
assertEquals("Channel check confirmed — your message got through.",
CompletionResolver.lastAssistantBlock(raw));
}
@Test
void cutsANextTurnPromptEchoAndTrailingTipsFromTheBlock() {
// The exact turn-2 leak: the scrape captured the settled answer, then a "✻ Cooked" spinner,
// then the NEXT turn's echoed prompt, then a "✶ Forming…" spinner and trailing tips/warnings
// whose lines (⎿, ⚠) are not themselves chrome-terminated. Stopping at the first boundary
// (the ✻ spinner) is what keeps every one of those interface lines out of the reply.
String raw = """
⏺ Channel confirmed — the bridge reply delivered successfully.
✻ Cooked for 9s
❯ Thanks. Now a small task: what is 17 * 23? Show just the number.
✶ Forming…
⎿ Tip: Name your conversations with /rename
⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY is set
""";
assertEquals("Channel confirmed — the bridge reply delivered successfully.",
CompletionResolver.lastAssistantBlock(raw));
}
// --- CB-115 misattribution guard: suppress a stale (unchanged) completion -------
@Test
void suppressesACompletionWhoseScrapeIsUnchangedFromDelivery() {
// Rapid back-to-back turn: the pane still shows the PREVIOUS turn's answer when this turn's
// (misattributed) completion boundary fires. The scrape == the delivery baseline, so the
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn); // scrape still "391" == baseline → suppress
assertFalse(waiter.isDone(), "a completion with no output change must not resolve the send");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
var turn = new CompletionResolver.InFlight(waiter, "391");
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(), "a completion with new output must resolve the send");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a");
herdr.readText("⏺ answer that /clear would erase\n❯ ");
resolver.resolveBeforePostAction("term_a");
assertTrue(waiter.isDone(), "the answer is captured before the adapter sends /clear");
assertEquals("answer that /clear would erase", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
resolver.resolve("term_a", turn); // scrape unchanged → clipped tail == baseline → suppress
assertFalse(waiter.isDone(),
"an unchanged >cap block must still be recognised as stale and suppressed");
assertTrue(rendezvous.isWaiting("term_a"), "the send stays waiting for a real reply");
}
@Test
void resolvesWhenThereIsNoBaseline() {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "with no baseline a completion resolves as before");
assertEquals("hello", waiter.getNow(null).text());
}
@Test
void resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent() {
// The most important branch of the CB-115 guard: a failed read means the resolver could not
// SEE the screen — "couldn't see", not "no change". It must still resolve the send (an empty
// tail beats hanging until the caller's timeout), even though a baseline was captured. The
// baseline here is "" (an empty pane at delivery), so without the !scrapeFailed clause the
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
resolver.resolve("term_a", turn);
assertTrue(waiter.isDone(),
"a failed scrape must still resolve the send, not hang until the caller's timeout");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind());
assertEquals("", waiter.getNow(null).text(), "the tail is empty because the screen was unreadable");
}
// --- CB-115/CB-116 fail guard: an already-done or absent waiter is left alone ---------
@Test
void failLeavesAnAlreadyResolvedWaiterUntouchedAndSkipsTheScrape() {
// The send was already resolved (e.g. by the worker's explicit reply) before fail fired.
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.fail("term_a", turn);
assertFalse(herdr.called("agent.read"),
"fail must not scrape a waiter that is already done");
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"fail must not overwrite the existing resolution");
assertEquals("already replied", waiter.getNow(null).text());
}
@Test
void failFallsBackToTheRegisteredWaiterWhenThereIsNoInFlightTurn() {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
assertTrue(waiter.isDone(), "fail falls back to the registered waiter when no turn is in flight");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind());
assertEquals("stuck on an error screen", waiter.getNow(null).text());
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
void aLateCompletionForOneTurnNeverResolvesTheNextTurnsWaiter() {
// The cross-turn stale reply the conversation test surfaced: turn N's completion fallback
// fires AFTER turn N was resolved by an explicit bridge_reply and turn N+1 has opened its own
// waiter on the same session. Resolving "whatever is waiting now" would hand turn N's stale
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
assertTrue(rendezvous.resolve("term_a", "N replied"));
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
assertFalse(waiterN1.isDone(), "turn N's late completion must not resolve turn N+1's waiter");
assertEquals(Rendezvous.Kind.REPLY, waiterN.getNow(null).kind(),
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
}
@@ -1,432 +0,0 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Golden-transcript tests for the status-gated injector (CB-103). Delivery rules are driven
* deterministically by feeding {@code onStatus}, so no timing or real polling is involved.
*/
class InjectorTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr();
private final Injector injector = new Injector(new AgentControl(herdr));
/**
* The logical messages delivered, in order. Under protocol 19 each delivery is one
* {@code agent.prompt} carrying the payload (it submits itself); the Enter nudge is a
* separate {@code agent.send_keys} and never appears here.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void deliversWhenIdle() {
CompletableFuture<Void> f = injector.enqueue(T, "hello");
assertFalse(f.isDone(), "not delivered until an injectable status arrives");
injector.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isDone());
assertEquals(List.of("hello"), sent());
}
@Test
void holdsDeliveryUntilTheWorkerIsAvailable() {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
ready.add(T); // the worker's Claude connects the bridge MCP
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send_keys"))
.count();
}
@Test
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
injector.onStatus(T, AgentStatus.IDLE); // still idle → re-nudge Enter
injector.onStatus(T, AgentStatus.IDLE); // and again
assertTrue(enterKeystrokes() > afterDeliver, "an unpicked-up delivery re-nudges Enter");
injector.onStatus(T, AgentStatus.WORKING); // worker finally starts
long atPickup = enterKeystrokes();
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(atPickup, enterKeystrokes(), "no more nudges once the worker has picked up");
}
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(List.of(), sent(), "must not inject mid-turn");
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("later"), sent());
}
@Test
void blockedIsInjectableButUnknownIsNot() {
injector.enqueue(T, "answer");
injector.onStatus(T, AgentStatus.UNKNOWN);
assertEquals(List.of(), sent(), "unknown status is not safe to inject");
injector.onStatus(T, AgentStatus.BLOCKED);
assertEquals(List.of("answer"), sent(), "blocked worker can be answered");
}
@Test
void twoRapidDeliveriesNeverInterleave() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
// First idle window delivers only m1, even if idle is observed twice before pickup.
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("m1"), sent(), "second message must wait for the turn to complete");
// Worker picks up m1 (works), then returns idle → m2 delivers.
injector.onStatus(T, AgentStatus.WORKING);
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("m1", "m2"), sent());
}
@Test
void transientUnknownDoesNotReleaseThePickupLatch() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.onStatus(T, AgentStatus.IDLE); // m1 sent, awaiting pickup
assertEquals(List.of("m1"), sent());
injector.onStatus(T, AgentStatus.UNKNOWN); // a detection glitch is NOT a pickup
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("m1"), sent(), "unknown must not let the next message interleave the turn");
injector.onStatus(T, AgentStatus.WORKING); // real pickup
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("m1", "m2"), sent());
}
@Test
void missedPickupEdgeIsReleasedByGraceSoTheQueueNeverWedges() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.onStatus(T, AgentStatus.IDLE); // m1 sent
assertEquals(List.of("m1"), sent());
// WORKING is never sampled (turn faster than the poll). The latch must release after
// the grace window so m2 is delivered rather than wedged forever.
for (int i = 0; i < 20; i++) injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("m1", "m2"), sent(), "safety valve must eventually deliver m2");
}
@Test
void fifoOrderAcrossManyTurns() {
injector.enqueue(T, "a");
injector.enqueue(T, "b");
injector.enqueue(T, "c");
for (int i = 0; i < 3; i++) {
injector.onStatus(T, AgentStatus.IDLE); // deliver one
injector.onStatus(T, AgentStatus.WORKING); // pickup
}
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("a", "b", "c"), sent());
}
@Test
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
assertEquals(Set.of(T), injector.activeTargets(),
"stays active so the poller can observe the worker pick the message up");
injector.onStatus(T, AgentStatus.WORKING); // pickup observed; now awaiting turn completion
assertEquals(Set.of(T), injector.activeTargets(),
"stays active after pickup so the working→idle completion boundary is observed");
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertTrue(injector.activeTargets().isEmpty(), "quiet once the delegated turn has completed");
}
@Test
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
assertEquals(List.of(), completed, "no completion until the turn returns to idle");
inj.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void postTurnResetSettlesBeforeNextDelegationWithoutBecomingATurn() {
AgentControl agents = new AgentControl(herdr);
class ResetListener implements TurnListener {
int completed;
@Override
public void onTurnComplete(String target) {
completed++;
}
@Override
public boolean hasPostTurnAction(String target) {
return true;
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completed++;
agents.send(target, "/clear"); // direct housekeeping, never Injector.enqueue
return true;
}
}
ResetListener listener = new ResetListener();
Injector inj = new Injector(agents, listener);
inj.enqueue(T, "first");
inj.enqueue(T, "second");
inj.onStatus(T, AgentStatus.IDLE); // first delegation
inj.onStatus(T, AgentStatus.WORKING);
inj.onStatus(T, AgentStatus.IDLE); // first complete; reset dispatched
assertEquals(List.of("first", "/clear"), sent(),
"same-tick completion must not let the queued delegation overtake reset");
inj.onStatus(T, AgentStatus.WORKING); // reset picked up, but this is not a bridge turn
inj.onStatus(T, AgentStatus.IDLE); // reset settled; second may now deliver
assertEquals(List.of("first", "/clear", "second"), sent());
assertEquals(1, listener.completed,
"reset settlement must not recursively emit another turn completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
// "the worker finished the task" signal, so the send should fall through to its timeout.
for (int i = 0; i < 15; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), completed, "no completion is synthesized from an unconfirmed turn");
}
/** Captures both turn-lifecycle callbacks so the CB-109 stall path can be asserted. */
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
completed.add(target);
}
@Override
public void onTurnFailed(String target) {
failed.add(target);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
private static final int STALL_SAMPLES = 130;
@Test
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < STALL_SAMPLES; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // then wedges
assertEquals(List.of(T), cap.failed, "a sustained unknown streak fails the outstanding send");
assertEquals(List.of(), cap.completed, "a wedge is a failure, not a completion");
assertTrue(inj.activeTargets().isEmpty(), "the wedged target is reclaimed, not polled forever");
}
@Test
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
for (int i = 0; i < 10; i++) inj.onStatus(T, AgentStatus.UNKNOWN); // brief glitch, well under grace
inj.onStatus(T, AgentStatus.IDLE); // working → idle: the real completion
assertEquals(List.of(), cap.failed, "a short unknown blip must not fail the turn");
assertEquals(List.of(T), cap.completed, "the streak reset, so the turn still completes");
}
@Test
void sendFailureDropsMessageAndFailsItsFuture() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
inj.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isCompletedExceptionally());
assertTrue(inj.activeTargets().isEmpty(), "poisoned message is dropped, not left blocking the queue");
}
@Test
void dropFailsPendingWaiters() {
CompletableFuture<Void> f = injector.enqueue(T, "orphan");
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a worker that vanishes mid-turn fails its in-flight send");
}
@Test
void dropFailsADeliveredTurnThatVanishesBeforePickupIsConfirmed() {
// Delivered but no WORKING sampled yet (awaitingCompletion=true, awaitingPickup still true,
// turnObserved=false) — a distinct state the other two drop tests don't cover. (Gap surfaced
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), cap.failed, "a delivery that vanishes before pickup still fails its send");
}
// ~60s of idle-but-not-ready at the 250ms prod poll interval; enough to trip the readiness grace.
private static final int READINESS_SAMPLES = 245;
@Test
void failsAQueuedMessageWhoseWorkerNeverBecomesReady() {
// CB-114: herdr keeps reporting the worker idle, but its Claude never connects the bridge MCP,
// so the readiness gate never opens. The message must not be held (and the target polled)
// forever — after the grace it fails, the caller unblocks via the worker-failure path, the
// target is reclaimed, and the never-set presence is cleared.
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of(), sent(), "a never-ready worker is never delivered to");
assertTrue(f.isCompletedExceptionally(), "the caller's future fails instead of hanging forever");
assertEquals(List.of(T), cap.failed, "the awaiting send resolves through the worker-failure path");
assertEquals(List.of(), cap.completed, "a never-ready worker is a failure, not a completion");
assertEquals(List.of(T), forgotten, "the never-ready worker's presence is cleared");
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
// available before the grace elapses, delivery proceeds as usual (the counter resets).
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
ready.add(T); // MCP connects before the grace elapses
inj.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent(), "a worker that connects within the grace is delivered to");
}
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
// linger past the worker's life (WorkerPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@Test
void pollerDeliversToAnIdleWorker() throws Exception {
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
FakeHerdr idle = new FakeHerdr().agentStatus("idle");
Injector inj = new Injector(new AgentControl(idle));
StatusPoller poller = new StatusPoller(new AgentControl(idle), inj, 10);
poller.start();
try {
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller");
delivered.get(2, TimeUnit.SECONDS); // completes when the poller drives the send
} finally {
poller.stop();
}
assertEquals(List.of("via-poller"), idle.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
})
.toList());
}
@Test
void deliveredFutureCarriesSendFailure() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
inj.onStatus(T, AgentStatus.IDLE);
ExecutionException ex = assertThrows(ExecutionException.class, f::get);
assertInstanceOf(HerdrException.class, ex.getCause());
}
}
@@ -1,201 +0,0 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-513 — the CB-505 authorization gate on the <strong>MCP</strong> entry path.
*
* <p>Why this file exists: CB-505 claimed authorization is "enforced on both entry paths", and it
* is — but only REST was ever tested ({@code BridgedAppAuthTest}). Coverage showed
* {@code BridgeMcp.deny()}, {@code principal()} and every tool-registration lambda at <em>zero</em>
* executed lines, because no test had ever constructed a {@code BridgeMcp} — the existing
* {@code BridgeMcpTest} calls only the static handler methods. An unexercised security control is
* a claim, not a control.
*
* <p>These tests construct a real {@code BridgeMcp} (which also exercises the constructor and the
* tool wiring) and drive the policy half of the gate directly.
*/
class BridgeMcpAuthzTest {
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private Metrics metrics;
private BridgeMcp mcp;
@AfterEach
void close() {
if (mcp != null) mcp.close();
}
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
private BridgeMcp mcp(boolean enforce) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
metrics = BridgedMetrics.create(sessions, new InMemoryReplyInbox());
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? new CallerResolver(identity) : null,
metrics);
return mcp;
}
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
// --- the table, enforced on THIS path too ---------------------------------------------------
@Test
void primaryMayOrchestrate() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN, Authz.Action.READ}) {
assertNull(m.denyFor(PRIMARY, a, "term_a"), a + " is the primary's to perform");
}
}
@Test
void aWorkerMayNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.SEND, Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
assertNotNull(denied, a + " must be refused to a worker");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_a"), "its own session is allowed");
assertNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_a"));
assertNotNull(m.denyFor(WORKER_A, Authz.Action.REPLY, "term_b"),
"worker A must not reply on worker B's session");
assertNotNull(m.denyFor(WORKER_A, Authz.Action.ASK, "term_b"));
}
@Test
void thePrimaryMayNotForgeAWorkerReplyOverMcp() {
BridgeMcp m = mcp(true);
// A forged reply would resolve the very rendezvous the primary is blocked on.
assertNotNull(m.denyFor(PRIMARY, Authz.Action.REPLY, "term_a"));
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
}
// --- CB-548: the architect on this path ------------------------------------------------
@Test
void anArchitectMaySendAndReadButNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.SEND, "term_a"),
"delegating a turn is the architect's job");
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.READ, null));
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(ARCH_DESIGN, a, null);
assertNotNull(denied, a + " must be refused to an architect");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.REPLY, "term_design"));
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.ASK, "term_design"));
assertNotNull(m.denyFor(ARCH_DESIGN, Authz.Action.REPLY, "term_a"),
"architect 'lead-designer' must not reply on worker term_a's session");
}
@Test
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
BridgeMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(ANON, Authz.Action.READ, null);
assertNotNull(denied, "authenticated as nothing ⇒ authorized for nothing");
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"),
"a missing credential is 401-shaped, not 403-shaped");
}
@Test
void aWrongRoleIsCountedAsForbiddenNotUnauthenticated() {
BridgeMcp m = mcp(true);
assertNotNull(m.denyFor(WORKER_A, Authz.Action.SPAWN, null));
assertEquals(1, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "forbidden"));
assertEquals(0, metrics.count(BridgedMetrics.AUTH_FAILURES, "reason", "unauthenticated"),
"the caller IS authenticated — it is just not the right role");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing BridgeMcpTest cases rely on no authorization being enforced.
BridgeMcp m = mcp(false);
assertNull(m.denyFor(ANON, Authz.Action.SPAWN, null),
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
}
// --- identity reconstruction from the transport context ------------------------------------
@Test
void principalIsRebuiltFromTheStashedRole() {
assertEquals(Role.WORKER, BridgeMcp.principalFrom("WORKER", "term_a", 7).role());
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
// CB-548: an architect round-trips through the same stash, carrying its slot name.
Principal arch = BridgeMcp.principalFrom("ARCHITECT", "term_design", 7, "lead-designer");
assertEquals(Role.ARCHITECT, arch.role());
assertEquals("lead-designer", arch.name());
assertEquals("term_design", arch.terminal());
}
@Test
void aMissingRoleFallsBackToTheHistoricalInterpretation() {
// Legacy path: no role stashed. A terminal means worker; its absence meant "the primary",
// which is exactly the pre-CB-501 default CB-501 inverted — preserved only here.
assertEquals(Role.WORKER, BridgeMcp.principalFrom(null, "term_a", 7).role());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom(null, null, 7).role());
}
}
@@ -1,487 +0,0 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Parity tests for the MCP tool adapters — they must produce the same outcomes as the REST routes,
* since both drive the same {@link MessageService}/{@link Rendezvous}. The MCP wire protocol itself
* is the SDK's concern; here we test the thin adapter logic directly.
*/
class BridgeMcpTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(allow), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> "tok");
}
private static SessionManager sessionManager(FakeHerdr h, String baseUrl, Set<String> allow) {
return new SessionManager(workerService(h, baseUrl, allow));
}
@Test
void sendThenReplyRoundTrips() throws Exception {
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "LGTM");
assertEquals("delivered", textOf(reply));
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("LGTM", textOf(res));
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
assertNotEquals(Boolean.TRUE, accepted.isError());
String out = textOf(accepted);
assertTrue(out.contains("ticket="), out);
String ticket = out.substring(out.indexOf("ticket=") + "ticket=".length()).trim();
// Wait until the send has opened its waiter before replying (CB-307: reply never errors,
// so the old retry-on-error pattern no longer works — it would queue instead of resolve).
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "async LGTM");
assertEquals("delivered", textOf(reply));
// Poll until the async send completes and reports the reply.
McpSchema.CallToolResult polled = BridgeMcp.poll(messages, ticket, null);
deadline = System.currentTimeMillis() + 3000;
while (!textOf(polled).contains("async LGTM") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(10);
polled = BridgeMcp.poll(messages, ticket, null);
}
assertEquals("async LGTM", textOf(polled));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown ticket"));
}
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@Test
void sendRejectsMissingArgs() {
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
}
@Test
void replyWithNoPendingSendIsQueuedNotError() {
// CB-307: a reply with no open send is now queued in the inbox, not an error.
McpSchema.CallToolResult res = BridgeMcp.reply(messages, "term_a", "orphan");
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
assertEquals("delivered", textOf(res));
// The reply is drainable by target.
var drained = messages.drainReplies("term_a");
assertEquals(1, drained.size());
assertEquals("orphan", drained.getFirst().content());
}
@Test
void bridgePollWithTargetDrainsReplies() {
// A reply with no open send queues it in the inbox.
BridgeMcp.reply(messages, "term_a", "queued-msg");
// bridge_poll with target drains the inbox.
McpSchema.CallToolResult res = BridgeMcp.poll(messages, null, "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
String text = textOf(res);
assertTrue(text.contains("queued-msg"), "the drained reply should appear in the result");
// Second drain returns empty.
McpSchema.CallToolResult empty = BridgeMcp.poll(messages, null, "term_a");
assertEquals("[]", textOf(empty));
}
@Test
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the send must be waiting for the ask to surface to");
// The worker asks mid-turn; the call blocks for the primary's answer.
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
// The primary's send unblocks with the question and a turnId to answer on.
McpSchema.CallToolResult q = send.get(6, TimeUnit.SECONDS);
assertNotEquals(Boolean.TRUE, q.isError());
String qt = textOf(q);
assertTrue(qt.contains("[question]"), qt);
String afterMarker = qt.substring(qt.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterMarker.substring(0, afterMarker.indexOf('"'));
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
// The resumed worker replies, resolving the answering send (wait for the reopened waiter).
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
McpSchema.CallToolResult reply = BridgeMcp.reply(messages, "term_a", "done");
assertEquals("delivered", textOf(reply));
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@Test
void askFromANonWorkerConnectionIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.ask(messages, null, "which config?", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("workers only"), textOf(res));
}
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.answer(messages, "term_a#999", "too late", 500L);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@Test
void spawnReturnsTheNewWorkersSessionAndPane() {
FakeHerdr h = new FakeHerdr();
SessionManager sm = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
// CB-519: the "paneId" wire field now carries the host-unique opaque id, not the herdr pane.
WorkerSession s = sm.roster().getFirst();
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertNotEquals("w9:pRoot_1", s.paneId(), "the id is decoupled from the herdr pane coordinate");
assertTrue(out.contains("\"status\":\"spawning\""), out);
}
@Test
void spawnRejectsAnOffAllowlistProfileWithoutTouchingHerdr() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "https://api.anthropic.com", Set.of("gx00.gw")), null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("subscription boundary"));
assertFalse(h.called("agent.start"), "the guard must block before any spawn");
}
@Test
void spawnRejectsAnUnknownProfileAsAnError() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res =
BridgeMcp.spawn(sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "nope");
assertTrue(res.isError());
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
}
@Test
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
@SuppressWarnings("unchecked")
Map<String, Object> create = (Map<String, Object>) h.lastCall("tab.create").params();
assertEquals("/req/dir", create.get("cwd"));
}
@Test
void profilesListsConfiguredProfilesAndDefault() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("ltms-local"), out);
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
}
@Test
void listReportsTrackedWorkers() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-304", null));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"state\":\"spawning\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions,
Map.of("term_me", "opus-5.0", "term_peer", "gpt-sol-5.6"), "term_me");
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"name\":\"opus-5.0\""), out);
assertTrue(out.contains("\"name\":\"gpt-sol-5.6\""), out);
assertTrue(out.contains("\"sessionId\":\"term_peer\""), out);
// The caller's own row is flagged, and only the caller's — a peer must be distinguishable
// from self without a second bridge_whoami call.
assertEquals(1, out.split("\"self\":true", -1).length - 1, out);
assertTrue(out.indexOf("term_me") < out.indexOf("\"self\":true"), out);
}
@Test
void listReportsBothHalvesEvenWhenEmpty() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
// An absent "leads" key is what made an empty worker roster read as "no peers" (CB-535).
String out = textOf(res);
assertTrue(out.contains("\"leads\":[]"), out);
assertTrue(out.contains("\"workers\":[]"), out);
}
@Test
void listReportsALeadHerdrCannotSeeAsUnknown() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions,
Map.of("term_ghost", "gone-away"), "term_me");
// Reported, not hidden: an unreachable peer is exactly what a would-be sender needs to see.
String out = textOf(res);
assertTrue(out.contains("\"name\":\"gone-away\""), out);
assertTrue(out.contains("\"status\":\"unknown\""), out);
}
@Test
void stopTearsDownAWorkerByPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), "w9:pW");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("stopped w9:pW", textOf(res));
assertTrue(h.called("pane.close"));
}
@Test
void stopRequiresAPaneId() {
FakeHerdr h = new FakeHerdr();
assertTrue(BridgeMcp.stop(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), " ").isError());
}
@Test
void bridgeAckReturnsConfirmationForValidArgs() {
McpSchema.CallToolResult res = BridgeMcp.ack(messages, "term_a", "msg-1");
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
}
@Test
void bridgeAckRejectsMissingArgs() {
assertTrue(BridgeMcp.ack(messages, null, "msg-1").isError());
assertTrue(BridgeMcp.ack(messages, "term_a", null).isError());
assertTrue(BridgeMcp.ack(messages, " ", "msg-1").isError());
}
@Test
void bridgeAckRemovesSpecificReply() {
// Queue a reply and capture its msgId.
BridgeMcp.reply(messages, "term_a", "orphan");
var before = messages.drainReplies("term_a");
assertEquals(1, before.size(), "one reply in the inbox");
String msgId = before.getFirst().msgId();
// Publish the same reply again and ack it via bridge_ack surface.
BridgeMcp.reply(messages, "term_a", "orphan-again");
var peeked = messages.drainReplies("term_a");
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
// ackReply works (no-op since published with a different UUID, but callable).
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = BridgeMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
@Test
void whoamiReportsThePrimaryAsPrimaryAndNothingElse() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.primary(100), sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"primary\""), out);
// The primary owns no session — leaking a sessionId here would invite it to reply as one.
assertFalse(out.contains("sessionId"), out);
}
@Test
void whoamiReportsAWorkerWithItsRegisteredSession() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-517", null));
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"" + s.terminalId() + "\""), out);
assertTrue(out.contains("\"profile\":\"ltms-local\""), out);
assertTrue(out.contains("\"worktree\":\"" + s.worktree() + "\""), out);
assertTrue(out.contains("\"branch\":\"" + s.branch() + "\""), out);
assertTrue(out.contains("\"owner\":\"term_primary\""), out);
}
/**
* A worker the registry has no record of — it outlived a daemon restart — must still learn the
* load-bearing fact. Degrading to "I don't know who you are" would put it back to guessing,
* which is the failure this tool exists to remove.
*/
@Test
void whoamiStillReportsWorkerRoleWhenTheSessionIsUnregistered() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker("term_orphan", 200),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"worker\""), out);
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
}
/**
* CB-548: an architect reports its role and which gateway-local slot its pane is bound to —
* the same shape as a lead, under the architect key, so it can tell a peer where to reach it.
*/
@Test
void whoamiReportsAnArchitectWithItsSlotAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.architect("lead-designer", "term_design", 400),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"architect\""), out);
assertTrue(out.contains("\"architect\":\"lead-designer\""), out);
assertTrue(out.contains("\"sessionId\":\"term_design\""), out);
}
}
@@ -1,46 +0,0 @@
package dev.ltms.bridged.mcp;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;
/** Connection → caller-identity resolution, with the OS peer-PID lookup faked. */
class ConnectionIdentityTest {
private final FakeHerdr herdr = new FakeHerdr();
private ConnectionIdentity with(PeerPidLookup pids) {
return new ConnectionIdentity(new PaneLocator(herdr), pids);
}
@Test
void resolvesWorkerFromLoopbackPeerPid() {
assertEquals("term_a", with(_ -> FakeHerdr.WORKER_PID).callerTerminal("127.0.0.1", 55555));
}
@Test
void nullForOffHostCaller() {
// A non-loopback peer can't be an on-host worker → treat as primary/unknown.
assertNull(with(_ -> FakeHerdr.WORKER_PID).callerTerminal("10.0.0.9", 55555));
}
@Test
void nullWhenPidOwnsNoPane() {
// e.g. the primary — its PID maps to no worker pane.
assertNull(with(_ -> 999_999).callerTerminal("127.0.0.1", 55555));
}
@Test
void resolvesTheCallersPidAndCwd() {
// CB-112: the primary maps to no pane, but its PID and cwd are still readable.
ConnectionIdentity id = new ConnectionIdentity(
new PaneLocator(herdr), _ -> 999_999, pid -> pid == 999_999 ? "/main/project" : null);
ConnectionIdentity.Caller c = id.resolve("127.0.0.1", 55555);
assertNull(c.terminal(), "the primary owns no worker pane");
assertEquals(999_999, c.pid());
assertEquals("/main/project", id.cwdForPid(c.pid()), "the primary's cwd is resolvable from its PID");
assertNull(id.cwdForPid(-1), "no cwd for an unresolved PID");
}
}
@@ -1,475 +0,0 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The message layer's resolution paths (CB-104 reply + CB-106 completion fallback). The turn is
* driven deterministically by feeding {@code onStatus} rather than running a real poller.
*/
class MessageServiceTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final Injector injector = new Injector(agents, completion);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
return CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 5000));
}
private void awaitWaiting() throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
}
@Test
void completionFallbackResolvesATurnThatNeverCalledBridgeReply() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
injector.onStatus(T, AgentStatus.IDLE); // deliver the task (baselines the pre-turn content)
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up and works
herdr.readText("BUILD GREEN: 391 files"); // the worker's turn produced new output
injector.onStatus(T, AgentStatus.IDLE); // working → idle: turn complete, no bridge_reply
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
"an unreplied but finished turn resolves via the completion fallback");
assertEquals("BUILD GREEN: 391 files", reply.text(), "the scraped transcript tail is returned");
assertTrue(reply.completed(), "a scraped completion still counts as completed");
}
@Test
void explicitBridgeReplyResolvesAsReplied() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker working
assertTrue(rendezvous.resolve(T, "LGTM ship it"), "an explicit reply resolves the send");
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, reply.outcome());
assertEquals("LGTM ship it", reply.text());
}
@Test
void aWedgedWorkerResolvesTheSendAsFailedWithTheErrorContext() throws Exception {
herdr.readText("API Error: Unable to connect to API (ENOTFOUND)");
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
for (int i = 0; i < 130; i++) injector.onStatus(T, AgentStatus.UNKNOWN); // then wedges (CB-109)
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertFalse(reply.completed(), "a wedge is terminal but not a successful completion");
assertTrue(reply.text().contains("ENOTFOUND"), "the error screen is carried as the failure reason");
}
@Test
void aWorkerThatVanishesMidTurnResolvesTheSendAsFailed() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts the turn
// The worker's pane crashes — the poller sees a *_not_found and drops it (CB-110).
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
MessageService.Reply reply = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome(),
"a delivered send whose worker vanishes fails instead of hanging to the timeout");
assertFalse(reply.completed());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
void askSurfacesAsAQuestionAndTheAnswerResumesTheSameTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// The worker asks mid-turn on its own thread; the call blocks for the primary's answer.
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's blocking send unblocks with the question and a turnId to answer on.
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "a question carries a turnId to answer on");
// The primary answers via bridge_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
// The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting(); // the answering send has (re)opened its forward waiter
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void duplicateAsksFromTheSameSessionCoalesceToOneTurn() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
// A transport retry: two concurrent bridge_ask calls from the same worker session.
CompletableFuture<MessageService.AskResult> ask1 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
CompletableFuture<MessageService.AskResult> ask2 =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
// The primary's single blocked send surfaces exactly ONE question (one turnId).
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertEquals("which config file?", q.text());
assertNotNull(q.turnId(), "only one turnId should be minted");
// The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a1.outcome());
assertEquals("config.yaml", a1.answer());
assertEquals(MessageService.AskOutcome.ANSWERED, a2.outcome());
assertEquals("config.yaml", a2.answer());
// The resumed worker finishes with a structured reply, resolving the answering send.
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"), "the worker's final reply resolves the answering send");
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
assertEquals("done", done.text());
}
@Test
void askWithNoOpenDelegationReturnsNoWaiter() {
MessageService.AskResult r = messages.ask(T, "anyone listening?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"a question with no blocked send has no primary to answer it");
}
@Test
void askTimesOutWhenThePrimaryNeverAnswers() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
MessageService.AskResult r = messages.ask(T, "still there?", 200); // primary never answers
assertEquals(MessageService.AskOutcome.TIMED_OUT, r.outcome());
// The send itself already unblocked with the question the instant the ask surfaced.
MessageService.Reply q = send.get(2, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
}
@Test
void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
// --- timeout, answer, poll, and lock-contention edges ----------------------------------
@Test
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working");
assertNull(r.text());
}
@Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
// No rendezvous.resolve(T, ...) — the reply future rides out its short timeout.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, r.outcome(),
"a delivered send whose worker never replies times out as still working");
}
@Test
void answerTimesOutWhenTheResumedWorkerNeverReplies() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertNotNull(q.turnId());
// The primary answers, unblocking the worker; but the worker never sends the follow-up
// bridge_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working");
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome());
assertEquals("config.yaml", a.answer());
}
@Test
void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
}
@Test
void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "async result"), "a reply resolves the async send");
// Wait for the background send to finish and publish a DONE view.
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket);
//noinspection BusyWait
Thread.sleep(5);
}
assertNotNull(view, "a resolved async send must become DONE");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("async result", view.reply(), "the completed ticket reports the reply");
assertEquals("reply", view.replySource(), "a structured bridge_reply is sourced from 'reply'");
}
@Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang");
assertNull(busy.text());
// Release the first send so it resolves cleanly and the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE); // deliver the first message
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
assertTrue(rendezvous.resolve(T, "first done"), "the first send resolves with a reply");
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
assertEquals("first done", firstReply.text());
}
// --- CB-307 reply inbox ----------------------------------------------------------------
@Test
void replyQueuesInInboxWhenNoSendIsOpen() {
// No send is open for this session — reply should queue in the inbox.
assertTrue(messages.reply(T, "queued-text"), "reply should succeed (queued)");
var drained = messages.drainReplies(T);
assertEquals(1, drained.size());
assertEquals("queued-text", drained.getFirst().content());
}
@Test
void replyResolvesOpenSendDoesNotQueue() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
// An explicit reply resolves the open send.
assertTrue(messages.reply(T, "send-resolved"), "reply should succeed (resolved live send)");
// The inbox should be empty — the reply went to the send, not the inbox.
assertTrue(messages.drainReplies(T).isEmpty(), "no reply in the inbox");
MessageService.Reply r = send.get(3, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, r.outcome());
assertEquals("send-resolved", r.text());
}
@Test
void drainRepliesReturnsAllPendingThenEmptyOnNextCall() {
messages.reply(T, "msg-1");
messages.reply(T, "msg-2");
var first = messages.drainReplies(T);
assertEquals(2, first.size());
var second = messages.drainReplies(T);
assertTrue(second.isEmpty(), "second drain should be empty (acked)");
}
@Test
void aQuestionIsNeverQueuedInTheInbox() {
// No send is open — bridge_ask with no delegation returns NO_WAITER,
// and the question text MUST NOT appear in the reply inbox.
// The inbox is only fed by MessageService.reply(), not by bridge_ask.
MessageService.AskResult r = messages.ask(T, "anyone there?", 500);
assertEquals(MessageService.AskOutcome.NO_WAITER, r.outcome(),
"bridge_ask with no open delegation must return NO_WAITER, never queued");
assertTrue(messages.drainReplies(T).isEmpty(), "questions must never be queued");
}
@Test
void completionFallbackIsNeverQueued() throws Exception {
// The fallback resolves a captured waiter, never the inbox.
CompletableFuture<MessageService.Reply> send = sendAsync();
awaitUninterruptibly(T);
injectDelivery();
// The worker never sends bridge_reply, but the turn completes.
herdr.readText("done-scraped");
completion.onTurnComplete(T); // The fallback arms and resolves the captured waiter.
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, r.outcome());
// The inbox should be empty — the reply went to the captured waiter.
assertTrue(messages.drainReplies(T).isEmpty(), "completion fallback must not queue");
}
// --- helpers ---------------------------------------------------------------------------
/** Like {@link #awaitWaiting()} but rethrows as unchecked. */
private void awaitUninterruptibly(String session) {
try {
awaitWaiting();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException(e);
}
}
/** Set up a delivered turn so the worker is working, ready for an ask or completion. */
private void injectDelivery() {
herdr.readText("$ prompt"); // pre-turn content baseline
injector.onStatus(T, AgentStatus.IDLE); // deliver the task
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up
}
// --- CB-516: a released session must not leave a send hanging ------------------------------
/**
* The bug this fixes: tearing a worker down left its rendezvous waiter open, so a blocking send
* kept blocking and an async one kept reporting PENDING until the 30-minute async timeout —
* even though the worker provably no longer existed.
*/
@Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, r.outcome(),
"an abandoned send fails rather than riding out its timeout");
assertEquals("session released", r.text(), "the caller is told why");
}
@Test
void abandonIsANoOpWhenNobodyIsWaiting() {
assertFalse(messages.abandon(T, "session released"),
"no open send ⇒ nothing to abandon");
}
@Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer"));
assertFalse(messages.abandon(T, "session released"),
"a send already answered by the worker must not be clobbered");
MessageService.Reply r = send.get(5, TimeUnit.SECONDS);
assertEquals("the real answer", r.text());
}
/** The async path is the one that hung: poll must report FAILED, not PENDING forever. */
@Test
void anAbandonedAsyncTaskPollsAsFailedNotPending() throws Exception {
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase());
messages.abandon(T, "session released");
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 3000;
while (System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
if (view.phase() != MessageService.Phase.PENDING) break;
Thread.sleep(10);
}
assertNotNull(view);
assertEquals(MessageService.Phase.FAILED, view.phase(),
"a delegation whose worker is gone must not keep reporting PENDING");
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
}
@@ -1,304 +0,0 @@
package dev.ltms.bridged.msg;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.Collections;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
* bounded reminders, and stop conditions.
*
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
* tests use a simple client with no concurrency concern.
*/
class ReplyPushLoopTest {
private static final String PRIMARY = "term_primary";
private static final String WORKER = "term_worker";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
private AgentControl agents;
private InMemoryReplyInbox inbox;
private ScheduledExecutorService scheduler;
@BeforeEach
void setUp() {
registry = new PrimaryRegistry(PRIMARY);
inbox = new InMemoryReplyInbox();
inbox.own(WORKER); // CB-520: the inbox only peeks/acks targets it owns
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@AfterEach
void tearDown() {
scheduler.shutdownNow();
}
// --- decide() logic ------------------------------------------------------------------------
@Test
void decideWithoutPrimaryIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
}
@Test
void decideWithEmptyInboxIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
@Test
void decideAtCapIsStop() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
}
@Test
void decideUnderCapWithInjectablePrimaryIsInject() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithBlockedPrimaryIsInject() {
agents = agentWithStatus("blocked");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"BLOCKED is injectable");
}
@Test
void decideUnderCapWithDonePrimaryIsInject() {
agents = agentWithStatus("done");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
"DONE is injectable");
}
@Test
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
agents = agentWithStatus("working");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
agents = agentWithStatus("unknown");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
}
@Test
void decideStopsAfterInboxIsEmptied() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
inbox.ack(WORKER, "m1");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
}
// --- onReplyQueued integration -------------------------------------------------------------
@Test
void injectablePrimaryCausesExactlyOneNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (1 agent.prompt call) should have been sent");
// Exactly one nudge = exactly 1 agent.prompt call (it submits itself)
assertEquals(1, rec.sendCount());
assertTrue(rec.sentParams().stream()
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
"nudge text should contain bridge_poll");
}
@Test
void onReplyQueuedIsIdempotentPerTarget() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(1, 100);
loop.onReplyQueued(WORKER);
loop.onReplyQueued(WORKER); // second call — should be a no-op
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"expected exactly one nudge (1 prompt)");
Thread.sleep(200);
assertEquals(1, rec.sendCount(),
"second onReplyQueued must not trigger another nudge");
}
@Test
void sendsUpToCapThenStops() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
cap + " nudges (" + cap + " prompts) should have fired");
Thread.sleep(300);
assertEquals(cap, rec.sendCount(),
"exactly " + cap + " agent.prompt calls (cap=" + cap + ")");
}
// --- nudge format --------------------------------------------------------------------------
@Test
void nudgeFormatIsCorrect() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
assertTrue(nudge.contains("Worker term_worker"));
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
}
// --- metrics (CB-512) ----------------------------------------------------------------------
@Test
void successfulNudgeIncrementsDelivered() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
loop(1, 50, metrics).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (1 agent.prompt call) should have been sent");
// The delivered count is bumped on the scheduler thread right after the send that releases
// the latch — settle briefly so the counter is published before we read it.
Thread.sleep(200);
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
"a successfully sent nudge must count as delivered");
}
@Test
void reminderCapIncrementsExhausted() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the reminder cap must count as exhausted");
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
return loop(5, 100);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs, Metrics metrics) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, metrics);
}
private static AgentControl agentWithStatus(String status) {
return new AgentControl(new FakeHerdrClient(status));
}
/** Non-recording (single-threaded) fake — safe for decide() tests. */
private static final class FakeHerdrClient implements HerdrClient {
private final String agentStatus;
FakeHerdrClient(String agentStatus) {
this.agentStatus = agentStatus;
}
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", agentStatus));
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
/**
* Thread-safe recording fake that counts agent.prompt calls (protocol 19: one nudge = one
* prompt). Uses synchronized access so the scheduler thread and test thread never race.
*/
private static final class RecordingHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(1);
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", "idle")); // recording double is always injectable
}
if ("agent.prompt".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
long sendCount() {
return calls.size();
}
List<Map.Entry<String, Object>> sentParams() {
return List.copyOf(calls);
}
@Override
public void close() {
}
}
private static RecordingHerdrClient recordingClient() {
return new RecordingHerdrClient();
}
}
@@ -1,183 +0,0 @@
package dev.ltms.bridged.placement;
import org.junit.jupiter.api.Test;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the placement policies. They run with no herdr and no launcher — pure selection
* logic exercised through the descriptor type so CB-308 host expansion will not need to rewrite
* these assertions.
*/
class PlacementPolicyTest {
private static Function<String, Integer> noSessions() {
return name -> 0;
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
return new PlacementContext("b", candidates, liveCount, unreachable);
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount) {
return ctx(candidates, liveCount, Set.of());
}
@Test
void fixedReturnsDefaultEvenIfOtherProfilesExist() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = ctx(List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b")), noSessions());
assertEquals("b", policy.select(ctx).profile());
}
@Test
void fixedFallsBackToFirstCandidateWhenNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of());
assertEquals("a", policy.select(ctx).profile());
}
@Test
void fixedThrowsWhenNoProfilesAndNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
assertThrows(PlacementException.class, () -> policy.select(ctx));
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"),
PlacementCandidate.profile("c"));
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("c", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
}
@Test
void roundRobinSkipsProfilesAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 2),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = Map.of("a", 2)::get;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void roundRobinThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = Map.of("a", 1, "b", 1)::get;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedAlternatesEvenlyWithEqualWeights() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.5f, null),
PlacementCandidate.profile("b", 0.5f, null));
int a = 0, b = 0;
for (int i = 0; i < 100; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(50, a, "equal weights should split 50/50");
assertEquals(50, b);
}
@Test
void weightedHoldsThreeToOneRatio() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.75f, null),
PlacementCandidate.profile("b", 0.25f, null));
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "0.75/0.25 should yield a 3:1 ratio over a multiple of 4");
assertEquals(10, b);
}
@Test
void weightedSkipsProfileAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void weightedThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = name -> 1;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedThrowsWhenAllUnreachable() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions(), Set.of("a", "b"))));
assertTrue(e.getMessage().contains("unreachable"), e.getMessage());
}
@Test
void mixedExclusionMessageNamesBothReasons() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
Set<String> unreachable = new HashSet<>();
unreachable.add("b");
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount, unreachable)));
assertTrue(e.getMessage().contains("1 at maxLoad"), e.getMessage());
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
}
}
@@ -1,489 +0,0 @@
package dev.ltms.bridged.rest;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.Worktrees;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import static org.junit.jupiter.api.Assertions.*;
/**
* REST acceptance tests — the feature contract over plain HTTP with a fake herdr, no
* live daemon and no Claude in the loop. This is the surface later MCP tools match by
* parity, and where the subscription boundary and worker placement are proven at the
* API edge.
*/
class BridgedAppTest {
private final ObjectMapper mapper = new ObjectMapper();
private final HttpClient http = HttpClient.newHttpClient();
private WorkerPresence presence;
private Javalin app;
private StatusPoller poller;
@AfterEach
void stop() {
if (poller != null) poller.stop();
if (app != null) app.stop();
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow) {
return start(herdr, workerBaseUrl, allow, "tab");
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement) {
return start(herdr, workerBaseUrl, allow, placement, new GitWorktrees());
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
"ltms-local", workerBaseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
placement, "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(allow),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, worktrees);
this.presence = sessions.asPresence();
Injector injector = new Injector(agents);
poller = new StatusPoller(agents, injector, 5); // delivers when the fake reports idle
poller.start();
Rendezvous rendezvous = new Rendezvous();
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
sessions.onAcquire(inbox::own);
// The static REST tests address "term_a" without acquiring it through SessionManager, so own
// it directly so the inbox contract holds for those endpoints.
inbox.own("term_a");
MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
}
private int startHealthy() {
return start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
}
private HttpResponse<String> postMessage(int port, String json) throws Exception {
return postJson(port, "/sessions/term_a/message", json);
}
private HttpResponse<String> postJson(int port, String path, String json) throws Exception {
HttpRequest r = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(json)).build();
return http.send(r, HttpResponse.BodyHandlers.ofString());
}
private HttpResponse<String> req(int port, String method, String path) throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path));
b = switch (method) {
case "POST" -> b.POST(HttpRequest.BodyPublishers.noBody());
case "DELETE" -> b.DELETE();
default -> b.GET();
};
return http.send(b.build(), HttpResponse.BodyHandlers.ofString());
}
@SuppressWarnings("unchecked")
private Map<String, Object> params(FakeHerdr herdr, String method) {
return (Map<String, Object>) herdr.lastCall(method).params();
}
@Test
void healthzOkWhenHerdrAnswers() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/healthz");
assertEquals(200, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("ok", body.get("status").asText());
assertEquals(19, body.get("herdr").get("protocol").asInt());
}
@Test
void healthzDegradedWhenHerdrDown() throws Exception {
FakeHerdr down = new FakeHerdr().healthy(false);
int port = start(down, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "GET", "/healthz");
assertEquals(503, res.statusCode());
assertEquals("degraded", mapper.readTree(res.body()).get("status").asText());
}
@Test
void sessionsMapsWorkspaceList() throws Exception {
int port = startHealthy();
JsonNode sessions = mapper.readTree(req(port, "GET", "/sessions").body()).get("sessions");
assertEquals(2, sessions.size());
assertEquals("w1", sessions.get(0).get("id").asText());
assertEquals("done", sessions.get(1).get("agentStatus").asText());
}
@Test
void agentsExposesSessionUuid() throws Exception {
int port = startHealthy();
JsonNode agents = mapper.readTree(req(port, "GET", "/agents").body()).get("agents");
assertEquals(1, agents.size());
assertEquals("sess-1111", agents.get(0).get("sessionId").asText());
assertEquals("idle", agents.get(0).get("status").asText());
}
@Test
void spawnWorkerLandsInOwnTabInWorkerSpaceAndInjectsBaseUrl() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers");
assertEquals(201, res.statusCode());
JsonNode body = mapper.readTree(res.body());
// CB-519: the responded paneId is a host-unique opaque UUID, not the herdr pane coordinate.
String id = body.get("paneId").asText();
assertNotEquals("w9:pRoot_1", id, "paneId is the host-unique id, not the herdr pane");
assertDoesNotThrow(() -> UUID.fromString(id), "paneId must be a UUID: " + id);
assertEquals("spawning", body.get("state").asText());
// Subscription boundary (protocol 19): tab.create carried base_url + token in its env
// map — the seed shell the agent starts into is what inherits them.
Map<String, Object> create = params(herdr, "tab.create");
@SuppressWarnings("unchecked")
Map<String, String> env = (Map<String, String>) create.get("env");
assertEquals("http://gx00.gw:8000", env.get("ANTHROPIC_BASE_URL"));
assertEquals("tok-abc", env.get("ANTHROPIC_AUTH_TOKEN"));
// Placement: worker space ensured, worker started INTO its tab's seed pane (which
// becomes the worker pane — nothing is dropped), and the tab given a friendly label.
Map<String, Object> start = params(herdr, "agent.start");
assertTrue(herdr.called("workspace.create"), "worker space must be found-or-created");
assertEquals("claude", start.get("kind"), "herdr resolves the executable from kind");
assertEquals("w9:pRoot_1", start.get("pane_id"), "worker must start into its tab's seed pane");
assertFalse(herdr.called("pane.close"), "the seed pane IS the worker pane — never dropped");
assertEquals("worker: ltms-local #1", params(herdr, "tab.rename").get("label"),
"tab label carries the worker number so siblings stay distinct");
}
@Test
void profilesEndpointListsConfiguredProfilesAndDefault() throws Exception {
int port = startHealthy();
JsonNode body = mapper.readTree(req(port, "GET", "/profiles").body());
assertEquals("ltms-local", body.get("default").asText());
assertEquals("ltms-local", body.get("profiles").get(0).asText());
}
@Test
void workersEndpointReturnsRegistryRosterWithLiveStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt"));
HttpResponse<String> spawn = req(port, "POST", "/workers?worktree=true&ticket=cb-304");
assertEquals(201, spawn.statusCode());
JsonNode spawned = mapper.readTree(spawn.body());
String paneId = spawned.get("paneId").asText();
HttpResponse<String> res = req(port, "GET", "/workers");
assertEquals(200, res.statusCode());
JsonNode workers = mapper.readTree(res.body()).get("workers");
assertEquals(1, workers.size());
JsonNode w = workers.get(0);
assertEquals(spawned.get("terminalId").asText(), w.get("sessionId").asText());
assertEquals(paneId, w.get("paneId").asText());
assertEquals("ltms-local", w.get("profile").asText());
assertEquals("spawning", w.get("state").asText());
assertTrue(w.has("worktree"), "worktree-backed session exposes worktree");
assertTrue(w.has("branch"), "worktree-backed session exposes branch");
assertEquals("unknown", w.get("liveStatus").asText(),
"liveStatus is unknown when herdr has no matching pane");
}
@Test
void spawnWithACwdParamRootsTheWorkerThere() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
assertEquals("/tmp/proj", params(herdr, "tab.create").get("cwd"), "the worker starts in cwd");
}
@Test
void spawnWithAnUnknownProfileIs400() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers?profile=nope");
assertEquals(400, res.statusCode());
assertEquals("unknown_profile", mapper.readTree(res.body()).get("error").asText());
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
}
@Test
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
// A space labelled "bridged-workers" already exists → no second workspace.create.
FakeHerdr herdr = new FakeHerdr().withWorkspace("w9", "bridged-workers");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers").statusCode());
assertFalse(herdr.called("workspace.create"), "existing worker space must be reused, not recreated");
assertTrue(herdr.called("tab.create"), "a fresh tab is still created for the worker");
}
@Test
@SuppressWarnings("unchecked")
void spawnRetriesUnderAFreshNameWhenAgentNameTaken() throws Exception {
// herdr rejects a duplicate agent name; the service must bump and retry.
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers").statusCode());
List<String> names = herdr.calls.stream()
.filter(c -> c.method().equals("agent.start"))
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
.toList();
assertEquals(3, names.size(), "2 rejected + 1 success");
assertEquals(3, Set.copyOf(names).size(), "each attempt must use a distinct name");
}
@Test
void spawnClosesTheCreatedTabWhenTheWorkerNeverStarts() throws Exception {
// Every agent.start attempt is rejected → spawn fails; the tab we created must not leak.
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(99);
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(500, req(port, "POST", "/workers").statusCode());
assertTrue(herdr.called("tab.create"), "a tab was created before the failed start");
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"), "orphaned tab must be closed");
}
@Test
void spawnWorkerRejectsOffAllowlistBaseUrlAndNeverTouchesHerdr() throws Exception {
FakeHerdr herdr = new FakeHerdr();
// base_url points at the subscription — guard must block before any herdr call.
int port = start(herdr, "https://api.anthropic.com", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers");
assertEquals(403, res.statusCode());
assertEquals("subscription_boundary", mapper.readTree(res.body()).get("error").asText());
assertFalse(herdr.called("agent.start"), "guard must stop the spawn before herdr");
assertFalse(herdr.called("workspace.create"), "guard must stop before provisioning a space");
}
@Test
void panePlacementSplitsFocusedTabWithoutADedicatedSpace() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "pane");
assertEquals(201, req(port, "POST", "/workers").statusCode());
Map<String, Object> start = params(herdr, "agent.start");
assertFalse(start.containsKey("tab_id"), "pane placement must not target a tab");
assertFalse(herdr.called("workspace.create"), "pane placement uses no dedicated space");
assertFalse(herdr.called("tab.create"));
}
@Test
void stopWorkerClosesPaneAndItsTab() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertTrue(herdr.called("pane.close"));
// Tab resolved from the pane (pane.get), then closed.
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"));
}
@Test
void messageReturnsTheWorkersStructuredReply() throws Exception {
// CB-104 (option C): the blocking send resolves on the worker's bridge_reply, not a scrape.
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
var send = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
try { return postMessage(port, "{\"content\":\"review this\",\"timeoutMs\":4000}"); }
catch (Exception e) { throw new RuntimeException(e); }
});
// Give the background send thread time to open its rendezvous waiter (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
Thread.sleep(200);
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"LGTM ship it\"}");
assertEquals(200, reply.statusCode());
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
assertEquals(200, res.statusCode());
assertEquals("LGTM ship it", mapper.readTree(res.body()).get("reply").asText());
// (injection via agent.prompt is covered deterministically by the timeout-working test)
}
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// CB-107 fire-and-poll: wait:false returns a ticket immediately; the result is polled.
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
assertEquals(202, accepted.statusCode());
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
assertFalse(ticket.isBlank(), "an async send must return a ticket");
// Give the background async send thread time to open its rendezvous waiter.
Thread.sleep(200);
HttpResponse<String> reply = postJson(port, "/sessions/term_a/reply", "{\"content\":\"async LGTM\"}");
assertEquals(200, reply.statusCode());
// Polling the ticket now reports the finished delegation and its reply.
JsonNode task;
long deadline = System.currentTimeMillis() + 3000;
do {
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
if ("done".equals(task.path("phase").asText())) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
assertEquals("done", task.get("phase").asText());
assertEquals("async LGTM", task.get("reply").asText());
assertEquals("reply", task.get("replySource").asText());
}
@Test
void pollUnknownTicketIs404() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/tasks/task-999");
assertEquals(404, res.statusCode());
assertEquals("unknown_ticket", mapper.readTree(res.body()).get("error").asText());
}
@Test
void replyWithNoPendingSendQueuesInsteadOfConflict() throws Exception {
// CB-307: a reply with no open send now queues in the inbox, not a 409 conflict.
int port = startHealthy();
HttpResponse<String> res = postJson(port, "/sessions/term_a/reply", "{\"content\":\"orphan\"}");
assertEquals(200, res.statusCode());
// The queued reply is drainable.
HttpResponse<String> drain = req(port, "GET", "/sessions/term_a/replies");
assertEquals(200, drain.statusCode());
JsonNode body = mapper.readTree(drain.body());
assertEquals(1, body.get("replies").size());
assertEquals("orphan", body.get("replies").get(0).get("content").asText());
}
@Test
void drainRepliesReturnsEmptyForNoReplies() throws Exception {
int port = startHealthy();
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/replies");
assertEquals(200, res.statusCode());
assertEquals(0, mapper.readTree(res.body()).get("replies").size());
}
@Test
void messageTimesOutQueuedWhenWorkerNeverInjectable() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("working"); // never injectable → never delivered
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
assertEquals(202, res.statusCode());
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
assertFalse(herdr.called("agent.prompt"), "no injection while the worker is mid-turn");
}
@Test
void messageTimesOutWorkingWhenDeliveredButNoReply() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // delivered, but nobody replies
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
assertEquals(202, res.statusCode());
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
assertTrue(herdr.called("agent.prompt"), "message was injected");
}
@Test
void messageRejectsBlankContent() throws Exception {
int port = startHealthy();
assertEquals(400, postMessage(port, "{}").statusCode());
}
@Test
void sessionStatusReportsLiveAgentStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "GET", "/sessions/term_a/status");
assertEquals(200, res.statusCode());
assertEquals("blocked", mapper.readTree(res.body()).get("status").asText());
}
@Test
void sessionStatusReportsReadinessFromMcpPresence() throws Exception {
int port = start(new FakeHerdr(), "http://gx00.gw:8000", Set.of("gx00.gw"));
// Not yet seen on the bridge MCP → not ready.
assertFalse(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
// Worker connects its MCP client → available.
presence.markPresent("term_a");
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
}
@Test
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "pane");
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertTrue(herdr.called("pane.close"));
assertFalse(herdr.called("tab.close"), "pane placement owns no tab to close");
assertFalse(herdr.called("pane.get"), "no tab resolution in pane placement");
}
@Test
void stopNeverClosesATabThatHoldsOtherPanes() throws Exception {
// The worker's tab has 2 panes (e.g. a pane-placement worker sharing a user tab).
FakeHerdr herdr = new FakeHerdr().withWorkerTabPaneCount(2);
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertTrue(herdr.called("pane.close"), "the worker's own pane is still closed");
assertFalse(herdr.called("tab.close"), "must not close a tab that holds the user's other panes");
}
@Test
void stopReportsFailureWhenPaneCloseFailsForARealReason() throws Exception {
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("herdr_busy");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
// A genuine teardown failure must surface, not be reported as a successful 204.
assertEquals(500, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertFalse(herdr.called("tab.close"), "tab is not removed when the pane close failed");
}
@Test
void stopToleratesAnAlreadyGonePane() throws Exception {
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("pane_not_found");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
// Already-gone is success; the (now-empty) tab is still cleaned up.
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertTrue(herdr.called("tab.close"));
}
}
@@ -1,130 +0,0 @@
package dev.ltms.bridged.session;
import java.util.Collections;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
public final class FakeWorktrees implements Worktrees {
public record AddCall(String repoRoot, String branch, String baseRef) {
}
public record RemoveCall(String repoRoot, String worktreePath) {
}
public record OverlayCall(String repoRoot, String worktreePath,
List<String> requested, List<String> copied, List<String> skipWorktree) {
}
public record RepoRootCall(String cwd) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private volatile RuntimeException addFailure;
private volatile String repoRoot = "/repo";
private volatile String prefix = "/worktrees";
public FakeWorktrees withRepoRoot(String root) {
this.repoRoot = root;
return this;
}
public FakeWorktrees withPrefix(String prefix) {
this.prefix = prefix;
return this;
}
/** Paths that exist in the primary repo and will be copied to the worktree. */
public FakeWorktrees exists(String... paths) {
Collections.addAll(existingPaths, paths);
return this;
}
/** Paths that exist AND are tracked, so overlayParity should --skip-worktree them. */
public FakeWorktrees track(String... paths) {
exists(paths);
Collections.addAll(trackedPaths, paths);
return this;
}
/** Make subsequent {@link #add} calls throw (simulates git worktree add failure). */
public FakeWorktrees failAdd(String message) {
this.addFailure = new WorktreeException(message);
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
addCalls.add(new AddCall(repoRoot, branch, baseRef));
if (addFailure != null) {
throw addFailure;
}
// The branch already carries a unique nonce, so the derived path is distinct per acquire
// without an extra counter — keep it a pure function of the branch the test can predict.
return prefix + "/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
removeCalls.add(new RemoveCall(repoRoot, worktreePath));
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
List<String> copied = new java.util.ArrayList<>();
List<String> skipped = new java.util.ArrayList<>();
for (String rel : overlay) {
if (!existingPaths.contains(rel)) {
continue; // missing source is silently skipped
}
copied.add(rel);
if (trackedPaths.contains(rel)) {
skipped.add(rel);
}
}
overlayCalls.add(new OverlayCall(repoRoot, worktreePath, List.copyOf(overlay),
List.copyOf(copied), List.copyOf(skipped)));
}
@Override
public String repoRoot(String cwd) {
repoRootCalls.add(new RepoRootCall(cwd));
return repoRoot;
}
public List<AddCall> addCalls() {
return List.copyOf(addCalls);
}
public List<RemoveCall> removeCalls() {
return List.copyOf(removeCalls);
}
public List<OverlayCall> overlayCalls() {
return List.copyOf(overlayCalls);
}
public List<RepoRootCall> repoRootCalls() {
return List.copyOf(repoRootCalls);
}
public AddCall lastAdd() {
return addCalls.isEmpty() ? null : addCalls.getLast();
}
public RemoveCall lastRemove() {
return removeCalls.isEmpty() ? null : removeCalls.getLast();
}
public OverlayCall lastOverlay() {
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
}
}
@@ -1,119 +0,0 @@
package dev.ltms.bridged.session;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-525 acceptance test for tool-surface isolation. This is one of the few tests that drives real
* {@code git} — the behaviour under test is precisely what {@link GitWorktrees} does to a checkout,
* so a fake would assert nothing. Everything happens inside a {@link TempDir} throwaway repo.
*/
class GitWorktreesTest {
/** A project MCP config with servers in it — what this repo actually commits. */
private static final String WITH_SERVERS = """
{
"mcpServers": {
"jetbrains": { "type": "sse", "url": "http://localhost:64342/sse" }
}
}
""";
private static Path initRepo(Path dir) throws Exception {
Files.createDirectories(dir);
git(dir, "init", "-q", "-b", "main");
git(dir, "config", "user.email", "test@example.invalid");
git(dir, "config", "user.name", "Test");
Files.writeString(dir.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(dir.resolve("README.md"), "seed\n");
git(dir, "add", ".mcp.json", "README.md");
git(dir, "commit", "-q", "-m", "seed");
return dir;
}
private static void git(Path cwd, String... args) throws Exception {
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
cmd.addAll(List.of(args));
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: " + String.join(" ", cmd));
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
}
/** Pending changes to {@code .mcp.json} in {@code cwd}, empty when git considers it unmodified. */
private static String mcpStatus(Path cwd) throws Exception {
Process p = new ProcessBuilder("git", "status", "--porcelain", "--", ".mcp.json")
.directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git status timed out");
return out;
}
/**
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
* through the primary's IDE servers edits the primary's tree while building its own.
*/
@Test
void aProvisionedWorktreeInheritsNoMcpServers(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-a", "HEAD");
Path mcp = Path.of(wt).resolve(".mcp.json");
assertTrue(Files.exists(mcp), ".mcp.json must still exist — present and explicitly empty");
String body = Files.readString(mcp);
assertFalse(body.contains("jetbrains"), "worktree inherited the primary's MCP servers:\n" + body);
assertTrue(body.replaceAll("\\s+", "").contains("\"mcpServers\":{}"),
"expected an explicitly empty server map, got:\n" + body);
}
/** Neutralizing must not look like work in progress, or a worker would commit it into its PR. */
@Test
void theNeutralizedConfigIsNotAPendingLocalModification(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-b", "HEAD");
assertEquals("", mcpStatus(Path.of(wt)),
"the neutralized .mcp.json shows as modified — --skip-worktree did not take");
}
/** Isolation is the worktree's business only; the primary's own checkout must be untouched. */
@Test
void thePrimaryCheckoutIsLeftAlone(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-525-c", "HEAD");
assertEquals(WITH_SERVERS, Files.readString(repo.resolve(".mcp.json")),
"the primary's .mcp.json was rewritten — isolation reached out of the worktree");
}
/** A repo that commits no {@code .mcp.json} still gets one, so nothing can be inherited later. */
@Test
void aRepoWithoutAnMcpConfigStillGetsANeutralOne(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-d", "HEAD");
// Untracked is the normal case here, so the --skip-worktree branch must be skipped rather
// than run and fail: `update-index --skip-worktree` on an unknown path exits non-zero.
String body = Files.readString(Path.of(wt).resolve(".mcp.json"));
assertTrue(body.replaceAll("\\s+", "").contains("\"mcpServers\":{}"), body);
}
}
@@ -1,496 +0,0 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-301 / CB-303 acceptance tests for the authoritative session registry, one-shot lifecycle FSM,
* and configurable lifecycle limits (idle TTL, context cap, drain).
* No live herdr — everything runs against the same {@link FakeHerdr} the rest of the project uses.
*/
class SessionManagerTest {
private SessionManager sessionManager(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock) {
return sessionManager(herdr, clock, 0);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
return sessionManager(herdr, clock, contextCap, false);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap,
boolean clearAfterTurn) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap, clearAfterTurn);
}
@Test
void primaryContactWithNoTerminalIsNotAReadinessSignal() {
// The MCP context extractor calls presence.markPresent(p.terminal()) on EVERY request,
// and the primary's terminal is null — the presence bridge must treat that as a no-op,
// not feed it into the READY transition (which NPEd on the first real primary contact).
SessionManager sessions = sessionManager(new FakeHerdr());
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null));
assertDoesNotThrow(() -> sessions.asPresence().markPresent(" "));
}
@Test
void acquireRegistersSpawningSessionWithDistinctPaneId() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
WorkerSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
assertEquals(WorkerSession.State.SPAWNING, a.state(), "fresh session starts spawning");
assertEquals("ltms-local", a.profile());
assertEquals("/work/a", a.cwd(), "explicit requested cwd is recorded");
assertEquals("term_primary", a.ownerTerminal());
assertTrue(a.spawnedAtNanos() > 0);
assertNotNull(a.paneId());
assertNotNull(a.terminalId());
assertNotEquals(a.paneId(), b.paneId(), "no pane reuse");
assertNotEquals(a.terminalId(), b.terminalId(), "no terminal reuse");
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and BridgeMcp's context extractor
// forwards that null into markPresent on EVERY MCP call. It only reached the registry scan
// once a session existed, so this NPE'd the primary's second spawn while the first passed.
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null),
"the primary's null terminal must not blow up an unrelated tool call");
assertDoesNotThrow(() -> sessions.onDelivered(null));
assertDoesNotThrow(() -> sessions.onTurnComplete(null));
assertDoesNotThrow(() -> sessions.onTurnFailed(null));
assertEquals(WorkerSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
"and must not transition any registered session");
}
@Test
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertEquals(WorkerSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"MCP presence moves SPAWNING → READY");
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
sessions.onDelivered(terminal);
assertEquals(WorkerSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
"delivery moves READY → BUSY");
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
"turn completion moves BUSY → DONE");
}
@Test
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", null);
String paneId = session.paneId();
sessions.release(paneId);
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
assertTrue(sessions.roster().isEmpty(), "released session is no longer in the roster");
assertDoesNotThrow(() -> sessions.release(paneId), "a second release is harmless");
}
@Test
void onTurnFailedMovesSessionToFailed() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnFailed(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
}
@Test
void recycleProducesNewPaneIdAndOldOneIsGone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String oldPane = oldSession.paneId();
String oldTerminal = oldSession.terminalId();
WorkerSession fresh = sessions.recycle(oldPane);
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
assertEquals(oldSession.profile(), fresh.profile(), "profile is preserved");
assertEquals(oldSession.cwd(), fresh.cwd(), "cwd is preserved");
assertEquals(oldSession.ownerTerminal(), fresh.ownerTerminal(), "owner is preserved");
assertTrue(sessions.get(oldPane).isEmpty(), "old pane is deregistered");
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
// The old session was the first spawn → pane w9:pRoot_1 (CB-519: the registry key is the
// uuid id, so teardown is asserted on the real pane coordinate).
long paneCloseCount = herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, paneCloseCount, "the old worker was torn down");
}
@Test
void rosterReflectsAcquiredMinusReleased() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
WorkerSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
assertEquals(2, sessions.roster().size());
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(a.paneId())));
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(b.paneId())));
sessions.release(a.paneId());
assertEquals(1, sessions.roster().size());
assertEquals(b.paneId(), sessions.roster().getFirst().paneId());
}
// --- CB-303 lifecycle limits ----------------------------------------------------
@Test
void reapIdleDoesNothingWhenNoSessions() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L);
assertEquals(0, sessions.reapIdle(10));
assertTrue(sessions.roster().isEmpty());
}
@Test
void readySessionPastIdleTtlIsReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 11;
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
assertTrue(sessions.get(session.paneId()).isEmpty(), "reaped session is removed from registry");
assertTrue(herdr.called("pane.close"), "reaped session tears the pane down");
}
@Test
void readySessionWithinIdleTtlSurvives() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 5;
assertEquals(0, sessions.reapIdle(10), "READY session within TTL is not reaped");
assertEquals(WorkerSession.State.READY,
sessions.get(session.paneId()).orElseThrow().state(),
"READY session survives");
}
@Test
void busySessionPastIdleTtlIsNotReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
clock[0] = 100;
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
assertEquals(WorkerSession.State.BUSY,
sessions.get(session.paneId()).orElseThrow().state(),
"BUSY session remains");
}
@Test
void doneSessionPastIdleTtlIsReaped() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
clock[0] = 21;
assertEquals(1, sessions.reapIdle(20), "DONE session past TTL is reaped");
assertTrue(sessions.get(session.paneId()).isEmpty(), "DONE session is removed");
}
@Test
void reapIdleReturnsCorrectCountAndSkipsBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
clock[0] = 50;
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "READY session is gone");
assertEquals(WorkerSession.State.BUSY,
sessions.get(busy.paneId()).orElseThrow().state(),
"BUSY session is still registered");
}
@Test
void contextCapDisabledSessionSurvivesMultipleTurns() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
long releaseCloseCount = paneCloseCallsFor(herdr, "w9:pRoot_1"); // the real pane coordinate
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
}
@Test
void contextCapTwoReleasesAfterSecondComplete() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 2);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE,
sessions.get(session.paneId()).orElseThrow().state(),
"first turn completes without release");
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
assertTrue(sessions.roster().isEmpty(), "released session leaves roster");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"forced release tears the worker pane down exactly once");
}
@Test
void clearAfterTurnResetsContextWithoutDoubleCountingTheTurn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, true);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
assertTrue(sessions.onTurnCompleteWithPostAction(session.terminalId()));
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(1, updated.turnCount(), "the reset is housekeeping, not a second delegation");
assertEquals(List.of("/clear"), promptTexts(herdr));
}
@Test
void contextCapReleaseWinsOverClearAfterTurn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 1, true);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
assertFalse(sessions.hasPostTurnAction(session.terminalId()),
"a session at its cap will be released, not reset for reuse");
assertFalse(sessions.onTurnCompleteWithPostAction(session.terminalId()));
assertTrue(sessions.get(session.paneId()).isEmpty());
assertTrue(promptTexts(herdr).isEmpty(), "never send /clear into a worker being torn down");
}
@Test
void clearAfterTurnFalsePreservesCompletionWithoutAControlPrompt() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, false);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onTurnComplete(session.terminalId());
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state());
assertTrue(promptTexts(herdr).isEmpty());
}
@Test
void drainAllReleasesBusyAndReadySessionsAndWaitsForBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(sessions.roster().isEmpty(), "drain clears the roster");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "ready session is released");
assertTrue(sessions.get(busy.paneId()).isEmpty(), "busy session is released after timeout");
// ready is the first spawn → pane w9:pRoot_1, busy the second → w9:pRoot_2 (FakeHerdr order).
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"ready worker pane is torn down");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_2"),
"busy worker pane is torn down");
}
private static long paneCloseCallsFor(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> "agent.prompt".equals(c.method()))
.map(c -> String.valueOf(((Map<?, ?>) c.params()).get("text")))
.toList();
}
// --- CB-306 spawn-readiness gate: no half-registered session on timeout ----------------
@Test
void acquireThrowsPeerUnreachableWhenGateTimesOutAndRegistersNoSession() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // never becomes injectable
long[] clock = {0};
// Gate-enabled launcher (1 ms timeout + no-op sleeper that advances clock past deadline)
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
1, () -> clock[0], () -> clock[0] += 10);
SessionManager sessions = new SessionManager(workers, new GitWorktrees(), () -> 0L, 0);
assertThrows(PeerUnreachableException.class,
() -> sessions.acquire("ltms-local", null, "/caller", "term_primary"),
"acquire must throw PeerUnreachableException when spawn times out");
// No half-registered session — the error happened inside spawn, before
// SessionManager could put() anything into the registry.
assertTrue(sessions.roster().isEmpty(),
"no session is registered when spawn times out (roster empty)");
}
// --- CB-516: release must notify, so a blocked send can be failed --------------------------
@Test
void releaseNotifiesTheListenerWithTheReleasedTerminal() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
sessions.release(s.paneId());
assertEquals(java.util.List.of(s.terminalId()), released,
"every teardown path funnels through release, so one hook must see the terminal");
}
@Test
void releasingAnUnknownPaneNotifiesNobody() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.release("w9:p404"); // idempotent teardown of something already gone
assertTrue(released.isEmpty(), "no session removed ⇒ no send was waiting on it");
}
@Test
void aThrowingReleaseListenerDoesNotBlockTheTeardown() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
sessions.onRelease(_ -> {
throw new IllegalStateException("listener blew up");
});
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a listener failure must never prevent the teardown it is reacting to");
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
}
}
@@ -1,227 +0,0 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-301-ext acceptance tests for worktree provisioning and config-parity overlay.
* No live git — every Worktrees call is handled by {@link FakeWorktrees} and every herdr
* call by {@link FakeHerdr}, matching the project's fake-based test style.
*/
class WorktreeSessionManagerTest {
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
private static String startCwd(FakeHerdr herdr) {
@SuppressWarnings("unchecked")
// Protocol 19: the worker's cwd rides on pane creation (tab.create), not agent.start.
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("tab.create").params();
Object cwd = start.get("cwd");
return cwd == null ? null : cwd.toString();
}
@Test
void sharedTreeAcquireMakesNoWorktreesCallsAndRecordsNullWorktree() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
assertTrue(worktrees.addCalls().isEmpty(), "shared-tree acquire never adds a worktree");
assertTrue(worktrees.repoRootCalls().isEmpty(), "shared-tree acquire never resolves a repo root");
assertTrue(worktrees.overlayCalls().isEmpty(), "shared-tree acquire never overlays parity");
assertNull(s.worktree(), "shared-tree session has no worktree");
assertNull(s.branch(), "shared-tree session has no branch");
assertEquals("/caller/proj", s.cwd(), "shared-tree cwd is the caller's cwd");
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
}
@Test
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-999", null));
assertEquals(1, worktrees.addCalls().size(), "one worktree was added");
FakeWorktrees.AddCall add = worktrees.lastAdd();
assertNotNull(add);
assertEquals("/repo", add.repoRoot());
assertTrue(add.branch().startsWith("worker/cb-999-"), "branch is worker/<slug>-<nonce>: " + add.branch());
assertNull(add.baseRef(), "null baseRef is passed through (HEAD default)");
String expectedPath = "/wt/" + add.branch().replace('/', '_');
assertEquals(expectedPath, s.worktree(), "session records the returned worktree path");
assertEquals(add.branch(), s.branch(), "session records the branch");
assertEquals(expectedPath, startCwd(herdr), "spawn receives the worktree path as cwd");
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
@Test
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".envrc")
.exists(".claude/settings.local.json");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-888", null));
assertEquals(1, worktrees.overlayCalls().size());
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay);
assertEquals("/repo", overlay.repoRoot());
assertEquals(List.of(".claude/settings.local.json", ".env", ".envrc"),
overlay.requested(), "default parity overlay is used when unset");
assertFalse(overlay.requested().contains(".mcp.json"),
"CB-525: replicating the primary's MCP config gives a worker the primary's IDE "
+ "servers, which navigate its edits out of its own worktree");
assertEquals(List.of(".claude/settings.local.json", ".envrc"), overlay.copied(),
"existing paths are copied; missing paths are skipped");
assertEquals(List.of(".envrc"), overlay.skipWorktree(),
"tracked copied paths are --skip-worktree'd");
}
@Test
void releaseRemovesWorktreeButDoesNotDeleteBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-666", null));
String paneId = s.paneId();
sessions.release(paneId);
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
assertEquals(1, worktrees.removeCalls().size(), "worktree session triggers one remove");
FakeWorktrees.RemoveCall remove = worktrees.lastRemove();
assertNotNull(remove);
assertEquals("/repo", remove.repoRoot());
assertEquals(s.worktree(), remove.worktreePath());
// The fake records no branch-delete calls because Worktrees.remove only removes the checkout.
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
}
@Test
void sharedTreeReleaseMakesNoWorktreesCalls() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
sessions.release(s.paneId());
assertTrue(herdr.called("pane.close"), "release tears the worker pane down");
assertTrue(worktrees.removeCalls().isEmpty(), "shared-tree release never removes a worktree");
}
@Test
void failedWorktreeAddUnwindsWithoutRegisteringSessionOrSpawning() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().failAdd("worktree add failed");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
assertThrows(WorktreeException.class, () ->
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-555", null)));
assertEquals(0, sessions.size(), "failed acquire leaves no registry entry");
assertFalse(herdr.called("agent.start"), "spawn is never reached when add fails");
assertTrue(worktrees.removeCalls().isEmpty(), "no worktree was added, so none is removed");
}
@Test
void twoWorktreeAcquiresYieldDistinctBranchesAndPaths() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
WorkerSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
assertNotEquals(a.branch(), b.branch(), "branches are distinct");
assertNotEquals(a.worktree(), b.worktree(), "paths are distinct");
assertEquals(2, worktrees.addCalls().size());
assertEquals(2, sessions.roster().size());
}
/**
* CB-507 regression. A plain REST spawn supplies neither a requested nor a caller cwd
* ({@code BridgedApp} hardcodes {@code callerCwd = null}), and the worktree branch used to
* resolve the repo root from just those two — yielding {@code null}, which the real
* {@code GitWorktrees} turns into {@code git -C null} and an NPE out of {@code ProcessBuilder}
* (HTTP 500).
*
* <p>Note this asserts on the <em>recorded</em> cwd rather than expecting a throw:
* {@link FakeWorktrees#repoRoot} only records its argument and returns a canned root, so a
* null flows through the fake harmlessly. That permissiveness is precisely why the whole
* suite stayed green while the feature was broken in production — so the assertion has to be
* "a usable cwd was passed down", not "an exception was raised".
*/
@Test
void worktreeAcquireWithNoRequestedOrCallerCwdStillResolvesANonNullRepoRoot() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507", null));
assertFalse(worktrees.repoRootCalls().isEmpty(),
"repoRoot should have been called to resolve the repo root");
String cwd = worktrees.repoRootCalls().getFirst().cwd();
assertNotNull(cwd, "a null cwd here becomes `git -C null` and NPEs in the real GitWorktrees");
assertFalse(cwd.isBlank(), "a blank cwd is as unusable as a null one");
}
/**
* The same line carried a second, quieter bug: it never consulted the profile's configured
* {@code cwd:}, so a worktree spawn silently ignored a pinned per-profile working directory.
* Routing through {@code effectiveCwd} honours it.
*/
@Test
void worktreeAcquireHonoursTheProfileConfiguredCwd() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
// Argument order matters: configDir is the 4th parameter, cwd the 11th (after mcpUrl).
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
SessionManager sessions = new SessionManager(launcher, worktrees);
sessions.acquire("ltms-local", null, null, null, new WorktreeRequest("cb-507b", null));
assertEquals(1, worktrees.repoRootCalls().size());
assertEquals("/pinned/dir", worktrees.repoRootCalls().getFirst().cwd(),
"the profile's configured cwd must reach repoRoot, not be ignored");
}
}
@@ -1,785 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
class ClaudeCodeLauncherTest {
private ClaudeCodeLauncher service(FakeHerdr herdr, List<String> argv, String mcpUrl) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
argv, "tab", "bridged-workers", "worker: {profile} #{n}", mcpUrl, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
/** The {@code args} of the last agent.start — protocol 19: everything after the executable. */
@SuppressWarnings("unchecked")
private List<String> spawnedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
}
@Test
void appendsBridgeMcpAndReplyCharterWhenMcpUrlSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), "http://127.0.0.1:8765/mcp").spawn();
List<String> args = spawnedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
assertTrue(args.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
"inline bridge MCP config present");
assertTrue(args.contains("--append-system-prompt"));
assertTrue(args.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
}
@Test
void startRetriesWhileTheSeedShellBoots() {
// tab.create returns before the seed shell reaches its prompt; herdr refuses agent.start
// into a not-ready pane with agent_pane_busy. The launcher must wait it out, not fail.
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
Map.of("ltms-local", new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null)),
"ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn succeeds once the shell is ready");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"two busy rejections, then the successful start");
}
@Test
void startResolvesTheExecutableFromKindAndDropsArgvZero() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
Map<?, ?> start = (Map<?, ?>) herdr.lastCall("agent.start").params();
assertEquals("claude", start.get("kind"), "herdr launches the canonical executable by kind");
// CB-533: the shared fixture pins model "coder", so the model flag is the whole args list.
// What this test guards is that argv[0] is NOT repeated — herdr supplies it from `kind`.
assertEquals(List.of("--model", "coder"), start.get("args"),
"the configured executable is not repeated in args");
}
@Test
void noBridgeFlagsWhenMcpUrlAbsent() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude", "--verbose"), null).spawn();
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--mcp-config"), "no bridge mount without mcpUrl");
assertFalse(args.contains("--append-system-prompt"), "no reply charter without mcpUrl");
// CB-533: the model flag is independent of the MCP mount — pinning the model is not part of
// "mount the bridge", so an unmounted worker still runs the model its profile names.
assertEquals(List.of("--verbose", "--model", "coder"), args,
"the operator's own args are preserved, in order, ahead of the model flag");
}
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
BridgedConfig.Worker gx10 = new BridgedConfig.Worker("gx10", "http://gx10.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
BridgedConfig.Worker ollama = new BridgedConfig.Worker("ollama", "http://ollama.ltms.dev", null,
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
Map.of("gx10", gx10, "ollama", ollama), "gx10", _ -> "tok");
}
@Test
void spawnPicksTheNamedProfilesBaseUrl() {
FakeHerdr herdr = new FakeHerdr();
multiProfile(herdr).spawn("ollama");
assertEquals("http://ollama.ltms.dev", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the named profile's base_url");
}
@Test
void spawnRejectsAnUnknownProfile() {
try (FakeHerdr herdr = new FakeHerdr()) {
assertThrows(IllegalArgumentException.class, () -> multiProfile(herdr).spawn("nope"));
}
}
@SuppressWarnings("unchecked")
private static String startCwd(FakeHerdr herdr) {
// Protocol 19: the worker's cwd is set at pane creation (tab.create), where the seed
// shell — which the agent starts into — is rooted.
Object v = ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("cwd");
return v == null ? null : v.toString();
}
@Test
void requestedCwdRootsTheWorker() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", "/work/proj", "/caller/home");
assertEquals("/work/proj", startCwd(herdr), "an explicit spawn cwd wins over everything");
}
@Test
void profileConfigCwdBeatsTheCallerCwd() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"w #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local", _ -> null);
svc.spawn("ltms-local", null, "/caller/home");
assertEquals("/pinned/dir", startCwd(herdr), "a profile-pinned cwd overrides the caller's");
}
@Test
void inheritsTheCallerCwdWhenNothingElseIsSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", null, "/primary/project");
assertEquals("/primary/project", startCwd(herdr), "no explicit/config cwd → inherit the primary's");
}
// --- CB-302 git-forge token injection (worker checkpoint grant) ------------
/** Protocol 19: the worker's env is injected at pane creation (tab.create), not agent.start. */
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
@Test
void injectsForgeTokenAndHostWhenProfileGrantsIt() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null); // parityOverlay null; gitHostEnv null → defaults to GITEA_HOST
Function<String, String> host = name -> switch (name) {
case "GITEA_ACCESS_TOKEN" -> "gt-secret";
case "GITEA_HOST" -> "git.ltms.dev";
default -> null;
};
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl", host).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("gt-secret", env.get("GITEA_TOKEN"), "the forge token is injected for a granting profile");
assertEquals("git.ltms.dev", env.get("GITEA_HOST"), "the paired forge host rides along with the token");
}
@Test
void noForgeTokenWhenProfileDoesNotGrantIt() {
FakeHerdr herdr = new FakeHerdr();
// gitTokenEnv unset (12-arg ctor); the env would resolve a token if asked, proving the gate
// is the profile config, not a missing env var.
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs"), "tab", "bridged-workers", "w #{n}",
null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local",
_ -> "would-be-secret").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("GITEA_TOKEN"), "no forge token when the profile does not opt in");
assertNull(env.get("GITEA_HOST"), "no forge host without a granted token");
}
// --- CB-117 orphan reap: the pure predicate --------------------------------
@Test
void isForeignWorkerMatchesOurSchemeWithANonSelfNonce() {
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-ollama-be09c2-2", "aaaaaa"),
"a bridge worker name with a different nonce is a prior daemon's orphan");
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-gx10-4127af-11", "aaaaaa"),
"profile and multi-digit seq are still parsed; foreign nonce ⇒ reap");
}
@Test
void isForeignWorkerSparesOurOwnLiveWorkersAndNonWorkers() {
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-abcdef-3", "abcdef"),
"a worker with THIS process's nonce is ours and live — never reap it");
assertFalse(ClaudeCodeLauncher.isForeignWorker(null, "abcdef"), "an unnamed agent is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude", "abcdef"), "a bare kind name is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("my-repl", "abcdef"), "a user's own label is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-XYZ123-2", "abcdef"),
"a non-hex nonce does not match our scheme");
}
// --- CB-117 orphan reap: the wiring through stop() -------------------------
private static long paneCloseCount(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
@Test
void reapsAForeignOrphanButSparesOurOwnWorkerAndUserSessions() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-be09c2-2", "term_orphan", "wQ:pF", "wQ:t8") // prior daemon's leak
.withAgent("claude-gx10-" + svc.nameNonce() + "-1", "term_mine", "wQ:pMine", "wQ:tMine"); // ours, live
// (the fake's default unnamed term_a stands in for a user's own Claude session)
int reaped = svc.reapOrphanWorkers();
assertEquals(1, reaped, "exactly the one foreign-nonce orphan is reaped");
assertEquals(1, paneCloseCount(herdr, "wQ:pF"), "the orphan's pane is closed");
assertEquals(0, paneCloseCount(herdr, "wQ:pMine"), "our own live worker's pane is left running");
assertEquals(0, paneCloseCount(herdr, "w2:p7"), "a user's own session is never touched");
assertTrue(herdr.called("tab.close"), "the orphan's now-empty dedicated tab is closed too");
}
@Test
void reapCountsAnAlreadyGoneOrphanAsReaped() {
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("pane_not_found");
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-0d856d-3", "term_gone", "wQ:pS", "wQ:tD");
assertEquals(1, svc.reapOrphanWorkers(),
"a pane that vanished between list and close is a successful reap, not a failure");
}
@Test
void reapIsSkippedWhenHerdrCannotBeListed() {
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.list throws
assertEquals(0, multiProfile(herdr).reapOrphanWorkers(), "a listing failure reaps nothing and does not throw");
}
// --- PeerHandle indirection ----------------------------------------------------------------
@Test
void spawnReturnsPeerHandleWithHostUniqueOpaqueId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn must return a non-null handle");
// CB-519: id() is a host-unique opaque UUID, decoupled from the herdr pane coordinate.
assertNotEquals("w9:pRoot_1", handle.id(),
"handle.id() must NOT be the herdr pane id");
assertDoesNotThrow(() -> UUID.fromString(handle.id()),
"handle.id() must be a UUID: " + handle.id());
}
@Test
void spawnReturnsPeerHandleWithCorrectTerminalId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, "/caller"));
assertEquals("term_new_1", handle.terminalId(), "handle.terminalId() must equal the agent's terminalId");
}
@Test
void capabilitiesIncludeMidTurnAskWorktreeOrphanReap() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.MID_TURN_ASK), "every Claude Code peer supports mid-turn ask");
assertTrue(caps.contains(Capability.WORKTREE), "every CLI peer supports worktree cwd");
assertTrue(caps.contains(Capability.ORPHAN_REAP), "every herdr launcher supports orphan reap");
assertTrue(caps.contains(Capability.CONTEXT_RESET), "Claude Code supports /clear");
}
@Test
void clearContextUsesTheClaudeCommandThroughTheOwningHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("claude"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertTrue(svc.clearContext(handle.id()));
Map<?, ?> prompt = (Map<?, ?>) herdr.lastCall("agent.prompt").params();
assertEquals("/clear", prompt.get("text"));
assertEquals("w9:pRoot_1", prompt.get("target"));
}
@Test
void capabilitiesIncludeSelfPrWhenProfileHasGitToken() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl",
_ -> "tok");
assertTrue(svc.capabilities().contains(Capability.SELF_PR),
"a profile with a git token grants SELF_PR");
}
@Test
void capabilitiesExcludeSelfPrWhenNoGitToken() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
assertFalse(svc.capabilities().contains(Capability.SELF_PR),
"no git token profile → no SELF_PR capability");
}
@Test
void effectiveCwdViaSpawnRequestMatchesExistingResolution() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
String cwd = svc.effectiveCwd(new SpawnRequest("ltms-local", "/work/proj", "/caller/home"));
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
// --- CB-547a: durable session identity (mint / resume / no-identity legacy) -----------------
@Test
void freshSpawnMintsASessionIdAndPassesTheName() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null, "my-session", null));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0 && flag + 1 < args.size(), "--session-id present: " + args);
String minted = args.get(flag + 1);
assertDoesNotThrow(() -> UUID.fromString(minted), "--session-id is a valid UUID: " + minted);
assertEquals("my-session", args.get(args.indexOf("-n") + 1), "the logical name rides as -n");
assertEquals(minted, handle.agentSessionId(),
"the resume handle is the minted id, known before the agent has written anything");
assertEquals("my-session", handle.sessionName(), "the handle carries the logical name");
}
@Test
void resumeSpawnPassesDashRAndNeverASessionId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null, "my-session", "cb-resume-1"));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "--session-id must NOT be passed on a resume (conflicts with -r)");
assertEquals("cb-resume-1", args.get(args.indexOf("-r") + 1), "-r carries the prior session id");
assertEquals("cb-resume-1", handle.agentSessionId(), "a resume adopts the prior id as its own");
assertEquals("my-session", handle.sessionName(), "the logical name survives a resume");
}
@Test
void noIdentitySpawnKeepsTheLegacyArgvAndCarriesNoSessionHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "no identity → no --session-id");
assertFalse(args.contains("-n"), "no identity → no -n");
assertFalse(args.contains("-r"), "no identity → no -r");
assertNull(handle.agentSessionId(), "no identity → no resume handle");
assertNull(handle.sessionName(), "no identity → no logical name");
}
@Test
void capabilitiesIncludeSessionNameAndSessionResume() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.SESSION_NAME), "Claude Code surfaces the bridge's logical name (-n)");
assertTrue(caps.contains(Capability.SESSION_RESUME), "Claude Code can relaunch onto a prior conversation (-r)");
}
@Test
void profilesViaPeerLauncherMatchesExistingApi() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals(Set.of("gx10", "ollama"), svc.profiles(), "profiles() via PeerLauncher must match");
}
@Test
void defaultProfileViaPeerLauncherMatches() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals("gx10", svc.defaultProfile(), "defaultProfile() via PeerLauncher must match");
}
@Test
void stopViaPeerLauncherTearsDownByHandleId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(handle.id());
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
}
// --- CB-519: host-unique id, decoupled from the pane coordinate ------------------------------
@Test
void twoSpawnsOnTheSamePaneNeverCollideOnHostUniqueId() {
// Two spawns may be placed on the same herdr pane coordinate (e.g. a pane that was reused
// or re-reported after a restart); the host-unique id must not collide even then.
FakeHerdr herdr = new FakeHerdr().pinNextStarts(2, "term_shared", "w9:pShared");
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle a = svc.spawn(new SpawnRequest(null, null, null));
PeerHandle b = svc.spawn(new SpawnRequest(null, null, null));
assertNotEquals(a.id(), b.id(),
"two spawns on the same pane coordinate get distinct host-unique ids");
assertNotEquals("w9:pShared", a.id(), "id is not the pane coordinate");
assertNotEquals("w9:pShared", b.id(), "id is not the pane coordinate");
}
@Test
void stopResolvesTheHostUniqueIdToThePaneThatSpawnedIt() {
// CB-519: id() != paneId, so stop(id) must tear down the exact pane the id names — and no
// other live peer's pane.
FakeHerdr herdr = new FakeHerdr(); // deterministic panes w9:pRoot_1, w9:pRoot_2 per spawn
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle a = svc.spawn(new SpawnRequest(null, null, null));
PeerHandle b = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(b.id());
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_2"), "stop(b.id()) closes only b's pane");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"), "a's pane is untouched");
}
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
private static Map<String, BridgedConfig.Worker> workerConfigMap(String profile, String mcpUrl) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
profile, "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers",
"worker: {profile} #{n}", mcpUrl, null, null);
return Map.of(cfg.profile(), cfg);
}
@Test
void spawnWaitsUntilInjectableThenReturnsHandle() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // first status call sees UNKNOWN
long[] clock = {0};
boolean[] firstSleep = {true};
// The sleeper: advance the fake clock, and on the first call flip the
// agent status to IDLE so the next poll succeeds.
Runnable sleeper = () -> {
clock[0] += 300;
if (firstSleep[0]) {
herdr.agentStatus("idle");
firstSleep[0] = false;
}
};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], sleeper);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when worker becomes injectable");
assertNotEquals("w9:pRoot_1", handle.id(),
"handle id is a host-unique opaque id, not the started pane");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane.close when worker becomes injectable before timeout");
}
@Test
void spawnThrowsPeerUnreachableWhenNeverInjectableAndReapsPane() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().contains("1000"),
"exception message references the timeout: " + ex.getMessage());
assertTrue(clock[0] >= 1000, "fake clock advanced past the timeout: " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"pane was closed on timeout (no orphan left behind)");
}
@Test
void spawnReturnsImmediatelyWhenGateIsDisabled() {
FakeHerdr herdr = new FakeHerdr();
// The default 6-arg constructor has spawnReadyTimeoutMs=0 (gate disabled).
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"),
"agent.get is never called when the gate is disabled (no polling)");
}
@Test
void spawnGateRespectsZeroTimeoutEvenWithFullConstructor() {
FakeHerdr herdr = new FakeHerdr();
long[] clock = {0};
// Explicit zero timeout with the full testability constructor — should
// skip polling entirely, just like the legacy default path.
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 1);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn still succeeds with zero timeout");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no orphan pane close from the gate path");
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
void workerInheritsTheDaemonPath() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/opt/tools/bin:/usr/bin" : null).spawn();
assertEquals("/opt/tools/bin:/usr/bin", startEnv(herdr).get("PATH"),
"a worker with no PATH cannot run the build it is asked to run");
}
@Test
void profileEnvIsInjectedIntoTheWorker() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/daemon/bin" : null).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("/opt/jdk", env.get("JAVA_HOME"), "profile env: is passed through");
assertEquals("/profile/bin", env.get("PATH"), "an explicit profile PATH overrides the daemon's");
}
/**
* The security-relevant ordering. {@code SubscriptionGuard} is checked against the profile's
* {@code baseUrl} only, so if a profile's {@code env:} could overwrite ANTHROPIC_BASE_URL a
* worker could be pointed at an unguarded host while the guard passed on a benign one.
*/
@Test
void profileEnvCannotOverrideGuardCheckedAnthropicVars() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertEquals("http://gx00.gw:8000", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
private ClaudeCodeLauncher serviceWithModel(FakeHerdr herdr, String model) {
BridgedConfig.Worker cfg = profileWithModel(model);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null);
}
private static BridgedConfig.Worker profileWithModel(String model) {
return new BridgedConfig.Worker("sonnet", "http://gx00.gw:8000", model, null,
"BRIDGED_WORKER_TOKEN", List.of("ccs", "sonnet"), "tab", "bridged-workers",
"w #{n}", "http://127.0.0.1:8765/mcp", null, null);
}
@Test
void aConfiguredModelIsPassedAsAModelFlagAsWellAsTheEnvVar() {
// ANTHROPIC_MODEL alone loses to `ccs`, which exports its own model family over whatever it
// inherited — so a profile that set model: was silently overruled by its own launcher.
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, "claude-sonnet-5").spawn("sonnet", null, null);
assertEquals("claude-sonnet-5", startEnv(herdr).get("ANTHROPIC_MODEL"));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--model");
assertTrue(flag >= 0, "the flag is what survives a wrapper argv like [ccs, sonnet]");
assertEquals("claude-sonnet-5", args.get(flag + 1));
}
@Test
void theModelFlagComesLastSoItOutranksTheOperatorsOwnArgv() {
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, "claude-sonnet-5").spawn("sonnet", null, null);
List<String> args = spawnedArgs(herdr);
assertEquals(args.size() - 2, args.indexOf("--model"));
}
@Test
void aProfileWithNoModelGetsNoModelFlag() {
// gx10 deliberately leaves model: unset so ccs owns selection; adding a flag would make
// this file a second source of truth for exactly the thing it declines to decide.
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, null).spawn("sonnet", null, null);
assertFalse(spawnedArgs(herdr).contains("--model"));
assertNull(startEnv(herdr).get("ANTHROPIC_MODEL"));
}
// --- CB-539: subscription-profile opt-in ----------------------------------------------------
/** A claude-code profile on the subscription: no baseUrl (by design), no off-sub endpoint. */
private static BridgedConfig.Worker subscriptionCfg(String profile, String baseUrl) {
return new BridgedConfig.Worker(
profile, baseUrl, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null, Map.of(), null, null, true);
}
@Test
void defaultRefusalIsPreservedForClaudeProfileWithNoBaseUrl() {
// Requirement 1: absent subscription:true ⇒ byte-identical refusal to today. A claude-code
// profile with no baseUrl and no subscription must still be refused (it would bill the sub).
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"ltms-local", null, "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers", "w #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
GuardException ex = assertThrows(GuardException.class, () -> svc.spawn("ltms-local", null, null));
assertTrue(ex.getMessage().contains("no ANTHROPIC_BASE_URL"),
"the refusal names the missing baseUrl: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the refusal");
}
@Test
void subscriptionProfileSpawnsWithoutInjectedAnthropicVars() {
// Requirement on subscription:true: no baseUrl is required (or injected), and neither
// ANTHROPIC_BASE_URL nor ANTHROPIC_AUTH_TOKEN is injected even though the token env would
// resolve one if asked.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = subscriptionCfg("sonnet", null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "no baseUrl injected for a subscription profile");
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"), "no auth token injected for a subscription profile");
assertEquals("sonnet", env.get("ANTHROPIC_MODEL"),
"the model alias is still injected; only the subscription-boundary vars are dropped");
}
@Test
void subscriptionPlusBaseUrlIsRefused() {
// Requirement 2: subscription:true + a baseUrl state opposite intents — refuse at spawn,
// naming the profile, rather than silently picking a winner.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = subscriptionCfg("sonnet", "http://gx00.gw:8000");
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> svc.spawn("sonnet", null, null));
assertTrue(ex.getMessage().contains("sonnet"), "refusal names the profile: " + ex.getMessage());
assertTrue(ex.getMessage().contains("subscription"), "refusal explains the contradiction: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the contradiction was refused");
}
@Test
void nonSubscriptionProfilesAreStillAllowlistChecked() {
// Requirement 3: the guard keeps its teeth for every other profile — a base_url whose host is
// not on the allowlist is still refused, whether or not any subscription profile exists.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker rogue = new BridgedConfig.Worker(
"rogue", "http://evil.example.com:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "rogue"), "tab", "bridged-workers", "w #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(rogue.profile(), rogue), rogue.profile(), _ -> null);
GuardException ex = assertThrows(GuardException.class, () -> svc.spawn("rogue", null, null));
assertTrue(ex.getMessage().contains("not on the"), "refusal cites the allowlist: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the allowlist refusal");
}
@Test
void aSubscriptionProfileHasAnEnvSuppliedAnthropicBindingStripped() {
// CB-542: even a subscription profile whose env: carries ANTHROPIC_BASE_URL (or AUTH_TOKEN)
// must not hand them to the worker — on the subscription path no guard would vet them. Config
// load refuses this loudly; this launcher-side strip is the belt-and-braces that makes the
// invariant hold for a profile built in code that never passed through that validation.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
"sonnet", null, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "sonnet"), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null,
Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com",
"ANTHROPIC_AUTH_TOKEN", "sk-ant-bad", "JAVA_HOME", "/opt/jdk"),
null, null, true);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"),
"the unguarded endpoint must not survive into the worker");
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"),
"the unguarded token must not survive into the worker");
assertEquals("/opt/jdk", env.get("JAVA_HOME"),
"only the Anthropic binding keys are stripped; the rest of env: still applies");
}
}
@@ -1,395 +0,0 @@
package dev.ltms.bridged.worker;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.HashMap;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/**
* The composite router: profile → owning adapter for spawn/cwd/parity, pane id → owner for stop,
* and fleet-wide union/dedup for list/reap/caps/profiles. Exercised through two real adapters —
* claude-code + opencode — over one FakeHerdr, so each call is observed reaching the right adapter
* (the started herdr agent name carries that adapter's {@code claude-}/{@code opencode-} prefix).
*/
class CompositePeerLauncherTest {
private ClaudeCodeLauncher claudeAdapter(FakeHerdr herdr) {
// 12-arg back-compat Worker ctor → kind defaults to claude-code.
BridgedConfig.Worker claude = new BridgedConfig.Worker("claude", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}",
null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("claude", claude), "claude", _ -> null);
}
private OpenCodeLauncher opencodeAdapter(FakeHerdr herdr) {
BridgedConfig.Worker gemini = new BridgedConfig.Worker("gemini", null, "google/gemini-2.5-pro",
null, "BRIDGED_WORKER_TOKEN", List.of("opencode"), "tab", "bridged-workers", "w #{n}",
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Worker.KIND_OPENCODE);
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", gemini), "gemini", _ -> "tok");
}
private CompositePeerLauncher composite(FakeHerdr herdr) {
return new CompositePeerLauncher(
List.of(claudeAdapter(herdr), opencodeAdapter(herdr)), "claude");
}
/**
* A minimal concrete HerdrPeerLauncher for policy tests. It either returns a fake handle for the
* requested profile or throws, depending on {@code failProfiles}. buildLaunch is a stub; only
* spawn/stop/list/caps/reap are exercised by the composite.
*/
private static final class StubLauncher extends HerdrPeerLauncher {
private final Set<String> failProfiles;
private final Map<String, Integer> spawnCounts = new HashMap<>();
StubLauncher(String prefix, FakeHerdr herdr,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Set<String> failProfiles) {
super(prefix, new AgentControl(herdr), new WorkspaceControl(herdr),
profiles, defaultProfile, _ -> null, 0L, System::currentTimeMillis, () -> { });
this.failProfiles = Set.copyOf(failProfiles);
}
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
return new Launch(Map.of(), List.of());
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String p = (req.profileName() == null || req.profileName().isBlank())
? defaultProfile() : req.profileName();
spawnCounts.merge(p, 1, Integer::sum);
if (failProfiles.contains(p)) {
throw new PeerUnreachableException(p + " is down");
}
return new PeerHandle() {
@Override public String id() { return "pane-" + p; }
@Override public String terminalId() { return "term-" + p; }
@Override public String profile() { return p; }
};
}
@Override
public void stop(String id) { }
@Override
public List<Agent> list() { return List.of(); }
@Override
public Set<Capability> capabilities() { return EnumSet.noneOf(Capability.class); }
@Override
public int reapOrphanWorkers() { return 0; }
int spawnCount(String profile) {
return spawnCounts.getOrDefault(profile, 0);
}
}
private static BridgedConfig.Worker stubWorker(String profile) {
return new BridgedConfig.Worker(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null, null, null);
}
private static BridgedConfig.Worker stubWorker(String profile, float weight, Integer maxLoad) {
return new BridgedConfig.Worker(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null,
weight, maxLoad);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
* so a {@code Map.of} would make "which profile is tried first" a coin flip per run and any
* assertion about the first attempt intermittently false.
*/
private static Map<String, BridgedConfig.Worker> ordered(String first, BridgedConfig.Worker a,
String second, BridgedConfig.Worker b) {
Map<String, BridgedConfig.Worker> m = new LinkedHashMap<>();
m.put(first, a);
m.put(second, b);
return m;
}
@SuppressWarnings("unchecked")
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("name");
}
@Test
void spawnRoutesEachProfileToItsOwningAdapter() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
composite.spawn(new SpawnRequest("gemini", null, null));
assertTrue(startedName(herdr).startsWith("opencode-"),
"the gemini profile is spawned by the opencode adapter: " + startedName(herdr));
composite.spawn(new SpawnRequest("claude", null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"the claude profile is spawned by the claude-code adapter: " + startedName(herdr));
}
@Test
void nullProfileResolvesTheDefaultAndRoutesToItsOwner() {
FakeHerdr herdr = new FakeHerdr();
composite(herdr).spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"a no-profile spawn resolves the default (claude) and routes to its adapter");
}
@Test
void unknownProfileIsRejected() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertThrows(IllegalArgumentException.class,
() -> composite.spawn(new SpawnRequest("nope", null, null)),
"a profile no adapter declares is an error");
}
@Test
void profilesAndDefaultAreExposedAcrossAdapters() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertEquals(Set.of("claude", "gemini"), composite.profiles(),
"profiles are the union of every adapter's profiles");
assertEquals("claude", composite.defaultProfile());
}
@Test
void capabilitiesAreTheUnionOfEveryAdapter() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher claude = claudeAdapter(herdr);
OpenCodeLauncher opencode = opencodeAdapter(herdr);
PeerLauncher composite = new CompositePeerLauncher(List.of(claude, opencode), "claude");
assertTrue(composite.capabilities().containsAll(claude.capabilities()),
"the fleet offers every claude-code capability");
assertTrue(composite.capabilities().containsAll(opencode.capabilities()),
"the fleet offers every opencode capability (incl. SELF_PR from its git-token profile)");
}
@Test
void listIsDeduplicatedByPaneIdAcrossAdaptersSharingHerdr() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
// Both adapters wrap the same herdr, so each list() returns the same global agent set;
// the composite must return each pane once, not once per adapter.
assertEquals(1, composite.list().size(),
"the single herdr-tracked pane appears once, not duplicated per adapter");
}
@Test
void reapSumsAcrossAdaptersAndEachAdapterReapsOnlyItsOwnPrefix() {
// One foreign opencode orphan + one foreign claude orphan, from a prior daemon (different nonce).
FakeHerdr herdr = new FakeHerdr()
.withAgent("opencode-gemini-ffffff-1", "term_o", "wQ:pO", "wQ:tO")
.withAgent("claude-claude-eeeeee-1", "term_c", "wQ:pC", "wQ:tC");
PeerLauncher composite = composite(herdr);
assertEquals(2, composite.reapOrphanWorkers(),
"both orphans are reaped — one by each adapter, summed by the composite");
}
@Test
void stopTearsDownAPaneSpawnedThroughTheComposite() {
// CB-519: handle.id() is a host-unique opaque UUID, not the herdr pane — stop(id) must
// resolve it through the owning adapter down to the actual pane coordinate it spawned.
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
assertNotEquals("w9:pRoot_1", handle.id(), "the id is decoupled from the pane coordinate");
composite.stop(handle.id());
assertTrue(herdr.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"stop routes to the spawning adapter and closes exactly that worker's pane");
}
@Test
void opencodeContextResetIsANoOpAndWarnsOnlyOnce() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
assertFalse(composite.clearContext(handle.id()));
assertFalse(composite.clearContext(handle.id()));
} finally {
logger.detachAppender(appender);
}
assertFalse(opencodeAdapter(herdr).capabilities().contains(Capability.CONTEXT_RESET));
assertTrue(herdr.calls.stream().noneMatch(c -> "agent.prompt".equals(c.method())),
"never type Claude's /clear into an opencode prompt");
assertEquals(1, appender.list.stream()
.filter(e -> e.getFormattedMessage().contains("context reset is unsupported"))
.count(), "unsupported reset is logged once per adapter, not once per turn");
}
@Test
void constructorRejectsAProfileClaimedByTwoAdapters() {
FakeHerdr herdr = new FakeHerdr();
// Two opencode adapters both declaring "gemini" — a profile-name collision.
OpenCodeLauncher a = opencodeAdapter(herdr);
OpenCodeLauncher b = opencodeAdapter(herdr);
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(a, b), "gemini"),
"a profile two adapters both claim is a configuration error");
}
@Test
void constructorRejectsAnEmptyAdapterList() {
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(), "claude"),
"at least one adapter must be configured");
}
@Test
void fixedDefaultIsNoOpForUnqualifiedSpawns() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = composite(herdr);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"fixed placement still routes an unqualified spawn to the default profile");
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
}
@Test
void weightedPolicyGatesProfileAtMaxLoad() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> "a".equals(name) ? 1 : 0);
for (int i = 0; i < 5; i++) {
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "profile a is at maxLoad, so every spawn must land on b");
}
}
@Test
void weightedPolicyDistributesAccordingToWeightRatio() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 0.75f, null),
"b", stubWorker("b", 0.25f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = composite.spawn(new SpawnRequest(null, null, null)).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "weighted distribution should hold the 3:1 ratio");
assertEquals(10, b);
}
@Test
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "the spawn must fail over from unreachable a to b");
assertEquals(1, adapter.spawnCount("a"), "a was tried once and failed");
assertEquals(1, adapter.spawnCount("b"), "b was tried once and succeeded");
}
/**
* Definition order — not hash order — decides an exact-weight tie. Paired with the test above
* (same two profiles, opposite declaration order, opposite expected first attempt) this pins the
* ordering contract from both sides: under a salted map one of the two must fail on every run.
*/
@Test
void reversingDefinitionOrderReversesWhichProfileIsTriedFirst() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"b", stubWorker("b"),
"a", stubWorker("a"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("b"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("a", h.profile(), "b is declared first and unreachable, so the spawn lands on a");
assertEquals(1, adapter.spawnCount("b"), "b, declared first, is the one tried first");
}
@Test
void failoverBoundedByCandidateCount() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a", "b"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
PeerUnreachableException e = assertThrows(PeerUnreachableException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)));
assertTrue(e.getMessage().contains("no reachable worker profile"), e.getMessage());
assertEquals(1, adapter.spawnCount("a"));
assertEquals(1, adapter.spawnCount("b"));
}
@Test
void emptyCandidateSetThrowsClearException() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Worker> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, 1));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 1);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
}
@@ -1,275 +0,0 @@
package dev.ltms.bridged.worker;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
* inline flags, no {@code ANTHROPIC_*}, no guard), the {@code -m} model flag, and the shared base
* transport (naming, reap, readiness gate) proving the {@link HerdrPeerLauncher} SPI is neutral.
*/
class OpenCodeLauncherTest {
private static BridgedConfig.Worker opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
return new BridgedConfig.Worker("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, gitTokenEnv, null, BridgedConfig.Worker.KIND_OPENCODE);
}
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot);
}
@SuppressWarnings("unchecked")
private static Map<String, Object> lastStart(FakeHerdr herdr) {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
}
/** Protocol 19: the worker's env is injected at pane creation (tab.create), not agent.start. */
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
Map<String, String> env =
(Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
return env == null ? Map.of() : env;
}
/** Protocol 19: agent.start carries only the args after the kind-resolved executable. */
@SuppressWarnings("unchecked")
private static List<String> startArgs(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("args");
}
@Test
void writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null))
.spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "opencode carries no ANTHROPIC_* / subscription boundary");
String cfgPath = env.get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "OPENCODE_CONFIG points the worker at the generated config file");
assertTrue(Path.of(cfgPath).startsWith(root), "config file is generated under the injected root");
// Assert on parsed structure, not substrings: the generated config is real JSON and its
// whitespace is the formatter's business, not the contract's.
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
JsonNode bridge = json.path("mcp").path("bridge");
assertEquals("remote", bridge.path("type").asText(), "bridge is mounted as a remote MCP server");
assertEquals("http://127.0.0.1:8765/mcp", bridge.path("url").asText(),
"the profile's bridge MCP url is present");
assertTrue(bridge.path("enabled").asBoolean(), "the bridge server is enabled");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter is mounted via instructions");
// The instructions entry is a real file path holding the reply charter.
Path charter = Path.of(cfgPath).resolveSibling("reply-charter.md");
assertTrue(Files.exists(charter), "the charter file the config references was written");
assertTrue(Files.readString(charter).contains("bridge_reply"),
"the charter instructs the worker to answer via bridge_reply");
}
@Test
void noConfigFileWhenMcpUrlAbsent(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
assertNull(startEnv(herdr).get("OPENCODE_CONFIG"),
"no bridge MCP url → no config file and no OPENCODE_CONFIG");
}
@Test
void passesTheModelAsDashMFlagAlongsideAutoApprove(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
List<String> args = startArgs(herdr);
assertTrue(args.contains("--auto"),
"--auto is present alongside -m so a spawned peer never blocks on approval");
int m = args.indexOf("-m");
assertTrue(m >= 0, "model is selected with -m");
assertEquals("google/gemini-2.5-pro", args.get(m + 1), "the provider/model selector follows -m");
}
@Test
void autoApproveIsUnconditionalWhenModelBlank(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
assertEquals(List.of("--auto"), startArgs(herdr),
"--auto is unconditional: a model-less worker still must never block on approval");
}
@Test
void injectsForgeTokenWhenProfileGrantsIt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN")).spawn();
assertEquals("tok", startEnv(herdr).get("GITEA_TOKEN"),
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
}
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP),
service(herdr, root, opencodeCfg(null, null, null)).capabilities(),
"no git token → no SELF_PR");
assertTrue(service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN"))
.capabilities().contains(Capability.SELF_PR),
"a git-token profile adds SELF_PR");
}
@Test
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
String nonce = "abc123";
assertTrue(OpenCodeLauncher.isForeignWorker("opencode-gemini-def456-1", nonce),
"an opencode pane from another process is foreign");
assertFalse(OpenCodeLauncher.isForeignWorker("opencode-gemini-" + nonce + "-1", nonce),
"our own opencode pane (same nonce) is not foreign");
assertFalse(OpenCodeLauncher.isForeignWorker("claude-ltms-local-def456-1", nonce),
"a claude pane is never reaped by the opencode adapter");
}
@Test
void productionConstructorsWireThroughToTheBase() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = opencodeCfg(null, null, null);
// 5-arg (gate disabled) and 7-arg (gate enabled) production constructors both expose the profile.
OpenCodeLauncher disabled = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
OpenCodeLauncher gated = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null, 5000, 100);
assertEquals(java.util.Set.of("gemini"), disabled.profiles());
assertEquals("gemini", gated.defaultProfile());
}
@Test
void spawnGateThrowsPeerUnreachableWhenNeverInjectable(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // never injectable
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root);
PeerUnreachableException ex = assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000, "the fake clock advanced past the timeout: " + clock[0]);
long closes = herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, closes, "the worker pane was reaped on timeout (no orphan)");
assertNotNull(ex.getMessage());
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
PeerHandle handle = service(herdr, root, opencodeCfg(null, null, null))
.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"), "no polling when the gate is disabled");
}
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
private static BridgedConfig.Worker pinnedCfg(String model, String baseUrl, String mcpUrl) {
return new BridgedConfig.Worker("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, null, null, BridgedConfig.Worker.KIND_OPENCODE);
}
@Test
void baseUrlDeclaresACustomOpenAiCompatibleProvider(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash", "http://127.0.0.1:8000", null))
.spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "a pinned endpoint needs a config file even with no bridge MCP url");
JsonNode provider = new ObjectMapper().readTree(Path.of(cfgPath).toFile())
.path("provider").path("local-vllm");
assertFalse(provider.isMissingNode(), "the provider id comes from the model selector");
assertEquals("@ai-sdk/openai-compatible", provider.path("npm").asText());
assertEquals("http://127.0.0.1:8000/v1", provider.path("options").path("baseURL").asText(),
"a bare host:port gets /v1 appended — that is where these servers mount the API");
assertFalse(provider.path("options").path("apiKey").asText().isBlank(),
"the AI SDK requires a non-empty key even when the server ignores it");
assertFalse(provider.path("models").path("deepseek-v4-flash").isMissingNode(),
"the model half of the selector is declared under the provider");
}
@Test
void aBaseUrlThatAlreadyCarriesAPathIsUsedVerbatim(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/m", "http://127.0.0.1:8000/openai/v1", null)).spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertEquals("http://127.0.0.1:8000/openai/v1",
json.path("provider").path("local-vllm").path("options").path("baseURL").asText(),
"an endpoint mounted on a custom path must not have /v1 bolted on");
}
@Test
void aPinnedEndpointRejectsAModelWithNoProviderPrefix(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher =
service(herdr, root, pinnedCfg("deepseek-v4-flash", "http://127.0.0.1:8000", null));
// Silently falling back to the default gateway would point the worker at the wrong LLM
// while looking healthy — the one failure mode worth being loud about.
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, launcher::spawn);
assertTrue(e.getMessage().contains("<provider>/<model>"), "the error says how to fix it");
}
@Test
void aPinnedEndpointAndTheBridgeMcpCoexistInOneConfig(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/deepseek-v4-flash",
"http://127.0.0.1:8000", "http://127.0.0.1:8766/mcp")).spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertEquals("remote", json.path("mcp").path("bridge").path("type").asText(),
"pinning an endpoint must not drop the bridge MCP mount");
assertFalse(json.path("provider").path("local-vllm").isMissingNode(),
"and the provider block is still declared alongside it");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter survives too");
}
@Test
void noBaseUrlDeclaresNoProviderSoTheDefaultGatewayIsUsed(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("opencode/some-free-model", "http://127.0.0.1:8766/mcp", null))
.spawn();
JsonNode json = new ObjectMapper()
.readTree(Path.of(startEnv(herdr).get("OPENCODE_CONFIG")).toFile());
assertTrue(json.path("provider").isMissingNode(),
"without a baseUrl opencode resolves its own provider as before");
}
}
-61
View File
@@ -1,61 +0,0 @@
# CB-504 — systemd unit for bridged (Linux).
#
# The macOS launchd agent (deploy/dev.ltms.bridged.plist) is the supervision target for the
# current single-host deployment. This unit exists for the per-host gateways CB-308 introduces,
# which will run on Linux.
#
# Install (user service — bridged drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/bridged.service ~/.config/systemd/user/
# # edit ExecStart / WorkingDirectory / Environment below, then:
# systemctl --user daemon-reload
# systemctl --user enable --now bridged
# journalctl --user -u bridged -f
[Unit]
Description=bridged — claude-bridge message server
Documentation=https://git.ltms.dev/lms/claude-bridge/wiki
# Ordering only: herdr is a user process and its socket may appear after us. This is advisory —
# bridged retries the herdr socket rather than exiting, which is what actually makes a late
# socket survivable. Do NOT add Requires=: a herdr restart must not take bridged down with it.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/src/claude-bridge/bridged
ExecStart=/usr/lib/jvm/temurin-25-jdk/bin/java -jar target/bridged.jar bridged.yaml
Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker it
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put the API/worker tokens in a private
# drop-in that systemd reads with restrictive permissions:
# systemctl --user edit bridged → [Service] / Environment=BRIDGED_API_TOKEN=...
# or point EnvironmentFile at a 0600 file:
# EnvironmentFile=%h/.config/bridged/env
Restart=on-failure
RestartSec=10s
# A bad config (e.g. a non-loopback bind without token auth) makes bridged fail fast by design.
# Give up rather than restart-loop on a permanent error.
StartLimitBurst=5
StartLimitIntervalSec=120
# The daemon reads the repo, writes worktrees, and talks to a Unix socket — it needs no more.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-write
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
StandardOutput=journal
StandardError=journal
SyslogIdentifier=bridged
[Install]
WantedBy=default.target
-80
View File
@@ -1,80 +0,0 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<!--
CB-504 — launchd agent for bridged (macOS).
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
CB-308 introduces.
Install:
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
# edit the paths + JAVA_HOME below to match this host, then:
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
launchctl list | grep bridged
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
on startup instead, so an agent that comes up before herdr converges rather than dying — that
retry is the actual fix; KeepAlive below is the backstop.
-->
<plist version="1.0">
<dict>
<key>Label</key>
<string>dev.ltms.bridged</string>
<key>ProgramArguments</key>
<array>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/CHANGEME/src/claude-bridge/bridged/target/bridged.jar</string>
<string>bridged.yaml</string>
</array>
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
<key>WorkingDirectory</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged</string>
<key>EnvironmentVariables</key>
<dict>
<key>JAVA_HOME</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home</string>
<key>HERDR_SOCKET_PATH</key>
<string>/Users/CHANGEME/.config/herdr/herdr.sock</string>
<!--
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
-->
<key>PATH</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin:/Users/CHANGEME/Tool/apache-maven-3.9.16/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<!--
Worker/API tokens are NOT set here: this file is committed. Export them from a private
launchd override or a wrapper script. bridged reads the API token from the env var named
by auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token.
-->
</dict>
<key>RunAtLoad</key>
<true/>
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
non-loopback bind without token auth — that is a config error, not a transient one). -->
<key>KeepAlive</key>
<dict>
<key>SuccessfulExit</key>
<false/>
</dict>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.out.log</string>
<key>StandardErrorPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.err.log</string>
<key>ProcessType</key>
<string>Background</string>
</dict>
</plist>
+127
View File
@@ -0,0 +1,127 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<!--
CB-504 / CB-594 — launchd agent for fleetd (macOS).
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
no systemd. A systemd unit ships alongside (deploy/fleetd.service) for the Linux gateways
CB-308 introduces.
Install:
cp deploy/dev.ltms.fleetd.plist ~/Library/LaunchAgents/
launchctl load -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist
launchctl list | grep fleetd
The paths below are already filled in for this host (resolved 2026-08-16 from
`/usr/libexec/java_home`... except that reported the system Applet-plugin JVM, not the jenv-
managed JDK 25 actually used to build/run fleetd, so JAVA_HOME here is the real one:
`JENV_VERSION=25.0.3 java -XshowSettings:properties -version 2>&1 | grep java.home`; `which mvn`;
`echo $HOME`). If this file is copied to a different host, re-resolve all three paths and check
no placeholder path is left behind; scripts/redeploy-fleetd.sh's check mode does not (and
cannot) check this file for you.
CB-594 — launchd cannot run a login shell (see the PATH comment on EnvironmentVariables below,
and scripts/fleetd-launchd-wrapper.sh for the fix): ProgramArguments below execs THAT wrapper,
not java directly, so WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN still get sourced from
${SHARED_ENV}/tools/secrets.sh even though launchd itself never sources anything.
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
does systemd in a way that survives a socket appearing late. fleetd retries the herdr socket
on startup instead, so an agent that comes up before herdr converges rather than dying — that
retry is the actual fix; KeepAlive below is the backstop.
CB-594 — KeepAlive vs. scripts/redeploy-fleetd.sh: a bare SIGTERM makes this JVM exit 143 even
with its shutdown hook running to completion (measured, see the CB-594 report), which
SuccessfulExit:false below reads as a crash and races to restart the OLD jar. The redeploy
script now detects a loaded agent and uses `launchctl unload`/`load` instead of a raw kill, so
only one supervisor ever touches the process at a time — read that script's own output on a
redeploy for the confirmation.
-->
<plist version="1.0">
<dict>
<key>Label</key>
<string>dev.ltms.fleetd</string>
<key>ProgramArguments</key>
<array>
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
<string>fleetd.yaml</string>
</array>
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
<key>WorkingDirectory</key>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd</string>
<key>EnvironmentVariables</key>
<dict>
<key>JAVA_HOME</key>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home</string>
<key>HERDR_SOCKET_PATH</key>
<string>/Users/dai.ha/.config/herdr/herdr.sock</string>
<!--
PATH matters more than it looks (CB-511): fleetd propagates its own PATH to every worker
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
source .zprofile/.zshrc, so without this the daemon (and therefore every worker) gets a
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
-->
<key>PATH</key>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin:/Users/dai.ha/Softwares/apache-maven/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<!--
Worker/API tokens are NOT set here: this file is committed. CB-594 —
scripts/fleetd-launchd-wrapper.sh (named in ProgramArguments above) is what supplies
them, by execing a login shell that sources ${SHARED_ENV}/tools/secrets.sh before the
daemon itself starts. fleetd also reads the API token from the env var named by
auth.tokenEnv (default FLEETD_API_TOKEN) and only in auth.mode: token — the wrapper
covers that one too, since it is the same login shell.
-->
</dict>
<key>RunAtLoad</key>
<true/>
<!--
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
most one per 10s; it does NOT cap how many times launchd retries. If fleetd fails fast on
every start — a bad fleetd.yaml, for example auth.mode: token with the token env var unset,
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
primitive, so this is not something a config change here can fix.
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist`,
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
more exits to restart). scripts/redeploy-fleetd.sh does not add a third way — it does not
make fleetd self-disable on a config error, on purpose: a fail-fast exit path that
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
this comment already names — so it buys nothing an operator watching for the crash loop
doesn't already have, at the cost of a new way to be silently down. Watch for it with
`launchctl list dev.ltms.fleetd` (a high restart count) or by tailing fleetd.out for the
same startup error repeating every ~10s.
-->
<key>KeepAlive</key>
<dict>
<key>SuccessfulExit</key>
<false/>
</dict>
<key>ThrottleInterval</key>
<integer>10</integer>
<!--
CB-594 — same file scripts/redeploy-fleetd.sh already tails ($BRIDGED/fleetd.out), and both
streams point at it, not two separate log files: the script's fresh-line / ERROR-count checks
after a restart read this one path regardless of whether launchd or the script started the
process, and a stdout/stderr split would make half of what happens during a launchd-driven
restart invisible to it.
-->
<key>StandardOutPath</key>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/fleetd.out</string>
<key>StandardErrorPath</key>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/fleetd.out</string>
<key>ProcessType</key>
<string>Background</string>
</dict>
</plist>
+85
View File
@@ -0,0 +1,85 @@
# CB-504 — systemd unit for fleetd (Linux).
#
# fleetd #360: the previous version of this file started clean and broke the daemon in three ways
# that nothing logs (see the DO NOT block and the ExecStart/PrivateTmp comments below for what and
# why). The unit below, plus its companion deploy/herdr.service, is the version that has actually
# run on fleet01 without those failures. Do not "improve" it back toward the old shape without
# re-reading why each line is the way it is.
#
# Install (user service — fleetd drives the user's herdr, not a system daemon):
# mkdir -p ~/.config/systemd/user
# cp deploy/fleetd.service deploy/herdr.service ~/.config/systemd/user/
# # edit WorkingDirectory / ExecStart below for your host's paths and java location
# systemctl --user daemon-reload
# systemctl --user enable --now herdr fleetd
# loginctl enable-linger $USER # REQUIRED -- see below
# journalctl --user -u fleetd -f
#
# `loginctl enable-linger` is not optional and is easy to miss, because leaving it out looks like
# success: `systemctl --user enable` reports "enabled" and both units run for as long as you stay
# logged in. A user manager without lingering starts at your first login and stops at your last
# logout, so the fleet simply does not come back after a reboot -- which is the whole reason to
# use systemd here rather than the setsid scripts these units replaced. Check it with
# `loginctl show-user $USER -p Linger`; the answer must be `Linger=yes`.
#
# Secrets (AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI, ...) are not set
# here and need no systemd drop-in: ExecStart runs a login shell, so they come from wherever your
# login shell already sources them (this host: ~/.fleet/secrets.sh via ~/.zprofile). If a token is
# missing there, fleetd still starts — the daemon reports every secret a configured profile
# references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)"; a missing one logs
# "startup secret NAME: MISSING" at WARN and the daemon starts anyway — the first visible symptom
# is a member that cannot open a pull request, hours later and in a different component.
[Unit]
Description=fleetd — fleet message server
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
# Ordering only. fleetd retries the herdr socket rather than exiting, which is what actually makes
# a late socket survivable. Do NOT add Requires=: a herdr restart must not take fleetd down too.
After=herdr.service
Wants=herdr.service
[Service]
Type=simple
WorkingDirectory=%h/LTMS/fleetd/fleetd
# A LOGIN shell, not java directly. Every secret this daemon needs (AI_GATEWAY_TOKEN,
# WORKER_GITEA_TOKEN, LAVINMQ_URI, COORD_AMQP_URI) lives in ~/.fleet/secrets.sh, which only
# ~/.zprofile sources. systemd runs no login shell. Started any other way the daemon boots fine
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
# unit) has to read them. A private /tmp turns the credential scrub into a silent no-op.
PrivateTmp=false
Restart=on-failure
RestartSec=10s
# A bad config makes fleetd fail fast by design. Give up rather than restart-loop forever.
StartLimitBurst=5
StartLimitIntervalSec=120
# DO NOT add ProtectSystem=, ProtectHome=, ProtectKernelTunables= or ProtectControlGroups=.
# Measured on fleet01 2026-09-05: each of those gives the unit its own mount namespace, and
# fleetd resolves a caller role by running lsof to find the loopback peer PID
# (mcp/LsofPeerPidLookup). Inside such a namespace lsof returns nothing, every caller falls back
# to ANONYMOUS, and the primary is refused every orchestration call with
# "unauthenticated: anonymous may not SPAWN".
# The daemon still starts, healthz still returns ok and the secrets still resolve - the only
# symptom is that the fleet cannot be driven at all. Verified by bisecting the directives:
# no sandbox 3 lsof lines | ProtectSystem=strict 0 | ProtectHome=read-only 0
# ProtectKernelTunables 0 | ProtectControlGroups 0 | RestrictSUIDSGID 3 | NoNewPrivileges 3
# The two below add no mount namespace and are safe.
NoNewPrivileges=true
RestrictSUIDSGID=true
StandardOutput=journal
StandardError=journal
SyslogIdentifier=fleetd
[Install]
WantedBy=default.target
+19
View File
@@ -0,0 +1,19 @@
#!/bin/zsh
# fleetd #360 — template for the script deploy/herdr.service's ExecStart wraps in a pty.
#
# `script -qfec <this> /dev/null` needs a real command to run, and that command has to be a LOGIN
# shell script: herdr itself needs the same secrets fleetd.service's login shell picks up (this
# host: ~/.fleet/secrets.sh via ~/.zprofile), because members it spawns inherit its environment.
# systemd's own Environment= lines in herdr.service are not enough for that -- they set TERM and a
# bare PATH so the pty starts at all, nothing more.
#
# Copy this file to the path deploy/herdr.service's ExecStart names
# (%h/LTMS/fleetd/fleetd-run/herdr-inner.sh by default) and `chmod +x` it. Not committed under
# that path itself because the session name below is host-specific.
# A 0x0 pty makes every pane spawn fail with "ghostty error -2" (see herdr-multi-instance-facts /
# fleet01-headless-herdr-standup) -- give it a real size before herdr ever touches it.
stty rows 50 cols 200
# -l: login shell, so herdr and everything it spawns gets the real secrets and PATH.
exec zsh -lc 'exec herdr --session <name>'
+40
View File
@@ -0,0 +1,40 @@
# fleetd #360 — systemd unit for herdr (Linux), the terminal multiplexer fleetd drives.
#
# This is fleetd.service's companion: fleetd.service's After=/Wants=herdr.service assumes this
# unit exists. Before this ticket it did not, so on a fresh host fleetd started against a herdr
# that systemd never supervised at all.
#
# Install: see deploy/fleetd.service's header comment (both units install the same way).
#
# ExecStart below runs deploy/herdr-inner.sh (copy the template of that name from this directory
# to the path in ExecStart, or point ExecStart at wherever you keep it, and make it executable).
# It is a separate file rather than an inline command because it must itself be a login shell (see
# its own header for why) and systemd's ExecStart does not run one.
[Unit]
Description=herdr terminal multiplexer (fleet session)
Documentation=https://git.ltms.dev/fleet/fleetd/wiki
[Service]
Type=simple
# script(1) gives herdr a real pty. Without it the client reports a 0x0 window and every pane
# spawn fails with "ghostty error -2" -- which surfaces as a fleetd spawn failure, not a herdr one.
ExecStart=/usr/bin/script -qfec %h/LTMS/fleetd/fleetd-run/herdr-inner.sh /dev/null
StandardInput=null
Environment=TERM=xterm-256color
Environment=PATH=%h/.local/bin:/usr/local/bin:/usr/bin:/bin
# PrivateTmp MUST stay false. fleetd writes the member ZDOTDIR scrub dir and the opencode config
# dir under its own java.io.tmpdir, and the member pane -- a child of THIS process -- has to read
# them. A private /tmp here silently breaks the credential scrub instead of failing loudly.
PrivateTmp=false
Restart=on-failure
RestartSec=5s
StandardOutput=journal
StandardError=journal
SyslogIdentifier=herdr
[Install]
WantedBy=default.target
+9 -9
View File
@@ -1,8 +1,8 @@
# LavinMQ — the AMQP broker behind bridged's durable ReplyInbox (CB-307 Stage 2).
# LavinMQ — the AMQP broker behind fleetd's durable ReplyInbox (CB-307 Stage 2).
#
# Why this file exists: the broker was previously run ad hoc and simply vanished from the host,
# which takes bridged down with it — AmqpReplyInbox.open throws on an unreachable broker and
# Bridged.java:187 does not guard it, so a missing broker is a hard startup failure, not a
# which takes fleetd down with it — AmqpReplyInbox.open throws on an unreachable broker and
# Fleetd.java:187 does not guard it, so a missing broker is a hard startup failure, not a
# degraded mode. This pins the version, keeps the data, and brings itself back after a reboot.
#
# Usage:
@@ -14,27 +14,27 @@
#
# Management UI: http://127.0.0.1:15672 (guest / guest)
#
# This is bridged's OWN broker. Do not point bridged at any other AMQP server on this host —
# This is fleetd's OWN broker. Do not point fleetd at any other AMQP server on this host —
# notably not the `local-rabbitmq` container, which belongs to a different project and would end
# up carrying this project's queues.
name: bridged-broker
name: fleetd-broker
services:
lavinmq:
# Pinned deliberately: :latest silently moves the broker under a running daemon.
image: cloudamqp/lavinmq:2.9.1
container_name: bridged-lavinmq
container_name: fleetd-lavinmq
# The failure this deployment exists to prevent — survive reboots and Docker restarts, but
# stay down if it was stopped on purpose.
restart: unless-stopped
# Loopback-bound on purpose. LavinMQ ships a default guest/guest account, which is only
# acceptable because nothing off-host can reach it. bridged connects over 127.0.0.1, and
# acceptable because nothing off-host can reach it. fleetd connects over 127.0.0.1, and
# binding 0.0.0.0 here would expose a broker with default credentials to the network.
ports:
- "127.0.0.1:5672:5672" # AMQP — bridged.yaml broker.uri points here
- "127.0.0.1:5672:5672" # AMQP — fleetd.yaml broker.uri points here
- "127.0.0.1:15672:15672" # HTTP management API + UI
# The whole point of Stage 2. Held-but-unacked replies live here; without a named volume a
@@ -57,4 +57,4 @@ services:
volumes:
lavinmq-data:
name: bridged-lavinmq-data
name: fleetd-lavinmq-data
+531
View File
@@ -0,0 +1,531 @@
# CB-201 and CB-227 refinement
Date: 2026-09-03
## Decision
#201 and #227 are one delivery program, but they are not one implementation unit.
#201 has a real seam: `CompletionResolver` can publish a typed backend-error event only after its
waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier
must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can
be built in parallel with the classifier.
I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all
composition files and lands after them. It also depends on the #234 defect 2 fix named in the task.
```mermaid
flowchart LR
U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"]
U2["Unit 2: credential outage policy"] --> U5
U3["Unit 3: lead outage nudge"] --> U5
U4["Unit 4: durable member outcome"] --> U5
D234["#234 defect 2: fail-loud target resolution"] --> U5
```
*Figure 1. Four file-disjoint foundations feed one composition unit.*
This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one
unit in this plan.
## Evidence checked in the current branch
I read both issue pages in full. Each page reports zero comments.
| Evidence | What the code says now |
|---|---|
| `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. |
| `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. |
| `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. |
| `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. |
| `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. |
| `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. |
| `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. |
| `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. |
| `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. |
| `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. |
| `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. |
| `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. |
| `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. |
| `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. |
| `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. |
| `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. |
| `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. |
| `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. |
| `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. |
| `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. |
| `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. |
| `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. |
| `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. |
I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`,
`CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`.
I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No
peer architect was named, so I did not exchange a design with one.
## Required behaviour
The policy should use these first values:
- Threshold: **2** classified backend errors.
- Window: **60 seconds**, measured from the first error to the second.
- Cool-off: **60 seconds**, starting when the threshold is reached.
- Correlation key: `credentialId`, never profile name and never error text.
- Incident rule: one active incident per credential. Errors during its cool-off do not extend it and
do not create more lead notices.
- Rearm rule: after cool-off ends, two fresh errors are needed for another incident.
Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window
fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without
turning a short backend fault into the default 1,800-second exhaustion quarantine.
A single classified error still fails its send and marks its member `backend_error`. It does not
cool a credential and does not send an outage notice. This is what “a single error changes nothing”
must mean at the credential level. It cannot mean that the failed member still looks successful.
```mermaid
sequenceDiagram
participant R1 as Resolver for member A
participant R2 as Resolver for member B
participant P as Outage policy
participant S as Spawn gate
participant N as Lead push loop
participant L as Lead pane
R1->>P: backend error for credential C
Note over P: Count 1, no cool-off
R2->>P: backend error for credential C within 60s
P->>P: Start one 60s incident
P->>S: Credential C is cooling off
P->>N: Queue one incident notice
N->>L: Inject when lead is idle, blocked, or done
L->>S: Request another spawn on credential C
S-->>L: Refuse and report remaining cool-off
```
*Figure 2. The second independent classification creates the fleet-level event.*
Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show
zero free capacity and both members as `backend_error`. The push loop would inject one outage notice
even if the lead had not polled either ticket yet. The design reports the outage. It does not recover
uncommitted work from the members.
## Unit 1 — Typed backend-error classification
### Scope
Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the
public send result as a failed send. The typed internal event is the seam #227 consumes.
The lookup returns the pattern for a target. The sink receives the target, matched line, and full
failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured
waiter. This copies the race rule already used by `ExhaustionSink`.
The classifier must run in all three current paths:
1. a normal non-empty assistant block;
2. the #211 raw scrape fallback;
3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure.
In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic
failure or completion. A fast turn still fails when no configured pattern matches.
Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the
operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name
profiles using this weaker legacy default.
### Files owned
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java`.
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java`.
- Change `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java`.
- Change `fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. A target-specific error pattern matches a normal assistant block and resolves the send as failed.
2. The same match calls `BackendErrorSink` exactly once after the waiter resolution wins.
3. A late classification which loses to `fleet_reply` does not call the sink.
4. An exhausted line that also matches the generic error pattern stays `BACKEND_EXHAUSTED`. It calls
only `ExhaustionSink`.
5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses
the #211 fallback and calls the backend-error sink.
6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast
turn stays a generic failure.
7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never
increments outage evidence.
8. A non-match keeps the existing completion result and text.
9. Constructors used by current callers keep compiling. They use the legacy default lookup and an
explicit inert sink until Unit 5 supplies the production objects.
10. Unit tests pass. The developer runs the focused test first, then `mvn clean install` from
`fleetd/`.
### Dependencies
None. Unit 1 can run with Units 2, 3, and 4.
Unit 5 depends on its new lookup, sink, and constructor.
### What to report back
- The exact classifier order in all three paths.
- The focused test command and result.
- The test which proves a losing waiter race does not publish an event.
- The test which proves a fast matching failure is typed.
- The final `mvn clean install` result.
- Any constructor kept only for transition and where Unit 5 replaces it.
## Unit 2 — Credential outage policy
### Scope
Build a small credential-keyed state machine. It accepts already-classified backend-error events.
It does not read pane text, profiles, sessions, or lead state.
Use an injected monotonic clock. A call records `credentialId`, target, and reason. It returns a new
incident only on the threshold crossing. The incident contains a stable event id, credential id,
the distinct affected targets, evidence count, window, and remaining cool-off.
This class owns both correlation and short cool-off. Keeping them together makes threshold crossing
and the cool-off deadline one atomic state change.
### Files owned
- Add `fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java`.
- Add `fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. One error creates no incident and no cool-off.
2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second
cool-off.
3. Two errors more than 60 seconds apart do not create an incident.
4. The exact 60-second boundary has a pinned result. Use inclusive `<= 60s` so scheduler delay does
not discard evidence at the boundary.
5. Different credentials never share evidence.
6. Different profiles which supply the same credential id do share evidence. The policy itself only
sees the credential id.
7. More errors during active cool-off do not extend its deadline and do not return another incident.
8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident.
9. Remaining seconds round up, matching `BackendQuarantine` reporting.
10. Concurrent second and third errors cannot return two incidents.
11. The focused tests and `mvn clean install` pass.
### Dependencies
None. Unit 2 can run with Units 1, 3, and 4.
Unit 5 depends on the policy API.
### What to report back
- The state transition table and locking method.
- The exact threshold, window, cool-off, and boundary rule.
- The test which proves one incident under concurrent calls.
- The focused test and `mvn clean install` results.
## Unit 3 — Lead outage nudge
### Scope
Add backend incidents as a fourth pending source in `ReplyPushLoop`. Do not create another scheduler
or call `AgentControl.send` from `Fleetd`. The existing combined per-lead schedule is the control
which prevents competing injected turns.
The entry point takes an incident id, affected worker targets, credential id, affected profile
names, and remaining cool-off. It resolves distinct owning leads through `PrimaryRegistry`.
Each `(incidentId, lead)` item is one-shot. It waits while the lead is not injectable. After one
successful `agents.send`, remove it. A send exception keeps it pending for a bounded retry. It never
uses the repeated reminder behaviour of an uncollected ticket.
Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It
uses `PrimaryRegistry.nudgeTargetFor(target)` and says that correlation could not run. If no lead is
known, log at `WARN`, not `DEBUG`.
### Files owned
- Change `fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java`.
- Change `fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. One incident affecting two workers owned by one lead causes one successful pane injection.
2. Two affected workers owned by two leads cause one successful injection per affected lead. This
is one notice per event per lead, not one notice per member.
3. Repeating the same incident id is idempotent.
4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable
or its attempt cap is reached.
5. After one successful injection, later ticks do not mention that incident again.
6. A failed `agents.send` is retried within the existing bound. A successful retry still gives only
one successful send.
7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two
competing turns.
8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the
lead to run `fleet_list`.
9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known,
the code logs a `WARN` naming the target and reason.
10. `stop()` clears incident state as it clears other push state.
11. Existing reply, ticket, and question tests stay green. The focused tests and
`mvn clean install` pass.
### Dependencies
None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel.
Unit 5 adapts the Unit 2 incident into this entry point.
### What to report back
- The exact one-shot and retry rules.
- The test showing one combined nudge with a failed ticket.
- The test showing one successful send for two affected workers.
- The focused test and `mvn clean install` results.
## Unit 4 — Durable member backend-error outcome
### Scope
Make a classified backend failure remain visible after its ticket is collected or expires.
Add `BACKEND_ERROR` to `MemberSession.State`. Add a nullable failure detail to `MemberSession` and
render it as `failureReason` in `SessionManager.rosterView`. Add
`SessionManager.onBackendError(target, reason)`.
The transition must handle both completion orderings:
- `BUSY -> BACKEND_ERROR` when classification wins before the normal completion state update;
- `DONE -> BACKEND_ERROR` when the async resolver runs after `SessionManager.onTurnComplete`.
It must use a compare-and-set retry or another atomic update. `RELEASED` must never return to the
roster. A backend-error member is terminal and cannot accept another delivery.
### Files owned
- Change `fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java`.
- Change `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java`.
- Change `fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java`.
No other unit may edit these files.
### Acceptance criteria
1. `onBackendError` moves a `BUSY` member to `BACKEND_ERROR` and stores the reason.
2. It also moves `DONE` to `BACKEND_ERROR`, covering the resolver race.
3. A later normal `onTurnComplete` cannot change `BACKEND_ERROR` back to `DONE`.
4. A released or unknown member is not recreated. The unknown case logs at `WARN` and returns an
explicit false result to its caller.
5. `onDelivered` refuses a `BACKEND_ERROR` member, just as it refuses generic `FAILED`.
6. `rosterView` reports `state: backend_error` and `failureReason` after the send ticket is gone.
7. Ordinary members do not gain a blank or invented `failureReason` field.
8. Existing constructors keep source compatibility for tests and adapters.
9. The focused tests and `mvn clean install` pass.
### Dependencies
None. Unit 4 can run with Units 1, 2, and 3.
Unit 5 calls the new session method from the production sink.
### What to report back
- The two race orderings and the tests for both.
- The exact roster JSON shape.
- The unknown-target result and log level.
- The focused test and `mvn clean install` results.
## Unit 5 — Production wiring, spawn gate, and fleet views
### Scope
Compose Units 1 to 4 in production. This is the only unit which edits `Fleetd.java`.
Add per-profile `errorPattern` config beside `exhaustedPattern`. Compile both once at startup. A
configured pattern wins over the legacy default. Report configured profiles and legacy-default
profiles separately at startup. A bad regex must stop startup with the profile and key in the
message.
Wire one production `BackendErrorSink` with this order:
1. mark the member `backend_error` with its reason;
2. resolve the profile and its current `effectiveCredentialId()` through the fail-loud #234 seam;
3. record the error in `BackendOutagePolicy`;
4. on a new incident, submit one event to `ReplyPushLoop`.
If target metadata cannot be resolved, do not end in `Optional.ifPresent`. Log an error and call the
Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588.
Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when
both states are active. Automatic placement needs a distinct `coolingOff` set so its refusal does
not say “exhausted”.
Extend the MCP (Model Context Protocol) views:
- A cooling profile has `free: 0`, `credentialId`, and `coolingOffForSeconds` in `fleet_list`.
- It does not have `quarantinedForSeconds` unless exhaustion quarantine is also active.
- `fleet_profiles` has a separate `coolingOff` map, not an entry in `quarantined`.
- A direct spawn refusal says the credential is cooling off after repeated backend errors and gives
the remaining seconds.
Do not change `FleetHealthMonitor.coverage`. It still describes the periodic health webhook path.
Lead-pane outage delivery is a separate capability.
### Files owned
- `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java`
- `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java`
- `fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java`
- `fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java`
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java`
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java`
- `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java`
- `fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java`
- `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java`
- Add `fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java`.
- `fleetd/fleetd.example.yaml`
- `CLAUDE.md`
No earlier unit edits these files.
The lead, not a worker, must update `wiki/7-Use-Cases.md`, `wiki/9-Implementation.md`, and
`wiki/11-Features.md`. Project rules forbid workers from committing `wiki/`. The portable block in
`CLAUDE.md` and `wiki/7-Use-Cases.md` must remain byte-identical.
### Acceptance criteria
1. `errorPattern` binds per profile. Blank uses the legacy default and is reported as degraded
coverage. A config reload which changes it is reported as deferred because patterns are compiled
at startup.
2. A malformed `errorPattern` stops startup and names `profiles.<name>.errorPattern`.
3. The real production sink never silently drops an unknown target. A test captures its error log
and the fallback notice call.
4. One real backend-error classification through `CompletionResolver` marks only its member. It does
not cool the credential and does not send an outage notice.
5. Two real classifications for one credential within 60 seconds start one incident.
6. The integration test then calls the real explicit-profile spawn gate. It is refused before any
adapter spawn call, with “cooling off” and remaining seconds in the message.
7. The same test calls an automatic placement path. A cooling candidate is skipped. If every
candidate is cooling, the error names cool-off rather than exhaustion.
8. Profiles sharing the credential are all blocked. A profile on another credential stays usable.
9. `fleet_list` from the same fixture shows both members as `backend_error`, preserves each failure
reason, and reports `free: 0`, the credential, and `coolingOffForSeconds`.
10. `fleet_profiles` reports cool-off separately from quarantine.
11. The real `ReplyPushLoop` receives one incident and makes one successful lead-pane send. Existing
failed-ticket notice content may share that same combined send.
12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine
also wins in spawn errors and fleet views.
13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors.
14. `healthCoverage` has the same value before and after this change for the same health config.
15. `fleetd.example.yaml` explains `errorPattern`, the legacy fallback, 2/60/60 policy, and the
difference between cool-off and exhaustion quarantine.
16. `CLAUDE.md` tells leads how `fleet_profiles` and `fleet_list` report cool-off. The lead later
applies the matching wiki updates and runs the documented byte-sync check.
17. The developer records the new end-to-end test failing before implementation, then passing. The
focused suites and `mvn clean install` pass.
### Dependencies
Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the
#234 defect 2 fix lands, because both areas touch the same target-resolution control path.
### What to report back
- The exact commits used for Units 1 to 4 and #234.
- The startup coverage line with one configured and one legacy-default profile.
- The red test output before the implementation and its green result after.
- The explicit and automatic spawn refusal text.
- Sample `fleet_list` and `fleet_profiles` JSON for cool-off and exhaustion.
- The number and text of lead-pane sends in the real-path test.
- The focused test commands and final `mvn clean install` result.
- The exact `CLAUDE.md` change and the wiki edits the lead must apply.
## File ownership summary
| Area | Unit | Shared edit risk |
|---|---:|---|
| Completion classification | 1 | Only Unit 1 edits `CompletionResolver` and its test. |
| Correlation and cool-off state | 2 | New files only. |
| Lead push scheduling | 3 | Only Unit 3 edits `ReplyPushLoop` and its test. |
| Member terminal state | 4 | Only Unit 4 edits `MemberSession`, `SessionManager`, and their test. |
| Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits `Fleetd`, `FleetConfig`, `CompositePeerLauncher`, placement context, `FleetMcp`, and `CLAUDE.md`. |
| Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. |
## What I would not build
1. **Do not reuse `BackendQuarantine` for outages.** Its repeat call restarts a long credential
quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage.
2. **Do not merge the exhaustion and generic error patterns.** Exhaustion must win because it has a
different policy and duration.
3. **Do not group by error string.** One outage can produce different text. The shared operational
limit is the credential.
4. **Do not mark a profile unusable until config changes.** The current classifier cannot safely
tell a permanent malformed request from a transient service fault. A permanent state would need
a stronger error taxonomy first.
5. **Do not quarantine on the first generic backend error.** That would turn one bad request or one
false pattern match into a fleet-wide capacity loss.
6. **Do not add this to `FleetHealthMonitor`.** The monitor samples slow member health. The exact
backend event already exists at completion resolution, and moving it to polling would lose type
and time.
7. **Do not add another direct lead injector.** `ReplyPushLoop` already owns status gating,
per-lead coalescing, retry bounds, and heartbeat stand-down.
8. **Do not change `healthCoverage` to `full`.** That field still means a webhook notification sink
exists for periodic health. A backend outage nudge does not make every health event visible.
9. **Do not persist incident history across daemon restart in this work.** Existing exhaustion
quarantine is also in memory. A 60-second state does not justify a new durable store.
10. **Do not build work recovery.** The PR-body survival story proves why checkpoint-first work is
useful, but these tickets are about detection, capacity, and signalling.
11. **Do not remove the legacy `API Error:` fallback in the first release.** Doing so would turn an
unedited config back into a false successful completion. Report it as degraded coverage instead.
12. **Do not reorder or add the old `visibleTurn` fallback from #201.** #211 already implemented the
narrow raw-scrape fallback at `CompletionResolver.classifyRawScrapeFallback`.
## Riskiest assumption and cheapest experiment
The riskiest assumption is that a configured error regex means “the backend failed this turn”. The
current code and test already show the counterexample: a worker may quote `API Error:` while writing
a valid report. Two such false matches on one credential would now remove capacity for 60 seconds.
The cheapest experiment is a replay corpus before Unit 5 ships:
1. Save the full pane text from the measured 2026-09-01 outage.
2. Produce one safe failure per backend with a disposable invalid endpoint or request.
3. Save one valid member report which quotes each error line.
4. Replay all samples through the real `CompletionResolver` test fixture.
5. Require outage samples to match and quoted-report samples not to match after assistant-block
extraction and baseline checks.
This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile
patterns before enabling correlation. Do not raise the threshold to hide a bad classifier.
## Sequencing with three developers
First wave:
1. Developer A: Unit 1, typed classification.
2. Developer B: Unit 2, credential outage policy.
3. Developer C: Unit 3, lead outage nudge.
As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and
review Units 1 to 4 independently.
Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict
integration branch, so no other active unit should touch its file list.
## Checks performed for this refinement
- Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments.
- Read the source and tests named in the evidence section.
- Ran `git status --short --branch`; the branch was clean before this document was added.
- Ran `git log --oneline -12` to identify the branch base.
- I did not run Maven because this change adds only a design document.
- Rendered both Mermaid blocks with `npx @mermaid-js/mermaid-cli`; both commands succeeded.
+10 -17
View File
@@ -2,7 +2,7 @@
**Status:** design spec for review → delegate implementation.
**Grounded in:** `WorkerService`, `Injector`/`StatusPoller`/`TurnListener`, `MessageService`,
`BridgeMcp`, `BridgedApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)).
`FleetMcp`, `FleetApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)).
## Problem
@@ -15,7 +15,7 @@ Consequences today:
since when?" without shelling to herdr for a raw agent list (no state, no ownership, no age).
- Cleanup of a worker that outlived its owning process depends entirely on the boot-time
name-nonce **reaper** (CB-117) — there is no live, authoritative roster during a run.
- `bridge_list` (CB-304) can only surface herdr's view, not a bridge-owned roster.
- `fleet_list` (CB-304) can only surface herdr's view, not a bridge-owned roster.
- There is no seam for per-session policy (checkpoint on teardown → CB-302; idle_ttl /
context_cap / drain → CB-303).
@@ -35,7 +35,7 @@ build on.
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
the release hook it will attach to.
"Recycle" under no-reuse is simply **release + fresh acquire** — a helper, not a pool operation.
Under no-reuse, a released session is terminal. A new `acquire` always creates a fresh session.
## Design
@@ -43,14 +43,9 @@ build on.
subscription-guarded spawn/teardown mechanics; `SessionManager` adds the registry, lifecycle, and
ownership on top.
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
**Package:** new `dev.ltms.fleet.session` — keeps the registry/lifecycle concern separate from
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
**`recycle` is IN SCOPE for CB-301** (decided): implement `recycle(paneId, …)` = `release` the old
session then `acquire` a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is
a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is
covered by a test from day one.
### `WorkerSession` (record or small mutable holder)
| Field | Source | Notes |
@@ -88,7 +83,6 @@ SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
final class SessionManager {
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
void release(String paneId); // deterministic teardown + deregister
WorkerSession recycle(String paneId, ...); // release + acquire (no-reuse convenience)
Optional<WorkerSession> get(String paneId);
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
@@ -102,13 +96,13 @@ final class SessionManager {
### Integration points
- **`Bridged.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the
- **`Fleetd.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the
`TurnListener` alongside `CompletionResolver` so it sees turn boundaries, and give it the
`WorkerPresence` signal for `READY`.
- **`BridgeMcp.spawn` / `BridgedApp.spawnWorker`** — route spawn through `SessionManager.acquire`
(carry `callerTerminal` as `ownerTerminal`). **`bridge_stop` / `DELETE /workers/{paneId}`** →
- **`FleetMcp.spawn` / `FleetApp.spawnWorker`** — route spawn through `SessionManager.acquire`
(carry `callerTerminal` as `ownerTerminal`). **`fleet_stop` / `DELETE /workers/{paneId}`** →
`SessionManager.release`.
- **`bridge_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`.
- **`fleet_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`.
- **`MessageService`** — no change required for one-shot; a later CB-303 auto-release hook can call
`release` from `onTurnComplete` under policy.
@@ -120,12 +114,11 @@ final class SessionManager {
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
a second `release` on the same paneId is a harmless no-op.
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
5. `recycle` produces a new paneId and the old one is gone (no-reuse invariant).
6. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
5. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
## Seams left open (deliberately)
- **CB-302** — attach a checkpoint step (`STATE.md` + commit) to the `release` path.
- **CB-303** — a policy loop over `roster()` using `spawnedAtNanos`/state to auto-`release` on
`idle_ttl`, or drain on `context_cap`.
- **CB-304** — `bridge_list` reads `roster()` for a bridge-owned roster + live join.
- **CB-304** — `fleet_list` reads `roster()` for a bridge-owned roster + live join.
+7 -7
View File
@@ -2,12 +2,12 @@
**Status:** ✅ shipped — implemented at commit `97ecc71` (per-worker git worktree + config-parity
overlay). As-built: `session/GitWorktrees.java` behind the `Worktrees` port, wired in
`Bridged.main` and configurable via `worktreeRoot` / per-profile `parityOverlay`
(see `bridged.example.yaml`). Branch/worktree surface in `bridge_list` landed with CB-304
`Fleetd.main` and configurable via `worktreeRoot` / per-profile `parityOverlay`
(see `fleetd.example.yaml`). Branch/worktree surface in `fleet_list` landed with CB-304
(`9fe04bf`); the worker-opened-PR checkpoint landed as CB-302 (`64e70ef`).
**Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`).
**Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md).
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`,
**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `FleetConfig.Worker`,
`inject/…LsofPeerPidLookup` (the `ProcessBuilder` exec pattern).
## Problem
@@ -43,7 +43,7 @@ provider.
### `WorktreeRequest` (new, nullable = "no worktree")
```java
package dev.ltms.bridged.session;
package dev.ltms.fleet.session;
/** Ask acquire() to provision an isolated worktree. null ⇒ run in the shared primary tree. */
public record WorktreeRequest(String ticketSlug, String baseRef) {
// ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo.
@@ -63,7 +63,7 @@ untouched.
### `Worktrees` seam (new)
```java
package dev.ltms.bridged.session;
package dev.ltms.fleet.session;
public interface Worktrees {
/** git -C <repoRoot> worktree add <path> -b <branch> <baseRef|HEAD>. Returns the worktree path. */
String add(String repoRoot, String branch, String baseRef);
@@ -112,7 +112,7 @@ public void release(String paneId) {
}
```
### Config — `BridgedConfig.Worker.parityOverlay` + a `worktreeRoot`
### Config — `FleetConfig.Worker.parityOverlay` + a `worktreeRoot`
- Add `List<String> parityOverlay` to the `Worker` record (12th field). Compact-constructor default
when null/empty: `[".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]` (missing paths are
@@ -122,7 +122,7 @@ public void release(String paneId) {
### Surface: MCP + REST
- `bridge_spawn` gains an optional `worktree` arg: `true`, or a ticket slug string. Truthy ⇒ build a
- `fleet_spawn` gains an optional `worktree` arg: `true`, or a ticket slug string. Truthy ⇒ build a
`WorktreeRequest(slug, null)` and call the 5-arg `acquire`.
- `POST /workers` gains `worktree` (+ optional `ticket`) in the body/query, same mapping.
- `workerView`/`view(WorkerSession)` include `worktree` and `branch` **when non-null** (omit for
+11 -11
View File
@@ -1,16 +1,16 @@
# CB-306 — Spawn-Readiness Gate (launcher-owned terminal readiness)
**Status:** design note / delegation spec (branch `worker/cb-306-readiness`)
**Issue:** gitea `lms/claude-bridge` #4
**Issue:** gitea `fleet/fleetd` #4
**Owner of the behaviour:** `ClaudeCodeLauncher` (the `PeerLauncher` adapter) — NOT core.
## 1. Problem
`bridge_spawn` today returns a session the instant the herdr pane is started. The pane is not
`fleet_spawn` today returns a session the instant the herdr pane is started. The pane is not
yet a usable Claude REPL — it may still be sitting at the folder-trust prompt, or the CLI may
never come up at all. Nothing blocks or times out on that. Consequences:
- A `bridge_send` to a not-yet-ready worker surfaces as a **~60 s MCP-client timeout** (the send
- A `fleet_send` to a not-yet-ready worker surfaces as a **~60 s MCP-client timeout** (the send
blocks waiting for a turn that can't start) instead of a fast, explicit spawn failure.
- A worker stuck at the folder-trust prompt lingers in `SPAWNING` forever; nothing fails it.
@@ -55,7 +55,7 @@ a crash, and a slow start all present as "never becomes injectable" and all corr
or `spawnReadyTimeoutMs` elapses.
3. **Ready** → return the `WorkerHandle(paneId, terminalId)` as today.
4. **Timeout** → the launcher **closes the pane it started** (and its tab, via the same path
`release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.bridged.peer`).
`release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.fleet.peer`).
No orphan pane is left behind — the launcher cleans up its own failed birth.
`spawnReadyTimeoutMs == 0` (or unset) **disables** the gate = legacy non-blocking behaviour, so the
@@ -79,8 +79,8 @@ Unit tests (add to the existing `ClaudeCodeLauncher` test):
## 5. Config
Add to the launcher-level config (a bridged-level knob, not per-profile) in `bridged.yaml` +
`BridgedConfig`:
Add to the launcher-level config (a fleetd-level knob, not per-profile) in `fleetd.yaml` +
`FleetConfig`:
```yaml
spawn_ready_timeout_ms: 20000 # 0 disables the gate (legacy non-blocking spawn)
@@ -88,7 +88,7 @@ spawn_ready_poll_ms: 300
```
Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sane defaults in code
(`20000` / `300`). Keep the names consistent with existing config field style in `BridgedConfig`.
(`20000` / `300`). Keep the names consistent with existing config field style in `FleetConfig`.
## 6. Core / MCP propagation
@@ -99,11 +99,11 @@ Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sa
`spawn` throws — verify the new exception flows through it (worktree removed, nothing registered).
- The **non-worktree** path registers the session only *after* `spawn` returns, so a throw means no
half-live `SPAWNING` session is ever registered — confirm this and add a test.
- `bridge_spawn` (MCP verb) must return an **error result** carrying the exception message, not a
success with a dead session. Trace `BridgeMcp`/`BridgedApp` spawn handlers and make sure the
- `fleet_spawn` (MCP verb) must return an **error result** carrying the exception message, not a
success with a dead session. Trace `FleetMcp`/`FleetApp` spawn handlers and make sure the
exception becomes a clean tool error, not an uncaught 500 with a stack trace.
**Out of scope (do NOT do here):** gating `bridge_send` on session `READY` (existing status-gate +
**Out of scope (do NOT do here):** gating `fleet_send` on session `READY` (existing status-gate +
this spawn gate already close the window), MCP-handshake-as-readiness signal, the CB-307 broker,
any config `kind:` discriminator, any second adapter.
@@ -111,7 +111,7 @@ any config `kind:` discriminator, any second adapter.
- `ClaudeCodeLauncher.spawn` blocks until injectable or throws `PeerUnreachableException` +
self-reaps the pane; gate disabled when timeout is 0.
- New `PeerUnreachableException` in `dev.ltms.bridged.peer`.
- New `PeerUnreachableException` in `dev.ltms.fleet.peer`.
- Config knobs wired (`spawn_ready_timeout_ms`, `spawn_ready_poll_ms`) with safe defaults.
- Existing `SPAWNING→READY` MCP-contact transition untouched.
- New unit tests (ready / timeout+reap / disabled) green; **all existing tests still pass unchanged**.
+20 -20
View File
@@ -13,28 +13,28 @@ Stage 2 (the AMQP/LavinMQ adapter behind the same port) is explicitly **out of s
## 1. The bug this fixes (grounded in current code)
The reverse (worker→primary) path is `Rendezvous` — a `ConcurrentHashMap<session, CompletableFuture<Resolution>>`
of **live blocking waiters only**. No queue, no store. When a worker calls `bridge_reply` and **no send
of **live blocking waiters only**. No queue, no store. When a worker calls `fleet_reply` and **no send
is currently open** for that worker:
- `Rendezvous.resolve(session, content)` → `complete(...)` → `waiters.get(session) == null` →
returns `false` (`msg/Rendezvous.java:212-215`).
- The `content` string is **never retained** — it is dropped. The worker is told it failed:
`BridgeMcp.reply` returns `error("no send is awaiting a reply for this worker")` (`mcp/BridgeMcp.java:270-272`);
REST returns `409 no_pending_send` (`rest/BridgedApp.java:339-345`).
`FleetMcp.reply` returns `error("no send is awaiting a reply for this worker")` (`mcp/FleetMcp.java:270-272`);
REST returns `409 no_pending_send` (`rest/FleetApp.java:339-345`).
This is the observed "communication break": a worker that finishes just after its `bridge_send` timed
This is the observed "communication break": a worker that finishes just after its `fleet_send` timed
out (the ~60s sync window) replies into the void. There is **no message-id, dedup, or ack** anywhere in
the message path today.
## 2. What to build
### 2.1 The port — `dev.ltms.bridged.msg.ReplyInbox`
### 2.1 The port — `dev.ltms.fleet.msg.ReplyInbox`
A thin interface owned by the `msg` layer. The in-memory adapter is Stage 1; the AMQP adapter (Stage 2)
implements the **same** interface, so keep it broker-agnostic.
```java
package dev.ltms.bridged.msg;
package dev.ltms.fleet.msg;
import java.util.List;
@@ -70,7 +70,7 @@ public interface ReplyInbox {
seen-set — your call; preserve insertion order).
- `peek` returns an immutable copy; `ack` removes by `msgId`. Thread-safe (concurrent publish vs. drain).
- **This is soft-state, NOT persistence.** Lost on a `java -jar` bounce — that is correct and consistent
with "bridged stays soft-state." Do **not** add any file/DB backing.
with "fleetd stays soft-state." Do **not** add any file/DB backing.
### 2.3 Publish seam — route reply through the service layer
@@ -79,7 +79,7 @@ in `MessageService`, which already owns the `Rendezvous` and will own the `Reply
- Add `MessageService.reply(String session, String content)`:
```java
/** Route a worker's explicit bridge_reply: resolve an open send, or queue it in the inbox if none. */
/** Route a worker's explicit fleet_reply: resolve an open send, or queue it in the inbox if none. */
public boolean reply(String session, String content) {
if (rendezvous.resolve(session, content)) {
return true; // a live send took it — unchanged fast path
@@ -89,12 +89,12 @@ in `MessageService`, which already owns the `Rendezvous` and will own the `Reply
}
```
- Repoint the two callers off the bare `rendezvous.resolve(...)` onto `messages.reply(...)`:
- `BridgeMcp.reply` (`mcp/BridgeMcp.java:262-273`) — on success return a normal ack; **remove** the
- `FleetMcp.reply` (`mcp/FleetMcp.java:262-273`) — on success return a normal ack; **remove** the
`error("no send is awaiting a reply…")` branch (that case is now a successful queue).
- `BridgedApp.replyMessage` (`rest/BridgedApp.java:330-346`) — return `200` (queued) instead of
- `FleetApp.replyMessage` (`rest/FleetApp.java:330-346`) — return `200` (queued) instead of
`409 no_pending_send`.
**DO NOT touch the QUESTION path.** `bridge_ask` / `rendezvous.resolveQuestion` must keep today's
**DO NOT touch the QUESTION path.** `fleet_ask` / `rendezvous.resolveQuestion` must keep today's
`NO_WAITER` behaviour — a mid-turn question is **interactive** (the worker blocks synchronously and cannot
consume a late answer), so it must **never** be queued. Only terminal `REPLY`s go to the inbox.
@@ -110,27 +110,27 @@ The primary re-checks a worker it delegated to. Expose a drain keyed by **worker
- Add `MessageService.drainReplies(String target)`: `peek` the inbox, `ack` each returned `msgId`, hand
back the `List<InboxMessage>` (or just the contents). At-least-once: peek→deliver→ack (ack only after
the caller has them, so an in-flight failure re-surfaces them).
- **DECISION (required default): expose via the existing poll verb, keyed by target.** Extend `bridge_poll`
- **DECISION (required default): expose via the existing poll verb, keyed by target.** Extend `fleet_poll`
to accept an optional `target` (worker session) and, when present, return that worker's drained replies —
alongside a matching REST route `GET /sessions/{id}/replies`. Do **not** change `send`/`answer` semantics
(do not drain inside `send` — that conflates "deliver to worker" with "collect its mail"). Keep the
existing ticket-based `bridge_poll(ticket)` path working unchanged. If you see a cleaner surface, still
existing ticket-based `fleet_poll(ticket)` path working unchanged. If you see a cleaner surface, still
ship this default and note the alternative for review.
## 3. Config
**None for Stage 1.** The in-memory adapter is the unconditional default — wire `new InMemoryReplyInbox()`
into `MessageService` in `Bridged.main`. Do **not** add a `broker:` config block (that arrives with the
into `MessageService` in `Fleetd.main`. Do **not** add a `broker:` config block (that arrives with the
Stage-2 AMQP adapter: absent → in-memory, present → AMQP).
## 4. Acceptance criteria (what the primary will verify)
1. New `ReplyInbox` + `InboxMessage` + `InMemoryReplyInbox` in `dev.ltms.bridged.msg`.
2. `bridge_reply` with **no open send** now **succeeds and queues** (no more `error` / `409`); the reply is
1. New `ReplyInbox` + `InboxMessage` + `InMemoryReplyInbox` in `dev.ltms.fleet.msg`.
2. `fleet_reply` with **no open send** now **succeeds and queues** (no more `error` / `409`); the reply is
later retrievable and identical.
3. The queued reply is drainable by the primary keyed by target; draining **acks** it (a second drain
returns nothing); dedup by `msgId` (re-publishing the same id does not double-queue).
4. **QUESTION path unchanged** — `bridge_ask` with no open send still returns `NO_WAITER` (add/keep a test
4. **QUESTION path unchanged** — `fleet_ask` with no open send still returns `NO_WAITER` (add/keep a test
proving a question is never queued).
5. Completion/failure fallbacks unchanged.
6. Unit tests covering: `InMemoryReplyInbox` publish/peek/ack/dedup/FIFO/concurrency; `MessageService.reply`
@@ -140,7 +140,7 @@ Stage-2 AMQP adapter: absent → in-memory, present → AMQP).
## 5. Build & verification (worker side)
- Build with Maven from the worktree's `bridged/` dir. **Capture the exit code without a masking pipe**
- Build with Maven from the worktree's `fleetd/` dir. **Capture the exit code without a masking pipe**
(`mvn clean install; echo "MVN_EXIT=$?"` — never `mvn … | tail`, which hides failures).
- Read the real test totals from `target/surefire-reports/TEST-*.xml`, not from stdout scroll.
- You do **not** have IDE MCP access — do not claim `ide_diagnostics` results. The **primary** runs the
@@ -152,7 +152,7 @@ Stage-2 AMQP adapter: absent → in-memory, present → AMQP).
- **`.mcp.json` is `--skip-worktree` in your worktree — never edit, `git add`, or commit it.**
- **`wiki/` is a submodule — never run git in it; never touch it.**
- Commit only your feature changes (the new port/adapter, the `msg`/`mcp`/`rest` wiring, tests, and if
you add config wiring in `Bridged.java`). Nothing else.
you add config wiring in `Fleetd.java`). Nothing else.
- Work only inside your assigned worktree on your feature branch. The primary fast-forwards `main` after
re-gating — do not touch `main`.
- Java 25 idioms are welcome (unnamed `_` params, records). Keep the diff minimal and match surrounding style.
@@ -171,7 +171,7 @@ removal, msgId dedup, cross-restart redelivery). It is `@Tag("contract")`, so th
untouched. Run it explicitly when Docker (or a broker) is available:
```bash
cd bridged
cd fleetd
mvn -Pcontract test -Dtest=AmqpReplyInboxContractTest # local: spins a RabbitMQ Testcontainers fixture
```
+17 -17
View File
@@ -18,7 +18,7 @@ The design rests on three pieces (the shape this ticket proposes):
1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker.
2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host
presence, not a central database.
3. **A per-host gateway** — each host runs a `bridged` that owns its local herdr, registers/manages
3. **A per-host gateway** — each host runs a `fleetd` that owns its local herdr, registers/manages
its own sessions, and proxies messages to/from other hosts over the broker.
## 2. What is single-host today (the assumptions to break)
@@ -27,7 +27,7 @@ The design rests on three pieces (the shape this ticket proposes):
flowchart TB
subgraph host["Single host (today)"]
primary["primary<br/>(MCP client)"]
daemon["bridged daemon<br/>127.0.0.1:8765"]
daemon["fleetd daemon<br/>127.0.0.1:8765"]
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
herdr["herdr<br/>(local unix-socket PTY mux)"]
w1["worker pane wQ:p1"]
@@ -48,21 +48,21 @@ Three concrete bake-ins assume one host:
|---|---|---|
| **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. |
| **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
| **loopback, no authn** | `rest/BridgedApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
| **loopback, no authn** | `rest/FleetApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
## 3. Target architecture
```mermaid
flowchart TB
subgraph hostA["HOST A"]
gA["gateway = bridged A"]
gA["gateway = fleetd A"]
regA["local registry + herdr"]
primary["primary (MCP client)"]
gA --- regA
primary --- gA
end
subgraph hostB["HOST B"]
gB["gateway = bridged B"]
gB["gateway = fleetd B"]
regB["local registry + herdr"]
wb["worker panes"]
gB --- regB
@@ -96,13 +96,13 @@ host's terminals.*
delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it.
- **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored
database. Per the persistence-boundary decision (bridged is soft-state; the broker owns *message*
database. Per the persistence-boundary decision (fleetd is soft-state; the broker owns *message*
durability, not *who/where/status*), each gateway announces its local agents `(globalId, host,
status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway
builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A
stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking).
- **Per-host gateway** = today's `bridged` daemon, evolved. It already registers/manages sessions
- **Per-host gateway** = today's `fleetd` daemon, evolved. It already registers/manages sessions
and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client
that consumes its agents' inboxes and injects into local herdr, and (b) presence announce +
union-roster assembly. Evolution, not rewrite.
@@ -111,7 +111,7 @@ host's terminals.*
```mermaid
flowchart LR
send["bridge_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
send["fleet_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
lookup -->|"yes"| local["inject via local herdr<br/>(today's Injector path)"]
lookup -->|"no"| pub["publish agent.&lt;id&gt;.inbox<br/>(broker routes to owning gateway)"]
pub --> consume["owning gateway consumes<br/>→ injects into its local herdr"]
@@ -123,7 +123,7 @@ broker. A sender is oblivious to which branch it took.*
## 4. What CB-307 already provides vs. what is net-new
**CB-307 delivers the transport half** and is independently valuable on a single host: the AMQP
broker fabric, the `bridged → broker` client/adapter, at-least-once + idempotent (dedup-by-id)
broker fabric, the `fleetd → broker` client/adapter, at-least-once + idempotent (dedup-by-id)
delivery, DLQ, and delayed-retry (remind). That *is* the "proxy cross-host message" backbone;
extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental.
@@ -158,17 +158,17 @@ sequenceDiagram
participant BR as broker
participant GA as gateway A
participant P as primary (host A, MCP client)
W->>GB: bridge_reply
W->>GB: fleet_reply
GB->>BR: publish primary-bound (durable, msg id)
BR->>GA: route to A's primary inbox
Note over GA: held durably until the primary pulls
P->>GA: blocking bridge_send resolves / bridge_poll
P->>GA: blocking fleet_send resolves / fleet_poll
GA-->>P: reply (then ACK to broker)
```
*Figure 4 — the broker makes the middle hop lossless, ordered, and idempotent; the **final** hop
into the primary is still a **pull** (gateway A holds the message until the primary's blocking
`bridge_send` or `bridge_poll`). Cross-host neither improves nor worsens this — it just spans hosts.
`fleet_send` or `fleet_poll`). Cross-host neither improves nor worsens this — it just spans hosts.
This is precisely the gap CB-307 closes on one host and CB-308 stretches across hosts.*
## 6. Staging & dependencies
@@ -207,8 +207,8 @@ they overlap (notably: the envelope is no longer optional, and dedup is split by
gateway's host. This extends the single-host invariant — *identity comes from the connection,
never an argument* — across the broker: cross-host, identity comes from the key. Complements
(not replaces) per-gateway broker logins over TLS.
2. **Profiles are owned by the worker's host.** `bridge_spawn(profile, host)` resolves the name in
the *target* gateway's `bridged.yaml`. Gateways advertise their profile names in presence
2. **Profiles are owned by the worker's host.** `fleet_spawn(profile, host)` resolves the name in
the *target* gateway's `fleetd.yaml`. Gateways advertise their profile names in presence
heartbeats, so a leader sees what each host offers before spawning; an unknown name is a clear
error from the target. Secrets (base URLs, tokens) never leave the host that uses them.
3. **Repo provisioning — clone from the forge, pinned.** A cross-host spawn names the repo URL and
@@ -259,7 +259,7 @@ they overlap (notably: the envelope is no longer optional, and dedup is split by
serialize acks fleet-wide. Ordering caveat: a *return* (unroutable) arrives **before** the
confirm, so "confirmed" ≠ "routed"; the sender checks the returned-set at confirm time.
`mandatory` is false only for `BROADCAST`, where an empty group is legal silence.
10. **Queue lifecycle is session lifecycle.** `bridge_stop`/reap deletes the worker's inbox queue
10. **Queue lifecycle is session lifecycle.** `fleet_stop`/reap deletes the worker's inbox queue
(its `broadcast.*` bindings die with it — no broadcasts to the dead); `x-expires` collects
queues orphaned by a crashed gateway (long for main/orchestrator inboxes, short for workers).
Queue names carry a version suffix (`.v2`): AMQP refuses to redeclare an existing durable
@@ -275,11 +275,11 @@ they overlap (notably: the envelope is no longer optional, and dedup is split by
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
- **Control authorization — THE GATE ON U4.** Signing (§7.1) settles *who sent it*; authorization
is *who may do what*. **Cross-host spawn must not land before the minimal version exists**: a
per-host allowlist in `bridged.yaml` — beside the peer public keys — of gateway ids permitted to
per-host allowlist in `fleetd.yaml` — beside the peer public keys — of gateway ids permitted to
publish control to this host, checked against the verified signature. A few lines of config and
check; without them, any principal holding broker credentials can start processes on every host
in the fleet.
- **Key distribution & rotation:** static config (host → public key in each `bridged.yaml`) is
- **Key distribution & rotation:** static config (host → public key in each `fleetd.yaml`) is
fine at the current 2–3 host scale; rotation is manual. A refinement, not a blocker.
- **Gateway death mid-turn:** the roster reaps it by missed heartbeat, and in-flight primary-bound
messages survive by broker durability; still open is reconciling *worker* state when the dead
+7 -7
View File
@@ -7,7 +7,7 @@ the core learning any one peer's environment.
## 1. Why
`claude-bridge` is a **communication bus between heterogeneous AI agents** — its stable surface is
the protocol (`bridge_spawn / send / poll / reply / ask / list / stop / status`), and that surface
the protocol (`fleet_spawn / send / poll / reply / ask / list / stop / status`), and that surface
should stay provider-neutral. Today the daemon can only materialize one kind of peer: an
off-subscription Claude Code CLI over herdr. Everything specific to *how that peer is set up*
(`ANTHROPIC_BASE_URL`, the subscription guard, `--mcp-config`/system-prompt flags, `claude-*`
@@ -60,7 +60,7 @@ Where Claude/herdr specifics actually live today:
| `claude-<profile>-<nonce>-<seq>` naming, `WORKER_NAME` regex, orphan reap (CB-117) | `WorkerService` | **→ adapter** (naming is a herdr-label detail) |
| `GITEA_TOKEN` / `GITEA_HOST` injection (CB-302 checkpoint) | `WorkerService.spawn` | **→ adapter** + a **capability** (§6) |
| tab/pane placement, worker space, tab labels | `WorkerService.spawnInTab/spawnAsPane` via herdr `WorkspaceControl` | **→ adapter** (herdr transport detail) |
| `BridgedConfig.Worker` profile shape (`baseUrl`, `model`, `configDir`, …) | `config` | **mostly adapter-shaped** — see §5 |
| `FleetConfig.Worker` profile shape (`baseUrl`, `model`, `configDir`, …) | `config` | **mostly adapter-shaped** — see §5 |
| FSM, registry, roster, `reapIdle`/`drainAll`/`contextCap`, `rosterView` | `SessionManager` | **stays core** |
| turn/completion detection (`TurnListener`, `CompletionResolver`, `StatusPoller`, `WorkerPresence`) | `inject/` | **stays core**, but reads herdr terminal output → transport-coupled (§4b) |
| message store & routing | `msg/` | **stays core** |
@@ -118,7 +118,7 @@ hard-wire "turns come from herdr".
## 5. Config shape
`BridgedConfig.Worker` is Claude-shaped (`baseUrl`, `model`, `configDir`, `tokenEnv`). Rather than
`FleetConfig.Worker` is Claude-shaped (`baseUrl`, `model`, `configDir`, `tokenEnv`). Rather than
break existing YAML, CB-401 keeps `workers:` exactly as-is and treats those fields as the
**ClaudeCodeLauncher's** profile schema. A future peer kind adds a `kind:` discriminator
(default `"claude-code"`) selecting the launcher; unknown-kind → clear config error. No migration of
@@ -126,7 +126,7 @@ existing configs. (Jackson already ignores unknown keys, so adding `kind` is bac
```mermaid
sequenceDiagram
participant MCP as bridge_spawn (MCP/REST)
participant MCP as fleet_spawn (MCP/REST)
participant SM as SessionManager
participant L as PeerLauncher (by profile.kind)
participant T as transport (herdr)
@@ -150,7 +150,7 @@ degrades gracefully when a launcher lacks one:
| Capability | Meaning | Claude Code | Codex (likely) | Human |
|---|---|---|---|---|
| `MID_TURN_ASK` | supports `bridge_ask` rendezvous | ✓ | ? | ✗ |
| `MID_TURN_ASK` | supports `fleet_ask` rendezvous | ✓ | ? | ✗ |
| `SELF_PR` | can open its own PR at checkpoint (CB-302) | ✓ (opt-in token) | ? | ✗ |
| `WORKTREE` | can run in a provisioned git worktree | ✓ | ✓ | ✗ |
| `ORPHAN_REAP` | spawner can reconcile orphaned peers on boot | ✓ | ? | ✗ |
@@ -185,13 +185,13 @@ messages, not verified facts.
Deliverable for CB-401 Stage A — mechanical, behaviour-preserving:
1. `PeerLauncher` interface + `PeerHandle` (opaque id) + `SpawnRequest` (profile, requestedCwd,
callerCwd) + `Capability` enum, new package `dev.ltms.bridged.peer`.
callerCwd) + `Capability` enum, new package `dev.ltms.fleet.peer`.
2. `ClaudeCodeLauncher implements PeerLauncher` = today's `WorkerService`, adapted: `spawn(...)`
returns a `PeerHandle` (id = paneId), `capabilities()` declares
`MID_TURN_ASK, SELF_PR(when token), WORKTREE, ORPHAN_REAP`.
3. `SessionManager` depends on `PeerLauncher`, not `WorkerService` concretely; routing keys on
`PeerHandle.id()` (== paneId today, so zero value change).
4. `Bridged.main` wires the concrete `ClaudeCodeLauncher` behind the interface.
4. `Fleetd.main` wires the concrete `ClaudeCodeLauncher` behind the interface.
5. **No behaviour change, no config change.** Full green gate: `ide_sync` → `ide_diagnostics`
(0 errors/0 warnings) → `mvn clean install` with `MVN_EXIT` captured (no masking pipe). All
existing tests pass unchanged; add tests only for the new `PeerHandle` indirection.
+21 -21
View File
@@ -11,7 +11,7 @@ All five increments of §4 are done, including increment 5 (the §5 live checkli
## 1. Goal
Prove the [`PeerLauncher`](../bridged/src/main/java/dev/ltms/bridged/peer/PeerLauncher.java) SPI
Prove the [`PeerLauncher`](../fleetd/src/main/java/dev/ltms/fleet/peer/PeerLauncher.java) SPI
actually holds for a **non-Claude** coding agent by shipping a second, first-class in-tree
adapter: **opencode** (`opencode` 1.1.31, a provider-agnostic terminal coding agent).
@@ -56,8 +56,8 @@ flowchart TB
*Figure 1 — the two concerns tangled inside today's single launcher; CB-402 splits them.*
There is also a **Stage-A deferral** to finish: `Bridged.main` still casts
`(ClaudeCodeLauncher) workers` at the `BridgeMcp` and `BridgedApp` constructors. Those two
There is also a **Stage-A deferral** to finish: `Fleetd.main` still casts
`(ClaudeCodeLauncher) workers` at the `FleetMcp` and `FleetApp` constructors. Those two
callers only invoke `profiles()`, `defaultProfile()`, and `list()` — **all already on the
`PeerLauncher` interface**. The cast survives for one reason only: `PeerLauncher.list()`
returns `List<?>` (element type erased) while the callers use `Agent` element methods in their
@@ -68,7 +68,7 @@ roster join. Finishing the migration is therefore small and contained (§4.D).
## 3. Target design
Template-Method base + two thin adapters + a routing composite that keeps the Stage-A seam
(one `PeerLauncher` reference held by `SessionManager` / `BridgeMcp` / `BridgedApp`) intact.
(one `PeerLauncher` reference held by `SessionManager` / `FleetMcp` / `FleetApp`) intact.
```mermaid
flowchart TB
@@ -114,7 +114,7 @@ Two adapter **hooks** (abstract):
protected abstract String namePrefix(); // "claude" | "opencode"
/** Build the peer-specific launch: env map + argv. Runs any pre-spawn guard here. */
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg, SpawnRequest req);
protected abstract Launch buildLaunch(FleetConfig.Worker cfg, SpawnRequest req);
record Launch(Map<String,String> env, List<String> argv) {}
```
@@ -126,7 +126,7 @@ so the opencode adapter never reaps a `claude-*` pane and vice-versa. The compos
### B. `kind:` config discriminator
Add one field to `BridgedConfig.Worker`:
Add one field to `FleetConfig.Worker`:
```java
String kind // "claude-code" (default) | "opencode"
@@ -138,7 +138,7 @@ String kind // "claude-code" (default) | "opencode"
so the record stays declarative.)
- Keep the existing back-compat constructors; `kind` is additive and optional.
`bridged.example.yaml` documents a two-kind `workers:` block.
`fleetd.example.yaml` documents a two-kind `workers:` block.
### C. `OpenCodeLauncher` — the adapter hooks for opencode
@@ -164,13 +164,13 @@ String kind // "claude-code" (default) | "opencode"
### D. `CompositePeerLauncher` + finish the Stage-A migration
- `Bridged.main` groups configured profiles by `kind`, instantiates one launcher per kind
- `Fleetd.main` groups configured profiles by `kind`, instantiates one launcher per kind
present, and wraps them in `CompositePeerLauncher implements PeerLauncher`.
- Routing methods (`spawn(req)`, `effectiveCwd(req)`, `parityOverlay(name)`) dispatch by the
profile's kind. Fan-out methods (`list()`, `reapOrphanWorkers()`, `capabilities()`,
`profiles()`, `defaultProfile()`) merge across sub-launchers. `stop(id)` tries each (teardown
only knows the pane id) — already best-effort/idempotent.
- **Migrate `BridgeMcp` + `BridgedApp` to the `PeerLauncher` interface**, dropping both
- **Migrate `FleetMcp` + `FleetApp` to the `PeerLauncher` interface**, dropping both
`(ClaudeCodeLauncher)` casts. Only friction is `list()`'s `List<?>`; resolve by giving the SPI
a typed roster element (small neutral `PeerAgent` view exposing `id()`/`name()`/status) that
the CB-304 roster join consumes — or, minimally, narrow at the callsite. Prefer the typed view.
@@ -179,12 +179,12 @@ String kind // "claude-code" (default) | "opencode"
sequenceDiagram
autonumber
participant P as Primary
participant M as BridgeMcp / REST
participant M as FleetMcp / REST
participant C as CompositePeerLauncher
participant O as OpenCodeLauncher
participant B as HerdrPeerLauncher (base)
participant H as herdr
P->>M: bridge_spawn(profile="oc-impl")
P->>M: fleet_spawn(profile="oc-impl")
M->>C: spawn(SpawnRequest)
C->>C: kind(profile)=="opencode"
C->>O: spawn(req)
@@ -208,7 +208,7 @@ sequenceDiagram
`ClaudeCodeLauncher` extend it with `namePrefix()="claude"` and `buildLaunch()` wrapping
today's guard+env+argv logic. Green build, identical tests — pure refactor. *(IDE
`refactor` where possible; the primary re-runs the gate workers can't.)*
2. **`kind:` discriminator.** Add the field + normalization + `bridged.example.yaml`. Default
2. **`kind:` discriminator.** Add the field + normalization + `fleetd.example.yaml`. Default
path unchanged (`kind=claude-code`).
3. **`OpenCodeLauncher`.** Implement the three hooks; unit-test `buildLaunch` (env has no
`ANTHROPIC_BASE_URL`; `OPENCODE_CONFIG` points at a file carrying the bridge MCP block +
@@ -226,11 +226,11 @@ until it is "major" (Stage-B whole), per the CB-401 bar.
- **opencode TUI ⇄ herdr injection.** herdr drives a pane by typing into a TUI. Must confirm
opencode's TUI accepts injected keystrokes/submit the way `claude` does, and reaches an
`injectable` status the CB-306 gate recognizes. *Validation:* spawn one opencode worker,
watch the readiness gate pass, `bridge_send` a trivial task.
watch the readiness gate pass, `fleet_send` a trivial task.
- **Bridge MCP visibility in opencode.** Confirm `OPENCODE_CONFIG` (or `opencode mcp add`)
actually surfaces the `bridge_*` tools inside the opencode session, and that `bridge_reply`
actually surfaces the `fleet_*` tools inside the opencode session, and that `fleet_reply`
is callable — the reply-charter is worthless if the tool isn't mounted. *Validation:* the
worker completes a task by calling `bridge_reply`; the reply lands via the CB-307 path.
worker completes a task by calling `fleet_reply`; the reply lands via the CB-307 path.
- **opencode MCP/config schema drift.** opencode is fast-moving (1.1.31 today). Pin the config
schema we generate against the installed version; treat the exact keys (`type: "remote"` vs
`"http"`, `instructions` shape) as a dogfood-verified fact, not an assumption.
@@ -244,7 +244,7 @@ until it is "major" (Stage-B whole), per the CB-401 bar.
- **Unit (hermetic):** base-extraction regression (existing `ClaudeCodeLauncher` tests pass
unchanged); `OpenCodeLauncher.buildLaunch` env/argv/config assertions; `kind` normalization
in `BridgedConfigTest`; `CompositePeerLauncher` routing + fan-out (merge of `profiles()`,
in `FleetConfigTest`; `CompositePeerLauncher` routing + fan-out (merge of `profiles()`,
summed `reapOrphanWorkers()`, per-kind reap isolation) with fake sub-launchers.
- **Live (dogfood, manual):** the §5 checklist on the running daemon.
- **Gate (primary):** IDE diagnostics 0/0 on every changed file, `mvn clean install` green with
@@ -269,7 +269,7 @@ until it is "major" (Stage-B whole), per the CB-401 bar.
## 8. As-built — live dogfood (2026-07-29)
Run against `bridged` on `127.0.0.1:8766` at main `19cdf8d`, with opencode **1.18.5** installed
Run against `fleetd` on `127.0.0.1:8766` at main `19cdf8d`, with opencode **1.18.5** installed
via Homebrew. Every §5 risk is now a verified fact rather than an assumption.
**The version-drift risk was the real one, and it did not bite.** This adapter was designed against
@@ -281,7 +281,7 @@ as a dogfood-verified fact for 1.18.5.
| §5 risk | Result |
|---|---|
| opencode TUI ⇄ herdr injection; CB-306 gate | ✅ `peer pane=wD:p3 reached injectable state` ~0.6s after `agent.start` |
| Bridge MCP visible + `bridge_reply` callable | ✅ MCP `initialize` from `Implementation[name=opencode, version=1.18.5]`; worker replied through the tool |
| Bridge MCP visible + `fleet_reply` callable | ✅ MCP `initialize` from `Implementation[name=opencode, version=1.18.5]`; worker replied through the tool |
| Config schema drift (1.1.31 → 1.18.5) | ✅ unchanged, see above |
| Provider credentials | ✅ free tier, zero credentials |
@@ -291,7 +291,7 @@ Full lifecycle exercised through the REST surface:
to `OpenCodeLauncher` (`spawning opencode profile=opencode-free`), pane `wD:p3`.
2. Readiness: `{"ready":true,"status":"idle"}`, roster state `ready`.
3. `POST /sessions/{id}/message` → **`{"replySource":"reply","reply":"391"}`** — a *structured*
`bridge_reply`, not the CB-115 completion-fallback transcript scrape. The clean path.
`fleet_reply`, not the CB-115 completion-fallback transcript scrape. The clean path.
4. `DELETE /workers/wD:p3` → `204`, roster empty, tolerant teardown (`tab_not_found` ignored —
opencode had already closed its own tab).
@@ -301,5 +301,5 @@ identity (loopback peer PID → herdr pane) classified an **opencode** process a
opencode-specific handling — confirming the identity model is peer-kind-agnostic, which is exactly
what CB-308 needs when it stretches the roster across hosts.
CB-502 counters for the same run: `bridged_sends_total{outcome="replied"} 1`,
`bridged_replies_total{path="rendezvous"} 1`, `bridged_inbox_depth{...} 0`.
CB-502 counters for the same run: `fleet_sends_total{outcome="replied"} 1`,
`fleet_replies_total{path="rendezvous"} 1`, `fleet_inbox_depth{...} 0`.
+60 -38
View File
@@ -1,6 +1,8 @@
# CB-500 — Multi-Tier Coordination (Stage 6)
**Status:** design note (proposal — ticket split deferred)
**Status:** design note. Developments A/B remain proposals; Development C (§6 and Figures 7–8) is
**SUPERSEDED** by the advisory-architect design in Gitea issue #16 and the `architects:` configuration
block (CB-548).
**Depends on:** CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn),
CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate),
CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304
@@ -16,8 +18,9 @@ workers) into a **multi-tier** one, along three axes the lead has asked for:
1. **Sandboxed workers** — each worker runs in a **separated, peer-owned sandbox** carrying its own
toolchain (Claude routed via `ANTHROPIC_BASE_URL`, a headless IDE, git, MCP, dev-tools), with
**per-role** sandboxes (a backend-agent image, a frontend-agent image).
2. **Main-agent pairs** — the "main" tier becomes a **pair** (on-subscription Opus + one cloud
module) collaborating, instead of a lone primary.
2. **Main-agent pairs** — **SUPERSEDED.** The considered model made the "main" tier a pair
(on-subscription Opus + one cloud module). The actual fleet is one human-driven lead plus two
short-lived advisory architects on different model families.
3. **An orchestrator tier** — a supervisor **above** the mains that owns their **session identity**
(naming, resume) and **curates context**, so every main→worker delegation carries the *exact*
slice of context it needs and nothing else.
@@ -35,14 +38,14 @@ never toolchain ownership (§7).
flowchart TB
human["human (types)"]
primary["PRIMARY (Opus)<br/>MCP client — pull-only"]
daemon["bridged daemon<br/>127.0.0.1:8765 (single host)"]
daemon["fleetd daemon<br/>127.0.0.1:8765 (single host)"]
comp["CompositePeerLauncher<br/>routes by kind"]
cc["ClaudeCodeLauncher"]
oc["OpenCodeLauncher"]
w1["worker pane (gx00 vLLM)"]
w2["worker pane (ollama)"]
human --> primary
primary -->|"bridge_send / spawn / ask"| daemon
primary -->|"fleet_send / spawn / ask"| daemon
daemon --> comp
comp --> cc
comp --> oc
@@ -63,6 +66,10 @@ Four concrete bake-ins assume a single tier:
## 3. Target multi-tier architecture
> **SUPERSEDED fleet sketch.** Figure 2 records the former two-main model. The actual fleet is one
> lead, two independent advisory architects, and N workers; architects are sideways peers, not leads
> and not a tier above the lead. See Gitea issue #16 and the `architects:` block.
```mermaid
flowchart TB
human["human"]
@@ -73,7 +80,7 @@ flowchart TB
m1["main A: Opus<br/>MCP client"]
m2["main B: cloud module<br/>MCP client"]
end
subgraph bus["bridged fabric (CB-307/308 substrate)"]
subgraph bus["fleetd fabric (CB-307/308 substrate)"]
chan["per-agent inbox channels<br/>agent.&lt;globalId&gt;.inbox"]
roster["federated roster (union view)"]
end
@@ -93,10 +100,9 @@ flowchart TB
chan --- roster
```
*Figure 2 — three tiers. Tier 0 owns the mains' session identity + context scope; Tier 1 is a
collaborating pair, each an MCP client with its own pull inbox; Tier 2 is peer-owned sandboxes the
bus launches into. The middle is CB-308's per-agent-channel + federated-roster substrate, now
carrying tier-to-tier traffic, not just host-to-host.*
*Figure 2 — **SUPERSEDED historical fleet sketch.** It proposed a collaborating pair of managed mains.
The actual fleet keeps one human-driven lead and uses two independent, short-lived advisory architects
on different model families, so agreement is evidence rather than correlated echo.*
The recursion is the key idea: **`orchestrator : mains :: main : workers`** — the same
spawn/name/resume/scope verbs at two levels.
@@ -140,11 +146,11 @@ env-manager" rule (§7).*
```mermaid
sequenceDiagram
participant M as main (delegator)
participant D as bridged
participant D as fleetd
participant SL as SandboxLauncher
participant SB as sandbox (peer-owned)
participant A as agent in sandbox
M->>D: bridge_spawn(profile=backend, role=backend)
M->>D: fleet_spawn(profile=backend, role=backend)
D->>SL: spawn(SpawnRequest)
SL->>SB: start entrypoint (image = backend role)
Note over SL,SB: bridge injects guarded ANTHROPIC_BASE_URL,<br/>mounts bridge MCP url + reply charter
@@ -166,6 +172,11 @@ gateway, because herdr keystroke-injection needs a locally-owned PTY.**
## 5. Development B — Main-agent pairs
> **SUPERSEDED — do not implement this model.** The two-main fleet was replaced by one human-driven
> lead and two independent advisory architects. They are deliberately different model families (Claude
> Sonnet 5 and GPT-5.6 through opencode), receive the same brief, and work independently so agreement
> is evidence rather than correlated echo. See Gitea issue #16 and `architects:`.
Both mains are MCP **clients**, so **neither can be called into** — each needs a **pull-based
per-agent inbox**, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery
that is singular today (single-slot `PrimaryRegistry`, a push-loop aimed at one terminal, "these
@@ -178,7 +189,7 @@ flowchart TB
m2["main B: cloud module<br/>MCP client (pull-only)"]
end
reg["PrimaryRegistry → multi-slot<br/>(terminal per main)"]
subgraph fabric["bridged"]
subgraph fabric["fleetd"]
ca["agent.A.inbox"]
cb["agent.B.inbox"]
push["ReplyPushLoop → N terminals"]
@@ -200,14 +211,14 @@ transport is already peer-neutral — what was missing is N pull endpoints).*
```mermaid
sequenceDiagram
participant MA as main A (Opus)
participant BR as bridged / broker
participant BR as fleetd / broker
participant MB as main B (cloud)
MA->>BR: bridge_send(to = main B, msg)
MA->>BR: fleet_send(to = main B, msg)
BR->>BR: publish agent.B.inbox (durable, msg id)
Note over BR: held until B pulls (B is a client too)
MB->>BR: blocking bridge_send / poll resolves
MB->>BR: blocking fleet_send / poll resolves
BR-->>MB: msg (then ACK)
MB->>BR: bridge_reply(to = main A)
MB->>BR: fleet_reply(to = main A)
BR->>BR: publish agent.A.inbox
MA->>BR: poll resolves
BR-->>MA: reply
@@ -221,6 +232,19 @@ push-loop fan-out; relax "orchestration tools only the primary calls" to "any re
## 6. Development C — Orchestrator tier
> **SUPERSEDED — do not implement this model.** The operator rejected a supervisor above the lead.
> The human continues to drive the pre-existing lead directly; fleetd neither spawns nor resumes that
> lead. What replaced this proposal is **one lead, two short-lived advisory architects, and N workers**:
> the lead engages architects sideways for a strong-model assessment, then discards them. Architect
> slots are declared in `architects:` (see Gitea issue #16), rather than making leads managed sessions.
> The two architects deliberately use different model families — Claude Sonnet 5 and GPT-5.6 through
> opencode — and receive the same brief independently. Agreement is evidence, not correlated echo
> from one provider or one conversation.
> **Historical alternative retained.** The text and figures below record the considered model and why it
> was rejected: it re-rooted the human-facing session above the lead, violating the still-true premise
> that configured leaders pre-exist, are recognised, and cannot be resumed by fleetd.
The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker*
sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**.
@@ -249,11 +273,10 @@ flowchart TB
m2 -->|"scoped delegation"| w
```
*Figure 7 — the recursion. Tiers 1 and 2 run the identical spawn/name/resume machinery; the
orchestrator merely operates it one level higher. **Re-rooting caveat:** today the primary IS the
human's live session; here the human drives the orchestrator, and the mains become managed,
resumable sessions. That moves the human-facing top up a tier — an intentional re-root, not an
add-on.*
*Figure 7 — **SUPERSEDED historical alternative.** The recursion re-rooted the human-facing session:
the human drove an orchestrator and the mains became managed, resumable sessions. The operator rejected
that re-root. The replacement keeps the human-driven, pre-existing lead and engages architects sideways
as short-lived advisory peers; see Gitea issue #16 and `architects:`.*
```mermaid
sequenceDiagram
@@ -264,16 +287,15 @@ sequenceDiagram
H->>O: high-level goal (large context)
O->>O: name/resume main A session
O->>MA: task + SCOPED context slice (not the whole history)
MA->>W: bridge_send(delegation, carrying only the relevant slice)
MA->>W: fleet_send(delegation, carrying only the relevant slice)
W-->>MA: result
MA-->>O: rollup
O->>O: fold into orchestrator context, pick next main/turn
```
*Figure 8 — context focus. The orchestrator holds the global context and hands each main only the
slice a given delegation needs, so the main→worker conversation stays on-point. Context *scoping* is
coordination (the bus already owns session/turn lifecycle) — it stays inside the identity boundary
(§7), unlike toolchain ownership which does not.*
*Figure 8 — **SUPERSEDED historical alternative.** This proposed an orchestrator holding global context
and slicing it for managed mains. The replacement has the human-driven lead send the same advisory brief
issue #16 and `architects:`.*
**Deltas:** a second, higher `SessionManager` instance whose "peers" are mains; the orchestrator
becomes the top MCP client; context-slice selection (new) layered on CB-303's `context_cap` +
@@ -315,7 +337,7 @@ flowchart LR
cb402["CB-401/402<br/>Peer Launcher SPI + composite<br/>(DONE / in-flight)"]
A["A · SandboxLauncher<br/>(placement-neutral, independent)"]
cb308["CB-308 substrate<br/>per-agent channels + global id<br/>+ federated roster"]
B["B · main-agent pair<br/>(multi-slot PrimaryRegistry)"]
B["B · main-agent pair (SUPERSEDED)<br/>(multi-slot PrimaryRegistry)"]
C["C · orchestrator tier<br/>(SessionManager recursed up)"]
cb402 --> A
cb402 --> cb308
@@ -333,7 +355,7 @@ flowchart LR
2. **A · SandboxLauncher** — independent; a second proof of the SPI (placement-neutral). Ships anytime.
3. **CB-308 substrate** — per-agent channels + global id + federated roster (the multi-host work,
promoted from host-to-host to tier-to-tier).
4. **B · main-agent pair** — multi-slot `PrimaryRegistry` + per-main inbox routing, on the substrate.
4. **B · main-agent pair** — **SUPERSEDED** by lead + two advisory architects.
5. **C · orchestrator tier** — the capstone; the recursive session manager + context scoping.
## 9. Open questions (to resolve at ticket-split)
@@ -341,8 +363,8 @@ flowchart LR
- **Sandbox mechanism:** container (`docker exec`) vs devcontainer — how the role→image mapping is
expressed on the profile. *(Topology **resolved** in §11: distributed = gateway-per-host × local
sandboxes; the remaining choice is only the local launch mechanism, not the shape.)*
- **Pair semantics:** are the two mains fully symmetric peers, or is one a co-primary that may also
delegate? Affects how `PrimaryRegistry` and the "orchestration tools" identity relax.
- **Pair semantics:** **SUPERSEDED.** The two-main question is replaced by the architect role's
least-privilege boundary: advisory architects can send/reply/ask/read but cannot spawn/stop/drain.
- **Orchestrator drivenness:** the mains become programmatically spawned/resumed — does the human
still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.)
- **Context-slice selection:** who decides the slice — orchestrator heuristics, explicit tool args,
@@ -363,7 +385,7 @@ ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C).
The follow-up question — *"clarify the architecture when we have distributed agents in sandboxes"* —
resolves the fork left open in §4 and §9. **Decision: Development A (sandbox launcher) and CB-308
(per-host federation) *compose*, not compete — each host runs a `bridged` gateway whose launcher
(per-host federation) *compose*, not compete — each host runs a `fleetd` gateway whose launcher
spawns agents into that host's *local* sandboxes.** A sandbox is never reached across the network; it
is reached by the gateway sitting next to it.
@@ -371,7 +393,7 @@ is reached by the gateway sitting next to it.
The bus delivers a turn by **herdr keystroke-injection** — `Injector → AgentControl.send` writes into
a PTY that its **local** herdr owns. The broker moves *messages and presence*, **never keystrokes**.
So an agent's PTY must live in a herdr that *some* `bridged` instance drives locally: a remote
So an agent's PTY must live in a herdr that *some* `fleetd` instance drives locally: a remote
container with no local herdr **cannot be injected into**. That rules out a central daemon reaching
remote PTYs, and collapses the design to a single identity:
@@ -381,7 +403,7 @@ remote PTYs, and collapses the design to a single identity:
flowchart TB
subgraph hostA["HOST A — gateway"]
mA["main / orchestrator<br/>MCP client → LOCAL gateway"]
gA["bridged A<br/>herdr + CompositePeerLauncher<br/>(incl. SandboxLauncher)"]
gA["fleetd A<br/>herdr + CompositePeerLauncher<br/>(incl. SandboxLauncher)"]
cBEa["sandbox: backend<br/>(local container)"]
cFEa["sandbox: frontend<br/>(local container)"]
mA --- gA
@@ -393,7 +415,7 @@ flowchart TB
roster["roster.* (federated presence)"]
end
subgraph hostB["HOST B — gateway"]
gB["bridged B<br/>herdr + SandboxLauncher"]
gB["fleetd B<br/>herdr + SandboxLauncher"]
cBEb["sandbox: backend<br/>(local container)"]
gB -->|"spawn → PTY in B's herdr"| cBEb
end
@@ -418,13 +440,13 @@ sequenceDiagram
participant BR as broker
participant GB as gateway B
participant SB as sandbox agent (host B, container)
MA->>GA: bridge_send(globalId on B, msg)
MA->>GA: fleet_send(globalId on B, msg)
GA->>GA: directory lookup - is globalId local? NO
GA->>BR: publish agent.ID.inbox (durable)
BR->>GB: route to the owning gateway
GB->>SB: inject via B's LOCAL herdr (keystrokes)
Note over GB,SB: SandboxLauncher already spawned the container -<br/>its PTY is in B's herdr, CB-306 readiness passed
SB-->>GB: bridge_reply (to B's LOCAL MCP endpoint)
SB-->>GB: fleet_reply (to B's LOCAL MCP endpoint)
GB->>BR: publish primary-bound (durable, msg id)
BR->>GA: route back to A
Note over GA: held until the main pulls (the main is a client)
@@ -448,7 +470,7 @@ unchanged — the sandbox is transparent to it.*
| Concern | Provided by |
|---|---|
| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `bridged`, evolved) |
| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `fleetd`, evolved) |
| Spawn into a local sandbox / role→image | **Development A** `SandboxLauncher` (§4), routed by `CompositePeerLauncher` |
| Addressing a remote sandboxed agent | **CB-308** global id + federated roster (host + role as metadata) |
| Orphan reap after a gateway restart | **CB-117** per-gateway, summed by the composite — each reaps only its **local** herdr |
+485
View File
@@ -0,0 +1,485 @@
# CB-591 — move the fleet onto the LLM and MCP gateway
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
with no terminator, and our third-party members cannot detect it (§7.2).
> **Superseded in part — 2026-09-13.** Two claims on this page are no longer true of the live fleet.
> I measured both on this host today.
>
> 1. **The model is named `acoder` now, not `deepseek-v4-flash`.** `acoder` is a stable alias, and
> the model behind it changed on 2026-08-28: it is Qwen3.8-27B, not DeepSeek. The old name is
> still served, so nothing broke — the gateway answers it and reports `"model": "acoder"` in the
> reply, which is how you can see for yourself that it is an alias. `fleetd.yaml` moved to
> `acoder` on 2026-09-13. Do not guess behaviour from the name; ask the gateway's own manifest,
> `GET https://llm.ltms.dev/v1/deployment`, and read its `generation` field.
> 2. **`local` sits at `weight: 0`, not 100.** Only `gx` is auto-selected today.
>
> §2 and §3 below are the plan as written in August. They are the record of the migration, so they
> stay as they are. If this note stops matching `fleetd.yaml`, re-measure and rewrite the note.
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
**member definition** in `fleetd.yaml`, because that is the part of this repo the change actually
touches.
---
## 1. What changed upstream
One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer.
| Surface | URL |
|---|---|
| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` |
| OpenAI models | `https://llm.ltms.dev/v1/models` |
| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` |
| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` |
Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must
never become a blanket proxy.
The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly
`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch.
---
## 2. Where claude-bridge sits today
We do **not** use the gateway. The `local` profile talks straight to the vLLM:
```yaml
local:
kind: claude-code
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
model: deepseek-v4-flash
configDir: /Users/dai.ha/.ccs/instances/gx10
```
Three facts about our side that decide the shape of this work:
1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes
`ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config).
`local` sets no `tokenEnv` today, because a direct vLLM needs no token.
2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is
`guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Fleetd.java:93` and handed to the
launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing
this makes every `local` spawn throw.
3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is
blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its
own consumer first."
Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in
`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members).
```mermaid
flowchart LR
subgraph now["Today"]
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
P1["primary"] --> CT1
end
subgraph after["Proposed"]
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
G --> V2["vLLM gx00.gw:8000"]
M4["local-direct<br/>weight 0, escape hatch"] --> V2
end
```
*The member definition is the only thing that moves. The model behind it does not.*
---
## 3. The member definition change
The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at
it**. That is the main opportunity here, and it is bigger than the `local` profile alone.
### 3a. `local` — claude-code, on `/anthropic`
| Key | Today | After | Note |
|---|---|---|---|
| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below |
| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` |
| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working |
| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** |
### 3b. A new opencode profile on `/v1` — no code needed
`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it
writes a custom provider block into the worker's opencode config:
- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that
already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged.
- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it).
- `model:` **must** be `<provider>/<model>` when `baseUrl` is set — a bare name is rejected loudly
rather than silently falling back to opencode's default gateway.
So the profile is pure config:
```yaml
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
tokenEnv: AI_GATEWAY_TOKEN
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
argv: ["opencode"]
mcpUrl: http://127.0.0.1:8765/mcp
gitTokenEnv: WORKER_GITEA_TOKEN
weight: 100 # same tier as `local` — free
maxLoad: 2
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
```
**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those
are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
a single point of failure rather than just adding capacity.
**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the
guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its
own provider credentials. So 3b needs **no allowlist change**; only 3a does.
**Both still need a restart, for a different reason.** `tokenEnv` is resolved by
`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process
environment**. The running `fleetd` inherited its environment when it started, so a variable added to
`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the
gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-fleetd.sh`
(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use
`scripts/redeploy-fleetd.sh --check` to confirm the name resolves before restarting anything.
### 3c. What this does to `ccs`
Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what
routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust**
and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error
recorded in `fleetd.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side
state*. Less load-bearing, not removable.
### Why `/anthropic` and never `/v1/chat/completions`
The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no
translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
directly.
Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not
send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code
speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice.
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
the API before any profile was switched:
| surface | request | result |
|---|---|---|
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
refuse an unauthenticated request here.
---
## 4. Decisions
### D1 — switch, but keep the direct path as an explicit profile · **recommended**
Switching buys four things we do not have:
- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires
a real single point of failure, not just a cost line.
- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real
measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/fleet/fleetd/issues/74) Gap 2 directly.
- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared
`legacy` token used by four systems.
- **It works off-LAN.** `gx00.gw` resolves on the LAN only.
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of
every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks,
nothing that matters is blocked."
So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`**
— never auto-selected, still spawnable with an explicit `fleet_spawn{profile: "local-direct"}`.
That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the
lead can actually reach during an incident.
### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call**
Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow.
- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful
for library work. But `mcpUrl` in `FleetConfig.Profile` is a **single `String`**, so a
claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
defensible; picking one is not mine to do.
### D3 — token scope
One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`,
referenced by name only. Never the literal value in `fleetd.yaml` — `tokenEnv` exists for this.
---
## 5. Units of work
```mermaid
flowchart TB
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
U2["U2 · profile + guard<br/>fleetd.yaml, restart"]
U3["U3 · verify live<br/>spawn, prove thinking survives"]
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
U1 --> U2 --> U3
U4 --> U5
U3 --> U5
```
| # | Scope | Who | Why |
|---|---|---|---|
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `fleetd.yaml` is gitignored, so a worker cannot see or edit it |
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
in that order separates the two causes instead of confusing them.
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
---
## 6. Traps carried over from the wiki
Each of these cost someone real debugging time upstream. They apply to us.
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
(`fleet_list` → `fleet_poll` anything wanted → `fleet_stop`), then rotate.
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
nothing else. Never reason as if the gateway authenticates.
3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
means the model match is a regex. Do not conflate them when diagnosing.
---
## 7. Verification — what would prove this works
Merging config is not proving it. The checks, in order:
1. `fleet_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
`fleet_reply`. This is the first proof of the token, the URL and the model name, and it risks
nothing the fleet depends on.
2. `fleet_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
for the same reason.
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
and the gateway.
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
`legacy`. That is the whole point of taking our own token.
6. `fleet_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
theoretical.
7. `fleet_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
---
## 7.1 What the live run actually found — 2026-08-15
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
### The blocker
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
surfaces:
```
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
```
32 KiB is far below one real agent turn.
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
request body before it can route on the model name — so that default is not a network tuning knob
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
```
listener default/llm/http per_connection_buffer_limit_bytes: 32768
```
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
answers 200.
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
system. Fixed in **systems/vms**, not here.
### The part worth remembering
Two members were spawned at the same moment with the same message:
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|---|---|---|
| READY → BUSY | 19:07:26 | 19:07:45 |
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
first turn that reads a file.
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
a real file. A liveness probe proves the token and the URL; it does not prove the path.
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
launcher's own generated config:
```
Error: Request Entity Too Large
...compacts context, retries...
Error: Request Entity Too Large
```
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
one turn; this one costs the whole task and is indistinguishable from a slow worker.
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
> `ls -dt /var/folders/*/*/T/fleetd-opencode-* | head -1`, check the provider block and the key's
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
### What checked out, and needs no re-testing
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
- `/v1/models` returned exactly `["deepseek-v4-flash"]` **on 2026-08-15**, so trap 3 was clear
then. It returns 6 ids now — `acoder`, `qwen3.8-27b-nvfp4`, `deepseek-v4-flash` and three
embedding names — measured on this host 2026-09-13. The exact-name rule still holds; the
one-item list does not.
- **Reasoning survives both surfaces** — see §3b above.
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
key rather than the `fleetd-local-noauth` placeholder.
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
spawned without throwing, which is the check that catches a missed restart.
### Resolution — both ceilings fixed, migration completed
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
| ceiling | was | now | our own check |
|---|---|---|---|
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
reaper for dead connections while putting the truncation timer out of practical reach.
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
Two configuration facts worth keeping, from their bisection:
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
shape as their `SecurityPolicy` caveat.
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
## 7.2 The risk we accepted, and why we could not remove it
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
resetting the connection, so the client receives what looks like a complete transfer
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
Measured on our side while the timeout was still 60s:
```
HTTP 200 61.07s 141992 bytes
message_stop 0 message_delta 0 error events 0
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
```
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
timeout, `max_stream_duration` — fails this same way.
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
So the honest statement of our position:
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
> trip the bug — **not** because we could detect it if it did.
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
handling. That ambiguity is now gone.
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
code.** That is the whole reason this section exists.
---
## 8. Related
- [CB-589 / #74](https://git.ltms.dev/fleet/fleetd/issues/74) — cost-first placement and a
gateway that reports live capacity. The per-consumer figures this migration unlocks are the first
input that ticket actually needs.
- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is
part of.
+18 -18
View File
@@ -12,12 +12,12 @@ not after it.
## 1. Why this stage is not optional bookkeeping
`bridged` today has **exactly one security control: the loopback bind**. Every other guarantee
`fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee
rests on it.
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `bridged` host); the
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the
token path is the split-host fallback."* The token path does not exist yet.
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
@@ -67,7 +67,7 @@ it will be disabled and the stage is wasted. So:
auth:
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
# mode: token # every non-worker caller must present a valid bearer token
# tokenEnv: BRIDGED_API_TOKEN # host env var holding the token; never the literal value
# tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value
```
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
@@ -112,9 +112,9 @@ The roadmap says "systemd unit". **This host is macOS — there is no systemd on
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
`java -jar`. Ship **both**:
- `deploy/dev.ltms.bridged.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
- `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
ordered start after herdr.
- `deploy/bridged.service` — systemd unit for the Linux gateways CB-308 introduces.
- `deploy/fleetd.service` — systemd unit for the Linux gateways CB-308 introduces.
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
@@ -158,10 +158,10 @@ configured. It already leaks nothing but herdr's version and up/down.
### 3.1 There are TWO entry paths, and only one of them has identity today
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
literally true, and the difference is security-relevant.** `BridgeMcp` calls `MessageService` /
literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` /
`SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
mounted as a raw servlet on Jetty's `ServletContextHandler`
(`BridgedApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
(`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
Javalin's `before` filters at all.
The current split is the mirror image of what you'd expect:
@@ -172,7 +172,7 @@ The current split is the mirror image of what you'd expect:
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
path, whereas the MCP `bridge_reply` derives the worker from the connection and refuses to read it
path, whereas the MCP `fleet_reply` derives the worker from the connection and refuses to read it
from an argument. Loopback-only bind is what makes this safe today.
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
@@ -188,15 +188,15 @@ Deliberately small; every one maps to a failure mode we have actually hit.
| Metric | Type | Why it exists |
|---|---|---|
| `bridged_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
| `bridged_send_duration_seconds` | histogram | delegated turn latency |
| `bridged_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `bridged_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `bridged_push_nudges_total{outcome}` | counter | outcome ∈ delivered\|exhausted — a rising `exhausted` means the primary is not draining |
| `bridged_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
| `bridged_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `bridged_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
| `bridged_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
| `fleet_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
| `fleet_send_duration_seconds` | histogram | delegated turn latency |
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it |
| `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
| `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
| `fleet_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
---
@@ -226,7 +226,7 @@ was answered from the forge itself. Kept here as the decision record.*
1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon;
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
`amqps://` URI. No keystore handling in `bridged`.
`amqps://` URI. No keystore handling in `fleetd`.
2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the
reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this
session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the
+972
View File
@@ -0,0 +1,972 @@
# M4 - Fleet health, recovery, routing, and capacity
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
the `fleet_list` capacity view; the remaining M4 units are not yet shipped. See
[Unit 2 — what has landed so far](#unit-2---what-has-landed-so-far) before planning Unit 2 work:
some of its criteria were met by separate CB tickets, and one of them contradicts the unit text.
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
human notification.
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
`inject/StatusRefiner`, `inject/CompletionResolver`, `inject/Injector`, `session/SessionManager`,
`msg/MessageService`, `msg/ReplyInbox`, `msg/ReplyPushLoop`, `msg/LeadHeartbeatLoop`,
`mcp/PrimaryRegistry`, and `herdr/AgentControl`.
## 1. Problem and decision boundary
The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken
communication. The bridge may read an agent pane from time to time. It must notify a person when
the fleet cannot move forward.
The four operator terms are not four equal health states. `IDLE` is a normal mode. An exception is
sometimes visible only as pane text. Stopped work may look the same as slow work. Broken
communication can occur on several links.
M4 uses this boundary:
- The bridge detects facts and joins evidence.
- The bridge repairs only mechanical failures with no judgement.
- The lead decides whether to stop, retry, replace, or reassign a member.
- A human is notified only when no healthy lead can act.
- n8n may route an outbound incident. It never classifies state or chooses recovery.
An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a
workflow engine lead authority is unsafe, while adding a new role is a separate authorization
design. An outbound sink needs no bridge role.
The bridge must never replay a delivered task. That task may already have changed files, pushed a
branch, opened a pull request, or changed external state. A replay can run those side effects twice.
This rule must remain true even if later code stores delivered prompt text.
## 2. Evidence model
A health state is mainly a comparison between two views:
- **herdr view:** current agents and raw live status from one `AgentControl.list()` call.
- **bridge view:** session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.
A strong fault often appears as a disagreement between those views. For example, `BUSY` in the
session FSM and `DONE` in herdr means the bridge missed a turn boundary. Pane reads support this
model, but they are not the main monitor.
`SessionManager.rosterView` already joins session state and live status. `AgentControl.list()`
already gets the whole live fleet in one call. M4 makes that join persistent and adds timers,
accepted-turn state, and incident state.
### 2.1 Real traces behind the design
The first trace was an architect that stopped making progress:
```text
profile=opus role=architect state=busy liveStatus=done
```
The session moved from `DONE` to `BUSY` for turn 2. Eighteen minutes later, the session still said
`BUSY`, herdr still said `DONE`, the async task still said `PENDING`, and no completion fallback had
run. This is `TURN_BOUNDARY_LOST`, not a general slow-turn guess.
The second trace had two async sends to the same pane, one second apart. The pane was then stopped.
One ticket became failed. The other stayed `pending - worker unknown`. Current
`MessageService.abandon` resolves only `Rendezvous.currentWaiter(target)`, while async tasks live in
a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent
post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`.
### 2.2 Corrections made during design
The first state table missed `BUSY` in fleetd plus `IDLE` or `DONE` in herdr. It would have found
the real trace only through a late, weak stall timer. The final model adds
`TURN_BOUNDARY_LOST` as a strong disagreement state.
The first notification design also required a webhook before `health.enabled` could turn on. That
removed useful local detection to avoid a narrower human-notification gap. The final design splits
detection from notification. Missing human escalation is shown as partial coverage instead of
disabling health.
## 3. Classification precedence
Evidence is applied in this order. A lower rule cannot hide a higher one.
1. **Control link:** failed fleet list plus failed ping becomes `CONTROL_LINK_DOWN`.
2. **Definitive target loss:** `_not_found` becomes `GONE` or `LEAD_UNREACHABLE` when the control
link is healthy.
3. **Startup and teardown invariants:** readiness expiry becomes `NEVER_READY`; surviving tasks
after teardown become `DELEGATION_ORPHANED`.
4. **Bridge/live disagreement:** `BUSY` plus stable raw `IDLE` or `DONE` becomes
`TURN_BOUNDARY_LOST`.
5. **Known screen evidence:** a tested fatal signature becomes `ERROR_ON_SCREEN`.
6. **Timed suspicion:** unchanged sparse pane probes may become `STALL_SUSPECTED`.
7. **Communication quality:** completion fallback becomes `MUTE`; an old inbox entry becomes
`REPLY_STRANDED`.
8. **Normal mode:** `STARTING`, `IDLE`, `WORKING`, `WORK_PENDING`, or `BLOCKED_AMBIGUOUS`.
The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are
reported beside the session FSM; most are not new FSM values.
```mermaid
flowchart TD
Registered["Member registered"] --> Starting["STARTING"]
Starting -->|"MCP presence"| Idle["IDLE"]
Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
Idle -->|"Accepted delivery"| Working["WORKING"]
Working -->|"Trusted turn boundary"| Idle
Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
Working -->|"Target not found"| Gone["GONE"]
Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
Pending -->|"Delivery or collection finishes"| Idle
Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
Lost -->|"Repair refused"| LeadDecision["Lead decision required"]
```
*Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an
automatic retry of the task.*
## 4. State model
### 4.1 Normal and transitional member states
| State | Exact evidence | Meaning and certainty |
|---|---|---|
| `STARTING` | Session is `SPAWNING`; MCP presence is absent | Normal inside the startup grace. MCP contact is the readiness signal. |
| `IDLE` | Session is `READY` or `DONE`; live status is `IDLE` or `DONE`; no open turn or inbox item exists | Normal. Idle is not a fault. |
| `WORKING` | Session is `BUSY`; raw live status is `WORKING`; the accepted turn is open | Certain that herdr sees work. It does not prove useful progress. |
| `WORK_PENDING` | Queued delivery or inbox content exists while the target is injectable | Transitional. Existing injector or push logic should move it. |
| `BLOCKED_AMBIGUOUS` | An open turn exists and raw live status is `BLOCKED` | The bridge cannot tell whether this is permission, input, or a settled screen. |
Idle may drive configured resource cleanup. It never opens an incident and never pages a person.
### 4.2 Member fault and quality states
| State | Exact evidence | Certainty and action |
|---|---|---|
| `NEVER_READY` | `SPAWNING`, no MCP presence, and an accepted delivery waits through the existing readiness grace | Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree. |
| `GONE` | Per-target herdr call returns `_not_found` while fleet list or ping works | Certain target loss. Fail all target work. Do not replay it. |
| `TURN_BOUNDARY_LOST` | Same session turn stays `BUSY`; same accepted task stays open; two raw snapshots show `IDLE` or `DONE` | Strong disagreement. Strict reconciliation may repair it. |
| `ERROR_ON_SCREEN` | Suspicious non-working state survives grace; `detection` matches a tested adapter-specific fatal signature | Certain only for the matched signature. A bare word such as `Exception` is not enough. |
| `STALL_SUSPECTED` | Open turn is older than the configured threshold; two normalised `recent_unwrapped` digests are unchanged; no boundary or reply occurs | Not certain. A long valid API call can look the same. Lead decides. |
| `MUTE` | Turn resolves through completion fallback instead of `fleet_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
| `REPLY_STRANDED` | Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
| `DELEGATION_ORPHANED` | Target is gone, failed, or released, but one or more tasks remain `PENDING` after reconciliation grace | Certain bridge invariant failure. This is not an inbox-drain fault. |
| `WORK_PRODUCT_AT_RISK` | Provisioned branch has commits after its recorded base; member is `DONE`, `FAILED`, or preserved after release; no turn or inbox item remains; long-idle threshold passed | A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
`MUTE` opens an incident only after a small fixed rate threshold for one target or profile, or when
it appears with another fault.
`WORK_PRODUCT_AT_RISK` must not become `WORK_PRODUCT_UNCOLLECTED`. The bridge does not know pull
request or merge state. If committed work appears with `REPLY_STRANDED` or
`DELEGATION_ORPHANED`, the existing incident gains `committedWorkAtRisk: true`.
### 4.3 Control-link state
| State | Exact evidence | Certainty and action |
|---|---|---|
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the fleetd-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global
link is a target fault, not a control-link fault.
### 4.4 Lead states
| State | Exact evidence | Meaning and action |
|---|---|---|
| `LEAD_IDLE` | Expected lead is present with raw injectable status; no actionable state waits | Normal. Existing heartbeat may run under its own policy. |
| `LEAD_WORKING` | Expected lead is present with raw `WORKING`; stall threshold is not met | Reachable and busy. Never inject into the live turn. |
| `LEAD_STATUS_UNKNOWN` | Expected lead is present with raw `UNKNOWN` | Neither dead nor a healthy routing target. Retain evidence and retry. |
| `LEAD_UNREACHABLE` | Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns `_not_found` with a healthy control link | Route to a healthy peer. If none exists, use human escalation. |
| `LEAD_UNRESPONSIVE` | Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected | Route to a healthy peer or a person. |
| `LEAD_STALL_SUSPECTED` | Lead stays `WORKING` past threshold; two sparse pane probes show no progress | Not certain. Never kill or restart automatically. Route to peer or person. |
The monitor retains the lead name and terminal, last successful sighting, raw status and age,
consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge
outcomes. Current heartbeat and push loops discard much of this history.
Expected lead identity comes from the same supplier used by `CallerResolver`. It is not liveness
evidence. `LeadTabScanner` keeps cached identity after a failed scan, so the health monitor compares
that identity with a fresh successful agent list. A dynamic identity also survives a two-successful-
snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.
### 4.5 Evidence limits
M4 cannot tell these cases apart with current evidence:
- A valid long call and a hung call may have the same status and pane digest.
- `BLOCKED` does not explain which input is needed.
- An idle prompt after failure may look like an idle prompt after success.
- A missing structured reply does not prove a broken MCP connection.
- An undrained inbox does not explain why the lead did not collect it.
- Arbitrary pane text cannot safely classify arbitrary exceptions.
- A branch ahead of its base does not prove that work was not collected.
Logs are outputs, not classifier inputs. The monitor never parses its own logs.
## 5. Automatic action and lead action
### 5.1 Actions the bridge may take
The bridge may:
- retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
- re-submit Enter after the existing paste/submit race;
- fail queued delivery after `NEVER_READY`;
- stop a never-ready process while preserving its provisioned worktree;
- fail all queued, accepted, and async tasks for a gone or released target;
- reconcile one lost boundary when every strict gate in Section 8 passes;
- hold typed messages, nudge the owning lead, and stop at the configured cap;
- use the existing bounded idle-lead heartbeat;
- deduplicate, route, update, and resolve incidents.
These actions do not choose new work and do not replay old work.
### 5.2 Decisions reserved for the lead
Only the lead may:
- stop or continue `BLOCKED_AMBIGUOUS`;
- stop, inspect, or wait on `ERROR_ON_SCREEN`;
- kill or continue `STALL_SUSPECTED`;
- spawn a replacement or reassign work;
- retry a delivered task;
- choose how to use partial work in a worktree;
- restart herdr or change network, model, credentials, backend, or configuration.
Reports include literal safe tool calls such as `fleet_status(sessionId="...")`,
`fleet_poll(ticket="...")`, `fleet_list()`, and optional `fleet_stop(paneId="...")`. A judgement
state never presents stop as the only action.
### 5.3 Release causes and worktree safety
| Release cause | Process action | Provisioned worktree |
|---|---|---|
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove only if clean; preserve a dirty worktree (CB-576) |
| `NEVER_READY` | Stop | Preserve |
| `GONE` | Best-effort stop | Preserve |
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
| `RELEASE_WITH_PENDING_TASKS` | Stop | Preserve |
| `SHUTDOWN` | Stop | Preserve |
Explicit stop is state-aware. `SPAWNING`, `BUSY`, `FAILED`, or any target with pending tasks uses a
preserving cause.
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
release cause, release time, state, and pending task ids. `fleet_list.preservedWorktrees` loads these
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
auto-deletes a preserved worktree.
## 6. Fleet health monitor
Add `FleetHealthMonitor`. Do not widen `StatusPoller` into a policy loop.
`StatusPoller` has a 250 ms delivery cadence and samples only injector targets with outstanding
work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One
loop cannot serve both cadences safely.
Build the monitor like `LeadHeartbeatLoop`:
- pure `decide(snapshot, priorState, now)` logic;
- a thin scheduler;
- an injected clock;
- edge-triggered state changes;
- no network work in the pure function;
- no sleeping in tests.
Each enabled fleet tick reads:
- one `AgentControl.list()` result for the whole fleet;
- one in-memory `SessionManager.roster()` snapshot;
- accepted turns and async task state;
- typed inbox depth, kind, and age;
- push and heartbeat outcomes;
- configured and discovered leads.
Existing failure paths publish structured evidence to the monitor. The monitor does not infer events
from log text.
### 6.1 Pane budget
Healthy idle members, recent working members, and quiet leads cause no pane reads.
A pane is eligible only for a stable lost boundary, sustained `BLOCKED` or `UNKNOWN`, work older
than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.
Compiled brakes apply even if config asks for more:
- per-target pane cooldown is at least 60 seconds;
- working age before the first progress probe is at least 300 seconds;
- at most two pane reads occur in one fleet tick;
- targets rotate fairly;
- only a normalised digest and optional clipped local excerpt are stored;
- no pane excerpt leaves fleetd in a human webhook.
Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress
comparison.
## 7. Typed inbox and routing
### 7.1 Semantic record
The typed inbox record carries:
```text
schemaVersion
kind: reply | health
msgId, target, subjectTerminal, recipientLead
severity, state, evidence
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
recoveryTried, suggestedToolCalls, content
```
A health message never calls `Rendezvous.resolve`. It cannot look like the member's task result.
Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages,
explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It
does not copy AMQP migration logic.
### 7.2 AMQP migration
The reader uses AMQP `content_type`, never body sniffing:
```text
Legacy v0: text/plain
Typed family: application/vnd.ltms.fleet.inbox-message+json
```
A legacy reply may begin with `{`. It remains plain text because its media type is `text/plain`.
Legacy text becomes `kind=reply` with exact UTF-8 content and absent typed metadata.
Typed JSON has required integer `schemaVersion: 1`. Version 1 ignores unknown optional fields.
Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch
are invalid data.
An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and
creates one operator-visible `unsupported_version` failure. A newer daemon may read it later.
Invalid known-format data is copied byte-for-byte to durable queue
`agent.<target>.inbox.quarantine`. A dedicated confirm-mode publisher confirms the persistent copy
before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged.
The raw body never enters logs.
Decode failure creates a redacted WARN, metric, `fleet_list` summary, and routed health incident.
One bad entry never escapes the consumer callback and never stops later valid messages.
Safe downgrade is not supported. The previous build ignores `content_type` and would show typed JSON
as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed
queues must be drained or preserved before an old jar runs.
The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key
ownership cases must run once against production LavinMQ before release, or the release must state
that LavinMQ was not checked.
### 7.3 Member routing
A member incident first goes to the exact lead that owns its accepted delegation.
`PrimaryRegistry` needs a no-fallback `delegatingLeadFor(memberTarget)` query. Health routing must not
use the old singular-primary fallback when several leads exist.
Publish the incident under the affected member target. Trigger the existing bounded push route. The
push waits until the owning lead is injectable, so it does not interrupt a live lead turn.
### 7.4 Peer lead routing
Figure 2 shows the route from incident to lead, peer, or person.
```mermaid
flowchart TD
Incident["Open incident"] --> Member{"Member incident?"}
Member -->|"yes"| Known{"Exact delegation owner known?"}
Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
Known -->|"yes"| Owner{"Owner lead healthy?"}
Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
Member -->|"no, lead incident"| PeerSet
PeerSet --> Peer{"Healthy peer exists?"}
Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
Peer -->|"no"| Sink
Sink -->|"yes"| Webhook["Send classified outbound incident"]
Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]
```
*Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.*
Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and
leads with an open unhealthy state. A reachable `WORKING` peer may be selected; its push waits for an
injectable window.
Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name,
then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A
routing generation marks a reassignment, and old pending assignments become superseded.
`fleet_list` lead rows show health, health age, assigned foreign incident count, and a bounded list
of incident id, subject, state, severity, age, and routing generation. The top-level view also shows
owner, recipient, and routing reason.
A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its
status-gated nudge names the failed lead and gives the exact
`fleet_poll(target="<recipient-terminal>")` call.
### 7.5 Lead inbox ownership
Add `LeadInboxRegistry`, driven by the same expected-lead supplier as `CallerResolver`.
It calls `replyInbox.own(leadTerminal)` at startup for configured leads, after successful discovery,
after config adds a lead, and before publication. Ownership is not an authorization side effect.
A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents
are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a
replacement terminal before moving messages from the old key. Never release a non-empty in-memory
lead key, because in-memory release clears local data.
### 7.6 Single-lead deployment
One lead and no peer is a normal mode, not an edge case.
An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead
has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the
failed lead's plan or context, and an uncertain relaunch could create two orchestrators.
With no webhook, only `fleet_list`, `/healthz`, metrics, WARN logs, and the incident journal remain.
These are passive surfaces. They are not a human notification.
## 8. Lost-boundary reconciliation
This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated
reply.
### 8.1 Why normal completion rules are not enough
Current `CompletionResolver.resolve` has two fail-open rules. It resolves when the delivery baseline
is missing. It also resolves an empty completion when the pane read fails. Those choices are valid
after a trusted `WORKING -> IDLE` boundary because the bridge knows the turn ran. They are unsafe
when health only guesses that a boundary was lost.
M4 gives each accepted send an internal `TurnToken`. It ties target, exact waiter, session turn,
delivery baseline, and task outcome together.
### 8.2 Delivery baseline
Capture the baseline immediately after prompt send and before the delivery future completes. Store:
```text
TurnToken
exact waiter identity
capture time and pane source
normalised assistant block clipped to MAX_SCRAPE_CHARS
whether a supported assistant marker was recognised
capture result: PRESENT | READ_FAILED | UNRECOGNISED
```
A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a
daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be
repaired.
Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current
extraction is Claude Code-specific and falls back to arbitrary raw text without `⏺`. That raw fallback
cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.
### 8.3 Strict gates and resolver result
Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.
The two raw snapshots must describe the same `TurnToken` and session turn. No `WORKING`,
`BLOCKED`, `UNKNOWN`, missing-agent, reply, failure, or new-delivery observation may occur between
them.
```mermaid
flowchart TD
Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
Read -->|"no"| Refuse
Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
Output -->|"no"| Refuse
Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
Resolve -->|"lost race"| Resnapshot
```
*Figure 3. Repair needs stronger evidence than a normal observed turn boundary.*
Refactor the current resolver into one guard core with two policies:
```text
resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)
```
The health monitor calls only:
```text
CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)
```
It returns `REPAIRED`, `ALREADY_RESOLVED`, `REFUSED_NO_CAPTURE`, `REFUSED_NO_BASELINE`,
`REFUSED_UNREADABLE`, `REFUSED_UNCHANGED`, `REFUSED_AMBIGUOUS_OUTPUT`, `STALE_TURN`, or
`RACE_LOST`.
Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move that turn from `BUSY` to `DONE`. A
per-target reconciliation gate stops a queued second send from being accepted between waiter
resolution and the FSM transition.
### 8.4 Lead-visible marker and refusal
A repaired result uses distinct `RECONCILED_COMPLETION` values in `Rendezvous`, `MessageService`,
task poll source, and metrics. The lead sees:
```text
[repaired completion - fleetd detected a lost turn boundary. The member did not call
fleet_reply; pane-derived text follows and may be partial]
```
Clipped text also keeps the existing clipped-tail marker.
A refused repair leaves `TURN_BOUNDARY_LOST` open and the ticket pending. The report states that no
reply was reconstructed and no task was replayed. `UNCHANGED`, `UNREADABLE`, and
`AMBIGUOUS_OUTPUT` get at most one delayed retry for the same token. Missing capture or baseline gets
no retry. After two refused scrapes, automatic repair stops for that token.
### 8.5 Target-wide teardown invariant
CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one
idempotent operation and checks this independent invariant after teardown:
- no injector entry exists for the target;
- no accepted turn or completion record exists;
- no rendezvous waiter or ask exists;
- every async task is terminal or was already terminal;
- no thread waiting for the target send lock can later accept it;
- new sends fail immediately;
- each old task has one terminal outcome and one metric count.
A violation becomes `DELEGATION_ORPHANED`. The monitor may call the same idempotent target-wide
failure operation once. It never recreates the task.
## 9. Capacity and utilisation
Capacity is a view, not a health state.
`fleet_list` adds one block per profile:
```text
profile, maxLoad, live, free, reclaimable
```
For an unlimited profile, `maxLoad` and `free` are null. `free` is
`max(0, maxLoad - live)` for a capped profile.
The view must use the exact live-count function used by placement. A second calculation could show a
free slot that placement then refuses. Member rows add `idleForSeconds` only when state is `READY` or
`DONE`, no accepted turn exists, and the inbox is empty. `reclaimable` means only that the member
holds capacity without open bridge work.
The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free,
and reclaimable counts, plus at most three long-idle members. Capacity does not make
`FleetState.hasPending()` true. A changed capacity fingerprint may re-arm one capped heartbeat
sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the
cap forever. Reply-push stand-down remains first.
The bridge must never:
- spawn a member because a slot is free;
- generate a task or acceptance criteria;
- move queued work to another member or profile;
- treat a free slot or idle member as an incident;
- stop an idle member only to improve utilisation.
The bridge knows capacity facts but has no work list. Only the lead has the plan, task context,
side-effect history, and acceptance criteria.
Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal
session edge, not every fleet tick.
This capacity design adds no automatic stop. The accepted `NEVER_READY` cleanup can still stop a
very slow startup after the existing grace, which is a known risk. Free capacity and long idle time
never trigger that path.
## 10. Human escalation and notification
### 10.1 Escalation rule
Notify a person only when no healthy lead can act:
- `CONTROL_LINK_DOWN` survives grace;
- a lead is unhealthy and no healthy peer can receive the incident;
- a member incident has no known owning lead;
- the only owning lead becomes unreachable, unresponsive, or stalled;
- incident publication or routing itself fails.
Do not page a person for a member fault while a healthy owning lead exists. An uncollected member
incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.
### 10.2 Detection and notification switches
`health.enabled` controls detection and bridge-local reporting. It does not require a webhook.
`health.notifications.mode` is `disabled` or `webhook`. Disabled is valid and is the default.
Webhook mode requires a resolved environment variable. Turning notification off stops outbound
attempts but keeps incidents. Turning it back on resumes still-open human incidents.
Without a sink, `fleet_list.healthCoverage` states that human escalation is unavailable. `/healthz`
keeps its existing HTTP liveness result and adds a nested `fleetHealth.status=partial` component.
Metrics and one startup or reload WARN expose the same limit.
### 10.3 Incident and delivery deduplication
One open incident uses this key:
```text
(scope, subjectStableId, state, causeFingerprint)
```
The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant
names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence
after resolution gets a new generation and incident id.
Each outbound event uses:
```text
Idempotency-Key = hash(incidentId, eventType, eventRevision)
```
Event types are `open`, `severity_changed`, `reminder`, and `resolved`. Transport retries keep the
same key.
An atomic owner-only journal beside the active config stores open incidents, routing, delivered
revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does
not stop detection, but notification coverage becomes degraded.
### 10.4 Retry, reminder, and resolve
Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx
with full-jitter exponential backoff:
```text
base: 5 seconds
factor: 3
maximum delay: 15 minutes
one outstanding attempt per event
```
Respect `Retry-After` up to 15 minutes. Other HTTP 4xx responses are permanent for that event until
config changes or a person requests replay.
Transport retry is not an incident reminder. `humanRepeatSeconds` creates a new reminder revision
for an unresolved critical incident after the last successful human event. Disabled mode does not
build an unbounded reminder queue.
Send `resolved` only if at least one human event for that incident was delivered. If an incident
resolves before its first successful delivery, cancel the pending open event and record local
resolution.
### 10.5 Outbound payload boundary
An outbound payload may contain incident id and event type, severity, state, scope, stable bridge
ids, role or profile, times, duration, structured evidence type and counts, recovery attempted,
routing reason, coverage, and safe tool calls.
It must never contain:
- raw pane text, pane excerpts, or pane digests;
- task briefs, prompts, or member reply content;
- source files, diffs, or worktree file content;
- worktree paths;
- environment values, tokens, credentials, headers, or webhook URL;
- raw exception messages or stack traces;
- arbitrary model output.
The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.
### 10.6 Metrics
M4 adds bounded-label series:
```text
fleet_health_incidents{scope,state,severity}
fleet_health_incidents_total{event}
fleet_health_notifications_total{event,outcome}
fleet_health_notification_queue_depth
fleet_health_notification_last_success_seconds
fleet_health_notification_capability{mode,status}
fleet_lead_health{lead,state}
fleet_lead_assigned_incidents{lead}
```
Metric labels never include terminal ids, incident ids, URLs, or error text.
## 11. Configuration
The optional `health:` block is absent or disabled by default. The dormant monitor scheduler does no
herdr or pane work while disabled. Every listed key is hot because the monitor reads `ConfigRef` on
each tick or notification.
| Key | Class | Default and hard bound | Purpose |
|---|---|---|---|
| `health.enabled` | Hot | `false` | Enable detection and bridge-local reporting. |
| `health.snapshotIntervalSeconds` | Hot | default 30, minimum 15 | Whole-fleet comparison cadence. |
| `health.workingSuspectAfterSeconds` | Hot | default 600, minimum 300 | Age before working-pane probes. |
| `health.paneProbeIntervalSeconds` | Hot | default 60, minimum 60 | Per-target pane cooldown. |
| `health.leadUnresponsiveAfterSeconds` | Hot | default 300, minimum 120 | Delay after exhausted actionable nudges before lead fault. |
| `health.humanRepeatSeconds` | Hot | default 3600, minimum 900 | Minimum repeat period for one open human incident. |
| `health.capacityLongIdleAfterSeconds` | Hot | default 900, minimum 300 | Long-idle threshold for capacity summaries. |
| `health.includePaneExcerpt` | Hot | `false` | Allow a clipped excerpt in local lead reports only. Human payloads still exclude it. |
| `health.notifications.mode` | Hot | `disabled` | Select `disabled` or `webhook`. |
| `health.notifications.webhookUrlEnv` | Hot | required in webhook mode | Name of the environment variable that holds the sink URL. |
| `health.notifications.requestTimeoutMs` | Hot | default 10000, range 1000-30000 | Whole webhook request limit. |
Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control
link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.
## 12. Delivery units and acceptance
### Unit 1 - Evidence model and fleet snapshot
Scope: health state model, fleet join, clocks, evidence retention, and pane budget.
Acceptance criteria:
1. One `agent.list` call covers one enabled fleet tick.
2. Pure decision tests cover every state and every evidence limit in Section 4.
3. `BUSY` plus stable raw `DONE` opens `TURN_BOUNDARY_LOST` after two snapshots.
4. Healthy fleet snapshots perform zero pane reads.
5. Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
6. Logs are outputs only; no log parsing exists.
7. Fleet snapshots expose the same profile live-count calculation that placement uses.
8. Capacity rows report cap, live, free, and reclaimable values without opening incidents.
### Unit 2 - Lost boundary and task reconciliation
Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved
worktree discovery.
Acceptance criteria:
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, and delivery
baseline.
**Corrected during implementation (2026-08-15).** This criterion first also required the session
turn number and the task outcome. That is not implementable at this layer, and the implementer
refused it three times rather than fabricate a value — correctly. The reason is an ordering fact
that is invisible from any single class: `MessageService` owns acceptance and holds the waiter and
the async `Task`, but it learns nothing about delivery, because the delivery event goes to
`CompletionResolver` through `TurnListener.onDelivered`. And `CompletionResolver.onDelivered` runs
*before* `SessionManager.onDelivered`, so the session turn number does not exist yet at the only
point where the token could capture it.
Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever
send happens to be waiting" ambiguity the token exists to remove — the same weak claim
`Rendezvous.currentWaiter` warns about. Injecting a turn counter into `MessageService` adds a
required cross-layer dependency to populate a field that nothing in this slice reads, which is
speculative coupling across a boundary already shown to be fragile.
So the token identifies the **accepted send**, and `SessionManager` keeps verifying its own
delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving
that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not
here. The token record carries a comment saying the field is deliberately absent.
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
observation, exact open waiter, successful baseline, and new recognised assistant output.
3. Missing, failed, late, or post-restart baseline never authorises repair.
4. Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback
without a recognised marker refuses repair.
5. Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and
exact-turn guards are not duplicated.
6. `reconcileLostBoundary` returns every typed result named in Section 8.3.
7. Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move the same turn to `DONE`.
8. A per-target reconciliation gate blocks a queued second send during repair and FSM update.
9. Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and
metric. Clipping keeps its extra marker.
10. Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or
baseline gets none.
11. Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.
12. Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure
operation.
13. The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.
14. A violated teardown invariant creates `DELEGATION_ORPHANED` and retries only the idempotent
failure operation.
15. `SPAWN_ROLLBACK` and normal `COMPLETED` remove worktrees. Abnormal and shutdown causes preserve
them.
16. Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.
17. Atomic preserved-worktree manifests reload after restart and appear in lead-only
`fleet_list.preservedWorktrees`.
18. Manifest failure preserves the worktree and opens an operator-visible health failure.
19. Provision records the base commit. Terminal, long-idle worktrees report
`WORK_PRODUCT_AT_RISK` only under the evidence in Section 4.2 and never auto-delete work.
20. No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is
available.
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
restart without capture, concurrent send and release, and preserved discovery after restart.
#### Unit 2 - what has landed so far
Checked against `main` at `e09cac6` on 2026-08-15. Unit 2 was written as one block, but parts of it
have since been built by separate CB tickets. Read this before planning the rest, or that work gets
done twice.
The check was a symbol survey of `fleetd/src/main/java` plus the merge history. It tells you whether
the machinery exists at all. It is **not** a line-by-line audit of whether each criterion is fully
met, and I did not run one.
| Criterion | Marker searched for | Found in main source | Reading |
|---|---|---|---|
| 1 | `TurnToken` | 8 files | **Done** — unit 2a, merged as `fec284e`. Criterion 1 was corrected first; see the note under it. |
| 2-5, 9 | `REPAIRED` | 0 files | Not started. The whole guarded-repair path is absent. |
| 6, 7, 10 | `reconcileLostBoundary` | 0 files | Not started. |
| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
| 15 | `SPAWN_ROLLBACK` | 0 files | **Contradicted — see below.** |
| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. |
| 17, 18 | `preservedWorktrees` | 0 files | Not started. No manifest, and no lead-only `fleet_list` field. |
| 19 | `WORK_PRODUCT_AT_RISK` | 0 files | Not started. |
**Criterion 15 no longer matches the code, and the code is right.** It says "normal `COMPLETED`
remove worktrees". Since CB-576 (`500bfa2`) that is false on purpose: a `COMPLETED` release now
preserves the worktree when it still holds uncommitted work, because deleting it destroys work
nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the
dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell"
must not be treated as "it is clean".
So criterion 15 should be rewritten as: `SPAWN_ROLLBACK` and a `COMPLETED` release with a **clean**
worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all
preserve it. `SPAWN_ROLLBACK` itself does not exist yet.
### Unit 3 - Typed inbox and member routing
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
`fleet_list`.
Acceptance criteria:
1. AMQP selects legacy or typed decoding only from `content_type`; it never sniffs the body.
2. Persistent `text/plain` from the old build becomes `kind=reply` with exact UTF-8 content,
including content beginning with `{`.
3. New entries use the vendor media type, `schemaVersion: 1`, UTF-8, persistent delivery, and AMQP
message ids.
4. Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
5. Unknown versions are not decoded or acked. They remain on the original queue and create one
deduplicated failure.
6. Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
7. Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the
original unacked.
8. Decode failures create redacted WARN, metric, `fleet_list` summary, and routed incident without
raw content.
9. Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
10. Lead keys require explicit ownership. Publication never claims a queue.
11. Unit codec tests cover legacy `{`, Unicode, malformed UTF-8, typed round trip, additive fields,
malformed JSON, missing fields, identity mismatch, media type, version, and dedup.
12. A live broker contract writes old wire data and reads it with the new adapter after reconnect.
13. Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery,
later progress past poison, property persistence, lead ownership, and ack removal.
14. Safe downgrade is documented as unsupported.
15. RabbitMQ contract tests pass with `mvn test -Pcontract`. The same cases run once on production
LavinMQ, or the release states that LavinMQ was not checked.
16. Member incidents route to the exact delegating lead and never resolve a task rendezvous.
17. `fleet_list` shows compact member health and capacity without pane content. Member
`idleForSeconds` is present only when no accepted turn or inbox item exists.
### Unit 4 - Lead health and peer routing
Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox
lifecycle.
Acceptance criteria:
1. Lead identity uses the `CallerResolver` supplier. Liveness uses successful current agent data.
2. Two successful-list absences with healthy ping become `LEAD_UNREACHABLE`; global link failure does
not.
3. Raw `WORKING`, raw `UNKNOWN`, first-seen time, failures, last success, and error class persist
across ticks.
4. Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
5. Dynamic lead identity survives a two-successful-snapshot retirement grace.
6. Member incidents first use exact delegation ownership with no singular-primary fallback.
7. Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
8. A selected working peer is not interrupted. Its push waits for an injectable window.
9. Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending
assignment.
10. `fleet_list` shows bounded foreign assignments, recipient, reason, and generation without pane
content.
11. `LeadInboxRegistry` owns configured and discovered lead keys before publication.
12. Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
13. Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
14. Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer,
several peers, reassignment, and no peer.
15. Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement.
LavinMQ is checked or named as unchecked.
16. A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive
evidence remains and every coverage surface says so.
### Unit 5 - Human sink, hot config, metrics, and operator coverage
Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and
operator documentation.
Acceptance criteria:
1. `health.enabled` works without a human sink.
2. Notification mode is hot, defaults to disabled, and supports disabled or webhook.
3. Webhook mode requires a resolved environment value. Bad notification config does not disable an
already valid detector.
4. Mode changes keep open incidents. Re-enable resumes eligible incidents.
5. `fleet_list`, `/healthz`, metrics, and one WARN show partial coverage without a sink. HTTP
liveness behavior stays unchanged.
6. One-lead, no-sink coverage states that lead failure has no active notification or recovery.
7. Incident and outbound dedupe use the stable keys in Section 10.3.
8. The owner-only local journal survives restart and contains no pane or task content.
9. Journal failure keeps detection running but marks notification coverage degraded.
10. Retry tests cover network failure, timeout, 408, 429, `Retry-After`, 5xx, permanent 4xx, jitter,
delay cap, config re-arm, and one outstanding attempt.
11. Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
12. Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale
open delivery.
13. Metrics use bounded labels and exclude ids, URLs, and error text.
14. Payload tests reject every content type forbidden in Section 10.5.
15. Webhook response bodies are ignored and cannot direct recovery.
16. Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup,
reassignment, disable/re-enable, and sink failure while local health continues.
17. `fleetd.example.yaml` documents all hot keys and compiled floors.
18. The operator Features wiki is updated separately. The portable `CLAUDE.md` block is checked and
changed only if shipped tool or inbox semantics make it untrue.
19. `mvn clean install` passes.
## 13. Not checked and release gates
These limits are part of the design, not optional follow-up notes.
- **OpenCode pane status and assistant markers were not checked.** OpenCode lost-boundary repair is
disabled until live fixtures exist.
- **Permission-prompt status was not checked** for Claude Code or OpenCode. `BLOCKED` remains
ambiguous and has no automatic action.
- **`recent_unwrapped` stability was not checked** across all supported agent kinds. If normalisation
is not stable, `STALL_SUSPECTED` must say its evidence is weaker.
- **The real `BUSY + DONE` trace was not replayed against live herdr.** The design uses the observed
production trace and current poller behavior.
- **CB-568 was not present when Unit 2 was designed.** Unit 2 must inspect the landed API and keep its
independent teardown invariant.
- **Production LavinMQ was not checked.** Existing durable-inbox contracts use RabbitMQ. Migration,
quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be
recorded as unchecked.
- **Live multi-lead routing was not checked.** Peer choice and reassignment are design rules backed by
fake-clock and adapter tests until a live exercise runs.
- **A live sole-lead failure with a webhook was not checked.** The no-peer path is a design result,
not a tested recovery.
- **No n8n, Slack, PagerDuty, or other receiver was checked.** The webhook remains generic and
outbound-only.
- **Deployment supervisor behavior for nested `/healthz` fields was not checked.** HTTP liveness
status stays unchanged to reduce this risk.
- **Incident-journal crash behavior was not checked** because the journal does not exist yet. Unit 5
must test atomic replacement and restart recovery.
- **Worktree merge state cannot be checked reliably** without forge or explicit collection evidence.
`WORK_PRODUCT_AT_RISK` stays a warning.
## 14. Locked exclusions
M4 does not expose `agent.read` as a bridge tool. It does not add a workflow engine, inbound n8n
authority, automatic task assignment, task replay, automatic lead replacement, or automatic member
spawn for free capacity.
The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the
place where judgement and work planning happen.
+129 -315
View File
@@ -1,296 +1,162 @@
# MCP Contract — `bridged`'s unified gateway
# MCP flows and error model — `fleetd`
> **Status:** 🟡 Design (2026-07-14). Greenfield — no MCP code exists yet; the pom carries
> only Javalin/Jackson. This page defines the tool surface that CB-104 and its followers
> implement. It supersedes nothing; it fills the "MCP server face" left open by the
> [Architecture](1-Architecture) page.
`bridged` is the **sole communication gateway** for every Claude session in the bridge. Both
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
mount the *same* MCP server with a single `claude mcp add` line, and talk only through its
tools. No Claude session ever addresses a broker, a peer, or the network directly.
This document defines every MCP tool that face must expose, who may call it, its blocking
semantics, and how it maps onto the code already in the tree.
> **What this page is.** The **flows**: how a delegation, a clarification, a detached task and a
> silent member each travel through `fleetd`. These shapes are what shipped, and they are hard to
> read off the code because they span the MCP face, the rendezvous registry, the `Injector` and
> herdr.
>
> **What this page is NOT: a tool reference.** It deliberately holds no tool catalogue, no
> parameter tables and no REST paths. **The live MCP schema is the authority** — each tool's own
> description and parameters, as mounted — with the intent→tool table in `CLAUDE.md` as the short
> form.
>
> That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to
> carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped;
> the page did not follow. By August it named two tools that do not exist, omitted five that do,
> had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve,
> and — worst — still described an identity model (*"any connection that does not map to a known
> worker is treated as a primary"*) that was a real privilege bug, fixed since by the ancestry
> walk in fleetd #161. Every one of those errors is the same error: **a second, hand-maintained
> copy of something the code already states**. So the second copy is gone rather than corrected.
> Only the flows remain, because a flow is a shape rather than a name, and shapes are what this
> page was ever good for.
>
> The names that do appear below are checked by `McpContractDocTest`, which fails if this page
> names a `fleet_*` tool the server does not register. That test is the whole reason it is safe to
> write a tool name here at all.
---
## 1. Design constraints (non-negotiable)
## 1. Rendezvous flows
These come from the project's core invariants and bound every decision below.
### 1.1 Delegation — happy path
1. **One server, both roles.** The primary and all workers mount an identical server. The
catalog must serve both, and `bridged` must decide *who is calling* from the connection —
never from a caller-supplied argument that could be spoofed.
2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards
`ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription.
Enforced today by [`SubscriptionGuard`](1-Architecture).
3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a
*single* MCP call that `bridged` holds open — never a cross-turn poll loop that would burn
subscription quota.
4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing
[`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most
one message per turn.
5. **`bridged` owns policy; herdr owns PTYs.** MCP tools express *intent*; `bridged`
translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls.
---
## 2. Topology
Both faces live in the one daemon. The **north face** is MCP (this document); the **south
face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients and dashboards.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["bridged — standalone daemon"]
MCP["MCP server (north face)<br/>bridge_send · bridge_reply<br/>bridge_ask · bridge_status · lifecycle"]
RDV["rendezvous registry<br/>(blocking-call waiters)"]
INJ["Injector + StatusPoller<br/>(status-gated writer)"]
SOCK["herdr socket client (south face)"]
MCP --> RDV
RDV --> INJ
INJ --> SOCK
MCP --> SOCK
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
OPUS -->|"bridge_send (blocks)"| MCP
W -.->|"bridge_reply / bridge_ask"| MCP
SOCK -->|"agent.start · agent.send<br/>agent.get · pane.close"| HERDR
HERDR -->|"drives PTY"| W
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
class OPUS,W ext
class MCP,RDV,INJ,SOCK core
```
---
## 3. Identity & addressing
Because the same server is mounted by everyone, `bridged` resolves the caller's role on every
request — this is the linchpin of the whole contract and has no code yet.
- **Workers are known.** `bridged` spawns every worker
([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on
the returned [`Agent`]. When a call arrives on a connection that maps to a known worker,
the caller is *that* worker — so **workers never pass a target**; routing is implicit.
- **The primary is "not a worker".** Any connection that does not map to a known worker is
treated as a primary. It addresses workers **explicitly** by `target` — a session UUID,
a `terminal_id`, or a friendly `profile` name.
- **Turn correlation.** A blocking `bridge_send` registers a *waiter* keyed by worker
identity. A worker's later `bridge_reply` / `bridge_ask` on the same identity resolves that
waiter. A `turn_id` is minted per exchange so a clarification round-trip
(§6.2) rejoins the right turn.
---
## 4. Transport
`bridged` is a long-lived daemon serving **multiple** concurrent clients (one primary + N
workers), so a per-client stdio child is the wrong shape. The recommended transport is
**streamable-HTTP / SSE** on the same bind as the REST face:
```bash
# identical on primary and every worker
claude mcp add --transport http bridged http://127.0.0.1:8080/mcp
```
This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions).
---
## 5. Tool catalog
| Tool | Caller | Blocks? | Backing (exists today?) |
|---|---|---|---|
| [`bridge_send`](#bridge_send) | primary | yes (default) | `Injector.enqueue` ✅ · rendezvous registry ❌ (CB-104) |
| [`bridge_reply`](#bridge_reply) | worker | no | rendezvous ❌ · pane injection via `Injector` ✅ |
| [`bridge_ask`](#bridge_ask) | worker | yes | reverse rendezvous ❌ |
| [`bridge_status`](#bridge_status) | either | no | `AgentControl.status` ✅ · `Injector.activeTargets` ✅ |
| [`bridge_spawn`](#lifecycle) | primary | no | `WorkerService.spawn` ✅ (`POST /workers`) |
| [`bridge_list`](#lifecycle) | either | no | `WorkerService.list` ✅ (`/agents`) |
| [`bridge_stop`](#lifecycle) | primary | no | `WorkerService.stop` ✅ (`DELETE /workers/{paneId}`) |
| [`bridge_read`](#bridge_read) | primary | no | `AgentControl.read` ✅ |
| [`bridge_cancel`](#bridge_cancel) | primary | no | — ❌ (future) |
### Core: delegation & rendezvous
#### `bridge_send`
*(primary → worker — the headline tool, CB-104)*
- **Params:** `message` (required); `target` (optional — defaults to the sole worker / default
profile); `timeout_seconds` (default 600); `block` (default `true`); `auto_spawn`
(default `true`); `turn_id` (optional — supplied when answering a worker's `bridge_ask`).
- **Blocking (`block:true`):** enqueue `message` via the `Injector`, then hold the call open
until exactly one of:
- worker calls `bridge_reply` → `{ outcome:"reply", text }`
- worker calls `bridge_ask` → `{ outcome:"question", text, turn_id }`
- worker's `agent_status` reaches done/idle with no reply → `{ outcome:"turn_done", text:<terminal tail> }`
- deadline elapses → `{ outcome:"timeout" }`
- worker gone → error `worker_gone`
- **Detached (`block:false`):** enqueue and return `{ outcome:"dispatched", dispatch_id }`
immediately. The eventual reply is injected into the primary's idle pane (§6.3), or drained
via `bridge_status` on a split-host primary.
#### `bridge_reply`
*(worker → primary)*
- **Params:** `text` (required); `final` (default `true`).
- **Behavior:** resolve the primary waiter registered against this worker with `text`. If no
waiter exists (detached delegation), `bridged` **injects the primary's idle pane** instead.
Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit.
#### `bridge_ask`
*(worker → primary — the reverse rendezvous)*
- **Params:** `question` (required); `timeout_seconds`.
- **Behavior:** blocks the *worker's* call. Surfaces the question to the primary (resolving its
open `bridge_send` with `outcome:"question"`, or injecting its pane). When the primary
answers — a `bridge_send` carrying the matching `turn_id` — that unblocks this call and
returns `{ answer }` to the worker, which continues **in the same turn**.
### Worker lifecycle
<a id="lifecycle"></a>
Thin adapters over [`WorkerService`](1-Architecture) — parity with the existing REST routes.
- **`bridge_spawn`** — `{ profile? }` → worker view (`sessionId`, `terminalId`, `paneId`,
`status`). Guard-checked; a boundary breach returns error `subscription_boundary` (the
REST `403`).
- **`bridge_list`** — no params → all workers + `agent_status`. Read-only, either role.
- **`bridge_stop`** — `{ target }` → tears down the pane and its dedicated tab. Idempotent.
### Observability
#### `bridge_status`
*(either role — the README's 4th named tool)*
- **Params:** `target?`.
- **Behavior:** per-worker `agent_status`, queue depth (`Injector.activeTargets`), whether a
rendezvous is open, and ids. For the *calling* session it also reports/drains **pending
messages addressed to me** — the path a split-host primary's `Stop`-hook uses to wake and
collect replies without being injectable. Read-only, non-blocking.
#### `bridge_read`
*(primary)*
- **Params:** `target`; `source` ∈ `visible | recent | recent_unwrapped | detection`.
- **Behavior:** returns the worker's terminal text so the primary can peek at a *detached*
worker's progress. Adapter over `AgentControl.read`.
### Control (future)
#### `bridge_cancel`
*(primary)*
- **Params:** `target`. Interrupt the worker's current turn / abandon the rendezvous. No
backing code yet.
---
## 6. Rendezvous flows
### 6.1 Delegation — happy path
One blocking call, zero polls.
One blocking call, zero polls. The lead's call is held open by `fleetd` until the member answers.
```mermaid
sequenceDiagram
participant P as Primary (Opus)
participant B as bridged (MCP + Injector)
participant P as "Lead (primary)"
participant B as "fleetd (MCP + Injector)"
participant H as herdr
participant W as Worker (Claude)
participant W as "Member"
P->>B: bridge_send("do X", target=w) — blocks
B->>B: register waiter(w)
B->>H: agent.send(w, "do X") (idle window)
H-->>W: prompt injected
W->>W: works the turn
W->>B: bridge_reply("result")
B->>B: resolve waiter(w)
B-->>P: { outcome:"reply", text:"result" }
P->>B: "fleet_send{sessionId, content} — blocks"
B->>B: "register waiter(sessionId)"
B->>H: "agent.send — only in an injectable window"
H-->>W: "prompt injected"
W->>W: "works the turn"
W->>B: "fleet_reply{content}"
B->>B: "resolve waiter"
B-->>P: "{ outcome: reply }"
```
### 6.2 Clarification — reverse rendezvous (`bridge_ask`)
**The cap that matters:** a blocking `fleet_send` is bounded by the *caller's own* MCP client
timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow
in §1.3, or the lead's call returns while the member is still working.
The worker pauses mid-turn to ask; the primary answers; the worker resumes in the same turn.
### 1.2 Clarification — reverse rendezvous
The member pauses mid-turn to ask, the lead answers, and the member resumes **the same turn** with
its context intact.
```mermaid
sequenceDiagram
participant P as Primary
participant B as bridged
participant W as Worker
participant P as "Lead"
participant B as fleetd
participant W as "Member"
P->>B: bridge_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>B: bridge_ask("which config?") — worker blocks
B-->>P: { outcome:"question", text:"which config?", turn_id }
P->>B: bridge_send("config.yaml", target=w, turn_id) — blocks again
B-->>W: resolve bridge_ask → { answer:"config.yaml" }
W->>W: resumes same turn
W->>B: bridge_reply("done")
B-->>P: { outcome:"reply", text:"done" }
P->>B: "fleet_send{sessionId, content} — blocks"
B-->>W: "content injected"
W->>B: "fleet_ask{question} — member blocks"
B-->>P: "{ outcome: question, turnId }"
P->>B: "fleet_send{turnId, content} — answers THIS turn"
B-->>W: "fleet_ask returns the answer"
W->>W: "resumes the same turn"
W->>B: "fleet_reply{content}"
B-->>P: "{ outcome: reply }"
```
### 6.3 Detached delegation — pane injection
**Answer with `turnId`, never `sessionId`.** A `sessionId` send starts a new turn; it does not
resolve the waiting `fleet_ask`.
The primary does not block; the reply arrives later in its idle pane.
**The window is about 55 seconds and no nudge extends it.** So never brief a member to "ask me":
decide before delegating, or give the member an explicit default to fall back on.
### 1.3 Detached delegation — the lead does not block
The lead gets a ticket immediately and collects the answer later. This is the flow for any real
task, because of the ~60s cap in §1.1.
```mermaid
sequenceDiagram
participant P as Primary
participant B as bridged
participant W as Worker
participant P as "Lead"
participant B as fleetd
participant W as "Member"
P->>B: bridge_send("do X", target=w, block=false)
B-->>P: { outcome:"dispatched", dispatch_id }
P->>P: continues its own work
W->>B: bridge_reply("result")
Note over B: no waiter → detached path
B->>B: Injector.enqueue(primary_pane, "result")
B-->>P: injected into idle pane (status-gated)
P->>B: "fleet_send{sessionId, content, wait:false}"
B-->>P: "accepted — ticket"
P->>P: "continues its own work"
W->>B: "fleet_reply{content}"
Note over B: "no waiter is blocked — the reply is held"
B->>B: "nudge the lead's own pane (status-gated)"
P->>B: "fleet_poll{ticket}"
B-->>P: "the member's report"
P->>B: "fleet_ack{target, msgId}"
```
### 6.4 Uncooperative worker — turn-done fallback
A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The
nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.
A worker that never calls `bridge_reply` still returns a result: `bridged` reads its terminal
tail when the turn completes.
### 1.4 The member never replies — turn-done fallback
A member that ends its turn without `fleet_reply` still produces something: `fleetd` reads its
pane tail. This is a **fallback, not a channel** — it is lossy in three separate ways, and every
one of them has produced a wrong answer in practice.
```mermaid
sequenceDiagram
participant P as Primary
participant B as bridged
participant W as Worker
participant P as "Lead"
participant B as fleetd
participant W as "Member"
P->>B: bridge_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>W: works, never calls bridge_reply
B->>B: StatusPoller sees agent_status → idle/done
B->>B: AgentControl.read(w, "recent")
B-->>P: { outcome:"turn_done", text:<terminal tail> }
P->>B: "fleet_send — blocks or detaches"
B-->>W: "content injected"
W->>W: "works, never calls fleet_reply"
B->>B: "StatusPoller sees the turn end"
B->>B: "read the pane tail"
B->>B: "classify: exhausted? echoed brief? real report?"
B-->>P: "{ outcome: turn_done } or a named failure"
```
The three ways it goes wrong, and what each looks like now:
| What happened | What the lead used to get | What it gets today |
|---|---|---|
| The report is longer than the scrape window | The **end** silently cut off | Still clipped, but marked partial |
| The member never started — spent credential | The lead's **own brief** echoed back as a report | A named failure: backend exhausted |
| The member is simply slow | A tail of work in progress | Unchanged — read it as a hint, not a result |
The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in
it from the member. It is suppressed now, but the general rule stands — **check the member's
worktree with `git log` before believing a report you did not watch arrive.**
---
## 7. Status gating
## 2. Status gating
Delivery only happens in a safe window. This is the state machine the `Injector` already
enforces via `AgentStatus.injectable()`; MCP `bridge_send` is simply its producer.
Delivery only happens in a safe window. `fleet_send` is a producer for the `Injector`, which
already enforces this through `AgentStatus.injectable()`.
```mermaid
stateDiagram-v2
[*] --> IDLE
IDLE --> WORKING: message delivered / picks up
WORKING --> IDLE: turn done
WORKING --> BLOCKED: awaits input
BLOCKED --> WORKING: input delivered
IDLE --> UNKNOWN: detection glitch
BLOCKED --> UNKNOWN: detection glitch
UNKNOWN --> IDLE: re-detected
IDLE --> WORKING: "message delivered, picked up"
WORKING --> IDLE: "turn done"
WORKING --> BLOCKED: "awaits input"
BLOCKED --> WORKING: "input delivered"
IDLE --> UNKNOWN: "detection glitch"
BLOCKED --> UNKNOWN: "detection glitch"
UNKNOWN --> IDLE: "re-detected"
note right of IDLE
injectable — deliver head of FIFO
@@ -306,69 +172,17 @@ stateDiagram-v2
end note
```
At most one message is delivered per turn: after a send the `Injector` waits for a `WORKING`
pickup before delivering the next, with a `PICKUP_GRACE_POLLS` fallback for turns faster than
the poll interval. A herdr `events.subscribe` stream can later replace the sampling without
touching this state machine.
**At most one message per turn.** After a send, the `Injector` waits for a `WORKING` pickup before
delivering the next, with a grace-poll fallback for turns that finish faster than the poll
interval.
---
Two consequences a lead feels directly:
## 8. Error model
- **A second send to a busy member never lands.** It reports as queued and times out. The member
is fine; the message simply waits, and then restarts the member when it next goes idle.
- **A spawned member is not deliverable until it has mounted the MCP.** Until then a send waits on
that gate for about 60 seconds and then fails without ever reaching the pane.
| Condition | `bridge_send` result | Notes |
|---|---|---|
| Worker replies | `{ outcome:"reply" }` | normal |
| Worker asks | `{ outcome:"question", turn_id }` | answer with `bridge_send(turn_id)` |
| Turn ends, no reply | `{ outcome:"turn_done" }` | terminal tail as text |
| Deadline elapsed | `{ outcome:"timeout" }` | message may still be queued/delivered |
| Worker vanished | error `worker_gone` | `Injector.drop` fails the queued future |
| Guard breach on spawn | error `subscription_boundary` | REST `403` parity |
| Delivery failed at herdr | error, message dropped | poisoned message not left blocking the FIFO |
`bridge_reply` from a worker with no open waiter is **not** an error — it falls through to
detached pane injection (§6.3).
---
## 9. Mapping to existing code
The MCP face is a thin adapter layer; nearly every capability already exists behind the REST
seam. Only the **rendezvous registry** and the **caller-identity resolver** are new.
| MCP tool | Existing collaborator | New work |
|---|---|---|
| `bridge_send` | `Injector.enqueue`, `AgentControl.send` | waiter registry, timeout, outcome mux (CB-104) |
| `bridge_reply` / `bridge_ask` | `Injector` (pane injection) | reverse rendezvous, identity resolver |
| `bridge_status` | `AgentControl.status`, `Injector.activeTargets` | pending-drain projection |
| `bridge_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only |
| `bridge_read` | `AgentControl.read` | MCP adapter only |
Because the REST routes in `BridgedApp` already exercise the collaborators, MCP tools are
validated by **parity** against those routes, not by re-testing behavior.
---
## 10. Open decisions
1. **`bridge_ask` direction.** This page defines it as *worker-asks-primary* (a genuine reverse
channel, matching the "inject the primary's pane" language). The alternative — a synonym for
a blocking primary→worker send — is weaker and produces different plumbing. **Recommend
worker-asks-primary.**
2. **Detached delivery shape.** A `block:false` param on `bridge_send` (keeps the catalog
small) vs. a separate `bridge_dispatch` tool. **Recommend the param.**
3. **Auto-spawn on send.** `bridge_send` provisions a worker per profile when none exists
(simplest primary UX) vs. requiring an explicit `bridge_spawn` first. **Recommend
auto-spawn, defaulting on.**
4. **Transport & SDK.** Streamable-HTTP/SSE co-located with the REST bind (recommended) vs.
stdio. Requires choosing a Java MCP server SDK and adding it to the pom.
---
## 11. Implementation staging
- **CB-104** — blocking `bridge_send` + rendezvous registry + caller-identity resolver
(the producer that finally drives the inert `StatusPoller`).
- **CB-1xx** — `bridge_reply` / `bridge_ask` reverse rendezvous + detached pane injection.
- **CB-1xx** — lifecycle + observability adapters (`bridge_spawn/list/stop/status/read`).
- **CB-1xx** — transport wiring + `claude mcp add` docs; parity tests vs. REST.
- **Later** — `bridge_cancel`; swap `StatusPoller` for herdr `events.subscribe`.
`UNKNOWN` is deliberately neither injectable nor a pickup. A pane whose status cannot be read is
not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an
unreadable pane as a ready one.
+7 -7
View File
@@ -1,8 +1,8 @@
# v1.0.0 — One leader, one host, complete
This is the first release of **`bridged`**.
This is the first release of **`fleetd`**.
`bridged` lets one main Claude Code session (the **leader**, on your Pro/Max subscription)
`fleetd` lets one main Claude Code session (the **leader**, on your Pro/Max subscription)
run a team of **workers** — extra Claude Code sessions on a cheaper or local model, and
non-Claude agents too. The leader's own session is never touched: it stays on subscription,
with a clean environment.
@@ -19,12 +19,12 @@ of this one.
## One gateway for all messages
- **Everyone talks through the same door.** The leader and every worker connect to the same
MCP server and use only its tools: `bridge_whoami` · `bridge_profiles` · `bridge_spawn` ·
`bridge_list` · `bridge_status` · `bridge_send` · `bridge_reply` · `bridge_ask` ·
`bridge_poll` · `bridge_ack` · `bridge_stop`.
MCP server and use only its tools: `fleet_whoami` · `fleet_profiles` · `fleet_spawn` ·
`fleet_list` · `fleet_status` · `fleet_send` · `fleet_reply` · `fleet_ask` ·
`fleet_poll` · `fleet_ack` · `fleet_stop`.
- **You are who your connection says you are.** The bridge finds out who is calling from the
connection itself, never from a name the caller sends. So a worker cannot pretend to be
someone else, and `bridge_whoami` tells each agent its own role — no guessing.
someone else, and `fleet_whoami` tells each agent its own role — no guessing.
- **The subscription line cannot be crossed.** Only a spawned worker gets
`ANTHROPIC_BASE_URL`; the leader never does. Each worker profile has a list of allowed
model hosts, checked before anything starts.
@@ -59,7 +59,7 @@ had to ask for its replies. That gap is now closed on a single machine:
- **The leader gets a tap on the shoulder.** When a reply lands, the bridge nudges the
leader's own pane — only when the leader is free, and only a few times. If the leader is on
another machine, this quietly falls back to pick-up mode; the reply still waits.
- **Workers can ask questions.** With `bridge_ask`, a worker can pause mid-task, ask the
- **Workers can ask questions.** With `fleet_ask`, a worker can pause mid-task, ask the
leader something, and continue the *same* task with the answer.
## More than one kind of worker
+21 -21
View File
@@ -1,9 +1,9 @@
# Team — lead orchestrating a mixed Claude + local-LLM fleet
The message server (`bridged`) delivers **one turn into one worker**. A **team** is the
The message server (`fleetd`) delivers **one turn into one worker**. A **team** is the
layer above it: a **Claude team-lead** that fans a job out across a **mixed fleet** of
workers — some on Claude, some on the remote local LLM — and reduces their replies. Same
`bridged` delivery, same subscription boundary; this doc is only about **orchestration** —
`fleetd` delivery, same subscription boundary; this doc is only about **orchestration** —
who the workers are, how the lead picks one, and how it runs many at once.
> Delivery mechanics (blocking `POST /message`, status-gated reply envelope) live in the
@@ -12,12 +12,12 @@ who the workers are, how the lead picks one, and how it runs many at once.
## The team
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
worker; a **thin client of `bridged`**. It plans, routes, dispatches, and integrates, and
worker; a **thin client of `fleetd`**. It plans, routes, dispatches, and integrates, and
never sets `ANTHROPIC_BASE_URL`.
- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
- **Workers** — a herd of `claude` panes in herdr, each an addressable `fleetd` session
with its **own model/env**:
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
- **Local workers** (`ANTHROPIC_BASE_URL=https://llm.ltms.dev/anthropic`) — bulk, cheap, or
embarrassingly parallel subtasks.
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
@@ -28,14 +28,14 @@ MCP) — only its model differs. Scale each kind horizontally by adding panes.
```mermaid
flowchart TB
LEAD["lead — Opus<br/>(Claude Code, env CLEAN)"]
BD["bridged<br/>message server + router"]
BD["fleetd<br/>message server + router"]
HERDR["herdr<br/>panes · agent-status"]
WC1["w-claude-1<br/>Sonnet · CLEAN"]
WC2["w-claude-2<br/>Sonnet · CLEAN"]
WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
ANT["api.anthropic.com<br/>(Pro/Max)"]
OLL["ollama.ltms.dev<br/>(local model)"]
OLL["llm.ltms.dev<br/>(gateway to the local model)"]
LEAD -->|"blocking POST /message (target role)"| BD
BD -->|"Unix socket · send_text · events.subscribe"| HERDR
@@ -62,14 +62,14 @@ flowchart TB
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
selection is **policy in the lead**, not a `bridged` concern — `bridged` just delivers to
selection is **policy in the lead**, not a `fleetd` concern — `fleetd` just delivers to
the session the lead names.
## Subscription boundary in a team
Unchanged from the base architecture, and it scales with the fleet: **only local-worker
panes** launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on
the subscription. `bridged` enforces which panes may carry the off-subscription env, so
the subscription. `fleetd` enforces which panes may carry the off-subscription env, so
adding workers never widens the boundary.
## Parallel fan-out (map / reduce)
@@ -80,7 +80,7 @@ different workers at once, then results are gathered.
```mermaid
sequenceDiagram
participant L as lead (Opus)
participant B as bridged
participant B as fleetd
participant WC as w-claude-1
participant WL as w-local-1
@@ -100,26 +100,26 @@ sequenceDiagram
```
- **Map:** the lead issues N concurrent blocking `POST /message` calls (one per subtask → its
chosen worker). Each call blocks only *that* request; `bridged` holds it open until the
chosen worker). Each call blocks only *that* request; `fleetd` holds it open until the
worker's turn completes (status-gated) and returns the reply envelope.
- **Reduce:** the lead collects the N envelopes and integrates. A slow local worker never
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
- **Detached / long jobs** use the async broker path instead of a held request (Channel 2 in
the base architecture), so the lead never busy-polls across turns.
Fan-out is bounded by the herd size (pane count) and `bridged`'s concurrency policy, not by
Fan-out is bounded by the herd size (pane count) and `fleetd`'s concurrency policy, not by
the lead.
## Knowing the roster
The lead discovers its team from `bridged` (session list / roles) rather than hard-coding
The lead discovers its team from `fleetd` (session list / roles) rather than hard-coding
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
```markdown
## Your team (via bridged)
You are the team-lead. Delegate through the bridged client — never launch workers yourself.
Roster: ask bridged for current sessions/roles.
## Your team (via fleetd)
You are the team-lead. Delegate through the fleetd client — never launch workers yourself.
Roster: ask fleetd for current sessions/roles.
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
@@ -133,22 +133,22 @@ tool instead of hand-rolling the HTTP request.
## What this layer does NOT change
- **Delivery** is still `bridged` → herdr `pane.send_text` + status events (Message-Server).
- **Delivery** is still `fleetd` → herdr `pane.send_text` + status events (Message-Server).
- **Completion timing** is still the worker status event; **reply content** still rides the
worker `Stop`-hook envelope.
- **Single-host** still applies: herdr's socket is local, so the whole herd lives on the
`bridged` host. The lead may be remote — it only needs HTTP to `bridged`.
`fleetd` host. The lead may be remote — it only needs HTTP to `fleetd`.
## Open questions
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `bridged` role-router
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `fleetd` role-router
(label-based). Start with the former; promote to the latter if routing logic grows.
- **Backpressure:** per-role concurrency caps in `bridged` so a fan-out can't exhaust the
- **Backpressure:** per-role concurrency caps in `fleetd` so a fan-out can't exhaust the
local gateway.
- **Result schema:** whether reply envelopes should carry structured metadata (worker, model,
tokens) to help the lead's reduce step.
## Status
🟡 Design (2026-07-11). Orchestration layer over the selected `bridged` server; inherits
🟡 Design (2026-07-11). Orchestration layer over the selected `fleetd` server; inherits
herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see the Message-Server design.
+10 -10
View File
@@ -20,7 +20,7 @@ requirement, not a nice-to-have.**
- **Worktree provisioned by the daemon** — `SessionManager` creates a dedicated git worktree +
branch per session, **hydrates it to full config parity** (below), and tears it down on release.
- **Worker opens its own PR** — the worker commits, pushes its branch, and opens the PR/MR itself,
returning the PR URL in its `bridge_reply`.
returning the PR URL in its `fleet_reply`.
## Why worktrees (the hazard being fixed)
@@ -76,7 +76,7 @@ worktree checks out anyway. Amber is the real gap — untracked local config the
3. **Never overlay the git plumbing** — the worktree's own `.git` file/branch is what gives
isolation; that's the *one* thing that must differ from the main tree.
The overlay set lives in config (`BridgedConfig.Worker.parityOverlay` — a list of repo-relative
The overlay set lives in config (`FleetConfig.Worker.parityOverlay` — a list of repo-relative
paths, with sane defaults) so it's auditable and per-repo tunable.
> **Trust note (deliberate).** Hydrating local config means the primary's local secrets/tokens
@@ -103,7 +103,7 @@ sequenceDiagram
Note over W: implement in the isolated worktree
W->>G: git commit + git push (SSH, same user)
W->>G: open PR (branch to main)
W-->>P: bridge_reply (prUrl, branch, summary, tests)
W-->>P: fleet_reply (prUrl, branch, summary, tests)
P->>SM: release(paneId)
SM->>G: git worktree remove wt
Note over G: branch + PR persist for review/merge
@@ -115,9 +115,9 @@ earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.*
## Infra facts (verified this session)
- **Remote:** `ssh://git@git.ltms.dev:2224/lms/claude-bridge.git` (gitea). Push is over **SSH** —
- **Remote:** `ssh://git@git.ltms.dev:2224/fleet/fleetd.git` (gitea). Push is over **SSH** —
a worker running as the same user with the same keys can `git push` **with no extra credential**.
- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `bridged`). The
- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `fleetd`). The
primary's gitea MCP comes from a global/user config, so **workers do not inherit it**. A worker
gets only the `bridge` MCP mounted (via `--mcp-config` launch flag).
- **No gitea CLI** (`tea`) installed; `glab` is present but is the GitLab CLI (wrong backend).
@@ -128,7 +128,7 @@ earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.*
| Option | Mechanism | Trade-off |
|---|---|---|
| **A. gitea REST + token** | Worker `curl`s `POST /api/v1/repos/lms/claude-bridge/pulls` with a scoped token injected by the daemon into the worker env | Minimal, no new server; token lives in the off-subscription worker's env (scope it tightly) |
| **A. gitea REST + token** | Worker `curl`s `POST /api/v1/repos/fleet/fleetd/pulls` with a scoped token injected by the daemon into the worker env | Minimal, no new server; token lives in the off-subscription worker's env (scope it tightly) |
| **B. mount gitea MCP into workers** | Add the gitea MCP to the worker's `--mcp-config` alongside `bridge` | Clean tool call, but the gitea MCP's own auth/token must be provisioned per worker; more moving parts |
| **C. install `tea` CLI** | Worker runs `tea pr create` with a token | Another dependency to install + configure; same token question as A |
@@ -140,7 +140,7 @@ and the token is a single scoped secret the daemon injects like it already injec
- Off-subscription workers already *could* push (SSH, same user). The **incremental grant is
PR-create**, i.e. a gitea API token.
- Scope the token **minimally**: the `lms/claude-bridge` repo, `write:repository` (create branch +
- Scope the token **minimally**: the `fleet/fleetd` repo, `write:repository` (create branch +
PR), **not** merge/admin/org. A leaked token can open PRs, not merge them — the primary/human is
still the merge gate.
- Inject via the daemon (env var, e.g. `GITEA_TOKEN`), never written to the worker's config dir —
@@ -153,11 +153,11 @@ and the token is a single scoped secret the daemon injects like it already injec
|---|---|---|
| Worktree provision/teardown | **CB-301 ext** — `SessionManager.acquire`/`release`; `WorkerSession` gains `worktree`, `branch` | daemon shells out to `git worktree add/remove` |
| **Config-parity overlay** | **CB-301 ext** — `SessionManager.acquire`, after `git worktree add` | symlink/copy the `parityOverlay` set into the worktree so the worker is a full peer; **this is what makes worktrees viable, not a dead-end** |
| Overlay config | `BridgedConfig.Worker.parityOverlay` — repo-relative paths, sane defaults | auditable, per-repo tunable; keep explicit + minimal (trust) |
| Overlay config | `FleetConfig.Worker.parityOverlay` — repo-relative paths, sane defaults | auditable, per-repo tunable; keep explicit + minimal (trust) |
| Branch naming | `worker/<ticket-slug>-<nonce>` off `main` (or a configured base) | one branch per session |
| Commit + push + PR handoff | **CB-302** — worker-driven, guided by the skill | push = SSH; PR = option A |
| Implementer skill | `.claude/skills/implementer/SKILL.md` | worktree-aware playbook (see below); mounts automatically since workers inherit repo cwd |
| gitea token injection | `WorkerService` env + `BridgedConfig` | repo-scoped, minimal perms |
| gitea token injection | `WorkerService` env + `FleetConfig` | repo-scoped, minimal perms |
| PR review + merge | Primary (has gitea MCP + judgment) | merge on green; the human/primary gate stays |
## Implementer skill (outline)
@@ -170,7 +170,7 @@ A worker-facing playbook (sibling to the existing `reviewer` skill):
3. **Push** your branch (`git push -u origin HEAD`).
4. **Open a PR** to `main` (option A `curl`, or the decided mechanism) with a title/body describing
the change and referencing the ticket.
5. **Reply** via `bridge_reply` with the **PR URL**, branch name, files changed, and test names —
5. **Reply** via `fleet_reply` with the **PR URL**, branch name, files changed, and test names —
that reply is the whole handoff.
6. Do **not** merge; do **not** touch `.mcp.json` or `wiki/`.
+11 -11
View File
@@ -1,6 +1,6 @@
# Worker startup: working directory & the folder-trust prompt
When `bridged` spawns a worker, the worker CLI may show an **interactive startup prompt** before it
When `fleetd` spawns a worker, the worker CLI may show an **interactive startup prompt** before it
is ready to accept a task — most importantly a *"Do you trust the files in this folder?"* dialog. An
unattended worker parked on that prompt never becomes injectable: the status-gated injector waits for
`idle`/`blocked`, the task is never delivered, and (worst case) a stray Enter answers the dialog
@@ -14,11 +14,11 @@ rule that **a worker inherits the primary's directory** (never `$HOME`), and how
```mermaid
flowchart TD
A["bridge_spawn / POST /workers"] --> B{"explicit cwd?<br/>(profile cwd or spawn arg)"}
A["fleet_spawn / POST /workers"] --> B{"explicit cwd?<br/>(profile cwd or spawn arg)"}
B -->|"yes — told otherwise"| C["use that cwd"]
B -->|"no"| D{"caller PID resolvable?<br/>(MCP peer PID)"}
D -->|"yes"| E["cwd = the primary's cwd<br/>lsof -a -p PID -d cwd"]
D -->|"no (REST / off-host)"| F["cwd = bridged daemon cwd<br/>(never $HOME by assumption)"]
D -->|"no (REST / off-host)"| F["cwd = fleetd daemon cwd<br/>(never $HOME by assumption)"]
C --> G["ensureWorkspace → tab.create → agent.start {cwd}"]
E --> G
F --> G
@@ -51,10 +51,10 @@ only affect the seed shell, which the bridge closes).
| # | Source | When |
|---|--------|------|
| 1 | Explicit `cwd` — a per-profile `cwd:` in config, or a spawn argument | "told otherwise" — pin a fixed workdir |
| 2 | The **primary's cwd**, auto-detected from the `bridge_spawn` caller | normal MCP spawn from the primary |
| 3 | The `bridged` daemon's own cwd | REST spawn / off-host caller — **never `$HOME`** |
| 2 | The **primary's cwd**, auto-detected from the `fleet_spawn` caller | normal MCP spawn from the primary |
| 3 | The `fleetd` daemon's own cwd | REST spawn / off-host caller — **never `$HOME`** |
The primary's cwd (source 2) is discoverable with no new plumbing: `bridged` already resolves the MCP
The primary's cwd (source 2) is discoverable with no new plumbing: `fleetd` already resolves the MCP
caller's loopback **peer PID** for connection identity (`ConnectionIdentity` → `LsofPeerPidLookup`);
the same PID yields its cwd via `lsof -a -p <pid> -d cwd -Fn` (the `n…` line). The primary maps to no
worker pane (it is not a worker), but its PID and cwd are still readable.
@@ -62,10 +62,10 @@ worker pane (it is not a worker), but its PID and cwd are still readable.
```mermaid
sequenceDiagram
participant P as "Primary (main)"
participant B as "bridged"
participant B as "fleetd"
participant O as "OS (lsof)"
participant H as "herdr"
P->>B: "bridge_spawn {profile} (no cwd)"
P->>B: "fleet_spawn {profile} (no cwd)"
B->>O: "peer PID for this connection's port"
O-->>B: "pid"
B->>O: "cwd of pid (lsof -d cwd)"
@@ -77,9 +77,9 @@ sequenceDiagram
*Figure 2 — a no-cwd spawn inherits the primary's directory from the caller's PID.*
> **Status:** implemented (CB-112). `bridged` threads the resolved `cwd` onto **`agent.start {cwd}`**
> **Status:** implemented (CB-112). `fleetd` threads the resolved `cwd` onto **`agent.start {cwd}`**
> (verified: the worker process is rooted there), keeping the single shared worker space. On an MCP
> `bridge_spawn` the primary's cwd is auto-detected from the caller's PID; over REST (no MCP caller)
> `fleet_spawn` the primary's cwd is auto-detected from the caller's PID; over REST (no MCP caller)
> it is the explicit `cwd` param else the daemon's cwd. Both placements (`tab` and legacy `pane`)
> carry it, since it rides `agent.start`.
@@ -133,5 +133,5 @@ unattended.
## See also
- `docs/MCP-Contract.md` — the tool surface (`bridge_spawn`, `bridge_profiles`, …).
- `docs/MCP-Contract.md` — the tool surface (`fleet_spawn`, `fleet_profiles`, …).
- `wiki/2-Message-Server.md` — the herdr `agent.*` / `workspace.*` schema (`workspace.create {cwd}`).
+208
View File
@@ -0,0 +1,208 @@
# Wiki audit for #168
**Source checked:** `.wiki-snapshot/` at `68e32c6` (2026-08-31). I did not use
`wiki/`. Code references below are from the current `fleetd` source tree. A quoted
line is a concrete claim that needs correction, unless the table says `KEEP`.
| Page | Verdict | One-line reason |
|---|---|---|
| `Home.md` | REVISE | Good overview, but it still names the retired product. |
| `_Sidebar.md` | REVISE | The heading still says `claude-bridge`. |
| `1-Architecture.md` | REBUILD | Its component contract mixes current names with removed tools, routes, and planned backends. |
| `2-Message-Server.md` | REBUILD | The claimed MCP schema, mount command, REST/SSE surface, and fallback paths are pre-build design. |
| `3-Approaches.md` | REVISE | Useful research history, but it presents unbuilt AgentAPI as a selectable fallback. |
| `4-Setup.md` | RETIRE | It is an intentional stub that only redirects to chapter 13. |
| `5-Operations.md` | RETIRE | It is an intentional stub that only redirects to chapter 13. |
| `6-Team.md` | REBUILD | It teaches role-addressed sends and a Claude-only team model that the shipped API does not have. |
| `7-Use-Cases.md` | REBUILD | Its flagship flow depends on removed `ccs` profiles and removed send parameters. |
| `8-Roadmap.md` | REBUILD | It is a historical plan, but it presents old implementation choices and planned work as the current stack. |
| `9-Implementation.md` | REBUILD | Its package, class, endpoint, and outcome map has drifted from the source. |
| `10-Cross-Host-Messaging.md` | REVISE | It labels most federation work proposed, but misses the shipped `coordinator:` lead channel. |
| `11-Features.md` | REVISE | It is the right catalogue, but code-path names are old and it misses the second-herdr-daemon capability. |
| `12-Claude-to-OpenCode.md` | REVISE | The porting guide is mostly current, but calls the product and spawned-member path a bridge. |
| `13-User-Guide.md` | REVISE | It is the best operator page, but needs the product rename and the second-herdr-daemon setup. |
## Pages needing work
### `Home.md` — REVISE
- Quote: `# claude-bridge` (line 1) and `` `claude-bridge` keeps`` (line 11).
The product is `fleet` / `fleetd`. The MCP server identifies itself as `fleet` in
`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:313-315`.
- Quote: `AgentAPI ... swappable fallback injector` (lines 73-76).
There is no AgentAPI implementation under `fleetd/src/main/java`; the actual
launchers are selected by `Profile.kind` in
`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:265-270`.
### `_Sidebar.md` — REVISE
- Quote: `### 📖 claude-bridge` (line 1).
Rename it to `fleet`. `FleetMcp` registers the current product-facing tool set at
`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:301-326`.
### `1-Architecture.md` — REBUILD
- Quote: `` `claude-bridge` lets`` (line 3). The product was renamed; the MCP
server name is `fleet` (`FleetMcp.java:313-315`).
- Quote: ``fleet_read`` in the tool list (line 102). No such tool is registered.
The complete registered list is `fleet_send` through `fleet_whoami` at
`FleetMcp.java:301-326`; `fleet_read` is absent.
- Quote: `SSE (GET /events)` (line 143). `FleetApp.build()` registers no `/events`
route; its routes are listed at `FleetApp.java:143-159`.
- Quote: `Redis Streams / NATS JetStream, or an embedded queue` (line 106).
The shipped durable inbox is AMQP, configured by `broker`, at
`FleetConfig.java:49-50` and `FleetConfig.java:655-714`.
- Quote: `AgentAPI (fallback)` (line 107). No AgentAPI adapter exists; shipped
launcher kinds are `claude-code` and `opencode` (`FleetConfig.java:265-270`).
### `2-Message-Server.md` — REBUILD
- Quote: `claude mcp add --transport http bridge http://127.0.0.1:8080/mcp`
(line 67). The daemon defaults to port `8765` in `FleetConfig.java:183-187`,
and identifies its server as `fleet` at `FleetMcp.java:313-315`.
- Quote: ``fleet_send(message, target?, {block, timeout_seconds, auto_spawn,
turn_id})`` (line 80). The real parameters are `sessionId`, `content`,
`timeoutMs`, `wait`, `turnId`, and `coordId` (`FleetMcp.java:1096-1108`).
- Quote: ``fleet_read(target, source)`` (line 85). It is not registered; see the
complete registration at `FleetMcp.java:301-326`.
- Quote: `docs/MCP-Contract.md ... normative` (lines 87-88). That is not a valid
reference: only §6 is current, as the current operator guide itself says at
`.wiki-snapshot/13-User-Guide.md:466`.
- Quote: `SSE (GET /events)` (line 45). No route exists in the built REST surface,
`FleetApp.java:143-159`.
### `3-Approaches.md` — REVISE
- Quote: `AgentAPI ... remains a swappable fallback injector` (lines 78-84).
It was never built. The shipped adapter selection is only `claude-code` or
`opencode` (`FleetConfig.java:265-270`). Keep it as discarded research, not an
operational fallback.
- Quote: `claude-bridge` (line 109). Rename the product to `fleet`; the runtime
package is `dev.ltms.fleet`, for example `FleetMcp.java:1`.
### `4-Setup.md` — RETIRE
It is a 25-line redirect and says its procedure was never written (lines 3-9).
Chapter 13 is the maintained install procedure. Keeping a second navigation page
adds no working documentation.
### `5-Operations.md` — RETIRE
It is a 35-line redirect and says its runbook was never written (lines 3-14).
Chapter 13 now owns run and recovery instructions.
### `6-Team.md` — REBUILD
- Quote: `fleet_send {role: w-claude, prompt: A}` (line 98). `fleet_send` accepts
`sessionId` and `content`, not `role` or `prompt` (`FleetMcp.java:1096-1108`).
- Quote: `some on Claude, some on the remote local LLM` (lines 3-5) and `Every
worker is ... Claude Code` (line 25). `opencode` is a first-class launcher kind,
not a Claude worker (`FleetConfig.java:265-270`).
- Quote: `fleetd's concurrency policy` (line 121). The configured capacity control
is per-profile `maxLoad` (`FleetConfig.java:251-264`), not the role routing model
described here.
### `7-Use-Cases.md` — REBUILD
- Quote: `ccs profile` (line 10), `ccs + herdr` (line 22), and `ccs-spawn`
(line 45). The configuration has `profiles` and `fleet`, not `ccs`:
`FleetConfig.java:34-58` and `FleetConfig.java:81-101`.
- Quote: `fleet_send({"to", "kind", "body", "block"})` (lines 55-62).
None of those are the shipped send parameters. The schema is
`FleetMcp.java:1096-1108`.
- Quote: `fleet_list() → { "profiles": ... }` (lines 74-80). `fleet_list` is a
roster view; `fleet_profiles` is the configured-backend view, as registered at
`FleetMcp.java:307-311` and described at `FleetMcp.java:1176-1182`.
### `8-Roadmap.md` — REBUILD
- Quote: `Java 21+` (line 43). The current project guidance and source use Java 25;
the `FleetConfig` source itself uses Java 25 unnamed lambda parameters, for
example `FleetConfig.java:102`.
- Quote: `herdr 0.7.0 / protocol 14` (line 46). The current REST health endpoint
reports the live protocol returned by herdr (`FleetApp.java:240-244`), while the
current operator guide records protocol 19 at
`.wiki-snapshot/13-User-Guide.md:76-85`.
- Quote: `ccs <profile> claude` and `ccs env <profile>` (lines 47-48). Shipped
configuration uses `Profile` records and launcher `kind`,
`FleetConfig.java:313-330` and `FleetConfig.java:265-270`.
- Quote: `Redis Streams via Lettuce` (line 50). The actual durable inbox is AMQP
`broker`, `FleetConfig.java:655-714`.
### `9-Implementation.md` — REBUILD
- Quote: `rest.FleetdApp` and `mcp.BridgeMcp` (lines 29-30). The classes are
`rest.FleetApp` and `mcp.FleetMcp` (`FleetApp.java:46`; `FleetMcp.java:67`).
- Quote: `dev.ltms.fleetd` (line 67). The source package is `dev.ltms.fleet`
(`FleetMcp.java:1`).
- Quote: `WorkerPresence` (line 110). The current class is `MemberPresence`, as
imported and used by `FleetMcp` at `FleetMcp.java:12` and `465-469`.
- Quote: the outcome list ending in `STALE_TURN` (lines 128-131). The code also
has `BACKEND_EXHAUSTED` (`FleetMcp.java:550-554`) and async `ASKING` handling
(`FleetMcp.java:664-668`).
- Quote: `FleetdApp` (line 207) and `FleetdConfig` (line 211). These names do not
resolve; current classes are `FleetApp` and `FleetConfig`.
### `10-Cross-Host-Messaging.md` — REVISE
- Quote: the chapter says the cross-host fabric is proposed except for the
single-host inbox (lines 3-8). Cross-host **lead-to-lead** delivery shipped:
`fleet_send` accepts `coordId` (`FleetMcp.java:1094-1107`) and publishes it at
`FleetMcp.java:616-641`; configuration has `coordinator` at
`FleetConfig.java:74-78` and `99-101`.
- Quote: `bridge.dlx` (line 90). This product name is stale. The shipped lead path
uses `LeadChannel`, not the proposed exchange flow (`FleetMcp.java:95-96` and
`616-641`). Keep the proposed federation design, but add a clear shipped/proposed
boundary for CB-637.
### `11-Features.md` — REVISE
- Quote: `mcp/BridgeMcp` (line 22), `config/FleetdConfig` (lines 25-27), and other
index references. These paths no longer resolve; the source classes are
`mcp/FleetMcp` (`FleetMcp.java:67`) and `config/FleetConfig`
(`FleetConfig.java:81`).
- Quote: `fleet_whoami` returns only `primary` or `worker` (lines 99-100).
It also returns `architect` (`FleetMcp.java:1235-1244`).
- The page needs the missing separate member-herdr-daemon feature listed below.
### `12-Claude-to-OpenCode.md` — REVISE
- Quote: `same bridge mount` (line 5) and `a bridge-spawned worker` (line 94).
Rename the product path to `fleet`. The daemon exposes the MCP server as `fleet`
(`FleetMcp.java:313-315`), and profiles select OpenCode with `kind: opencode`
(`FleetConfig.java:332-335`).
- Quote: the sample mount name is `fleetd` (line 67). The server name is `fleet`;
update the sample to avoid teaching a second product name.
### `13-User-Guide.md` — REVISE
- Quote: `The bridge is the only channel` (line 63). The invariant is correct, but
the product term needs the `fleet` rename. The daemon's MCP server name is
`fleet` (`FleetMcp.java:313-315`).
- Quote: it describes one herdr socket (lines 72-85). It needs the optional
`memberHerdrSocket` setup and two-daemon health meaning. The config key is in
`FleetConfig.java:34-37`, and `/healthz` checks both daemons when configured at
`FleetApp.java:210-245`.
## MISSING
`11-Features.md` has a body section for **routing members through a separate herdr daemon**
(`## memberHerdrSocket`, line 2174), but **no row in the index table** at the top of the page
(lines 20-95). That table is how the page is meant to be read, so a capability absent from it is
effectively undiscoverable. Lead note: this is my own omission — I added the section on 2026-08-31
and did not add the matching row. Fixed in the wiki at `68e32c6`'s successor.
The original audit stated the feature had no entry at all. That was wrong: the section exists. The
gap is the index row. Recorded here rather than silently corrected, because the difference matters —
"undocumented" and "documented but unindexed" are different jobs.
Evidence for the feature itself: `FleetConfig.java:34-37` and `FleetApp.java:103-115`, `210-245`,
and `247-263`.
## Audit method and coverage
I checked all 15 pages. I checked concrete tool, route, config, class, file, and
product-name claims claim-by-claim on 11 pages: Home, Sidebar, 1, 2, 4, 5, 6, 7, 9,
11, and 13. I skimmed the remaining four long historical or research pages (3, 8, 10,
12), then checked their concrete claims that affect the verdict. This is an audit of
the supplied snapshot, not a wiki rewrite.
+6 -6
View File
@@ -2,11 +2,11 @@
A standard, repeatable **live** end-to-end test of the two-way channel: it drives a real
multi-turn conversation between a primary and an off-subscription worker **through the
running `bridged` daemon**, captures the full transcript, and grades the channel.
running `fleetd` daemon**, captures the full transcript, and grades the channel.
This is the committed form of the ad-hoc channel test that discovered the CB-115 gaps
(herdr `unknown` misclassification wedging delivery, dirty completion scrapes, and workers
never calling `bridge_reply` in conversation). Run it after any change to the injector,
never calling `fleet_reply` in conversation). Run it after any change to the injector,
status handling, completion/failure paths, or the worker reply charter.
## What it exercises
@@ -20,7 +20,7 @@ construction** — it only calls the bridge's loopback REST face.
```mermaid
sequenceDiagram
participant T as conversation_test.py
participant B as bridged (REST)
participant B as fleetd (REST)
participant W as worker (off-sub)
T->>B: POST /workers (spawn)
T->>B: GET /sessions/{id}/status (await ready)
@@ -28,7 +28,7 @@ sequenceDiagram
T->>B: POST /sessions/{id}/message {wait:false}
B-->>T: ticket
B->>W: inject prompt (status-gated)
W-->>B: bridge_reply
W-->>B: fleet_reply
T->>B: GET /tasks/{ticket} (poll)
B-->>T: done + reply
end
@@ -37,7 +37,7 @@ sequenceDiagram
## Prerequisites
- `bridged` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
- `fleetd` is running (default REST on `http://127.0.0.1:8765`) with at least one worker
profile configured and its backend reachable.
- herdr is up (the daemon needs it).
- Python 3 (standard library only — no pip installs).
@@ -68,7 +68,7 @@ Per-turn grade:
| Grade | Meaning |
|------------|---------------------------------------------------------------------|
| `OK` | delivered and resolved by an explicit `bridge_reply` (`source=reply`) |
| `OK` | delivered and resolved by an explicit `fleet_reply` (`source=reply`) |
| `DEGRADED` | delivered and answered, but resolved via completion-scrape fallback |
| `EMPTY` | turn completed but the reply was empty |
| `FAILED` | the worker's turn ended in failure (`phase=failed`) |
+8 -8
View File
@@ -1,6 +1,6 @@
#!/usr/bin/env python3
"""Sustained back-and-forth bridge test — ONE primary, ONE worker, many dependent turns
over a fixed wall-clock window (default 5 minutes), through the running `bridged` daemon.
over a fixed wall-clock window (default 5 minutes), through the running `fleetd` daemon.
Where conversation_test.py proves a handful of turns work and issue_hunt_test.py proves
fan-out isolation, this proves the channel stays healthy under a *sustained, stateful*
@@ -24,7 +24,7 @@ Usage:
--turn-timeout per-turn max wait, seconds (default 150)
--keep-worker do not stop the worker at the end
Exit code: 0 if every turn in the window resolved via a clean bridge_reply with no channel
Exit code: 0 if every turn in the window resolved via a clean fleet_reply with no channel
break; 1 otherwise. A live per-turn log streams to stdout so the run can be watched.
"""
import argparse
@@ -47,12 +47,12 @@ STEPS = [7, 3, 11, 5, 9, 4, 13, 6, 8, 2]
RULES = (
"Let's play a running-total game across several messages. The total starts at 0. "
"In each message I'll tell you to add a number; keep the running total yourself and "
"reply via bridge_reply with ONLY the current total as a plain integer — no words, no "
"reply via fleet_reply with ONLY the current total as a plain integer — no words, no "
"punctuation, just the number. Do not restate the arithmetic. First move: add {step}."
)
NEXT = ("Add {step}. Reply via bridge_reply with only the new running total.")
NEXT = ("Add {step}. Reply via fleet_reply with only the new running total.")
REANCHOR = ("Let's re-sync — the running total is {total}. Now add {step}. Reply via "
"bridge_reply with only the new running total.")
"fleet_reply with only the new running total.")
def parse_int(reply):
@@ -142,15 +142,15 @@ def main():
print("=" * 72)
print(f"SUSTAINED CONVERSATION SUMMARY — 1 primary <-> 1 worker over {dur}s (~{dur/60:.1f} min)")
print(f" turns: {turns}")
print(f" clean bridge_reply exchanges: {oks + drifts}/{turns} (channel breaks: {breaks})")
print(f" clean fleet_reply exchanges: {oks + drifts}/{turns} (channel breaks: {breaks})")
print(f" arithmetic correct (continuity held): {oks}/{turns} (drifts: {drifts})")
print(f" latency: avg {avg}s over {turns} turns")
ok = breaks == 0 and turns >= 2
if ok and drifts == 0:
print(" RESULT: PASS — every turn resolved via bridge_reply and the worker held the "
print(" RESULT: PASS — every turn resolved via fleet_reply and the worker held the "
"running total across the whole window.")
elif ok:
print(f" RESULT: PASS (channel) — every turn resolved via bridge_reply for the full "
print(f" RESULT: PASS (channel) — every turn resolved via fleet_reply for the full "
f"window; {drifts} arithmetic drift(s) (worker recovered after re-anchor).")
else:
print(" RESULT: FAIL — the channel broke on at least one turn (see CHANNEL BREAK above).")
+5 -5
View File
@@ -1,10 +1,10 @@
#!/usr/bin/env python3
"""Standard bridge conversation test — a multi-turn primary↔worker exchange through
the running `bridged` daemon, fully captured, with automatic gap analysis.
the running `fleetd` daemon, fully captured, with automatic gap analysis.
This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 gaps
(herdr `unknown` misclassification, dirty completion scrape, workers not calling
bridge_reply). It drives a real off-subscription worker over the live gateway exactly
fleet_reply). It drives a real off-subscription worker over the live gateway exactly
as a primary Opus session would (async fire-and-poll), records every turn, and grades
the channel.
@@ -130,9 +130,9 @@ def grade(rec):
phase, source, reply = rec["phase"], rec["source"], rec["reply"]
has_reply = bool(reply and reply.strip())
if phase == "done" and source == "reply" and has_reply:
return "OK", "clean explicit bridge_reply"
return "OK", "clean explicit fleet_reply"
if phase == "done" and has_reply:
return "DEGRADED", f"resolved via {source} (worker did not call bridge_reply)"
return "DEGRADED", f"resolved via {source} (worker did not call fleet_reply)"
if phase == "done" and not has_reply:
return "EMPTY", "turn completed but reply was empty"
if phase == "failed":
@@ -206,7 +206,7 @@ def main():
ok = all(g in ("OK", "DEGRADED") for g in grades)
reply_clean = all(g == "OK" for g in grades)
if reply_clean:
print(" RESULT: PASS — every turn delivered and got a clean bridge_reply.")
print(" RESULT: PASS — every turn delivered and got a clean fleet_reply.")
elif ok:
print(" RESULT: PASS (with notes) — every turn delivered & replied, but some via fallback.")
else:
@@ -1,16 +1,16 @@
#!/usr/bin/env python3
"""Live bridge_ask test — the REVERSE rendezvous (CB-205), watched end to end.
"""Live fleet_ask test — the REVERSE rendezvous (CB-205), watched end to end.
Every other harness drives the forward path: primary `bridge_send` → worker `bridge_reply`.
Every other harness drives the forward path: primary `fleet_send` → worker `fleet_reply`.
This drives the one that runs the other way. A worker is told to pause its delegated turn,
ask the primary a question via `bridge_ask`, and only finish once it has the answer — so the
ask the primary a question via `fleet_ask`, and only finish once it has the answer — so the
turn round-trips primary→worker→primary→worker inside a SINGLE delegation.
The mechanics that only this path exercises:
• a worker's mid-turn question surfacing on the primary's *own* blocked send (Outcome.QUESTION),
• the `turnId` correlation that lets the primary answer the exact paused turn,
• the answer resuming that same turn and the worker's final `bridge_reply` landing on the
• the answer resuming that same turn and the worker's final `fleet_reply` landing on the
re-opened forward waiter (never a stale or cross-wired one).
It is two blocking REST calls, no polling:
@@ -24,7 +24,7 @@ Like the rest of the suite it talks ONLY to the bridge's REST face on loopback
ANTHROPIC_BASE_URL and never touches herdr, so it is subscription-safe by construction.
Usage:
python3 bridge_ask_test.py [--base URL] [--profile NAME] [--repo DIR] [--out DIR]
python3 fleet_ask_test.py [--base URL] [--profile NAME] [--repo DIR] [--out DIR]
[--send-timeout SECS] [--answer-timeout SECS] [--keep-worker]
--base bridge REST base URL (default http://127.0.0.1:8765)
@@ -57,14 +57,14 @@ ANSWER_COLOR = "blue"
# A task that CANNOT be completed without asking: the worker is not told which color to choose,
# only that the primary will name one when asked. So a correct final reply is only reachable by
# actually calling bridge_ask and using the answer.
# actually calling fleet_ask and using the answer.
TASK_PROMPT = (
"You are a bridge worker in a quick coordination game. You do NOT know which color to pick — "
"only the primary does. Do exactly this, in order:\n"
"1. Call the `bridge_ask` tool with EXACTLY this question: \"PICK A COLOR: red or blue?\"\n"
"1. Call the `fleet_ask` tool with EXACTLY this question: \"PICK A COLOR: red or blue?\"\n"
"2. The primary will answer with one color word. Take that color and uppercase it.\n"
"3. Call `bridge_reply` with EXACTLY one line: CHOSEN=<COLOR> (e.g. CHOSEN=GREEN if told green).\n"
"Do not guess a color. Do not call bridge_reply before bridge_ask has returned an answer. "
"3. Call `fleet_reply` with EXACTLY one line: CHOSEN=<COLOR> (e.g. CHOSEN=GREEN if told green).\n"
"Do not guess a color. Do not call fleet_reply before fleet_ask has returned an answer. "
"Do nothing else — no file reads, no other tools."
)
@@ -134,7 +134,7 @@ def run(base, profile, repo, send_timeout, answer_timeout):
rec.update(tid=tid, pane=pane, spawned=True)
await_ready(base, tid)
# 1) Delegate the ask-forcing task. This blocks until the worker calls bridge_ask, at which
# 1) Delegate the ask-forcing task. This blocks until the worker calls fleet_ask, at which
# point our own send unblocks carrying the question and the turnId to answer on.
print(f"[{now()}] delegating task (blocks until the worker asks; up to {send_timeout}s)…")
t0 = time.time()
@@ -159,7 +159,7 @@ def run(base, profile, repo, send_timeout, answer_timeout):
rec.update(phase="no_turnid", detail="question surfaced without a turnId to answer on")
return rec
# 2) Answer on that exact turn. This blocks again until the resumed worker calls bridge_reply.
# 2) Answer on that exact turn. This blocks again until the resumed worker calls fleet_reply.
print(f"[{now()}] answering '{ANSWER_COLOR}' on turn {rec['turnId']} (blocks until reply; up to {answer_timeout}s)…")
t1 = time.time()
rec["phase"] = "awaiting_reply"
@@ -188,7 +188,7 @@ def run(base, profile, repo, send_timeout, answer_timeout):
def grade(rec):
"""PASS only if the worker asked, the turn resumed, and the reply reflects the answer."""
if rec["phase"] == "no_question":
return "NO_ASK", "the worker finished/stalled without ever calling bridge_ask"
return "NO_ASK", "the worker finished/stalled without ever calling fleet_ask"
if rec["phase"] in ("send_error", "answer_error", "spawn"):
return "ERROR", rec.get("detail") or "transport error before the round-trip completed"
if rec["phase"] == "no_turnid":
@@ -203,26 +203,26 @@ def grade(rec):
return "OK", "asked, resumed the same turn, and the reply reflected the primary's answer"
if reflected:
return "DEGRADED", f"reply reflected the answer but resolved via {rec['replySource']} " \
"(worker did not call bridge_reply cleanly)"
"(worker did not call fleet_reply cleanly)"
return "WRONG_ANSWER", f"the worker replied but did not reflect '{ANSWER_COLOR}' — " \
f"the answer may not have reached the resumed turn: {rec['reply']!r}"
return "WEDGE", f"unexpected terminal phase {rec['phase']}: {rec.get('detail')}"
def write_transcript(out_dir, rec, meta):
path = out_dir / "bridge_ask_transcript.md"
path = out_dir / "fleet_ask_transcript.md"
g, note = grade(rec)
with path.open("w") as f:
f.write(f"# Live bridge_ask — reverse rendezvous — {datetime.now():%Y-%m-%d %H:%M}\n\n")
f.write(f"# Live fleet_ask — reverse rendezvous — {datetime.now():%Y-%m-%d %H:%M}\n\n")
f.write(f"One worker paused its delegated turn to ask the primary, then resumed with the "
f"answer (profile `{meta['profile']}`). Result: **`{g}`**.\n\n")
f.write("## Round-trip\n\n")
f.write(f"1. **primary → worker** (delegation): the ask-forcing task.\n")
f.write(f"2. **worker → primary** (`bridge_ask`, {rec.get('ask_latency')}s): "
f.write(f"2. **worker → primary** (`fleet_ask`, {rec.get('ask_latency')}s): "
f"{rec.get('question')!r} — surfaced on the primary's blocked send as a "
f"`question` with `turnId={rec.get('turnId')}`.\n")
f.write(f"3. **primary → worker** (answer on that turn): `{ANSWER_COLOR}`.\n")
f.write(f"4. **worker → primary** (`bridge_reply`, {rec.get('answer_latency')}s, "
f.write(f"4. **worker → primary** (`fleet_reply`, {rec.get('answer_latency')}s, "
f"source={rec.get('replySource')}): {rec.get('reply')!r}\n\n")
f.write(f"> **{g}:** {note}\n")
if rec.get("detail"):
@@ -231,7 +231,7 @@ def write_transcript(out_dir, rec, meta):
def main():
ap = argparse.ArgumentParser(description="Live bridge_ask reverse-rendezvous test (CB-205)")
ap = argparse.ArgumentParser(description="Live fleet_ask reverse-rendezvous test (CB-205)")
ap.add_argument("--base", default="http://127.0.0.1:8765")
ap.add_argument("--profile", default=None)
ap.add_argument("--repo", default=str(REPO_ROOT))
@@ -241,7 +241,7 @@ def main():
ap.add_argument("--keep-worker", action="store_true")
args = ap.parse_args()
print(f"[{now()}] live bridge_ask: 1 primary, 1 worker "
print(f"[{now()}] live fleet_ask: 1 primary, 1 worker "
f"(profile={args.profile or 'default'}, repo={args.repo})\n")
rec = {"spawned": False, "pane": None}
@@ -262,7 +262,7 @@ def main():
print()
print("=" * 72)
print("LIVE bridge_ask SUMMARY — reverse rendezvous (CB-205)")
print("LIVE fleet_ask SUMMARY — reverse rendezvous (CB-205)")
print(f" asked: {rec.get('question')!r} (turnId={rec.get('turnId')}, {rec.get('ask_latency')}s)")
print(f" answered: {ANSWER_COLOR!r}")
print(f" replied: {rec.get('reply')!r} (source={rec.get('replySource')}, {rec.get('answer_latency')}s)")

Some files were not shown because too many files have changed in this diff Show More