lead comms: wire LeadMailbox into the daemon (fleet_send{coordId} + receive loop) #167

Closed
agent wants to merge 0 commits from worker/lead-comms-wiring-c014b9-7 into main
Member

Unit 2 of 2 for lead-to-lead messaging. The sibling ticket landed LeadMailbox, LeadMessage and the coordinator: config block; nothing opened, sent through, or read that mailbox. This wires it in.

What changed

LeadChannel (new) — a small interface LeadMailbox now implements: publish / peek / ack plus a selfCoordId() accessor. It is a test seam: LeadMailbox owns a live AMQP connection, so without it every test of the code around it would need a broker. LeadMailbox's AMQP logic is untouched — its whole diff is the implements clause, four @Override marks and the accessor.

Fleetd.openLeadMailbox — opens this daemon's mailbox right after the reply inbox is selected, with the same env-injected seam selectReplyInbox uses. Every "off" path returns null and the daemon still starts:

case log
no coordinator: block silent (opt-in)
uriEnv does not resolve INFO
broker configured, selfId unset WARN — a mailbox is the queue named after the coord-id that owns it
broker refuses at boot WARN, credentials stripped, carry on

Closed in the ordered shutdown hook, after the loop that reads it has stopped.

fleet_send{coordId} — publishes LeadMessage(from=selfCoordId, to=coordId) to the peer's mailbox and returns the broker-confirmed receipt. coordId is mutually exclusive with sessionId/turnId and is rejected by name rather than resolved by precedence, so a caller who meant the other route learns it. An unroutable / nacked / timed-out publish becomes a tool error naming the coordId, never a crash. The existing worker send/reply path is not touched.

LeadCoordLoop (new) — the receive half. Each tick peeks the mailbox, resolves the local lead pane, and — only at a turn boundary (AgentStatus.injectable(), exactly like ReplyPushLoop) — injects "[lead <from>] <text>" and then acks. Anything not delivered (no pane, lead mid-turn, herdr threw) stays unacked and is retried, so a message is never dropped. One message per tick, so every delivery is gated on a status read that already saw the previous one.

fleet_list — adds {"coordinator": {"selfId": ..., "configured": true}} when coordination is on. There is no peer discovery yet, so this answers the one question an operator cannot answer otherwise: which coord-id a peer must use to reach me. Omitted entirely when off.

Tests

20 new hermetic tests, no broker anywhere (the broker round trip stays the sibling ticket's @Tag("contract") LeadMailboxTest):

  • FleetMcpLeadCoordTest (6) — envelope from/to/content, both mutual-exclusion rejections, the not-configured error, an unroutable peer surfacing as a tool error.
  • LeadCoordLoopTest (8) — delivered-then-acked; not acked when mid-turn, when no lead pane is known, or when herdr refuses; lead resolved by name among several; held when ambiguous; one-per-tick FIFO; an empty mailbox never touches herdr.
  • FleetdLeadMailboxSelectionTest (6) — each on/off path, and that the URI password never reaches a log line.
  • FleetMcpTest +2 — the fleet_list coordinator row on and off.

mvn clean install: Tests run: 944, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

For the reviewer

  • Two design calls I made rather than asking: one message per tick (a second paste would act on a status read taken before the first injection), and hold rather than guess when several leads are known and none is named after coordinator.selfId (guessing types a peer's message into the wrong pane and then acks it away).
  • The delivered-but-ack-failed window can duplicate a message. That direction is deliberate and documented — a peer lead seeing a message twice is a nuisance; never seeing it is the failure this path exists to remove.
  • Not done, needs the lead: CLAUDE.md's intent-to-tool table still says a peer lead is messaged with fleet_send{sessionId} only, and wiki/11-Features.md has no entry for this. Both are outside a worker's reach — CLAUDE.md must stay byte-identical with the wiki template, and wiki/ is a submodule I must not commit.
Unit 2 of 2 for lead-to-lead messaging. The sibling ticket landed `LeadMailbox`, `LeadMessage` and the `coordinator:` config block; nothing opened, sent through, or read that mailbox. This wires it in. ## What changed **`LeadChannel` (new)** — a small interface `LeadMailbox` now implements: `publish` / `peek` / `ack` plus a `selfCoordId()` accessor. It is a test seam: `LeadMailbox` owns a live AMQP connection, so without it every test of the code around it would need a broker. `LeadMailbox`'s AMQP logic is untouched — its whole diff is the `implements` clause, four `@Override` marks and the accessor. **`Fleetd.openLeadMailbox`** — opens this daemon's mailbox right after the reply inbox is selected, with the same env-injected seam `selectReplyInbox` uses. Every "off" path returns `null` and the daemon still starts: | case | log | |---|---| | no `coordinator:` block | silent (opt-in) | | `uriEnv` does not resolve | INFO | | broker configured, `selfId` unset | WARN — a mailbox is the queue named after the coord-id that owns it | | broker refuses at boot | WARN, credentials stripped, carry on | Closed in the ordered shutdown hook, after the loop that reads it has stopped. **`fleet_send{coordId}`** — publishes `LeadMessage(from=selfCoordId, to=coordId)` to the peer's mailbox and returns the broker-confirmed receipt. `coordId` is mutually exclusive with `sessionId`/`turnId` and is rejected **by name** rather than resolved by precedence, so a caller who meant the other route learns it. An unroutable / nacked / timed-out publish becomes a tool error naming the coordId, never a crash. The existing worker send/reply path is not touched. **`LeadCoordLoop` (new)** — the receive half. Each tick peeks the mailbox, resolves the local lead pane, and — only at a turn boundary (`AgentStatus.injectable()`, exactly like `ReplyPushLoop`) — injects `"[lead <from>] <text>"` and then acks. Anything not delivered (no pane, lead mid-turn, herdr threw) stays **unacked** and is retried, so a message is never dropped. One message per tick, so every delivery is gated on a status read that already saw the previous one. **`fleet_list`** — adds `{"coordinator": {"selfId": ..., "configured": true}}` when coordination is on. There is no peer discovery yet, so this answers the one question an operator cannot answer otherwise: which coord-id a peer must use to reach me. Omitted entirely when off. ## Tests 20 new hermetic tests, no broker anywhere (the broker round trip stays the sibling ticket's `@Tag("contract")` `LeadMailboxTest`): - `FleetMcpLeadCoordTest` (6) — envelope `from`/`to`/content, both mutual-exclusion rejections, the not-configured error, an unroutable peer surfacing as a tool error. - `LeadCoordLoopTest` (8) — delivered-then-acked; **not** acked when mid-turn, when no lead pane is known, or when herdr refuses; lead resolved by name among several; held when ambiguous; one-per-tick FIFO; an empty mailbox never touches herdr. - `FleetdLeadMailboxSelectionTest` (6) — each on/off path, and that the URI password never reaches a log line. - `FleetMcpTest` +2 — the `fleet_list` coordinator row on and off. `mvn clean install`: **Tests run: 944, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.** ## For the reviewer - Two design calls I made rather than asking: **one message per tick** (a second paste would act on a status read taken before the first injection), and **hold rather than guess** when several leads are known and none is named after `coordinator.selfId` (guessing types a peer's message into the wrong pane and then acks it away). - The delivered-but-ack-failed window can duplicate a message. That direction is deliberate and documented — a peer lead seeing a message twice is a nuisance; never seeing it is the failure this path exists to remove. - **Not done, needs the lead:** `CLAUDE.md`'s intent-to-tool table still says a peer lead is messaged with `fleet_send{sessionId}` only, and `wiki/11-Features.md` has no entry for this. Both are outside a worker's reach — `CLAUDE.md` must stay byte-identical with the wiki template, and `wiki/` is a submodule I must not commit.
agent added 7 commits 2026-08-24 17:42:36 +02:00
Adds `ideMcpUrl` to FleetConfig.Profile (default off). When set, the
Claude Code launcher mounts the IDE Index MCP as a second inline
--mcp-config server named `intellij`, and appends an IDE charter that
pins every ide_* call to the member's own worktree (spec.cwd()). The
charter order is role -> ide -> reply, one --append-system-prompt-file,
reply last (CB-618). The mount gate now fires on ideMcpUrl alone, not
only mcpUrl. ConfigRef treats an ideMcpUrl change as deferred, like the
other launch flags.

Never touches .mcp.json or CLAUDE.md — the mount and the rule arrive as
launch flags, so a project's own config is untouched.

Not yet done (see fleetd #162): the bridged-owned IDE lifecycle
(open on provision, close before worktree removal), and the opencode
adapter (separate ticket). fleetd.example.yaml documents ideMcpUrl and
fixes the stale parityOverlay default.

911 tests green.
Move the IDE guidance text to PeerLauncher.ideOverlayText (shared by both
launchers). ClaudeCodeLauncher drops it from the reply-charter file and writes
CLAUDE.local.md into a provisioned worktree instead, gated on a .git FILE
(safety: never writes into the primary's real .git-DIRECTORY checkout) and
registers it in info/exclude. OpenCodeLauncher mounts the intellij server and
adds the rules file to the instructions array.
git reads info/exclude from the common dir for a linked worktree (only
info/sparse-checkout is per-worktree), so the entry written into
<common>/worktrees/<name>/info/exclude was never honoured and CLAUDE.local.md
showed as untracked -- at risk of being swept into a worker's PR. Derive the
common dir (<common>/worktrees/<name> -> <common>) and write there. Found by
dogfooding a real spawn on fleet01; the test now uses the real worktree layout
and asserts the entry lands in the common dir, not the per-worktree gitdir.
The overlay pinned project_path to the worktree root. For a repo whose Maven
module is a subdir (this repo's pom is in `bridged/`, not at the root), opening
the root imports no module and every ide_* call resolves nothing. Pin and open
the module dir instead.

Two new opt-in per-Profile keys, both read only when ideMcpUrl is set:
- ideProjectDir: repo-relative module dir the IDE opens and the overlay pins;
  blank keeps the old worktree-root behaviour.
- ideOpenCommand: host command that opens that dir in the IDE at spawn, with
  {dir} substituted and run through /bin/sh -c so env (e.g. DISPLAY) can be set
  inline. Best-effort and non-fatal — a failure never fails the spawn. Blank
  keeps the manual-open behaviour. No close half yet (deferred).

Shared helpers PeerLauncher.ideProjectPath / openInIde back both launchers.
The two Profile fields ride a back-compat constructor, so every existing call
site and YAML compiles and behaves unchanged.

Tests: overlay content pins the module dir when ideProjectDir is set;
ideProjectPath resolution; openInIde no-op on a blank command. 918 tests green.
LeadMailbox: durable leader-to-leader mailbox over a shared coordination vhost
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m35s
7b98cca967
Adds the broker-side mechanism for lead-to-lead messages across daemons/hosts
(unit 1 of 2): a LeadMessage envelope carrying from/to coord-ids, an
AMQP-backed LeadMailbox modeled closely on AmqpReplyInbox (consume-and-hold,
deferred manual ack, confirm-mode publish, recovery handling), and a new
optional coordinator: config block (separate vhost from broker:, leader
traffic only). Config parsing + accessors only — FleetMcp/Fleetd/Injector/
MessageService and the send path are untouched; wiring is a separate ticket.
lead comms: wire LeadMailbox into the daemon so lead-to-lead messages flow
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Failing after 1m22s
6058b8472b
The sibling ticket landed the mechanism (LeadMailbox, LeadMessage, the
coordinator: config block) but nothing opened it, nothing sent through it, and
nothing read it. This is the wiring.

- LeadChannel: a small interface LeadMailbox now implements (publish/peek/ack
  plus a selfCoordId() accessor). It exists so FleetMcp and the receive loop can
  be tested with a fake instead of a live broker. LeadMailbox's AMQP logic is
  untouched — the diff is the implements clause, four @Override marks and the
  accessor.

- Fleetd.openLeadMailbox: opens this daemon's mailbox after the reply inbox is
  selected, with the same env-injected seam selectReplyInbox uses. Every "off"
  path returns null and the daemon still starts: no coordinator block (silent),
  a uriEnv that does not resolve (INFO), a configured broker with no selfId
  (WARN — a mailbox is named after the coord-id that owns it), or a broker that
  refuses at boot (WARN, credentials stripped). Closed in the ordered shutdown
  hook, after the loop that reads it has stopped.

- fleet_send{coordId}: publishes a LeadMessage(from=selfCoordId, to=coordId) to
  the peer's mailbox and returns the broker-confirmed receipt. coordId is
  mutually exclusive with sessionId/turnId and is rejected by name rather than
  resolved by precedence. An unroutable/nacked/timed-out publish comes back as a
  tool error naming the coordId, never a crash. The worker send/reply path is
  not touched.

- LeadCoordLoop: the receive half. Each tick peeks the mailbox, resolves the
  local lead pane, and — only at a turn boundary — injects "[lead <from>] <text>"
  and acks. Anything not delivered stays unacked and is retried, so a message is
  never dropped; one message per tick, so every delivery is gated on a status
  read that already saw the previous one.

- fleet_list reports {selfId, configured} when coordination is on, so an
  operator can find the coord-id a peer must use to reach them. Omitted
  entirely when it is off.

Tests: 20 new hermetic tests (no broker) across routing, delivery and startup
selection. mvn clean install: Tests run: 944, Failures: 0, Errors: 0, Skipped: 0
— BUILD SUCCESS.
Owner

Landed on main via the cb-634-ide-mcp integration merge (450a5ed), not this PR. The wiring (fleet_send{coordId} + LeadCoordLoop + Fleetd openLeadMailbox) is now in main and live: lead coordination: ON as coord-id opus, and the send/unroutable/mutual-exclusion + status-gated receive loop were all verified live. The CB-635 comment collision was relabeled to CB-637. Closing as redundant.

Landed on `main` via the `cb-634-ide-mcp` integration merge (`450a5ed`), not this PR. The wiring (`fleet_send{coordId}` + `LeadCoordLoop` + Fleetd `openLeadMailbox`) is now in `main` and live: `lead coordination: ON as coord-id opus`, and the send/unroutable/mutual-exclusion + status-gated receive loop were all verified live. The CB-635 comment collision was relabeled to CB-637. Closing as redundant.
ltms closed this pull request 2026-08-24 19:13:01 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Failing after 1m22s

Pull request closed

Sign in to join this conversation.