LeadMailbox: durable leader-to-leader mailbox over a shared coordination vhost #166

Closed
agent wants to merge 0 commits from worker/lead-mailbox-c19577-6 into main
Member

Unit 1 of 2 (see ticket): the broker-side mechanism for lead-to-lead messages across daemons/hosts, over a shared AMQP coordination vhost. A second ticket (owned by the primary) wires this into fleet_send/fleet_list.

Adds

  • LeadMessage (msg/LeadMessage.java): wire envelope (msgId, from, to, content) — carries an explicit sender, unlike ReplyInbox.InboxMessage, because a lead needs the sender's coord-id to reply.
  • LeadMailbox (msg/LeadMailbox.java): AMQP-backed durable mailbox modeled closely on AmqpReplyInbox — same connection/channel/confirm/recovery style, consume-and-hold with deferred manual ack, confirm-mode publish (persistent, mandatory, 10s confirm timeout), automatic + topology recovery. Simpler than AmqpReplyInbox: single-target (this daemon owns exactly one queue, lead.<selfCoordId>.inbox, from construction) rather than per-worker multiplexed. Publish never implies ownership (CB-308 federation).
  • FleetConfig.Coordinator record: optional coordinator: block (uri/uriEnv/selfId/prefetch), mirroring Broker's effectiveUri/hasUriEnv/prefetchOrDefault. A SEPARATE vhost from broker: — member inboxes stay per-fleet, this is leader-only traffic. Added as the new 19th canonical record field with a same-shaped back-compat 18-arg constructor (coordinator=null) so existing call sites/tests keep compiling. Added to KNOWN_TOP_LEVEL_KEYS. Config parsing + accessors only — nothing wires it into a live LeadMailbox yet.
  • fleetd.example.yaml: commented coordinator: block documenting the split from broker:.
  • Tests: LeadMailboxTest (@Tag("contract"), mirrors AmqpReplyInboxContractTest's Testcontainers/AMQP_URI strategy — excluded from the default build, runs under -Pcontract with Docker or CI's AMQP_URI) proving publish→peek→ack, dedup, redelivery-after-restart, federation (publish without owning), and unroutable-publish failure. FleetConfigTest additions proving the coordinator: block parses, effectiveUri honors uriEnv, and an absent block leaves coordinator() null.

Out of scope (per ticket) — not touched: FleetMcp, Fleetd wiring, Injector, MessageService, the send path.

Build: mvn clean install — BUILD SUCCESS, Tests run: 922, Failures: 0, Errors: 0, Skipped: 0 (a first full run stalled on an unrelated, pre-existing flake in MessageServiceTest/Injector — confirmed unrelated to this change: that single test and the whole class pass cleanly in isolation, and a second full clean run went green in 28s with all 922 tests passing).

Unit 1 of 2 (see ticket): the broker-side mechanism for lead-to-lead messages across daemons/hosts, over a shared AMQP coordination vhost. A second ticket (owned by the primary) wires this into fleet_send/fleet_list. **Adds** - `LeadMessage` (`msg/LeadMessage.java`): wire envelope `(msgId, from, to, content)` — carries an explicit sender, unlike `ReplyInbox.InboxMessage`, because a lead needs the sender's coord-id to reply. - `LeadMailbox` (`msg/LeadMailbox.java`): AMQP-backed durable mailbox modeled closely on `AmqpReplyInbox` — same connection/channel/confirm/recovery style, consume-and-hold with deferred manual ack, confirm-mode publish (persistent, mandatory, 10s confirm timeout), automatic + topology recovery. Simpler than `AmqpReplyInbox`: single-target (this daemon owns exactly one queue, `lead.<selfCoordId>.inbox`, from construction) rather than per-worker multiplexed. Publish never implies ownership (CB-308 federation). - `FleetConfig.Coordinator` record: optional `coordinator:` block (`uri`/`uriEnv`/`selfId`/`prefetch`), mirroring `Broker`'s `effectiveUri`/`hasUriEnv`/`prefetchOrDefault`. A SEPARATE vhost from `broker:` — member inboxes stay per-fleet, this is leader-only traffic. Added as the new 19th canonical record field with a same-shaped back-compat 18-arg constructor (coordinator=null) so existing call sites/tests keep compiling. Added to `KNOWN_TOP_LEVEL_KEYS`. Config parsing + accessors only — nothing wires it into a live `LeadMailbox` yet. - `fleetd.example.yaml`: commented `coordinator:` block documenting the split from `broker:`. - Tests: `LeadMailboxTest` (`@Tag("contract")`, mirrors `AmqpReplyInboxContractTest`'s Testcontainers/AMQP_URI strategy — excluded from the default build, runs under `-Pcontract` with Docker or CI's AMQP_URI) proving publish→peek→ack, dedup, redelivery-after-restart, federation (publish without owning), and unroutable-publish failure. `FleetConfigTest` additions proving the `coordinator:` block parses, `effectiveUri` honors `uriEnv`, and an absent block leaves `coordinator()` null. **Out of scope (per ticket)** — not touched: `FleetMcp`, `Fleetd` wiring, `Injector`, `MessageService`, the send path. **Build**: `mvn clean install` — BUILD SUCCESS, Tests run: 922, Failures: 0, Errors: 0, Skipped: 0 (a first full run stalled on an unrelated, pre-existing flake in `MessageServiceTest`/`Injector` — confirmed unrelated to this change: that single test and the whole class pass cleanly in isolation, and a second full clean run went green in 28s with all 922 tests passing).
agent added 6 commits 2026-08-24 17:24:07 +02:00
Adds `ideMcpUrl` to FleetConfig.Profile (default off). When set, the
Claude Code launcher mounts the IDE Index MCP as a second inline
--mcp-config server named `intellij`, and appends an IDE charter that
pins every ide_* call to the member's own worktree (spec.cwd()). The
charter order is role -> ide -> reply, one --append-system-prompt-file,
reply last (CB-618). The mount gate now fires on ideMcpUrl alone, not
only mcpUrl. ConfigRef treats an ideMcpUrl change as deferred, like the
other launch flags.

Never touches .mcp.json or CLAUDE.md — the mount and the rule arrive as
launch flags, so a project's own config is untouched.

Not yet done (see fleetd #162): the bridged-owned IDE lifecycle
(open on provision, close before worktree removal), and the opencode
adapter (separate ticket). fleetd.example.yaml documents ideMcpUrl and
fixes the stale parityOverlay default.

911 tests green.
Move the IDE guidance text to PeerLauncher.ideOverlayText (shared by both
launchers). ClaudeCodeLauncher drops it from the reply-charter file and writes
CLAUDE.local.md into a provisioned worktree instead, gated on a .git FILE
(safety: never writes into the primary's real .git-DIRECTORY checkout) and
registers it in info/exclude. OpenCodeLauncher mounts the intellij server and
adds the rules file to the instructions array.
git reads info/exclude from the common dir for a linked worktree (only
info/sparse-checkout is per-worktree), so the entry written into
<common>/worktrees/<name>/info/exclude was never honoured and CLAUDE.local.md
showed as untracked -- at risk of being swept into a worker's PR. Derive the
common dir (<common>/worktrees/<name> -> <common>) and write there. Found by
dogfooding a real spawn on fleet01; the test now uses the real worktree layout
and asserts the entry lands in the common dir, not the per-worktree gitdir.
The overlay pinned project_path to the worktree root. For a repo whose Maven
module is a subdir (this repo's pom is in `bridged/`, not at the root), opening
the root imports no module and every ide_* call resolves nothing. Pin and open
the module dir instead.

Two new opt-in per-Profile keys, both read only when ideMcpUrl is set:
- ideProjectDir: repo-relative module dir the IDE opens and the overlay pins;
  blank keeps the old worktree-root behaviour.
- ideOpenCommand: host command that opens that dir in the IDE at spawn, with
  {dir} substituted and run through /bin/sh -c so env (e.g. DISPLAY) can be set
  inline. Best-effort and non-fatal — a failure never fails the spawn. Blank
  keeps the manual-open behaviour. No close half yet (deferred).

Shared helpers PeerLauncher.ideProjectPath / openInIde back both launchers.
The two Profile fields ride a back-compat constructor, so every existing call
site and YAML compiles and behaves unchanged.

Tests: overlay content pins the module dir when ideProjectDir is set;
ideProjectPath resolution; openInIde no-op on a blank command. 918 tests green.
LeadMailbox: durable leader-to-leader mailbox over a shared coordination vhost
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m35s
7b98cca967
Adds the broker-side mechanism for lead-to-lead messages across daemons/hosts
(unit 1 of 2): a LeadMessage envelope carrying from/to coord-ids, an
AMQP-backed LeadMailbox modeled closely on AmqpReplyInbox (consume-and-hold,
deferred manual ack, confirm-mode publish, recovery handling), and a new
optional coordinator: config block (separate vhost from broker:, leader
traffic only). Config parsing + accessors only — FleetMcp/Fleetd/Injector/
MessageService and the send path are untouched; wiring is a separate ticket.
Owner

Landed on main via the cb-634-ide-mcp integration merge (450a5ed), not this PR. LeadMailbox + LeadMessage are now in main; the 5 contract tests pass against a real RabbitMQ. Closing as redundant.

Landed on `main` via the `cb-634-ide-mcp` integration merge (`450a5ed`), not this PR. `LeadMailbox` + `LeadMessage` are now in `main`; the 5 contract tests pass against a real RabbitMQ. Closing as redundant.
ltms closed this pull request 2026-08-24 19:12:59 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m35s

Pull request closed

Sign in to join this conversation.