Part of #145 (CB-632). Documentation only, plus one internal literal. Unit 1 renamed the package and classes, which left every doc describing classes that no longer exist. This fixes the prose across README.md, docs/ and bridged/docs/ -- 18 files. Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and "bridged" where it names the daemon as a product rather than a path. Also renamed two literals, because a doc that disagrees with the code is worse than one that is out of date: - bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey OpenCodeLauncher sends when a profile resolves no token, to a local endpoint that does not check it. No test asserts the old string. - the vnd.ltms.bridged.* media type in the M4 design doc. It appears in no Java file, so nothing implements it yet. Deliberately NOT renamed, because each is still literally true today and changes only at the cutover: - paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar, .bridged-worktrees, deploy/dev.ltms.bridged.plist, scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh - bridged_* metric names -- renaming these after the monitoring is wired would break dashboard continuity, so they move before it is - bridge_* MCP tool names, which answer alongside fleet_* on purpose - BRIDGED_* environment variables, read by a file outside this repo Method note: perl, not sed. BSD sed has no \b and no lookaround, and a word-boundary expression there fails silently. The prose replace uses (?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an identifier, then every remaining hit was read by hand. Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
18 KiB
CB-308 — Multi-Host Federation (Stage 5)
Status: design note (proposal) — core design decisions resolved 2026-08-10 (§7)
Depends on: CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built on
CB-307's broker fabric.
Relates to: CB-401 (PeerHandle opaque id), CB-304 (rosterView), CB-306 (spawn-readiness),
CB-303 (lifecycle limits), CB-117 (orphan reap).
1. Goal
Let claude-bridge coordinate agents that live on more than one host — a primary on host A
delegating to workers on hosts B, C, … — without any host learning another host's terminals. The
bus stays a provider-neutral communication fabric; multi-host is an addressing + routing
concern, not a new kind of peer.
The design rests on three pieces (the shape this ticket proposes):
- Dedicated per-agent channels — every agent has its own addressable inbox on the broker.
- A federated agent directory — a global "who/where/status" lookup, assembled from per-host presence, not a central database.
- A per-host gateway — each host runs a
fleetdthat owns its local herdr, registers/manages its own sessions, and proxies messages to/from other hosts over the broker.
2. What is single-host today (the assumptions to break)
flowchart TB
subgraph host["Single host (today)"]
primary["primary<br/>(MCP client)"]
daemon["fleetd daemon<br/>127.0.0.1:8765"]
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
herdr["herdr<br/>(local unix-socket PTY mux)"]
w1["worker pane wQ:p1"]
w2["worker pane wQ:p2"]
primary --> daemon
daemon --> reg
daemon --> herdr
herdr --> w1
herdr --> w2
end
Figure 1 — everything is co-located and loopback.
Three concrete bake-ins assume one host:
| Assumption | Where | Why it blocks multi-host |
|---|---|---|
| herdr is local | herdr/ unix socket ~/.config/herdr/herdr.sock |
You cannot drive another host's PTYs → each host must own its herdr. This is why a per-host gateway is mandatory. |
registry is in-process, keyed by paneId |
session/SessionManager |
paneId (e.g. wQ:p2B) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
| loopback, no authn | rest/FleetApp binds 127.0.0.1:8765 |
Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
3. Target architecture
flowchart TB
subgraph hostA["HOST A"]
gA["gateway = fleetd A"]
regA["local registry + herdr"]
primary["primary (MCP client)"]
gA --- regA
primary --- gA
end
subgraph hostB["HOST B"]
gB["gateway = fleetd B"]
regB["local registry + herdr"]
wb["worker panes"]
gB --- regB
gB --- wb
end
subgraph broker["BROKER (LavinMQ / AMQP) — CB-307 fabric"]
inbox["agent.<id>.inbox queues"]
roster["roster.* presence topic"]
dlq["DLQ · delayed-retry (remind)"]
end
gA -->|"publish to agent.<id>.inbox"| inbox
gB -->|"publish to agent.<id>.inbox"| inbox
inbox -->|"owning gateway consumes"| gA
inbox -->|"owning gateway consumes"| gB
gA -->|"announce local agents"| roster
gB -->|"announce local agents"| roster
roster -->|"union view"| gA
roster -->|"union view"| gB
Figure 2 — each gateway owns its local herdr + registry, consumes only its own agents' inboxes, and announces its agents onto a shared presence topic. The broker routes; no host sees another host's terminals.
3.1 Component mapping (the three pieces)
-
Dedicated per-agent channels = a per-agent AMQP routing key / queue, e.g.
agent.<globalId>.inbox. The agent's owning gateway is the only consumer of its inbox. Senders publish toagent.<id>.inboxand never need to know the agent's host — the broker routes to whichever gateway holds it. LavinMQ additionally gives durability, DLX, and a native delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it. -
Federated agent directory = a soft-state, bridge-owned roster, not a broker-stored database. Per the persistence-boundary decision (fleetd is soft-state; the broker owns message durability, not who/where/status), each gateway announces its local agents
(globalId, host, status, capabilities)on aroster.*presence topic with periodic heartbeats. Every gateway builds an eventually-consistent union view — literally CB-304'srosterView, federated. A stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking). -
Per-host gateway = today's
fleetddaemon, evolved. It already registers/manages sessions and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client that consumes its agents' inboxes and injects into local herdr, and (b) presence announce + union-roster assembly. Evolution, not rewrite.
3.2 Routing rule
flowchart LR
send["fleet_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
lookup -->|"yes"| local["inject via local herdr<br/>(today's Injector path)"]
lookup -->|"no"| pub["publish agent.<id>.inbox<br/>(broker routes to owning gateway)"]
pub --> consume["owning gateway consumes<br/>→ injects into its local herdr"]
Figure 3 — one fork: local agents keep today's in-process inject path; remote agents go over the broker. A sender is oblivious to which branch it took.
4. What CB-307 already provides vs. what is net-new
CB-307 delivers the transport half and is independently valuable on a single host: the AMQP
broker fabric, the fleetd → broker client/adapter, at-least-once + idempotent (dedup-by-id)
delivery, DLQ, and delayed-retry (remind). That is the "proxy cross-host message" backbone;
extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental.
Net-new for CB-308 (multi-host), five items:
- Global agent id — decouple the routing key from
paneId. CB-401'sPeerHandlealready abstracts the routing id; make it host-unique (e.g.<host>/<paneId>or a UUID minted at spawn). The registry and all verbs route on the global id. - Federated directory — presence announce + heartbeat + union roster over
roster.*(§3.1). - Gateway routing — the
local ? inject : publishfork (§3.2), plus each gateway consuming its own agents' inbox queues and injecting into local herdr. - Cross-host spawn —
spawn on host B= publish a control request to B's control channel → gateway B runsClaudeCodeLauncher.spawnlocally (CB-306's readiness gate becomes more valuable here: the far side wants a positive "agent ready" before anyone sends) → announces the new agent into the federated roster. - Trust — the broker connection is now the security boundary. A gateway injects env/tokens at daemon privilege (the CB-401 Stage-C concern), so a remote-triggered spawn/send needs authn/authz: who may act on which host, and which control channels a gateway will honour. Authenticity is resolved — signed messages, §7.1; authorization (who may do what) remains open — §8.
5. The one thing the broker does NOT dissolve
The MCP asymmetry survives the network. The primary is an MCP client to its local gateway; it cannot be called into. A worker on B replying to a primary on A flows:
sequenceDiagram
participant W as worker (host B)
participant GB as gateway B
participant BR as broker
participant GA as gateway A
participant P as primary (host A, MCP client)
W->>GB: fleet_reply
GB->>BR: publish primary-bound (durable, msg id)
BR->>GA: route to A's primary inbox
Note over GA: held durably until the primary pulls
P->>GA: blocking fleet_send resolves / fleet_poll
GA-->>P: reply (then ACK to broker)
Figure 4 — the broker makes the middle hop lossless, ordered, and idempotent; the final hop
into the primary is still a pull (gateway A holds the message until the primary's blocking
fleet_send or fleet_poll). Cross-host neither improves nor worsens this — it just spans hosts.
This is precisely the gap CB-307 closes on one host and CB-308 stretches across hosts.
6. Staging & dependencies
flowchart LR
cb307["CB-307<br/>broker-based reliable delivery<br/>(single host first)"] --> cb308["CB-308<br/>multi-host federation<br/>(this note)"]
cb308 --> a["global agent id"]
cb308 --> b["federated directory"]
cb308 --> c["gateway routing"]
cb308 --> d["cross-host spawn"]
cb308 --> e["cross-host trust model"]
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
class e gate
Figure 5 — CB-307 is the foundation; CB-308's five items build on it. The trust model (e) is the gating concern before any host accepts remote control.
Recommendation: keep CB-307 scoped to single-host broker reliability (foundation, independently
useful), and build CB-308's items on top once the broker fabric exists. Choose CB-307's broker /
channel naming multi-host-ready now (per-agent routing keys, a roster.* topic namespace) so
CB-308 doesn't have to repaint the topology.
7. Resolved design decisions (2026-08-10)
Settled in a design review of this note + wiki chapter 10. The broker-level operational rules (inbox caps, TLS + private broker, schema versioning, trace id, exclusive consumers, U8 broadcast) are recorded in wiki 10 §10; the CB-308-side decisions are below. Entries 1–6 are the first-pass decisions; 7–10 came out of the adversarial second-pass review (same day) and supersede 1–6 where they overlap (notably: the envelope is no longer optional, and dedup is split by path).
-
Sender authenticity — sign every message. Each gateway holds its own signing key and signs what it publishes (sender gid,
msgId, timestamp). The receiving gateway verifies the signature and checks against the roster that the claimed sender lives on the signing gateway's host. This extends the single-host invariant — identity comes from the connection, never an argument — across the broker: cross-host, identity comes from the key. Complements (not replaces) per-gateway broker logins over TLS. -
Profiles are owned by the worker's host.
fleet_spawn(profile, host)resolves the name in the target gateway'sbridged.yaml. Gateways advertise their profile names in presence heartbeats, so a leader sees what each host offers before spawning; an unknown name is a clear error from the target. Secrets (base URLs, tokens) never leave the host that uses them. -
Repo provisioning — clone from the forge, pinned. A cross-host spawn names the repo URL and the exact commit. The target gateway clones from the forge into a local cache (first spawn only), then cuts a per-worker worktree — the CB-301-ext flow with a clone step in front, covered by the same repo-scoped forge token (CB-302). Git stays the only channel code moves through.
-
Asks are live-only, with expiry.
ASK/ANSWER(U2) traverse the broker as short-lived (TTL'd) messages carrying theturn_id, and are never held durably — the single-host rule kept. An answer arriving after its turn ended is not injected; it is dropped and the leader gets aTOO_LATEnotice, so the one failure case is loud rather than weird. Only terminal replies are durable. Walkthrough: wiki 10 §7.4. -
Spawn dedup — a spawn id, remembered on the target. The control queue redelivers like any queue; a replayed
SpawnRequestmust not double-spawn. Requests carry a unique spawn id; the target gateway keeps a short memory of handled ids and answers a redelivery with the existingPeerHandle. CB-117's orphan reap stays as the backstop. -
Broker down — local unaffected, remote fails fast. The routing fork (§3.2) means same-host traffic never touches the broker; that is now a written promise. A send to a remote agent while the broker is unreachable fails immediately with a clear error — the gateway never buffers on the broker's behalf (it stays soft-state, so a crash cannot lose messages it claimed to deliver). Gateways auto-reconnect; remote hosts read as unknown in the roster meanwhile. Broker HA is a later ops choice, not a design requirement.
-
Turn state — split by where the signals are. The worker's gateway owns the turn record (turnId minting, ask coalescing, STALE_TURN, the completion/failure fallbacks, CB-516 abandon): every input to those decisions — pane status, injection, teardown — is local to it. The sender's gateway owns only the waiter. The two are stitched by terminal-outcome envelope kinds (
REPLY/FAILED/ABANDONED) published to the sender's inbox: a worker dying on B fails A's waiter fast because gateway B sees the death synchronously and says so.ABANDONEDis belt-and-braces over the waiter's own timeout and roster expiry, never a replacement — the case where the waiter hangs longest is gateway B itself dying, which is exactly when B can publish nothing. -
Dual ack model + spawn idempotence by construction. Forward path (a brief into a worker): ack before the inject — at-most-once, duplicates structurally impossible; the loss window is closed by an
INJECTEDconfirmation published after the inject lands (noINJECTEDwithin a bound = loud fast failure at the sender, not a silent send-timeout). Reply/pull path keeps ack-after-drain — a duplicate reply is benign, deduped bymsgId. Spawn: the requester mints spawn id = the new worker's gid; the target checks it against the live pane registry, and the gid is stored in the herdr pane itself (label/env, readable back), so a restarted gateway rebuilds gid↔pane from herdr and the check survives restarts with no persisted ledger — this storage point is the load-bearing detail of the no-ledger position. An in-flight reservation set, entered before the launcher call, absorbs a redelivery arriving while the first spawn is still inside CB-306's readiness gate; a crash mid-spawn leaves a half-built pane, which is exactly what CB-117 reaps. -
Publish is enforced, not fire-and-forget. Publisher confirms + the
mandatoryflag + a return listener, on a publish channel separate from the consume/ack channel — synchronous confirms on the single shared channel would hold its lock across a broker round trip and serialize acks fleet-wide. Ordering caveat: a return (unroutable) arrives before the confirm, so "confirmed" ≠ "routed"; the sender checks the returned-set at confirm time.mandatoryis false only forBROADCAST, where an empty group is legal silence. -
Queue lifecycle is session lifecycle.
fleet_stop/reap deletes the worker's inbox queue (itsbroadcast.*bindings die with it — no broadcasts to the dead);x-expirescollects queues orphaned by a crashed gateway (long for main/orchestrator inboxes, short for workers). Queue names carry a version suffix (.v2): AMQP refuses to redeclare an existing durable queue with new arguments (PRECONDITION_FAILED— a crash loop on an in-place upgrade from v1.0.0), and the suffix keeps old sender-keyed and new recipient-keyed queues apart during the keying migration (wiki 10 §3 footnote).
8. Still open
- Directory ground-truth: pure soft-state presence (heartbeats) vs. also treating broker queue existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if split-brain roster views cause mis-routing.
- Gateway discovery: how gateways find the broker and each other (static config vs. discovery).
- Control authorization — THE GATE ON U4. Signing (§7.1) settles who sent it; authorization
is who may do what. Cross-host spawn must not land before the minimal version exists: a
per-host allowlist in
bridged.yaml— beside the peer public keys — of gateway ids permitted to publish control to this host, checked against the verified signature. A few lines of config and check; without them, any principal holding broker credentials can start processes on every host in the fleet. - Key distribution & rotation: static config (host → public key in each
bridged.yaml) is fine at the current 2–3 host scale; rotation is manual. A refinement, not a blocker. - Gateway death mid-turn: the roster reaps it by missed heartbeat, and in-flight primary-bound messages survive by broker durability; still open is reconciling worker state when the dead gateway's host comes back (orphaned panes vs. still-valid sessions).
(Resolved and moved up: the global id scheme — an opaque UUID minted by the spawn requester as the spawn id, host carried as roster metadata; §7.8.)
9. Implementation order (each step verifiable single-host)
- As-built fixes, independent of CB-308 (v1.0.x tickets):
basicQosprefetch on the AMQP consumer (today the queue drains into gateway heap, so any cap would guard an empty queue); publisher confirms +mandatory(§7.9); thedrainRepliesjavadoc that claims "the ack is local" — false for the AMQP adapter. - Envelope + signing (wiki 10 §2.1) — testable against the single-host broker.
- Recipient-keyed queue migration (
.v2names, drain-by-from,ReplyPushLooprekeyed). - Global id + queue lifecycle (§7.8, §7.10).
- Roster: host-level heartbeat + signed presence; then the routing fork (§3.2).
- U2 cross-host with the terminal-outcome kinds (§7.4, §7.7).
- U4 cross-host spawn — gated on the control allowlist (§8).
- U8 broadcast last — it is the feature that punishes an unfinished queue lifecycle.