da5a987df0
- docs/CB-307-Reliable-Delivery.md: ReplyInbox port + in-memory adapter spec (Stage 1, no broker); publish at the Rendezvous no-waiter drop seam, drain by target. - docs/CB-308-Multi-Host-Federation.md: per-host gateway + per-agent broker channels + federated roster proposal (gitea #6), wiki-ready with theme-safe mermaid.
206 lines
10 KiB
Markdown
206 lines
10 KiB
Markdown
# CB-308 — Multi-Host Federation (Stage 5)
|
|
|
|
**Status:** design note (proposal)
|
|
**Depends on:** CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built *on*
|
|
CB-307's broker fabric.
|
|
**Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness),
|
|
CB-303 (lifecycle limits), CB-117 (orphan reap).
|
|
|
|
## 1. Goal
|
|
|
|
Let `claude-bridge` coordinate agents that live on **more than one host** — a primary on host A
|
|
delegating to workers on hosts B, C, … — without any host learning another host's terminals. The
|
|
bus stays a **provider-neutral communication fabric**; multi-host is an addressing + routing
|
|
concern, not a new kind of peer.
|
|
|
|
The design rests on three pieces (the shape this ticket proposes):
|
|
|
|
1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker.
|
|
2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host
|
|
presence, not a central database.
|
|
3. **A per-host gateway** — each host runs a `bridged` that owns its local herdr, registers/manages
|
|
its own sessions, and proxies messages to/from other hosts over the broker.
|
|
|
|
## 2. What is single-host today (the assumptions to break)
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
subgraph host["Single host (today)"]
|
|
primary["primary<br/>(MCP client)"]
|
|
daemon["bridged daemon<br/>127.0.0.1:8765"]
|
|
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
|
|
herdr["herdr<br/>(local unix-socket PTY mux)"]
|
|
w1["worker pane wQ:p1"]
|
|
w2["worker pane wQ:p2"]
|
|
primary --> daemon
|
|
daemon --> reg
|
|
daemon --> herdr
|
|
herdr --> w1
|
|
herdr --> w2
|
|
end
|
|
```
|
|
|
|
*Figure 1 — everything is co-located and loopback.*
|
|
|
|
Three concrete bake-ins assume one host:
|
|
|
|
| Assumption | Where | Why it blocks multi-host |
|
|
|---|---|---|
|
|
| **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. |
|
|
| **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
|
|
| **loopback, no authn** | `rest/BridgedApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
|
|
|
|
## 3. Target architecture
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
subgraph hostA["HOST A"]
|
|
gA["gateway = bridged A"]
|
|
regA["local registry + herdr"]
|
|
primary["primary (MCP client)"]
|
|
gA --- regA
|
|
primary --- gA
|
|
end
|
|
subgraph hostB["HOST B"]
|
|
gB["gateway = bridged B"]
|
|
regB["local registry + herdr"]
|
|
wb["worker panes"]
|
|
gB --- regB
|
|
gB --- wb
|
|
end
|
|
subgraph broker["BROKER (LavinMQ / AMQP) — CB-307 fabric"]
|
|
inbox["agent.<id>.inbox queues"]
|
|
roster["roster.* presence topic"]
|
|
dlq["DLQ · delayed-retry (remind)"]
|
|
end
|
|
gA -->|"publish to agent.<id>.inbox"| inbox
|
|
gB -->|"publish to agent.<id>.inbox"| inbox
|
|
inbox -->|"owning gateway consumes"| gA
|
|
inbox -->|"owning gateway consumes"| gB
|
|
gA -->|"announce local agents"| roster
|
|
gB -->|"announce local agents"| roster
|
|
roster -->|"union view"| gA
|
|
roster -->|"union view"| gB
|
|
```
|
|
|
|
*Figure 2 — each gateway owns its local herdr + registry, consumes only its own agents' inboxes,
|
|
and announces its agents onto a shared presence topic. The broker routes; no host sees another
|
|
host's terminals.*
|
|
|
|
### 3.1 Component mapping (the three pieces)
|
|
|
|
- **Dedicated per-agent channels** = a per-agent AMQP routing key / queue, e.g.
|
|
`agent.<globalId>.inbox`. The agent's **owning gateway is the only consumer** of its inbox.
|
|
Senders publish to `agent.<id>.inbox` and never need to know the agent's host — the broker
|
|
routes to whichever gateway holds it. LavinMQ additionally gives durability, DLX, and a native
|
|
delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it.
|
|
|
|
- **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored
|
|
database. Per the persistence-boundary decision (bridged is soft-state; the broker owns *message*
|
|
durability, not *who/where/status*), each gateway announces its local agents `(globalId, host,
|
|
status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway
|
|
builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A
|
|
stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking).
|
|
|
|
- **Per-host gateway** = today's `bridged` daemon, evolved. It already registers/manages sessions
|
|
and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client
|
|
that consumes its agents' inboxes and injects into local herdr, and (b) presence announce +
|
|
union-roster assembly. Evolution, not rewrite.
|
|
|
|
### 3.2 Routing rule
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
send["bridge_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
|
|
lookup -->|"yes"| local["inject via local herdr<br/>(today's Injector path)"]
|
|
lookup -->|"no"| pub["publish agent.<id>.inbox<br/>(broker routes to owning gateway)"]
|
|
pub --> consume["owning gateway consumes<br/>→ injects into its local herdr"]
|
|
```
|
|
|
|
*Figure 3 — one fork: local agents keep today's in-process inject path; remote agents go over the
|
|
broker. A sender is oblivious to which branch it took.*
|
|
|
|
## 4. What CB-307 already provides vs. what is net-new
|
|
|
|
**CB-307 delivers the transport half** and is independently valuable on a single host: the AMQP
|
|
broker fabric, the `bridged → broker` client/adapter, at-least-once + idempotent (dedup-by-id)
|
|
delivery, DLQ, and delayed-retry (remind). That *is* the "proxy cross-host message" backbone;
|
|
extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental.
|
|
|
|
**Net-new for CB-308 (multi-host), five items:**
|
|
|
|
1. **Global agent id** — decouple the routing key from `paneId`. CB-401's `PeerHandle` already
|
|
abstracts the routing id; make it host-unique (e.g. `<host>/<paneId>` or a UUID minted at spawn).
|
|
The registry and all verbs route on the global id.
|
|
2. **Federated directory** — presence announce + heartbeat + union roster over `roster.*`
|
|
(§3.1).
|
|
3. **Gateway routing** — the `local ? inject : publish` fork (§3.2), plus each gateway consuming
|
|
its own agents' inbox queues and injecting into local herdr.
|
|
4. **Cross-host spawn** — `spawn on host B` = publish a control request to B's control channel →
|
|
gateway B runs `ClaudeCodeLauncher.spawn` **locally** (CB-306's readiness gate becomes *more*
|
|
valuable here: the far side wants a positive "agent ready" before anyone sends) → announces the
|
|
new agent into the federated roster.
|
|
5. **Trust** — the broker connection is now the security boundary. A gateway injects env/tokens at
|
|
daemon privilege (the CB-401 Stage-C concern), so a **remote-triggered spawn/send** needs
|
|
authn/authz: who may act on which host, and which control channels a gateway will honour.
|
|
|
|
## 5. The one thing the broker does NOT dissolve
|
|
|
|
The MCP asymmetry survives the network. The primary is an MCP **client** to its **local** gateway;
|
|
it cannot be called into. A worker on B replying to a primary on A flows:
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
participant W as worker (host B)
|
|
participant GB as gateway B
|
|
participant BR as broker
|
|
participant GA as gateway A
|
|
participant P as primary (host A, MCP client)
|
|
W->>GB: bridge_reply
|
|
GB->>BR: publish primary-bound (durable, msg id)
|
|
BR->>GA: route to A's primary inbox
|
|
Note over GA: held durably until the primary pulls
|
|
P->>GA: blocking bridge_send resolves / bridge_poll
|
|
GA-->>P: reply (then ACK to broker)
|
|
```
|
|
|
|
*Figure 4 — the broker makes the middle hop lossless, ordered, and idempotent; the **final** hop
|
|
into the primary is still a **pull** (gateway A holds the message until the primary's blocking
|
|
`bridge_send` or `bridge_poll`). Cross-host neither improves nor worsens this — it just spans hosts.
|
|
This is precisely the gap CB-307 closes on one host and CB-308 stretches across hosts.*
|
|
|
|
## 6. Staging & dependencies
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
cb307["CB-307<br/>broker-based reliable delivery<br/>(single host first)"] --> cb308["CB-308<br/>multi-host federation<br/>(this note)"]
|
|
cb308 --> a["global agent id"]
|
|
cb308 --> b["federated directory"]
|
|
cb308 --> c["gateway routing"]
|
|
cb308 --> d["cross-host spawn"]
|
|
cb308 --> e["cross-host trust model"]
|
|
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
|
class e gate
|
|
```
|
|
|
|
*Figure 5 — CB-307 is the foundation; CB-308's five items build on it. The trust model (e) is the
|
|
gating concern before any host accepts remote control.*
|
|
|
|
**Recommendation:** keep CB-307 scoped to single-host broker reliability (foundation, independently
|
|
useful), and build CB-308's items on top once the broker fabric exists. Choose CB-307's broker /
|
|
channel naming **multi-host-ready** now (per-agent routing keys, a `roster.*` topic namespace) so
|
|
CB-308 doesn't have to repaint the topology.
|
|
|
|
## 7. Open questions
|
|
|
|
- **Directory ground-truth:** pure soft-state presence (heartbeats) vs. also treating broker queue
|
|
existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if
|
|
split-brain roster views cause mis-routing.
|
|
- **Global id scheme:** `<host>/<paneId>` (human-legible, leaks host) vs. opaque UUID (clean, needs
|
|
the directory to resolve host). Probably UUID in the protocol, host as directory metadata.
|
|
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
|
|
- **Trust model shape:** per-host shared secret vs. mTLS on the broker vs. a capability token per
|
|
control action — ties into CB-401 Stage-C.
|
|
- **Failure semantics:** a host/gateway dies mid-turn — how the federated roster reaps it (missed
|
|
heartbeat) and whether in-flight primary-bound messages survive (broker durability = yes).
|