2e138a199b
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
304 lines
18 KiB
Markdown
304 lines
18 KiB
Markdown
# CB-308 — Multi-Host Federation (Stage 5)
|
||
|
||
**Status:** design note (proposal) — core design decisions resolved 2026-08-10 (§7)
|
||
**Depends on:** CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built *on*
|
||
CB-307's broker fabric.
|
||
**Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness),
|
||
CB-303 (lifecycle limits), CB-117 (orphan reap).
|
||
|
||
## 1. Goal
|
||
|
||
Let `claude-bridge` coordinate agents that live on **more than one host** — a primary on host A
|
||
delegating to workers on hosts B, C, … — without any host learning another host's terminals. The
|
||
bus stays a **provider-neutral communication fabric**; multi-host is an addressing + routing
|
||
concern, not a new kind of peer.
|
||
|
||
The design rests on three pieces (the shape this ticket proposes):
|
||
|
||
1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker.
|
||
2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host
|
||
presence, not a central database.
|
||
3. **A per-host gateway** — each host runs a `fleetd` that owns its local herdr, registers/manages
|
||
its own sessions, and proxies messages to/from other hosts over the broker.
|
||
|
||
## 2. What is single-host today (the assumptions to break)
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
subgraph host["Single host (today)"]
|
||
primary["primary<br/>(MCP client)"]
|
||
daemon["fleetd daemon<br/>127.0.0.1:8765"]
|
||
reg["in-process registry<br/>keyed by PeerHandle.id() == paneId"]
|
||
herdr["herdr<br/>(local unix-socket PTY mux)"]
|
||
w1["worker pane wQ:p1"]
|
||
w2["worker pane wQ:p2"]
|
||
primary --> daemon
|
||
daemon --> reg
|
||
daemon --> herdr
|
||
herdr --> w1
|
||
herdr --> w2
|
||
end
|
||
```
|
||
|
||
*Figure 1 — everything is co-located and loopback.*
|
||
|
||
Three concrete bake-ins assume one host:
|
||
|
||
| Assumption | Where | Why it blocks multi-host |
|
||
|---|---|---|
|
||
| **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. |
|
||
| **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. |
|
||
| **loopback, no authn** | `rest/FleetApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. |
|
||
|
||
## 3. Target architecture
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
subgraph hostA["HOST A"]
|
||
gA["gateway = fleetd A"]
|
||
regA["local registry + herdr"]
|
||
primary["primary (MCP client)"]
|
||
gA --- regA
|
||
primary --- gA
|
||
end
|
||
subgraph hostB["HOST B"]
|
||
gB["gateway = fleetd B"]
|
||
regB["local registry + herdr"]
|
||
wb["worker panes"]
|
||
gB --- regB
|
||
gB --- wb
|
||
end
|
||
subgraph broker["BROKER (LavinMQ / AMQP) — CB-307 fabric"]
|
||
inbox["agent.<id>.inbox queues"]
|
||
roster["roster.* presence topic"]
|
||
dlq["DLQ · delayed-retry (remind)"]
|
||
end
|
||
gA -->|"publish to agent.<id>.inbox"| inbox
|
||
gB -->|"publish to agent.<id>.inbox"| inbox
|
||
inbox -->|"owning gateway consumes"| gA
|
||
inbox -->|"owning gateway consumes"| gB
|
||
gA -->|"announce local agents"| roster
|
||
gB -->|"announce local agents"| roster
|
||
roster -->|"union view"| gA
|
||
roster -->|"union view"| gB
|
||
```
|
||
|
||
*Figure 2 — each gateway owns its local herdr + registry, consumes only its own agents' inboxes,
|
||
and announces its agents onto a shared presence topic. The broker routes; no host sees another
|
||
host's terminals.*
|
||
|
||
### 3.1 Component mapping (the three pieces)
|
||
|
||
- **Dedicated per-agent channels** = a per-agent AMQP routing key / queue, e.g.
|
||
`agent.<globalId>.inbox`. The agent's **owning gateway is the only consumer** of its inbox.
|
||
Senders publish to `agent.<id>.inbox` and never need to know the agent's host — the broker
|
||
routes to whichever gateway holds it. LavinMQ additionally gives durability, DLX, and a native
|
||
delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it.
|
||
|
||
- **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored
|
||
database. Per the persistence-boundary decision (fleetd is soft-state; the broker owns *message*
|
||
durability, not *who/where/status*), each gateway announces its local agents `(globalId, host,
|
||
status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway
|
||
builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A
|
||
stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking).
|
||
|
||
- **Per-host gateway** = today's `fleetd` daemon, evolved. It already registers/manages sessions
|
||
and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client
|
||
that consumes its agents' inboxes and injects into local herdr, and (b) presence announce +
|
||
union-roster assembly. Evolution, not rewrite.
|
||
|
||
### 3.2 Routing rule
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
send["fleet_send(globalId, msg)"] --> lookup{"directory:<br/>is globalId local?"}
|
||
lookup -->|"yes"| local["inject via local herdr<br/>(today's Injector path)"]
|
||
lookup -->|"no"| pub["publish agent.<id>.inbox<br/>(broker routes to owning gateway)"]
|
||
pub --> consume["owning gateway consumes<br/>→ injects into its local herdr"]
|
||
```
|
||
|
||
*Figure 3 — one fork: local agents keep today's in-process inject path; remote agents go over the
|
||
broker. A sender is oblivious to which branch it took.*
|
||
|
||
## 4. What CB-307 already provides vs. what is net-new
|
||
|
||
**CB-307 delivers the transport half** and is independently valuable on a single host: the AMQP
|
||
broker fabric, the `fleetd → broker` client/adapter, at-least-once + idempotent (dedup-by-id)
|
||
delivery, DLQ, and delayed-retry (remind). That *is* the "proxy cross-host message" backbone;
|
||
extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental.
|
||
|
||
**Net-new for CB-308 (multi-host), five items:**
|
||
|
||
1. **Global agent id** — decouple the routing key from `paneId`. CB-401's `PeerHandle` already
|
||
abstracts the routing id; make it host-unique (e.g. `<host>/<paneId>` or a UUID minted at spawn).
|
||
The registry and all verbs route on the global id.
|
||
2. **Federated directory** — presence announce + heartbeat + union roster over `roster.*`
|
||
(§3.1).
|
||
3. **Gateway routing** — the `local ? inject : publish` fork (§3.2), plus each gateway consuming
|
||
its own agents' inbox queues and injecting into local herdr.
|
||
4. **Cross-host spawn** — `spawn on host B` = publish a control request to B's control channel →
|
||
gateway B runs `ClaudeCodeLauncher.spawn` **locally** (CB-306's readiness gate becomes *more*
|
||
valuable here: the far side wants a positive "agent ready" before anyone sends) → announces the
|
||
new agent into the federated roster.
|
||
5. **Trust** — the broker connection is now the security boundary. A gateway injects env/tokens at
|
||
daemon privilege (the CB-401 Stage-C concern), so a **remote-triggered spawn/send** needs
|
||
authn/authz: who may act on which host, and which control channels a gateway will honour.
|
||
*Authenticity* is resolved — signed messages, §7.1; *authorization* (who may do what) remains
|
||
open — §8.
|
||
|
||
## 5. The one thing the broker does NOT dissolve
|
||
|
||
The MCP asymmetry survives the network. The primary is an MCP **client** to its **local** gateway;
|
||
it cannot be called into. A worker on B replying to a primary on A flows:
|
||
|
||
```mermaid
|
||
sequenceDiagram
|
||
participant W as worker (host B)
|
||
participant GB as gateway B
|
||
participant BR as broker
|
||
participant GA as gateway A
|
||
participant P as primary (host A, MCP client)
|
||
W->>GB: fleet_reply
|
||
GB->>BR: publish primary-bound (durable, msg id)
|
||
BR->>GA: route to A's primary inbox
|
||
Note over GA: held durably until the primary pulls
|
||
P->>GA: blocking fleet_send resolves / fleet_poll
|
||
GA-->>P: reply (then ACK to broker)
|
||
```
|
||
|
||
*Figure 4 — the broker makes the middle hop lossless, ordered, and idempotent; the **final** hop
|
||
into the primary is still a **pull** (gateway A holds the message until the primary's blocking
|
||
`fleet_send` or `fleet_poll`). Cross-host neither improves nor worsens this — it just spans hosts.
|
||
This is precisely the gap CB-307 closes on one host and CB-308 stretches across hosts.*
|
||
|
||
## 6. Staging & dependencies
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
cb307["CB-307<br/>broker-based reliable delivery<br/>(single host first)"] --> cb308["CB-308<br/>multi-host federation<br/>(this note)"]
|
||
cb308 --> a["global agent id"]
|
||
cb308 --> b["federated directory"]
|
||
cb308 --> c["gateway routing"]
|
||
cb308 --> d["cross-host spawn"]
|
||
cb308 --> e["cross-host trust model"]
|
||
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
|
||
class e gate
|
||
```
|
||
|
||
*Figure 5 — CB-307 is the foundation; CB-308's five items build on it. The trust model (e) is the
|
||
gating concern before any host accepts remote control.*
|
||
|
||
**Recommendation:** keep CB-307 scoped to single-host broker reliability (foundation, independently
|
||
useful), and build CB-308's items on top once the broker fabric exists. Choose CB-307's broker /
|
||
channel naming **multi-host-ready** now (per-agent routing keys, a `roster.*` topic namespace) so
|
||
CB-308 doesn't have to repaint the topology.
|
||
|
||
## 7. Resolved design decisions (2026-08-10)
|
||
|
||
Settled in a design review of this note + wiki chapter 10. The broker-level operational rules
|
||
(inbox caps, TLS + private broker, schema versioning, trace id, exclusive consumers, U8 broadcast)
|
||
are recorded in wiki 10 §10; the CB-308-side decisions are below. Entries 1–6 are the first-pass
|
||
decisions; 7–10 came out of the adversarial second-pass review (same day) and supersede 1–6 where
|
||
they overlap (notably: the envelope is no longer optional, and dedup is split by path).
|
||
|
||
1. **Sender authenticity — sign every message.** Each gateway holds its own signing key and signs
|
||
what it publishes (sender gid, `msgId`, timestamp). The receiving gateway verifies the
|
||
signature **and** checks against the roster that the claimed sender lives on the signing
|
||
gateway's host. This extends the single-host invariant — *identity comes from the connection,
|
||
never an argument* — across the broker: cross-host, identity comes from the key. Complements
|
||
(not replaces) per-gateway broker logins over TLS.
|
||
2. **Profiles are owned by the worker's host.** `fleet_spawn(profile, host)` resolves the name in
|
||
the *target* gateway's `fleetd.yaml`. Gateways advertise their profile names in presence
|
||
heartbeats, so a leader sees what each host offers before spawning; an unknown name is a clear
|
||
error from the target. Secrets (base URLs, tokens) never leave the host that uses them.
|
||
3. **Repo provisioning — clone from the forge, pinned.** A cross-host spawn names the repo URL and
|
||
the exact commit. The target gateway clones from the forge into a local cache (first spawn
|
||
only), then cuts a per-worker worktree — the CB-301-ext flow with a clone step in front,
|
||
covered by the same repo-scoped forge token (CB-302). Git stays the only channel code moves
|
||
through.
|
||
4. **Asks are live-only, with expiry.** `ASK`/`ANSWER` (U2) traverse the broker as short-lived
|
||
(TTL'd) messages carrying the `turn_id`, and are never held durably — the single-host rule
|
||
kept. An answer arriving after its turn ended is **not** injected; it is dropped and the leader
|
||
gets a `TOO_LATE` notice, so the one failure case is loud rather than weird. Only terminal
|
||
replies are durable. Walkthrough: wiki 10 §7.4.
|
||
5. **Spawn dedup — a spawn id, remembered on the target.** The control queue redelivers like any
|
||
queue; a replayed `SpawnRequest` must not double-spawn. Requests carry a unique spawn id; the
|
||
target gateway keeps a short memory of handled ids and answers a redelivery with the existing
|
||
`PeerHandle`. CB-117's orphan reap stays as the backstop.
|
||
6. **Broker down — local unaffected, remote fails fast.** The routing fork (§3.2) means same-host
|
||
traffic never touches the broker; that is now a written promise. A send to a remote agent while
|
||
the broker is unreachable **fails immediately** with a clear error — the gateway never buffers
|
||
on the broker's behalf (it stays soft-state, so a crash cannot lose messages it claimed to
|
||
deliver). Gateways auto-reconnect; remote hosts read as unknown in the roster meanwhile. Broker
|
||
HA is a later ops choice, not a design requirement.
|
||
|
||
7. **Turn state — split by where the signals are.** The *worker's* gateway owns the turn record
|
||
(turnId minting, ask coalescing, STALE_TURN, the completion/failure fallbacks, CB-516 abandon):
|
||
every input to those decisions — pane status, injection, teardown — is local to it. The
|
||
*sender's* gateway owns only the waiter. The two are stitched by terminal-outcome envelope
|
||
kinds (`REPLY` / `FAILED` / `ABANDONED`) published to the sender's inbox: a worker dying on B
|
||
fails A's waiter fast because gateway B sees the death synchronously and says so.
|
||
**`ABANDONED` is belt-and-braces over the waiter's own timeout and roster expiry, never a
|
||
replacement** — the case where the waiter hangs longest is gateway B itself dying, which is
|
||
exactly when B can publish nothing.
|
||
8. **Dual ack model + spawn idempotence by construction.** Forward path (a brief into a worker):
|
||
ack **before** the inject — at-most-once, duplicates structurally impossible; the loss window
|
||
is closed by an `INJECTED` confirmation published after the inject lands (no `INJECTED` within
|
||
a bound = loud fast failure at the sender, not a silent send-timeout). Reply/pull path keeps
|
||
ack-after-drain — a duplicate reply is benign, deduped by `msgId`. Spawn: the requester mints
|
||
**spawn id = the new worker's gid**; the target checks it against the **live pane registry**,
|
||
and the gid is **stored in the herdr pane itself** (label/env, readable back), so a restarted
|
||
gateway rebuilds gid↔pane from herdr and the check survives restarts with *no persisted
|
||
ledger* — this storage point is the load-bearing detail of the no-ledger position. An
|
||
**in-flight reservation set**, entered before the launcher call, absorbs a redelivery arriving
|
||
while the first spawn is still inside CB-306's readiness gate; a crash mid-spawn leaves a
|
||
half-built pane, which is exactly what CB-117 reaps.
|
||
9. **Publish is enforced, not fire-and-forget.** Publisher confirms + the `mandatory` flag + a
|
||
return listener, on a **publish channel separate from the consume/ack channel** — synchronous
|
||
confirms on the single shared channel would hold its lock across a broker round trip and
|
||
serialize acks fleet-wide. Ordering caveat: a *return* (unroutable) arrives **before** the
|
||
confirm, so "confirmed" ≠ "routed"; the sender checks the returned-set at confirm time.
|
||
`mandatory` is false only for `BROADCAST`, where an empty group is legal silence.
|
||
10. **Queue lifecycle is session lifecycle.** `fleet_stop`/reap deletes the worker's inbox queue
|
||
(its `broadcast.*` bindings die with it — no broadcasts to the dead); `x-expires` collects
|
||
queues orphaned by a crashed gateway (long for main/orchestrator inboxes, short for workers).
|
||
Queue names carry a version suffix (`.v2`): AMQP refuses to redeclare an existing durable
|
||
queue with new arguments (`PRECONDITION_FAILED` — a crash loop on an in-place upgrade from
|
||
v1.0.0), and the suffix keeps old sender-keyed and new recipient-keyed queues apart during
|
||
the keying migration (wiki 10 §3 footnote).
|
||
|
||
## 8. Still open
|
||
|
||
- **Directory ground-truth:** pure soft-state presence (heartbeats) vs. also treating broker queue
|
||
existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if
|
||
split-brain roster views cause mis-routing.
|
||
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
|
||
- **Control authorization — THE GATE ON U4.** Signing (§7.1) settles *who sent it*; authorization
|
||
is *who may do what*. **Cross-host spawn must not land before the minimal version exists**: a
|
||
per-host allowlist in `fleetd.yaml` — beside the peer public keys — of gateway ids permitted to
|
||
publish control to this host, checked against the verified signature. A few lines of config and
|
||
check; without them, any principal holding broker credentials can start processes on every host
|
||
in the fleet.
|
||
- **Key distribution & rotation:** static config (host → public key in each `fleetd.yaml`) is
|
||
fine at the current 2–3 host scale; rotation is manual. A refinement, not a blocker.
|
||
- **Gateway death mid-turn:** the roster reaps it by missed heartbeat, and in-flight primary-bound
|
||
messages survive by broker durability; still open is reconciling *worker* state when the dead
|
||
gateway's host comes back (orphaned panes vs. still-valid sessions).
|
||
|
||
*(Resolved and moved up: the global id scheme — an opaque UUID minted by the spawn requester as
|
||
the spawn id, host carried as roster metadata; §7.8.)*
|
||
|
||
## 9. Implementation order (each step verifiable single-host)
|
||
|
||
1. **As-built fixes, independent of CB-308** (v1.0.x tickets): `basicQos` prefetch on the AMQP
|
||
consumer (today the queue drains into gateway heap, so any cap would guard an empty queue);
|
||
publisher confirms + `mandatory` (§7.9); the `drainReplies` javadoc that claims "the ack is
|
||
local" — false for the AMQP adapter.
|
||
2. Envelope + signing (wiki 10 §2.1) — testable against the single-host broker.
|
||
3. Recipient-keyed queue migration (`.v2` names, drain-by-`from`, `ReplyPushLoop` rekeyed).
|
||
4. Global id + queue lifecycle (§7.8, §7.10).
|
||
5. Roster: host-level heartbeat + signed presence; then the routing fork (§3.2).
|
||
6. U2 cross-host with the terminal-outcome kinds (§7.4, §7.7).
|
||
7. U4 cross-host spawn — **gated on the control allowlist (§8)**.
|
||
8. U8 broadcast **last** — it is the feature that punishes an unfinished queue lifecycle.
|