CB-308: multi-host federation — per-host gateway + per-agent broker channels + federated roster (Stage 5) #6

Open
opened 2026-07-18 20:20:46 +02:00 by ltms · 1 comment
Owner

Goal

Let claude-bridge coordinate agents across more than one host (a primary on host A delegating to workers on hosts B, C, …) without any host learning another host's terminals. Multi-host is an addressing + routing concern layered on CB-307's broker fabric — not a new kind of peer. The bus stays a provider-neutral communication fabric.

Full design note (wiki-ready, with diagrams): docs/CB-308-Multi-Host-Federation.md.

The owner's proposed shape (three pieces)

  1. Dedicated per-agent channels — every agent has its own addressable broker inbox (agent.<globalId>.inbox); the agent's owning gateway is the sole consumer. Senders publish to the id; the broker routes to whichever host holds it.
  2. Federated agent directory — a soft-state, bridge-owned "who/where/status" lookup assembled from per-host presence heartbeats on a roster.* topic (the union of per-gateway rosters = CB-304 rosterView, federated). Not a broker-stored DB — preserves the persistence boundary.
  3. Per-host gateway = evolved bridged — each host runs a daemon that owns its local herdr + local registry, consumes its agents' inboxes and injects locally, and announces presence. Evolution, not rewrite.

Single-host bake-ins to break

Assumption Where Consequence
herdr is local ~/.config/herdr/herdr.sock can't drive remote PTYs → per-host gateway mandatory
registry in-process, keyed by paneId session/SessionManager paneId is host-local → need a host-unique global id
loopback, no authn rest/BridgedApp 127.0.0.1:8765 a second host talking to a gateway is a trust boundary

Net-new work (five items, all on top of CB-307)

  1. Global agent id — decouple routing key from paneId; CB-401's PeerHandle already abstracts it (make host-unique / UUID).
  2. Federated directory — presence announce + heartbeat + union roster over roster.* (reap by missed heartbeat, reuses CB-303 TTL thinking).
  3. Gateway routing — the local ? inject : publish agent.<id>.inbox fork; each gateway consumes its own agents' inboxes.
  4. Cross-host spawn — publish a control request to the far gateway → it runs ClaudeCodeLauncher.spawn locally (CB-306 readiness gate is more valuable remotely) → announces into the federated roster.
  5. Cross-host trust/authz — the broker connection is the security boundary; remote-triggered spawn injects env/tokens at daemon privilege (ties into CB-401 Stage-C).

What does NOT change

The bridged → primary last hop is still the MCP asymmetry (primary is an MCP client; nothing can call into it). The broker makes cross-host delivery lossless + idempotent, but gateway A still holds a primary-bound message until the local primary pulls. CB-308 stretches CB-307's single-host guarantee across hosts; it does not dissolve the pull.

Dependency & staging

Depends on CB-307 (#5) — the broker fabric is the foundation. Recommendation: pick CB-307's channel naming multi-host-ready now (per-agent routing keys, roster.* namespace) so this ticket doesn't repaint the topology.

Relates to: CB-401 (PeerHandle opaque id), CB-304 (rosterView), CB-306 (spawn-readiness), CB-303 (lifecycle TTL).
Layer: session + peer + new gateway/routing + msg broker fabric. Type: architecture / scale. Stage: 5 (multi-host).

## Goal Let `claude-bridge` coordinate agents across **more than one host** (a primary on host A delegating to workers on hosts B, C, …) without any host learning another host's terminals. Multi-host is an **addressing + routing** concern layered on CB-307's broker fabric — not a new kind of peer. The bus stays a provider-neutral communication fabric. Full design note (wiki-ready, with diagrams): **`docs/CB-308-Multi-Host-Federation.md`**. ## The owner's proposed shape (three pieces) 1. **Dedicated per-agent channels** — every agent has its own addressable broker inbox (`agent.<globalId>.inbox`); the agent's owning gateway is the sole consumer. Senders publish to the id; the broker routes to whichever host holds it. 2. **Federated agent directory** — a soft-state, bridge-owned "who/where/status" lookup assembled from per-host presence heartbeats on a `roster.*` topic (the union of per-gateway rosters = CB-304 `rosterView`, federated). Not a broker-stored DB — preserves the persistence boundary. 3. **Per-host gateway = evolved `bridged`** — each host runs a daemon that owns its local herdr + local registry, consumes its agents' inboxes and injects locally, and announces presence. Evolution, not rewrite. ## Single-host bake-ins to break | Assumption | Where | Consequence | |---|---|---| | herdr is local | `~/.config/herdr/herdr.sock` | can't drive remote PTYs → **per-host gateway mandatory** | | registry in-process, keyed by `paneId` | `session/SessionManager` | `paneId` is host-local → need a **host-unique global id** | | loopback, no authn | `rest/BridgedApp` `127.0.0.1:8765` | a second host talking to a gateway is a **trust boundary** | ## Net-new work (five items, all on top of CB-307) 1. **Global agent id** — decouple routing key from `paneId`; CB-401's `PeerHandle` already abstracts it (make host-unique / UUID). 2. **Federated directory** — presence announce + heartbeat + union roster over `roster.*` (reap by missed heartbeat, reuses CB-303 TTL thinking). 3. **Gateway routing** — the `local ? inject : publish agent.<id>.inbox` fork; each gateway consumes its own agents' inboxes. 4. **Cross-host spawn** — publish a control request to the far gateway → it runs `ClaudeCodeLauncher.spawn` locally (CB-306 readiness gate is *more* valuable remotely) → announces into the federated roster. 5. **Cross-host trust/authz** — the broker connection is the security boundary; remote-triggered spawn injects env/tokens at daemon privilege (ties into CB-401 Stage-C). ## What does NOT change The `bridged → primary` last hop is still the MCP asymmetry (primary is an MCP client; nothing can call into it). The broker makes cross-host delivery lossless + idempotent, but gateway A still **holds** a primary-bound message until the local primary pulls. CB-308 stretches CB-307's single-host guarantee across hosts; it does not dissolve the pull. ## Dependency & staging **Depends on CB-307** (#5) — the broker fabric is the foundation. Recommendation: pick CB-307's channel naming multi-host-ready now (per-agent routing keys, `roster.*` namespace) so this ticket doesn't repaint the topology. **Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness), CB-303 (lifecycle TTL). **Layer:** `session` + `peer` + new gateway/routing + `msg` broker fabric. **Type:** architecture / scale. **Stage:** 5 (multi-host).
ltms added this to the 2.0 — one operation centre, many hosts milestone 2026-08-16 17:36:21 +02:00
Author
Owner

Status update from 2026-08-23. Two of this ticket's assumptions can now be replaced with measurements, and one of them turns a design preference into a constraint.

The dependency is satisfied: the broker fabric is live and shared

CB-307 was the stated blocker. As of today a single LavinMQ 2.9.1 instance serves both fleets, running on fleet01, with one vhost per fleet (/mac, /fleet01) and one user per fleet. Isolation is enforced, not agreed: the Mac's credential passes the AMQP contract suite 8/8 against its own vhost and fails 8/8 against the other one.

Two changes landed today to make that survivable, both of which this ticket needs:

  • #151 — broker.uriEnv keeps the AMQP password out of every host's config file.
  • #152 — an unreachable broker at boot no longer stops the daemon. Before this, one broker outage would have taken every federated host down at once. That was acceptable when each fleet had its own loopback broker; it is not acceptable for a shared fabric.

The routing choice is now forced, not preferred

This ticket proposes per-agent broker channels over host-to-host HTTP. I measured whether the HTTP option even exists:

from → to result
fleet01 → Mac 10.10.3.14:8765 unreachable
fleet01 → Mac tailscale 100.104.246.88:8765 unreachable
Mac → fleet01 10.10.20.13:8765 unreachable

Both daemons bind 127.0.0.1 only, and on top of that fleet01 cannot reach the Mac at all at the network level, in either direction of the connection attempt. The Mac reaches fleet01 over ssh; nothing goes the other way.

So today the only bidirectional channel between the two hosts is the broker. Both hosts can reach 10.10.20.13:5672; neither can reach the other's REST or MCP port. The "single-host bake-ins to break" table lists loopback-plus-no-authn as a trust boundary to solve — in practice there is no host-to-host HTTP path to secure, because there is none to begin with.

That is worth stating plainly in the design: the broker is not the better transport here, it is the available one. A per-host gateway that expects to be dialled by a peer cannot be built on this network without new firewall and bind-address work that nobody has asked for.

The as-built input

I wrote up the whole fleet01 stand-up in #156 — measured, not recalled. It is meant as the input for this ticket. The parts that bear directly on the five net-new work items here:

  • Item 3 (gateway routing) — see above; the fork has only one viable branch today.
  • Item 2 (federated directory) — LeadLauncher.ensureLeads() runs once, inline in main(), with no supervision loop. A dead lead is never restored and nothing logs it. Presence heartbeats across hosts will be built on top of a local layer that cannot currently notice a local death. Related: #150 (leads always report ready: false).
  • Item 5 (cross-host trust) — fleet01 holds only four credentials and keeps them in ~/.zprofile, which member panes never read because a herdr pane on Linux is a non-login shell. That is a better posture than the Mac's and is worth preserving as hosts are added. But #157 shows a token leaking through a git remote URL, which no env-var control can catch.

The prerequisite this ticket does not name

Nothing supervises fleetd or herdr on fleet01 — no systemd unit, no user unit, no cron. I checked. A reboot or a crash ends that fleet silently, and deploy/bridged.service is not installed and carries the login-shell secret gap tracked in #103.

Federating hosts that nothing keeps alive multiplies the failure rather than containing it. I would treat #103 as a hard dependency of this ticket, not as unrelated ops work.

Status update from 2026-08-23. Two of this ticket's assumptions can now be replaced with measurements, and one of them turns a design preference into a constraint. ## The dependency is satisfied: the broker fabric is live and shared CB-307 was the stated blocker. As of today a **single LavinMQ 2.9.1 instance serves both fleets**, running on `fleet01`, with one vhost per fleet (`/mac`, `/fleet01`) and one user per fleet. Isolation is enforced, not agreed: the Mac's credential passes the AMQP contract suite 8/8 against its own vhost and fails 8/8 against the other one. Two changes landed today to make that survivable, both of which this ticket needs: - **#151** — `broker.uriEnv` keeps the AMQP password out of every host's config file. - **#152** — an unreachable broker at boot no longer stops the daemon. Before this, one broker outage would have taken **every** federated host down at once. That was acceptable when each fleet had its own loopback broker; it is not acceptable for a shared fabric. ## The routing choice is now forced, not preferred This ticket proposes per-agent broker channels over host-to-host HTTP. I measured whether the HTTP option even exists: | from → to | result | |---|---| | fleet01 → Mac `10.10.3.14:8765` | unreachable | | fleet01 → Mac tailscale `100.104.246.88:8765` | unreachable | | Mac → fleet01 `10.10.20.13:8765` | unreachable | Both daemons bind `127.0.0.1` only, and on top of that **fleet01 cannot reach the Mac at all** at the network level, in either direction of the connection attempt. The Mac reaches fleet01 over ssh; nothing goes the other way. So today the **only bidirectional channel between the two hosts is the broker**. Both hosts can reach `10.10.20.13:5672`; neither can reach the other's REST or MCP port. The "single-host bake-ins to break" table lists loopback-plus-no-authn as a trust boundary to solve — in practice there is no host-to-host HTTP path to secure, because there is none to begin with. That is worth stating plainly in the design: the broker is not the *better* transport here, it is the *available* one. A per-host gateway that expects to be dialled by a peer cannot be built on this network without new firewall and bind-address work that nobody has asked for. ## The as-built input I wrote up the whole `fleet01` stand-up in **#156** — measured, not recalled. It is meant as the input for this ticket. The parts that bear directly on the five net-new work items here: - **Item 3 (gateway routing)** — see above; the fork has only one viable branch today. - **Item 2 (federated directory)** — `LeadLauncher.ensureLeads()` runs **once**, inline in `main()`, with no supervision loop. A dead lead is never restored and nothing logs it. Presence heartbeats across hosts will be built on top of a local layer that cannot currently notice a local death. Related: #150 (leads always report `ready: false`). - **Item 5 (cross-host trust)** — `fleet01` holds only four credentials and keeps them in `~/.zprofile`, which member panes never read because a herdr pane on Linux is a non-login shell. That is a better posture than the Mac's and is worth preserving as hosts are added. But #157 shows a token leaking through a git remote URL, which no env-var control can catch. ## The prerequisite this ticket does not name **Nothing supervises `fleetd` or `herdr` on fleet01** — no systemd unit, no user unit, no cron. I checked. A reboot or a crash ends that fleet silently, and `deploy/bridged.service` is not installed and carries the login-shell secret gap tracked in #103. Federating hosts that nothing keeps alive multiplies the failure rather than containing it. I would treat #103 as a hard dependency of this ticket, not as unrelated ops work.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#6