Files
fleetd/docs/CB-500-Multi-Tier-Coordination.md
T
Dai Ha 84c8a2d2f0 CB-500 §11: resolve distributed-sandbox topology (gateway-per-host × local sandboxes)
Clarifies the open fork from §4/§9: Development A (sandbox launcher) and CB-308
(per-host federation) COMPOSE — each host runs a bridged gateway whose launcher
spawns agents into that host's LOCAL sandboxes; the broker moves messages +
presence, never keystrokes.

The forcing fact: delivery is herdr keystroke-injection into a locally-owned PTY,
so "a sandboxed agent on another host" ≡ "a sandbox spawned by that host's
gateway" (a remote container with no local herdr can't be injected into). Rules
out a central daemon reaching remote PTYs.

Adds Figure 11 (composed topology) + Figure 12 (remote-delegation sequence: the
local?inject:publish fork with a sandboxed far side — both injection points stay
local, only the middle hop crosses the broker), the two forced reachability
changes (host-routable mcpUrl; PTY in the local gateway's herdr), and a
maps-to-existing-seams table (CB-308 gateway × Dev-A launcher, CB-117 reap,
CB-303 container lifecycle, CB-308 #5 trust). No new pillars. Both diagrams
mmdc-validated; §4/§9 updated to point at §11.
2026-07-28 16:26:10 +02:00

461 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CB-500 — Multi-Tier Coordination (Stage 6)
**Status:** design note (proposal — ticket split deferred)
**Depends on:** CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn),
CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate),
CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304
(`rosterView`).
**Relates to:** the bus-identity boundary — see §7. This note stays a **proposal**; no code until the
staging in §6 is reviewed and the arc is split into tickets.
## 1. Goal
Grow `claude-bridge` from a **single-tier** coordinator (one human-driven primary → a flat pool of
workers) into a **multi-tier** one, along three axes the lead has asked for:
1. **Sandboxed workers** — each worker runs in a **separated, peer-owned sandbox** carrying its own
toolchain (Claude routed via `ANTHROPIC_BASE_URL`, a headless IDE, git, MCP, dev-tools), with
**per-role** sandboxes (a backend-agent image, a frontend-agent image).
2. **Main-agent pairs** — the "main" tier becomes a **pair** (on-subscription Opus + one cloud
module) collaborating, instead of a lone primary.
3. **An orchestrator tier** — a supervisor **above** the mains that owns their **session identity**
(naming, resume) and **curates context**, so every main→worker delegation carries the *exact*
slice of context it needs and nothing else.
The through-line: **this is not a new pillar.** It is the existing `PeerLauncher` and
`SessionManager` patterns extended one tier up, riding the **same CB-308 substrate** that multi-host
already needs. Sandbox = a placement-neutral spawn target (CB-402 pattern). Pair + orchestrator =
per-agent channels + a recursive session manager (CB-308 pattern). The bus stays a
**provider-neutral communication fabric**; every addition is addressing, launch, or session scoping —
never toolchain ownership (§7).
## 2. Single-tier today (the assumptions to break)
```mermaid
flowchart TB
human["human (types)"]
primary["PRIMARY (Opus)<br/>MCP client — pull-only"]
daemon["bridged daemon<br/>127.0.0.1:8765 (single host)"]
comp["CompositePeerLauncher<br/>routes by kind"]
cc["ClaudeCodeLauncher"]
oc["OpenCodeLauncher"]
w1["worker pane (gx00 vLLM)"]
w2["worker pane (ollama)"]
human --> primary
primary -->|"bridge_send / spawn / ask"| daemon
daemon --> comp
comp --> cc
comp --> oc
cc --> w1
oc --> w2
```
*Figure 1 — one human-driven primary, one daemon, a flat pool of bare herdr-pane workers.*
Four concrete bake-ins assume a single tier:
| Assumption | Where (verified) | Why it blocks the direction |
|---|---|---|
| **Exactly one primary** | `mcp/PrimaryRegistry` — an `AtomicReference<String>`, "single-slot registry for the primary's terminal" | A *pair* needs N addressable mains, each with its own pull inbox. |
| **Workers are bare panes** | `worker/*Launcher` spawn a herdr pane via `argv:["ccs", …]` into a pre-existing env | A *sandbox* is a richer launch target (container/devcontainer) — a new placement, not a new provider. |
| **`SpawnRequest` is flat** | `peer/SpawnRequest(profileName, requestedCwd, callerCwd)` | A sandbox/role selection needs a spawn-target dimension the record does not carry. |
| **No tier above the primary** | there is no manager of the *primary's own* session — `SessionManager` manages *workers* only | An orchestrator that names/resumes/scopes the mains is a wholly new (but pattern-reusable) tier. |
## 3. Target multi-tier architecture
```mermaid
flowchart TB
human["human"]
subgraph orch["TIER 0 — orchestrator"]
osm["OrchestratorSessionManager<br/>(SessionManager, recursed up)<br/>names · resumes · scopes context"]
end
subgraph mains["TIER 1 — main pair"]
m1["main A: Opus<br/>MCP client"]
m2["main B: cloud module<br/>MCP client"]
end
subgraph bus["bridged fabric (CB-307/308 substrate)"]
chan["per-agent inbox channels<br/>agent.&lt;globalId&gt;.inbox"]
roster["federated roster (union view)"]
end
subgraph workers["TIER 2 — sandboxed workers"]
sbBE["backend sandbox<br/>Claude via ANTHROPIC_BASE_URL<br/>+ headless IDE · git · MCP · dev-tools"]
sbFE["frontend sandbox<br/>(role-specific image)"]
end
human --> osm
osm -->|"spawn / name / resume"| m1
osm -->|"spawn / name / resume"| m2
m1 <-->|"pull inbox"| chan
m2 <-->|"pull inbox"| chan
m1 -->|"scoped delegation"| bus
m2 -->|"scoped delegation"| bus
bus --> sbBE
bus --> sbFE
chan --- roster
```
*Figure 2 — three tiers. Tier 0 owns the mains' session identity + context scope; Tier 1 is a
collaborating pair, each an MCP client with its own pull inbox; Tier 2 is peer-owned sandboxes the
bus launches into. The middle is CB-308's per-agent-channel + federated-roster substrate, now
carrying tier-to-tier traffic, not just host-to-host.*
The recursion is the key idea: **`orchestrator : mains :: main : workers`** — the same
spawn/name/resume/scope verbs at two levels.
## 4. Development A — Sandboxed, role-specific workers
A "sandbox" is a **placement**, not a provider — so it slots into the CB-401 SPI exactly the way
CB-402's opencode adapter slotted in as a new *provider*. CB-402 proved the SPI is
provider-neutral; a `SandboxLauncher` proves it is **placement-neutral**.
```mermaid
flowchart TB
req["SpawnRequest<br/>(profileName, cwd, + sandbox/role)"]
comp["CompositePeerLauncher<br/>routes by kind"]
cc["ClaudeCodeLauncher<br/>kind: claude-code"]
oc["OpenCodeLauncher<br/>kind: opencode"]
sb["SandboxLauncher (NEW)<br/>kind: sandbox"]
subgraph owned["bridge OWNS (launch + inject boundary)"]
launch["run sandbox entrypoint<br/>(docker/devcontainer up → agent)"]
inject["inject + guard baseUrl,<br/>mount bridge MCP + charter"]
end
subgraph peer["peer OWNS (inside the sandbox)"]
img["image = backend|frontend role<br/>headless IDE · git · dev-tools · MCP"]
end
req --> comp
comp --> cc
comp --> oc
comp --> sb
sb --> launch --> inject
inject -.->|"launches into, never builds"| img
classDef line fill:#b7791f,stroke:#7b341e,color:#ffffff;
class inject line
```
*Figure 3 — the ownership line (amber). The bridge runs the sandbox entrypoint and injects the same
boundary it owns today (guarded `baseUrl`, mounted MCP + reply charter). Everything inside the image
— the IDE, git, dev-tools — is the peer's. The bridge references the image/role; it never provisions
tools. This is what keeps "give the worker a headless IDE" on the right side of the "bus, not
env-manager" rule (§7).*
```mermaid
sequenceDiagram
participant M as main (delegator)
participant D as bridged
participant SL as SandboxLauncher
participant SB as sandbox (peer-owned)
participant A as agent in sandbox
M->>D: bridge_spawn(profile=backend, role=backend)
D->>SL: spawn(SpawnRequest)
SL->>SB: start entrypoint (image = backend role)
Note over SL,SB: bridge injects guarded ANTHROPIC_BASE_URL,<br/>mounts bridge MCP url + reply charter
SB->>A: launch Claude (headless IDE, git, MCP ready — peer's own)
A-->>SL: MCP connects → readiness gate (CB-306) passes
SL-->>D: PeerHandle(globalId)
D-->>M: spawned, injectable
```
*Figure 4 — spawn into a peer-owned sandbox. Identical control flow to today's pane spawn (incl. the
CB-306 readiness gate); only the launcher's `buildLaunch` differs — exactly the CB-402 seam.*
**Deltas:** a new `kind: sandbox` adapter (extends the same `HerdrPeerLauncher`/`PeerLauncher` base);
a spawn-target/role dimension on `SpawnRequest` and the `Worker` profile; optionally a `SANDBOX`
`Capability`. Per-role = two profiles → two images; `CompositePeerLauncher` already routes them. If a
sandbox is a *separate host/container*, it reuses CB-308's global id + per-host gateway wholesale —
**the distributed case is resolved in §11: a sandbox on another host is one spawned by that host's
gateway, because herdr keystroke-injection needs a locally-owned PTY.**
## 5. Development B — Main-agent pairs
Both mains are MCP **clients**, so **neither can be called into** — each needs a **pull-based
per-agent inbox**, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery
that is singular today (single-slot `PrimaryRegistry`, a push-loop aimed at one terminal, "these
tools only the primary calls") generalizes from a singleton to a **set**.
```mermaid
flowchart TB
subgraph pair["TIER 1 — collaborating pair"]
m1["main A: Opus<br/>MCP client (pull-only)"]
m2["main B: cloud module<br/>MCP client (pull-only)"]
end
reg["PrimaryRegistry → multi-slot<br/>(terminal per main)"]
subgraph fabric["bridged"]
ca["agent.A.inbox"]
cb["agent.B.inbox"]
push["ReplyPushLoop → N terminals"]
end
m1 <-->|"peer-to-peer message"| m2
m1 -->|"register terminal"| reg
m2 -->|"register terminal"| reg
ca -->|"pull / nudge"| m1
cb -->|"pull / nudge"| m2
reg --> push
push --> ca
push --> cb
```
*Figure 5 — the pair. Each main owns an addressable inbox; `PrimaryRegistry` becomes multi-slot; the
push loop nudges each main's terminal. Mains message each other as equals over the same bus (the
transport is already peer-neutral — what was missing is N pull endpoints).*
```mermaid
sequenceDiagram
participant MA as main A (Opus)
participant BR as bridged / broker
participant MB as main B (cloud)
MA->>BR: bridge_send(to = main B, msg)
BR->>BR: publish agent.B.inbox (durable, msg id)
Note over BR: held until B pulls (B is a client too)
MB->>BR: blocking bridge_send / poll resolves
BR-->>MB: msg (then ACK)
MB->>BR: bridge_reply(to = main A)
BR->>BR: publish agent.A.inbox
MA->>BR: poll resolves
BR-->>MA: reply
```
*Figure 6 — main↔main is the CB-307 asymmetry applied on both ends: two clients, so both hops are
pull. This is why Part B **depends on** the per-agent-channel substrate, not just a config flag.*
**Deltas:** `PrimaryRegistry` single-slot → keyed-by-main; per-main inbox routing (CB-308 #1);
push-loop fan-out; relax "orchestration tools only the primary calls" to "any registered main."
## 6. Development C — Orchestrator tier
The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker*
sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**.
```mermaid
flowchart TB
human["human"]
subgraph t0["TIER 0 — orchestrator (new top MCP client)"]
osm["OrchestratorSessionManager<br/>= SessionManager pattern"]
idm["session identity<br/>name · resume · idle-ttl (CB-303)"]
ctx["context scoper<br/>(CB-303 context-cap + turn_id)"]
end
subgraph t1["TIER 1 — mains (now MANAGED sessions)"]
m1["main A"]
m2["main B"]
end
subgraph t2["TIER 2 — workers"]
w["sandboxed workers"]
end
human --> osm
osm --> idm
osm --> ctx
idm -->|"spawn / name / resume"| m1
idm -->|"spawn / name / resume"| m2
ctx -->|"inject exact context slice"| m1
m1 -->|"scoped delegation"| w
m2 -->|"scoped delegation"| w
```
*Figure 7 — the recursion. Tiers 1 and 2 run the identical spawn/name/resume machinery; the
orchestrator merely operates it one level higher. **Re-rooting caveat:** today the primary IS the
human's live session; here the human drives the orchestrator, and the mains become managed,
resumable sessions. That moves the human-facing top up a tier — an intentional re-root, not an
add-on.*
```mermaid
sequenceDiagram
participant H as human
participant O as orchestrator
participant MA as main A
participant W as worker
H->>O: high-level goal (large context)
O->>O: name/resume main A session
O->>MA: task + SCOPED context slice (not the whole history)
MA->>W: bridge_send(delegation, carrying only the relevant slice)
W-->>MA: result
MA-->>O: rollup
O->>O: fold into orchestrator context, pick next main/turn
```
*Figure 8 — context focus. The orchestrator holds the global context and hands each main only the
slice a given delegation needs, so the main→worker conversation stays on-point. Context *scoping* is
coordination (the bus already owns session/turn lifecycle) — it stays inside the identity boundary
(§7), unlike toolchain ownership which does not.*
**Deltas:** a second, higher `SessionManager` instance whose "peers" are mains; the orchestrator
becomes the top MCP client; context-slice selection (new) layered on CB-303's `context_cap` +
`turn_id` scoping; mains gain a resumable session id in the federated roster.
## 7. The identity boundary — the one clause to hold
The bus is a **communication fabric, not an env/toolchain manager**. This direction is compatible
**only** with the ownership split below; the amber line in Figure 3 is where it must hold.
```mermaid
flowchart LR
subgraph ok["STAYS A BUS (owned)"]
a["launch INTO a sandbox<br/>(opaque image/role reference)"]
b["inject + guard baseUrl,<br/>mount MCP + charter"]
c["session identity + context scope<br/>(name/resume/turn_id)"]
d["per-agent addressing + roster"]
end
subgraph drift["BECOMES ENV-MANAGER (forbidden)"]
e["build images / install IDE<br/>or dev-tools"]
f["wire the bridge's OWN IDE MCP<br/>into a worker"]
g["enumerate 'what a frontend<br/>agent needs'"]
end
ok -.->|"red flag: any feature that only<br/>makes sense for ONE kind of peer"| drift
classDef bad fill:#9b2c2c,stroke:#742a2a,color:#ffffff;
class e,f,g bad
```
*Figure 9 — the guardrail. A worker having a headless IDE **inside its own sandbox** is the peer
owning its toolchain (left) — the opposite of the bridge reaching into the peer (right). Sandbox
specs are peer-owned references (like `argv`/image id); the instant the bridge builds or installs
them, it has drifted. This resolves the apparent contradiction between "do NOT give workers IDE MCP
access" and "give workers a sandboxed IDE" — different owners.*
## 8. Staging & dependencies
```mermaid
flowchart LR
cb402["CB-401/402<br/>Peer Launcher SPI + composite<br/>(DONE / in-flight)"]
A["A · SandboxLauncher<br/>(placement-neutral, independent)"]
cb308["CB-308 substrate<br/>per-agent channels + global id<br/>+ federated roster"]
B["B · main-agent pair<br/>(multi-slot PrimaryRegistry)"]
C["C · orchestrator tier<br/>(SessionManager recursed up)"]
cb402 --> A
cb402 --> cb308
cb308 --> B
cb308 --> C
B --> C
A -.->|"if sandbox = separate host/container,<br/>reuses CB-308 global id"| cb308
classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
class cb308 gate
```
*Figure 10 — the substrate (amber) is the shared enabler for B and C. Recommended order:*
1. **Finish CB-402** (merge the opencode adapter branch).
2. **A · SandboxLauncher** — independent; a second proof of the SPI (placement-neutral). Ships anytime.
3. **CB-308 substrate** — per-agent channels + global id + federated roster (the multi-host work,
promoted from host-to-host to tier-to-tier).
4. **B · main-agent pair** — multi-slot `PrimaryRegistry` + per-main inbox routing, on the substrate.
5. **C · orchestrator tier** — the capstone; the recursive session manager + context scoping.
## 9. Open questions (to resolve at ticket-split)
- **Sandbox mechanism:** container (`docker exec`) vs devcontainer — how the role→image mapping is
expressed on the profile. *(Topology **resolved** in §11: distributed = gateway-per-host × local
sandboxes; the remaining choice is only the local launch mechanism, not the shape.)*
- **Pair semantics:** are the two mains fully symmetric peers, or is one a co-primary that may also
delegate? Affects how `PrimaryRegistry` and the "orchestration tools" identity relax.
- **Orchestrator drivenness:** the mains become programmatically spawned/resumed — does the human
still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.)
- **Context-slice selection:** who decides the slice — orchestrator heuristics, explicit tool args,
or the main pulling on demand? This is the genuinely new responsibility; keep it *scoping*, not
content authorship, to stay inside the boundary.
- **Trust:** every new tier boundary that accepts spawn/send is a trust edge (the CB-308 item #5 /
CB-401 Stage-C concern) — orchestrator→main and main→sandbox both need authz.
## 10. Ticket-split guidance (deferred)
This note is deliberately one arc; when split, the natural tickets are **A** (SandboxLauncher +
role/spawn-target), **the CB-308 substrate** (likely already its own ticket), **B** (multi-primary
pair), and **C** (orchestrator tier + context scoping) — with the identity clause (§7) as an
acceptance criterion on **A** specifically. Sequence per §8; nothing here is a new pillar, so each
ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C).
## 11. Distributed sandboxes — the resolved topology
The follow-up question — *"clarify the architecture when we have distributed agents in sandboxes"* —
resolves the fork left open in §4 and §9. **Decision: Development A (sandbox launcher) and CB-308
(per-host federation) *compose*, not compete — each host runs a `bridged` gateway whose launcher
spawns agents into that host's *local* sandboxes.** A sandbox is never reached across the network; it
is reached by the gateway sitting next to it.
### 11.1 The one fact that fixes the shape
The bus delivers a turn by **herdr keystroke-injection** — `Injector → AgentControl.send` writes into
a PTY that its **local** herdr owns. The broker moves *messages and presence*, **never keystrokes**.
So an agent's PTY must live in a herdr that *some* `bridged` instance drives locally: a remote
container with no local herdr **cannot be injected into**. That rules out a central daemon reaching
remote PTYs, and collapses the design to a single identity:
> **"a sandboxed agent on another host" ≡ "a sandbox spawned by that host's gateway."**
```mermaid
flowchart TB
subgraph hostA["HOST A — gateway"]
mA["main / orchestrator<br/>MCP client → LOCAL gateway"]
gA["bridged A<br/>herdr + CompositePeerLauncher<br/>(incl. SandboxLauncher)"]
cBEa["sandbox: backend<br/>(local container)"]
cFEa["sandbox: frontend<br/>(local container)"]
mA --- gA
gA -->|"spawn (docker/devcontainer)<br/>→ PTY in A's herdr"| cBEa
gA --> cFEa
end
subgraph broker["BROKER (AMQP) — CB-307/308 fabric"]
inbox["agent.ID.inbox queues"]
roster["roster.* (federated presence)"]
end
subgraph hostB["HOST B — gateway"]
gB["bridged B<br/>herdr + SandboxLauncher"]
cBEb["sandbox: backend<br/>(local container)"]
gB -->|"spawn → PTY in B's herdr"| cBEb
end
gA <-->|"messages + presence<br/>(NOT keystrokes)"| inbox
gB <-->|"messages + presence"| inbox
gA --- roster
gB --- roster
classDef line fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
class inbox,roster line
```
*Figure 11 — the composed topology. Each gateway owns its local herdr and runs a `SandboxLauncher`
(the §4 adapter) that spawns role-specific containers **on its own host**; the broker (blue) carries
only messages + roster between gateways. Keystroke-injection stays strictly local to each gateway.*
### 11.2 How a delegation reaches a sandboxed agent on another host
```mermaid
sequenceDiagram
participant MA as main (host A)
participant GA as gateway A
participant BR as broker
participant GB as gateway B
participant SB as sandbox agent (host B, container)
MA->>GA: bridge_send(globalId on B, msg)
GA->>GA: directory lookup - is globalId local? NO
GA->>BR: publish agent.ID.inbox (durable)
BR->>GB: route to the owning gateway
GB->>SB: inject via B's LOCAL herdr (keystrokes)
Note over GB,SB: SandboxLauncher already spawned the container -<br/>its PTY is in B's herdr, CB-306 readiness passed
SB-->>GB: bridge_reply (to B's LOCAL MCP endpoint)
GB->>BR: publish primary-bound (durable, msg id)
BR->>GA: route back to A
Note over GA: held until the main pulls (the main is a client)
MA->>GA: poll / blocking send resolves
GA-->>MA: reply
```
*Figure 12 — the `local ? inject : publish` fork (CB-308 §3.2) with a sandboxed far side. Only the
**middle** hop crosses the network via the broker; **both** injection points (into the sandbox on B,
and the drain-nudge back into the main on A) are local herdr writes. This is CB-308's routing rule
unchanged — the sandbox is transparent to it.*
### 11.3 Two reachability changes any sandbox forces
| Change | Today | Under sandboxes |
|---|---|---|
| **`mcpUrl`** | `http://127.0.0.1:8765/mcp` (loopback) | must be **host-routable from inside the container** (e.g. `host.docker.internal` or the gateway's LAN IP) — the worker connects to **its own gateway's** MCP, never a remote one. |
| **PTY ownership** | pane in the daemon's herdr | pane is the **container's** attached PTY, in the **local** gateway's herdr (via `docker exec`/devcontainer) — non-negotiable per §11.1. |
### 11.4 Everything maps to an existing seam (nothing new invented)
| Concern | Provided by |
|---|---|
| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `bridged`, evolved) |
| Spawn into a local sandbox / role→image | **Development A** `SandboxLauncher` (§4), routed by `CompositePeerLauncher` |
| Addressing a remote sandboxed agent | **CB-308** global id + federated roster (host + role as metadata) |
| Orphan reap after a gateway restart | **CB-117** per-gateway, summed by the composite — each reaps only its **local** herdr |
| Container up/down | tied to **CB-303** session lifecycle — `SandboxLauncher.stop` tears the container down with the pane |
| Cross-gateway spawn/send trust | **CB-308 item #5** / CB-401 Stage-C — each gateway edge is a trust boundary |
*The net: distributed sandboxes add **zero** new pillars — they are `CB-308 gateway × Development-A
launcher` at every host, with the §7 ownership line (peer owns the image; the bridge only launches
into it) holding at each gateway.*