From 9b8d55bc1821e471004628ec5742dc35186d89d8 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Tue, 28 Jul 2026 15:20:16 +0200 Subject: [PATCH] =?UTF-8?q?CB-500:=20design=20note=20=E2=80=94=20multi-tie?= =?UTF-8?q?r=20coordination=20(Stage=206)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Scopes the lead's next direction on top of the peer-launcher arc, as a proposal (ticket split deferred): - A · sandboxed, role-specific workers — a sandbox is a *placement*, so it slots into the CB-401 SPI as a new kind: sandbox adapter exactly the way CB-402's opencode slotted in as a new provider (SPI proven placement-neutral, not just provider-neutral). Ownership line held: bridge launches INTO a peer-owned image, never provisions the IDE/dev-tools inside it. - B · main-agent pairs (Opus + cloud) — both mains are MCP clients so neither can be called into; each needs a pull inbox → depends on CB-308's per-agent channels. PrimaryRegistry single-slot → multi-slot. - C · orchestrator tier — SessionManager recursed one tier up (orchestrator:mains :: main:workers) + context scoping; re-roots the human from a live primary to the orchestrator. 10 mermaid diagrams (component + sequence per development, ownership guardrail, staging graph), all mmdc-validated and theme-safe. §7 pins the bus-vs-env-manager boundary as an acceptance criterion on A; §8 stages A → CB-308 substrate → B → C. --- docs/CB-500-Multi-Tier-Coordination.md | 357 +++++++++++++++++++++++++ 1 file changed, 357 insertions(+) create mode 100644 docs/CB-500-Multi-Tier-Coordination.md diff --git a/docs/CB-500-Multi-Tier-Coordination.md b/docs/CB-500-Multi-Tier-Coordination.md new file mode 100644 index 0000000..c9cd5e9 --- /dev/null +++ b/docs/CB-500-Multi-Tier-Coordination.md @@ -0,0 +1,357 @@ +# CB-500 — Multi-Tier Coordination (Stage 6) + +**Status:** design note (proposal — ticket split deferred) +**Depends on:** CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn), +CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate), +CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304 +(`rosterView`). +**Relates to:** the bus-identity boundary — see §7. This note stays a **proposal**; no code until the +staging in §6 is reviewed and the arc is split into tickets. + +## 1. Goal + +Grow `claude-bridge` from a **single-tier** coordinator (one human-driven primary → a flat pool of +workers) into a **multi-tier** one, along three axes the lead has asked for: + +1. **Sandboxed workers** — each worker runs in a **separated, peer-owned sandbox** carrying its own + toolchain (Claude routed via `ANTHROPIC_BASE_URL`, a headless IDE, git, MCP, dev-tools), with + **per-role** sandboxes (a backend-agent image, a frontend-agent image). +2. **Main-agent pairs** — the "main" tier becomes a **pair** (on-subscription Opus + one cloud + module) collaborating, instead of a lone primary. +3. **An orchestrator tier** — a supervisor **above** the mains that owns their **session identity** + (naming, resume) and **curates context**, so every main→worker delegation carries the *exact* + slice of context it needs and nothing else. + +The through-line: **this is not a new pillar.** It is the existing `PeerLauncher` and +`SessionManager` patterns extended one tier up, riding the **same CB-308 substrate** that multi-host +already needs. Sandbox = a placement-neutral spawn target (CB-402 pattern). Pair + orchestrator = +per-agent channels + a recursive session manager (CB-308 pattern). The bus stays a +**provider-neutral communication fabric**; every addition is addressing, launch, or session scoping — +never toolchain ownership (§7). + +## 2. Single-tier today (the assumptions to break) + +```mermaid +flowchart TB + human["human (types)"] + primary["PRIMARY (Opus)
MCP client — pull-only"] + daemon["bridged daemon
127.0.0.1:8765 (single host)"] + comp["CompositePeerLauncher
routes by kind"] + cc["ClaudeCodeLauncher"] + oc["OpenCodeLauncher"] + w1["worker pane (gx00 vLLM)"] + w2["worker pane (ollama)"] + human --> primary + primary -->|"bridge_send / spawn / ask"| daemon + daemon --> comp + comp --> cc + comp --> oc + cc --> w1 + oc --> w2 +``` + +*Figure 1 — one human-driven primary, one daemon, a flat pool of bare herdr-pane workers.* + +Four concrete bake-ins assume a single tier: + +| Assumption | Where (verified) | Why it blocks the direction | +|---|---|---| +| **Exactly one primary** | `mcp/PrimaryRegistry` — an `AtomicReference`, "single-slot registry for the primary's terminal" | A *pair* needs N addressable mains, each with its own pull inbox. | +| **Workers are bare panes** | `worker/*Launcher` spawn a herdr pane via `argv:["ccs", …]` into a pre-existing env | A *sandbox* is a richer launch target (container/devcontainer) — a new placement, not a new provider. | +| **`SpawnRequest` is flat** | `peer/SpawnRequest(profileName, requestedCwd, callerCwd)` | A sandbox/role selection needs a spawn-target dimension the record does not carry. | +| **No tier above the primary** | there is no manager of the *primary's own* session — `SessionManager` manages *workers* only | An orchestrator that names/resumes/scopes the mains is a wholly new (but pattern-reusable) tier. | + +## 3. Target multi-tier architecture + +```mermaid +flowchart TB + human["human"] + subgraph orch["TIER 0 — orchestrator"] + osm["OrchestratorSessionManager
(SessionManager, recursed up)
names · resumes · scopes context"] + end + subgraph mains["TIER 1 — main pair"] + m1["main A: Opus
MCP client"] + m2["main B: cloud module
MCP client"] + end + subgraph bus["bridged fabric (CB-307/308 substrate)"] + chan["per-agent inbox channels
agent.<globalId>.inbox"] + roster["federated roster (union view)"] + end + subgraph workers["TIER 2 — sandboxed workers"] + sbBE["backend sandbox
Claude via ANTHROPIC_BASE_URL
+ headless IDE · git · MCP · dev-tools"] + sbFE["frontend sandbox
(role-specific image)"] + end + human --> osm + osm -->|"spawn / name / resume"| m1 + osm -->|"spawn / name / resume"| m2 + m1 <-->|"pull inbox"| chan + m2 <-->|"pull inbox"| chan + m1 -->|"scoped delegation"| bus + m2 -->|"scoped delegation"| bus + bus --> sbBE + bus --> sbFE + chan --- roster +``` + +*Figure 2 — three tiers. Tier 0 owns the mains' session identity + context scope; Tier 1 is a +collaborating pair, each an MCP client with its own pull inbox; Tier 2 is peer-owned sandboxes the +bus launches into. The middle is CB-308's per-agent-channel + federated-roster substrate, now +carrying tier-to-tier traffic, not just host-to-host.* + +The recursion is the key idea: **`orchestrator : mains :: main : workers`** — the same +spawn/name/resume/scope verbs at two levels. + +## 4. Development A — Sandboxed, role-specific workers + +A "sandbox" is a **placement**, not a provider — so it slots into the CB-401 SPI exactly the way +CB-402's opencode adapter slotted in as a new *provider*. CB-402 proved the SPI is +provider-neutral; a `SandboxLauncher` proves it is **placement-neutral**. + +```mermaid +flowchart TB + req["SpawnRequest
(profileName, cwd, + sandbox/role)"] + comp["CompositePeerLauncher
routes by kind"] + cc["ClaudeCodeLauncher
kind: claude-code"] + oc["OpenCodeLauncher
kind: opencode"] + sb["SandboxLauncher (NEW)
kind: sandbox"] + subgraph owned["bridge OWNS (launch + inject boundary)"] + launch["run sandbox entrypoint
(docker/devcontainer up → agent)"] + inject["inject + guard baseUrl,
mount bridge MCP + charter"] + end + subgraph peer["peer OWNS (inside the sandbox)"] + img["image = backend|frontend role
headless IDE · git · dev-tools · MCP"] + end + req --> comp + comp --> cc + comp --> oc + comp --> sb + sb --> launch --> inject + inject -.->|"launches into, never builds"| img + classDef line fill:#b7791f,stroke:#7b341e,color:#ffffff; + class inject line +``` + +*Figure 3 — the ownership line (amber). The bridge runs the sandbox entrypoint and injects the same +boundary it owns today (guarded `baseUrl`, mounted MCP + reply charter). Everything inside the image +— the IDE, git, dev-tools — is the peer's. The bridge references the image/role; it never provisions +tools. This is what keeps "give the worker a headless IDE" on the right side of the "bus, not +env-manager" rule (§7).* + +```mermaid +sequenceDiagram + participant M as main (delegator) + participant D as bridged + participant SL as SandboxLauncher + participant SB as sandbox (peer-owned) + participant A as agent in sandbox + M->>D: bridge_spawn(profile=backend, role=backend) + D->>SL: spawn(SpawnRequest) + SL->>SB: start entrypoint (image = backend role) + Note over SL,SB: bridge injects guarded ANTHROPIC_BASE_URL,
mounts bridge MCP url + reply charter + SB->>A: launch Claude (headless IDE, git, MCP ready — peer's own) + A-->>SL: MCP connects → readiness gate (CB-306) passes + SL-->>D: PeerHandle(globalId) + D-->>M: spawned, injectable +``` + +*Figure 4 — spawn into a peer-owned sandbox. Identical control flow to today's pane spawn (incl. the +CB-306 readiness gate); only the launcher's `buildLaunch` differs — exactly the CB-402 seam.* + +**Deltas:** a new `kind: sandbox` adapter (extends the same `HerdrPeerLauncher`/`PeerLauncher` base); +a spawn-target/role dimension on `SpawnRequest` and the `Worker` profile; optionally a `SANDBOX` +`Capability`. Per-role = two profiles → two images; `CompositePeerLauncher` already routes them. If a +sandbox is a *separate host/container*, it reuses CB-308's global id + per-host gateway wholesale. + +## 5. Development B — Main-agent pairs + +Both mains are MCP **clients**, so **neither can be called into** — each needs a **pull-based +per-agent inbox**, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery +that is singular today (single-slot `PrimaryRegistry`, a push-loop aimed at one terminal, "these +tools only the primary calls") generalizes from a singleton to a **set**. + +```mermaid +flowchart TB + subgraph pair["TIER 1 — collaborating pair"] + m1["main A: Opus
MCP client (pull-only)"] + m2["main B: cloud module
MCP client (pull-only)"] + end + reg["PrimaryRegistry → multi-slot
(terminal per main)"] + subgraph fabric["bridged"] + ca["agent.A.inbox"] + cb["agent.B.inbox"] + push["ReplyPushLoop → N terminals"] + end + m1 <-->|"peer-to-peer message"| m2 + m1 -->|"register terminal"| reg + m2 -->|"register terminal"| reg + ca -->|"pull / nudge"| m1 + cb -->|"pull / nudge"| m2 + reg --> push + push --> ca + push --> cb +``` + +*Figure 5 — the pair. Each main owns an addressable inbox; `PrimaryRegistry` becomes multi-slot; the +push loop nudges each main's terminal. Mains message each other as equals over the same bus (the +transport is already peer-neutral — what was missing is N pull endpoints).* + +```mermaid +sequenceDiagram + participant MA as main A (Opus) + participant BR as bridged / broker + participant MB as main B (cloud) + MA->>BR: bridge_send(to = main B, msg) + BR->>BR: publish agent.B.inbox (durable, msg id) + Note over BR: held until B pulls (B is a client too) + MB->>BR: blocking bridge_send / poll resolves + BR-->>MB: msg (then ACK) + MB->>BR: bridge_reply(to = main A) + BR->>BR: publish agent.A.inbox + MA->>BR: poll resolves + BR-->>MA: reply +``` + +*Figure 6 — main↔main is the CB-307 asymmetry applied on both ends: two clients, so both hops are +pull. This is why Part B **depends on** the per-agent-channel substrate, not just a config flag.* + +**Deltas:** `PrimaryRegistry` single-slot → keyed-by-main; per-main inbox routing (CB-308 #1); +push-loop fan-out; relax "orchestration tools only the primary calls" to "any registered main." + +## 6. Development C — Orchestrator tier + +The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker* +sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**. + +```mermaid +flowchart TB + human["human"] + subgraph t0["TIER 0 — orchestrator (new top MCP client)"] + osm["OrchestratorSessionManager
= SessionManager pattern"] + idm["session identity
name · resume · idle-ttl (CB-303)"] + ctx["context scoper
(CB-303 context-cap + turn_id)"] + end + subgraph t1["TIER 1 — mains (now MANAGED sessions)"] + m1["main A"] + m2["main B"] + end + subgraph t2["TIER 2 — workers"] + w["sandboxed workers"] + end + human --> osm + osm --> idm + osm --> ctx + idm -->|"spawn / name / resume"| m1 + idm -->|"spawn / name / resume"| m2 + ctx -->|"inject exact context slice"| m1 + m1 -->|"scoped delegation"| w + m2 -->|"scoped delegation"| w +``` + +*Figure 7 — the recursion. Tiers 1 and 2 run the identical spawn/name/resume machinery; the +orchestrator merely operates it one level higher. **Re-rooting caveat:** today the primary IS the +human's live session; here the human drives the orchestrator, and the mains become managed, +resumable sessions. That moves the human-facing top up a tier — an intentional re-root, not an +add-on.* + +```mermaid +sequenceDiagram + participant H as human + participant O as orchestrator + participant MA as main A + participant W as worker + H->>O: high-level goal (large context) + O->>O: name/resume main A session + O->>MA: task + SCOPED context slice (not the whole history) + MA->>W: bridge_send(delegation, carrying only the relevant slice) + W-->>MA: result + MA-->>O: rollup + O->>O: fold into orchestrator context, pick next main/turn +``` + +*Figure 8 — context focus. The orchestrator holds the global context and hands each main only the +slice a given delegation needs, so the main→worker conversation stays on-point. Context *scoping* is +coordination (the bus already owns session/turn lifecycle) — it stays inside the identity boundary +(§7), unlike toolchain ownership which does not.* + +**Deltas:** a second, higher `SessionManager` instance whose "peers" are mains; the orchestrator +becomes the top MCP client; context-slice selection (new) layered on CB-303's `context_cap` + +`turn_id` scoping; mains gain a resumable session id in the federated roster. + +## 7. The identity boundary — the one clause to hold + +The bus is a **communication fabric, not an env/toolchain manager**. This direction is compatible +**only** with the ownership split below; the amber line in Figure 3 is where it must hold. + +```mermaid +flowchart LR + subgraph ok["STAYS A BUS (owned)"] + a["launch INTO a sandbox
(opaque image/role reference)"] + b["inject + guard baseUrl,
mount MCP + charter"] + c["session identity + context scope
(name/resume/turn_id)"] + d["per-agent addressing + roster"] + end + subgraph drift["BECOMES ENV-MANAGER (forbidden)"] + e["build images / install IDE
or dev-tools"] + f["wire the bridge's OWN IDE MCP
into a worker"] + g["enumerate 'what a frontend
agent needs'"] + end + ok -.->|"red flag: any feature that only
makes sense for ONE kind of peer"| drift + classDef bad fill:#9b2c2c,stroke:#742a2a,color:#ffffff; + class e,f,g bad +``` + +*Figure 9 — the guardrail. A worker having a headless IDE **inside its own sandbox** is the peer +owning its toolchain (left) — the opposite of the bridge reaching into the peer (right). Sandbox +specs are peer-owned references (like `argv`/image id); the instant the bridge builds or installs +them, it has drifted. This resolves the apparent contradiction between "do NOT give workers IDE MCP +access" and "give workers a sandboxed IDE" — different owners.* + +## 8. Staging & dependencies + +```mermaid +flowchart LR + cb402["CB-401/402
Peer Launcher SPI + composite
(DONE / in-flight)"] + A["A · SandboxLauncher
(placement-neutral, independent)"] + cb308["CB-308 substrate
per-agent channels + global id
+ federated roster"] + B["B · main-agent pair
(multi-slot PrimaryRegistry)"] + C["C · orchestrator tier
(SessionManager recursed up)"] + cb402 --> A + cb402 --> cb308 + cb308 --> B + cb308 --> C + B --> C + A -.->|"if sandbox = separate host/container,
reuses CB-308 global id"| cb308 + classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff; + class cb308 gate +``` + +*Figure 10 — the substrate (amber) is the shared enabler for B and C. Recommended order:* + +1. **Finish CB-402** (merge the opencode adapter branch). +2. **A · SandboxLauncher** — independent; a second proof of the SPI (placement-neutral). Ships anytime. +3. **CB-308 substrate** — per-agent channels + global id + federated roster (the multi-host work, + promoted from host-to-host to tier-to-tier). +4. **B · main-agent pair** — multi-slot `PrimaryRegistry` + per-main inbox routing, on the substrate. +5. **C · orchestrator tier** — the capstone; the recursive session manager + context scoping. + +## 9. Open questions (to resolve at ticket-split) + +- **Sandbox mechanism:** container (`docker exec`) vs devcontainer vs a remote host (then it *is* + CB-308). How the role→image mapping is expressed on the profile. +- **Pair semantics:** are the two mains fully symmetric peers, or is one a co-primary that may also + delegate? Affects how `PrimaryRegistry` and the "orchestration tools" identity relax. +- **Orchestrator drivenness:** the mains become programmatically spawned/resumed — does the human + still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.) +- **Context-slice selection:** who decides the slice — orchestrator heuristics, explicit tool args, + or the main pulling on demand? This is the genuinely new responsibility; keep it *scoping*, not + content authorship, to stay inside the boundary. +- **Trust:** every new tier boundary that accepts spawn/send is a trust edge (the CB-308 item #5 / + CB-401 Stage-C concern) — orchestrator→main and main→sandbox both need authz. + +## 10. Ticket-split guidance (deferred) + +This note is deliberately one arc; when split, the natural tickets are **A** (SandboxLauncher + +role/spawn-target), **the CB-308 substrate** (likely already its own ticket), **B** (multi-primary +pair), and **C** (orchestrator tier + context scoping) — with the identity clause (§7) as an +acceptance criterion on **A** specifically. Sequence per §8; nothing here is a new pillar, so each +ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C).