Files
fleetd/docs/CB-500-Multi-Tier-Coordination.md
Dai Ha ecc590f344 CB-632 unit 5: rename the daemon and classes in the docs prose
Part of #145 (CB-632). Documentation only, plus one internal literal.

Unit 1 renamed the package and classes, which left every doc describing
classes that no longer exist. This fixes the prose across README.md,
docs/ and bridged/docs/ -- 18 files.

Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and
"bridged" where it names the daemon as a product rather than a path.

Also renamed two literals, because a doc that disagrees with the code is
worse than one that is out of date:

  - bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey
    OpenCodeLauncher sends when a profile resolves no token, to a local
    endpoint that does not check it. No test asserts the old string.
  - the vnd.ltms.bridged.* media type in the M4 design doc. It appears
    in no Java file, so nothing implements it yet.

Deliberately NOT renamed, because each is still literally true today and
changes only at the cutover:

  - paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar,
    .bridged-worktrees, deploy/dev.ltms.bridged.plist,
    scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh
  - bridged_* metric names -- renaming these after the monitoring is
    wired would break dashboard continuity, so they move before it is
  - bridge_* MCP tool names, which answer alongside fleet_* on purpose
  - BRIDGED_* environment variables, read by a file outside this repo

Method note: perl, not sed. BSD sed has no \b and no lookaround, and a
word-boundary expression there fails silently. The prose replace uses
(?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an
identifier, then every remaining hit was read by hand.

Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
2026-08-23 06:46:34 +02:00

24 KiB
Raw Permalink Blame History

CB-500 — Multi-Tier Coordination (Stage 6)

Status: design note. Developments A/B remain proposals; Development C (§6 and Figures 7–8) is SUPERSEDED by the advisory-architect design in Gitea issue #16 and the architects: configuration block (CB-548). Depends on: CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn), CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate), CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304 (rosterView). Relates to: the bus-identity boundary — see §7. This note stays a proposal; no code until the staging in §6 is reviewed and the arc is split into tickets.

1. Goal

Grow claude-bridge from a single-tier coordinator (one human-driven primary → a flat pool of workers) into a multi-tier one, along three axes the lead has asked for:

  1. Sandboxed workers — each worker runs in a separated, peer-owned sandbox carrying its own toolchain (Claude routed via ANTHROPIC_BASE_URL, a headless IDE, git, MCP, dev-tools), with per-role sandboxes (a backend-agent image, a frontend-agent image).
  2. Main-agent pairs — SUPERSEDED. The considered model made the "main" tier a pair (on-subscription Opus + one cloud module). The actual fleet is one human-driven lead plus two short-lived advisory architects on different model families.
  3. An orchestrator tier — a supervisor above the mains that owns their session identity (naming, resume) and curates context, so every main→worker delegation carries the exact slice of context it needs and nothing else.

The through-line: this is not a new pillar. It is the existing PeerLauncher and SessionManager patterns extended one tier up, riding the same CB-308 substrate that multi-host already needs. Sandbox = a placement-neutral spawn target (CB-402 pattern). Pair + orchestrator = per-agent channels + a recursive session manager (CB-308 pattern). The bus stays a provider-neutral communication fabric; every addition is addressing, launch, or session scoping — never toolchain ownership (§7).

2. Single-tier today (the assumptions to break)

flowchart TB
    human["human (types)"]
    primary["PRIMARY (Opus)<br/>MCP client — pull-only"]
    daemon["fleetd daemon<br/>127.0.0.1:8765 (single host)"]
    comp["CompositePeerLauncher<br/>routes by kind"]
    cc["ClaudeCodeLauncher"]
    oc["OpenCodeLauncher"]
    w1["worker pane (gx00 vLLM)"]
    w2["worker pane (ollama)"]
    human --> primary
    primary -->|"fleet_send / spawn / ask"| daemon
    daemon --> comp
    comp --> cc
    comp --> oc
    cc --> w1
    oc --> w2

Figure 1 — one human-driven primary, one daemon, a flat pool of bare herdr-pane workers.

Four concrete bake-ins assume a single tier:

Assumption Where (verified) Why it blocks the direction
Exactly one primary mcp/PrimaryRegistry — an AtomicReference<String>, "single-slot registry for the primary's terminal" A pair needs N addressable mains, each with its own pull inbox.
Workers are bare panes worker/*Launcher spawn a herdr pane via argv:["ccs", …] into a pre-existing env A sandbox is a richer launch target (container/devcontainer) — a new placement, not a new provider.
SpawnRequest is flat peer/SpawnRequest(profileName, requestedCwd, callerCwd) A sandbox/role selection needs a spawn-target dimension the record does not carry.
No tier above the primary there is no manager of the primary's own session — SessionManager manages workers only An orchestrator that names/resumes/scopes the mains is a wholly new (but pattern-reusable) tier.

3. Target multi-tier architecture

SUPERSEDED fleet sketch. Figure 2 records the former two-main model. The actual fleet is one lead, two independent advisory architects, and N workers; architects are sideways peers, not leads and not a tier above the lead. See Gitea issue #16 and the architects: block.

flowchart TB
    human["human"]
    subgraph orch["TIER 0 — orchestrator"]
        osm["OrchestratorSessionManager<br/>(SessionManager, recursed up)<br/>names · resumes · scopes context"]
    end
    subgraph mains["TIER 1 — main pair"]
        m1["main A: Opus<br/>MCP client"]
        m2["main B: cloud module<br/>MCP client"]
    end
    subgraph bus["fleetd fabric (CB-307/308 substrate)"]
        chan["per-agent inbox channels<br/>agent.&lt;globalId&gt;.inbox"]
        roster["federated roster (union view)"]
    end
    subgraph workers["TIER 2 — sandboxed workers"]
        sbBE["backend sandbox<br/>Claude via ANTHROPIC_BASE_URL<br/>+ headless IDE · git · MCP · dev-tools"]
        sbFE["frontend sandbox<br/>(role-specific image)"]
    end
    human --> osm
    osm -->|"spawn / name / resume"| m1
    osm -->|"spawn / name / resume"| m2
    m1 <-->|"pull inbox"| chan
    m2 <-->|"pull inbox"| chan
    m1 -->|"scoped delegation"| bus
    m2 -->|"scoped delegation"| bus
    bus --> sbBE
    bus --> sbFE
    chan --- roster

Figure 2 — SUPERSEDED historical fleet sketch. It proposed a collaborating pair of managed mains. The actual fleet keeps one human-driven lead and uses two independent, short-lived advisory architects on different model families, so agreement is evidence rather than correlated echo.

The recursion is the key idea: orchestrator : mains :: main : workers — the same spawn/name/resume/scope verbs at two levels.

4. Development A — Sandboxed, role-specific workers

A "sandbox" is a placement, not a provider — so it slots into the CB-401 SPI exactly the way CB-402's opencode adapter slotted in as a new provider. CB-402 proved the SPI is provider-neutral; a SandboxLauncher proves it is placement-neutral.

flowchart TB
    req["SpawnRequest<br/>(profileName, cwd, + sandbox/role)"]
    comp["CompositePeerLauncher<br/>routes by kind"]
    cc["ClaudeCodeLauncher<br/>kind: claude-code"]
    oc["OpenCodeLauncher<br/>kind: opencode"]
    sb["SandboxLauncher (NEW)<br/>kind: sandbox"]
    subgraph owned["bridge OWNS (launch + inject boundary)"]
        launch["run sandbox entrypoint<br/>(docker/devcontainer up → agent)"]
        inject["inject + guard baseUrl,<br/>mount bridge MCP + charter"]
    end
    subgraph peer["peer OWNS (inside the sandbox)"]
        img["image = backend|frontend role<br/>headless IDE · git · dev-tools · MCP"]
    end
    req --> comp
    comp --> cc
    comp --> oc
    comp --> sb
    sb --> launch --> inject
    inject -.->|"launches into, never builds"| img
    classDef line fill:#b7791f,stroke:#7b341e,color:#ffffff;
    class inject line

Figure 3 — the ownership line (amber). The bridge runs the sandbox entrypoint and injects the same boundary it owns today (guarded baseUrl, mounted MCP + reply charter). Everything inside the image — the IDE, git, dev-tools — is the peer's. The bridge references the image/role; it never provisions tools. This is what keeps "give the worker a headless IDE" on the right side of the "bus, not env-manager" rule (§7).

sequenceDiagram
    participant M as main (delegator)
    participant D as fleetd
    participant SL as SandboxLauncher
    participant SB as sandbox (peer-owned)
    participant A as agent in sandbox
    M->>D: fleet_spawn(profile=backend, role=backend)
    D->>SL: spawn(SpawnRequest)
    SL->>SB: start entrypoint (image = backend role)
    Note over SL,SB: bridge injects guarded ANTHROPIC_BASE_URL,<br/>mounts bridge MCP url + reply charter
    SB->>A: launch Claude (headless IDE, git, MCP ready — peer's own)
    A-->>SL: MCP connects → readiness gate (CB-306) passes
    SL-->>D: PeerHandle(globalId)
    D-->>M: spawned, injectable

Figure 4 — spawn into a peer-owned sandbox. Identical control flow to today's pane spawn (incl. the CB-306 readiness gate); only the launcher's buildLaunch differs — exactly the CB-402 seam.

Deltas: a new kind: sandbox adapter (extends the same HerdrPeerLauncher/PeerLauncher base); a spawn-target/role dimension on SpawnRequest and the Worker profile; optionally a SANDBOX Capability. Per-role = two profiles → two images; CompositePeerLauncher already routes them. If a sandbox is a separate host/container, it reuses CB-308's global id + per-host gateway wholesale — the distributed case is resolved in §11: a sandbox on another host is one spawned by that host's gateway, because herdr keystroke-injection needs a locally-owned PTY.

5. Development B — Main-agent pairs

SUPERSEDED — do not implement this model. The two-main fleet was replaced by one human-driven lead and two independent advisory architects. They are deliberately different model families (Claude Sonnet 5 and GPT-5.6 through opencode), receive the same brief, and work independently so agreement is evidence rather than correlated echo. See Gitea issue #16 and architects:.

Both mains are MCP clients, so neither can be called into — each needs a pull-based per-agent inbox, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery that is singular today (single-slot PrimaryRegistry, a push-loop aimed at one terminal, "these tools only the primary calls") generalizes from a singleton to a set.

flowchart TB
    subgraph pair["TIER 1 — collaborating pair"]
        m1["main A: Opus<br/>MCP client (pull-only)"]
        m2["main B: cloud module<br/>MCP client (pull-only)"]
    end
    reg["PrimaryRegistry → multi-slot<br/>(terminal per main)"]
    subgraph fabric["fleetd"]
        ca["agent.A.inbox"]
        cb["agent.B.inbox"]
        push["ReplyPushLoop → N terminals"]
    end
    m1 <-->|"peer-to-peer message"| m2
    m1 -->|"register terminal"| reg
    m2 -->|"register terminal"| reg
    ca -->|"pull / nudge"| m1
    cb -->|"pull / nudge"| m2
    reg --> push
    push --> ca
    push --> cb

Figure 5 — the pair. Each main owns an addressable inbox; PrimaryRegistry becomes multi-slot; the push loop nudges each main's terminal. Mains message each other as equals over the same bus (the transport is already peer-neutral — what was missing is N pull endpoints).

sequenceDiagram
    participant MA as main A (Opus)
    participant BR as fleetd / broker
    participant MB as main B (cloud)
    MA->>BR: fleet_send(to = main B, msg)
    BR->>BR: publish agent.B.inbox (durable, msg id)
    Note over BR: held until B pulls (B is a client too)
    MB->>BR: blocking fleet_send / poll resolves
    BR-->>MB: msg (then ACK)
    MB->>BR: fleet_reply(to = main A)
    BR->>BR: publish agent.A.inbox
    MA->>BR: poll resolves
    BR-->>MA: reply

Figure 6 — main↔main is the CB-307 asymmetry applied on both ends: two clients, so both hops are pull. This is why Part B depends on the per-agent-channel substrate, not just a config flag.

Deltas: PrimaryRegistry single-slot → keyed-by-main; per-main inbox routing (CB-308 #1); push-loop fan-out; relax "orchestration tools only the primary calls" to "any registered main."

6. Development C — Orchestrator tier

SUPERSEDED — do not implement this model. The operator rejected a supervisor above the lead. The human continues to drive the pre-existing lead directly; fleetd neither spawns nor resumes that lead. What replaced this proposal is one lead, two short-lived advisory architects, and N workers: the lead engages architects sideways for a strong-model assessment, then discards them. Architect slots are declared in architects: (see Gitea issue #16), rather than making leads managed sessions. The two architects deliberately use different model families — Claude Sonnet 5 and GPT-5.6 through opencode — and receive the same brief independently. Agreement is evidence, not correlated echo from one provider or one conversation.

Historical alternative retained. The text and figures below record the considered model and why it was rejected: it re-rooted the human-facing session above the lead, violating the still-true premise that configured leaders pre-exist, are recognised, and cannot be resumed by fleetd.

The orchestrator is SessionManager recursed one tier up: today it spawns/names/reaps worker sessions; the orchestrator does the same for main sessions, and adds context scoping.

flowchart TB
    human["human"]
    subgraph t0["TIER 0 — orchestrator (new top MCP client)"]
        osm["OrchestratorSessionManager<br/>= SessionManager pattern"]
        idm["session identity<br/>name · resume · idle-ttl (CB-303)"]
        ctx["context scoper<br/>(CB-303 context-cap + turn_id)"]
    end
    subgraph t1["TIER 1 — mains (now MANAGED sessions)"]
        m1["main A"]
        m2["main B"]
    end
    subgraph t2["TIER 2 — workers"]
        w["sandboxed workers"]
    end
    human --> osm
    osm --> idm
    osm --> ctx
    idm -->|"spawn / name / resume"| m1
    idm -->|"spawn / name / resume"| m2
    ctx -->|"inject exact context slice"| m1
    m1 -->|"scoped delegation"| w
    m2 -->|"scoped delegation"| w

Figure 7 — SUPERSEDED historical alternative. The recursion re-rooted the human-facing session: the human drove an orchestrator and the mains became managed, resumable sessions. The operator rejected that re-root. The replacement keeps the human-driven, pre-existing lead and engages architects sideways as short-lived advisory peers; see Gitea issue #16 and architects:.

sequenceDiagram
    participant H as human
    participant O as orchestrator
    participant MA as main A
    participant W as worker
    H->>O: high-level goal (large context)
    O->>O: name/resume main A session
    O->>MA: task + SCOPED context slice (not the whole history)
    MA->>W: fleet_send(delegation, carrying only the relevant slice)
    W-->>MA: result
    MA-->>O: rollup
    O->>O: fold into orchestrator context, pick next main/turn

Figure 8 — SUPERSEDED historical alternative. This proposed an orchestrator holding global context and slicing it for managed mains. The replacement has the human-driven lead send the same advisory brief issue #16 and architects:.

Deltas: a second, higher SessionManager instance whose "peers" are mains; the orchestrator becomes the top MCP client; context-slice selection (new) layered on CB-303's context_cap + turn_id scoping; mains gain a resumable session id in the federated roster.

7. The identity boundary — the one clause to hold

The bus is a communication fabric, not an env/toolchain manager. This direction is compatible only with the ownership split below; the amber line in Figure 3 is where it must hold.

flowchart LR
    subgraph ok["STAYS A BUS (owned)"]
        a["launch INTO a sandbox<br/>(opaque image/role reference)"]
        b["inject + guard baseUrl,<br/>mount MCP + charter"]
        c["session identity + context scope<br/>(name/resume/turn_id)"]
        d["per-agent addressing + roster"]
    end
    subgraph drift["BECOMES ENV-MANAGER (forbidden)"]
        e["build images / install IDE<br/>or dev-tools"]
        f["wire the bridge's OWN IDE MCP<br/>into a worker"]
        g["enumerate 'what a frontend<br/>agent needs'"]
    end
    ok -.->|"red flag: any feature that only<br/>makes sense for ONE kind of peer"| drift
    classDef bad fill:#9b2c2c,stroke:#742a2a,color:#ffffff;
    class e,f,g bad

Figure 9 — the guardrail. A worker having a headless IDE inside its own sandbox is the peer owning its toolchain (left) — the opposite of the bridge reaching into the peer (right). Sandbox specs are peer-owned references (like argv/image id); the instant the bridge builds or installs them, it has drifted. This resolves the apparent contradiction between "do NOT give workers IDE MCP access" and "give workers a sandboxed IDE" — different owners.

8. Staging & dependencies

flowchart LR
    cb402["CB-401/402<br/>Peer Launcher SPI + composite<br/>(DONE / in-flight)"]
    A["A · SandboxLauncher<br/>(placement-neutral, independent)"]
    cb308["CB-308 substrate<br/>per-agent channels + global id<br/>+ federated roster"]
    B["B · main-agent pair (SUPERSEDED)<br/>(multi-slot PrimaryRegistry)"]
    C["C · orchestrator tier<br/>(SessionManager recursed up)"]
    cb402 --> A
    cb402 --> cb308
    cb308 --> B
    cb308 --> C
    B --> C
    A -.->|"if sandbox = separate host/container,<br/>reuses CB-308 global id"| cb308
    classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
    class cb308 gate

Figure 10 — the substrate (amber) is the shared enabler for B and C. Recommended order:

  1. Finish CB-402 (merge the opencode adapter branch).
  2. A · SandboxLauncher — independent; a second proof of the SPI (placement-neutral). Ships anytime.
  3. CB-308 substrate — per-agent channels + global id + federated roster (the multi-host work, promoted from host-to-host to tier-to-tier).
  4. B · main-agent pair — SUPERSEDED by lead + two advisory architects.
  5. C · orchestrator tier — the capstone; the recursive session manager + context scoping.

9. Open questions (to resolve at ticket-split)

  • Sandbox mechanism: container (docker exec) vs devcontainer — how the role→image mapping is expressed on the profile. (Topology resolved in §11: distributed = gateway-per-host × local sandboxes; the remaining choice is only the local launch mechanism, not the shape.)
  • Pair semantics: SUPERSEDED. The two-main question is replaced by the architect role's least-privilege boundary: advisory architects can send/reply/ask/read but cannot spawn/stop/drain.
  • Orchestrator drivenness: the mains become programmatically spawned/resumed — does the human still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.)
  • Context-slice selection: who decides the slice — orchestrator heuristics, explicit tool args, or the main pulling on demand? This is the genuinely new responsibility; keep it scoping, not content authorship, to stay inside the boundary.
  • Trust: every new tier boundary that accepts spawn/send is a trust edge (the CB-308 item #5 / CB-401 Stage-C concern) — orchestrator→main and main→sandbox both need authz.

10. Ticket-split guidance (deferred)

This note is deliberately one arc; when split, the natural tickets are A (SandboxLauncher + role/spawn-target), the CB-308 substrate (likely already its own ticket), B (multi-primary pair), and C (orchestrator tier + context scoping) — with the identity clause (§7) as an acceptance criterion on A specifically. Sequence per §8; nothing here is a new pillar, so each ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C).

11. Distributed sandboxes — the resolved topology

The follow-up question — "clarify the architecture when we have distributed agents in sandboxes" — resolves the fork left open in §4 and §9. Decision: Development A (sandbox launcher) and CB-308 (per-host federation) compose, not compete — each host runs a fleetd gateway whose launcher spawns agents into that host's local sandboxes. A sandbox is never reached across the network; it is reached by the gateway sitting next to it.

11.1 The one fact that fixes the shape

The bus delivers a turn by herdr keystroke-injection — Injector → AgentControl.send writes into a PTY that its local herdr owns. The broker moves messages and presence, never keystrokes. So an agent's PTY must live in a herdr that some fleetd instance drives locally: a remote container with no local herdr cannot be injected into. That rules out a central daemon reaching remote PTYs, and collapses the design to a single identity:

"a sandboxed agent on another host" ≡ "a sandbox spawned by that host's gateway."

flowchart TB
    subgraph hostA["HOST A — gateway"]
        mA["main / orchestrator<br/>MCP client → LOCAL gateway"]
        gA["fleetd A<br/>herdr + CompositePeerLauncher<br/>(incl. SandboxLauncher)"]
        cBEa["sandbox: backend<br/>(local container)"]
        cFEa["sandbox: frontend<br/>(local container)"]
        mA --- gA
        gA -->|"spawn (docker/devcontainer)<br/>→ PTY in A's herdr"| cBEa
        gA --> cFEa
    end
    subgraph broker["BROKER (AMQP) — CB-307/308 fabric"]
        inbox["agent.ID.inbox queues"]
        roster["roster.* (federated presence)"]
    end
    subgraph hostB["HOST B — gateway"]
        gB["fleetd B<br/>herdr + SandboxLauncher"]
        cBEb["sandbox: backend<br/>(local container)"]
        gB -->|"spawn → PTY in B's herdr"| cBEb
    end
    gA <-->|"messages + presence<br/>(NOT keystrokes)"| inbox
    gB <-->|"messages + presence"| inbox
    gA --- roster
    gB --- roster
    classDef line fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    class inbox,roster line

Figure 11 — the composed topology. Each gateway owns its local herdr and runs a SandboxLauncher (the §4 adapter) that spawns role-specific containers on its own host; the broker (blue) carries only messages + roster between gateways. Keystroke-injection stays strictly local to each gateway.

11.2 How a delegation reaches a sandboxed agent on another host

sequenceDiagram
    participant MA as main (host A)
    participant GA as gateway A
    participant BR as broker
    participant GB as gateway B
    participant SB as sandbox agent (host B, container)
    MA->>GA: fleet_send(globalId on B, msg)
    GA->>GA: directory lookup - is globalId local? NO
    GA->>BR: publish agent.ID.inbox (durable)
    BR->>GB: route to the owning gateway
    GB->>SB: inject via B's LOCAL herdr (keystrokes)
    Note over GB,SB: SandboxLauncher already spawned the container -<br/>its PTY is in B's herdr, CB-306 readiness passed
    SB-->>GB: fleet_reply (to B's LOCAL MCP endpoint)
    GB->>BR: publish primary-bound (durable, msg id)
    BR->>GA: route back to A
    Note over GA: held until the main pulls (the main is a client)
    MA->>GA: poll / blocking send resolves
    GA-->>MA: reply

Figure 12 — the local ? inject : publish fork (CB-308 §3.2) with a sandboxed far side. Only the middle hop crosses the network via the broker; both injection points (into the sandbox on B, and the drain-nudge back into the main on A) are local herdr writes. This is CB-308's routing rule unchanged — the sandbox is transparent to it.

11.3 Two reachability changes any sandbox forces

Change Today Under sandboxes
mcpUrl http://127.0.0.1:8765/mcp (loopback) must be host-routable from inside the container (e.g. host.docker.internal or the gateway's LAN IP) — the worker connects to its own gateway's MCP, never a remote one.
PTY ownership pane in the daemon's herdr pane is the container's attached PTY, in the local gateway's herdr (via docker exec/devcontainer) — non-negotiable per §11.1.

11.4 Everything maps to an existing seam (nothing new invented)

Concern Provided by
Per-host gateway (owns local herdr + sessions) CB-308 (today's fleetd, evolved)
Spawn into a local sandbox / role→image Development A SandboxLauncher (§4), routed by CompositePeerLauncher
Addressing a remote sandboxed agent CB-308 global id + federated roster (host + role as metadata)
Orphan reap after a gateway restart CB-117 per-gateway, summed by the composite — each reaps only its local herdr
Container up/down tied to CB-303 session lifecycle — SandboxLauncher.stop tears the container down with the pane
Cross-gateway spawn/send trust CB-308 item #5 / CB-401 Stage-C — each gateway edge is a trust boundary

The net: distributed sandboxes add zero new pillars — they are CB-308 gateway × Development-A launcher at every host, with the §7 ownership line (peer owns the image; the bridge only launches into it) holding at each gateway.