Files
fleetd/docs/CB-401-Peer-Launcher-SPI.md
T
Dai Ha ecc590f344 CB-632 unit 5: rename the daemon and classes in the docs prose
Part of #145 (CB-632). Documentation only, plus one internal literal.

Unit 1 renamed the package and classes, which left every doc describing
classes that no longer exist. This fixes the prose across README.md,
docs/ and bridged/docs/ -- 18 files.

Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and
"bridged" where it names the daemon as a product rather than a path.

Also renamed two literals, because a doc that disagrees with the code is
worse than one that is out of date:

  - bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey
    OpenCodeLauncher sends when a profile resolves no token, to a local
    endpoint that does not check it. No test asserts the old string.
  - the vnd.ltms.bridged.* media type in the M4 design doc. It appears
    in no Java file, so nothing implements it yet.

Deliberately NOT renamed, because each is still literally true today and
changes only at the cutover:

  - paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar,
    .bridged-worktrees, deploy/dev.ltms.bridged.plist,
    scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh
  - bridged_* metric names -- renaming these after the monitoring is
    wired would break dashboard continuity, so they move before it is
  - bridge_* MCP tool names, which answer alongside fleet_* on purpose
  - BRIDGED_* environment variables, read by a file outside this repo

Method note: perl, not sed. BSD sed has no \b and no lookaround, and a
word-boundary expression there fails silently. The prose replace uses
(?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an
identifier, then every remaining hit was read by hand.

Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
2026-08-23 06:46:34 +02:00

10 KiB

CB-401 — Peer Launcher SPI (Stage 4: pluggable peers)

Status: design note (feature branch feature/peer-launcher-spi) Stage: 4 — makes the bridge scalable to heterogeneous peers (Claude Code, Codex, …) without the core learning any one peer's environment.

1. Why

claude-bridge is a communication bus between heterogeneous AI agents — its stable surface is the protocol (fleet_spawn / send / poll / reply / ask / list / stop / status), and that surface should stay provider-neutral. Today the daemon can only materialize one kind of peer: an off-subscription Claude Code CLI over herdr. Everything specific to how that peer is set up (ANTHROPIC_BASE_URL, the subscription guard, --mcp-config/system-prompt flags, claude-* naming, git-token injection) is baked directly into the core spawn path.

The goal of CB-401 is a seam — a PeerLauncher SPI — so that "how to bring a peer of kind X to life" lives in a swappable adapter the bus delegates to, while the bus itself owns only transport, session/turn lifecycle, and routing. This both unlocks a second peer kind (Codex, a human, another Claude) and retroactively gives the CB-301-ext / CB-302 environment features a principled home (an adapter) instead of sitting in core.

Non-goal for CB-401: dynamic/external plugin loading (arbitrary jars). That is Stage C and carries a security model of its own (§7). CB-401 delivers a first-party, in-tree, config-selected SPI with exactly one implementation, proving the seam.

2. The boundary

flowchart LR
    subgraph core["BRIDGE CORE — provider-neutral"]
        proto["Protocol verbs<br/>spawn / send / poll / reply / ask / list / stop"]
        sm["SessionManager<br/>FSM · registry · roster · lifecycle"]
        msg["Message store / routing"]
        life["Lifecycle limits<br/>idle_ttl · context_cap · drain"]
    end
    subgraph adapters["PEER ADAPTERS — env-specific"]
        cc["ClaudeCodeLauncher<br/>(today's WorkerService)"]
        cx["CodexLauncher<br/>(future, CB-402)"]
        hu["HumanLauncher<br/>(future)"]
    end
    sm -->|"delegates spawn/release"| spi{{"PeerLauncher SPI"}}
    spi --> cc
    spi --> cx
    spi --> hu
    cc -.->|"herdr transport"| herdr["herdr (terminal multiplexer)"]

Figure 1 — the core delegates peer materialization to a launcher chosen by profile; the core never learns a peer's env.

3. Coupling audit (as-built, main @ 0efb65c)

Where Claude/herdr specifics actually live today:

Concern Location Verdict
ANTHROPIC_BASE_URL / ANTHROPIC_MODEL / CLAUDE_CONFIG_DIR / ANTHROPIC_AUTH_TOKEN env WorkerService.spawn → adapter
SubscriptionGuard.assertWorker(baseUrl) (billing boundary) WorkerService.spawn → guard → adapter (it guards an ANTHROPIC_* concept)
--mcp-config + --append-system-prompt REPLY_CHARTER (Claude Code CLI flags) WorkerService.argvWithBridge → adapter
claude-<profile>-<nonce>-<seq> naming, WORKER_NAME regex, orphan reap (CB-117) WorkerService → adapter (naming is a herdr-label detail)
GITEA_TOKEN / GITEA_HOST injection (CB-302 checkpoint) WorkerService.spawn → adapter + a capability (§6)
tab/pane placement, worker space, tab labels WorkerService.spawnInTab/spawnAsPane via herdr WorkspaceControl → adapter (herdr transport detail)
FleetConfig.Worker profile shape (baseUrl, model, configDir, …) config mostly adapter-shaped — see §5
FSM, registry, roster, reapIdle/drainAll/contextCap, rosterView SessionManager stays core
turn/completion detection (TurnListener, CompletionResolver, StatusPoller, WorkerPresence) inject/ stays core, but reads herdr terminal output → transport-coupled (§4b)
message store & routing msg/ stays core
Agent / paneId handle herdr/ generalize — see §4a

Finding: the extraction is tractable because ~90% of the coupling is already funnelled through one class (WorkerService). Renaming/adapting it to ClaudeCodeLauncher implements PeerLauncher and having SessionManager depend on the interface is the bulk of Stage A.

4. Two friction points

4a. paneId is a herdr handle, not a peer-neutral id

SessionManager keys its registry by paneId, MCP/REST route by paneId, and WorkerSession stores it. paneId is a herdr pane handle — meaningless for a peer that isn't a herdr pane.

Decision: introduce an opaque PeerHandle the launcher returns. It carries a launcher-assigned id (the registry/routing key) plus launcher-private coordinates (for herdr: paneId, tabId, terminalId). Stage A keeps id == paneId for the Claude adapter so nothing downstream changes value, but the type stops being "a herdr pane" — the core routes on PeerHandle.id().

classDiagram
    class PeerLauncher {
        <<interface>>
        +Set~Capability~ capabilities()
        +PeerHandle spawn(SpawnRequest req)
        +void release(PeerHandle h)
        +String effectiveCwd(SpawnRequest req)
        +List~String~ parityOverlay(String profile)
        +int reapOrphans()
    }
    class PeerHandle {
        <<interface>>
        +String id()
    }
    class ClaudeCodeLauncher {
        herdr AgentControl/WorkspaceControl
        SubscriptionGuard
    }
    PeerLauncher <|.. ClaudeCodeLauncher
    ClaudeCodeLauncher ..> PeerHandle : returns

Figure 2 — the SPI the core sees. ClaudeCodeLauncher is today's WorkerService, adapted.

4b. Turn/completion detection reads herdr output

inject/ (turn listener, completion resolver, status poller, presence) infers turn boundaries from herdr terminal scraping. That is genuinely peer-transport-specific — a Codex peer would signal turns differently. For CB-401 this stays in core (it's the Claude/herdr transport's detector), but §6's capability model is what lets a future non-herdr peer bring its own turn-signalling without the core assuming terminal scraping. Out of scope for Stage A; noted so the SPI doesn't accidentally hard-wire "turns come from herdr".

5. Config shape

FleetConfig.Worker is Claude-shaped (baseUrl, model, configDir, tokenEnv). Rather than break existing YAML, CB-401 keeps workers: exactly as-is and treats those fields as the ClaudeCodeLauncher's profile schema. A future peer kind adds a kind: discriminator (default "claude-code") selecting the launcher; unknown-kind → clear config error. No migration of existing configs. (Jackson already ignores unknown keys, so adding kind is backward-safe.)

sequenceDiagram
    participant MCP as fleet_spawn (MCP/REST)
    participant SM as SessionManager
    participant L as PeerLauncher (by profile.kind)
    participant T as transport (herdr)
    MCP->>SM: acquire(profile, cwd, owner)
    SM->>L: spawn(SpawnRequest)
    L->>L: build env + guard + argv (adapter-private)
    L->>T: start(name, argv, env, cwd)
    T-->>L: handle (paneId…)
    L-->>SM: PeerHandle(id)
    SM->>SM: register session keyed by handle.id()
    SM-->>MCP: session view

Figure 3 — spawn delegation. The core's acquire is unchanged in shape; only the thing it calls becomes an interface.

6. Capabilities

Peers are not uniform. Let each launcher declare a capability set; the protocol is the union and degrades gracefully when a launcher lacks one:

Capability Meaning Claude Code Codex (likely) Human
MID_TURN_ASK supports fleet_ask rendezvous ✓ ? ✗
SELF_PR can open its own PR at checkpoint (CB-302) ✓ (opt-in token) ? ✗
WORKTREE can run in a provisioned git worktree ✓ ✓ ✗
ORPHAN_REAP spawner can reconcile orphaned peers on boot ✓ ? ✗

A verb invoked against a peer that lacks the capability returns a clean "unsupported for this peer" rather than a crash. This keeps the protocol honest as peers diversify and prevents the core from assuming "every peer is a Claude in a worktree" (the drift signal from the identity note).

7. Staging & the Stage-C security gate

flowchart TD
    A["Stage A — CB-401<br/>extract PeerLauncher SPI in-tree<br/>ClaudeCodeLauncher = adapted WorkerService<br/>one impl, config-selected"] --> B["Stage B — CB-402+<br/>2nd in-tree adapter (Codex/human)<br/>proves the SPI held"]
    B --> C["Stage C<br/>dynamic external plugin loading<br/>ServiceLoader / jar discovery"]
    C -.requires.-> G["Trust & capability model<br/>what env/tokens a plugin may inject"]
    classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff;
    class G gate

Figure 4 — deliver A now; B when a real second peer exists; C only if third parties must ship adapters, and only behind a trust model.

Security note (Stage C, not now): a launcher runs at daemon privilege and touches process spawning and env/token injection into peers — the most sensitive path in the system. "Extra plugins" must mean first-party, in-tree, config-selected for the foreseeable future. A third-party plugin that can inject env into a peer needs a genuine trust/capability model before it may exist. Regardless of stage, the primary remains the merge/verify gate — self-reports over the bus are messages, not verified facts.

8. Stage A scope (this branch, delegation-ready)

Deliverable for CB-401 Stage A — mechanical, behaviour-preserving:

  1. PeerLauncher interface + PeerHandle (opaque id) + SpawnRequest (profile, requestedCwd, callerCwd) + Capability enum, new package dev.ltms.fleet.peer.
  2. ClaudeCodeLauncher implements PeerLauncher = today's WorkerService, adapted: spawn(...) returns a PeerHandle (id = paneId), capabilities() declares MID_TURN_ASK, SELF_PR(when token), WORKTREE, ORPHAN_REAP.
  3. SessionManager depends on PeerLauncher, not WorkerService concretely; routing keys on PeerHandle.id() (== paneId today, so zero value change).
  4. Fleetd.main wires the concrete ClaudeCodeLauncher behind the interface.
  5. No behaviour change, no config change. Full green gate: ide_sync → ide_diagnostics (0 errors/0 warnings) → mvn clean install with MVN_EXIT captured (no masking pipe). All existing tests pass unchanged; add tests only for the new PeerHandle indirection.

Explicitly out of scope for Stage A: kind: config discriminator, any second adapter, capability enforcement at the verb layer (declare only), touching inject/ turn detection, dynamic loading.