From 06cceeee7cecbfaa720b7dd53276a5c469f49d59 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Mon, 31 Aug 2026 10:29:05 +0700 Subject: [PATCH] #168: rebuild chapters 1, 2, 7, 8 and 9 against the source MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The audit marked all five REBUILD. Every factual claim on them is now checked in the code and carries a file:line reference. What was wrong and is now fixed: - Ch.1 named a `fleet_read` tool and an SSE `GET /events` route. Neither exists. It also named Redis Streams and NATS JetStream as the queue; the shipped inbox is AMQP. The subscription boundary is back as its own section, sourced from SubscriptionGuard, and the REST list now matches FleetApp.build(). - Ch.2 described tool parameters that were never shipped. The tool table now comes from each tool's own schema method. - Ch.7 was built on `ccs` profiles and on send parameters that do not exist. Every flow now uses the real tools. The portable CLAUDE.md block is unchanged, byte for byte, and the sync check still passes. - Ch.8 presented old plans as the current stack. It is now a delivery record in four states, and "built, not switched on" means no host enables it today — AMQP and the coordinator mailbox are both on here, so both moved to live. - Ch.9 had drifted from the source in its package, class and endpoint map. Also: chapters 1, 2 and 8 had "I checked this in the code" written on the page itself. That belongs in a worker's report, not in a reference page. The pages now state the fact and cite the line. Every Mermaid diagram was rendered with mmdc before this commit. --- 1-Architecture.md | 469 +++++++++++--------------- 2-Message-Server.md | 806 ++++++++++---------------------------------- 7-Use-Cases.md | 489 ++++++++++++++------------- 8-Roadmap.md | 743 +++++++++------------------------------- 9-Implementation.md | 572 ++++++++++++++++--------------- 5 files changed, 1079 insertions(+), 2000 deletions(-) diff --git a/1-Architecture.md b/1-Architecture.md index 265bdfe..8db089f 100644 --- a/1-Architecture.md +++ b/1-Architecture.md @@ -1,322 +1,233 @@ # 1. Architecture -`claude-bridge` lets a **primary** Claude Code session — Opus 4.8 on your Pro/Max -subscription — drive one or more **secondary worker** Claude Code sessions running a -*different, cheaper/local* model, **without ever putting a proxy on the primary**. A -standalone daemon, **`fleetd`**, sits between them: it drives [herdr](https://herdr.dev) (an -agent multiplexer) to inject turns and read live agent-status, and exposes an **MCP server** -that every Claude session mounts. +**fleet** lets a **primary** Claude Code session — the one on your own subscription — drive +one or more **members**: separate Claude Code or OpenCode sessions that can run a different, +cheaper or local model. A standalone daemon, **`fleetd`**, sits between them. It drives +[herdr](https://herdr.dev), a terminal multiplexer, and exposes an **MCP** server that every +session mounts. MCP (Model Context Protocol) is the client contract every session uses to +talk to `fleetd`; see [Message Server](2-Message-Server) for the full tool schema. -Two invariants define the whole design; everything else follows from them. +A member is either a **worker** (does assigned work, replies, never delegates) or an +**architect** (can delegate to workers but never spawns or tears one down). `fleetd` resolves +its own name as `fleet` to every MCP client +(`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:314`). -> **1 — Subscription boundary.** Only a *worker* process ever sets `ANTHROPIC_BASE_URL`; the -> primary never does, so it stays on Pro/Max. `fleetd` is not a `claude` process and holds -> zero Anthropic quota. -> -> **2 — Sole gateway.** Every Claude session talks *only* to `fleetd`, over the MCP tools it -> mounts. No session addresses a broker, a peer session, or the network directly. +One invariant and two facts explain almost everything else in this page. They are easy to miss +because the old design docs got them wrong. + +## The subscription boundary — why fleet exists at all + +This is the whole reason `fleetd` exists. The primary stays on the operator's subscription. A +worker can run on a different, often cheaper, backend. There is no proxy on the primary, so +there is no accidental bill. `fleetd` enforces this with one class, `SubscriptionGuard` +(`fleetd/src/main/java/dev/ltms/fleet/guard/SubscriptionGuard.java:9-21`, its class-level +javadoc): + +- **A worker must egress to an off-subscription host.** Before `fleetd` spawns a worker, it + calls `guard.assertWorker(baseUrl)`. This checks the profile's `ANTHROPIC_BASE_URL` against a + configured allowlist of hostnames. It throws if the URL is missing, cannot be parsed, or is not + on that list (`SubscriptionGuard.java:32-45`). The allowlist is a plain config value — an empty + list unless the operator sets one — read from `guard.offSubscriptionHosts` in `fleetd.yaml` + (`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:1153-1163`, the `Guard` record). + The configured hostnames differ per host, so this page names none. +- **The primary must never carry `ANTHROPIC_BASE_URL`.** `fleetd` checks its own startup + environment against this rule once, at boot, and refuses to start if the variable is set + (`SubscriptionGuard.java:49-54`, called from `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:131-132`). + This guards `fleetd`'s own process, not a worker's. The primary Claude session's own + environment is a separate concern, upstream of `fleetd`. +- **A profile may opt out on purpose, with `subscription: true`.** Some backends have no + off-subscription endpoint to point at. For those, a profile can be marked `subscription: true` + to run deliberately on the operator's subscription instead. When set, the launcher writes no + `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN` for that worker at all, and skips the guard's + `assertWorker` check for that profile only. Every other profile keeps the hard refusal + (`FleetConfig.java:271-280`, the `subscription` field's javadoc; enforced at + `fleetd/src/main/java/dev/ltms/fleet/member/ClaudeCodeLauncher.java:189-234`). Setting both `subscription: true` and a `baseUrl` on the same profile is a + contradiction, and `fleetd` refuses it loudly at two different points: once at spawn time, as an `IllegalStateException` that names the profile + (`ClaudeCodeLauncher.java:198-205`), and once at config load time, if the profile's `env:` + block tries to smuggle `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN` back in past the skipped + guard (`FleetConfig.java:2040-2070`, `validateSubscriptionProfiles()`). + +## Fact 1 — fleetd never forks a member process + +`fleetd` does not run `claude` or `opencode` itself, and it does not fork a child process for +a member. Spawning a member is two herdr calls over a Unix domain socket: + +1. `tab.create` (or `pane.split`), carrying the member's working directory and its environment + map — this is where a worker's `ANTHROPIC_BASE_URL` and token are set + (`fleetd/src/main/java/dev/ltms/fleet/herdr/WorkspaceControl.java:84-94`). +2. `agent.start`, carrying the agent's name, kind (`claude` or `opencode`), and its argument + list, targeting the pane just created + (`fleetd/src/main/java/dev/ltms/fleet/herdr/AgentControl.java:101-108`). + +herdr is the one that forks the pseudo-terminal (PTY) and starts the process, and it does so +as **its own OS user** — the user herdr itself runs as, not `fleetd`'s user +(`fleetd/src/main/java/dev/ltms/fleet/herdr/UnixSocketHerdrClient.java:1-30` — the client +connects over a Unix domain socket, one connection per call). + +This single fact explains three things that would otherwise look arbitrary: + +- **The credential model.** A member's environment comes from the `env` map `fleetd` hands to + herdr at `tab.create`, not from a process `fleetd` controls directly. Whatever the shell that + seeds that pane already has access to, the member can reach too — this is why the credential + scrub work in the project's history exists at all. +- **Why `fleetd` and its members share one filesystem.** `fleetd`, herdr, and every member pane + must sit on the same machine and the same filesystem, because the Unix socket and the + worktree paths `fleetd` hands to herdr are plain local paths, not a network protocol. +- **Why running a member as a different OS user needs a second herdr daemon.** herdr forks as + its own user, so the only way to change a member's user is to run a second herdr process + under that other user and point `fleetd` at its socket. `fleetd` supports this today as + `memberHerdrSocket` — a second herdr socket for members, separate from the lead's socket + (`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:36,84`; + wired up in `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:153-154`). With no + `memberHerdrSocket` configured, the lead and its members share the one herdr daemon. + +## Fact 2 — identity comes from the connection, never an argument + +When a session calls an MCP tool, `fleetd` does not read who is calling from a parameter in +the call. It resolves the caller from the network connection itself: the loopback peer PID +(from the operating system) mapped to a herdr pane (from herdr), or the primary if no pane +matches (`fleetd/src/main/java/dev/ltms/fleet/mcp/ConnectionIdentity.java:16-48`). `fleetd` +builds this resolved identity into every MCP call's context before the tool handler ever runs +(`FleetMcp.java:157-172`). + +This is why a worker can never reply as another worker, or as the primary: `fleet_reply`'s +caller is always the connection it arrived on, never a field in the request +(`FleetMcp.java:215` — "`fleet_reply`'s identity is the CONNECTION, never an argument"). The +resolved identity is one of three roles — `primary`, `worker`, or `architect` — reported by +`fleet_whoami` and readable in code as `Principal.Role` +(`fleetd/src/main/java/dev/ltms/fleet/auth/Principal.java:16-95`). + +`Authz` is the table that decides what each role may do, checked on every MCP call and every +REST call (`fleetd/src/main/java/dev/ltms/fleet/auth/Authz.java:44-72`): + +| Action | Who may do it | Why | +|---|---|---| +| `SPAWN`, `STOP`, `DRAIN` | primary only | fleet lifecycle is the primary's alone; an architect coordinates members but never stands them up or tears them down | +| `SEND` | primary or architect | delegating a turn; a worker sending would be it escalating its own role | +| `REPLY`, `ASK` | only the caller that owns that session | a caller may act only as the pane it occupies — this is what stops one member forging a reply for another | +| `READ`, `METRICS` | primary, worker, or architect | read-only status carries no secret and is open to every authenticated role | ## System overview ```mermaid flowchart TB - subgraph herd["herdr — agent multiplexer (Claude sessions run as panes)"] - PP["primary pane · Opus 4.8
env CLEAN · MCP client"] - WP["worker pane(s) · claude
ANTHROPIC_BASE_URL set · MCP client"] + subgraph herd["herdr — owns the PTYs and panes"] + PP["primary pane
MCP client"] + WP["worker pane(s)
MCP client"] + AP["architect pane(s)
MCP client"] end - subgraph BD["fleetd — standalone daemon · THE gateway (no Anthropic quota)"] - SRV["SERVER face
MCP server · REST/SSE · policy brain"] - CLI["CLIENT face
status-gated injector · herdr socket client"] - SRV --> CLI + subgraph FD["fleetd — the daemon, no Anthropic quota of its own"] + MCP["MCP server (FleetMcp)
+ REST surface (FleetApp)"] + ID["ConnectionIdentity + Authz
who is calling, what they may do"] + MSG["MessageService
rendezvous + status-gated injector"] + MCP --> ID + ID --> MSG end - Q["broker / queue
(fleetd-owned · below the gateway)"] - M["worker model
ollama.ltms.dev · GX10 vLLM"] + BR["broker (optional)
AMQP — LavinMQ in production"] - PP -->|"MCP tools"| SRV - WP -->|"MCP tools"| SRV - CLI -->|"Unix socket · send_text · events.subscribe"| herd - SRV -.->|"durability · cross-host (internal)"| Q - WP -->|"inference"| M - - classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff; - classDef core fill:#2f855a,stroke:#22543d,color:#ffffff; - classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff; - class PP ext - class SRV,CLI core - class Q,M warn + PP -->|"fleet_send / fleet_spawn / fleet_status"| MCP + WP -->|"fleet_reply / fleet_ask"| MCP + AP -->|"fleet_send / fleet_reply"| MCP + MSG -->|"Unix domain socket
tab.create, agent.start, agent.prompt"| herd + MSG -.->|"durable reply inbox
lead-to-lead coordination"| BR ``` -*Figure: the Claude sessions are **herdr panes**; they reach *up* into `fleetd`'s SERVER face -over MCP, while `fleetd`'s CLIENT face drives them *down* through herdr's socket. `fleetd` is -the only thing any session connects to — the broker sits below the gateway line, owned by -`fleetd`, touched by no session. Only worker panes carry `ANTHROPIC_BASE_URL`.* - -## The two invariants - -### Subscription boundary — why the bridge exists - -The point of the bridge is to keep the **primary** on Pro/Max while a **worker** runs a -cheaper/local model, with no policy-violating proxy on the primary. - -- The **primary** `claude` **never** sets `ANTHROPIC_BASE_URL`. It authenticates to - `api.anthropic.com` on your subscription and reaches the worker *only* through `fleetd`'s - MCP tools — never by re-pointing its own endpoint. -- Only a **worker** `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev` - (+ `ANTHROPIC_AUTH_TOKEN`) or a GX10 vLLM URL. Because model choice is just per-process env, - this also sidesteps Claude Code's lack of per-subagent provider routing — the worker is a - *separate process*, not a subagent. -- **`fleetd` holds no quota**, so it may hold a permanent herdr event subscription and poll - its own internal queue with no policy concern. Its **subscription guard** refuses to spawn a - worker pane whose resolved `ANTHROPIC_BASE_URL` host isn't on an off-subscription allowlist, - and refuses to ever set that var on a pane tagged *primary*. (The guard validates the - resolved host and confirms egress post-spawn — a substring check on the launch string is not - enough; see [Message Server](2-Message-Server).) - -> **Rule:** anything that sets `ANTHROPIC_BASE_URL` is, by definition, a worker. If you are -> ever tempted to set it on the primary, stop — that is the subscription line. - -### Sole gateway — how sessions communicate - -Because every session mounts `fleetd` over MCP, `fleetd` is the single chokepoint for *all* -agent traffic. This is a deliberate simplification with real payoffs: - -- **No session-side transport.** A Claude session calls MCP tools and nothing else — no - `curl`, no broker client, no hand-rolled hook polling a queue. There is no shell step that - could leak env, so mounting the bridge is subscription-safe by construction. -- **Async is a push, not a poll.** When a message arrives for an idle session, `fleetd` - **injects it into that session's pane over herdr**, gated on the live `agent_status`. The - session is never asked to busy-poll anything; the old "primary must never perpetual-poll" - footgun disappears because there is nothing for it to poll. -- **Guardrails are centrally enforced.** Round/turn budgets (anti ping-pong), rate limits, - per-session authz, and audit all live in `fleetd` — one place — instead of cooperative - sentinels each session must honour. -- **The broker is infrastructure, below the line.** If `fleetd` needs durability or a host - hop it owns a broker/queue for that. No Claude session sees it. +*Figure: every session — primary, worker, or architect — reaches `fleetd` only through its MCP +tools. `fleetd` resolves the caller's identity from the connection, checks it against `Authz`, +then either answers directly or drives herdr over the Unix socket described in Fact 1. The +broker is optional infrastructure `fleetd` owns; no session ever talks to it directly.* ## Components | Component | Role | Notes | |---|---|---| -| **`fleetd` — SERVER face** | The gateway: **MCP server** (the contract every session mounts) + REST/SSE for non-Claude clients, over the **policy brain** — session tracker, subscription guard, reply rendezvous. | `fleet_send` · `fleet_reply` · `fleet_ask` · `fleet_status` · `fleet_spawn` · `fleet_list` · `fleet_stop` · `fleet_read`. | -| **`fleetd` — CLIENT face** | Drives herdr: a **status-gated injector** (per-pane FIFO, delivers only when `agent_status ∈ {idle, blocked}`) over a **herdr socket client** (NDJSON, id-correlated, live event stream). | Single writer per pane → no injector-vs-injector races. | -| **herdr** | Agent multiplexer. Owns the PTYs, panes, persistence, and — crucially — **`agent_status_changed` events**. Claude sessions run here as panes. | Socket is **local-only**; young/single-dev → injector kept pluggable. | -| **Worker `claude`** | A *real* Claude Code process (inherits `CLAUDE.md`, hooks, skills, MCP), pointed at a different model. Recyclable, not immortal. | Only these carry `ANTHROPIC_BASE_URL`. Materialized by a swappable `PeerLauncher` adapter (`ClaudeCodeLauncher` today); the core bus is peer-neutral. | -| **Broker / queue** *(optional, internal)* | `fleetd`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `fleetd` will later inject. | Redis Streams / NATS JetStream, or an embedded queue for a single host. | -| **AgentAPI** *(fallback)* | Swappable injector behind the CLIENT face if herdr is unavailable. | Screen-stability heuristic instead of events — see [Approaches](3-Approaches). | +| **`fleetd`** | The daemon: an MCP server (`FleetMcp`) plus a REST surface (`FleetApp`) over one shared `MessageService`, so both are validated by parity rather than by re-implementing behaviour. Holds no Anthropic subscription quota itself. | Registered MCP server name is `fleet` (`FleetMcp.java:314`). Listens on `127.0.0.1:8765` by default (`FleetConfig.java:183-186`). | +| **MCP tools** | The full registered set, all on `FleetMcp`: `fleet_send`, `fleet_reply`, `fleet_ask`, `fleet_status`, `fleet_poll`, `fleet_ack`, `fleet_spawn`, `fleet_list`, `fleet_stop`, `fleet_profiles`, `fleet_whoami`. | `FleetMcp.java:301-327`. There is no `fleet_read` tool — an earlier design doc named one, but it was never built. | +| **REST routes** | `/healthz`, `/metrics`, `/sessions`, `/agents`, `/members` (GET/POST), `/members/{paneId}` (DELETE), `/profiles`, `/sessions/{id}/message`, `/sessions/{id}/reply`, `/sessions/{id}/replies`, `/sessions/{id}/ask`, `/sessions/{id}/status`, `/tasks/{ticket}`, plus the `/mcp` servlet. | `FleetApp.java:143-158`, `build()`, read in full. There is **no `/events` route** — an earlier design doc described an SSE (`GET /events`) status stream; it was never built. | +| **herdr** | An external terminal multiplexer. Owns the PTYs, the panes, and the live `agent_status` events. Forks every member process as its own OS user, over its own Unix socket. | Socket is local-only per Fact 1 above. | +| **Worker** | A member that does the assigned work and replies. Runs as a real `claude` or `opencode` process, inheriting the project's `CLAUDE.md`, hooks, skills and MCP mounts — never orchestrates or spawns. | Launcher kind is `claude-code` (default) or `opencode`, set per profile (`FleetConfig.java:265-266`, `333`, `335`). | +| **Architect** | A member that can delegate to workers (`fleet_send`) but has no lifecycle rights — it cannot `fleet_spawn` or `fleet_stop` a peer. | `Authz.java:53-58`. | +| **Broker (optional)** | An external AMQP broker `fleetd` owns for durable reply delivery across a restart, and for lead-to-lead coordination across daemons. Its presence swaps the in-memory reply inbox for the AMQP-backed one; absent, `fleetd` stays soft-state. | `FleetConfig.java:655-673`. Production default is LavinMQ; a stock RabbitMQ works too, since both speak AMQP 0-9-1. Not Redis Streams or NATS JetStream — an earlier design doc named those; the shipped adapter is AMQP only. | -## Traffic: two modes across the gateway - -Both modes are pure MCP from the Claude side; they differ only in whether a caller waits on -the connection. - -### Mode 1 — blocking request/response (the main path) - -The primary delegates with **one** `fleet_send` tool call. `fleetd` holds it open (the -primary is idle-waiting, spending no quota) and **resolves it on whichever lands first**: the -worker's structured `fleet_reply`, or the worker's turn-done status edge -(`agent_status: working → idle` — herdr has no `done` status). The reply comes -back as the tool result — so worker → primary rides `fleetd`'s state and needs **no keystroke -into the primary pane**, even single-host. +## How a delegation travels ```mermaid sequenceDiagram - participant P as "Primary (Opus) — MCP client" - participant S as "fleetd (gateway)" + autonumber + participant P as "Primary — MCP client" + participant F as "fleetd" participant H as "herdr" - participant W as "Worker (other model)" + participant M as "Member pane (worker or architect)" - P->>S: "fleet_send(worker, task) — tool call PARKS" - S->>S: "await pane agent_status = idle" - S->>H: "pane.send_text + send_keys (inject)" - H->>W: "new turn" - activate W - H-->>S: "event: agent_status = working" - W->>S: "fleet_reply(result) — structured (preferred)" - H-->>S: "event: agent_status working → idle (turn done)" - deactivate W - S-->>P: "tool result = reply (unparks the call)" - Note over P,W: "resolves on fleet_reply or the idle edge — whichever lands first.
a blocked worker returns as the tool result, then the primary re-answers" + P->>F: "fleet_send(sessionId, content)" + F->>F: "resolve caller from the connection, check Authz.SEND" + F->>F: "MessageService opens a rendezvous, injector waits for the pane to go idle" + F->>H: "agent.prompt (inject the turn)" + H->>M: "new turn" + activate M + Note over M: "member works the turn" + M->>F: "fleet_reply(content)" + Note over F: "caller resolved from the connection —
a member can only reply as itself" + H-->>F: "agent_status: working -> idle (turn done)" + deactivate M + F-->>P: "tool result = the reply (unparks the fleet_send call)" ``` -*Figure: a single parked tool call, not a busy-poll. SSE (`GET /events`) carries status to -*observers* (a human, a dashboard) in parallel; the primary never has to hold it. For work that -may outrun a sane request timeout, use Mode 2.* +*Figure: one blocking delegation. `fleet_send` parks until the rendezvous resolves — on the +member's `fleet_reply`, or on the pane going idle again, whichever comes first. A member's +`fleet_spawn` step (the two herdr calls from Fact 1) already happened before this exchange; this +diagram covers only the turn itself. `fleetd` never asks the primary to poll — the reply +comes back as the tool result of the original call.* -> **Late replies are no longer lost (CB-307).** If the worker finishes *after* the primary's -> blocking `fleet_send` has already timed out (or was never opened), its `fleet_reply` finds no -> live waiter to resolve. `fleetd` now **holds that reply in a per-worker inbox** rather than -> discarding it, and the primary collects it later keyed by target (`fleet_poll(target)` / -> `GET /sessions/{id}/replies`). With a broker configured the reply is **durable** across a daemon -> restart (Stage 2, `AmqpReplyInbox` on LavinMQ); without one it is soft-state in-memory. And -> delivery is **active, not just pull** (Stage 3): the moment a reply lands with no open send, -> `fleetd` injects a bounded, status-gated *drain nudge* into a same-host primary's own herdr pane -> — so the primary need not be polling to notice. An off-host / non-herdr primary keeps the pull -> path. See [Roadmap](8-Roadmap#delivery-reliability--multi-host-cb-306--cb-308). +## Two traffic modes: blocking, or fire-and-poll -### Mode 2 — asynchronous delivery (fleetd-mediated) +`fleet_send` has two modes, both handled by the same class, `MessageService` +(`fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java`, class-level javadoc read in +full). -For traffic with no caller waiting — a detached progress report, an out-of-band question, a -webhook injecting work — the recipient must be *woken*. `fleetd` does the waking by -**injecting the recipient's idle pane**, driven by the live `agent_status`, so delivery is -event-driven rather than a timed poll. If durability or a host hop is needed, `fleetd` -enqueues internally first; the recipient still receives by injection. +- **Blocking (the default).** `MessageService.send` delivers the content and blocks the caller + until the member's `fleet_reply` resolves the rendezvous, or the call times out + (`MessageService.java:498-519`, the `send` overloads). This is the diagram above. +- **Fire-and-poll (`wait:false`).** The MCP tool schema for `fleet_send` documents `wait` as + "Block for the reply (default true); false returns a ticket to poll" + (`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1100`, the `sendTool()` schema). + `MessageService.sendAsync` runs that same blocking `send` on a background virtual thread and + returns a ticket immediately (`MessageService.java:710-753`); the caller collects the result + later with `fleet_poll(ticket)`, which reads back `PENDING`, `DONE` (with the reply), or + `FAILED` (`MessageService.java:765-793`, the `poll` method). -```mermaid -sequenceDiagram - participant SRC as "Source: worker fleet_reply/ask · webhook · bus" - participant S as "fleetd (gateway)" - participant Q as "queue (internal)" - participant H as "herdr" - participant R as "Recipient pane (idle Claude)" +**Why the async path exists at all** — stated directly in the code, not something I inferred: +"A caller's MCP client caps a blocking call at ~60s, but a real delegated task runs for minutes" +(`MessageService.java:40-41`, the class javadoc). A blocking `fleet_send` is capped by the +*calling client's own* MCP timeout, not by anything `fleetd` controls — so any task expected to +run longer than about a minute should use `wait:false` and `fleet_poll`, never a long blocking +call. - SRC->>S: "MCP fleet_reply / fleet_ask · or REST ingress" - opt durability / cross-host - S->>Q: "enqueue (ack + visibility timeout)" - end - S->>H: "await recipient agent_status = idle" - H-->>S: "event: idle" - S->>H: "pane.send_text + send_keys (inject)" - H->>R: "new turn = the message" - Note over S,R: "recipient polled nothing — fleetd pushed on the idle edge" -``` +## What this page does not cover -*Figure: `fleetd` mediates async in both directions. Sources reach it over the gateway (MCP -for Claude, REST for external), it optionally parks the message on its internal queue, waits -for the idle event, and injects.* +This page covers the daemon, the MCP and REST surface, the identity model, the subscription +boundary, the broker, and the two traffic modes. Every claim above carries a `file:line` +reference into the source. -> **Where a hook is used.** A `Stop`-hook appears in exactly the spots where neither an MCP -> tool call nor a pane injection can serve — and in **every** case it targets `fleetd`, never -> a broker: -> - a **split-host primary that is not a herdr pane** (e.g. Opus on your Mac) — the one session -> `fleetd` cannot inject into (a *same-host* primary **is** a herdr pane and gets the CB-307 -> Stage-3 drain nudge) — runs a `Stop`-hook that **long-polls `fleetd`** for queued messages; -> and -> - a **non-MCP (herdr-only) worker** may run a `Stop`-hook that **POSTs its reply to -> `fleetd`** at turn end, a structured alternative to scraping the pane (see -> [Message Server](2-Message-Server) → *Reply envelope*). -> -> The gateway invariant holds in every topology: a hook is just a transport adapter to -> `fleetd` for a session MCP/injection can't reach. - -## Worker lifecycle — the Ralph loop - -`fleetd` treats a worker as a **recyclable** resource, not one immortal session: a long-lived -pane fills its context window. On a context/idle cap it checkpoints and respawns fresh, with -continuity carried by **artifacts on disk** (git commits, a `STATE.md`/task file the worker is -told to keep current) — **not** `claude --resume`, which would reload the context you are -trying to shed. - -```mermaid -stateDiagram-v2 - [*] --> Spawning - Spawning --> Ready: "claude prompt detected" - Ready --> Working: "turn injected" - Working --> Blocked: "permission / question" - Blocked --> Working: "fleetd answers (send_input)" - Working --> Ready: "agent_status → idle (turn done)" - Ready --> Recycling: "context / idle cap hit" - Recycling --> Spawning: "state persisted to disk" - Working --> Failed: "pane.exited (crash)" - Failed --> Spawning: "auto-restart + replay unacked" - Ready --> [*]: "drain / shutdown" -``` - -*Figure: `blocked`, the turn-done `working → idle` edge, and `pane.exited` are real herdr -signals, not heuristics — the reason herdr-centric beats screen scraping. Full lifecycle -detail is in [Message Server](2-Message-Server) → *Worker session lifecycle*.* - -## Topologies - -- **Same-host (default).** Primary, `fleetd`, herdr, and workers on one off-subscription box. - Every session is a herdr pane, so `fleetd` can inject *either* direction; a broker is - optional. Simplest to run and the focus of the design. -- **Split-host.** Primary Opus on your Mac; `fleetd` + herdr + workers on the GPU host near - the model. herdr's socket stays local, so the Mac reaches the worker host **only over - `fleetd`'s MCP/HTTP endpoint**. The primary isn't a herdr pane, so async wake-ups use the - `Stop`-hook-polls-`fleetd` path above. Deployment diagrams: [Message Server](2-Message-Server) - → *Deployment model*. -- **Multi-host federation (planned — CB-308).** Not one `fleetd` fronting remote workers, but - **many gateways, one per host**, meshed over a broker. Each host runs its own `fleetd` owning - its local herdr + registry; agents get a **host-unique global id** and a per-agent broker inbox - `agent..inbox` whose *owning* gateway is the sole consumer; a **federated roster** - (soft-state presence on a `roster.*` topic) gives every gateway an eventually-consistent - who/where/status view. Sending is a one-line fork — **local → inject via herdr today; remote → - publish to the broker**, which routes to the owning gateway. Crucially **invariant #2 holds - per host**: a session still talks only to its *local* gateway, so the MCP pull-asymmetry (a - primary is pulled-from, never called-into) survives the network unchanged. This builds directly - on the CB-307 broker fabric; the gating concern is the cross-host **trust model** (the broker - link becomes the security boundary). Design: **`docs/CB-308-Multi-Host-Federation.md`**; - staging in [Roadmap → Delivery reliability & multi-host](8-Roadmap#delivery-reliability--multi-host-cb-306--cb-308). - -```mermaid -flowchart TB - subgraph hostA["Host A"] - gA["gateway = fleetd A
local herdr + registry"] - P["primary (MCP client)"] - P --- gA - end - subgraph hostB["Host B"] - gB["gateway = fleetd B
local herdr + registry"] - WB["worker panes"] - gB --- WB - end - subgraph BR["broker (LavinMQ / AMQP) — CB-307 fabric"] - IN["agent.<id>.inbox queues"] - RO["roster.* presence topic"] - end - gA -->|"remote send → publish"| IN - IN -->|"owning gateway consumes"| gB - gB -->|"reply → publish (durable, msg id)"| IN - IN -->|"held until primary pulls"| gA - gA -->|"announce local agents"| RO - gB -->|"announce local agents"| RO - RO -->|"union view"| gA - RO -->|"union view"| gB - - classDef core fill:#2f855a,stroke:#22543d,color:#ffffff; - classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff; - classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff; - class gA,gB core - class P,WB ext - class IN,RO warn -``` - -*Figure: multi-host is an addressing + routing concern, not a new kind of peer. Each gateway owns -its own herdr and consumes only its own agents' inboxes; the broker routes between them; the -federated roster is soft-state (fleetd owns who/where/status, the broker owns message durability). -No host ever sees another host's terminals.* - -**Security:** `fleetd` is an agent-control surface — `fleet_send` runs arbitrary prompts and -key-passthrough sends raw keystrokes into a live agent. Bind it to `localhost` + SSH tunnel, or -front it with a bearer token + TLS; never expose the port unauthenticated. One `fleetd` is -today **one trust domain** (no per-session authz yet — an open item in -[Message Server](2-Message-Server)). - -## Failure modes & single points of failure - -Making `fleetd` the sole gateway buys a clean model at the cost of a real SPOF. Degrade -deliberately: - -| What dies | Effect | Recovery | -|---|---|---| -| **`fleetd`** | **All** agent comms stop — sync *and* async — since it is the only gateway; in-flight blocking calls error out. | herdr + workers keep running (state on disk / queue). systemd restarts `fleetd`; it re-attaches to existing panes via `workspace.list`/`pane.list` and drains its queue. This restart path is load-bearing — harden it. | -| **herdr** | No pane control; all delivery (sync + async injection) dead. | PTYs die with the *server* (only client detach survives). Respawn workers from persisted state (Ralph loop); replay unacked queue items. | -| **Broker / queue** (internal) | Durability + cross-host async degrade; **same-host async still works** (idle-injection needs no queue). | `fleetd` delivers locally without it; ack + visibility timeout re-deliver on recovery. Nothing silently dropped. | -| **Model endpoint** | Workers stall or error mid-turn. | herdr status shows `working` stuck / `blocked`; `fleetd` times out the blocking call and surfaces the error. | -| **All of the above** | Full bridge outage. | The **primary is never downstream** of any bridge component — it stays fully usable on its own subscription. The bridge is additive, never on the primary's critical path. | +It does not cover the worker lifecycle, the idle reaper, or how a reply is recycled — those are +in [Message Server](2-Message-Server). It also names no hostname on the subscription-boundary +allowlist, because that list is a plain config value and differs per host. ## Related pages -- **[Message Server](2-Message-Server)** — the `fleetd` design in depth: MCP contract, herdr control, reply rendezvous, lifecycle, API, tech stack, milestones. -- **[Approaches](3-Approaches)** — why herdr-centric, and the full transport comparison (AgentAPI, Agent SDK, queue+Stop-hook, tmux). -- **[Team](6-Team)** — a team-lead orchestrating a mixed Claude + local-LLM worker fleet over the same gateway. -- **[Setup](4-Setup)** · **[Operations](5-Operations)** — bring-up and the day-2 runbook. +- **[Message Server](2-Message-Server)** — the `fleetd` design in depth: MCP tool schema, herdr + control, reply rendezvous, worker lifecycle, REST API. +- **[Approaches](3-Approaches)** — why herdr, and the transport alternatives considered and + rejected (including the AgentAPI adapter, which was never built). +- **[Team](6-Team)** — a primary orchestrating a mixed fleet of Claude and OpenCode members. +- **[13 User Guide](13-User-Guide)** — the written procedure: install, run, and what to do + when it breaks. [Setup](4-Setup) and [Operations](5-Operations) are older pages that now + point here. ## Sources - [herdr — socket API](https://herdr.dev/docs/socket-api/) · [agents / state detection](https://herdr.dev/docs/agents/) · [GitHub](https://github.com/ogulcancelik/herdr) -- [coder/agentapi](https://github.com/coder/agentapi) — fallback injector -- [Model Context Protocol](https://modelcontextprotocol.io) — the client contract both sessions mount -- [Hooks reference — Claude Code Docs](https://code.claude.com/docs/en/hooks) — the split-host `Stop`-hook adapter +- [Model Context Protocol](https://modelcontextprotocol.io) — the client contract every session mounts +- [Hooks reference — Claude Code Docs](https://code.claude.com/docs/en/hooks) diff --git a/2-Message-Server.md b/2-Message-Server.md index dc09b74..35ad20d 100644 --- a/2-Message-Server.md +++ b/2-Message-Server.md @@ -1,702 +1,250 @@ -# 2. Herdr Message Server (`fleetd`) +# 2. Message Server (`fleetd`) -> **Status:** 🟢 Proposed primary approach (2026-07-11) — supersedes AgentAPI as the -> centric transport. AgentAPI is retained only as a *fallback injector* (see [Approaches](3-Approaches)). +`fleetd` is the always-on daemon that runs the fleet. It controls +[herdr](https://herdr.dev), an agent multiplexer, over herdr's Unix-socket API. It gives every +Claude Code session — the **lead** (the orchestrating session, called "primary" in the code) +and every **member** (a spawned worker) — one clean way to send tasks and get replies back. -`fleetd` is a small, always-on **message server that controls [herdr](https://herdr.dev)** -and exposes a clean 2-way messaging API between a **primary** Claude Code session (Opus 4.8, -on Pro/Max) and one or more **secondary worker** sessions running a different/cheaper model. -It replaces the hand-rolled terminal emulation of AgentAPI by standing on herdr's structured -socket API: herdr owns the PTYs, multiplexing, persistence, and — crucially — **agent-status -events**; `fleetd` owns the *policy* (subscription boundary, session lifecycle, delivery -gating) and the *client-facing contract*. That contract is an **MCP server that both the -primary and the workers mount** (one unified Claude setup — see -[The client contract — MCP](#the-client-contract--mcp-unified-for-primary--workers)), with -REST/SSE retained for non-Claude clients and an optional broker for async/cross-host duplex. +`fleetd` has two faces: -## Why herdr-centric (vs AgentAPI) +- a **SERVER face** — an **MCP server** at `/mcp`, plus a REST API, sitting over the policy + layer (session tracking, the subscription guard, and the reply rendezvous); and +- a **CLIENT face** — a herdr socket client that starts members, sends text into their panes, + and reads their live status. -AgentAPI re-implements, per process, an in-memory terminal emulator and a *screen-stability -heuristic* to guess when the agent is done. herdr already provides all of that as a -persistent service, and adds three things AgentAPI cannot: +The two faces are separate in the code. The MCP server is built and mounted in +`dev.ltms.fleet.mcp.FleetMcp` (`FleetMcp.java:150-174`, `FleetMcp.java:461-463`), the REST +routes are built in `dev.ltms.fleet.rest.FleetApp.build()` (`FleetApp.java:126-160`), and both +sit on top of the same `dev.ltms.fleet.msg.MessageService` (`FleetMcp.java:127`, +`FleetApp.java:62`). -| Capability | AgentAPI | herdr (via `fleetd`) | -|---|---|---| -| Inject a turn into a **worker** | ✅ terminal emulation | ✅ `pane.send_text` + `pane.send_keys` | -| Inject a turn into the **primary** | ❌ (only wraps worker) | ◐ same primitive — **only when the primary is a herdr pane** (single-host); split-host wakes via the primary's `Stop`-hook polling `fleetd` | -| "Done / blocked" signal | ⚠ screen-stability heuristic | ✅ `events.subscribe(pane.agent_status_changed)` | -| Worker self-reports state | ❌ | ✅ `pane.report_agent` (via herdr `SKILL.md`) | -| Multiplex a *herd* of workers + attach/observe | ❌ one server per session | ✅ native workspaces/tabs/panes | -| Persistence / detach-reattach over SSH | ❌ | ✅ headless server | +> **A note on this page's sources.** Everything below with a `file:line` reference was read +> directly from the code in this worktree. `docs/MCP-Contract.md` is **not** a normative +> reference for tool or route names — only its §6 (the rendezvous flows) is current; the rest +> is a pre-build design document whose names never caught up with the shipped code. -### How the primary actually consumes a reply (blocking call vs return-and-reinvoke) +## Why herdr, not a hand-rolled terminal reader -An earlier draft claimed "no synchronous call-and-await." That over-stated it. The precise -model has **two** shapes, and the distinction is *cross-turn busy-polling* (forbidden), not -*blocking* (fine): - -- **Blocking request/response (default, short/medium tasks).** The primary issues **one** - MCP tool call — `fleet_send(target, task)` (blocking by default) — and `fleetd` **holds it - open** until it observes the turn-done edge (`agent_status: working → idle` — herdr has no - `done` status) or a worker `fleet_reply`, then returns - the collected reply as the tool result. From the primary's view this is a single tool call - parked on a result, exactly like any long-running `Bash` command: it consumes **no** - Anthropic quota (the primary isn't looping, it's idle-waiting) and doesn't freeze anything - the primary needs. SSE (`GET /events`) is a *parallel observer channel* for - humans/dashboards — the primary never has to hold it. -- **Return-and-reinvoke (long/detached/async tasks).** When a task may outrun a sane request - timeout, or is fire-and-forget, the primary's call returns immediately and the reply comes - back later **through `fleetd`** — `fleetd` injects it into the primary's idle pane over - herdr (same-host), or a split-host primary's `Stop`-hook long-polls **`fleetd`** for it. In - neither case does the primary touch a broker: if `fleetd` needs durability it queues the - message internally (below the gateway) and still delivers by the same route. This is the - correct shape for work that outlives a connection. - -What `fleetd` does **not** offer is a *held-open bidirectional conversation* — each exchange -is one request in, one reply out. That is a feature for a subscription-safe bridge, not a -limitation. See **Trade-offs** below. - -## The client contract — MCP (unified for primary + workers) - -Both Claude sessions — the primary Opus **and** every worker — reach `fleetd` the same way: -they **mount `fleetd` as an MCP server**. Claude Code speaks MCP natively, so the bridge -becomes a set of first-class tools instead of a `curl` the model must be told to run. One -config line, identical on both sides, wires the whole mesh: - -```bash -claude mcp add --transport http bridge http://127.0.0.1:8080/mcp -# or a project .mcp.json / CLAUDE.md entry that every session on the host inherits -``` - -That is the point of the MCP SERVER face: **one unified Claude setup**. Primary and workers -load the *same* server and differ only in which tools they call — no shell step that could -leak env, no hand-rolled HTTP client, no per-session bespoke wiring. REST/SSE (below) stays -for *non-Claude* callers (webhooks, dashboards, a human CLI); Claude ↔ Claude goes over MCP. - -### Tool surface - -| Caller | Tool | Blocks? | Does | -|---|---|---|---| -| **Primary** | `fleet_send(message, target?, {block, timeout_seconds, auto_spawn, turn_id})` | `block:true` (default) → yes · `block:false` → no | Deliver a turn to a worker. Blocking form returns the outcome as the tool result (`reply` \| `question` \| `turn_done` \| `timeout`); detached form returns a `dispatch_id` and the reply is **injected into the primary's idle pane** when ready. | -| **Primary** | `fleet_status(target?)` | no | Worker's live `agent_status` (`idle`\|`working`\|`blocked`\|`unknown`), queue depth, open rendezvous — and for the *calling* session it also **reports/drains pending messages addressed to it**: how a split-host / non-pane primary pulls replies injection can't serve (subsumes the earlier `fleet_poll`). | -| **Worker** | `fleet_reply(text, {final})` | no | Emit a **structured** reply/payload to whoever awaits this turn. | -| **Worker** | `fleet_ask(question)` | yes | Worker-initiated question up the chain (true 2-way); parks the worker until the primary answers. | -| both | `fleet_list()` | no | List workers + status, for orchestration or a human (the earlier `fleet_sessions`). | -| **Primary** | `fleet_spawn({profile?})` · `fleet_stop(target)` · `fleet_read(target, source)` | no | Worker lifecycle (guard-checked spawn, idempotent teardown) and peeking at a detached worker's terminal. | - -> The normative tool-by-tool surface (parameters, outcomes, error model) is -> **`docs/MCP-Contract.md`** in the main repo (2026-07-14); this table mirrors it. - -### The rendezvous — why worker → primary needs no keystrokes - -`fleetd` is the meeting point. When the primary is parked in a blocking `fleet_send`, -`fleetd` resolves that pending tool call the instant **either** signal arrives: - -- herdr's status stream shows the pane's turn-done edge — `agent_status_changed: working → - idle` (**passive** — fires even for an uncooperative worker; herdr has no `done` status), - **or** -- the worker calls `fleet_reply(result)` (**active** — a structured payload; preferred). - -Because the reply flows back through `fleetd`'s own state, the old "herdr types the answer -into the primary's pane" path is **no longer needed, even single-host** — the primary reads -its answer as an ordinary MCP tool result. herdr keystroke-injection into the *primary* pane -survives only as a degraded fallback for a herdr-only (non-MCP) primary. - -```mermaid -sequenceDiagram - participant P as "Primary (Opus) — MCP client" - participant S as "fleetd (MCP server + rendezvous)" - participant H as "herdr" - participant W as "Worker claude — MCP client" - - P->>S: "fleet_send(worker, task, block) — tool call PARKS" - S->>H: "send_text + send_keys (gated on idle)" - H->>W: "inject turn" - activate W - H-->>S: "event: agent_status_changed = working" - par active payload - W->>S: "fleet_reply(result) — structured" - and passive signal - H-->>S: "event: agent_status_changed → idle (turn done)" - end - deactivate W - S-->>P: "tool result = reply (unparks fleet_send)" - Note over P,W: "worker → primary rode fleetd's state,
not a keystroke into the primary pane" -``` - -*Figure: the reply resolves on whichever of the two signals lands first; `fleet_reply` -carries the structured payload, the herdr event guarantees the timing even if the worker -never cooperates.* - -### Worker mounts MCP too — graceful tiers - -Mounting the bridge MCP on the *worker* is what makes the loop symmetric and structured, yet -it degrades cleanly if a worker is left unmodified: - -| Tier | Worker setup | Reply channel | Worker can ask back? | -|---|---|---|---| -| **Unified (recommended)** | mounts `bridge` MCP (same one line) | structured `fleet_reply` | yes — `fleet_ask` | -| **Hooked (no MCP)** | a `Stop`-hook installed | structured envelope POSTed to `fleetd` (see [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply)) | no | -| **Unmodified (last resort)** | stock `claude` | `pane.read` scrape on the turn-done `working→idle` edge (lossy) | no | - -The **primary-side contract is identical** in both tiers; only the worker's reply fidelity -changes. Ship the unified setup — one MCP line on every session — and keep herdr-only as the -zero-coupling escape hatch. MCP tool calls never touch `ANTHROPIC_BASE_URL`, so mounting the -bridge on either side is subscription-safe by construction (see -[Subscription boundary](#subscription-boundary-enforced-not-just-documented)). - -## Architecture - -`fleetd` is a **standalone daemon** — one component, two faces: - -- a **SERVER face** — an **MCP server** the Claude Code sessions mount, plus REST/SSE for - non-Claude clients, sitting over the session tracker, subscription guard, and reply - rendezvous (the policy brain); and -- a **CLIENT face** — a herdr socket client that injects turns (status-gated) and subscribes - to agent-status. - -The Claude sessions themselves live **as panes inside herdr**. Each pane reaches *up* to -`fleetd`'s MCP server (to send/reply); `fleetd`'s herdr client reaches *down* through -herdr's socket to drive those same panes and read their status. herdr, `fleetd`, the panes, -and the model are colocated on one host (same-host scenario); a broker is optional for async -duplex. +The project's earlier approach, `AgentAPI`, would have re-implemented a terminal emulator and +guessed when an agent was done from screen stability. herdr already solves that as a running +service: it owns the panes, tells `fleetd` the instant an agent's status changes +(`idle`/`working`/`blocked`), and survives a detach/reattach over SSH. `fleetd`'s job is the +policy on top: which panes may be written to, when, and how a reply gets back to whoever sent +the task. ```mermaid flowchart TB - subgraph herd["herdr — agent multiplexer (Claude sessions run here, same host)"] - PP["primary pane · Opus
env CLEAN · MCP client"] - WP["worker pane(s) · claude
ANTHROPIC_BASE_URL set · MCP client"] + subgraph herd["herdr — agent multiplexer"] + PP["lead pane
MCP client"] + WP["member pane(s)
ANTHROPIC_BASE_URL set
MCP client"] end - subgraph fleetd["fleetd — standalone daemon (NOT a claude process)"] + subgraph fleetd["fleetd — standalone daemon"] subgraph srv["SERVER face"] - MCP["MCP server
fleet_send · reply · ask · status"] - REST["REST / SSE
(non-Claude clients)"] - POL["policy brain
session tracker · subscription guard
· reply rendezvous"] + MCP["MCP server (/mcp)
fleet_send · fleet_reply · fleet_ask · ..."] + REST["REST API"] + POL["policy: SessionManager
SubscriptionGuard · MessageService/Rendezvous"] end subgraph cli["CLIENT face"] - INJ["injector
status-gated"] - HCL["herdr socket client
send_text · events · pane.read"] + INJ["Injector
status-gated FIFO per pane"] + HCL["herdr socket client"] end MCP --> POL REST --> POL POL --> INJ --> HCL - HCL -->|"events · replies"| POL end - MODEL["ollama.ltms.dev / GX10 vLLM
(worker model)"] - BROKER["broker / queue (optional, internal)
Redis / NATS — durability · cross-host"] - PP -->|"MCP tools"| MCP WP -->|"MCP tools"| MCP - HCL -->|"Unix socket · drive + status"| herd - WP -->|"inference"| MODEL - POL -.->|"async: enqueue / cross-host"| BROKER + HCL -->|"Unix socket"| herd classDef core fill:#2f855a,stroke:#22543d,color:#ffffff; classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff; - classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff; class MCP,REST,POL,INJ,HCL core - class PP ext - class MODEL,BROKER warn + class PP,WP ext ``` -*Figure: `fleetd` is one standalone daemon with a **SERVER** face (the MCP endpoint the -Claude panes mount, over the policy brain) and a **CLIENT** face (the herdr socket client). -The Claude sessions are herdr **panes**: they call *up* into the MCP server, while `fleetd`'s -client drives them *down* through herdr's socket and gates every injection on live -agent-status. Only worker panes carry `ANTHROPIC_BASE_URL`; `fleetd` holds no quota, so it -subscribes freely. The broker is **`fleetd`-internal** (durability / cross-host) — no pane -ever addresses it; async delivery is `fleetd` injecting an idle pane. (Split-host moves the -primary out of herdr — see [Deployment model](#deployment-model).)* +*Figure: `fleetd` is one daemon with a SERVER face (MCP + REST, over the policy layer) and a +CLIENT face (the herdr socket client). Verified: `FleetMcp.java:150-174` (MCP build), +`FleetApp.java:126-160` (REST routes), `FleetMcp.java:127` / `FleetApp.java:62` +(both share one `MessageService`).* -### Components +## Mounting the MCP server -| Component | Responsibility | -|---|---| -| **herdr socket client** | NDJSON over `~/.config/herdr/herdr.sock`; correlates responses by `id`; maintains a long-lived `events.subscribe` stream. | -| **Session manager** | Maps a logical session → herdr `workspace/tab/pane` id. Spawns the worker `claude` (env-prefixed launch line into a fresh pane's shell), health-checks, and **recycles on context ceiling** (Ralph loop, see below). | -| **Injector** | Per-pane FIFO queue. Delivers `send_text` + `send_keys "enter"` **only when** that pane's `agent_status ∈ {idle, blocked}` — never mid-run. | -| **Reply rendezvous** | Resolves an awaiting `fleet_send` on whichever lands first: a worker `fleet_reply` (structured, **preferred**), the turn-done `working → idle` status edge (timing guarantee), or — worker-side hook path — a `Stop`-hook envelope; last-resort `pane.read {source:"recent-unwrapped"}` scrape. See [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply). | -| **Subscription guard** | Refuses to spawn a *worker* pane without `ANTHROPIC_BASE_URL`; refuses to *ever* set it on a pane designated *primary*; can assert egress host via `pane.process_info`. | -| **SERVER API** | **MCP server** — the Claude-facing contract both primary and workers mount (`fleet_send`/`reply`/`ask`/`status`/`poll`/`sessions`). Plus **REST + SSE** (OpenAPI, AgentAPI-shaped) for non-Claude clients — webhooks, dashboards, a human CLI. | -| **Broker connector** *(optional, internal)* | `fleetd`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `fleetd` will later inject into an idle pane. No Claude session ever connects to it. | +Both the lead and every member reach `fleetd` the same way: they mount it as an MCP server. +The daemon listens by default on `127.0.0.1:8765` and serves the MCP endpoint at `/mcp`. The +port comes from `FleetConfig.Bind`, which defaults to `8765` when the config does not set one +(`FleetConfig.java:183-187`); the endpoint path is set at `FleetMcp.java:152` +(`.mcpEndpoint("/mcp")`). The server reports its own name as `fleet` at +MCP initialize time (`FleetMcp.java:313-315`: `.serverInfo("fleet", "0.1.0")`). -## The herdr control contract (what `fleetd` drives) - -> **Verified against the running herdr 0.7.0 (protocol 14), not just docs.** A spike hit the -> live socket. `ping` returns `{version:"0.7.0", protocol:14, capabilities:{live_handoff:true}}`; -> `workspace.list` / `pane.list` return the pane inventory (with `agent_status`); `events.subscribe` -> exists. **There is no `session.snapshot`** (the earlier draft was wrong — use -> `workspace.list`/`pane.list` to enumerate/re-attach). Crucially, herdr 0.7.0 exposes a -> **native `agent.*` namespace** — `agent.start`, `agent.send`, `agent.read`, `agent.list`, -> `agent.get`, `agent.focus`, plus `server.agent_manifests` and `pane.report_agent` — so herdr -> already models "agents", not just panes. (Also available: `worktree.*` for git-isolated workers.) -> -> **CB-102 — decided: `fleetd` drives the native `agent.*` path** (the pane + `send_text` -> workaround below is the documented fallback only). The spike proved against the live daemon that -> `agent.start` accepts a **first-class `env` map** that reaches the process environment — so a -> worker's `ANTHROPIC_BASE_URL` is injected cleanly and guard-checked, with no shell-prefix -> parsing and `fleetd`'s own env untouched. herdr also tracks each worker's **Claude session -> UUID** (`agent_session.value`), grounding the ID contract in herdr's own identity. Observed -> `agent.*` schema: -> -> | Call | Params | Returns | -> |---|---|---| -> | `agent.start` | `{name, argv:[...], env:{...}, tab_id?}` (8 fields; `tab_id` **honored** — worker lands in that tab, else it splits the focused tab) | `{agent:{terminal_id, pane_id, workspace_id, tab_id, agent_status}}` | -> | `agent.send` | `{target, text}` | ack | -> | `agent.read` | `{target, source}` — `source ∈ visible\|recent\|recent_unwrapped\|detection` | `{read:{text}}` | -> | `agent.get` | `{target}` | `{agent:{…, agent_session:{value:}, agent_status}}` | -> | `agent.list` | `{}` | `{agents:[…]}` — discovery, each with its session UUID | -> -> `target` is the `terminal_id`. There is **no `agent.stop`** — tear a worker down with -> `pane.close {pane_id}`. `agent_status ∈ idle\|working\|blocked\|unknown` drives the -> status-gated injector (inject only when idle/blocked). -> -> **CB-108 — worker placement: a tab per worker in a dedicated "worker space."** By default -> `agent.start` (no `tab_id`) *splits the currently-focused tab*, so it would clutter — and could -> co-tenant — the user's real work tabs. Instead `fleetd` puts every worker in its **own tab** -> inside a dedicated **workspace** (`worker.workspace`, default `fleetd-workers`), found-or-created -> once and shared (a future per-session layout is just a distinct label). The recipe, all verified -> live: `workspace.create {label}` (idempotent find-first) → `tab.create {workspace_id}` → -> `agent.start {…, tab_id}` → `pane.close {root_pane}` (drop herdr's seed shell so the tab holds -> only the worker) → `tab.rename {tab_id, label}`. Teardown: resolve the pane's tab via `pane.get`, -> `pane.close`, then `tab.close` **only if the tab holds one pane** (never a shared user tab). The -> shared space is persistent (its seed tab is a stable anchor); `worker.placement: pane` restores -> the legacy split behaviour. -> -> **Two more `agent.*` facts the CB-108 build pinned:** (1) an agent's **`name` must be unique** -> among running agents — a 2nd `agent.start {name:"claude"}` fails `agent_name_taken`, so `fleetd` -> names workers `claude---` (the per-process `nonce` survives a daemon restart -> where old workers linger). (2) herdr detects an agent's **kind and status from terminal output** -> (braille spinner / `❯` prompt patterns in the remote `*.toml` manifests), **not** from `name` — so -> the unique name is a pure label; the detected kind arrives later in `agent.get`/`agent.list`'s -> `agent` field (absent, hence kind unknown, at start-time). -> -> **Two wire facts the Stage-1 client (CB-101) pinned by contract test — both bit the first -> build:** (1) the request `id` **must be a JSON string** — an integer id is rejected with -> `invalid_request`. (2) herdr serves **one request/response per connection, then closes it** — -> a second write on the same socket gets a broken pipe. So the client is **connection-per-call** -> (open → one frame → one line → close), which as a bonus needs no locking. A long-lived socket is -> only for the streaming `events.subscribe` path, not request/response. - -The pane-based control path — the **fallback** now that CB-102 chose `agent.*`; shown here for -reference and for herdr builds without the `agent.*` namespace: - -```jsonc -// spawn: create a pane, then launch the worker in its shell (env stays worker-only) -{"id":"1","method":"workspace.create","params":{"cwd":"~/work","label":"worker-a"}} -{"id":"2","method":"pane.send_text","params":{"pane_id":"w1:p1", - "text":"ANTHROPIC_BASE_URL=https://ollama.ltms.dev ANTHROPIC_AUTH_TOKEN=… claude"}} -{"id":"3","method":"pane.send_keys","params":{"pane_id":"w1:p1","keys":"enter"}} - -// deliver a turn (gated on status=idle|blocked) -{"id":"4","method":"pane.send_text","params":{"pane_id":"w1:p1","text":""}} -{"id":"5","method":"pane.send_keys","params":{"pane_id":"w1:p1","keys":"enter"}} - -// readiness / completion — push, not polling -{"id":"6","method":"events.subscribe","params":{"subscriptions":[ - {"type":"pane.agent_status_changed","pane_id":"w1:p1"}]}} - -// answer a blocked worker / interrupt -{"id":"7","method":"pane.send_input","params":{"pane_id":"w1:p1","keys":"ctrl+c"}} - -// fallback content read -{"id":"8","method":"pane.read","params":{"pane_id":"w1:p1","source":"recent-unwrapped","lines":200}} +```bash +claude mcp add --transport http fleet http://127.0.0.1:8765/mcp ``` -> **Implementation note (verify in the CLI reference):** herdr's socket may or may not expose -> a direct "spawn command + env" primitive. `fleetd` uses the robust path — create pane → -> `send_text` the env-prefixed launch line — which guarantees `ANTHROPIC_BASE_URL` lands in -> the **worker pane's shell only**. If a native spawn call exists, prefer it and pass env -> explicitly; the guard invariant is unchanged. +The port in the command above is the *default*. A real deployment's actual port comes from its +`bind:` block in `fleetd.yaml` (or `bridged.yaml`, depending on the host) — check that file +before assuming 8765. -## Reply envelope (how a worker emits a structured reply) +## Tool reference -With the worker **mounting the bridge MCP** (the unified setup), the cleanest path is direct: -the worker simply **calls `fleet_reply(result)`** before finishing — a first-class tool call -that hands `fleetd` a structured payload with no transcript scraping. Prefer this whenever -the worker is MCP-mounted. +The registered tool set is built once, in `FleetMcp`'s constructor +(`FleetMcp.java:301-327`). The parameters listed below come from each tool's own schema method +(`FleetMcp.java:1088-1246`), not from any older document. -The hook path below remains the **fallback** for a herdr-only worker (no MCP) — a Claude Code -worker cannot otherwise `XADD` to a broker or POST a callback, so "the worker writes an -envelope" resolves to **a worker-side hook that runs our code at turn end**: +| Tool | Who calls it | Parameters | Returns | +|---|---|---|---| +| `fleet_send` | lead | `content` (required), `sessionId`, `timeoutMs`, `wait` (default `true`), `turnId`, `coordId` | Delegates `content` to the member named by `sessionId` and, by default, blocks for its reply. `wait:false` returns a ticket to poll with `fleet_poll` instead of blocking. Passing `turnId` (instead of `sessionId`) answers a member's open `fleet_ask` question. Passing `coordId` (instead of `sessionId`/`turnId`) sends to a peer lead's mailbox on another daemon — this is coordination between leads, not a task. `sessionId`, `turnId` and `coordId` are mutually exclusive. | +| `fleet_reply` | member | `content` (required) | Ends a delegated turn with a structured answer. The member's identity comes from its connection, never an argument, so a member can only ever reply as itself. | +| `fleet_ask` | member | `question` (required), `timeoutMs` | Pauses the member's current turn to ask the lead a question, and blocks until the lead answers (the lead answers with `fleet_send{turnId, content}`). Default timeout is 55 seconds and the cap is 115 seconds — kept under a typical MCP client's own ~60s call cap so the tool returns a clean timeout instead of the client just severing the connection. | +| `fleet_status` | lead | `sessionId` (required) | The member's live status: `idle`, `working`, `blocked`, or `unknown`. If the member is paused mid-turn in an async `fleet_ask`, the result also carries the open question and its `turnId`. | +| `fleet_poll` | lead | `ticket`, `target` | With `ticket`: the state of a `fleet_send{wait:false}` delegation — pending, done (with the reply), asking, or failed. With `target` instead: drains that member's reply inbox (replies that arrived when no `fleet_send` was open waiting for them). | +| `fleet_ack` | lead | `target` (required), `msgId` (required) | Removes one specific reply from a member's inbox, leaving any others queued. | +| `fleet_spawn` | lead | `role`, `profile`, `cwd`, `worktree`, `ticket`, `sessionName`, `resumeSessionId` | Starts a new member. `role` picks the contract (`dev`, `reviewer`, or `architect`; default `dev`); `profile` picks the backend (default is the configured default profile). `worktree:true` (with `ticket`) or `worktree:` provisions an isolated git worktree. Returns the member's `sessionId` (for `fleet_send`) and `paneId` (for `fleet_stop`). | +| `fleet_list` | lead | none | The whole fleet: `leads` (peer orchestrators, each with `sessionId`, `name`, live `status`, and `self:true` on the caller's own row) and `members` (each with `sessionId`, `paneId`, `role`, `profile`, `state`, and — when set — `worktree`/`branch`/`owner`). When capacity facts are configured, also a `capacity` row per profile. | +| `fleet_stop` | lead | `paneId` (required) | Tears a member down by its pane id. | +| `fleet_profiles` | lead | none | The configured backend profiles, the default one, and (when CB-578's stage B quarantine has tripped) which profiles are currently refusing new spawns and for how long. | +| `fleet_whoami` | any | none | The caller's own resolved role (`primary`, `architect`, or `worker`) and identity — the caller never has to guess its own role from a side channel. | -1. A **`Stop`-hook** on the worker fires when its turn ends. The hook reads the **last - assistant message** from the session transcript (`~/.claude/projects//.jsonl`, - the path Claude Code exposes to hooks) and POSTs `{session_id, turn_id, status, text, - artifacts}` to **`fleetd`'s callback endpoint** — the gateway, not a broker (the worker-side - hook never writes the broker directly; `fleetd` queues internally if it must). -2. `fleetd` correlates that envelope to the open blocking request by `session_id`/`turn_id` - and returns it as the response body. The turn-done `working → idle` status edge is the - *timing* signal; the hook payload is the *content*. -3. If no hook is installed, `fleetd` falls back to `pane.read {source:"recent-unwrapped"}` - and best-effort parses the last assistant block (reuse AgentAPI's `msgfmt`). This is - lossy and is the reason the envelope path is preferred. +`fleet_read` does not exist. The full registration list is at `FleetMcp.java:301-327`, and it +has no tool by that name. Older pages named one; it was never built. -> **Envelope contract — now specified.** A first-cut envelope (fields `from`/`to`/`session`/ -> `turn`/`corr`/`kind`/`body`, and the `kind` verb vocabulary that tells a recipient what to do -> next) is defined in [Use Cases → The ID contract](7-Use-Cases#mechanism-4--the-id-contract-envelope). -> Landing it is **Stage 2** ([Roadmap](8-Roadmap), ticket `CB-201`). The MCP-mounted worker -> emits its reply via `fleet_reply` (structured); the `Stop`-hook envelope remains the fallback -> for a non-MCP worker. Still open: exactly how tool/diff artifacts attach to `review.reply`. +## The rendezvous — a blocking call, not polling -## Delivery gating & races (the injector is a single writer) - -The injector gates on cached `agent_status` (updated by the `events.subscribe` stream) and -then calls `send_text`. That read-then-send is a **TOCTOU window**: the pane could leave -`idle` between the status read and the keystrokes landing. Mitigations, and the residual gap: - -- **Single writer per pane.** `fleetd` is the *only* automated injector into a worker pane; - the per-pane FIFO queue serializes deliveries so two turns never interleave. This removes - injector-vs-injector races, not injector-vs-agent ones. -- **Serialize send within the event loop.** Do the status check and the `send_text`/`send_keys` - pair as one non-preemptible unit on the same event-loop virtual thread that consumes events, - so a status change can't be processed mid-send. -- **Initial readiness needs prompt detection, not just status.** On spawn there may be no - `agent_status` event until the first turn. Gate the *first* injection on an - `output_matched`/`pane.read {source:"detection"}` prompt-ready signal (the `Ready` state - below), not on an absent status. -- **Residual race (accepted).** A human typing into the same worker pane, or the agent - self-transitioning to `working` in the millisecond after the gate, can still collide. The - cost is a corrupted turn, not a subscription breach; recovery is `ctrl+c` + re-inject. - Treat a worker pane as **fleetd-owned** (don't hand-drive it) to avoid this. - -## Message flow - -### Blocking delegation (primary → worker → primary) +When the lead calls `fleet_send` and blocks, `fleetd` does not make the lead poll. It parks +the MCP call and resolves it the instant one of two things happens: the member calls +`fleet_reply`, or (a fallback, CB-106) the member's turn ends without ever calling +`fleet_reply`, in which case `fleetd` returns the scraped transcript tail instead, clearly +flagged as such. The outcome list is `MessageService.Outcome` (`MessageService.java:65-103`), +and `FleetMcp.formatReply` renders it back to the tool caller (`FleetMcp.java:541-564`). ```mermaid sequenceDiagram - participant P as "Primary (Opus)" - participant S as "fleetd" - participant H as "herdr" - participant W as "Worker claude" + autonumber + participant L as Lead — MCP client + participant F as "fleetd (MessageService + Injector)" + participant M as Member — MCP client - P->>S: "fleet_send(id, task, block) — tool call PARKS" - S->>S: "await pane status = idle" - S->>H: "pane.send_text + send_keys enter" - H->>W: "inject turn" - activate W - H-->>S: "event: agent_status_changed = working" - S-->>P: "SSE: status working (observers only)" - W->>S: "fleet_reply(result) — or Stop-hook envelope (fallback)" - H-->>S: "event: agent_status_changed → idle (turn done)" - deactivate W - S->>S: "resolve reply (fleet_reply, event, or pane.read fallback)" - S-->>P: "tool result = assistant reply (unparks call)" - Note over P: "review diff / result, merge" + L->>F: fleet_send(sessionId, content) — call blocks + F->>F: Injector delivers content when the pane is idle + activate M + M->>M: works the turn + M->>F: fleet_reply(content) + deactivate M + F-->>L: tool result = the member's reply + Note over L,M: if the member's turn ends with no fleet_reply,
fleetd returns the scraped transcript tail instead ``` -*Figure: the primary's request blocks; `fleetd` gates injection on `idle`, streams status -transitions over SSE for observers, and returns the reply — from the worker's `Stop`-hook -envelope, falling back to a `recent-unwrapped` scrape — as the blocking call's response body.* +*Figure: the send/reply round trip. Verified against `MessageService.java:65-103` (the outcome +enum) and `FleetMcp.java:486-501` / `541-564` (`fleet_send`'s handler and reply rendering).* -### Async duplex — fleetd-mediated (event bus / worker → recipient) +## `fleet_ask` — a member asking the lead back + +`fleet_ask` is the reverse direction: a member pauses its own turn to ask the lead a question, +and the lead answers by calling `fleet_send` again with `turnId` set instead of `sessionId`. +The member then resumes the same turn. See `FleetMcp.ask` (`FleetMcp.java:521-538`), and the +`turnId` branch in `fleet_send` (`FleetMcp.java:198-203`, calling `FleetMcp.answer` at +`FleetMcp.java:508-514`). ```mermaid sequenceDiagram - participant SRC as "Source: webhook/bus (REST) · worker fleet_reply/ask (MCP)" - participant S as "fleetd (gateway)" - participant Q as "broker / queue (internal)" - participant H as "herdr" - participant R as "Recipient pane (idle Claude)" + autonumber + participant L as Lead + participant F as fleetd + participant M as Member - SRC->>S: "REST ingress · or MCP fleet_reply / fleet_ask" - opt durability / cross-host - S->>Q: "enqueue (group + XACK)" - Q-->>S: "dequeue when ready" - end - S->>H: "await recipient agent_status = idle" - H-->>S: "event: idle" - S->>H: "pane.send_text + send_keys (inject)" - H->>R: "new turn = the message" - Note over S,R: "recipient polled nothing — fleetd pushed on the idle edge.
split-host primary: its Stop-hook polls fleetd, never the broker" + L->>F: fleet_send(sessionId, content) — call blocks + F->>M: deliver content + activate M + M->>F: fleet_ask(question) — member's own call blocks + F-->>L: tool result = QUESTION(text=question, turnId) + L->>F: fleet_send(turnId, content=answer) — call blocks again + F-->>M: unblocks fleet_ask with the answer + M->>M: resumes the same turn + M->>F: fleet_reply(content) + deactivate M + F-->>L: tool result = the member's reply ``` -*Figure: `fleetd` mediates async in both directions. Every source reaches it over the gateway -(MCP for Claude, REST for external), it optionally parks the message on its **internal** queue, -waits for the recipient's idle event, and injects. The queue is a `fleetd` implementation -detail — no Claude session touches it, and herdr scrollback is never the source of truth.* +*Figure: the reverse rendezvous. `QUESTION` is one of `MessageService.Outcome`'s values +(`MessageService.java:91`); `FleetMcp.formatReply` turns it into a tool result that names +the `turnId` and tells the lead how to answer (`FleetMcp.java:556-558`).* -## Worker session lifecycle +## Detached delegation (`wait:false`) -`fleetd` treats a worker as a **recyclable** resource, not one immortal session — a -long-lived pane fills its context window. When idle-cycle or token caps trip, `fleetd` -kills the pane and respawns fresh (**Ralph loop**). +A blocking `fleet_send` call is capped by the caller's own MCP client timeout — usually +around 60 seconds — but a real task can run for minutes. `fleet_send{wait:false}` runs the +same delegation on a background thread and returns a ticket right away; the lead checks on it +later with `fleet_poll{ticket}`. See the class-level doc comment on `MessageService` +(`MessageService.java:40-44`, describing `sendAsync`) and the `fleet_poll` tool handler +(`FleetMcp.java:645-670`). -**What "state on disk" actually means (be precise — this is easy to hand-wave).** Recycling -deliberately **sheds the conversation transcript**; it is *not* `claude --resume`, which would -reload the full context you are trying to drop. Continuity is instead carried by **artifacts -the worker externalizes as it works**: +A finished ticket is kept for **10 minutes after it completes** (`MessageService.java:62`, +`TICKET_TTL_NANOS`), then pruned. So the ten minutes are the window to *collect* the report, whatever +the task's own runtime was. -- Git commits / a working branch (the real output). -- A durable **task/progress file** (a scratchpad, `TODO`/`STATE.md`, or the broker's own - record of the outstanding work item) that the worker is instructed — via `CLAUDE.md` / - skill — to keep current. -- The fresh worker is spawned with a prompt that says *"here is the task and the state file; - continue from it,"* not with the old messages. +This was a real bug until 2026-08-31 (#197): the TTL was measured from when the ticket was +**created**, so the window was ten minutes minus however long the task ran. Any delegation lasting +more than ten minutes had its report destroyed the moment it arrived, and `fleet_poll{target}` +returned nothing rather than holding it. It is fixed, but a daemon that has not been redeployed since +still behaves the old way. Separately, if a member calls +`fleet_reply` while no `fleet_send` is open waiting for it, `fleetd` does not drop the reply — +it queues it in that member's inbox, and a background push loop (`ReplyPushLoop`) nudges the +lead's own pane to go check when this happens (`MessageService.java:370-383`, +`MessageService.reply`). The mechanism that carries the nudge into the pane lives in +`ReplyPushLoop`; `MessageService.reply` calls `pushLoop.onReplyQueued(session)` when a reply +arrives with no waiter. -This only works if the worker is disciplined about writing that state **before** a recycle -boundary. `fleetd` can enforce a checkpoint (inject "commit and update STATE.md" before it -kills the pane), but a worker that ignores it loses in-flight context. **Open question:** -whether to also snapshot the raw transcript (`--resume` the *same* session on crash-restart, -vs. a clean context on a planned recycle) — the two restart reasons may want different -policies. Treat cross-recycle continuity as a pattern to prove in M2, not a guarantee. +## REST API -```mermaid -stateDiagram-v2 - [*] --> Spawning - Spawning --> Ready: "claude prompt detected" - Ready --> Working: "turn injected" - Working --> Blocked: "permission / question" - Blocked --> Working: "fleetd answers (send_input)" - Working --> Ready: "agent_status → idle (turn done)" - Ready --> Recycling: "context / idle cap hit" - Recycling --> Spawning: "state persisted to disk" - Working --> Failed: "pane.exited (crash)" - Failed --> Spawning: "auto-restart + replay unacked" - Ready --> [*]: "drain / shutdown" -``` - -*Figure: the state machine `fleetd` drives per worker. `blocked`, the turn-done -`working → idle` edge, and `pane.exited` are real herdr signals, not heuristics — the reason -herdr-centric beats screen scraping.* - -## Subscription boundary (enforced, not just documented) - -The invariant is unchanged from [Architecture](1-Architecture) — **anything that sets -`ANTHROPIC_BASE_URL` is, by definition, the worker** — but here it is *enforced in code*: - -- `fleetd` is **not** a `claude` process. It consumes zero Anthropic quota, so it may - busy-poll the broker and hold a permanent herdr event subscription with no policy concern. -- The **subscription guard** blocks any spawn of a *worker* pane whose launch does not - resolve an **off-subscription** `ANTHROPIC_BASE_URL`, and blocks any attempt to set that - var on a pane tagged *primary*. -- **The guard's reach is limited — be honest about it.** A substring check on the launch - *string* is necessary but not sufficient: - - The var may arrive from a **profile / `.envrc` / direnv / systemd `Environment=`**, not - the launch line — so string-absence does **not** prove the worker is on-subscription, and - string-presence does **not** prove it points off-subscription (`ANTHROPIC_BASE_URL=https://api.anthropic.com` - would pass a naive check while burning subscription-adjacent auth). The guard must - validate the **resolved value's host** against an allowlist and, after spawn, confirm - egress via `pane.process_info` / a health call to the worker model — not trust the string. - - The guard only constrains panes **`fleetd` spawns**. A split-host primary on your Mac is - a process `fleetd` never sees; it cannot inspect that env. There, subscription safety - rests on the operator (the Mac `claude` simply is never given the var) plus the fact that - the *only* thing crossing to the worker host is **MCP/HTTP traffic to `fleetd`**, never an - endpoint swap and never a broker connection. -- Injecting keystrokes into the **primary** pane is subscription-safe: it is simulated - typing, identical to the human at the keyboard — the primary still talks to - `api.anthropic.com` on Pro/Max. `fleetd` never re-points the primary's endpoint. (This - path exists only single-host, where the primary is a herdr pane.) -- A startup self-check asserts any **locally-hosted** primary pane's env has **no** - `ANTHROPIC_BASE_URL` and logs each worker's resolved egress host. It cannot self-check a - remote primary. - -## API surface (SERVER face) - -**Two faces over one core — and REST is the testability surface.** Every feature is implemented -as a **REST route**; the **MCP tools are thin adapters over those routes** — e.g. `fleet_send` -→ `POST /sessions/{id}/message`, `fleet_status` → `GET /sessions/{id}/status`, `fleet_reply` → -`POST /sessions/{id}/reply` (rendezvous). The REST routes are also the surface non-Claude clients -use (webhooks, dashboards, a human CLI), kept AgentAPI-shaped for drop-in migration. Because the -logic lives in REST, **each feature is acceptance-tested by an HTTP call with no Claude/MCP in -the loop**, and MCP is verified by a **parity test** (tool result == REST result). See -[Roadmap → Testability](8-Roadmap#testability--the-rest-api-is-the-contract-surface). +Every feature is also reachable over plain HTTP, so it can be tested and driven without an +MCP client. `FleetApp.build()` (`FleetApp.java:126-160`) registers every route the daemon has; +the table below is that complete list. | Method + path | Purpose | |---|---| -| `POST /sessions` | Create a worker session `{model, base_url, cwd, label}` → returns `session_id` | -| `GET /sessions` | List sessions + live `agent_status` | -| `POST /sessions/{id}/message` | Deliver a turn `{content, type:"user"\|"raw"}` (queued, status-gated) | -| `GET /sessions/{id}/events` | **SSE**: `status`, `message`, `blocked`, `exited` | -| `GET /sessions/{id}/messages` | Conversation history | -| `GET /sessions/{id}/status` | `idle` \| `working` \| `blocked` \| `unknown` | -| `POST /sessions/{id}/keys` | Raw keys passthrough `{keys:"ctrl+c"}` — answer/interrupt | -| `DELETE /sessions/{id}` | Drain + recycle | -| `GET /healthz` · `GET /metrics` | Liveness + Prometheus | +| `GET /healthz` | Liveness. Also confirms herdr is reachable (and, if a second `memberHerdrSocket` is configured, that the member daemon is too). | +| `GET /metrics` | Prometheus scrape (only registered when a metrics registry is configured). | +| `GET /sessions` | Herdr workspaces, one row per workspace, with live agent status. | +| `GET /agents` | Every agent herdr tracks, keyed by its Claude session id. | +| `GET /members` | The fleet's own session roster, merged with live herdr status — this is what `fleet_list` is built from. | +| `GET /profiles` | Configured backend profiles and the default one. | +| `POST /members` | Spawn a member (`fleet_spawn`'s REST equivalent). | +| `DELETE /members/{paneId}` | Tear a member down (`fleet_stop`'s REST equivalent). | +| `POST /sessions/{id}/message` | Deliver a turn, blocking by default (`fleet_send`'s REST equivalent); `{"wait": false}` in the body returns a ticket instead. | +| `POST /sessions/{id}/reply` | A member's structured reply (`fleet_reply`'s REST equivalent). | +| `GET /sessions/{id}/replies` | Drain a member's reply inbox. | +| `POST /sessions/{id}/ask` | A member's mid-turn question (`fleet_ask`'s REST equivalent). | +| `GET /sessions/{id}/status` | Live status, plus `ready` (whether the injector can currently deliver to this target) and any open question. | +| `GET /tasks/{ticket}` | Poll an async (`wait:false`) delegation. | -## Use cases +There is no `GET /events` route and no Server-Sent Events route of any kind. The table above +lists every `app.get`, `app.post` and `app.delete` call in `FleetApp.build()`, and that is the +whole set. An older design doc described an SSE status stream; it was never built. -1. **Subscription-safe delegation (the core case).** Primary Opus offloads bulk/routine - work — codegen, refactors, test writing, log triage — to a worker on a cheap/local model, - keeping Opus's context clean and its quota for review/merge decisions. -2. **A herd of specialized workers.** One `fleetd` + herdr multiplexes several workers - (e.g. a DeepSeek coder, a fast summarizer, a long-context reader), each its own pane, - each addressable by `session_id`. Rolls up to a single status sidebar. -3. **Async event-bus automation.** A webhook/CI/NATS event hits `fleetd`'s **REST ingress**; - `fleetd` wakes an idle worker by injection, and the result flows back to the primary (or a - Slack/Telegram bridge) — all mediated by `fleetd`, no human in the loop. The source never - addresses a worker or a broker directly. -4. **Human co-pilot from anywhere.** Because herdr persists and detaches, the same worker is - reachable from a phone/chat bridge **posting to `fleetd`** while you're away, and from the - attached TUI when you're back. -5. **Long-running "perpetual" workers.** The Ralph-loop lifecycle lets a worker run for hours - across many context recycles without a human respawning it, state carried on disk. +## The subscription boundary -## Deployment model +The one invariant the whole design protects: **whatever sets `ANTHROPIC_BASE_URL` is a +member, never the lead.** This is enforced in code, not only documented. `SubscriptionGuard` +holds both checks (`SubscriptionGuard.java`): -### Single-host (default — simplest, recommended to start) +- `assertWorker(baseUrl)` refuses to spawn a member whose `ANTHROPIC_BASE_URL` is missing, or + whose host is not on the configured off-subscription allowlist. +- `assertPrimaryClean(env)` refuses to let the lead's own environment carry + `ANTHROPIC_BASE_URL` at all. -Primary, `fleetd`, herdr, and workers all on one off-subscription box. `fleetd` can inject -into **both** the primary and the worker panes (both local to one herdr), so the broker is -optional. Best for a workstation or a single dev box. - -### Split-host (primary local, workers remote near the model) - -Primary Opus runs on your Mac; `fleetd` + herdr + workers run on the GPU host next to -`ollama.ltms.dev` / GX10 vLLM. herdr's socket is **local-only**, so the Mac reaches the worker -host **only over `fleetd`'s MCP/HTTP endpoint** — never a remote herdr socket and never the -broker (the broker, if any, stays `fleetd`-internal on the worker host). - -```mermaid -flowchart LR - subgraph mac["Your Mac (subscription)"] - OPUS["Primary Opus
(Claude Code)"] - end - subgraph host["Worker host (off-subscription, near model)"] - direction TB - BD["fleetd
:8080 MCP · REST/SSE"] - HS["herdr server"] - W2["worker claude pane(s)"] - BR["broker / queue
(fleetd-owned, internal)"] - BD -->|"Unix socket"| HS --> W2 - BD -.->|"durability / cross-host"| BR - end - ML["ollama.ltms.dev / GX10 vLLM"] - - OPUS -->|"MCP/HTTP — the only link (sync + Stop-hook async)"| BD - W2 --> ML - - classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff; - classDef core fill:#2f855a,stroke:#22543d,color:#ffffff; - classDef warn fill:#b7791f,stroke:#7b341e,color:#ffffff; - class OPUS ext - class BD,HS core - class BR warn -``` - -*Figure: the Mac's **only** link to the worker host is `fleetd`'s MCP/HTTP endpoint — it -carries sync replies, and (since the Mac primary isn't a herdr pane) its `Stop`-hook polls that -same endpoint for async wake-ups. herdr's socket and the broker stay local and `fleetd`-owned. -The subscription boundary tracks the host boundary — nothing on the Mac ever sets -`ANTHROPIC_BASE_URL`.* - -**Security:** bind `fleetd`'s HTTP to `localhost` and reach it over an SSH tunnel, or front -it with a bearer token + TLS. Never expose the port unauthenticated — it is an agent-control -surface: `POST /message` runs arbitrary prompts, and `POST /keys` sends raw keystrokes -(including `ctrl+c`) into a live agent. (This closes the gap left by AgentAPI's open `:3284`.) - -**Multi-tenancy is an open item.** A single `fleetd` fronting a *herd* of workers today has -**one shared token = full control of every session**; there is no per-session authorization. -That is acceptable for a single-operator box but not for shared/multi-user use. Before that, -add per-session scoping (a capability token per `session_id`) and an audit log of injected -turns. Until then, treat one `fleetd` as one trust domain. - -## Proposed tech stack - -| Layer | Choice | Why | Alternative | -|---|---|---|---| -| **Server core** | **Java 21+ (virtual threads)** | Loom virtual threads fit the socket + MCP + queue + SSE **blocking fan-in** as cleanly as goroutines — one blocking thread per pane/call, no callback soup; mature libraries; the official MCP Java SDK exists. | **Kotlin** (same JVM, terser); Go (single static binary, smaller RSS); Rust (matches herdr, slower to build). | -| **Deploy artifact** | Runnable **JAR** on a JRE, or **GraalVM `native-image`** | native-image restores the "single binary → `scp` + systemd, fast cold start, small RSS" story the JVM otherwise gives up. | plain JRE + fat JAR (simplest); jlink custom runtime. | -| **herdr transport** | JDK **`UnixDomainSocketAddress` + `SocketChannel`** (native UDS, no dep), **NDJSON** via Jackson, `id`-correlated + a persistent events stream | Native herdr contract; a dedicated **virtual thread** blocks on the event stream. | — | -| **SERVER API — Claude** | **Official MCP Java SDK** (streamable-HTTP transport) on an embedded server (Jetty / Spring Boot) | The unified contract both primary and workers mount; native to Claude Code, no shell/`curl`, subscription-safe by construction | stdio MCP adapter (per-session subprocess) if a long-lived HTTP endpoint is undesirable | -| **SERVER API — others** | **REST + SSE** via **Javalin** (light) or Spring MVC | Drop-in for AgentAPI-shaped/non-Claude clients; SSE streams status cheaply | JAX-RS (Helidon/Quarkus); gRPC if callers are all code | -| **Internal queue** *(optional)* | **Redis Streams** via **Lettuce** (consumer groups, `XACK`, visibility timeout) — `fleetd`-owned, below the gateway | Durability + cross-host for async; satisfies the guardrails in [Architecture](1-Architecture). Same-host can start with an in-JVM queue and add this only when durability/cross-host is needed | NATS JetStream for multi-host scale; embedded H2/SQLite for a single host | -| **Config** | **YAML via Jackson** (`jackson-dataformat-yaml`) + env overrides | 12-factor; secrets via env only | MicroProfile Config; Spring config if on Spring Boot | -| **Observability** | **SLF4J + Logback**; **Micrometer** → Prometheus `/metrics`; `/healthz` | Ops from day one | OpenTelemetry traces | -| **Process supervision** | **systemd** unit (`java -jar` or the native-image binary), ordered after herdr | Restart-on-crash; ordered start (herdr before `fleetd`) | Docker Compose colocating herdr + `fleetd`; k8s (overkill for one host) | -| **Testing** | **JUnit 5** + a **mock UDS socket** server + golden transcripts; fake `ccs`/`claude` stubs | Deterministic CI without a real TTY | Testcontainers (Redis) + a real herdr for e2e | - -**Recommendation: Java 21+ with virtual threads.** The daemon is almost entirely -blocking-I/O fan-in (herdr socket, MCP calls, queue, SSE) — the exact shape Loom makes trivial: -one blocking virtual thread per pane and per in-flight call, no reactive plumbing. Ship a -**GraalVM `native-image`** build to recover Go's small-footprint/fast-start deploy. Use -**Kotlin** instead only if the team prefers it — same JVM, same libraries. (`agentapi`'s Go -`msgfmt` reply parser is trivial to reimplement; it isn't a reason to stay on Go.) - -### Interface sketch (Java 21) - -```java -// herdr socket client — one call method, id-correlated; events on a separate virtual thread. -interface Herdr { - JsonNode call(String method, Object params) throws IOException; // blocking, id-correlated - BlockingQueue subscribe(List subs) throws IOException; // fed by an event-loop vthread -} - -// Injector: deliver ONLY when the pane is safe to type into (status-gated, single writer per pane). -final class Injector { - void deliver(String pane, String text) throws IOException { - Status st = status.get(pane); - if (st != Status.IDLE && st != Status.BLOCKED) { - queue.computeIfAbsent(pane, k -> new ArrayDeque<>()).add(text); // hold until next idle event - return; - } - herdr.call("pane.send_text", Map.of("pane_id", pane, "text", text)); - herdr.call("pane.send_keys", Map.of("pane_id", pane, "keys", "enter")); - } -} - -// Subscription guard: the boundary, in code. The base_url comes from `ccs env ` -// (see Use Cases → ccs spawn), NOT a hand-built env — validate the RESOLVED host against an -// off-subscription allowlist, then confirm egress post-spawn via pane.process_info. -final class Guard { - private static final Set OFF_SUB_HOSTS = Set.of("ollama.ltms.dev" /* + GX10 vLLM host */); - - void assertWorker(String baseUrl) { - if (baseUrl == null || baseUrl.isBlank()) - throw new GuardException("refusing to spawn worker without ANTHROPIC_BASE_URL"); - String host = URI.create(baseUrl).getHost(); // api.anthropic.com must be rejected - if (host == null || !OFF_SUB_HOSTS.contains(host)) - throw new GuardException("worker base_url host %s is not an approved off-subscription host".formatted(host)); - } - - // Only meaningful for a primary fleetd itself hosts (single-host). A remote/Mac primary is a - // process fleetd never sees — its cleanliness is the operator's. - void assertLocalPrimaryClean(Map env) { - if (env.containsKey("ANTHROPIC_BASE_URL")) - throw new GuardException("primary env is tainted — this is the subscription line"); - } -} -``` - -## Build plan (milestones) - -| Milestone | Deliverable | Proves | -|---|---|---| -| **M0 — Spike** | herdr client + spawn one worker pane + one `send_text`/`send_keys` round-trip | herdr socket drives a real `claude` | -| **M1 — Status gate + MCP** | `events.subscribe` → Injector delivers only on `idle`/`blocked`; **MCP `fleet_send`/`status` mounted on the primary**, blocking reply via the rendezvous; SSE status out | No mid-run corruption; real completion signal; primary drives over MCP | -| **M2 — Boundary + lifecycle + worker MCP** | Subscription guard + Ralph-loop recycle + worker `fleet_reply`/`fleet_ask` (unified mount) + envelope fallback | Subscription-safe; survives context ceiling; symmetric 2-way | -| **M3 — Async + durability + split-host** | Idle-injection async delivery; `fleetd`-internal queue (Redis Streams) for durability/cross-host; split-host `Stop`-hook adapter that polls `fleetd` | Detached/long work, cross-host, gateway-only (no Claude↔broker) | -| **M4 — Harden** | Auth/TLS, metrics, mock-socket CI, systemd unit | Production shape | - -## Trade-offs & risks - -| Risk | Mitigation | -|---|---| -| **No held-open conversation** — each exchange is one request in, one reply out | By design. Short/medium tasks use a **single blocking call** (fine — no quota burn); long/detached tasks use **return-and-reinvoke**, the reply delivered later **through `fleetd`** (idle-pane injection, or a split-host `Stop`-hook polling `fleetd`). What's excluded is a persistent bidirectional stream the primary must babysit. See *How the primary actually consumes a reply*. | -| **Blocking call can outlive its timeout** on a very long task | Set a request deadline; on timeout `fleetd` returns "still working, await async" and the reply lands via `fleetd`'s async path (idle-injection) instead of erroring the delegation. Pick async delivery (Mode 2) up front for known-long work. | -| **herdr is young / single-dev** — betting transport on it | Durability lives in `fleetd`'s **internal queue** (mature Redis/NATS), not herdr — herdr carries only ephemeral delivery + status. The injector is a **pluggable interface** — fall back to `tmux send-keys` or AgentAPI without touching the queue or the gateway contract. | -| **herdr socket is local-only** | `fleetd`'s **MCP/HTTP** is the sole cross-host link; herdr and the queue stay per-host and `fleetd`-owned. | -| **Mid-run interrupt still unsolved** | Same as AgentAPI. Injection gates on status; `ctrl+c` via `pane.send_input` is the only (disruptive) interrupt. | -| **Spawn-with-env uncertainty in socket API** | Launch via `send_text` of the env-prefixed command → env is provably worker-only; verify native spawn in the CLI reference and prefer it if present. | -| **Reply-scrape fragility (fallback path)** | Prefer the structured **envelope** path; scrape `recent-unwrapped` only as a last resort. | -| **herdr socket API is unversioned + single-dev churn** | Pin the herdr version in the systemd/Compose unit; keep the socket client behind the `Herdr` interface; probe `ping` (assert `protocol: 14`) on connect and enumerate via `workspace.list`/`pane.list`, failing fast on an unexpected schema. Don't build against `UNCERTAIN` primitives (e.g. native spawn-with-env) until confirmed in the running CLI. | -| **SPOF (fleetd / herdr / queue)** | Documented in [Architecture](1-Architecture) → *Failure modes*. Key property: the **primary is never downstream** of a bridge component, so a total outage costs workers only, never the subscription session. | -| **Injection TOCTOU / shared pane** | Single-writer injector + serialized send; worker panes are fleetd-owned. Residual collision corrupts a turn (recoverable), never the subscription boundary. See *Delivery gating & races*. | +Both throw `GuardException` on violation, which `fleet_spawn` turns into a +`subscription boundary: …` tool error (`FleetMcp.java:825-826`) and the REST route into an +HTTP `403` (`FleetApp.java:384-385`). ## Related pages -- **[Architecture](1-Architecture)** — the two-invariant / two-mode model this refines; subscription boundary -- **[Approaches](3-Approaches)** — transport comparison; AgentAPI now the *fallback injector* -- **[Home](Home)** — project overview +- **[Architecture](1-Architecture)** — the two-invariant model this page refines. +- **[Approaches](3-Approaches)** — why herdr was chosen; `AgentAPI` as discarded research. +- **[Home](Home)** — project overview. ## Sources -- [herdr — socket API](https://herdr.dev/docs/socket-api/) · [agent guide](https://herdr.dev/agent-guide.md) · [SKILL.md](https://raw.githubusercontent.com/ogulcancelik/herdr/master/SKILL.md) · [GitHub](https://github.com/ogulcancelik/herdr) -- [coder/agentapi](https://github.com/coder/agentapi) — fallback injector; reference for `msgfmt` reply parsing -- [Issue #27441 — inter-agent message injection](https://github.com/anthropics/claude-code/issues/27441) · [#24947 — `claude inject`](https://github.com/anthropics/claude-code/issues/24947) -- [Redis Streams consumer groups](https://redis.io/docs/latest/develop/data-types/streams/) · [NATS JetStream](https://docs.nats.io/nats-concepts/jetstream) - - +- [herdr — socket API](https://herdr.dev/docs/socket-api/) · [GitHub](https://github.com/ogulcancelik/herdr) +- `docs/MCP-Contract.md` §6 only — the rest of that file is pre-build design and does not + match the shipped tool names or routes. diff --git a/7-Use-Cases.md b/7-Use-Cases.md index 6c2093f..d780066 100644 --- a/7-Use-Cases.md +++ b/7-Use-Cases.md @@ -1,291 +1,308 @@ # 7. Use Cases -Concrete scenarios on the finalized architecture ([sole gateway](1-Architecture) + -[MCP-unified](2-Message-Server) + ccs-spawned workers). The **flagship** is a code-review -*conversation* between the primary Opus and a worker running the **GX10 vLLM** model — it -exercises every mechanism the system needs (trigger, discovery, ccs spawn, the ID contract, -and worker lifecycle), so we design it in full, then catalogue the rest. +This page shows what a lead actually does with fleet, using the tools `fleetd` really ships. +Every flow below uses real tool names and real parameters, taken from `FleetMcp.java` — the file +where every `fleet_*` tool is registered and its schema is defined. +For the full tool reference, see [Message Server](2-Message-Server). -Everything below holds the invariants: the primary stays env-CLEAN on Pro/Max, the worker's -model/account come from a **ccs profile**, and both talk only to `fleetd` over MCP. +The whole page rests on two invariants: the lead never sets `ANTHROPIC_BASE_URL` (it stays on +its own subscription), and both the lead and every member talk to each other only through +`fleetd`'s MCP tools. See [Architecture](1-Architecture) for those invariants in full. -## Flagship — a review conversation (Opus ↔ gx00 reviewer) +## Flagship — spawning a reviewer and having a conversation -You're working in Opus on code in a workspace and want a second pair of eyes — an actual -back-and-forth **review conversation**, not a one-shot lint — with a cheap worker on the GX10 -vLLM model. Opus never leaves its subscription; it just calls MCP tools. +You are working in the lead session and want a second opinion on a diff — an actual +back-and-forth, not a one-shot lint — from a member running on a different backend. The lead +never leaves its own session; it only calls MCP tools. ```mermaid sequenceDiagram - participant O as "Opus (primary, env CLEAN)" - participant B as "fleetd (gateway)" - participant C as "ccs + herdr" - participant W as "worker · gx00-vllm" + autonumber + participant L as Lead + participant F as fleetd + participant M as Member (reviewer) - O->>B: "fleet_list() — what workers can I use?" - B-->>O: "profiles:[gx00-vllm=DeepSeek, …] · sessions:[]" - O->>B: "fleet_send(to: reviewer@gx00-vllm, {kind: review.request, diff, focus})" - Note over B: "guard: ccs env gx00-vllm → base_url host on allowlist ✓" - B->>C: "spawn: ccs gx00-vllm claude (new herdr pane)" - C->>W: "worker Ready → inject review.request" - activate W - W->>B: "fleet_reply({kind: review.reply, findings:[…]})" - deactivate W - B-->>O: "tool result = findings" - Note over O: "reads findings, wants to dig into one" - O->>B: "fleet_send(session s_7f3a, {kind: question — why finding 3 high-sev})" - B->>W: "inject into the SAME reviewer pane (context still warm)" - activate W - W->>B: "fleet_reply({kind: answer, …})" - deactivate W - B-->>O: "tool result = answer" - Note over O,W: "same reviewer session reused across turns → the diff stays in its context" + L->>F: fleet_spawn(role="reviewer", profile="terra") + F-->>L: { sessionId, paneId, role: "reviewer", profile: "terra", status: "ready" } + L->>F: fleet_send(sessionId, content="Load the reviewer skill. Review PR #42 ...") — call blocks + F->>M: deliver content when the pane is idle + activate M + M->>M: reviews the diff + M->>F: fleet_reply(content="3 findings: ...") + deactivate M + F-->>L: tool result = the review + Note over L: reads the findings, wants to dig into one + L->>F: fleet_send(sessionId, content="why is finding 2 high severity?") — call blocks + F->>M: deliver — same pane, same warm context + activate M + M->>F: fleet_reply(content="because ...") + deactivate M + F-->>L: tool result = the answer + Note over L,M: same member session reused across turns —
the diff stays in its context until the lead calls fleet_stop ``` -*Figure: discovery → trigger → ccs-spawn (guarded) → structured reply → follow-up on the same -warm session. Every arrow from Opus is an MCP tool call to `fleetd`; the worker's provider -routing lives entirely in its ccs profile.* +*Figure: spawn → delegate → structured reply → a follow-up on the same warm session. The shape +follows the real schemas of `fleet_spawn` and `fleet_send` (`FleetMcp.java:1148-1174`, +`1088-1109`) and of `fleet_reply` (`FleetMcp.java:1213-1223`). +`sessionId` is the value `fleet_spawn` returns, and it is what identifies "the same member" on +the follow-up call — there is no separate reuse mechanism.* -The rest of this page is the five mechanisms this scenario needs. +The rest of this page walks through the pieces this flow is built from. -## Mechanism 1 — triggering a subtask (`fleet_send`) +## Mechanism 1 — delegating work (`fleet_send`) -Opus delegates with **one** tool call: +The lead delegates with one tool call. `fleet_send`'s schema (`FleetMcp.java:1088-1109`) has +these parameters: ```jsonc fleet_send({ - "to": "reviewer@gx00-vllm", // role@profile, or a live session id - "kind": "review.request", - "body": { "workspace": "/repo", "base": "main", "head": "HEAD", - "focus": ["correctness","security"], "instructions": "…" }, - "block": true // true (default) → reply as tool result; false → dispatch_id + "sessionId": "term_a7", // the member to delegate to (from fleet_spawn / fleet_list) + "content": "Load the reviewer skill. Review PR #42 for correctness and security.", + "wait": true, // default true → blocks for the reply; false → returns a ticket + "timeoutMs": 25000 // optional; how long to block (default), capped at 120000 }) ``` -`fleetd` resolves the target (spawn-or-reuse, below), injects the turn into the worker's -herdr pane gated on `agent_status`, and — with `block:true` (the default) — holds the call -open until the reply lands (worker `fleet_reply` or the turn-done `working→idle` edge), -returning it as the tool result. -Review turns are short, so they block; a long/detached job would use `block:false` and come -back via idle-pane injection ([Mode 2](1-Architecture#traffic-two-modes-across-the-gateway)). +There is no `to`, `kind`, `body`, or `block` field. `content` is plain text — the lead writes +clear instructions into it, the same way a person would write a task in chat. There is no +structured envelope underneath it; see *Ideas that were never built*, below. -## Mechanism 2 — worker discovery (knowing your choices) +`fleetd` holds the call open until the member replies with `fleet_reply`, or (a fallback) its +turn ends without ever calling `fleet_reply` — in which case the tool result is the scraped +transcript tail instead, clearly marked as such. A short review turn can block; a long or +detached job passes `wait:false` and is collected later with `fleet_poll`. Both are covered in +full, with the exact outcome list, in [Message Server](2-Message-Server). -Opus shouldn't hard-code pane ids or guess what's available. `fleet_list()` returns both -what's **runnable** and what's **live**: +To answer a member's `fleet_ask` question, or to message a peer lead, `fleet_send` takes +`turnId` or `coordId` instead of `sessionId` — also covered on that page. -```jsonc -fleet_list() → { - "profiles": [ // spawnable = the ccs worker roster (Mechanism 3) - { "name": "gx00-vllm", "model": "DeepSeek-V3", "host": "gx00.ltms.dev", "status": "available" }, - { "name": "ollama-local","model": "llama3.1", "host": "ollama.ltms.dev","status": "available" } - ], - "sessions": [ // live worker sessions right now - { "id": "s_7f3a", "role": "reviewer", "profile": "gx00-vllm", "agent_status": "idle" } - ] -} -``` +## Mechanism 2 — knowing your choices (`fleet_profiles` and `fleet_list`) -This is the "let Opus know its worker choices" surface: the catalogue is `fleetd`'s configured -roster of ccs worker profiles, and the live list is what it's already running. +The old version of this page said `fleet_list()` returns a `profiles` array. That call does not +exist. The two tools answer different questions: -## Mechanism 3 — spawning via ccs profiles (the flexibility) +- **`fleet_profiles`** (`FleetMcp.java:874-895`) answers "what backends can I spawn onto?" — the + configured **profiles**, which one is the default, and which are currently quarantined after a + usage-limit refusal: -**A worker's identity *is* a ccs profile.** `ccs [claude-args…]` launches `claude` -with that profile's account and provider routing, so `fleetd` never hand-assembles env — it -just picks a profile: + ```jsonc + fleet_profiles() → { + "profiles": ["terra", "sonnet", "local-llama"], + "default": "sonnet", + "quarantined": { // present only if something is actually quarantined + "terra": { "credentialId": "terra-key", "quarantinedForSeconds": 900 } + } + } + ``` -- **Config.** `fleetd.yaml` lists worker profiles by name; each maps to a ccs profile - (`ccs api` profile pointing at GX10 vLLM / Ollama), an expected model, and an allowlisted - `base_url` host. -- **Spawn.** `fleetd` tells herdr to open a pane and `send_text`: `ccs gx00-vllm claude` - (plus flags — workspace dir, an injected reviewer system prompt). No `ANTHROPIC_BASE_URL=…` - prefix; the profile carries it. -- **Guard (subscription boundary, in ccs terms).** Before spawning, `fleetd` runs - `ccs env ` and validates the **resolved** `ANTHROPIC_BASE_URL` host against the - off-subscription allowlist. A profile that resolves to `api.anthropic.com` (a subscription - profile) is **refused as a worker** — that would burn your quota. The primary Opus is *your* - session on *your* subscription profile; `fleetd` never spawns it. -- **Swap = repoint.** Changing the reviewer's model is choosing a different ccs profile — no - `fleetd` code change. Add a profile → it appears in `fleet_list().profiles`. +- **`fleet_list`** (`FleetMcp.java:920-973`) answers "who is actually running right now?" — the + live roster, split into `leads` (peer orchestrators) and `members` (spawned sessions): + + ```jsonc + fleet_list() → { + "leads": [ { "sessionId": "term_p1", "name": "opus", "status": "idle", "self": true } ], + "members": [ + { "sessionId": "term_a7", "paneId": "w9:pW", "role": "reviewer", "profile": "terra", + "state": "ready", "worktree": "/wt/cb-42", "branch": "worker/cb-42-3f2a" } + ], + "healthCoverage": "off" + } + ``` + + I confirmed the `members` row's field names (`sessionId`, `paneId`, `profile`, `role`, + `state`, and — when set — `worktree`/`branch`/`owner`/`agentSessionId`) in + `SessionManager.rosterView` (`SessionManager.java:552-582`), which `fleet_list` builds each + row from. + +`fleet_profiles` is the catalogue of what you *could* spawn; `fleet_list` is what is *actually* +running. An empty `members` array means no member is spawned right now — it says nothing about +which profiles exist. + +## Mechanism 3 — spawning a member (`fleet_spawn` and `profiles:`) + +**A member's backend is a configured profile, not a `ccs` profile.** There is no `ccs` anywhere +in the shipped configuration. The `FleetConfig` record (`FleetConfig.java:81-101`) has a +`profiles:` field instead — a `Map`, documented at `FleetConfig.java:37-41`. + +- **Config.** `fleetd.yaml` (or `bridged.yaml`, depending on the host) has a `profiles:` block. + Each named profile carries its own `baseUrl`, `model`, `tokenEnv` (the host env var holding + the auth token), and `argv` (the launch command — `claude` by default). The full field list is + `FleetConfig.Profile` (`FleetConfig.java:314-330`). +- **Spawn.** `fleet_spawn` takes `role` (what the member is for — `dev`, `reviewer`, or + `architect`; default `dev`) and `profile` (which backend — omit it for the default). These + are independent: a reviewer can run on the same profile as the developer it reviews. The + tool's own description says so (`FleetMcp.java:1148-1174`). +- **Guard.** Before spawning, `fleetd` checks the resolved `ANTHROPIC_BASE_URL` host against an + allowlist. A profile that resolves to a subscription host is refused — that would burn the + lead's own quota. The check is `SubscriptionGuard.assertWorker` + (`SubscriptionGuard.java`), called from `fleet_spawn`'s handler + (`FleetMcp.java:825-826`, turning a `GuardException` into a `subscription boundary: …` tool + error). +- **Add a backend = add a config entry.** Adding a profile to `profiles:` makes it appear in + `fleet_profiles`'s output — no code change. ```mermaid flowchart LR - CFG["fleetd.yaml
worker profiles"] --> PICK["pick profile
gx00-vllm"] - PICK --> GUARD{"ccs env host
on allowlist?"} - GUARD -->|"no (api.anthropic.com)"| REJ["refuse — would burn subscription"] - GUARD -->|"yes (gx00.ltms.dev)"| SPAWN["herdr: ccs gx00-vllm claude"] - SPAWN --> PANE["worker pane
Ready"] + CFG["fleetd.yaml
profiles: block"] --> PICK["fleet_spawn(profile)"] + PICK --> GUARD{"resolved baseUrl host
on the allowlist?"} + GUARD -->|"no"| REJ["refused — subscription boundary"] + GUARD -->|"yes"| SPAWN["herdr starts the profile's argv"] + SPAWN --> PANE["member pane
ready"] classDef ok fill:#2f855a,stroke:#22543d,color:#ffffff; classDef bad fill:#b7791f,stroke:#7b341e,color:#ffffff; class SPAWN,PANE ok class REJ bad ``` -*Figure: ccs profile selection is the spawn contract; `ccs env` is how the subscription guard -sees the resolved endpoint before committing.* +*Figure: `profiles:` is the spawn contract; the subscription guard checks the resolved host +before a pane is ever started. Verified against `FleetConfig.java:314-330` (the `Profile` +record) and `SubscriptionGuard.java` (read in full).* -## Mechanism 4 — the ID contract (envelope) +**Resuming a member, instead of respawning cold.** `fleet_spawn` also takes `resumeSessionId` — +a prior member's `agentSessionId` (shown by `fleet_list`) — to relaunch onto that same +conversation instead of starting fresh. This only works for a profile whose backend adapter +declares `Capability.SESSION_RESUME`; otherwise the spawn is refused rather than silently +starting cold (`SessionManager.java:220-235`). -Every message across the gateway is one envelope. It answers two questions each side needs: -**whose message is this** (`from`/`to`) and **what do I do next** (`kind`). +## Mechanism 4 — the member's own lifecycle limits -| Field | Meaning | -|---|---| -| `v` | envelope version (`1`) | -| `from` | sender identity — `primary:opus`, or `role@profile` / session id for a worker | -| `to` | recipient — `role@profile` or a live session id | -| `session` | `fleetd` session id (the worker session), e.g. `s_7f3a` | -| `turn` | monotonic counter within the session | -| `corr` | correlation id (`session#turn`) — `fleetd`'s rendezvous matches a reply to its request | -| `kind` | **verb.noun** telling the recipient what to do (vocabulary below) | -| `body` | kind-specific payload | +A deployment can bound how long a member session lives. `FleetConfig.Lifecycle` has two +independent knobs (`FleetConfig.java:637-639`): -**`kind` vocabulary (the "what to do next"):** +- **`idleTtlSeconds`** — a `SessionReaper` (`SessionReaper.java`) tears down a member session + that has sat idle (no delegated work) longer than this. +- **`contextCap`** — a member is force-released once it has served this many delegated turns, + win or lose. In `SessionManager.completeTurn` (`SessionManager.java:635-645`), hitting the cap + calls `release(paneId)` — the pane is torn + down, not checkpointed or automatically respawned. -| kind | direction | body | -|---|---|---| -| `review.request` | primary → worker | `{workspace, base, head\|diff, files?, focus[], instructions}` | -| `review.reply` | worker → primary | `{summary, findings:[{file,line,severity,issue,suggestion}], verdict}` | -| `question` / `answer` | either way | free-form follow-up tied to the same `session` | -| `ask` | worker → primary | worker-initiated blocker (needs a decision) — via `fleet_ask` | -| `ack` / `status` | control | delivery/liveness, no new turn | +I looked for an automatic "checkpoint state to disk, then respawn and continue the same task" +mechanism (the old page called this a "Ralph loop") and did not find one in the session +lifecycle code. What the code actually does when a member's session ends is closer to nothing +automatic at all: continuity, when it exists, comes from what the member itself left behind — +a git commit on its branch, a pull request — or from an explicit `fleet_spawn{resumeSessionId}` +call the lead makes on purpose (Mechanism 3, above). There is no automatic recycle step between +turns, so do not plan around one. -The worker learns this contract from a **reviewer skill / `CLAUDE.md` snippet** injected at -spawn ("you are a reviewer; requests arrive as `review.request`; reply with `fleet_reply` -`kind: review.reply`"). So both ends know the sender and the required next action without a -held-open conversation — one request in, one structured reply out. +## Ideas that were never built -## Mechanism 5 — worker lifecycle +The earlier version of this page described a structured message envelope — fields like `v`, +`from`, `to`, `session`, `turn`, `corr`, `kind`, `body` — as if every `fleet_send` carried one. +It does not exist. `fleet_send`'s `content` parameter is a plain string +(`FleetMcp.java:1096-1108`: `"content", stringProp(...)`), and there is no envelope-parsing code +behind it. In practice, structure comes from what the lead writes into +`content` — for example, naming a skill on the first line ("Load the reviewer skill.") so the +member knows the procedure to follow. This is documented practice (see the primary-side +directive, below), not a wire format. -For a review conversation the reviewer is **persistent within a work session** so the diff -stays warm across follow-ups, and recyclable so it never outgrows its context window: +If a structured envelope is ever built, it belongs back on this page as a real mechanism with +its own `file:line` citations. Until then, this section is a marker for readers who remember +the old design, not a thing to build against. -- **Spawn-on-demand** — first `review.request` for a workspace spawns the profile's worker. -- **Reuse** — subsequent turns (`question`, next-file review) target the same `session`; its - context carries the code under review. -- **Recycle (Ralph loop)** — on a context/idle cap `fleetd` checkpoints (commit + `STATE.md`) - and respawns fresh — **not** `claude --resume`. See - [Architecture → Worker lifecycle](1-Architecture#worker-lifecycle--the-ralph-loop). -- **Drain** — an `idle_ttl` (e.g. 20 min idle) or workspace close tears the pane down. +## Primary-side directive — when to delegate -Policy knobs in `fleetd.yaml`: `idle_ttl`, `context_cap`, `max_workers_per_profile`. - -## Primary-side directive — *when* to delegate (a `CLAUDE.md` snippet) - -Mechanism 4 gives the **worker** a `CLAUDE.md` snippet so it knows how to answer. The **primary** -needs the mirror image: a standing reminder to *reach for the bridge in the first place* instead of -spending subscription tokens on work a cheaper worker could do. The trigger is an environment signal -— the `fleetd` MCP tools being connected (e.g. a `FLEETD_MCP_URL` marker in the primary's env). -When that property is set, this instance is a **bridge primary** and should delegate by default. - -The decision the directive encodes: +A lead on a metered subscription should reach for a member before doing bulk or mechanical work +itself. The shipped version of this reminder lives in the project's own `CLAUDE.md`, not in an +environment variable — a lead confirms it is a lead (as opposed to a spawned member) by calling +`fleet_whoami`, not by checking a marker string. ```mermaid flowchart TD - T["a task arrives"] --> Q{"bridge available?
(FLEETD_MCP_URL set /
fleetd MCP connected)"} - Q -->|"no"| SELF["do it on the primary"] + T["a task arrives"] --> Q{"is the fleet MCP
mounted at all?"} + Q -->|"no"| SELF["do it in this session"] Q -->|"yes"| J{"needs YOUR judgment,
or bulk / mechanical / parallel?"} J -->|"judgment / interactive"| SELF - J -->|"bulk / mechanical / parallel"| DEL["fleet_send → worker"] + J -->|"can I write a brief a member
could succeed on?"| DEL["fleet_spawn + fleet_send"] classDef self fill:#2f855a,stroke:#22543d,color:#ffffff; classDef del fill:#2b6cb0,stroke:#1a365d,color:#ffffff; class SELF self class DEL del ``` -*Figure: an env property (bridge available) flips the default from "do it myself" to "delegate unless -it needs my judgment."* +*Figure: whether the fleet MCP is mounted flips the default from "do it myself" to "delegate +unless it needs my judgment or I cannot write a brief the member can succeed on."* -This is a *suggestion*, not wiring: `fleetd` never edits an agent's `CLAUDE.md` (that would cross -the subscription boundary in the wrong direction). The operator writes it. +### As-built — the `CLAUDE.md` "Bridge communication" section -### As-built — the `CLAUDE.md` **Bridge communication** section +This reminder is **shipped** as the first section of the project's own `CLAUDE.md`, ahead of +everything else — because a member runs in a git worktree of the same repository and inherits +that file verbatim. One file serves both roles, so the section starts by making the reader +establish which role it is, rather than assuming. -**Shipped** in the repo's own `CLAUDE.md` as the first section, ahead of the IDE workflow. The -design-era sketch above imagined a primary-only reminder keyed on an env marker; what shipped is -broader, because a worker runs in a **git worktree of the same repo** and therefore inherits the -same tracked `CLAUDE.md` verbatim. One file, both roles — so the section is **role-split**, and the -first thing it does is make the reader establish which role it is. +The section is split into layers, each reaching a different audience: -Why in `CLAUDE.md` and not somewhere else — the four layers, each with a different reach: +| Layer | Carries | Reaches | +|---|---|---| +| the launcher's reply charter | the one rule that must survive with no repo checkout: end every turn with `fleet_reply` | every spawned member, at launch, whatever its backend | +| `CLAUDE.md` → Bridge communication | protocol invariants and orchestration policy | the lead and every member that reads the repo | +| role playbook skills (e.g. `implementer`, `reviewer`) | per-job procedure | a member told to load one | +| the bridge's own docs (`docs/MCP-Contract.md` §6, this wiki) | design detail, flows | anyone who goes looking | -| Layer | Carries | Reaches | Cost to the reader | -|---|---|---|---| -| `REPLY_CHARTER` (`ClaudeCodeLauncher` / `OpenCodeLauncher`) | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every worker, at launch, both peer kinds | always in the system prompt | -| **`CLAUDE.md` → Bridge communication** | protocol invariants + orchestration policy | primary **and** every claude-code worker — tracked in git, so worktrees get it free | always in context | -| `.claude/skills/{implementer,reviewer}` | per-job procedure: commit/push/PR recipe, finding format | a worker told to load it | on demand | -| `docs/MCP-Contract.md` | design detail, flows, error model | anyone who goes looking | on demand | +A rule lives in exactly one of these layers — the outermost one that must obey it. Putting a +rule in the wrong layer is how a skill and the shipped section end up disagreeing. -The rule that keeps them from drifting: **a rule lives in exactly one layer — the outermost one that -must obey it.** Duplicating a rule into a skill is how the skill and the charter end up disagreeing. +What the section actually tells a reader to do: -What each part of the section pins down: +- **Confirm your role with `fleet_whoami`** before anything else, rather than guessing from a + side channel. +- **Never set or forward `ANTHROPIC_BASE_URL`.** Only the daemon puts a member off-subscription, + at spawn. +- **The bridge is the only channel.** Text printed to a terminal reaches nobody. +- **A lead**: split work into units with clear acceptance criteria, spawn every delegated unit + first, then send them all with `wait:false`, poll for results, verify a member's claim by + re-running the build rather than trusting a "clean" report, and never delegate the merge. +- **A member**: load the named skill, stay in the assigned scope, use `fleet_ask` only for a + decision that is genuinely the lead's, and end every turn with exactly one `fleet_reply`. -- **Role identification** — **`fleet_whoami`** (below) answers it authoritatively; the section - tells the reader to call it rather than infer. A fallback ladder remains for when it is - unreachable: the charter in the system prompt (reliable — the launcher appends it in the same - branch that mounts the MCP, so bridge tools without a charter is not a reachable state); the mount - name (`mcp__fleetd__*` for the primary's `.mcp.json` vs `mcp__bridge__*` for a worker's inline - config); `ANTHROPIC_BASE_URL` (one-way — Claude-model workers run clean, so absence proves - nothing); then **fail toward worker**. The two errors are asymmetric: a primary acting as a worker - gets refused by the authz gate — loud and self-correcting — while a worker acting as the primary - ends its turn silently and the sender receives nothing. -- **Invariants (both roles)** — never set/forward `ANTHROPIC_BASE_URL`; the bridge is the only - channel (terminal text reaches nobody); identity comes from the connection, never an argument - (mirroring [`Authz`](9-Implementation)); delivery is status-gated, one message per turn; never - touch herdr directly. -- **Primary** — **delegate-by-default**, then an intent→tool table over the shipped tools. The - default answer to "who does this?" is a worker: the test is not *"could I do this faster myself?"* - (usually yes) but *"can I write a brief good enough for a worker to succeed?"* — a wasted worker - turn costs a worker turn, while doing it yourself costs the primary's context and subscription. - Independent units fan out (one worktree worker each, all dispatched `wait:false`, then poll) - rather than serializing. Plus the policy the tool descriptions can't carry: pass `profile:` - explicitly; prefer `wait:false` + `fleet_poll`, since a blocking `fleet_send` is capped by the - *caller's own* MCP client timeout (~60s) long before a real task finishes; make every delegation - self-contained; **name the worker's skill in the first line of `content`** — that instruction is - what turns an opt-in skill into a reliable one; you are the merge gate; verify what a worker - claims rather than trusting a "clean" report. Delegating work never delegates responsibility. -- **Worker** — the turn contract: load the named skill, stay in scope, `fleet_ask` only for a - decision that is genuinely the lead's, end with exactly one `fleet_reply`, report only what you - actually ran, never merge, never commit `.mcp.json` or `wiki/`. - -**Known gap:** *opencode* workers never read `CLAUDE.md` — they receive `REPLY_CHARTER` as an -instructions file and nothing else. Any rule a non-Claude peer must obey belongs in the charter, not -in this section. The charter currently carries only the reply rule. +An opencode member never reads `CLAUDE.md` at all — it only receives the reply charter as an +instructions file. Any rule a non-Claude peer must obey has to live in the charter, not in this +section. ### `fleet_whoami` — asking instead of guessing -**Shipped.** The daemon always knew the answer: `ConnectionIdentity` maps a call's loopback peer PID -to a herdr pane, and every tool call is already gated on the `Principal` it yields. What was missing -was any way for an agent to *ask* — so an agent's own role had to be inferred from side channels the -daemon does not control, with a silent failure mode when the inference went the wrong way. +**Shipped.** `fleetd` already resolves every caller's identity for its own authorization checks +— `fleet_whoami` (no parameters) just reports that same resolution back as data, instead of +making the caller infer its own role from a side channel. Its handler is `FleetMcp.java:739-786`, +and it returns: -`fleet_whoami` (no params, `READ` in the [authz table](9-Implementation)) returns that same -resolved identity as data: +```jsonc +// called by a member +{ "role": "worker", "sessionId": "term_a7", "paneId": "w9:pW", "profile": "terra", + "state": "ready", "worktree": "/wt/cb-42", "branch": "worker/cb-42-3f2a" } -```json -{"role":"worker","sessionId":"term_a7","paneId":"w9:pW","profile":"ollama", - "state":"ready","worktree":"/wt/cb-517","branch":"worker/cb-517-3f2a","owner":"term_primary"} +// called by the lead +{ "role": "primary" } ``` -The primary gets `{"role":"primary"}` and nothing more — deliberately: handing it a `sessionId` it -does not own would invite exactly the forged `fleet_reply` that `Authz` refuses. A worker the -session registry has no record of — one that outlived a daemon restart — still gets `role` and -`sessionId`, which is the load-bearing part; the registry fields are simply absent rather than -invented. +The lead's own row deliberately carries no `sessionId` beyond what a lead session already has — +handing it one it does not own would invite the same forged-identity problem the connection-based +resolution exists to prevent. A member the session registry has no record of (one that outlived +a daemon restart) still gets `role` and `sessionId` — the fields that matter — with the registry +fields simply absent rather than invented. -Two properties worth keeping if this is ever reimplemented: +## More use cases (catalogue) -- **It reports, it does not decide.** The value comes from being the *same* resolution the - authorization gate uses, not a parallel one that could disagree with it. -- **It degrades toward the useful answer.** Never "unknown" when the role is known. +Same mechanisms, different tasks. Every row below uses only the tools covered above. -### The portable `CLAUDE.md` block — copy as-is +| Use case | Shape | +|---|---| +| **Delegated refactor or codegen** | `fleet_spawn(role="dev", profile=…, worktree=true, ticket=…)`, then `fleet_send(sessionId, content)`. The member edits, commits on its branch, and opens its own PR (see the `implementer` skill); the reply is a summary. The lead reviews and merges. | +| **Bulk test-writing or log triage** | Cheap, parallelizable work on a local profile: `fleet_spawn` several members, `fleet_send{wait:false}` each one, then `fleet_poll` each ticket. | +| **Parallel multi-file review** | One reviewer member per file or area, each `fleet_spawn{role="reviewer"}`, dispatched with `wait:false` and collected with `fleet_poll` — see [Team](6-Team). | +| **A member resumed onto its prior conversation** | `fleet_spawn{profile, resumeSessionId}` with the `agentSessionId` `fleet_list` reported for a member that already finished a turn — only on a profile whose adapter supports it (Mechanism 3). | +| **Cross-host lead coordination** | `fleet_send{coordId, content}` to a peer lead on another daemon, over a configured `coordinator:` broker — coordination between leads, never a task. See [Message Server](2-Message-Server). | -This is the **canonical text**, verbatim. Drop it into any project whose agents mount the bridge MCP; -it needs no editing — every project-specific detail was deliberately pushed out of it (see the two -notes after the block). Improvements land *here* first, then propagate to each project's `CLAUDE.md`. +## The portable `CLAUDE.md` block + +This is the **canonical text**, verbatim. Copy it into any project whose agents mount the fleet MCP +server. It needs no editing — every project-specific detail was deliberately pushed out of it, into +the two notes below the block. Improvements land *here* first, then go out to each project's +`CLAUDE.md`. + +The block still calls its own section "Bridge communication", and it still names the legacy +`mcp__bridge__*` mount in the role ladder. Both are kept on purpose: the text must stay +byte-identical with the copy in the `fleet/fleetd` repo, and a member spawned before the rename +really does report the old mount name. ```markdown ## Bridge communication (enforced — read this first) @@ -475,38 +492,26 @@ must obey belongs in the charter, not here. Two rules keep this portable, and both were learned by getting them wrong first: 1. **Nothing repo-local inside the block.** The first draft named `auth/Authz.java`, the - `.mcp.json`/`wiki/` commit exclusions, and the `implementer`/`reviewer` skills by name — all - meaningless in another project. Each moved to a **Project addendum** section that sits *below* - the block and never interleaves with it, so the block can be replaced wholesale without reading it. -2. **Every fallback signal must be one-way.** The role ladder originally read the MCP mount name in - both directions — `mcp__fleetd__*` ⇒ primary, `mcp__bridge__*` ⇒ worker. Only the second half is - real: the launcher hard-codes `bridge` for a worker's inline config, while a primary's mount is - named by whoever wrote that project's `.mcp.json`. A two-way reading of a one-way signal is a - confident wrong answer, so the block states only the direction that holds. + `.mcp.json` and `wiki/` commit exclusions, and the `implementer` and `reviewer` skills by name — + all meaningless in another project. Each moved to a **Project addendum** section that sits + *below* the block and never mixes into it, so the block can be replaced whole without reading it. +2. **Every fallback signal must be one-way.** The role ladder first read the MCP mount name in both + directions: one name meant primary, another meant worker. Only one half is real. The launcher + hard-codes the member's mount name (`PeerLauncher.java:34`, `MCP_MOUNT_NAME = "fleet"`), but a + primary's mount is named by whoever wrote that project's `.mcp.json`, so it can be anything. A + two-way reading of a one-way signal is a confident wrong answer, so the block states only the + direction that holds. -In the `claude-bridge` repo itself the block is treated as **shipped surface, not documentation**: -its `CLAUDE.md` carries a change-checklist mapping each part of the code (tool catalog, `Authz`, -`ConnectionIdentity`, `REPLY_CHARTER`, injector, worktree overlay, skills) to the part of the block -that change can invalidate, plus a sync check that fails if this template and that copy have drifted. -A code change that silently falsifies the block is an incomplete change — the agents reading it have -no other source. +In the `fleet/fleetd` repo itself the block is treated as **shipped surface, not documentation**. +That repo's `CLAUDE.md` carries a change-checklist mapping each part of the code — the tool +catalogue, `Authz`, `ConnectionIdentity`, `REPLY_CHARTER`, the injector, the worktree overlay, the +skills — to the part of the block that change can invalidate. It also carries a check that fails if +this template and that copy have drifted apart. A code change that silently makes the block false is +an incomplete change, because the agents reading it have no other source. -## More use cases (catalogue) - -Same machinery, different `kind`/lifecycle: - -| Use case | Shape | -|---|---| -| **Delegated refactor / codegen** | `fleet_send(kind: task.request)`, blocking; worker edits + commits on a branch; reply = diff summary. Opus reviews/merges. | -| **Test writing / log triage** | Bulk, cheap, parallelizable — a local worker profile; fire several concurrently and reduce. | -| **Parallel multi-file review** | A **fleet** of workers, one per area, fan-out/gather — see [Team](6-Team). | -| **Perpetual worker** | Long-running across many Ralph recycles, state on disk; woken by async injection ([Mode 2](1-Architecture)). | -| **Event-bus automation** | A webhook hits `fleetd`'s REST ingress → injects an idle worker → result back to Opus or a chat bridge. No human in the loop. | - ## Related -- **[Architecture](1-Architecture)** — invariants, two modes, lifecycle state machine. -- **[Message Server](2-Message-Server)** — the `fleetd` design these mechanisms live in. -- **[Roadmap](8-Roadmap)** — the staged plan + tickets that build this review scenario first. -- **[Team](6-Team)** — the fleet/orchestration layer above a single review. +- **[Architecture](1-Architecture)** — the invariants every flow on this page holds to. +- **[Message Server](2-Message-Server)** — the full `fleet_*` tool reference and REST surface. +- **[Team](6-Team)** — running more than one member at once. diff --git a/8-Roadmap.md b/8-Roadmap.md index 6a3fe9a..20fc095 100644 --- a/8-Roadmap.md +++ b/8-Roadmap.md @@ -1,651 +1,220 @@ -# 8. Roadmap & Delivery +# 8. Roadmap and delivery record -A **walking-skeleton-first** plan: Stage 1 delivers the flagship -[review scenario](7-Use-Cases#flagship--a-review-conversation-opus--gx00-reviewer) end-to-end -(thin but real — Opus gets a review back from a gx00 worker over MCP), then later stages add -reply fidelity, the subscription guard, lifecycle, pluggable peers, and hardening. This re-scopes the -[Message Server](2-Message-Server#build-plan-milestones) M0–M4 milestones around the scenario, -so there's something usable after every stage. +> **Read this page as a delivery record, not as a system guide.** +> Each item appears in one status group only. The groups separate code that exists +> from work that was planned or rejected. For the current operator surface, use +> [Features](11-Features) and [User Guide](13-User-Guide). -## Stages +## Status at a glance -```mermaid -gantt - title Indicative delivery sequence (relative durations, not committed dates) - dateFormat YYYY-MM-DD - axisFormat %b %d - section Skeleton - Stage 1 — review walking skeleton :s1, 2026-07-14, 10d - section Fidelity - Stage 2 — envelope + guard + reply :s2, after s1, 8d - section Scale - Stage 3 — lifecycle + discovery + fleet :s3, after s2, 10d - section Pluggable peers - Stage 4 — PeerLauncher SPI Stage A :s4, after s3, 10d - section Harden - Stage 5 — auth · metrics · CI · systemd :s5, after s4, 8d -``` +This page was checked against the source tree on 2026-08-31. A claim marked +**built and live** has code evidence below. “Live” means it is in the built daemon +path, not that this page checked a particular host deployment. -| Stage | Goal | Delivers (usable outcome) | -|---|---|---| -| **1 — Walking skeleton** | One review, happy path | Opus mounts `fleetd` (MCP), calls `fleet_send` with a diff, gets a review back from a real `ccs gx00-vllm claude` worker. Single hardcoded profile, same host, no guard/lifecycle. | -| **2 — Contract + guard** | Trust the reply, trust the boundary | Structured [envelope](7-Use-Cases#mechanism-4--the-id-contract-envelope) + worker `fleet_reply`; reply rendezvous; subscription guard via `ccs env`; reviewer skill. | -| **3 — Lifecycle + discovery** | Reuse, recycle, choose | Session manager (spawn/reuse/recycle Ralph loop, `idle_ttl`); `fleet_list` roster+live; multiple profiles. (`fleet_ask` landed early in Stage 2.) | -| **4 — Pluggable peers** | PeerLauncher SPI + two adapters | ✅ Extract `PeerLauncher` SPI in-tree; `ClaudeCodeLauncher` (Stage A) **and** `OpenCodeLauncher` (Stage B, CB-402) routed by a `kind:` discriminator through `CompositePeerLauncher`. Stage C (dynamic external plugin loading) remains future work, gated by a trust/capability model. | -| **5 — Harden** | Production shape | ✅ Bearer auth + fail-fast exposure guard, `/metrics` + `/healthz`, mock-socket CI on the Gitea runner, launchd + systemd units, per-session authz + audit log. TLS deliberately terminates at a reverse proxy, not in the daemon. | - -## Tech stack - -Consolidated from [Message Server](2-Message-Server#proposed-tech-stack); the ccs pieces are new. - -| Concern | Choice | Note | -|---|---|---| -| **Language** | **Java 21+ (virtual threads)** | Loom fits the blocking socket + MCP + queue + SSE fan-in; **GraalVM `native-image`** recovers the `scp`+systemd single-binary deploy. Kotlin OK (same JVM). | -| **REST API core** | **Javalin** (or Spring MVC) + SSE | **The implementation & testability surface** — every feature is a REST endpoint (see [Testability](#testability--the-rest-api-is-the-contract-surface)). | -| **MCP server** | **Official MCP Java SDK**, streamable-HTTP (Jetty/Spring) | A **thin adapter over the REST core** — the Claude-facing face of the same features. | -| **herdr client** | JDK `UnixDomainSocketAddress` + `SocketChannel`, NDJSON (Jackson) | Native UDS, no dep; event stream on a virtual thread. Pinned to herdr **0.7.0 / protocol 14** ([verified API](2-Message-Server#the-herdr-control-contract-what-fleetd-drives)). | -| **Worker spawn** | **`ccs claude`** into a herdr pane | Profile = worker identity; no env-prefix. | -| **Subscription guard** | **`ccs env `** → resolved `base_url` host allowlist | Boundary check in ccs terms. | -| **Config** | **YAML via Jackson** + env | `fleetd.yaml`: worker profiles, allowlist, bind addr, lifecycle knobs. | -| **Internal queue** *(future)* | **Redis Streams via Lettuce** (ack + visibility) | Below the gateway; in-JVM queue OK single-host. | -| **Observability** | **SLF4J+Logback**; **Micrometer**→Prometheus `/metrics`; `/healthz` | Stage 5. | -| **Supervision** | systemd unit (`java -jar` or native-image), ordered after herdr | Stage 5. | -| **Testing** | **JUnit 5**; **mock UDS herdr** server; **REST contract tests per feature**; **contract test vs real herdr 0.7.0**; fake `ccs`/`claude` stubs | See [Testability](#testability--the-rest-api-is-the-contract-surface). | - -## Testability — the REST API is the contract surface - -Every feature is implemented as a **REST endpoint first**; the **MCP tools are thin adapters -over those endpoints** (already the design in [Message Server](2-Message-Server#api-surface-server-face)). -That inversion is what makes the system testable *against expectations*: - -- **Each feature = one endpoint = one acceptance test.** You verify behaviour by calling REST - with an input and asserting the response — **no Claude session, no MCP client, no TTY in the - loop.** Deterministic and CI-friendly. -- **MCP is validated by parity.** For each tool, one test asserts the MCP call and the REST call - return the same result for the same input. If REST is green and parity holds, MCP is correct by - construction — we don't re-test business logic through the MCP layer. -- **Expectations are pinned to reality, not docs.** herdr behaviour is asserted by a **contract - test against the running herdr 0.7.0** (`ping.protocol == 14`, `workspace.list`/`pane.list` - shapes), so the socket client can't silently drift from the actual server. - -```mermaid -flowchart TB - subgraph tests["Test layers (all hit the REST surface or below)"] - FT["feature/acceptance tests
HTTP → REST endpoint"] - PT["parity tests
MCP tool == REST route"] - UT["unit tests
mock UDS herdr · fake ccs/claude"] - CT["contract test
vs real herdr 0.7.0"] - end - MCPC["MCP client (Claude)"] -->|"thin adapter"| REST["REST API (feature core)"] - REST --> CORE["session mgr · injector · guard · rendezvous"] - CORE --> HERDR["herdr socket client"] - FT --> REST - PT --> MCPC - PT --> REST - UT --> CORE - CT --> HERDR - classDef core fill:#2f855a,stroke:#22543d,color:#ffffff; - classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff; - class REST,CORE,HERDR core - class MCPC ext -``` - -*Figure: features live in the REST core; MCP wraps it; tests target REST (feature), MCP-vs-REST -(parity), the core with a mocked socket (unit), and the real herdr (contract).* - -### Feature ⇄ endpoint ⇄ test map (the acceptance surface) - -| Feature | REST endpoint | MCP tool | Acceptance test | -|---|---|---|---| -| Deliver a turn (blocking) | `POST /sessions/{id}/message` | `fleet_send` | reply returned; `working→idle` unblocks; timeout → 202 | -| Worker status | `GET /sessions/{id}/status` | `fleet_status` | matches herdr `agent_status` | -| Detached dispatch + drain | `POST …/message?block=false` · `GET /sessions/{id}/status` (pending drain) | `fleet_send(block:false)` · `fleet_status` | `dispatch_id` issued; reply retrievable | -| Worker reply | `POST /sessions/{id}/reply` | `fleet_reply` | resolves the awaiting request by `corr` | -| Worker question | `POST /sessions/{id}/ask` | `fleet_ask` | surfaces to primary; parks worker | -| Discovery | `GET /sessions` | `fleet_list` | roster + live match config/herdr | -| Spawn (guarded) | `POST /sessions` | *(internal)* | rejects on-subscription profile; accepts allowlisted | -| Health | `GET /healthz` · `GET /metrics` | — | liveness + Prometheus | - -Every ticket's **Acceptance** below is written to be executed against these endpoints. - -## Tickets by stage (2–5) - -Compact scope; expand into detailed tickets when a stage starts (as Stage 1 is below). - -| Stage | Tickets | +| Status | Meaning | |---|---| -| **2** | `CB-201` envelope schema + codec · `CB-202` worker `fleet_reply` tool + reviewer skill · `CB-203` reply rendezvous (corr match; resolve on reply *or* the `working→idle` edge) · `CB-204` subscription guard via `ccs env` + allowlist · `CB-205` blocked-worker path (`fleet_ask`) | -| **3** | `CB-301` ✅ session manager (spawn/reuse/recycle) · `CB-301-ext` ✅ per-worker git worktree + config-parity overlay (`97ecc71`) · `CB-302` ✅ worker checkpoint — **shipped as commit→push→**_**worker-opened PR**_ (`64e70ef`: repo-scoped forge-token injection + the implementer skill), which **supersedes** this row's original `STATE.md`-file framing; see `docs/Worker-Git-Workflow.md` · `CB-303` ✅ `idle_ttl`/`context_cap`/drain · `CB-304` ✅ `fleet_list` roster+live (`9fe04bf`) · `CB-305` ✅ multi-profile routing | -| **4** | `CB-401` ✅ PeerLauncher SPI Stage A — extracted in-tree, one adapter (`ClaudeCodeLauncher`), core uses the `PeerLauncher` interface, main @ `3aa69a9`. `CB-402` ✅ **Stage B complete** (`ded226a`, dogfooded 2026-07-29) — `OpenCodeLauncher` as the SPI-proving second adapter: shares none of Claude's private seams (no `ANTHROPIC_BASE_URL`, no `SubscriptionGuard`), mounts the bridge MCP via a generated `OPENCODE_CONFIG`, and is routed by `kind:` through `CompositePeerLauncher`. Live spawn→send→`fleet_reply`→teardown verified against opencode 1.18.5. Stage C — dynamic external plugin loading, future, gated by trust/capability model. | -| **5** | ✅ **CB-501–505 landed** (the stage as originally scoped): `CB-501` bearer auth + non-loopback-bind fail-fast (TLS at a proxy, not in-JVM — see `docs/CB-5xx-Hardening.md` D3) · `CB-502` `/metrics` (zero-dependency Prometheus renderer, D4) + `/healthz` · `CB-503` mock-socket CI (`.gitea/workflows/ci.yml`) · `CB-504` launchd agent + systemd unit + herdr-socket startup wait · `CB-505` per-session authz table + audit log. **The 5xx line did not stop there** — `CB-506`…`CB-525` shipped after this row was written; see [Stage 5 continued](#cb-506525--stage-5-continued-as-built) below and [Features](11-Features) for the operator-facing ones. | +| **Built and live** | The current daemon creates or registers the feature. | +| **Built, not switched on** | The code is complete, but no host enables it today. | +| **Planned** | The old roadmap proposed it. It is not current daemon behaviour. | +| **Tried and dropped** | Research or an old design that was not shipped. | -## CB-401 — Peer Launcher SPI (Stage 4) +## Built and live -Stage A has landed on main: +### The MCP fleet surface -- ✅ **SPI extracted in-tree** (`dev.ltms.fleetd.peer`): `PeerLauncher`, `PeerHandle`, `SpawnRequest`, `Capability`. -- ✅ **One adapter** — `worker.ClaudeCodeLauncher` implements `PeerLauncher`; it is the renamed/adapted `WorkerService` and still performs guard-checked spawn, orphan reap, and teardown. -- ✅ **Core decoupled** — `session.SessionManager` now depends on the `PeerLauncher` interface and keys its registry on `PeerHandle.id()` (equal to herdr `paneId` for the Claude adapter, so no value change). -- ✅ **Behaviour-preserving** — full green gate on main @ `3aa69a9`, **183 tests**. +`FleetMcp` registers these tools: `fleet_ack`, `fleet_ask`, `fleet_list`, +`fleet_poll`, `fleet_profiles`, `fleet_reply`, `fleet_send`, `fleet_spawn`, +`fleet_status`, `fleet_stop`, and `fleet_whoami`. +`fleet_read` is not registered. The `fleet_send` schema also supports `coordId` +for peer-lead coordination. See `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1088-1245`. -**Stage B / CB-402 — landed on main (`ded226a`) and live-dogfooded 2026-07-29 (issue #7 closed):** +The current tool set covers a delegation, a worker question, an asynchronous +ticket, inbox acknowledgement, member lifecycle, profile discovery, and caller +identity. The tool descriptions are the current contract. See +`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1088-1245`. -- ✅ `HerdrPeerLauncher` base extracted; `OpenCodeLauncher` implements the three divergent hooks. -- ✅ `kind:` discriminator on worker profiles; `CompositePeerLauncher` routes spawn/stop/reap/list by kind. -- ✅ The SPI is proven provider-neutral: opencode uses **none** of Claude Code's private launch seams. -- ✅ **Live dogfood complete.** Spawn → CB-306 readiness gate → `fleet_send` → structured - `fleet_reply` (`replySource: "reply"`, not the completion fallback) → teardown, all through the - REST surface against opencode **1.18.5**. The schema-drift risk did not materialise: the adapter - was designed against 1.1.31 and its generated `OPENCODE_CONFIG` still validates unchanged. - Provider question resolved — opencode's gateway serves **free-tier models with zero credentials**, - so no key was needed and no `guard` entry applies. Full run: `docs/CB-402-OpenCode-Adapter.md` §8. +### Two launcher kinds -**Stage C (future):** dynamic external plugin loading (`ServiceLoader`/jar discovery). Gated by a trust/capability model — a launcher runs at daemon privilege and can inject env/tokens into peers, so third-party plugins are not enabled without that model. +Profiles support `claude-code` and `opencode`. `Fleetd` builds launcher adapters +for the configured kinds and puts them behind `CompositePeerLauncher`. See +`fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:159-201` and +`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:265-270,314-345`. -## CB-5xx — Stage 5 hardening (as-built) +This is the shipped replacement for the old Claude-only plan. A member’s role and +profile stay separate. The role states whether it is an architect, developer, or +reviewer. See `fleetd/src/main/java/dev/ltms/fleet/peer/MemberRole.java:6-53`. -The single-host close-out, done **before** cross-host rather than after, because CB-308's own -gating concern is the trust model and it inherits whatever identity shape lands here. Full design -and decision record: **`docs/CB-5xx-Hardening.md`**. +### Member worktrees and developer pull requests -The finding that shaped the stage: `fleetd` had **exactly one security control — the loopback -bind**. `ConnectionIdentity` resolves a worker from its connection (unforgeable), but *any* caller -that was not a recognised worker pane — including, had the bind ever widened, an arbitrary remote -client — was treated as **the primary**, the most privileged role on the bus. CB-501 inverts that -default: `ANONYMOUS` is now the fallback and `PRIMARY` must be established. +`fleet_spawn` can request an isolated worktree. The tool states that a developer +member opens its own pull request. See +`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1148-1173`. -- **CB-501 ✅ auth.** `auth.mode: loopback-trust` (default, the historical behaviour named honestly) - or `token` (bearer required of every non-worker caller). Worker identity is *never* token-gated, - so enabling auth cannot lock the fleet out of `fleet_reply`. Constant-time token comparison. - **The highest-value line in the stage is a startup check:** a non-loopback `bind.host` under - `loopback-trust` now *refuses to start* rather than silently promoting every reachable client to - primary. TLS terminates at a reverse proxy by design (D3), not in the JVM. -- **CB-502 ✅ `/metrics`.** Zero new dependencies — a ~150-line Prometheus text renderer instead of - Micrometer (D4), because this pom already hand-reconciles Jackson 2/3 and a Jetty BOM and carries - four accepted-CVE advisories, and the project's mandated dependency CVE gate could not be run. - Instrumented at `MessageService` — the single funnel both surfaces share. -- **CB-503 ✅ CI.** `.gitea/workflows/ci.yml` on the (already-running) Gitea Actions runner. Needs - no contract-test flag: the pom's `default-excludes` profile already excludes `@Tag("contract")`, - so a plain `mvn -B clean install` *is* the mock-socket surface. -- **CB-504 ✅ supervision.** launchd agent (the real target — this host is macOS, there is no - systemd) **and** a systemd unit for the Linux gateways CB-308 adds. Ordering directives are - advisory; the actual fix is that fleetd now **waits up to 30s for the herdr socket and then - serves degraded** instead of crashing into a restart loop on a boot-order race. -- **CB-505 ✅ authz + audit.** The role table enforced on **both** entry paths — and that plural is - the point. The wiki has long described MCP as "a thin adapter over the REST core"; at code level - it is not. `BridgeMcp` calls the service layer *directly*, and `/mcp` is a raw servlet on Jetty's - context handler that never passes through Javalin's `before` filter. Enforcing only at REST would - have left `/mcp` wide open. The load-bearing rule is *own-session-only*: a worker may reply or ask - only as itself (already structurally true over MCP, newly true over REST, which had simply trusted - the session id in the URL path). Audit lines are JSON to a dedicated appender and **never carry - message content**. +`Fleetd` gives `SessionManager` a `GitWorktrees` instance. `GitWorktrees` creates +the linked worktree, configures its credential helper, and neutralizes inherited +tool configuration. See `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:218-229` +and `fleetd/src/main/java/dev/ltms/fleet/session/GitWorktrees.java:111-132`. -## CB-506–525 — Stage 5 continued (as-built) +The developer role’s contract says it provisions a worktree, commits, pushes, and +opens its own pull request. See +`fleetd/src/main/java/dev/ltms/fleet/peer/MemberRole.java:37-44`. -The 5xx line kept running after the stage's original five tickets. Two things drove it: **dogfooding -the bridge on its own development** surfaced defects the test suite could not (a primary that -classified itself as a worker, a worker that edited the wrong checkout), and a **coverage push** -turned "it works when I try it" into guarded behaviour. Operator-facing entries are catalogued in -[Features](11-Features); this section is the ticket-level record. +### Durable replies use AMQP when configured -**Capabilities** — what an operator gained. +The current durable inbox choice is AMQP. `broker` is the configuration block for +an external AMQP broker. When it is absent, the inbox is in memory. See +`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:49-50,656-714`. -| | | -|---|---| -| `CB-508` | an opencode profile can pin its own OpenAI-compatible endpoint | -| `CB-511` | workers get a real toolchain — the daemon's `PATH` is propagated, plus a per-profile `env:` map. Before this a worker inherited whatever `PATH` the herdr *server* was started with, which on a long-lived herdr can predate your toolchain entirely and leave workers unable to run `mvn` at all | -| `CB-517` | `fleet_whoami` — a session asks the daemon for its own role instead of inferring it; and the `CLAUDE.md` bridge block became a portable charter copied verbatim into every project that mounts the bridge. LavinMQ pinned as a durable, self-restarting broker | -| `CB-518` | weighted placement (`placement: weighted`, per-profile `weight` / `maxLoad`), smooth weighted round-robin with failover to the next candidate | -| `CB-521` | herdr adapter ported to **protocol 19** (herdr 0.8.0) — a hard version coupling, see the note below | -| `CB-522` | `primary.terminal` — the primary may run *inside* a herdr pane. Without the pin the pane lookup reads it as a worker and refuses every orchestration verb, and the failure is self-locking: the learned terminal is populated by the very calls being refused | -| `CB-525` | a provisioned worktree's tool surface is isolated to what its launcher mounts | +At startup, `Fleetd` selects the reply inbox from `broker` and can open +`AmqpReplyInbox`. See `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:399,811` +and `fleetd/src/main/java/dev/ltms/fleet/msg/AmqpReplyInbox.java:79-145`. -**Contracts** — internal shape changes a maintainer needs to know. +This deployment enables AMQP reply durability through its configured LavinMQ +broker. Without a usable `broker` configuration, the daemon falls back to the +in-memory inbox. See `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:656-714,1915`. -| | | -|---|---| -| `CB-507` | fixed an NPE (HTTP 500) on a worktree spawn with no cwd, plus regression tests | -| `CB-516` | a delegation now *fails* when its worker session is released, rather than hanging | -| `CB-519` | `PeerHandle.id()` is a host-unique opaque UUID, deliberately decoupled from the herdr pane id: the id is the registry/routing key and must never collide across daemon processes on one host, while the pane id stays a launcher-private placement/teardown coordinate | -| `CB-520` | `ReplyInbox` split into explicit own/release and publish halves | -| `CB-524` | placement is reproducible across restarts — `Map.of`/`Map.copyOf` iteration order is salted per JVM run, which silently discarded YAML definition order and made equal-weight placement differ run to run | +### Cross-host lead-to-lead coordination -**Quality infrastructure** — invisible, but it is why the above is trustworthy. +Lead-to-lead coordination works through a separate `coordinator:` block. A lead +uses `fleet_send` with `coordId`; the tool publishes to the peer lead’s durable +mailbox. See `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:597-641,1088-1108` +and `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:74-78,718-760`. -`CB-505 fix` audit lines are valid JSON · `CB-506` the test suite stays out of the production audit -log · `CB-509` JaCoCo coverage reporting · `CB-510` SessionReaper 0 → 86.7% · `CB-512` -`fleetd_push_nudges_total` increments wired · `CB-513` the MCP-side authorization gate, 27.4 → -57.4% · `CB-514` MessageService timeout/answer/poll/lock edges · `CB-515` the turn-attribution -guards regression-protected · `CB-521` the AMQP contract test runnable both locally and in CI. +This feature is only for coordination between peer leads. It is not a way to +delegate work sideways. The restriction is in the tool description at +`fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1093-1107`. -> **Version coupling (CB-521).** `fleetd` speaks one herdr wire protocol; the adapter is ported -> wholesale on a bump, with no negotiation or compat shim. The daemon can therefore be perfectly -> healthy — `/healthz` ok, `fleet_whoami` resolving, profiles listed — while **every spawn fails**, -> because health only pings herdr and never checks that the adapter and the binary agree. Symptom: -> `invalid_request: missing field `. After any restart onto a jar carrying an adapter change, -> verify with a real `fleet_spawn`, not with `/healthz`. +This deployment has a `coordinator:` block with `selfId`, so its peer-lead +mailbox is enabled. An absent or empty block opens no mailbox. See +`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:746-760,1944-1948`. -## Release 1.1 — the single-host close-out (open) +### Current delivery path -`v1.0.0` was tagged **2026-08-10**, message *"One leader, one host, complete"*. Since then **211 -commits** have landed on `main` with **no tag**. So the stage tables above are true and still leave -the obvious question unanswered: what is actually left before the single-host story can be called -finished? - -**The answer, as of 2026-08-17: nothing. Release 1 is finished.** The milestone is closed — 20 of 20, -every ticket merged, deployed, and verified on the running daemon. CB-596 was the last one, and it -was closed on a measurement taken inside a live member pane, not on a merge. - -**`v1.1.0` is tagged and pushed**, on `7d41ccc`, with the build re-run on the tagged commit: 870 -tests, 0 failures. **217 commits since `v1.0.0`.** Gitea milestone 16 is closed and no pull requests -are left open. The next line of work is milestone 17 — **2.0, one operation centre, many hosts.** - -Gitea milestone: **`1.1 — single-host close-out`**. The admission rule is one sentence — **if it -would still be broken with exactly one host, it belongs in 1.1.** Read strictly, that rule sent four -tickets to 2.0, not one — see *Deferred to 2.0* below. The distinction it turns on: a capability that -is **missing** is not the same as one that is **broken**. Cost-first placement is missing, and there -is a workaround; a config key that is accepted and silently does nothing is broken. +The current source builds the delivery path around the following dependencies. +This is a code map, not a promise about a host’s running configuration. ```mermaid flowchart LR - subgraph r1["release 1.1 — one fleetd per host"] - sup["supervision
CB-594 · CB-600"] - dur["durability claim
CB-527 · CB-528"] - sec["credential scope
CB-593"] - bugs["nudge scheduling
CB-590 · CB-598"] - cfg["config accuracy
CB-597 · CB-599 · CB-602"] - flaky["green build
CB-601 · CB-603"] - ask["fleet_ask reach
CB-582"] - docs["catalogue debt
CB-595"] - val["config validation
CB-604 · CB-606"] - op["needs the operator
CB-596"] - wip["snapshot pruning
CB-586"] - end - subgraph r2["release 2 — one centre, many hosts"] - fed["CB-308 federation"] - defer["CB-589 · CB-548 · CB-605"] - found["found by CB-596
CB-607 · CB-608"] - end - r1 --> r2 - classDef done fill:#2f855a,stroke:#22543d,color:#ffffff; - classDef todo fill:#b7791f,stroke:#7b341e,color:#ffffff; - class sup,dur,sec,bugs,cfg,flaky,ask,docs,val,wip,op done - class fed,defer,found todo + P["Profile kind"] --> L["Launcher adapter"] + L --> S["SessionManager"] + S --> W["GitWorktrees (when requested)"] + M["FleetMcp tools"] --> S + B["broker config"] --> R["ReplyInbox"] + C["coordinator config"] --> H["LeadChannel"] + M --> H ``` -*Figure: the 1.1 milestone by theme, as of 2026-08-17. Everything on the left is green — 20 of 20 -closed, merged, deployed and verified on the running daemon. The right is release 2, and it grew by -two: CB-596's own gap detector found both of them on its first live spawn.* +*Figure: current code links profiles, member sessions, optional worktrees, reply delivery, and peer-lead coordination.* -### Done — merged and verified +Evidence for the launcher and session links is +`fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:159-229`. Evidence for the MCP +links is `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1088-1245`. -| Theme | Tickets | What it was | -|---|---|---| -| **Supervision** | CB-594 (#80) · CB-600 (#91) | The launchd unit shipped in `deploy/` and had never been installed, and launchd does not source a login shell — so a supervised daemon got no `WORKER_GITEA_TOKEN` and no `AI_GATEWAY_TOKEN`. The operator had to pick supervision *or* a working fleet. CB-600 then closed the gaps that only bite once the agent is loaded: a log path the script and the plist could silently disagree about, and a failed `load` leaving the agent stopped **and** persistently disabled. | -| **Durability** | CB-527 (#10) · CB-528 (#11) | No `basicQos`, so the backlog lived in JVM heap rather than on the broker; no publisher confirms, so a publish to a missing queue was a silent black hole. The wiki promised more durability than the code delivered. Both closed, with contract tests running against a real broker in CI. | -| **Credential scope** | CB-593 (#79) | The decision about which inherited credentials a member may keep is now recorded, so a deliberate choice no longer looks like an oversight. The canonical block also stopped claiming a member mounts only the bridge. | -| **Nudge scheduling** | CB-590 (#75) · CB-598 (#87) · CB-582 (#61) | Two schedules could inject into one lead pane at once. Then the reminder budget turned out to be a counter carried forward with no memory of *which item* it counted, so work arriving during the backoff window inherited an already-capped count and was never nudged once. CB-582 added the third source: a worker paused in `fleet_ask` now nudges the lead itself, and `fleet_status` and REST both show the open question. **The ~55s window is closed, not removed** — still do not brief a worker to "ask me". | -| **Config accuracy** | CB-597 (#85) · CB-599 (#89) · CB-602 (#96) · CB-604 (#102) | Two documented knobs that are read by nothing; a capacity refusal that surfaced as a bare HTTP 500; no test at all in the code→example direction, so a brand-new key could ship undocumented; and an unknown `kind:` accepted silently and routed to the wrong adapter, which now refuses at config load. | -| **Catalogue debt** | CB-595 (#81) | `wiki/11-Features.md` had fallen about fourteen entries behind, worst on the entries that changed what a config key *means*. `fleetd.yaml` is gitignored, so that page is the only place an operator could learn them. Cleared — and it turned up two defects on the way (CB-604, CB-606). | -| **Green build** | CB-601 (#95) · CB-603 (#100) | Two flaky tests, same root: a test that observes an asynchronous loop must be safe against that loop's thread, and neither the compiler nor a green build will say it is not. | -| **Config validation** | CB-604 (#102) · CB-606 (#106) | Four config fields shared one shape: lower-cased in a compact constructor, then compared against exactly **one** string, so a typo fell through to the other branch in silence. The worst was `auth.mode` — a typo of `token` behaved as `loopback-trust`, and `validateAuthExposure()` only fires on a **non-loopback** bind, so the common loopback bind hid it end to end and the daemon authenticated nobody while the config said otherwise. All four now refuse at load, naming the field, the value, the accepted set, and what would have happened. | -| **Snapshot pruning** | CB-586 (#67) | Nothing pruned `refs/wip/*`, so CB-578 stage C's snapshots pinned their whole trees forever. The rule that landed needs **both** conditions: the tree is already reachable from `main`, and the ref is older than 24h. Reachability is the floor — a snapshot exists because the work was committed nowhere else, so a plain TTL would delete the only copy. `/members` now reports `wipRefs{count,costBytes}`. The sweep shipped **dead**: a `Long.MIN_VALUE` "never yet" sentinel overflowed the interval gate, which returned before the assignment that would have fixed it, so it never ran once — and every unit test passed, because they all called the seam directly and walked around the gate. | +## Built, not switched on -### The last one — CB-596, closed on a measurement +### Separate herdr daemon for members -| Theme | Ticket | The point | -|---|---|---| -| **Credential scope** | CB-596 (#82) | CB-592 blocks one credential name. A member's pane runs a login shell that re-sources the operator's whole secret store, so the other names come straight back. Merged `ac40de1`, deployed on jar `e11160695fbe`: `memberCredentials:` in config, deny-by-default, **34 known / 7 allowed / 29 blocked**, with a startup summary and a per-spawn gap warning. The operator applied the matching guarded block to the secret store on 2026-08-17, and it was then **verified in a live member pane: 29 of 29 blocked names hold the sentinel; the 5 allow-listed names present in that pane keep their real values.** The operator's own shell is unchanged. | +`memberHerdrSocket` lets the daemon connect members through a second herdr daemon. +When the key is blank or absent, members use the lead socket instead. See +`fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:148-158` and +`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:34-37`. -**Why the measurement had to happen at all.** The launcher writes the member's environment, and *then* -the pane's login shell re-sources the secret store and overwrites it. So the config half proves -nothing on its own — it has to be paired with a `BRIDGED_MEMBER`-guarded block that re-applies the -same sentinel, and only a reading taken inside a real member pane shows which one won. This is the -whole reason CB-592 was built the way it was; CB-596 inherits it. +The current repository does not switch this option on anywhere. Therefore this +page does not describe it as active on a host. An operator must set +`memberHerdrSocket` and restart the daemon for it to take effect. -**What the measurement actually said.** Two results, and they point in opposite directions. +## Planned -The reassuring one: **CB-592 works.** The member holds the sentinel -`blocked-by-fleetd-cb592-see-gitea-issue-77`, not the admin token — confirmed by hashing the -sentinel, which is a hardcoded non-secret string, and matching it against what the member reported. -The guarded-`export` mechanism does beat the login shell, which is the whole reason it was built that -way. +### External launcher plugins -The other one: **`GITLAB_PERSONAL_ACCESS_TOKEN` is a full personal access token for a second forge, -and nothing blocks it.** Alongside it sit Cloudflare, Tailscale, Grafana, Home Assistant, Telegram and -Confluence credentials — 24 live secrets with no bridge purpose. The rule that produced CB-592 was -*a member must not hold the operator's admin forge credential*. That argument never stopped at Gitea; -only the implementation did. +The old roadmap proposed dynamic external plugin loading for launcher adapters. +The current code confirms only two accepted kinds: `claude-code` and `opencode`. +See `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:1688-1731`. -This is the CB-604 → CB-606 shape for the third time in this milestone: **the instance that got fixed -was not the worst instance.** The operator chose deny-by-default — an allow-list of four names, with -every other known name replaced by the sentinel — because a deny-list has exactly the defect this -ticket is about: it is silently wrong the moment a new secret is added, and nothing reports it. +There is no plugin discovery and no trust model for third-party launchers in the +current source. This stays planned; it is not an extension point today. -**A note on the probe's own output.** It prints `NAME | STATE | LEN | SHA256-12`. That digest is a -safe way to compare two readings on one machine and an unsafe thing to publish: `GRAFANA_ADMIN_USER` -has length 5 and hashes to `8c6976e5b541`, which is SHA-256 of `admin`. A truncated hash of a -low-entropy value is not an anonymiser. The finding was written up with names and lengths only. +### Multi-host member federation -**What the fix found that no list contained.** The gap detector warns, on every spawn, about any -credential-shaped variable on neither `known:` nor `allow:`. On its very first spawn it named two: -`CLAUDE_CODE_MESSAGING_TOKEN` and `SSH_AUTH_SOCK`. Neither is in the secret store, so no list written -by reading that file could ever have held them. `SSH_AUTH_SOCK` is the serious one — a member needs -it to push, because worktree remotes are `ssh://`, but it lets a member sign with every key the -operator's agent holds. That is strictly broader than the repo-scoped forge token CB-302 built to -avoid exactly this. Allowed for now, with the reasoning written into the config; filed as **CB-607 -(#110)** on 2.0, fixed by pushing over HTTPS with `WORKER_GITEA_TOKEN`. +The roadmap once described cross-host member spawning, a federated roster, and +member routing between gateway daemons. None of those member-federation features are +in the current source. Do not infer them from the shipped `coordId` lead channel, +which carries lead-to-lead coordination only. -The general lesson is worth more than the instance: **the enumeration was of a file, and the exposure -is of an environment.** Anything granted by a handle rather than by a value is invisible to a list -built by reading `secrets.sh` — which is the argument for having the daemon report the gap instead of -trusting the list. +The limited feature that is built is peer-lead coordination through `coordinator:`. +See `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:597-641`. -**And the verification tool had the same defect.** The probe reported 26 blocked; the policy blocks -29. Both numbers were right and counting different sets — `scripts/probe-member-credentials.sh` -carries its **own** hardcoded list of 31 names and never reads the config, so it skipped the three -`N8N_*` names added to `known:` after it was written. Those three were then measured directly and are -blocked, so the result stands at 29 of 29. But note where the failure pointed: the script exited 0 and -its table looked complete. **A checker that under-reports fails in the worst direction, because a -clean run is taken as evidence.** Filed as **CB-608 (#111)**. +### Exact old stages, ticket counts, versions, and deployment claims -That makes three copies of one list — Java (CB-592), config (CB-596), and the probe — and only the -gap between the first two is guarded. Also recorded from the same run: -`CLAUDE_CODE_MESSAGING_TOKEN` is **unset in a member pane**, though it is present in the daemon's own -environment. Allow-listing it was harmless, but not for the reason it was allowed. +The earlier page gave dates, ticket ranges, test counts, herdr protocol versions, +release tags, and host deployment results. Those are historical records, and this +rewrite checked claims against the source tree rather than against source control +history or a running host. They are not repeated here as current facts. -### Deferred to 2.0 — the cut decision +For the same reason, this page does not claim that a launchd or systemd unit is +installed, that a broker is running, or that a feature was dogfooded. Those need +deployment evidence, not only source evidence. -Made 2026-08-16, by reading the admission rule strictly. +## Tried and dropped -| Ticket | Why it moved | -|---|---| -| CB-308 (#6) | Federation. It is what 2.0 *is*. | -| CB-589 (#74) | Cost-first placement. `weighted` spreads by ratio with no idea which profile costs money, so paid spawns happen while the free box sits idle — but that is working behaviour that costs too much, with a decisive workaround (`local.weight: 100`), not something broken on one host. A missing **explanation** of the workaround *was* broken, because `fleetd.yaml` is gitignored and a fresh host starts without it. **That half shipped in 1.1**; the policy did not. | -| CB-548 (#16) | Architect slots. A new capability, and genuinely blocked: `MemberRegistry.bind` is never called. | -| CB-605 (#103) | The systemd unit carries launchd's login-shell secret gap. Bites only when the first Linux gateway is stood up. | -| CB-607 (#110) | A member holds `SSH_AUTH_SOCK`, so it can sign with every key the operator's agent holds — broader than the repo-scoped token CB-302 built. Allowed **on purpose** today, because worktree remotes are `ssh://` and blocking it stops members pushing. A documented, deliberate scope reduction, not a break. | -| CB-608 (#111) | The credential probe keeps its own hardcoded name list and under-reported by three. The policy it checks is correct; what drifts is the checker. | +### AgentAPI -### Before the tag +AgentAPI was a research fallback in older plans. It was never built. The current +profile validation accepts only `claude-code` and `opencode`. See +`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:1688-1731`. -1. ~~Land or defer CB-606 and CB-586.~~ Both merged and deployed. -2. ~~Land CB-596.~~ Merged `ac40de1`, 870 tests, BUILD SUCCESS. -3. ~~Redeploy the daemon.~~ Done twice. The second, 2026-08-17, put CB-596 live: jar - `e11160695fbe`, and the startup line reads - `memberCredentials: 34 known name(s), 5 allowed — blocking 29 on every spawn`. A merge is not a - deployment, and this change alters what a member's environment contains. -4. ~~Confirm `fleet_whoami` still answers `primary`.~~ Done: `primary`. Note the trap — the redeploy - cuts the lead's own bridge MCP mount. It reconnected by itself the second time and did not the - first, so do not rely on either; if the tools are gone, ask the operator to run `/mcp`. -5. ~~Apply the secret-store half of CB-596.~~ Done by the operator on 2026-08-17: `secrets.sh` backed - up, then the `BRIDGED_MEMBER`-guarded block appended. `zsh -n` and `bash -n` both parse it, and the - operator's own shell is unchanged. -6. ~~Run the verification probe.~~ Done, inside a live member pane. **29 of 29 blocked names hold the - sentinel.** The probe itself covered 26 of them — see CB-608 — and the other three were measured - directly. -7. ~~Tag `v1.1.0`.~~ **Done, 2026-08-17.** The operator authorised it; the lead cut it. Annotated tag - on `7d41ccc`, pushed to `origin`, **217 commits since `v1.0.0`**, build re-run on the tagged - commit at 870 tests and 0 failures. Gitea milestone 16 closed, 0 open pull requests. No Gitea - *release* object was created — the tag is the artifact; a release is a separate publish. -8. ~~Write the operator guide.~~ **Done, 2026-08-17** — [13 User Guide](13-User-Guide). Chapters 4 - and 5 turned out never to have been written past their scope note, so they redirect to it now. +Do not present AgentAPI as an adapter, a fallback, or an operator choice. -**Not part of the tag, but do not lose it:** `GITEA_ACCESS_TOKEN` still needs rotating. A lead leaked -about 31 characters of it into a transcript while dumping the *structure* of the secret store — the -redaction covered `export NAME=…` lines, and that token sits on a line beginning with a guard, so it -fell through. Reported at once; the operator chose to rotate later. It was never a CB-596 criterion, -which is exactly how it could go missing with that ticket. +### Redis Streams, NATS JetStream, and an embedded durable queue -**On "Stage 5 ✅ production shape".** That row was never wrong — CB-504 did build a launchd agent and -a systemd unit. But a unit in `deploy/` that no host has loaded is not supervision. CB-594 and CB-600 -closed the gap between *built* and *safe to install*; the agent is still **not installed**, because -installing it is the operator's decision. Read the Stage 5 table as *built*, not as *running*. +Older plans considered Redis Streams, NATS JetStream, and an embedded queue. They +were not the shipped durable inbox. The current configuration and startup path use +an AMQP `broker` and `AmqpReplyInbox`; otherwise they use an in-memory inbox. See +`fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:49-50,656-714,1915` +and `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:399,811`. -## Delivery reliability & multi-host (CB-306 – CB-308) +Neither Redis Streams nor NATS JetStream appears anywhere in the current source. +They are discarded options, not supported configuration choices. -A cross-cutting track that came out of a **"communication break" review** of the reverse -(worker → primary) path. The bridge *pushes* to workers (the status-gated injector) but only -*pulls* to the primary — a worker reply resolves only an **already-open** blocking `fleet_send`; -MCP is client-initiated (bridge = server, primary = client), so the server can't call into the -primary. That asymmetry is the root of all three tickets. +### REST-first and SSE roadmap -```mermaid -flowchart LR - cb306["CB-306 ✅
spawn-readiness gate"] --> cb307["CB-307 ✅
reliable worker→primary delivery"] - cb307 --> s1["Stage 1 ✅
ReplyInbox port + in-memory adapter"] - cb307 --> s2["Stage 2 ✅
AMQP / LavinMQ durable adapter"] - cb307 --> s3["Stage 3 ✅
active push-to-primary + reminder"] - s2 --> cb308["CB-308 ⏳
multi-host federation"] - classDef done fill:#2f855a,stroke:#22543d,color:#ffffff; - classDef todo fill:#b7791f,stroke:#7b341e,color:#ffffff; - class cb306,s1,s2,s3 done - class cb308 todo -``` +The old page described every feature as a REST endpoint with MCP as a thin wrapper, +and it proposed Server-Sent Events (SSE). The daemon registers no SSE route, and MCP +and REST are two faces over one shared `MessageService` rather than one wrapping the +other. Treat that old shape as a design idea, not as the current system. -*Figure: CB-307's broker fabric is the foundation CB-308 stretches across hosts; CB-306 makes a -spawn fail fast instead of stalling.* +## Historical delivery sequence -- **CB-306 — spawn-readiness gate ✅ (main `7dd6c46`).** `ClaudeCodeLauncher.spawn` blocks until the - new worker's Claude has connected the bridge MCP (`AgentStatus.injectable()`) or a - `spawn_ready_timeout_ms` (default 20 000) elapses → on timeout it self-reaps the pane and throws - `PeerUnreachableException` (MCP tool error / HTTP 502). Opt-out with `timeout == 0`. Turns the old - pre-REPL folder-trust stall (surfacing as a ~60s send timeout) into a fast, explicit failure. +The project moved from a small single-host delegation plan toward the current fleet +daemon. The useful history is the decision path below. It does not claim that every +old stage, ticket, or test result is still current. -- **CB-307 — reliable worker → primary delivery.** Behind one `ReplyInbox` port so the adapter is - swappable (see [Implementation → `msg`](9-Implementation#msg--the-service-core-4-classes)). - - **Stage 1 ✅ (main `ba6b4a5`).** `ReplyInbox` port + `InMemoryReplyInbox` (soft-state, dedup by - `msgId`). A `fleet_reply` with no open send is now **held** instead of dropped; the primary - drains it by target via `fleet_poll(target)` / `GET /sessions/{id}/replies`. No broker, no new - dependency. **Only terminal replies are queued** — `fleet_ask` (interactive) and the - completion/failure fallbacks are deliberately *not* (would risk double-delivery). Live on the - running daemon. - - **Stage 2 ✅ (main `2bc5f3a`).** `AmqpReplyInbox` behind the *same* port for cross-restart - durability — **LavinMQ** default (single Crystal binary; `com.rabbitmq:amqp-client` works - unchanged; native DLX + delayed-message exchange), RabbitMQ interchangeable by URI. Mapping = - consume-and-hold with deferred manual ack: each target owns a durable queue `agent..inbox`; - a manual-ack consumer pulls persistent messages into an in-memory held map but does *not* ack until - the primary drains, so a `java -jar` bounce leaves them on the broker for redelivery. `broker:` - config absent → in-memory, present → AMQP. Contract test `@Tag("contract")` runs against a RabbitMQ - container. Dogfooded live (reply survived a daemon bounce, redelivered + acked exactly once). - **`fleetd` stays soft-state — the broker owns message durability, not the bus.** - - **Stage 3 — active push-to-primary + reminder ✅ (main `d4c9704`, gitea #5 closed).** The durable - inbox is a *landing zone* but delivery was still **pull** (the primary had to poll). Stage 3 makes - it **active**: a `ReplyPushLoop` nudges the primary the moment a reply lands with no open send, and - reminds (bounded) until drained. Push channel resolved the crux design question — the same-host - primary **is** a herdr pane (its `terminal_id` resolves on every MCP call), so the loop injects a - *drain nudge* (not the payload) into the primary's own pane via `AgentControl.send`, **status-gated** - (only when `injectable()`, never mid-turn) and **bounded** (`primary.push_reminders`=5, - `push_backoff_ms`=15000). **Ack = drain**: stop when `inbox.peek(target).isEmpty()`. A single-slot - `PrimaryRegistry` learns the primary from orchestration-side tools; an off-host / non-herdr primary - leaves it empty → the loop is a no-op and delivery degrades to pull (reply never lost). Optional - per-`msgId` `fleet_ack` tool for finer control than drain-all. Dogfooded live end-to-end. +1. The first plan focused on one delegated review flow. +2. Later work added member lifecycle, profiles, and worktree isolation. +3. The launcher boundary became provider-neutral enough for `claude-code` and + `opencode`. +4. Reply delivery gained an optional AMQP-backed inbox. +5. Peer leads gained a separate coordination channel. +6. A second herdr daemon became an optional member-routing setup. -- **CB-308 — multi-host federation ⏳ (design note, gitea #6; depends on CB-307 Stage 2).** A primary - on host A delegating to workers on hosts B, C… with no host learning another's terminals. Built on - CB-307's broker fabric: a **per-host gateway** (evolved `fleetd` owning its local herdr + - registry), **per-agent broker channels** `agent..inbox` (owning gateway = sole consumer), - a **federated roster** (soft-state presence on a `roster.*` topic = [CB-304](9-Implementation)'s - `rosterView`, federated), and a one-line **routing fork** (`local ? inject : publish`). Five - net-new deltas: global agent id, federated directory, gateway routing, cross-host spawn, and the - cross-host **trust model** (the gating concern — the broker link becomes the security boundary). - The MCP pull-asymmetry survives the network: the final hop into the primary is still a pull from - its *local* gateway. Full design: **`docs/CB-308-Multi-Host-Federation.md`**; topology sketch in - [Architecture → Topologies](1-Architecture#topologies). +The current code supports steps 2 through 6 as described in the status sections. +The original review demo details, old `ccs` commands, and the AgentAPI fallback do +not describe the current system. -## Stage 1 — detailed tickets +## What this rewrite did not promote -**Goal:** Opus, from its own subscription session, mounts `fleetd` and gets a code review -back from a real worker spawned under the `ccs ltms-local` profile (routing to the gx00 vLLM at -`http://gx00.gw:8000`) — same host, one hardcoded profile, reply via a pane/agent read (envelope -comes in Stage 2). This is the thinnest end-to-end vertical slice. +These are not current, shipped behaviour, so they stay planned or are left out: +external launcher plugins, multi-host member federation, a REST-first core, SSE, +exact release and test figures, live deployment results, old herdr versions, and old +`ccs` profile commands. -**Definition of done for the stage:** `CB-107` demo passes. +The build targets Java 25 (`fleetd/pom.xml:16`, `maven.compiler.release`), and the +source uses Java 25 features such as unnamed lambda parameters. An older version of +this page said Java 21. -> **Build status — Stage 1 COMPLETE** (in `fleetd/`, Maven · Java 25 · **307 unit/acceptance tests -> green** as of the CB-5xx close-out; the live-herdr and broker contract tests run separately via -> `mvn test -Pcontract`). The `CB-107` end-to-end demo gate passes and the -> bridge is **dogfooded**: an Opus primary delegates real tasks to off-subscription workers that -> reply through it (delegated code reviews have produced committed bug fixes). -> ✅ **CB-101** herdr client — connection-per-call, contract-tested vs live 0.7.0. -> ✅ **CB-102** worker spawn — native `agent.*` with env injection, guard-checked, live-verified. -> ✅ **CB-103** status-gated injector — per-target FIFO, one message per turn, TOCTOU closed. -> ✅ **CB-104** blocking `fleet_send` — `POST /sessions/{id}/message` + reply rendezvous. -> ✅ **CB-105** MCP adapter — `fleet_send`/`fleet_reply`/`fleet_status` over the REST core, with -> connection-based caller identity (loopback peer PID → herdr pane). -> ✅ **CB-106** config + logging — Jackson YAML + Logback. -> ✅ **CB-107** e2e review demo (stage gate) — runs green end-to-end; the worker is verifiably the -> off-subscription process, the primary's env has no `ANTHROPIC_BASE_URL`. -> ✅ **CB-108** worker placement — one tab per worker in a dedicated, shared **worker space** -> (`worker.placement/workspace/tabLabel`); unique per-worker names; single-pane-guarded, tolerant -> teardown. Reviewed (high-effort multi-agent) and live-verified. -> Two herdr wire facts pinned by contract test: **string ids**, and **one request per -> connection**. `agent.start` env is the subscription-safe injection point (no shell prefix). -> -> **Beyond Stage 1 — shipped delivery-reliability hardening.** The implementation continued the -> `CB-1xx` sequence past the skeleton for work that overlaps Stage 2–4 below. ⚠️ **These commit -> numbers (CB-106…CB-118 in git) reuse the CB-106/107/108 slots this plan assigned above to -> config/e2e-demo/placement — identify the items below by name, not number.** Shipped: -> completion fallback (a confirmed `working→idle` turn resolves a send), async fire-and-poll -> (`wait:false` + `fleet_poll`, beats the caller's MCP call timeout), fleet MCP tools -> (`fleet_spawn`/`fleet_list`/`fleet_stop`/`fleet_profiles`/`fleet_poll`), failure detection -> for wedged (`unknown`), vanished, and never-ready workers, multi-profile workers with -> per-profile base_url guards, worker cwd inheritance (never `$HOME`), and a readiness gate that -> holds delivery until the worker's Claude has connected the bridge MCP. -> -> **Reliability follow-ups (CB-115…CB-118).** Turn-completion made robust: a clean transcript -> scrape that strips TUI chrome to the last assistant block (CB-115), waiter-identity so a late -> completion fallback can never cross into the next turn's send (CB-116), startup reconciliation -> that reaps orphaned worker panes leaked by a prior daemon process (CB-117), and a -> completion-baseline clip fix so the CB-115 misattribution guard holds for assistant blocks over -> the scrape cap (CB-118). Verified by a single-worker conversation harness, a **1-primary / -> N-worker fan-out issue-hunt**, and a **sustained 5-minute stateful back-and-forth** (30 turns, -> every one a clean `fleet_reply`, running total held) — all under `e2e/`. -> -> **Stage 2 — rich message semantics COMPLETE.** The contract-and-guard stage has landed (the -> +13 tests above are its coverage): -> ✅ **CB-205** `fleet_ask` reverse rendezvous — a worker pauses its delegated turn to ask the -> primary and resumes the **same** turn with the answer. The question surfaces on the primary's own -> blocked `fleet_send` carrying a `turn_id`, and the primary answers by sending on that `turn_id`. -> Live e2e (`e2e/fleet_ask_test.py`): the worker asked in ~9s and, after the primary answered, -> resumed and replied via a clean `fleet_reply`. -> ✅ **CB-202** reviewer-role skill (`.claude/skills/reviewer/SKILL.md`) — the playbook a worker -> loads to review a scoped assignment, ask the lead via `fleet_ask` when the call is genuinely -> theirs, and report exactly one structured finding via `fleet_reply`. Pairs the already-shipped -> `fleet_reply`/`fleet_ask` tools with the role guidance for using them. -> ⚠️ **CB-201 descoped** — the planned `{from,to,corr}` envelope schema + codec was **not** built. -> Connection identity (loopback peer PID → herdr pane) already routes every reply and question to -> the right session robustly, so CB-201 shipped as a *lightweight* addition only where the reverse -> path needs it: a `QUESTION` message kind + `turn_id` correlation. No heavyweight envelope. -> ✅ **CB-203 / CB-204** (reply rendezvous · subscription guard) shipped earlier in the CB-1xx -> sequence — completion-fallback resolution on the `working→idle` edge, and the per-profile -> `base_url` allowlist checked before any herdr call. **All five Stage 2 tickets are done.** +## Related pages ---- - -### CB-101 — herdr socket client -**Scope.** Connect to `~/.config/herdr/herdr.sock` (`HERDR_SOCKET_PATH`) via -`UnixDomainSocketAddress`; implement `call(method, params) → JsonNode` (NDJSON, `id`-correlated) -and `subscribe(subs) → BlockingQueue` (fed by an event-loop virtual thread) for -`pane.agent_status_changed`. Probe `ping` on connect (assert `protocol: 14`); enumerate via -`workspace.list`/`pane.list`; fail fast on schema mismatch. **Build against the -[verified 0.7.0 API](2-Message-Server#the-herdr-control-contract-what-fleetd-drives), not the docs.** -**Acceptance.** JUnit test against a **mock UDS socket** replays a `workspace.create` round-trip -and one event; a **contract test vs the real herdr** asserts `ping.protocol == 14` and the -`workspace.list`/`pane.list` shapes. -**Deps.** none. - -### CB-102 — spawn a worker (native `agent.*`) ✅ -**Decided (spike done).** herdr's `agent.start {name, argv, env}` carries the worker command and -a **first-class `env` map that reaches the process** (proven via `agent.read`), so `fleetd` -injects `ANTHROPIC_BASE_URL` there — guard-checked before the call — with no shell prefix and its -own env untouched. Chosen over the pane + `send_text` fallback. herdr tracks each worker's Claude -**session UUID** (`agent.list`/`agent.get`), which grounds the ID contract. Teardown is -`pane.close` (no `agent.stop`). -**Built.** `AgentControl` (start/send/read/get/status/list/close) + `ClaudeCodeLauncher` (builds env -from config, `assertWorker` **before** any herdr call, implements `PeerLauncher`) + REST `POST /workers`, `GET /agents`, -`DELETE /workers/{paneId}`. -**Acceptance (met).** Unit: `POST /workers` with an off-allowlist base_url → **403 and herdr is -never touched**; a good base_url → `agent.start` env carries `ANTHROPIC_BASE_URL`. Contract: a -live probe spawn proves env propagation and cleans up its pane. (Real `ccs ltms-local claude` -worker spawn is exercised by the CB-107 demo.) -**Deps.** CB-101. - -### CB-103 — status-gated injector -**Scope.** Per-pane FIFO; `deliver(pane, text)` sends `send_text` + `send_keys "enter"` **only** -when `agent_status ∈ {idle, blocked}`, else queues until the next idle event. Single writer per -pane; status-check + send serialized on the event-loop virtual thread (close the TOCTOU window). -**Acceptance.** A delivery issued while the pane is `working` lands only after the `idle` event; -two rapid deliveries never interleave (golden transcript). -**Deps.** CB-101. - -### CB-104 — blocking `fleet_send` as a **REST endpoint** + reply capture -**Scope.** Implement the feature as **`POST /sessions/{id}/message`** (`{content}`, blocking) → -resolve to the (single, hardcoded) worker pane → inject → wait for the turn-done `working→idle` edge → return -`pane.read {source:"recent-unwrapped"}` of the last assistant block in the response body. (No -envelope yet; scrape is acceptable for Stage 1.) This REST route is the feature; MCP wraps it in -CB-105. -**Acceptance.** A **REST contract test** (`POST /sessions/{id}/message`, **no Claude in the -loop**) returns a non-empty reply from the pane; a `working→idle` transition unblocks it; a -timeout returns a typed "still working" (HTTP 202-style) response. -**Deps.** CB-102, CB-103. - -### CB-105 — MCP adapter over the REST core (SERVER face) -**Scope.** Streamable-HTTP MCP server exposing `fleet_send`/`fleet_status` as **thin adapters -over the CB-104 REST routes** (`POST /sessions/{id}/message`, `GET /sessions/{id}/status`); bind -`127.0.0.1:8080`. One-line mount: `claude mcp add --transport http bridge -http://127.0.0.1:8080/mcp`. -**Acceptance.** **Parity test** — `fleet_send` via MCP and `POST …/message` via REST produce -identical results for the same input; `claude mcp list` shows `bridge` connected; `fleet_status` -returns the live `agent_status`. -**Deps.** CB-104. - -### CB-106 — config + wiring -**Scope.** `fleetd.yaml` load (**Jackson YAML**): one worker profile (`gx00-vllm` → ccs profile, -model, `base_url` host), bind address, workspace root. **SLF4J/Logback** startup line logging the -resolved worker command (secrets redacted). Run as a foreground process (systemd deferred to -Stage 5). -**Acceptance.** Bad/missing profile fails fast with a clear message; a valid config boots and -logs the resolved (redacted) spawn command. -**Deps.** none (parallel with CB-101). - -### CB-107 — end-to-end review demo (stage gate) -**Scope.** Scripted demo: start herdr → start `fleetd` → mount MCP on a primary `claude` → -from the primary, `fleet_send` a real diff with `kind:"review.request"` (body carried as text -for now) → assert a review comes back as the tool result. Document the exact steps in -[Setup](4-Setup). -**Acceptance.** The demo runs green on one host end-to-end; the worker is verifiably the -`ccs gx00-vllm` process (not the subscription); the primary's env has no `ANTHROPIC_BASE_URL`. -**Deps.** CB-105, CB-106. - -```mermaid -flowchart LR - CB101["CB-101
herdr client"] --> CB102["CB-102
ccs spawn"] - CB102 --> CB103["CB-103
injector"] - CB103 --> CB104["CB-104
fleet_send"] - CB104 --> CB105["CB-105
MCP server"] - CB106["CB-106
config"] --> CB107 - CB105 --> CB107["CB-107
e2e demo (gate)"] - classDef gate fill:#2f855a,stroke:#22543d,color:#ffffff; - class CB107 gate -``` - -*Figure: Stage 1 dependency order. CB-106 runs in parallel; everything converges on the CB-107 -end-to-end gate.* - -## Open decisions (surface before/while building) - -- **ccs worker profiles** — which existing ccs profiles (or new `ccs api` profiles) back the - `gx00-vllm` / local workers, and their exact `base_url` hosts for the allowlist. -- **Reviewer system prompt** — ✅ **Decided (CB-202):** shipped as a loadable **skill** - (`.claude/skills/reviewer/SKILL.md`), not a `CLAUDE.md` snippet. Workers inherit the repo cwd, so - the skill is available to every reviewer worker without polluting the developer-facing project - `CLAUDE.md`. -- **Envelope-as-text (Stage 1) → structured (Stage 2)** — confirm the Stage-1 shortcut (request - body inlined as prompt text, reply scraped) is acceptable before the envelope lands. - -## Related - -- **[Use Cases](7-Use-Cases)** — the review scenario and the five mechanisms these stages build. -- **[Message Server](2-Message-Server)** — component design the tickets implement. -- **[Setup](4-Setup)** · **[Operations](5-Operations)** — bring-up and day-2, fleshed out as stages land. +- [Features](11-Features) lists operator-facing capabilities. +- [User Guide](13-User-Guide) gives the maintained operator procedure. +- [Cross-Host Messaging](10-Cross-Host-Messaging) gives the wider design history. +- [Implementation](9-Implementation) maps the current source structure. diff --git a/9-Implementation.md b/9-Implementation.md index 00b6711..8813a4b 100644 --- a/9-Implementation.md +++ b/9-Implementation.md @@ -1,152 +1,184 @@ # 9. Implementation Architecture (as-built) -> **Scope.** This is the *as-built* code map of the `fleetd` module — the actual packages, -> classes, flows, and state machines in the source tree, as a companion to the design-level -> [1. Architecture](1-Architecture) and [2. Message Server](2-Message-Server). Every enum, -> constant, and route below was verified against source at main `3aa69a9`; the `msg`-layer -> **reply-inbox** (CB-307 Stage 1 `ba6b4a5`), its **AMQP durable adapter** (Stage 2 `2bc5f3a`), and -> the **active push-to-primary loop** (Stage 3 `d4c9704`) are folded in below. +> **Scope.** This is the *as-built* code map of the `fleetd` module: the real packages, classes, +> flows, and state machines in the source tree. It is a companion to the design-level +> [1. Architecture](1-Architecture) and [2. Message Server](2-Message-Server) pages. Every class +> name, enum value, and route below was checked against the source at `main` `a49e968` +> (2026-08-31). Where this page and the code ever disagree again, trust the code, not this page. -`fleetd` is a single-host Java 25 / Maven daemon: the sole gateway between an on-subscription -**primary** (Opus) and off-subscription **workers**, speaking to the **herdr** PTY manager -(protocol 19, herdr 0.8.0 — ported in CB-521) over a Unix-domain socket. It presents two equivalent faces — a REST -server (the testability seam) and an MCP server — over one shared service core. +`fleetd` is a single-host Java 25 / Maven daemon. It is the one gateway between an on-subscription +**primary** (a lead, running Opus) and off-subscription **members** (workers and architects), +speaking to the **herdr** PTY manager over a Unix-domain socket. It presents two equivalent faces — +a REST server (`rest.FleetApp`) and an MCP server (`mcp.FleetMcp`) — over one shared service core +(`msg`). The Java package root is `dev.ltms.fleet`. ## Component map -The daemon is layered. The two north faces (REST + MCP) are thin adapters over one service core -(`msg`); the core drives delivery through `inject`, which speaks to the outside world only through -`herdr`. `guard` sits on the spawn path; `config` and the `Fleetd` entry point wire it all. - ```mermaid flowchart TB - primary["Primary (Opus)
on-subscription"] - worker["Worker (Claude Code)
off-subscription"] + lead["Lead (primary)
on-subscription"] + peerlead["Peer lead
another daemon"] + worker["Worker / architect
off-subscription"] subgraph fleetd["fleetd daemon"] direction TB subgraph faces["North faces (thin adapters)"] - rest["rest.FleetdApp
REST · Javalin"] - mcp["mcp.BridgeMcp
MCP · /mcp servlet"] + rest["rest.FleetApp
REST · Javalin"] + mcp["mcp.FleetMcp
MCP · /mcp servlet"] end - msg["msg.MessageService + Rendezvous + ReplyInbox
service core · rendezvous + held-reply inbox"] - inject["inject.Injector + StatusPoller
single-writer, status-gated delivery"] + authz["auth.CallerResolver + auth.Authz
identity + the one decision table"] + msg["msg.MessageService + Rendezvous
+ ReplyInbox + ReplyPushLoop
service core"] + inject["inject.Injector + StatusPoller
+ CompletionResolver
single-writer, status-gated delivery"] + health["health.FleetHealthMonitor
slow whole-fleet evidence loop"] + session["session.SessionManager
+ GitWorktrees
member registry + worktrees"] + peerspi["peer.PeerLauncher SPI
member.CompositePeerLauncher"] + adapters["member.ClaudeCodeLauncher
member.OpenCodeLauncher"] guard["guard.SubscriptionGuard
boundary enforcement"] - ccl["member.ClaudeCodeLauncher
first Claude adapter · implements PeerLauncher"] - herdr["herdr.AgentControl / WorkspaceControl
JSON-RPC over UNIX socket"] + herdr["herdr.AgentControl / WorkspaceControl
via herdr.HerdrRouter"] + leadmsg["msg.LeadChannel / LeadMailbox
lead ↔ lead over a shared broker"] end herdrd["herdr daemon
(PTY manager)"] - primary -->|"fleet_send / answer"| rest - primary -->|"MCP tool calls"| mcp + lead -->|"fleet_send / fleet_spawn / answer"| rest + lead -->|"MCP tool calls"| mcp worker -->|"fleet_reply / fleet_ask"| mcp - rest --> msg - mcp --> msg - mcp --> ccl - ccl --> guard - ccl --> herdr + rest --> authz + mcp --> authz + authz --> msg + mcp --> session + session --> peerspi + peerspi --> adapters + adapters --> guard + adapters --> herdr msg --> inject + health --> session + health --> msg inject --> herdr herdr --> herdrd herdrd -.->|"drives PTYs"| worker + mcp --> leadmsg + leadmsg <-.->|"shared broker, cross-host"| peerlead classDef core fill:#2b6cb0,stroke:#1a365d,color:#ffffff; classDef gate fill:#b7791f,stroke:#7b341e,color:#ffffff; class msg,inject core - class guard gate + class guard,authz gate ``` -*Figure 1 — component layers. The service core (`msg`) is reached identically from either face; -delivery reaches workers only via `inject → herdr`. Peer materialization is delegated to a -`PeerLauncher` adapter (`ClaudeCodeLauncher` today); the subscription guard (amber) lives inside -the adapter and gates spawn.* +*Figure 1 — component layers. Both north faces resolve identity through `auth` before reaching the +service core (`msg`); delivery to a member reaches it only via `inject → herdr`. Spawning a member is +delegated through the `peer.PeerLauncher` SPI, today implemented by `member.CompositePeerLauncher` +routing to `member.ClaudeCodeLauncher` or `member.OpenCodeLauncher`; the subscription guard (amber) +lives inside those adapters and gates spawn. `health` watches the same state the other layers +produce but never drives delivery itself. `msg.LeadChannel` is the separate cross-host path a lead +uses to reach a peer lead directly (see [10. Cross-Host Messaging](10-Cross-Host-Messaging)).* ## Package & class reference -Fourteen packages under `dev.ltms.fleetd`: `auth`, `config`, `guard`, `herdr`, `inject`, `lead`, -`mcp`, `member`, `metrics`, `msg`, `peer`, `placement`, `rest`, `session`. Below, per layer: the -classes, their kind, and their role. Method signatures are abbreviated; see source for full -contracts. The sections below do not yet cover every package — `auth`, `herdr`, `inject`, `msg`, -`peer`, `mcp` and the edge are written up; `lead`, `metrics`, `placement` and `session` are not. +Sixteen packages under `dev.ltms.fleet`: `auth`, `config`, `guard`, `health`, `herdr`, `inject`, +`lead`, `logging`, `mcp`, `member`, `metrics`, `msg`, `peer`, `placement`, `rest`, `session`. Below, +per package: what it owns and its few most important classes, each with a `file:line` you can check +yourself. This is not every class in every package — it is the ones a maintainer needs to find first. -### `herdr` — the wire layer (11 classes) +### `herdr` — the wire layer -The only speaker of the herdr protocol. Encodes newline-delimited JSON-RPC (string request ids, -**one request per connection**) and projects herdr's workspace/tab/pane/agent nodes into typed -Java records. +The only speaker of the herdr protocol. Encodes newline-delimited JSON-RPC and projects herdr's +workspace/tab/pane/agent nodes into typed Java records. -| Class | Kind | Role | -|---|---|---| -| `HerdrClient` | interface | The client face — the only herdr speaker in `fleetd` (`call(method, params)`). | -| `UnixSocketHerdrClient` | class | JDK Unix-socket transport; fresh connection per `call`, one frame, one line, close. | -| `HerdrCodec` | class | Encodes/decodes JSON-RPC frames (`encode`, `decodeResult`). | -| `HerdrException` | class | Runtime failure carrying an optional herdr protocol `code()`. | -| `AgentControl` | class | Domain wrapper over `agent.*`: `start`, `send`, `read`, `status`, `close`. | -| `Agent` | record | Projects an `agent` node into a worker handle (status/session metadata). | -| `AgentStatus` | enum | Worker lifecycle state (see [state machine](#state-machine-agentstatus)). | -| `WorkspaceControl` | class | Placement over `workspace.*`/`tab.*`: `ensureWorkspace` (synchronized), `createTab`, `renameTab`, `closeTab`. | -| `Workspace` | record | Projects a herdr workspace node. | -| `Tab` | record | Projects a tab plus the placeholder pane seeded at creation. | -| `PaneLocator` | class | Resolves a PID to a herdr `terminal_id` (`terminalForPid`). | +| Class | Role | +|---|---| +| `HerdrClient` (`herdr/HerdrClient.java:1`) | The client face — one `call(method, params)`. | +| `UnixSocketHerdrClient` (`herdr/UnixSocketHerdrClient.java:1`) | JDK Unix-socket transport; fresh connection per call. | +| `AgentControl` (`herdr/AgentControl.java:25`) | Domain wrapper over `agent.*`: start/send/read/status/close a member's agent. | +| `WorkspaceControl` (`herdr/WorkspaceControl.java:1`) | Placement over `workspace.*`/`tab.*`. | +| `AgentStatus` (`herdr/AgentStatus.java:8`) | A member's lifecycle state — see the [state machine](#state-machine-agentstatus) below. | +| `HerdrRouter` (`herdr/HerdrRouter.java:7`) | Routes a call to the lead herdr daemon or the separate member herdr daemon (CB-185's `memberHerdrSocket`), by target id. | +| `LeadTabScanner` (`herdr/LeadTabScanner.java:16`) | Discovers which panes host a lead by scanning herdr for operator-labelled tabs, feeding `auth.CallerResolver`. | +| `PaneLocator` (`herdr/PaneLocator.java:1`) | Resolves a PID to a herdr `terminal_id`. | -**Invariant:** a prompt is submitted only by a standalone `"\r"` `agent.send` after the message -text (`AgentControl.SUBMIT_KEY`) — never a shell prefix. `closeTab`/`PaneLocator` swallow only -"already gone" errors; other failures propagate. +`AgentStatus.injectable()` is only ever `IDLE`, `BLOCKED`, or `DONE` — never `WORKING`. -### `inject` — status-gated delivery + turn lifecycle (6 classes) +### `inject` — status-gated delivery and turn lifecycle -The single-writer delivery layer. Polls worker status, queues per target, injects only when safe, -and **synthesizes turn boundaries** so a blocked `fleet_send` resolves even when a worker never -calls `fleet_reply`. This is where the turn state machine lives. +The single-writer delivery layer. Polls a member's status, queues per target, injects only when +safe, and synthesizes turn boundaries so a blocked `fleet_send` resolves even when a member never +calls `fleet_reply`. -| Class | Kind | Role | -|---|---|---| -| `StatusPoller` | class | Virtual-thread loop sampling active workers; feeds refined status to `Injector` (`start`/`stop`, idempotent). | -| `StatusRefiner` | class | Reclassifies `UNKNOWN` by reading pane content — pane text is authoritative (`refine`, `classify`). | -| `Injector` | class | Queues + delivers at most one message per turn, serialized per target (`enqueue`, `onStatus`, `drop`, `activeTargets`). | -| `CompletionResolver` | class | `TurnListener` impl; scrapes the pane to resolve `Rendezvous` on completion/failure. | -| `TurnListener` | interface | Turn-lifecycle callbacks: `onDelivered`, `onTurnComplete`, `onTurnFailed`. | -| `WorkerPresence` | class | Tracks which workers have connected the bridge MCP (`markPresent`, `isPresent`, `forget`). | +| Class | Role | +|---|---| +| `Injector` (`inject/Injector.java:48`) | Queues and delivers at most one message per turn, serialized per target. | +| `StatusPoller` (`inject/StatusPoller.java:21`) | Virtual-thread loop sampling active members; feeds `Injector.onStatus`. | +| `StatusRefiner` (`inject/StatusRefiner.java:26`) | Reclassifies `UNKNOWN` by reading the pane's real content. | +| `CompletionResolver` (`inject/CompletionResolver.java:1`) | `TurnListener` that scrapes the pane to resolve `Rendezvous` on turn completion, failure, or usage-limit exhaustion. | +| `TurnListener` (`inject/TurnListener.java:13`) | Turn-lifecycle callbacks: `onDelivered`, `onTurnComplete`, `onTurnFailed`. | +| `MemberPresence` (`inject/MemberPresence.java:18`) | Tracks which members have connected the bridge MCP — the real readiness signal, not herdr's `idle`. Imported by `FleetMcp` (`mcp/FleetMcp.java:12`, used at `465-469`). This is the class the wiki used to call `WorkerPresence`; that name is gone. | +| `ExhaustedPatternLookup`, `ExhaustionSink` | Seams for the CB-578 usage-limit classification (see `Outcome.BACKEND_EXHAUSTED` below). | -### `msg` — the service core (6 classes) +### `msg` — the service core -Owns the forward rendezvous (`fleet_send` → `fleet_reply`) and the reverse rendezvous -(`fleet_ask` → answer), plus async fire-and-poll — and, since CB-307, the **reply inbox** that -holds a worker's terminal reply when *no* forward send is open (instead of dropping it) and the -**push loop** that actively nudges the primary to drain it. +Owns the forward rendezvous (`fleet_send` → `fleet_reply`), the reverse rendezvous (`fleet_ask` → +answer), async fire-and-poll, the reply inbox that holds a member's reply when no send is open, the +push loop that nudges the primary to drain it, and — new since the wiki last described this page — +the lead-to-lead channel. -| Class | Kind | Role | -|---|---|---| -| `MessageService` | class | Orchestrates send/reply/ask/answer + async dispatch (`send`, `answer`, `ask`, `sendAsync`, `poll`), and routes `fleet_reply` through **`reply`** (resolve an open send, else publish to the inbox) + **`drainReplies`** (peek-then-ack a target's held replies). | -| `Rendezvous` | class | Low-level registry of forward waiters + reverse-ask futures (`open`, `resolve`, `resolveQuestion`, `openAsk`, `answerAsk`, `closeAsk`, `resolveCompletion`, `resolveFailure`). **Untouched by CB-307** — it stays a pure synchronization primitive; a `false` from `resolve` (no live waiter) is what triggers the inbox publish, one layer up in `MessageService`. | -| `ReplyInbox` | interface | The port (CB-307): `publish(target, msgId, content)` (idempotent, dedup by `msgId`), `peek(target)` (non-destructive FIFO snapshot), `ack(target, msgId)`. Nested `InboxMessage(msgId, target, content)` record. Two adapters implement the *same* port — soft-state in-memory and durable AMQP. | -| `InMemoryReplyInbox` | class | Stage-1 default adapter — per-target FIFO in a `ConcurrentHashMap>`, thread-safe, dedup by `msgId`. **Soft-state, not persistence** — undrained replies are lost on a `java -jar` bounce, consistent with "fleetd stays soft-state; the broker owns durability." Selected when `broker:` config is absent. | -| `AmqpReplyInbox` | class | Stage-2 durable adapter (CB-307 `2bc5f3a`) — one durable queue `agent..inbox` per target; **consume-and-hold with deferred manual ack** (a manual-ack consumer pulls persistent messages into an in-memory held map but doesn't ack until the primary drains, so a bounce leaves them on the broker for redelivery). Dedup keys the held map by `msgId`; automatic connection + topology recovery re-declares queues and clears stale delivery-tags. LavinMQ default, RabbitMQ by URI swap. Selected when `broker:` config is present. | -| `ReplyPushLoop` | class | Stage-3 active push (CB-307 `d4c9704`) — a dedicated status-gated scheduled loop (mechanism (b), **not** the worker `Injector`, so it stays decoupled from `WorkerPresence`). `onReplyQueued(target)` (called on the no-waiter branch of `reply`) starts a bounded reminder loop: at each tick `decide(target, count)` returns `INJECT` / `WAIT_BUSY` / `STOP`, injecting a *drain nudge* into the primary's own pane via `AgentControl.send` only when `status(primary).injectable()`. Stops when `peek(target)` is empty (ack = drain) or the reminder cap is reached; idempotent per target. | +| Class | Role | +|---|---| +| `MessageService` (`msg/MessageService.java:46`) | Orchestrates send/reply/ask/answer + async dispatch. | +| `Rendezvous` (`msg/Rendezvous.java:26`) | Low-level registry of forward waiters and reverse-ask futures. | +| `ReplyInbox` (`msg/ReplyInbox.java:19`) | The port: `publish`, `peek`, `ack`, plus `own`/`release` for federation. | +| `InMemoryReplyInbox` / `AmqpReplyInbox` | The two adapters — soft-state, and cross-restart durable over AMQP. | +| `ReplyPushLoop` (`msg/ReplyPushLoop.java:50`) | Status-gated loop that nudges a lead's own pane about queued replies, finished tickets, and open questions. | +| `LeadChannel` (`msg/LeadChannel.java:22`) | This daemon's lead-to-lead mailbox port: `publish`, `peek`, `ack`, `selfCoordId`. | +| `LeadMailbox` | The production `LeadChannel` implementation, over a shared AMQP broker. | +| `LeadCoordLoop` (`msg/LeadCoordLoop.java:15`) | Takes what arrived in this daemon's own mailbox and types it into the local lead's pane. | +| `LeadHeartbeatLoop` (`msg/LeadHeartbeatLoop.java:20`) | Opt-in nudge for a lead that has sat idle too long with nothing driving it (CB-551). | -**Enums (verbatim).** `Rendezvous.Kind`: `REPLY`, `COMPLETION`, `FAILED`, `QUESTION`. -`MessageService.Outcome`: `REPLIED`, `COMPLETED_UNREPLIED`, `WORKER_FAILED`, `QUESTION`, -`TIMED_OUT_WORKING`, `TIMED_OUT_QUEUED`, `BUSY`, `STALE_TURN`. `AskOutcome`: `ANSWERED`, -`NO_WAITER`, `TIMED_OUT`. `Phase`: `PENDING`, `DONE`, `FAILED`. +**`MessageService.Outcome`** (`msg/MessageService.java:65`): `REPLIED`, `COMPLETED_UNREPLIED`, +`WORKER_FAILED`, `BACKEND_EXHAUSTED`, `QUESTION`, `TIMED_OUT_WORKING`, `TIMED_OUT_QUEUED`, `BUSY`, +`STALE_TURN`. This is nine values, not the six the wiki used to list — `BACKEND_EXHAUSTED` (a +backend refused on a usage limit, CB-578) and `QUESTION`/`STALE_TURN` (the `fleet_ask` reverse +rendezvous) were missing before. -### `auth` — who is calling, and what may they do (7 classes) +**`MessageService.Phase`** (an async ticket's lifecycle, `msg/MessageService.java:141`): +`PENDING`, `ASKING`, `DONE`, `FAILED`. `ASKING` is new: it means the member paused mid-turn in +`fleet_ask` and `fleet_poll` shows the open question instead of a stale "still pending". -The role axis. Every request arrives on a connection, and this package turns that connection into -a `Principal` and then decides. Identity is **never** taken from a tool argument, so a caller -cannot claim to be someone else. +**`Rendezvous.Kind`** (`msg/Rendezvous.java:29`): `REPLY`, `COMPLETION`, `FAILED`, +`BACKEND_EXHAUSTED`, `QUESTION`. -| Class | Kind | Role | -|---|---|---| -| `Role` | enum | `PRIMARY` (a lead orchestrator), `WORKER` (a spawned member with its own pane), `ARCHITECT` (a spawned member bound to a configured slot). Anything unrecognised is anonymous, not a role. | -| `Principal` | record | `(role, terminal, pid, name)` plus the factories `anonymous`, `primary`, `leader`, `worker`, `architect`. Predicates: `isPrimary`, `isWorker`, `isArchitect`, `isSpawnedMember` (worker **or** architect — CB-560), `isAnonymous`, `ownsSession`. | -| `CallerResolver` | class | Connection → `Principal`. Since CB-561 the **only** public way to build one is `withLeadsAndMembers(identity, tokenMode, token, leadTerminals, memberRegistry)`; the older map- and supplier-form constructors were deleted because they produced a resolver that could never return `ARCHITECT`, which failed silently. The remaining constructors are package-private and exist for tests. | -| `Authz` | class | The one decision table, `permits(caller, action, targetSession)`. | -| `AuditLog` | class | `allowed` / `denied` / `failed` — one line per decision, so a refusal is visible rather than a mystery. | -| `MemberLifecycle` | interface | `acquired(role, profile, terminal)` / `released(terminal)`. The seam the session lifecycle calls; a no-op default keeps tests free of the registry. | -| `MemberRegistry` | class | Flattens every `fleet:` pool into slots keyed `architect:opus`, and owns the live `terminal → slot` bindings. `acquired` binds **only** when `role == ARCHITECT`, to the first free slot with a matching profile; `bind`/`unbind` are compare-safe, so a stale unbind cannot remove a replacement. | +**`MessageService.AskOutcome`**: `ANSWERED`, `NO_WAITER`, `TIMED_OUT`. -**The decision table** (`Authz.Action` → who): +**Ticket retention (issue #197, fixed 2026-08-31).** `MessageService.pruneTerminalTickets` +(`msg/MessageService.java`) drops a finished ticket once `TICKET_TTL_NANOS` +(`msg/MessageService.java:62`, 10 minutes) has passed **since it completed**. `Task` stamps +`completedNanos` from a `whenComplete` hook registered in its constructor, so every completion path +stamps it — a reply, the completion fallback, a timeout, a failure, or an abandon on teardown. + +Until 2026-08-31 the age was measured from `createdNanos` instead, so the window to collect was ten +minutes **minus however long the task ran**. Any delegation lasting longer than the TTL had its reply +destroyed on the first sweep after it finished, and the reply lives only in `Task.future`, so nothing +could recover it. Worth knowing when reading older behaviour reports, and when running a daemon that +has not been redeployed since the fix. + +### `auth` — who is calling, and what they may do + +Every request arrives on a connection, and this package turns that connection into a `Principal` +and then decides. Identity is never taken from a tool argument. + +| Class | Role | +|---|---| +| `Role` (`auth/Role.java:12`) | `PRIMARY`, `WORKER`, `ARCHITECT`, `ANONYMOUS`. | +| `Principal` (`auth/Principal.java:16`) | `(role, terminal, pid, name)` plus factories `anonymous`, `primary`, `leader`, `worker`, `architect`. | +| `CallerResolver` (`auth/CallerResolver.java:42`) | Connection → `Principal`. Resolution order: a pane named by `leaders:`/the tab scanner ⇒ `PRIMARY`; a pane bound to an architect slot ⇒ `ARCHITECT`; any other on-host pane ⇒ `WORKER`; otherwise a bearer token (token mode) or loopback trust ⇒ `PRIMARY`; otherwise `ANONYMOUS`. | +| `Authz` (`auth/Authz.java:12`) | The one decision table, `permits(caller, action, targetSession)`. | +| `MemberRegistry` (`auth/MemberRegistry.java:35`) | The architect-slot registry: configured slots plus the live `terminal → slot` bindings an architect spawn creates. | +| `AuditLog` (`auth/AuditLog.java:22`) | One JSON line per allow/deny decision — never the message content. | + +**`Authz.Action`** (`auth/Authz.java:18`): `SPAWN`, `STOP`, `SEND`, `REPLY`, `ASK`, `DRAIN`, `READ`, +`METRICS` — eight actions, not the five the wiki used to group them into. + +**The decision table** (`auth/Authz.java:44`): | Action | Permitted to | |---|---| @@ -155,78 +187,122 @@ cannot claim to be someone else. | `REPLY`, `ASK` | the caller that owns the target session — only ever itself | | `READ`, `METRICS` | primary, worker, or architect | -**Two consequences worth stating.** First, a member's role reaches the principal only through -`MemberRegistry`, and only architects bind; a developer and a reviewer are both `Role.WORKER` at -this layer, and the difference between them lives in the roster (`fleet_list`), not in the -principal. Second, `SEND` is the one action an architect gains over a worker, which is what lets -two architects talk to each other without the lead relaying every message. +An architect gets `SEND` but never `SPAWN`/`STOP`/`DRAIN`: it can delegate turns to workers, but +fleet lifecycle stays the primary's alone. -### `peer` — the launcher SPI (4 classes) +### `peer` — the launcher SPI -The seam that keeps the core peer-neutral. The bus delegates spawn/teardown to a launcher -implementation while the core owns transport, session lifecycle, and routing. The first adapter -is `member.ClaudeCodeLauncher`; future adapters (e.g. Codex) implement the same SPI. +The seam that keeps the core peer-neutral. The core delegates spawn/teardown to a `PeerLauncher` +implementation and knows nothing about how a peer's process is actually built. -| Class | Kind | Role | -|---|---|---| -| `PeerLauncher` | interface | SPI for materializing a connected peer: `spawn`, `stop`, `effectiveCwd`, `parityOverlay`, `profiles`, `defaultProfile`, `list`, `reapOrphanWorkers`, `capabilities`. | -| `PeerHandle` | interface | Opaque handle returned by `spawn`. `id()` is the registry/routing key; `terminalId()` is the transport-level session id (herdr terminal UUID today). | -| `SpawnRequest` | record | Spawn parameters: `profileName`, `requestedCwd`, `callerCwd`. A null/blank profile means "use the default"; a null/blank cwd means "inherit from config or caller". | -| `Capability` | enum | Declared launcher capabilities: `MID_TURN_ASK`, `SELF_PR`, `WORKTREE`, `ORPHAN_REAP`. Stage A only *declares* them (advisory); verb-layer enforcement — a verb against a peer lacking a capability returning a clean "unsupported" rather than crashing — is planned, not yet wired. | +| Class | Role | +|---|---| +| `PeerLauncher` (`peer/PeerLauncher.java:22`) | SPI: `spawn`, `stop`, `effectiveCwd`, `parityOverlay`, `profiles`, `defaultProfile`, `list`, `reapOrphanWorkers`, `capabilities`, `capabilitiesFor`, `clearContext`. | +| `PeerHandle` (`peer/PeerHandle.java:12`) | Returned by `spawn`: `id()` is the registry/routing key, `terminalId()` the transport session id. | +| `SpawnRequest` (`peer/SpawnRequest.java:16`) | Spawn parameters, including `role` (CB-557: which contract) and `resumeSessionId`. | +| `Capability` (`peer/Capability.java:9`) | Declared adapter capabilities: `MID_TURN_ASK`, `SELF_PR`, `WORKTREE`, `CONTEXT_RESET`, `ORPHAN_REAP`, `SESSION_NAME`, `SESSION_RESUME`. | +| `MemberRole` (`peer/MemberRole.java:24`) | What a member is *for*: `ARCHITECT`, `DEV`, `REVIEWER` — independent of which backend (profile) it runs on. | +| `CharterReceipt` (`peer/CharterReceipt.java:21`) | A SHA-256 fingerprint of the exact charter text a member was launched with, for audit — never the charter text itself. | -### `mcp` — the MCP north face (7 classes) +**Shipped kinds are `claude-code` and `opencode` only** (`config/FleetConfig.java:333,335`). +AgentAPI was never built — no adapter for it exists in this source tree. + +### `member` — the peer adapters + +| Class | Role | +|---|---| +| `HerdrPeerLauncher` (`member/HerdrPeerLauncher.java:63`) | Abstract base for a peer materialized as a herdr agent — tab/pane placement, the CB-306 spawn-readiness gate, unique naming, orphan reap, teardown, listing, cwd resolution. Everything transport-shared lives here. | +| `ClaudeCodeLauncher` (`member/ClaudeCodeLauncher.java:43`) | The Claude Code adapter: builds the worker env (`ANTHROPIC_BASE_URL`), asserts the subscription guard, mounts the bridge MCP inline. | +| `OpenCodeLauncher` (`member/OpenCodeLauncher.java:54`) | The opencode adapter: no subscription guard (opencode is not on the subscription), writes an ephemeral `opencode.json` instead of inline flags, model as a `-m` flag. | +| `CompositePeerLauncher` (`member/CompositePeerLauncher.java:69`) | The `PeerLauncher` the core actually holds when more than one adapter is configured — a router in front of one `HerdrPeerLauncher` per peer kind, routing by profile, by pane id, or fanning out fleet-wide. | +| `MemberEnvAllowList` / `EnvAllowListScrub` | The credential-guard allow-list a spawned member's environment is scrubbed against (CB-596/CB-611 family). | + +### `mcp` — the MCP north face Exposes `fleetd` as a Streamable-HTTP MCP endpoint and resolves caller identity from the -**connection**, not from tool arguments (unspoofable). +**connection**, never from tool arguments. -| Class | Kind | Role | -|---|---|---| -| `BridgeMcp` | class | Builds the MCP server, registers `fleet_*` tools, holds thin tool adapters (`servlet`, `send`, `answer`, `ask`, `reply`, `spawn`). | -| `ConnectionIdentity` | class | Resolves *who is calling*: peer PID → herdr pane → `terminal_id` (`resolve`, `callerTerminal`, `cwdForPid`). | -| `PrimaryRegistry` | class | Single-slot thread-safe holder of the primary's `terminal_id` (CB-307 Stage 3). `record(id)` learns it from orchestration-side tools (`fleet_send`/`fleet_spawn`) when the caller resolves to a non-null terminal that is *not* a registered worker; `primaryTerminal()` / `isKnown()` feed the push loop. Optional constructor-pin (`primary.terminal` config) for an operator override; stays empty (→ push degrades to pull) when the primary is off-host or non-herdr. | -| `PeerPidLookup` | interface | Abstracts OS peer-PID lookup for a loopback source port. | -| `LsofPeerPidLookup` | class | `lsof` impl; excludes fleetd's own PID. | -| `ProcessCwdLookup` | interface | Abstracts PID → cwd lookup. | -| `LsofProcessCwdLookup` | class | `lsof`-based cwd lookup for spawn inheritance. | +| Class | Role | +|---|---| +| `FleetMcp` (`mcp/FleetMcp.java:67`) | Builds the MCP server, registers the `fleet_*` tools, holds the thin tool adapters. | +| `ConnectionIdentity` (`mcp/ConnectionIdentity.java:16`) | Resolves *who is calling* from the OS peer PID and herdr's pane map — unforgeable. | +| `PrimaryRegistry` (`mcp/PrimaryRegistry.java:22`) | Tracks the primary's terminal id, and (CB-532) which lead delegated to which worker, so a reply nudge reaches the right lead. | +| `LsofPeerPidLookup` / `LsofProcessCwdLookup` | `lsof`-based OS lookups behind the identity resolution. | -**Tool surface:** `fleet_send` (delegate + block; with `turnId`, answers an ask) · `fleet_reply` -(worker → structured answer; if no send is open it is now **held in the reply inbox**, not -errored) · `fleet_ask` (worker pauses to ask primary) · `fleet_status` · `fleet_poll` (async -ticket; with an optional `target`, **drains that worker's held replies** from the inbox) · -`fleet_ack` (ack one held reply by `msgId` — finer than drain-all) · `fleet_spawn` · -`fleet_list` · `fleet_stop` · `fleet_profiles`. -**Invariant:** `fleet_reply`/`fleet_ask` accept *no* identity argument; it comes only from the -transport context. **CB-307 scope:** only a *terminal* `fleet_reply` with no open send is queued — -`fleet_ask` (interactive; the worker blocks and can't consume a late answer) and the injector's -completion/failure fallbacks (they target a *captured* waiter, CB-116) are **never** queued. +**The registered tool set is exactly eleven tools** (`mcp/FleetMcp.java:301-326`): `fleet_send`, +`fleet_reply`, `fleet_ask`, `fleet_status`, `fleet_poll`, `fleet_ack`, `fleet_spawn`, `fleet_list`, +`fleet_stop`, `fleet_profiles`, `fleet_whoami`. **There is no `fleet_read` tool.** A `fleet_send` +carrying a `coordId` (instead of a `sessionId`/`turnId`) routes to a peer lead on another daemon +through `msg.LeadChannel`, not to a worker session — see `sendToLead` +(`mcp/FleetMcp.java:616-642`). -### `rest` · `member` · `lead` · `guard` · `config` — the edge +### `session` — the member registry and worktrees -| Class | Kind | Role | -|---|---|---| -| `rest.FleetdApp` | class | Javalin routes; validates bodies, maps `Outcome` → HTTP status. | -| `member.ClaudeCodeLauncher` | class | Guard-checked spawn, orphan-pane reaping at boot, teardown. The first-class `PeerLauncher` adapter for Claude Code over herdr (`spawn`, `reapOrphanWorkers`, `stop`, `list`, `profiles`, `capabilities`). | -| `guard.SubscriptionGuard` | class | Host-allowlist + primary-cleanliness enforcement (`assertWorker`, `assertPrimaryClean`). | -| `guard.GuardException` | class | Thrown on any subscription-boundary violation. | -| `config.FleetdConfig` | record | YAML config with defaults; single legacy worker or named `workers` map (`load`, `workerProfiles`, `defaultProfile`). | -| `Fleetd` | class | Static `main` — wires real collaborators and starts the server. | +| Class | Role | +|---|---| +| `SessionManager` (`session/SessionManager.java:42`) | Authoritative in-daemon registry of every member this process spawned. Implements `TurnListener` so the injector's turn boundaries drive `SPAWNING → READY → BUSY → DONE`/`FAILED`. One-shot: a finished member is torn down, never reused. | +| `MemberSession` (`session/MemberSession.java:35`) | The immutable record: pane id, terminal id, profile, `MemberRole`, cwd, owner, state, worktree/branch, `CharterReceipt`, `agentSessionId`. | +| `GitWorktrees` (`session/GitWorktrees.java:34`) | Production `Worktrees` implementation, shelling `git`. Neutralizes `.mcp.json` and opencode's repo-level config in every worktree so a member never inherits the primary's credentialed MCP mounts. | +| `Worktrees` (`session/Worktrees.java:7`) | The seam: `add`, `remove`, `hasUncommitted`, `overlayParity`, `repoRoot`, plus `snapshot`/`wipRefs`/`pruneWipRefs` for the CB-586 `refs/wip/*` retention sweep. | +| `SessionReaper` (`session/SessionReaper.java:13`) | Virtual-thread loop tearing down idle `READY`/`DONE` sessions past their TTL, and sweeping stale `refs/wip/*` snapshots. | -**REST routes** (the acceptance surface — every capability is reachable here without MCP): +### `health` — fleet health evidence (new package) + +Not in the daemon the wiki last described this page against. A pure classifier plus a slow +collection loop, deliberately separate from the delivery poller in `inject`. + +| Class | Role | +|---|---| +| `FleetHealth` (`health/FleetHealth.java:11`) | Pure classifier: `HealthSnapshot` + prior tick's `HealthPrior` → `HealthDecision`. No I/O. | +| `HealthState` (`health/HealthState.java:4`) | `STARTING`, `IDLE`, `WORKING`, `WORK_PENDING`, `BLOCKED_AMBIGUOUS`, `NEVER_READY`, `GONE`, `TURN_BOUNDARY_LOST`, `ERROR_ON_SCREEN`, `STALL_SUSPECTED`, `MUTE`, `REPLY_STRANDED`, `DELEGATION_ORPHANED`, `CONTROL_LINK_DOWN`. | +| `FleetHealthMonitor` (`health/FleetHealthMonitor.java:23`) | The scheduled loop that collects a `HealthSnapshot` per member and calls `FleetHealth.decide`. | +| `MuteCounter` (`health/MuteCounter.java:11`) | Counts turns that ended via the completion fallback instead of a structured `fleet_reply` — an observation, never a classifier state on its own. | + +`ERROR_ON_SCREEN` is declared but not decided yet — the class javadoc says it needs a bounded pane +read and an adapter-specific fatal signature that do not exist yet. + +### `lead` — starting configured leads + +| Class | Role | +|---|---| +| `LeadLauncher` (`lead/LeadLauncher.java:50`) | Starts the leads `fleet.leaders:` declares, when none is already running. Deliberately does **not** reuse `HerdrPeerLauncher`: a lead must never get the reply charter, the idle reaper, or `ANTHROPIC_BASE_URL` — all three are true of every member spawn path. | + +### `placement`, `metrics`, `guard`, `config`, `logging`, `rest` — the edge + +| Class | Role | +|---|---| +| `placement.PlacementPolicy` (`placement/PlacementPolicy.java:7`) | How an unqualified spawn picks a profile: `fixed` (the historical default-profile behaviour) or `weighted` (smooth weighted round-robin with `maxLoad` gating). | +| `placement.BackendQuarantine` (`placement/BackendQuarantine.java:25`) | Keyed by credential id, not profile: puts a credential on cooldown after a `BACKEND_EXHAUSTED` classification (CB-578 stage B), so a fresh spawn does not walk straight back onto an account that just refused. | +| `metrics.Metrics` (`metrics/Metrics.java:25`) | The registry and Prometheus text renderer — dependency-free by design. | +| `metrics.FleetMetrics` (`metrics/FleetMetrics.java:21`) | The named series: `fleet_sends_total`, `fleet_replies_total`, `fleet_push_nudges_total`, `fleet_lead_heartbeat_nudges_total`, `fleet_spawns_total`, `fleet_herdr_calls_total`, `fleet_auth_failures_total`, `fleet_sessions`, `fleet_inbox_depth`. | +| `guard.SubscriptionGuard` (`guard/SubscriptionGuard.java:23`) | `assertWorker` (the member's base_url must be on the off-subscription allowlist) and `assertPrimaryClean` (the primary's own env must carry none). | +| `config.FleetConfig` (`config/FleetConfig.java:81`) | The YAML config record: `memberHerdrSocket` (CB-185, optional second herdr daemon for members), `Coordinator` (CB-637 lead-to-lead broker), `Leader`, `Fleet` (the `architects`/`developers`/`reviewers` pools). | +| `logging.McpCancelledNotificationFilter` | Suppresses one noisy SDK warning for a normal MCP cancellation notice — cosmetic, not a control. | +| `rest.FleetApp` (`rest/FleetApp.java:46`) | Javalin routes; validates bodies, maps `Outcome` to an HTTP status. | +| `Fleetd` (`Fleetd.java`) | Static `main` — wires every real collaborator above and starts both faces. | + +## REST routes + +The acceptance surface — every capability MCP exposes is reachable here too, without Claude in the +loop (`rest/FleetApp.java:143-158`). **There is no `/events` route**, and the worker endpoints are +`/members`, not `/workers`. | Method + path | Purpose | |---|---| -| `GET /healthz` | Liveness + herdr reachability | -| `GET /sessions` | herdr workspace list | -| `GET /agents` | herdr-tracked agents | -| `GET /profiles` | Configured worker profiles + default | -| `POST /workers` | Spawn a worker (`?profile=`, `?cwd=`, or JSON body) | -| `DELETE /workers/{paneId}` | Stop a worker | -| `POST /sessions/{id}/message` | `fleet_send` — blocking, `wait:false`, or answer via `turnId` | -| `POST /sessions/{id}/reply` | `fleet_reply` (worker) | -| `POST /sessions/{id}/ask` | `fleet_ask` (worker → primary) | -| `GET /sessions/{id}/status` | Worker lifecycle + MCP readiness | -| `GET /sessions/{id}/replies` | Drain a worker's **held replies** from the inbox (CB-307): peek-then-ack, second call returns `[]` | -| `GET /tasks/{ticket}` | Poll an async (`wait:false`) send | +| `GET /healthz` | Liveness. Checks both herdr daemons when `memberHerdrSocket` is configured, and reports `protocolMismatch` if their protocol versions differ. | +| `GET /metrics` | Prometheus text exposition (present only when a metrics registry is wired). | +| `GET /sessions` | herdr workspace list, merged across both daemons. | +| `GET /agents` | herdr-tracked agents. | +| `GET /members` | Registry roster + live herdr status (CB-304). | +| `GET /profiles` | Configured backend profiles + the default. | +| `POST /members` | Spawn a member (`?role=&profile=&cwd=&worktree=&ticket=`, or a JSON body). | +| `DELETE /members/{paneId}` | Stop a member. | +| `POST /sessions/{id}/message` | `fleet_send` — blocking, `wait:false`, or an answer via `turnId`. | +| `POST /sessions/{id}/reply` | `fleet_reply` (member). | +| `GET /sessions/{id}/replies` | Drain a member's held replies from the inbox (peek + ack). | +| `POST /sessions/{id}/ask` | `fleet_ask` (member → primary). | +| `GET /sessions/{id}/status` | `fleet_status`, plus `ready` (the injector's own readiness gate). | +| `GET /tasks/{ticket}` | Poll an async (`wait:false`) send. | ## Bootstrap wiring @@ -234,25 +310,32 @@ completion/failure fallbacks (they target a *captured* waiter, CB-116) are **nev ```mermaid flowchart LR - cfg["load config
+ assertPrimaryClean"] --> guard["SubscriptionGuard"] - guard --> herdr["connect UnixSocketHerdrClient
→ AgentControl / WorkspaceControl"] - herdr --> ccl["member.ClaudeCodeLauncher
behind PeerLauncher SPI
→ reapOrphanWorkers()"] - ccl --> rv["Rendezvous → CompletionResolver
→ Injector + StatusPoller"] - rv --> ms["MessageService"] - ms --> mcp["BridgeMcp
(connection identity)"] - ms --> app["FleetdApp
(start Javalin)"] - mcp --> app + cfg["load config
+ assertPrimaryClean"] --> router["herdr.HerdrRouter
(lead daemon + optional member daemon)"] + router --> adapters["member.ClaudeCodeLauncher
member.OpenCodeLauncher
→ member.CompositePeerLauncher"] + adapters --> reap["reapOrphanWorkers()"] + reap --> sessions["session.SessionManager
+ session.GitWorktrees"] + sessions --> leads["lead.LeadLauncher
.ensureLeads()"] + leads --> rv["msg.Rendezvous → inject.Injector
+ inject.StatusPoller"] + rv --> push["msg.ReplyPushLoop
+ msg.LeadHeartbeatLoop"] + push --> ms["msg.MessageService"] + ms --> fh["health.FleetHealthMonitor"] + fh --> auth["auth.MemberRegistry
+ auth.CallerResolver.withLeadsAndMembers"] + auth --> mcp["mcp.FleetMcp"] + mcp --> app["rest.FleetApp
(start Javalin)"] ``` -*Figure 2 — startup wiring in `Fleetd.main`. The guard asserts the primary env is clean before -anything else; `PeerLauncher` is wired as an interface, with `ClaudeCodeLauncher` as the first -adapter; `reapOrphanWorkers()` clears stale panes from a prior daemon restart.* +*Figure 2 — startup wiring in `Fleetd.main`. The guard asserts the primary's own environment is +clean before anything else; the two adapters plug into one `CompositePeerLauncher`; +`reapOrphanWorkers()` clears stale panes left by a prior daemon restart; `CallerResolver` is built +last among the core services because it needs the live lead and architect bindings that everything +above it produces.* ## Flow: forward rendezvous (`fleet_send` → `fleet_reply`) -The main path. A primary's send blocks on a per-session waiter that resolves on the worker's -explicit reply — **or**, as a fallback, on a confirmed `working → idle` turn boundary the injector -observes (so a worker that finishes without calling `fleet_reply` still unblocks the caller). +The main path. A primary's send blocks on a per-session waiter that resolves on the member's +explicit reply — or, as a fallback, on a confirmed `working → injectable` turn boundary the +injector observes, so a member that finishes without calling `fleet_reply` still unblocks the +caller. ```mermaid sequenceDiagram @@ -261,60 +344,56 @@ sequenceDiagram participant MS as MessageService participant RV as Rendezvous participant IJ as Injector - participant W as Worker + participant W as Member P->>MS: send(target, content) Note over MS: acquire per-session ReentrantLock + MS->>RV: open(target) — register the waiter first MS->>IJ: enqueue(target, content) - MS->>RV: open(target) → CompletableFuture IJ->>W: inject when injectable (one msg/turn) IJ-->>MS: onDelivered (capture waiter + baseline) - alt worker replies explicitly + alt member replies explicitly W->>RV: fleet_reply → resolve(REPLY) else turn completes without reply IJ-->>RV: onTurnComplete → resolveCompletion(COMPLETION) - else worker fails / wedges + else member wedges (CB-109) IJ-->>RV: onTurnFailed → resolveFailure(FAILED) + else scrape matches a usage-limit refusal + IJ-->>RV: resolveExhausted(BACKEND_EXHAUSTED) end RV-->>MS: Resolution - MS-->>P: Reply (REPLIED / COMPLETED_UNREPLIED / WORKER_FAILED / TIMED_OUT_*) + MS-->>P: Reply (REPLIED / COMPLETED_UNREPLIED / WORKER_FAILED / BACKEND_EXHAUSTED / TIMED_OUT_*) ``` *Figure 3 — forward send. If the timeout fires first, the outcome is `TIMED_OUT_WORKING` -(delivered) or `TIMED_OUT_QUEUED` (never delivered).* +(delivered) or `TIMED_OUT_QUEUED` (never delivered). `BACKEND_EXHAUSTED` (CB-578) is a healthy pane +whose account refused on a usage limit — kept apart from `WORKER_FAILED` so the primary sees the +real cause.* -**The stranded-reply case (CB-307).** The figure is the happy path — a live waiter exists. The -failure that motivated CB-307 is the *reverse*: a worker finishes and calls `fleet_reply` **after** -its send has already timed out (the ~60s client window) or was never opened, so `Rendezvous.resolve` -finds no waiter. Before CB-307 that reply was silently discarded (the worker got an error / REST -`409`). Now `MessageService.reply` publishes it to the `ReplyInbox` under a fresh `msgId`, and the -primary collects it later keyed by target — `fleet_poll(target)` or `GET /sessions/{id}/replies` -(peek → deliver → ack, so an in-flight failure re-surfaces it). With `broker:` configured the -`AmqpReplyInbox` gives this **cross-restart durability** (Stage 2) — the held reply survives a -`java -jar` bounce and is redelivered; without it the in-memory adapter is soft-state (undrained -replies clear on restart). +**The stranded-reply case.** The figure is the happy path — a live waiter exists. The case that +motivated the reply inbox is the reverse: a member finishes and calls `fleet_reply` after its send +has already timed out, or one was never opened, so `Rendezvous.resolve` finds no waiter. Before +that inbox existed, the reply was silently lost. Now `MessageService.reply` +(`msg/MessageService.java:370`) publishes it into the `ReplyInbox` under a fresh id, and the +primary collects it later — `fleet_poll(target=…)` or `GET /sessions/{id}/replies`. With `broker:` +configured, `AmqpReplyInbox` gives this cross-restart durability; without it, the in-memory adapter +is soft-state and undrained replies are lost on a restart. -Delivery is no longer purely pull. When the reply lands with no waiter, `reply` also calls -`ReplyPushLoop.onReplyQueued(target)`, which **actively nudges the primary to drain** (Stage 3): it -injects a *"run `fleet_poll(target=…)`"* turn into the primary's own herdr pane — the same -`AgentControl.send` primitive that delivers to workers, pointed at the primary — but only when the -primary is `injectable()` (never mid-turn), bounded to `primary.push_reminders` nudges on a -`push_backoff_ms` schedule, and stopping the instant a drain empties the inbox. The nudge carries -*no payload* (the drain response is the clean transport), so re-nudging is idempotent, and if every -push fails the durable inbox is still the backstop. An off-host / non-herdr primary leaves -`PrimaryRegistry` empty, so the loop is a no-op and delivery cleanly degrades to pull. +Delivery is not purely pull. When a reply lands with no waiter, `reply` also calls +`ReplyPushLoop.onReplyQueued(target)`, which nudges the primary's own herdr pane to run +`fleet_poll` — but only when the primary is `injectable()` (never mid-turn), bounded to a +configured reminder count, and stopping the instant a drain empties the inbox. ## Flow: reverse rendezvous (`fleet_ask` → answer) -A worker pauses its own turn to ask the primary a question; the question surfaces on the primary's -open forward send, and the answer resumes the *same* worker turn. Duplicate asks from one session -coalesce onto a single `turnId` (only the `fresh` owner surfaces the question and tears the turn -down). +A member pauses its own turn to ask the primary a question; the question surfaces on the primary's +open forward send, and the answer resumes the *same* member turn. Duplicate asks from one session +coalesce onto a single `turnId`. ```mermaid sequenceDiagram autonumber - participant W as Worker + participant W as Member participant MS as MessageService participant RV as Rendezvous participant P as Primary @@ -326,45 +405,45 @@ sequenceDiagram RV->>RV: resolveQuestion(session, question, turnId) RV-->>P: forward send returns QUESTION + turnId P->>MS: answer(turnId, content) - MS->>RV: answerAsk(turnId, content) → unblock worker - MS->>RV: open(workerSession) — new forward waiter - W-->>RV: resumes turn → fleet_reply + MS->>RV: answerAsk(turnId, content) — unblock the member + MS->>RV: open(memberSession) — a new forward waiter + W-->>RV: resumes the turn → fleet_reply RV-->>P: Reply else no forward send open RV-->>W: AskOutcome.NO_WAITER (closeAsk) end ``` -*Figure 4 — reverse ask. Worker session is derived from `turnId` (`askSession`), never a caller -argument. If the `turnId` is unknown, `answer` returns `Outcome.STALE_TURN`.* +*Figure 4 — reverse ask. The member session is derived from `turnId` (`Rendezvous.askSession`), +never a caller argument. If the `turnId` is unknown, `answer` returns `Outcome.STALE_TURN`.* ## State machine: `AgentStatus` -The worker's herdr-reported lifecycle. `fromWire` maps `idle|working|blocked|done` and folds any -unrecognized string to `UNKNOWN`. A status is **injectable** (safe to deliver into) only when -`IDLE`, `BLOCKED`, or `DONE` — never mid-turn (`WORKING`). `UNKNOWN` both wedges delivery and is -reclassified by `StatusRefiner` from pane content. +The member's herdr-reported lifecycle (`herdr/AgentStatus.java:8`). `fromWire` maps +`idle|working|blocked|done` and folds any unrecognized string to `UNKNOWN`. A status is +**injectable** only when `IDLE`, `BLOCKED`, or `DONE` — never mid-turn (`WORKING`). `UNKNOWN` both +wedges delivery and is reclassified by `StatusRefiner` from the pane's real content. ```mermaid stateDiagram-v2 [*] --> IDLE IDLE --> WORKING: message injected, turn starts WORKING --> IDLE: turn completes - WORKING --> BLOCKED: worker asks / awaits input - WORKING --> DONE: worker exits turn + WORKING --> BLOCKED: member asks / awaits input + WORKING --> DONE: member exits the turn BLOCKED --> WORKING: unblocked DONE --> IDLE: next turn IDLE --> UNKNOWN: herdr can't classify WORKING --> UNKNOWN: herdr can't classify - UNKNOWN --> IDLE: StatusRefiner reads pane - UNKNOWN --> WORKING: StatusRefiner reads pane + UNKNOWN --> IDLE: StatusRefiner reads the pane + UNKNOWN --> WORKING: StatusRefiner reads the pane note right of WORKING not injectable end note note right of DONE - injectable; working to done - is a turn-boundary like IDLE + injectable; a turn-boundary + equivalent to IDLE end note ``` @@ -372,11 +451,10 @@ stateDiagram-v2 ## State machine: `Injector` per-target turn lifecycle -The heart of delivery reliability. Each target moves through a lifecycle driven by -`Injector.onStatus`; a turn is judged complete **only** from a confirmed `WORKING → injectable` -boundary (a real `WORKING` sample seen, then an injectable one). Three grace counters bound the -failure paths: `PICKUP_GRACE_POLLS = 8`, `TURN_STALL_GRACE_POLLS = 120`, -`READINESS_GRACE_POLLS = 240`. +The heart of delivery reliability (`inject/Injector.java:48`). A turn is judged complete only from +a confirmed `WORKING → injectable` boundary. Three grace counters bound the failure paths, all +still `PICKUP_GRACE_POLLS = 8`, `TURN_STALL_GRACE_POLLS = 120`, `READINESS_GRACE_POLLS = 240` +(`inject/Injector.java:58,68,81`). ```mermaid stateDiagram-v2 @@ -401,43 +479,11 @@ stateDiagram-v2 ``` *Figure 6 — per-target turn lifecycle in `Injector`. `Complete` fires `TurnListener.onTurnComplete` -(→ `CompletionResolver` scrapes the pane and resolves the captured waiter); `Failed` fires +(→ `CompletionResolver` scrapes the pane and resolves the captured waiter, or classifies it +`BACKEND_EXHAUSTED` when the scrape matches a configured usage-limit pattern); `Failed` fires `onTurnFailed`. `CompletionResolver` captures the exact waiter at delivery time and suppresses any -scrape matching the pre-turn baseline, so a late completion for turn N can never resolve turn -N+1's waiter (CB-116 safety).* - -## Subscription boundary - -The one non-negotiable invariant, enforced in code. `SubscriptionGuard.assertWorker(baseUrl)` runs -in `ClaudeCodeLauncher.spawn()` **before any herdr call**: the worker's `ANTHROPIC_BASE_URL` host must -be on the allowlist (Stage-1: `gx00.gw`, `ollama.ltms.dev`). `assertPrimaryClean` (called at -startup) hard-stops if the primary env carries any `ANTHROPIC_BASE_URL`. Spawn injects -`ANTHROPIC_BASE_URL`/`ANTHROPIC_MODEL`/`CLAUDE_CONFIG_DIR`/`ANTHROPIC_AUTH_TOKEN` into the -**worker's** env only — `fleetd`'s own env is never mutated. CWD resolves -`requestedCwd → profile cwd → caller cwd → daemon user.dir → "."`. - -## Concurrency model (at a glance) - -- **`herdr`** — lock-free: each `UnixSocketHerdrClient.call` owns its own socket; safe to call - concurrently from virtual threads. `WorkspaceControl.ensureWorkspace` is `synchronized`. -- **`inject`** — one virtual-thread poller loop; each target has its own monitor so - check-and-send is serialized; futures/listeners complete *after* the monitor is released to - keep the poller unblocked. -- **`msg`** — per-session `ReentrantLock` serializes forward sends (at most one forward waiter - per session); `Rendezvous` uses `ConcurrentHashMap` + `CompletableFuture` for cross-thread - resolution; async sends run on a dedicated virtual-thread executor. At most one open reverse-ask - turn per session (coalesced via `openAsksBySession`). -- **`mcp`** — identity resolution shells out to `lsof` once per worker-identity request; shared - services own their own concurrency. -- **`member`** — `AtomicLong` name sequence + per-process nonce avoid herdr name collisions - without locks; the orphan reaper is best-effort and never aborts startup. - -## Related pages - -- [1. Architecture](1-Architecture) — system, invariants, the two modes (design level). -- [2. Message Server](2-Message-Server) — the `fleetd` design rationale, herdr control contract. -- [8. Roadmap](8-Roadmap) — stages, tickets, and the feature ⇄ endpoint ⇄ test map. - ---- -*Generated from a fan-out code audit (one worker per layer) and verified against source at main -`3aa69a9`.* +scrape matching the pre-turn baseline, so a late completion for turn N can never resolve turn N+1's +waiter (CB-116 safety). Not shown: an optional post-turn action a `TurnListener` can request after +`Complete`, which the injector waits to settle before the next delivery.
+`Complete`/`Failed` state names are a naming device for this diagram — the code tracks the same +lifecycle as boolean flags and counters on a per-target `Track` record, not a literal Java enum. \ No newline at end of file