CB-634: rename bridged -> fleetd across the wiki
Match the code cutover: daemon name, config (fleetd.yaml), scripts, launchd/ systemd units, module dir, and MCP tool prefix bridge_* -> fleet_*. Kept: the BRIDGED_MEMBER security marker, mcp__bridge__ (historical mount name), and the .bridged-worktrees on-disk path. The portable CLAUDE.md block stays byte-identical with the repo's CLAUDE.md.
+60
-60
@@ -3,17 +3,17 @@
|
||||
`claude-bridge` lets a **primary** Claude Code session — Opus 4.8 on your Pro/Max
|
||||
subscription — drive one or more **secondary worker** Claude Code sessions running a
|
||||
*different, cheaper/local* model, **without ever putting a proxy on the primary**. A
|
||||
standalone daemon, **`bridged`**, sits between them: it drives [herdr](https://herdr.dev) (an
|
||||
standalone daemon, **`fleetd`**, sits between them: it drives [herdr](https://herdr.dev) (an
|
||||
agent multiplexer) to inject turns and read live agent-status, and exposes an **MCP server**
|
||||
that every Claude session mounts.
|
||||
|
||||
Two invariants define the whole design; everything else follows from them.
|
||||
|
||||
> **1 — Subscription boundary.** Only a *worker* process ever sets `ANTHROPIC_BASE_URL`; the
|
||||
> primary never does, so it stays on Pro/Max. `bridged` is not a `claude` process and holds
|
||||
> primary never does, so it stays on Pro/Max. `fleetd` is not a `claude` process and holds
|
||||
> zero Anthropic quota.
|
||||
>
|
||||
> **2 — Sole gateway.** Every Claude session talks *only* to `bridged`, over the MCP tools it
|
||||
> **2 — Sole gateway.** Every Claude session talks *only* to `fleetd`, over the MCP tools it
|
||||
> mounts. No session addresses a broker, a peer session, or the network directly.
|
||||
|
||||
## System overview
|
||||
@@ -25,13 +25,13 @@ flowchart TB
|
||||
WP["worker pane(s) · claude<br/>ANTHROPIC_BASE_URL set · MCP client"]
|
||||
end
|
||||
|
||||
subgraph BD["bridged — standalone daemon · THE gateway (no Anthropic quota)"]
|
||||
subgraph BD["fleetd — standalone daemon · THE gateway (no Anthropic quota)"]
|
||||
SRV["SERVER face<br/>MCP server · REST/SSE · policy brain"]
|
||||
CLI["CLIENT face<br/>status-gated injector · herdr socket client"]
|
||||
SRV --> CLI
|
||||
end
|
||||
|
||||
Q["broker / queue<br/>(bridged-owned · below the gateway)"]
|
||||
Q["broker / queue<br/>(fleetd-owned · below the gateway)"]
|
||||
M["worker model<br/>ollama.ltms.dev · GX10 vLLM"]
|
||||
|
||||
PP -->|"MCP tools"| SRV
|
||||
@@ -48,10 +48,10 @@ flowchart TB
|
||||
class Q,M warn
|
||||
```
|
||||
|
||||
*Figure: the Claude sessions are **herdr panes**; they reach *up* into `bridged`'s SERVER face
|
||||
over MCP, while `bridged`'s CLIENT face drives them *down* through herdr's socket. `bridged` is
|
||||
*Figure: the Claude sessions are **herdr panes**; they reach *up* into `fleetd`'s SERVER face
|
||||
over MCP, while `fleetd`'s CLIENT face drives them *down* through herdr's socket. `fleetd` is
|
||||
the only thing any session connects to — the broker sits below the gateway line, owned by
|
||||
`bridged`, touched by no session. Only worker panes carry `ANTHROPIC_BASE_URL`.*
|
||||
`fleetd`, touched by no session. Only worker panes carry `ANTHROPIC_BASE_URL`.*
|
||||
|
||||
## The two invariants
|
||||
|
||||
@@ -61,13 +61,13 @@ The point of the bridge is to keep the **primary** on Pro/Max while a **worker**
|
||||
cheaper/local model, with no policy-violating proxy on the primary.
|
||||
|
||||
- The **primary** `claude` **never** sets `ANTHROPIC_BASE_URL`. It authenticates to
|
||||
`api.anthropic.com` on your subscription and reaches the worker *only* through `bridged`'s
|
||||
`api.anthropic.com` on your subscription and reaches the worker *only* through `fleetd`'s
|
||||
MCP tools — never by re-pointing its own endpoint.
|
||||
- Only a **worker** `claude` launches with `ANTHROPIC_BASE_URL=https://ollama.ltms.dev`
|
||||
(+ `ANTHROPIC_AUTH_TOKEN`) or a GX10 vLLM URL. Because model choice is just per-process env,
|
||||
this also sidesteps Claude Code's lack of per-subagent provider routing — the worker is a
|
||||
*separate process*, not a subagent.
|
||||
- **`bridged` holds no quota**, so it may hold a permanent herdr event subscription and poll
|
||||
- **`fleetd` holds no quota**, so it may hold a permanent herdr event subscription and poll
|
||||
its own internal queue with no policy concern. Its **subscription guard** refuses to spawn a
|
||||
worker pane whose resolved `ANTHROPIC_BASE_URL` host isn't on an off-subscription allowlist,
|
||||
and refuses to ever set that var on a pane tagged *primary*. (The guard validates the
|
||||
@@ -79,31 +79,31 @@ cheaper/local model, with no policy-violating proxy on the primary.
|
||||
|
||||
### Sole gateway — how sessions communicate
|
||||
|
||||
Because every session mounts `bridged` over MCP, `bridged` is the single chokepoint for *all*
|
||||
Because every session mounts `fleetd` over MCP, `fleetd` is the single chokepoint for *all*
|
||||
agent traffic. This is a deliberate simplification with real payoffs:
|
||||
|
||||
- **No session-side transport.** A Claude session calls MCP tools and nothing else — no
|
||||
`curl`, no broker client, no hand-rolled hook polling a queue. There is no shell step that
|
||||
could leak env, so mounting the bridge is subscription-safe by construction.
|
||||
- **Async is a push, not a poll.** When a message arrives for an idle session, `bridged`
|
||||
- **Async is a push, not a poll.** When a message arrives for an idle session, `fleetd`
|
||||
**injects it into that session's pane over herdr**, gated on the live `agent_status`. The
|
||||
session is never asked to busy-poll anything; the old "primary must never perpetual-poll"
|
||||
footgun disappears because there is nothing for it to poll.
|
||||
- **Guardrails are centrally enforced.** Round/turn budgets (anti ping-pong), rate limits,
|
||||
per-session authz, and audit all live in `bridged` — one place — instead of cooperative
|
||||
per-session authz, and audit all live in `fleetd` — one place — instead of cooperative
|
||||
sentinels each session must honour.
|
||||
- **The broker is infrastructure, below the line.** If `bridged` needs durability or a host
|
||||
- **The broker is infrastructure, below the line.** If `fleetd` needs durability or a host
|
||||
hop it owns a broker/queue for that. No Claude session sees it.
|
||||
|
||||
## Components
|
||||
|
||||
| Component | Role | Notes |
|
||||
|---|---|---|
|
||||
| **`bridged` — SERVER face** | The gateway: **MCP server** (the contract every session mounts) + REST/SSE for non-Claude clients, over the **policy brain** — session tracker, subscription guard, reply rendezvous. | `bridge_send` · `bridge_reply` · `bridge_ask` · `bridge_status` · `bridge_spawn` · `bridge_list` · `bridge_stop` · `bridge_read`. |
|
||||
| **`bridged` — CLIENT face** | Drives herdr: a **status-gated injector** (per-pane FIFO, delivers only when `agent_status ∈ {idle, blocked}`) over a **herdr socket client** (NDJSON, id-correlated, live event stream). | Single writer per pane → no injector-vs-injector races. |
|
||||
| **`fleetd` — SERVER face** | The gateway: **MCP server** (the contract every session mounts) + REST/SSE for non-Claude clients, over the **policy brain** — session tracker, subscription guard, reply rendezvous. | `fleet_send` · `fleet_reply` · `fleet_ask` · `fleet_status` · `fleet_spawn` · `fleet_list` · `fleet_stop` · `fleet_read`. |
|
||||
| **`fleetd` — CLIENT face** | Drives herdr: a **status-gated injector** (per-pane FIFO, delivers only when `agent_status ∈ {idle, blocked}`) over a **herdr socket client** (NDJSON, id-correlated, live event stream). | Single writer per pane → no injector-vs-injector races. |
|
||||
| **herdr** | Agent multiplexer. Owns the PTYs, panes, persistence, and — crucially — **`agent_status_changed` events**. Claude sessions run here as panes. | Socket is **local-only**; young/single-dev → injector kept pluggable. |
|
||||
| **Worker `claude`** | A *real* Claude Code process (inherits `CLAUDE.md`, hooks, skills, MCP), pointed at a different model. Recyclable, not immortal. | Only these carry `ANTHROPIC_BASE_URL`. Materialized by a swappable `PeerLauncher` adapter (`ClaudeCodeLauncher` today); the core bus is peer-neutral. |
|
||||
| **Broker / queue** *(optional, internal)* | `bridged`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `bridged` will later inject. | Redis Streams / NATS JetStream, or an embedded queue for a single host. |
|
||||
| **Broker / queue** *(optional, internal)* | `fleetd`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `fleetd` will later inject. | Redis Streams / NATS JetStream, or an embedded queue for a single host. |
|
||||
| **AgentAPI** *(fallback)* | Swappable injector behind the CLIENT face if herdr is unavailable. | Screen-stability heuristic instead of events — see [Approaches](3-Approaches). |
|
||||
|
||||
## Traffic: two modes across the gateway
|
||||
@@ -113,31 +113,31 @@ the connection.
|
||||
|
||||
### Mode 1 — blocking request/response (the main path)
|
||||
|
||||
The primary delegates with **one** `bridge_send` tool call. `bridged` holds it open (the
|
||||
The primary delegates with **one** `fleet_send` tool call. `fleetd` holds it open (the
|
||||
primary is idle-waiting, spending no quota) and **resolves it on whichever lands first**: the
|
||||
worker's structured `bridge_reply`, or the worker's turn-done status edge
|
||||
worker's structured `fleet_reply`, or the worker's turn-done status edge
|
||||
(`agent_status: working → idle` — herdr has no `done` status). The reply comes
|
||||
back as the tool result — so worker → primary rides `bridged`'s state and needs **no keystroke
|
||||
back as the tool result — so worker → primary rides `fleetd`'s state and needs **no keystroke
|
||||
into the primary pane**, even single-host.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as "Primary (Opus) — MCP client"
|
||||
participant S as "bridged (gateway)"
|
||||
participant S as "fleetd (gateway)"
|
||||
participant H as "herdr"
|
||||
participant W as "Worker (other model)"
|
||||
|
||||
P->>S: "bridge_send(worker, task) — tool call PARKS"
|
||||
P->>S: "fleet_send(worker, task) — tool call PARKS"
|
||||
S->>S: "await pane agent_status = idle"
|
||||
S->>H: "pane.send_text + send_keys (inject)"
|
||||
H->>W: "new turn"
|
||||
activate W
|
||||
H-->>S: "event: agent_status = working"
|
||||
W->>S: "bridge_reply(result) — structured (preferred)"
|
||||
W->>S: "fleet_reply(result) — structured (preferred)"
|
||||
H-->>S: "event: agent_status working → idle (turn done)"
|
||||
deactivate W
|
||||
S-->>P: "tool result = reply (unparks the call)"
|
||||
Note over P,W: "resolves on bridge_reply or the idle edge — whichever lands first.<br/>a blocked worker returns as the tool result, then the primary re-answers"
|
||||
Note over P,W: "resolves on fleet_reply or the idle edge — whichever lands first.<br/>a blocked worker returns as the tool result, then the primary re-answers"
|
||||
```
|
||||
|
||||
*Figure: a single parked tool call, not a busy-poll. SSE (`GET /events`) carries status to
|
||||
@@ -145,33 +145,33 @@ sequenceDiagram
|
||||
may outrun a sane request timeout, use Mode 2.*
|
||||
|
||||
> **Late replies are no longer lost (CB-307).** If the worker finishes *after* the primary's
|
||||
> blocking `bridge_send` has already timed out (or was never opened), its `bridge_reply` finds no
|
||||
> live waiter to resolve. `bridged` now **holds that reply in a per-worker inbox** rather than
|
||||
> discarding it, and the primary collects it later keyed by target (`bridge_poll(target)` /
|
||||
> blocking `fleet_send` has already timed out (or was never opened), its `fleet_reply` finds no
|
||||
> live waiter to resolve. `fleetd` now **holds that reply in a per-worker inbox** rather than
|
||||
> discarding it, and the primary collects it later keyed by target (`fleet_poll(target)` /
|
||||
> `GET /sessions/{id}/replies`). With a broker configured the reply is **durable** across a daemon
|
||||
> restart (Stage 2, `AmqpReplyInbox` on LavinMQ); without one it is soft-state in-memory. And
|
||||
> delivery is **active, not just pull** (Stage 3): the moment a reply lands with no open send,
|
||||
> `bridged` injects a bounded, status-gated *drain nudge* into a same-host primary's own herdr pane
|
||||
> `fleetd` injects a bounded, status-gated *drain nudge* into a same-host primary's own herdr pane
|
||||
> — so the primary need not be polling to notice. An off-host / non-herdr primary keeps the pull
|
||||
> path. See [Roadmap](8-Roadmap#delivery-reliability--multi-host-cb-306--cb-308).
|
||||
|
||||
### Mode 2 — asynchronous delivery (bridged-mediated)
|
||||
### Mode 2 — asynchronous delivery (fleetd-mediated)
|
||||
|
||||
For traffic with no caller waiting — a detached progress report, an out-of-band question, a
|
||||
webhook injecting work — the recipient must be *woken*. `bridged` does the waking by
|
||||
webhook injecting work — the recipient must be *woken*. `fleetd` does the waking by
|
||||
**injecting the recipient's idle pane**, driven by the live `agent_status`, so delivery is
|
||||
event-driven rather than a timed poll. If durability or a host hop is needed, `bridged`
|
||||
event-driven rather than a timed poll. If durability or a host hop is needed, `fleetd`
|
||||
enqueues internally first; the recipient still receives by injection.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant SRC as "Source: worker bridge_reply/ask · webhook · bus"
|
||||
participant S as "bridged (gateway)"
|
||||
participant SRC as "Source: worker fleet_reply/ask · webhook · bus"
|
||||
participant S as "fleetd (gateway)"
|
||||
participant Q as "queue (internal)"
|
||||
participant H as "herdr"
|
||||
participant R as "Recipient pane (idle Claude)"
|
||||
|
||||
SRC->>S: "MCP bridge_reply / bridge_ask · or REST ingress"
|
||||
SRC->>S: "MCP fleet_reply / fleet_ask · or REST ingress"
|
||||
opt durability / cross-host
|
||||
S->>Q: "enqueue (ack + visibility timeout)"
|
||||
end
|
||||
@@ -179,30 +179,30 @@ sequenceDiagram
|
||||
H-->>S: "event: idle"
|
||||
S->>H: "pane.send_text + send_keys (inject)"
|
||||
H->>R: "new turn = the message"
|
||||
Note over S,R: "recipient polled nothing — bridged pushed on the idle edge"
|
||||
Note over S,R: "recipient polled nothing — fleetd pushed on the idle edge"
|
||||
```
|
||||
|
||||
*Figure: `bridged` mediates async in both directions. Sources reach it over the gateway (MCP
|
||||
*Figure: `fleetd` mediates async in both directions. Sources reach it over the gateway (MCP
|
||||
for Claude, REST for external), it optionally parks the message on its internal queue, waits
|
||||
for the idle event, and injects.*
|
||||
|
||||
> **Where a hook is used.** A `Stop`-hook appears in exactly the spots where neither an MCP
|
||||
> tool call nor a pane injection can serve — and in **every** case it targets `bridged`, never
|
||||
> tool call nor a pane injection can serve — and in **every** case it targets `fleetd`, never
|
||||
> a broker:
|
||||
> - a **split-host primary that is not a herdr pane** (e.g. Opus on your Mac) — the one session
|
||||
> `bridged` cannot inject into (a *same-host* primary **is** a herdr pane and gets the CB-307
|
||||
> Stage-3 drain nudge) — runs a `Stop`-hook that **long-polls `bridged`** for queued messages;
|
||||
> `fleetd` cannot inject into (a *same-host* primary **is** a herdr pane and gets the CB-307
|
||||
> Stage-3 drain nudge) — runs a `Stop`-hook that **long-polls `fleetd`** for queued messages;
|
||||
> and
|
||||
> - a **non-MCP (herdr-only) worker** may run a `Stop`-hook that **POSTs its reply to
|
||||
> `bridged`** at turn end, a structured alternative to scraping the pane (see
|
||||
> `fleetd`** at turn end, a structured alternative to scraping the pane (see
|
||||
> [Message Server](2-Message-Server) → *Reply envelope*).
|
||||
>
|
||||
> The gateway invariant holds in every topology: a hook is just a transport adapter to
|
||||
> `bridged` for a session MCP/injection can't reach.
|
||||
> `fleetd` for a session MCP/injection can't reach.
|
||||
|
||||
## Worker lifecycle — the Ralph loop
|
||||
|
||||
`bridged` treats a worker as a **recyclable** resource, not one immortal session: a long-lived
|
||||
`fleetd` treats a worker as a **recyclable** resource, not one immortal session: a long-lived
|
||||
pane fills its context window. On a context/idle cap it checkpoints and respawns fresh, with
|
||||
continuity carried by **artifacts on disk** (git commits, a `STATE.md`/task file the worker is
|
||||
told to keep current) — **not** `claude --resume`, which would reload the context you are
|
||||
@@ -214,7 +214,7 @@ stateDiagram-v2
|
||||
Spawning --> Ready: "claude prompt detected"
|
||||
Ready --> Working: "turn injected"
|
||||
Working --> Blocked: "permission / question"
|
||||
Blocked --> Working: "bridged answers (send_input)"
|
||||
Blocked --> Working: "fleetd answers (send_input)"
|
||||
Working --> Ready: "agent_status → idle (turn done)"
|
||||
Ready --> Recycling: "context / idle cap hit"
|
||||
Recycling --> Spawning: "state persisted to disk"
|
||||
@@ -229,16 +229,16 @@ detail is in [Message Server](2-Message-Server) → *Worker session lifecycle*.*
|
||||
|
||||
## Topologies
|
||||
|
||||
- **Same-host (default).** Primary, `bridged`, herdr, and workers on one off-subscription box.
|
||||
Every session is a herdr pane, so `bridged` can inject *either* direction; a broker is
|
||||
- **Same-host (default).** Primary, `fleetd`, herdr, and workers on one off-subscription box.
|
||||
Every session is a herdr pane, so `fleetd` can inject *either* direction; a broker is
|
||||
optional. Simplest to run and the focus of the design.
|
||||
- **Split-host.** Primary Opus on your Mac; `bridged` + herdr + workers on the GPU host near
|
||||
- **Split-host.** Primary Opus on your Mac; `fleetd` + herdr + workers on the GPU host near
|
||||
the model. herdr's socket stays local, so the Mac reaches the worker host **only over
|
||||
`bridged`'s MCP/HTTP endpoint**. The primary isn't a herdr pane, so async wake-ups use the
|
||||
`Stop`-hook-polls-`bridged` path above. Deployment diagrams: [Message Server](2-Message-Server)
|
||||
`fleetd`'s MCP/HTTP endpoint**. The primary isn't a herdr pane, so async wake-ups use the
|
||||
`Stop`-hook-polls-`fleetd` path above. Deployment diagrams: [Message Server](2-Message-Server)
|
||||
→ *Deployment model*.
|
||||
- **Multi-host federation (planned — CB-308).** Not one `bridged` fronting remote workers, but
|
||||
**many gateways, one per host**, meshed over a broker. Each host runs its own `bridged` owning
|
||||
- **Multi-host federation (planned — CB-308).** Not one `fleetd` fronting remote workers, but
|
||||
**many gateways, one per host**, meshed over a broker. Each host runs its own `fleetd` owning
|
||||
its local herdr + registry; agents get a **host-unique global id** and a per-agent broker inbox
|
||||
`agent.<globalId>.inbox` whose *owning* gateway is the sole consumer; a **federated roster**
|
||||
(soft-state presence on a `roster.*` topic) gives every gateway an eventually-consistent
|
||||
@@ -253,12 +253,12 @@ detail is in [Message Server](2-Message-Server) → *Worker session lifecycle*.*
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph hostA["Host A"]
|
||||
gA["gateway = bridged A<br/>local herdr + registry"]
|
||||
gA["gateway = fleetd A<br/>local herdr + registry"]
|
||||
P["primary (MCP client)"]
|
||||
P --- gA
|
||||
end
|
||||
subgraph hostB["Host B"]
|
||||
gB["gateway = bridged B<br/>local herdr + registry"]
|
||||
gB["gateway = fleetd B<br/>local herdr + registry"]
|
||||
WB["worker panes"]
|
||||
gB --- WB
|
||||
end
|
||||
@@ -285,31 +285,31 @@ flowchart TB
|
||||
|
||||
*Figure: multi-host is an addressing + routing concern, not a new kind of peer. Each gateway owns
|
||||
its own herdr and consumes only its own agents' inboxes; the broker routes between them; the
|
||||
federated roster is soft-state (bridged owns who/where/status, the broker owns message durability).
|
||||
federated roster is soft-state (fleetd owns who/where/status, the broker owns message durability).
|
||||
No host ever sees another host's terminals.*
|
||||
|
||||
**Security:** `bridged` is an agent-control surface — `bridge_send` runs arbitrary prompts and
|
||||
**Security:** `fleetd` is an agent-control surface — `fleet_send` runs arbitrary prompts and
|
||||
key-passthrough sends raw keystrokes into a live agent. Bind it to `localhost` + SSH tunnel, or
|
||||
front it with a bearer token + TLS; never expose the port unauthenticated. One `bridged` is
|
||||
front it with a bearer token + TLS; never expose the port unauthenticated. One `fleetd` is
|
||||
today **one trust domain** (no per-session authz yet — an open item in
|
||||
[Message Server](2-Message-Server)).
|
||||
|
||||
## Failure modes & single points of failure
|
||||
|
||||
Making `bridged` the sole gateway buys a clean model at the cost of a real SPOF. Degrade
|
||||
Making `fleetd` the sole gateway buys a clean model at the cost of a real SPOF. Degrade
|
||||
deliberately:
|
||||
|
||||
| What dies | Effect | Recovery |
|
||||
|---|---|---|
|
||||
| **`bridged`** | **All** agent comms stop — sync *and* async — since it is the only gateway; in-flight blocking calls error out. | herdr + workers keep running (state on disk / queue). systemd restarts `bridged`; it re-attaches to existing panes via `workspace.list`/`pane.list` and drains its queue. This restart path is load-bearing — harden it. |
|
||||
| **`fleetd`** | **All** agent comms stop — sync *and* async — since it is the only gateway; in-flight blocking calls error out. | herdr + workers keep running (state on disk / queue). systemd restarts `fleetd`; it re-attaches to existing panes via `workspace.list`/`pane.list` and drains its queue. This restart path is load-bearing — harden it. |
|
||||
| **herdr** | No pane control; all delivery (sync + async injection) dead. | PTYs die with the *server* (only client detach survives). Respawn workers from persisted state (Ralph loop); replay unacked queue items. |
|
||||
| **Broker / queue** (internal) | Durability + cross-host async degrade; **same-host async still works** (idle-injection needs no queue). | `bridged` delivers locally without it; ack + visibility timeout re-deliver on recovery. Nothing silently dropped. |
|
||||
| **Model endpoint** | Workers stall or error mid-turn. | herdr status shows `working` stuck / `blocked`; `bridged` times out the blocking call and surfaces the error. |
|
||||
| **Broker / queue** (internal) | Durability + cross-host async degrade; **same-host async still works** (idle-injection needs no queue). | `fleetd` delivers locally without it; ack + visibility timeout re-deliver on recovery. Nothing silently dropped. |
|
||||
| **Model endpoint** | Workers stall or error mid-turn. | herdr status shows `working` stuck / `blocked`; `fleetd` times out the blocking call and surfaces the error. |
|
||||
| **All of the above** | Full bridge outage. | The **primary is never downstream** of any bridge component — it stays fully usable on its own subscription. The bridge is additive, never on the primary's critical path. |
|
||||
|
||||
## Related pages
|
||||
|
||||
- **[Message Server](2-Message-Server)** — the `bridged` design in depth: MCP contract, herdr control, reply rendezvous, lifecycle, API, tech stack, milestones.
|
||||
- **[Message Server](2-Message-Server)** — the `fleetd` design in depth: MCP contract, herdr control, reply rendezvous, lifecycle, API, tech stack, milestones.
|
||||
- **[Approaches](3-Approaches)** — why herdr-centric, and the full transport comparison (AgentAPI, Agent SDK, queue+Stop-hook, tmux).
|
||||
- **[Team](6-Team)** — a team-lead orchestrating a mixed Claude + local-LLM worker fleet over the same gateway.
|
||||
- **[Setup](4-Setup)** · **[Operations](5-Operations)** — bring-up and the day-2 runbook.
|
||||
|
||||
+10
-10
@@ -46,7 +46,7 @@ flowchart LR
|
||||
main["human-driven lead<br/>(MCP client)"]
|
||||
wrk["worker / sandboxed agent<br/>(herdr pane)"]
|
||||
end
|
||||
gw["gateway = bridged<br/>(one per host)"]
|
||||
gw["gateway = fleetd<br/>(one per host)"]
|
||||
arch -->|"delivered by local MCP pull"| gw
|
||||
main -->|"delivered by local MCP pull"| gw
|
||||
wrk -->|"delivered by local herdr inject"| gw
|
||||
@@ -112,7 +112,7 @@ CB-308 is a **migration, not a rename**. Migrated queues take a version-suffixed
|
||||
(`agent.<gid>.inbox.v2`): AMQP refuses to redeclare an existing durable queue with new arguments
|
||||
(`PRECONDITION_FAILED` — a crash loop on an in-place upgrade from v1.0.0), and the suffix keeps
|
||||
old sender-keyed and new recipient-keyed queues apart while both exist. The per-worker drain
|
||||
surface (`bridge_poll(target)`) survives by filtering on the envelope's `from` field (§2.1). This
|
||||
surface (`fleet_poll(target)`) survives by filtering on the envelope's `from` field (§2.1). This
|
||||
is also where CB-201's story completes: descoping the envelope was right on one host (connection
|
||||
identity routes everything) and wrong across hosts (a broker hop has none) — the envelope returns
|
||||
as §2.1.*
|
||||
@@ -171,7 +171,7 @@ co-located with that entity.** Gateways additionally own a control queue and a r
|
||||
|
||||
Why the roster queue is **transient** while inboxes are **durable**: the broker owns *message*
|
||||
durability, **not** who/where/status. Presence is rebuilt from heartbeats on reconnect; a missed
|
||||
heartbeat expires an entry (CB-303 TTL thinking). This preserves the persistence boundary — bridged
|
||||
heartbeat expires an entry (CB-303 TTL thinking). This preserves the persistence boundary — fleetd
|
||||
stays soft-state; only *messages* are durable. See [1. Architecture](1-Architecture).
|
||||
|
||||
## 5. Invariants
|
||||
@@ -186,7 +186,7 @@ stays soft-state; only *messages* are durable. See [1. Architecture](1-Architect
|
||||
globalId; the broker routes to whichever gateway holds that inbox. Senders are oblivious to the
|
||||
recipient's host — the roster resolves *existence*, the broker resolves *location*.
|
||||
3. **The final hop is a pull for MCP clients.** For a primary/main/orchestrator the broker makes the
|
||||
*middle* hop lossless + ordered + idempotent, but the *last* hop is still `bridge_poll` / drain —
|
||||
*middle* hop lossless + ordered + idempotent, but the *last* hop is still `fleet_poll` / drain —
|
||||
the gateway **holds** the reply until the client pulls (consume-and-hold). The broker does not
|
||||
dissolve the MCP asymmetry (§ CB-308 "the one thing the broker does NOT dissolve").
|
||||
4. **Keystrokes never traverse the broker.** Only messages + presence. Injection into an agent is a
|
||||
@@ -255,22 +255,22 @@ sequenceDiagram
|
||||
participant BR as bridge.msg
|
||||
participant GB as gateway B
|
||||
participant W as worker (host B)
|
||||
MA->>GA: bridge_send(gidW, content)
|
||||
MA->>GA: fleet_send(gidW, content)
|
||||
GA->>GA: roster - gidW local? NO
|
||||
GA->>BR: publish(key = gidW)
|
||||
BR->>GB: route to agent.gidW.inbox
|
||||
GB->>W: inject via B's LOCAL herdr
|
||||
W-->>GB: bridge_reply(to gidMA)
|
||||
W-->>GB: fleet_reply(to gidMA)
|
||||
GB->>BR: publish(key = gidMA, durable)
|
||||
BR->>GA: route to agent.gidMA.inbox
|
||||
Note over GA: held until MA pulls (MA is an MCP client)
|
||||
MA->>GA: blocking send resolves / bridge_poll
|
||||
MA->>GA: blocking send resolves / fleet_poll
|
||||
GA-->>MA: reply
|
||||
```
|
||||
|
||||
*Figure 4 — U1. Both injection points (into W on B, and the drain into MA on A) are local; only the
|
||||
two middle hops cross the broker. U3 (stranded reply) is the same picture where step 8's "held"
|
||||
outlives A's send window and MA collects it later by `bridge_poll(gidW)` / drain.*
|
||||
outlives A's send window and MA collects it later by `fleet_poll(gidW)` / drain.*
|
||||
|
||||
### 7.2 U4 — cross-host spawn (the control plane)
|
||||
|
||||
@@ -282,7 +282,7 @@ sequenceDiagram
|
||||
participant CX as bridge.control
|
||||
participant GB as gateway B
|
||||
participant SL as SandboxLauncher (B)
|
||||
MA->>GA: bridge_spawn(profile, host = B)
|
||||
MA->>GA: fleet_spawn(profile, host = B)
|
||||
GA->>CX: publish(key = hostB, SpawnRequest)
|
||||
CX->>GB: route to control.hostB
|
||||
GB->>SL: spawn locally (into a local sandbox)
|
||||
@@ -342,7 +342,7 @@ sequenceDiagram
|
||||
participant BR as bridge.msg
|
||||
participant GA as gateway A
|
||||
participant MA as main (host A)
|
||||
W->>GB: bridge_ask(question)
|
||||
W->>GB: fleet_ask(question)
|
||||
GB->>BR: publish ASK (key = gidMA, expiresAt, turn_id)
|
||||
BR->>GA: route
|
||||
GA->>MA: surfaces on the open send / drain (local pull)
|
||||
|
||||
+147
-147
@@ -6,7 +6,7 @@ capability, so a feature that shipped six weeks ago is still findable without re
|
||||
or a commit log.
|
||||
|
||||
**What belongs here.** A capability an operator can *use, configure, or observe*: an MCP tool, a
|
||||
`bridged.yaml` knob, an endpoint, or a behaviour visible from outside the daemon. Internal contract
|
||||
`fleetd.yaml` knob, an endpoint, or a behaviour visible from outside the daemon. Internal contract
|
||||
changes go to [Implementation](9-Implementation); test and coverage work is a
|
||||
[Roadmap](8-Roadmap) line. If a change adds none of those, it has no entry here — that is a normal
|
||||
outcome, not an omission.
|
||||
@@ -19,28 +19,28 @@ six weeks, and the table alone will not carry it.
|
||||
|
||||
| Capability | Turn it on with | Since | Code |
|
||||
|---|---|---|---|
|
||||
| [Ask the bridge who you are](#ask-the-bridge-who-you-are) | `bridge_whoami` | CB-517 | `mcp/BridgeMcp` |
|
||||
| [Ask the bridge who you are](#ask-the-bridge-who-you-are) | `fleet_whoami` | CB-517 | `mcp/BridgeMcp` |
|
||||
| [Primary inside a herdr pane](#primary-inside-a-herdr-pane) | `primary.terminal:` | CB-522 | `auth/CallerResolver` |
|
||||
| [More than one lead](#more-than-one-lead) | `fleet.leaders:` | CB-530 | `auth/CallerResolver` |
|
||||
| [Unknown config keys are named](#unknown-config-keys-are-named) | (always on) | CB-530 | `config/BridgedConfig` |
|
||||
| [Unknown config keys are named](#unknown-config-keys-are-named) | (always on) | CB-530 | `config/FleetdConfig` |
|
||||
| [Find leads by tab name](#find-leads-by-tab-name) | `fleet.leaders.<n>.tabPrefix` | CB-531 | `herdr/LeadTabScanner` |
|
||||
| [One fleet block, role as the key](#one-fleet-block-role-as-the-key) | `fleet:` | CB-557 | `config/BridgedConfig` |
|
||||
| [One fleet block, role as the key](#one-fleet-block-role-as-the-key) | `fleet:` | CB-557 | `config/FleetdConfig` |
|
||||
| [Role pools decide the backend](#role-pools-decide-the-backend) | `fleet.developers:` etc. | CB-557 | `member/CompositePeerLauncher` |
|
||||
| [Tab labels name the role](#tab-labels-name-the-role) | `fleet.tabLabel:` | CB-557 | `member/HerdrPeerLauncher` |
|
||||
| [Launch a lead when none is live](#launch-a-lead-when-none-is-live) | `fleet.leaders.<n>.profile` + `instances` | CB-558 | `lead/LeadLauncher` |
|
||||
| [Re-read the config without a restart](#re-read-the-config-without-a-restart) | `configReload.enabled: true` | CB-559 | `config/ConfigRef` |
|
||||
| [Set a role's launch charter in config](#set-a-roles-launch-charter-in-config) | `fleet.charters:` | CB-566 | `config/BridgedConfig` |
|
||||
| [Set a role's launch charter in config](#set-a-roles-launch-charter-in-config) | `fleet.charters:` | CB-566 | `config/FleetdConfig` |
|
||||
| [Leads talk to each other](#leads-talk-to-each-other) | (always on, two leads) | CB-532 | `auth/Principal` |
|
||||
| [A lead can be delivered to](#a-lead-can-be-delivered-to) | automatic | CB-534 | `Bridged.deliverableTo` |
|
||||
| [A lead can be delivered to](#a-lead-can-be-delivered-to) | automatic | CB-534 | `Fleetd.deliverableTo` |
|
||||
| [Know when a completion fallback is partial](#know-when-a-completion-fallback-is-partial) | automatic | CB-563 | `inject/CompletionResolver` |
|
||||
| [Leads are visible in bridge_list](#leads-are-visible-in-bridge_list) | automatic | CB-535 | `mcp/BridgeMcp.listFleet` |
|
||||
| [Leads are visible in fleet_list](#leads-are-visible-in-fleet_list) | automatic | CB-535 | `mcp/BridgeMcp.listFleet` |
|
||||
| [Reply nudges follow the delegating lead](#reply-nudges-follow-the-delegating-lead) | automatic (retires `primary:`) | CB-532 | `mcp/PrimaryRegistry` |
|
||||
| [Advisory architect slots](#advisory-architect-slots) | `fleet.architects:` | CB-548 | `auth/MemberRegistry` |
|
||||
| [Weighted worker placement](#weighted-worker-placement) | `placement: weighted` + `weight` / `maxLoad` | CB-518 | `placement/` |
|
||||
| [Give workers a toolchain](#give-workers-a-toolchain) | per-profile `env:` | CB-511 | `worker/HerdrPeerLauncher` |
|
||||
| [Run a worker on the subscription](#run-a-worker-on-the-subscription) | profile `subscription: true` | CB-539 | `worker/ClaudeCodeLauncher` |
|
||||
| [Keep a worker conversation](#keep-a-worker-conversation) | session name + resume id on spawn | CB-547 | `peer/SpawnRequest` |
|
||||
| [Isolated worktree per worker](#isolated-worktree-per-worker) | `bridge_spawn{worktree, ticket}` | CB-301-ext | `session/GitWorktrees` |
|
||||
| [Isolated worktree per worker](#isolated-worktree-per-worker) | `fleet_spawn{worktree, ticket}` | CB-301-ext | `session/GitWorktrees` |
|
||||
| [Worker tool-surface isolation](#worker-tool-surface-isolation) | automatic | CB-525 | `session/GitWorktrees` |
|
||||
| [Worktree-hostile config isolation](#worktree-hostile-config-isolation) | automatic | CB-543 | `session/GitWorktrees` |
|
||||
| [Worker opens its own PR](#worker-opens-its-own-pr) | `gitTokenEnv:` / `gitHostEnv:` | CB-302 | `worker/HerdrPeerLauncher` |
|
||||
@@ -64,13 +64,13 @@ six weeks, and the table alone will not carry it.
|
||||
| [Tell a usage-limit refusal from a real reply](#tell-a-usage-limit-refusal-from-a-real-reply) | profile `exhaustedPattern:` | CB-578 | `inject/CompletionResolver` |
|
||||
| [Stop spawning onto an exhausted account](#stop-spawning-onto-an-exhausted-account) | `quarantineCooldownSeconds:` + profile `credentialId:` | CB-578 | `placement/BackendQuarantine` |
|
||||
| [See which charter a member got](#see-which-charter-a-member-got) | automatic | CB-571 | `peer/CharterReceipt` |
|
||||
| [Redeploy the daemon safely](#redeploy-the-daemon-safely) | run the script | — | `scripts/redeploy-bridged.sh` |
|
||||
| [Redeploy the daemon safely](#redeploy-the-daemon-safely) | run the script | — | `scripts/redeploy-fleetd.sh` |
|
||||
|
||||
Nearly every knob above lives in one file, on one profile:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Y["bridged.yaml"] --> G["bind / auth / guard"]
|
||||
Y["fleetd.yaml"] --> G["bind / auth / guard"]
|
||||
Y --> B["broker"]
|
||||
Y --> L["lifecycle"]
|
||||
Y --> W["profiles:"]
|
||||
@@ -96,7 +96,7 @@ so the arrows meet: the same profile may serve several roles.*
|
||||
|
||||
## Ask the bridge who you are
|
||||
|
||||
**What.** `bridge_whoami` returns `{"role":"primary"}` or `{"role":"worker", sessionId, profile,
|
||||
**What.** `fleet_whoami` returns `{"role":"primary"}` or `{"role":"worker", sessionId, profile,
|
||||
worktree, branch}`.
|
||||
|
||||
**On.** Always available; no configuration.
|
||||
@@ -105,7 +105,7 @@ worktree, branch}`.
|
||||
`CLAUDE.md` — a worker runs in a worktree of the same repo, so it inherits the file verbatim. Before
|
||||
this, a session had to *infer* its role from side-channels (the mount name, `ANTHROPIC_BASE_URL`,
|
||||
the system prompt), each of which is one-way and some of which are absent for Claude-model workers.
|
||||
The daemon already resolves the role from the connection for its authorization gate; `bridge_whoami`
|
||||
The daemon already resolves the role from the connection for its authorization gate; `fleet_whoami`
|
||||
just exposes that same answer, so guessing is never necessary.
|
||||
|
||||
**Gotcha.** The answer comes from the connection and cannot be forged or overridden by an argument.
|
||||
@@ -116,13 +116,13 @@ If it disagrees with what you expect, the daemon is right and your assumption is
|
||||
|
||||
**What.** Lets the orchestrating session run inside a herdr pane instead of an outside terminal.
|
||||
|
||||
**On.** `primary: terminal: term_<id>` — read the id from `bridge_whoami`, re-pin whenever the
|
||||
**On.** `primary: terminal: term_<id>` — read the id from `fleet_whoami`, re-pin whenever the
|
||||
primary moves panes.
|
||||
|
||||
**Why.** Caller identity resolves a loopback PID to its herdr pane, and the pane scan covers *every*
|
||||
pane, not just bridged-spawned ones. So a primary living in a pane classified itself as a worker and
|
||||
pane, not just fleetd-spawned ones. So a primary living in a pane classified itself as a worker and
|
||||
was refused spawn/send/stop — every verb it exists to call. The failure is **self-locking**: the
|
||||
daemon can also *learn* the primary's terminal, but only from `bridge_send`/`bridge_spawn`, the
|
||||
daemon can also *learn* the primary's terminal, but only from `fleet_send`/`fleet_spawn`, the
|
||||
exact calls being refused. Only an operator-set pin breaks the cycle, which is why the pinned value
|
||||
is consulted and the learned one deliberately is not.
|
||||
|
||||
@@ -141,7 +141,7 @@ reproduces the historical always-the-default-profile behaviour.
|
||||
**Why.** Profiles differ in model and cost, not tier. Without weights, every unqualified spawn piles
|
||||
onto one backend regardless of what it costs or how loaded it is.
|
||||
|
||||
**Gotcha.** An explicit `profile:` on `bridge_spawn` bypasses the policy entirely — placement only
|
||||
**Gotcha.** An explicit `profile:` on `fleet_spawn` bypasses the policy entirely — placement only
|
||||
governs *unqualified* spawns. Equal-weight candidates tie-break on **YAML definition order**, so
|
||||
that order is load-bearing config, not cosmetics (CB-524).
|
||||
|
||||
@@ -199,7 +199,7 @@ architects:
|
||||
|
||||
**Why.** The rejected alternative put an orchestrator above the lead and made the lead a managed,
|
||||
resumable session. That breaks the actual control boundary: the human drives a lead that pre-exists,
|
||||
is recognised, and cannot be resumed by bridged. Two advisory model families, Claude Sonnet 5 and
|
||||
is recognised, and cannot be resumed by fleetd. Two advisory model families, Claude Sonnet 5 and
|
||||
GPT-5.6 through opencode, receive the same brief independently so agreement is evidence rather than
|
||||
correlated echo.
|
||||
|
||||
@@ -212,7 +212,7 @@ Shipping the binding is not the same as shipping the role. For one day an archit
|
||||
resolved as `architect`, and could not receive a single message: the presence map that opens the
|
||||
readiness gate was keyed on `Role.WORKER`, so an architect was never marked available. It built clean
|
||||
and passed two reviewers. When a change adds a role, check every place that assumes a member *is* a
|
||||
worker — `bridge_status` tells you whether a member ever became ready.
|
||||
worker — `fleet_status` tells you whether a member ever became ready.
|
||||
|
||||
## Keep a worker conversation
|
||||
|
||||
@@ -232,7 +232,7 @@ core from assuming Claude Code's flags are a universal protocol.
|
||||
|
||||
**Gotcha.** OpenCode has no name flag, and it writes its session record only after it persists a
|
||||
conversation. Its id is therefore discovered lazily and may be absent immediately after spawn; its slot
|
||||
name remains only in bridged's roster.
|
||||
name remains only in fleetd's roster.
|
||||
|
||||
## Give workers a toolchain
|
||||
|
||||
@@ -240,7 +240,7 @@ name remains only in bridged's roster.
|
||||
|
||||
**On.** Automatic for `PATH`; add `env: {JAVA_HOME: ..., ...}` on a profile to extend or override.
|
||||
|
||||
**Why.** A worker's environment does **not** come from your shell. `bridged` hands herdr an explicit
|
||||
**Why.** A worker's environment does **not** come from your shell. `fleetd` hands herdr an explicit
|
||||
env map and herdr merges it into *its own* process env — so before this, a worker inherited whatever
|
||||
`PATH` the herdr server happened to be started with. On a long-lived herdr that can predate your
|
||||
toolchain entirely, leaving workers unable to run `mvn` or `java` at all.
|
||||
@@ -251,14 +251,14 @@ the `SubscriptionGuard`. A `subscription: true` profile is the exception: it ski
|
||||
injects neither key, so an `ANTHROPIC_BASE_URL`/`ANTHROPIC_AUTH_TOKEN` in its `env:` would survive
|
||||
unguarded — that configuration is refused at load (see [Run a worker on the subscription](#run-a-worker-on-the-subscription)).
|
||||
Since the default `PATH` is the *daemon's* own, start the daemon with a good one (see the
|
||||
`PATH` lines in `deploy/dev.ltms.bridged.plist` and `deploy/bridged.service`).
|
||||
`PATH` lines in `deploy/dev.ltms.fleetd.plist` and `deploy/fleetd.service`).
|
||||
|
||||
## Isolated worktree per worker
|
||||
|
||||
**What.** Provisions a git worktree on its own branch per worker, and copies a configurable set of
|
||||
local config files in ("parity overlay") so the worker sees the same local setup.
|
||||
|
||||
**On.** `bridge_spawn{worktree: true, ticket: "cb-123"}`; overlay list via `parityOverlay:`
|
||||
**On.** `fleet_spawn{worktree: true, ticket: "cb-123"}`; overlay list via `parityOverlay:`
|
||||
(default `[.claude/settings.local.json, .env, .envrc]`).
|
||||
|
||||
**Why.** Parallel workers editing one checkout collide. A worktree gives each its own branch and
|
||||
@@ -322,7 +322,7 @@ works.
|
||||
PR but must not be able to merge — the primary is the gate, and a worker that can merge is not
|
||||
gated.
|
||||
|
||||
**Where the value comes from.** `gitTokenEnv:` names a variable, and `bridged` reads it from **its own
|
||||
**Where the value comes from.** `gitTokenEnv:` names a variable, and `fleetd` reads it from **its own
|
||||
process environment** — so the token must be exported in the shell that launches the daemon, not
|
||||
stored in the repo. Keep it beside the operator's other credentials in one sourced file; a per-repo
|
||||
copy is a second copy of the same secret, and the copy you forget is the one that leaks or goes
|
||||
@@ -343,17 +343,17 @@ delegation N+1 from inheriting delegation N's conversation.
|
||||
**Gotcha.** `clearAfterTurn` defaults to `false`. It uses Claude Code's `/clear` command without
|
||||
counting that housekeeping as a delegated turn, and waits for it to settle before delivering the
|
||||
next task. Peer kinds without a known context-reset operation (currently opencode) treat the knob as
|
||||
a no-op and log that once; bridged never guesses a command. Lifecycle config is read at boot, so
|
||||
a no-op and log that once; fleetd never guesses a command. Lifecycle config is read at boot, so
|
||||
changes need a daemon restart.
|
||||
|
||||
## Durable reply inbox
|
||||
|
||||
**What.** A worker's reply survives with no waiter attached: it is queued and collected later by
|
||||
`bridge_poll` / `bridge_ack`, over a real AMQP broker when one is configured.
|
||||
`fleet_poll` / `fleet_ack`, over a real AMQP broker when one is configured.
|
||||
|
||||
**On.** `broker:` pointing at an AMQP URI; omit it for an in-memory inbox.
|
||||
|
||||
**Why.** A blocking `bridge_send` is capped by the *caller's* MCP client timeout (~60s), far below a
|
||||
**Why.** A blocking `fleet_send` is capped by the *caller's* MCP client timeout (~60s), far below a
|
||||
real task's runtime. Without a durable inbox, a reply arriving after that window lands nowhere.
|
||||
|
||||
**Gotcha.** The broker is **LavinMQ**, not RabbitMQ. Point it at a stray local RabbitMQ and you are
|
||||
@@ -364,10 +364,10 @@ the session is still BUSY currently discards the later completion rather than pa
|
||||
|
||||
**What.** `broker.uriEnv:` names an environment variable that holds the AMQP URI, instead of writing
|
||||
the URI into the config file. An AMQP URI carries `user:password@` inline, so the old `broker.uri:`
|
||||
put a live password in clear text in `bridged.yaml`.
|
||||
put a live password in clear text in `fleetd.yaml`.
|
||||
|
||||
**On.** `broker: { uriEnv: LAVINMQ_URI }`. The variable must be on the **daemon's own** environment,
|
||||
so start the daemon from a login shell — `scripts/redeploy-bridged.sh --check` reports whether the
|
||||
so start the daemon from a login shell — `scripts/redeploy-fleetd.sh --check` reports whether the
|
||||
named variable resolves, without ever printing its value.
|
||||
|
||||
**Why.** The same reason `auth.tokenEnv` and `Profile.tokenEnv` exist: a config file gets read,
|
||||
@@ -382,7 +382,7 @@ who moved to the secret store must never be silently returned to clear text. Bot
|
||||
|
||||
## An unreachable broker does not stop the daemon
|
||||
|
||||
**What.** If the broker cannot be reached **at boot**, bridged logs a loud warning and starts on the
|
||||
**What.** If the broker cannot be reached **at boot**, fleetd logs a loud warning and starts on the
|
||||
in-memory reply inbox for that process lifetime, instead of failing to start at all.
|
||||
|
||||
**On.** Always on. There is no background retry: fix the broker and restart to get durability back.
|
||||
@@ -427,7 +427,7 @@ path — the endpoint you name is the endpoint it uses.
|
||||
|
||||
## Onboard a project with the plugin
|
||||
|
||||
**What.** A Claude Code plugin that makes any project bridge-ready: it mounts the `bridged` MCP
|
||||
**What.** A Claude Code plugin that makes any project bridge-ready: it mounts the `fleetd` MCP
|
||||
gateway and ships a `/claude-bridge:setup` skill that runs preflight, applies standard project
|
||||
settings, names the environment variables the operator must export, and verifies the session
|
||||
resolves as the primary.
|
||||
@@ -444,7 +444,7 @@ the artifact is public-safe: every secret is referenced by environment-variable
|
||||
never enters a file.
|
||||
|
||||
**Gotcha.** The plugin is client-side setup only — it mounts a daemon, it does not install one.
|
||||
`bridged` and `herdr` remain separate services, and the setup skill deliberately refuses to install
|
||||
`fleetd` and `herdr` remain separate services, and the setup skill deliberately refuses to install
|
||||
them (guessing at a system-service install is how you get two daemons on one socket). Note also the
|
||||
plugin root is `plugin/`, **not** the repo root: an installed plugin's `.mcp.json` is a committed
|
||||
file, while this repo's root `.mcp.json` is local-only and `--skip-worktree`, so rooting the plugin
|
||||
@@ -494,7 +494,7 @@ leaders:
|
||||
```
|
||||
|
||||
`terminal` is the only field identity depends on; `kind`/`model` are descriptive and are echoed back
|
||||
by `bridge_whoami` as `leader: <name>`. `role` still reads `primary` — a lead **is** a primary for
|
||||
by `fleet_whoami` as `leader: <name>`. `role` still reads `primary` — a lead **is** a primary for
|
||||
authorization, so nothing keying on the role breaks.
|
||||
|
||||
**Why.** `primary.terminal` is singular by construction: one pane is the lead and every other pane
|
||||
@@ -531,14 +531,14 @@ merge, and an explicit `leaders:` entry outranks a label for the same terminal.
|
||||
|
||||
**Why.** A lead is never spawned — a human opens a tab and starts an agent in it — so unlike a
|
||||
worker, the daemon cannot learn its `terminal_id` at creation. `leaders:` therefore costs a
|
||||
four-step ritual per lead: start the session, ask it `bridge_whoami` for its id, edit config,
|
||||
four-step ritual per lead: start the session, ask it `fleet_whoami` for its id, edit config,
|
||||
restart. Naming the tab is one step, taken at the moment the operator is already there. The label
|
||||
also survives what the id does not: close and reopen the tab and the `terminal_id` changes, while
|
||||
the label is retyped as-is.
|
||||
|
||||
**Gotcha.** The direction of trust is what makes this safe, and it is one-way: bridged **reads**
|
||||
**Gotcha.** The direction of trust is what makes this safe, and it is one-way: fleetd **reads**
|
||||
lead tab labels and never writes them, so what is in the tab bar is always what a human typed.
|
||||
Two guards keep that from eroding — the configured worker spaces (where bridged *does* write
|
||||
Two guards keep that from eroding — the configured worker spaces (where fleetd *does* write
|
||||
labels, via `tabLabel`) are excluded from the scan wholesale, so nothing the bridge places can land
|
||||
in a matching tab; and startup **refuses** a `tabPrefix` that any worker `tabLabel` also matches,
|
||||
because overlapping those two namespaces would have the daemon label its own workers as leads and
|
||||
@@ -548,9 +548,9 @@ registry — a herdr hiccup must not demote a live lead mid-session.
|
||||
|
||||
## Leads talk to each other
|
||||
|
||||
**What.** A lead can message another lead and *be answered*. `bridge_send{sessionId: <peer's
|
||||
terminal>}` reaches a peer, and the peer closes the exchange with `bridge_reply` — the same
|
||||
rendezvous a worker uses. `bridge_whoami` now reports a lead's own `sessionId`, which is how a lead
|
||||
**What.** A lead can message another lead and *be answered*. `fleet_send{sessionId: <peer's
|
||||
terminal>}` reaches a peer, and the peer closes the exchange with `fleet_reply` — the same
|
||||
rendezvous a worker uses. `fleet_whoami` now reports a lead's own `sessionId`, which is how a lead
|
||||
learns the address to give a peer.
|
||||
|
||||
**On.** Nothing to configure; it applies as soon as two panes resolve as leads (via
|
||||
@@ -558,7 +558,7 @@ learns the address to give a peer.
|
||||
|
||||
**Why.** CB-530 and CB-531 widened *recognition* — both leads are seen — but nothing had widened
|
||||
*addressing*, so collaboration was one-way and silently so: the send was accepted, the peer's
|
||||
`bridge_reply` was refused, and the sender waited out its timeout. The cause was that a lead's
|
||||
`fleet_reply` was refused, and the sender waited out its timeout. The cause was that a lead's
|
||||
`Principal` carried no terminal, so `ownsSession()` could never be true for it and the `REPLY`/`ASK`
|
||||
rules excluded it by construction. A lead now carries the pane it was matched by, and the rule it
|
||||
must satisfy is unchanged: **you may act as the pane you occupy, and as no other**. That was always
|
||||
@@ -570,15 +570,15 @@ whose turn it is not, and never as a way to end its own turn. Widening who may r
|
||||
what they may reply *as*: a lead still cannot act for another pane, and an unnamed primary (token
|
||||
mode, or off-host, with no pane at all) owns nothing and remains a sender only.
|
||||
|
||||
## Leads are visible in bridge_list
|
||||
## Leads are visible in fleet_list
|
||||
|
||||
**What.** `bridge_list` returns `leads` alongside `workers`. Each lead row carries its `sessionId`
|
||||
(the address to `bridge_send` to), its `name`, its live status, and `self: true` on the caller's own
|
||||
**What.** `fleet_list` returns `leads` alongside `workers`. Each lead row carries its `sessionId`
|
||||
(the address to `fleet_send` to), its `name`, its live status, and `self: true` on the caller's own
|
||||
row.
|
||||
|
||||
**On.** Automatic, wherever a pane resolves as a lead.
|
||||
|
||||
**Why.** A lead had no way to discover a peer. `bridge_list` enumerated the worker roster alone, so a
|
||||
**Why.** A lead had no way to discover a peer. `fleet_list` enumerated the worker roster alone, so a
|
||||
lead asking "who else is here?" got an empty array — which reads as *no peers* but only ever meant
|
||||
*no workers spawned*. The peer lead on this bridge drew exactly that wrong conclusion and reported
|
||||
itself alone in a two-lead fleet. Addresses had to be carried between panes by a human, which is not
|
||||
@@ -623,29 +623,29 @@ It was also intermittently masked: presence is a sticky set, so a pane that was
|
||||
|
||||
## Reply nudges follow the delegating lead
|
||||
|
||||
**What.** When a worker's reply lands with no `bridge_send` open, the CB-307 nudge goes to the lead
|
||||
**What.** When a worker's reply lands with no `fleet_send` open, the CB-307 nudge goes to the lead
|
||||
that delegated that worker — not to a globally-configured "the primary". This is what retires
|
||||
`primary.terminal:`, which now logs a deprecation warning at startup.
|
||||
|
||||
**On.** Automatic. Delete `primary:` from `bridged.yaml`; keep it only if you still want
|
||||
**On.** Automatic. Delete `primary:` from `fleetd.yaml`; keep it only if you still want
|
||||
`pushReminders`/`pushBackoffMs`, or a fallback nudge destination across restarts.
|
||||
|
||||
**Why.** `PrimaryRegistry` held one slot answering "who is the primary" — a question with no correct
|
||||
answer once two leads drive one fleet. Whichever lead called `bridge_send` first captured *every*
|
||||
answer once two leads drive one fleet. Whichever lead called `fleet_send` first captured *every*
|
||||
nudge thereafter, so the other lead's results were announced to the wrong pane. The binding that
|
||||
actually matters is per-delegation and is known exactly where it is created: at `bridge_send`, where
|
||||
actually matters is per-delegation and is known exactly where it is created: at `fleet_send`, where
|
||||
the target is the argument and the lead is resolved from the connection.
|
||||
|
||||
**Gotcha.** A restart loses the delegation map while the durable inbox keeps the reply. With one
|
||||
lead the old pin (or the first lead to send) is an unambiguous fallback; with several and no
|
||||
recorded delegation the daemon nudges **nobody** rather than guessing, and delivery degrades to
|
||||
`bridge_poll`. That is the correct degradation — interrupting the wrong lead with someone else's
|
||||
`fleet_poll`. That is the correct degradation — interrupting the wrong lead with someone else's
|
||||
result is worse than a quiet inbox — but it does mean a post-restart reply may need an explicit
|
||||
poll.
|
||||
|
||||
## Unknown config keys are named
|
||||
|
||||
**What.** A top-level key in `bridged.yaml` that this build does not understand is logged as a WARN
|
||||
**What.** A top-level key in `fleetd.yaml` that this build does not understand is logged as a WARN
|
||||
naming it, at load.
|
||||
|
||||
**On.** Always on; nothing to configure.
|
||||
@@ -711,7 +711,7 @@ have one answer for a fleet that has three kinds of member.
|
||||
|
||||
**Gotcha.** A role with **no** pool is unconstrained, not blocked — it falls back to every profile,
|
||||
so a config that pools some roles and not others keeps working. And an **explicit** profile is not
|
||||
confined to the pool: `bridge_spawn{profile:"opus"}` carries no role, so it defaults to `dev`, and
|
||||
confined to the pool: `fleet_spawn{profile:"opus"}` carries no role, so it defaults to `dev`, and
|
||||
judging it against the dev pool would refuse a spawn the operator asked for by name. `maxLoad` still
|
||||
applies to it.
|
||||
|
||||
@@ -750,7 +750,7 @@ daemon restart after a reboot left a fleet with no orchestrator and no sign of w
|
||||
|
||||
**Gotcha — three, and they are the whole design.** An auto-launched lead is **not** a member: it
|
||||
gets no worker reply charter (that text tells its reader it is an off-subscription worker who must
|
||||
end every turn with `bridge_reply` — the opposite of an orchestrator), it is never registered with
|
||||
end every turn with `fleet_reply` — the opposite of an orchestrator), it is never registered with
|
||||
`SessionManager` (the idle reaper would kill it for being idle, which is a lead's normal state), and
|
||||
`ANTHROPIC_BASE_URL`/`AUTH_TOKEN` are stripped from its env whatever the profile says.
|
||||
|
||||
@@ -767,7 +767,7 @@ placed in one would never be found again and would be relaunched on every boot.
|
||||
|
||||
## Re-read the config without a restart
|
||||
|
||||
**What.** `bridged` watches `bridged.yaml`'s modified time and re-reads the file when it changes.
|
||||
**What.** `fleetd` watches `fleetd.yaml`'s modified time and re-reads the file when it changes.
|
||||
Consumers read the live config at the point of use, so a change reaches the next spawn without
|
||||
rebuilding anything.
|
||||
|
||||
@@ -835,7 +835,7 @@ after the daemon started resolving architects as well.
|
||||
|
||||
**Gotcha — the type is a map on purpose, and blank is not the same as absent.**
|
||||
|
||||
`BridgedConfig.Fleet` is `@JsonIgnoreProperties(ignoreUnknown = true)`. A typed record field named
|
||||
`FleetdConfig.Fleet` is `@JsonIgnoreProperties(ignoreUnknown = true)`. A typed record field named
|
||||
`architetc` would be dropped in silence, so the operator would see a clean "config reloaded" and no
|
||||
charter. A `Map<String,String>` keeps every key the operator wrote, which lets `validateCharters()`
|
||||
see the bad one and refuse the whole config, naming the key and listing the valid ones.
|
||||
@@ -855,10 +855,10 @@ world-readable.
|
||||
|
||||
## Know when a completion fallback is partial
|
||||
|
||||
**What.** If a member ends a turn without `bridge_reply`, bridged scrapes its pane and resolves the
|
||||
**What.** If a member ends a turn without `fleet_reply`, fleetd scrapes its pane and resolves the
|
||||
waiting send with that tail, so the sender is not left hanging. The tail is capped at
|
||||
`MAX_SCRAPE_CHARS` (4000). When the cap bites, the returned text now ends with
|
||||
`[Pane tail clipped: member did not call bridge_reply.]` and bridged logs a WARN with the original
|
||||
`[Pane tail clipped: member did not call fleet_reply.]` and fleetd logs a WARN with the original
|
||||
length and the cap.
|
||||
|
||||
**On.** Automatic, whenever the completion fallback reads more than 4000 characters.
|
||||
@@ -870,7 +870,7 @@ deliberate and unchanged; the defect was silence, not the number.
|
||||
|
||||
**Gotcha.** The marker does not recover the missing text, and it is not a licence to skip the reply.
|
||||
The fallback is a liveness net, not a channel: it returns only the pane tail, stripped to the last
|
||||
assistant block. A member must still end every delegated turn with exactly one `bridge_reply`. The
|
||||
assistant block. A member must still end every delegated turn with exactly one `fleet_reply`. The
|
||||
marker is appended **after** the CB-115 misattribution guard compares the scrape to its baseline, so
|
||||
marking cannot make an unchanged pane look like new output.
|
||||
|
||||
@@ -878,9 +878,9 @@ marking cannot make an unchanged pane look like new output.
|
||||
|
||||
## Reject a profile name as a send target
|
||||
|
||||
**What.** `bridge_send` refuses a `sessionId` that exactly matches a configured profile name, on both
|
||||
**What.** `fleet_send` refuses a `sessionId` that exactly matches a configured profile name, on both
|
||||
the blocking and the `wait:false` path, before any ticket is issued. The error names the value and
|
||||
points the caller at `bridge_list`.
|
||||
points the caller at `fleet_list`.
|
||||
|
||||
**On.** Always on; no configuration.
|
||||
|
||||
@@ -888,8 +888,8 @@ points the caller at `bridge_list`.
|
||||
it and returned `accepted — task delegated`, so the lead believed the work was dispatched. About 60
|
||||
seconds later the injector logged `sol is gone, dropping its queue`, and twenty minutes after that the
|
||||
ticket still reported `pending — worker unknown`. The whole delegation was lost and nothing told the
|
||||
sender. Confusing a profile for a session id is the single easiest mistake to make with `bridge_send`,
|
||||
because both are short names the operator sees side by side in `bridge_profiles` and `bridge_list`.
|
||||
sender. Confusing a profile for a session id is the single easiest mistake to make with `fleet_send`,
|
||||
because both are short names the operator sees side by side in `fleet_profiles` and `fleet_list`.
|
||||
|
||||
**Gotcha.** The check is deliberately narrow: it rejects *only* a value the bridge can prove is a
|
||||
profile. A target missing from the member roster is still accepted, because it may be a peer lead's
|
||||
@@ -903,7 +903,7 @@ CB-561 disabled architect resolution while still compiling and passing its tests
|
||||
|
||||
## See free fleet capacity
|
||||
|
||||
**What.** `bridge_list` returns a top-level `capacity` block — for each profile its `maxLoad`, `live`,
|
||||
**What.** `fleet_list` returns a top-level `capacity` block — for each profile its `maxLoad`, `live`,
|
||||
`free` and `reclaimable` count — and adds `idleForSeconds` and `reclaimable` to every member row.
|
||||
`free` is `null` when the profile is uncapped.
|
||||
|
||||
@@ -921,7 +921,7 @@ capacity facts but no work list, and choosing work needs authority it does not h
|
||||
stays with the lead.
|
||||
|
||||
Two more things to know. `live` is read through the same `liveCountRef` function that placement
|
||||
consumes, so an advertised free slot cannot drift from what `bridge_spawn` will actually accept —
|
||||
consumes, so an advertised free slot cannot drift from what `fleet_spawn` will actually accept —
|
||||
a second count would eventually disagree, and a capacity view that lies is worse than none. And
|
||||
`lifecycle.idleTtlSeconds` defaults to 1800, so a finished member holds its slot for 30 minutes
|
||||
before the reaper takes it. That is far slower than slots turn over during active orchestration,
|
||||
@@ -935,21 +935,21 @@ wall-clock meaning across a restart.
|
||||
|
||||
## Answer a question on an async delegation
|
||||
|
||||
**What.** When a member calls `bridge_ask` during a `wait:false` delegation, `bridge_poll` on that
|
||||
**What.** When a member calls `fleet_ask` during a `wait:false` delegation, `fleet_poll` on that
|
||||
ticket returns a non-terminal `ASKING` phase carrying the question text and its `turnId`. The lead
|
||||
answers with `bridge_send{turnId, content}`; the member resumes the same turn and its real reply
|
||||
answers with `fleet_send{turnId, content}`; the member resumes the same turn and its real reply
|
||||
still arrives on the original ticket.
|
||||
|
||||
**On.** Always on; no configuration.
|
||||
|
||||
**Why.** `bridge_ask` did not work at all on an async delegation. `Outcome.QUESTION` is deliberately
|
||||
**Why.** `fleet_ask` did not work at all on an async delegation. `Outcome.QUESTION` is deliberately
|
||||
non-terminal, but the poll path tested "did this complete?", so a question fell into the failure
|
||||
branch: the ticket was marked `FAILED`, and the question text and `turnId` were both discarded. The
|
||||
member blocked for 55 seconds, gave up, and had to abandon its task and report the ambiguity in its
|
||||
final reply instead. That is the mechanism working backwards — `bridge_ask` exists so a member can
|
||||
final reply instead. That is the mechanism working backwards — `fleet_ask` exists so a member can
|
||||
resolve a decision *without* losing its turn. It also made the guidance self-contradictory: leads are
|
||||
told to prefer `wait:false` for anything non-trivial **and** to answer an ask with
|
||||
`bridge_send{turnId, content}`, and both could not be followed at once.
|
||||
`fleet_send{turnId, content}`, and both could not be followed at once.
|
||||
|
||||
**Gotcha.** The ask still times out after 55 seconds by default (115s maximum). Those caps are not a
|
||||
policy choice and raising them does not help: they exist because the *member's own* MCP client call
|
||||
@@ -986,10 +986,10 @@ was queued but not yet delivered sat until its timeout, which is 30 minutes on a
|
||||
|
||||
**What.** An opt-in background observer. Each tick it takes **one** `AgentControl.list()` for the whole
|
||||
fleet and one in-memory roster snapshot, joins them, and classifies every member. A member entering a
|
||||
fault state logs one WARN; recovering logs one INFO. `bridge_list` reports `healthCoverage`, which is
|
||||
fault state logs one WARN; recovering logs one INFO. `fleet_list` reports `healthCoverage`, which is
|
||||
`off`, `detection-only`, or `full`.
|
||||
|
||||
**On.** A `health:` block in `bridged.yaml` with `enabled: true`. `intervalSeconds` defaults to 30 and
|
||||
**On.** A `health:` block in `fleetd.yaml` with `enabled: true`. `intervalSeconds` defaults to 30 and
|
||||
is floored at 15. With no block at all nothing is constructed and no herdr call is ever made.
|
||||
|
||||
**Why.** A fault is usually a *disagreement between two views*, not a value you can read from one of
|
||||
@@ -1014,7 +1014,7 @@ classifier `false` means "no fault" rather than "not known yet".
|
||||
|
||||
## Keep a worktree that still holds work
|
||||
|
||||
**What.** Before a finished session's git worktree is deleted, bridged checks whether it still holds
|
||||
**What.** Before a finished session's git worktree is deleted, fleetd checks whether it still holds
|
||||
uncommitted changes. If it does, the directory is kept and a WARN names its path, the pane and the
|
||||
release cause. A clean worktree is removed as before.
|
||||
|
||||
@@ -1032,12 +1032,12 @@ file counts as dirty. That is deliberate: the work at risk in the original incid
|
||||
was never `git add`ed, and ignoring untracked files would have missed exactly it. The cost is that a
|
||||
profile whose parity overlay ever copies an untracked, non-gitignored file would make *every* release
|
||||
preserve, and worktrees would pile up silently. Inert today — tracked overlay files carry
|
||||
`--skip-worktree` so `--porcelain` cannot see them, and `bridged.yaml` is gitignored — but it is a real
|
||||
`--skip-worktree` so `--porcelain` cannot see them, and `fleetd.yaml` is gitignored — but it is a real
|
||||
constraint on `overlayParity`, tracked in CB-581.
|
||||
|
||||
Second gotcha: `hasUncommitted` tolerates a worktree that is already gone and reports it clean. It has
|
||||
to. It runs inside `SessionManager.release()` *after* the registry entry is dropped and *before* the
|
||||
pane is stopped, so throwing there would orphan a live pane and strand a `bridge_send` caller on a
|
||||
pane is stopped, so throwing there would orphan a live pane and strand a `fleet_send` caller on a
|
||||
rendezvous nothing resolves. Anything added to that window needs the same tolerance.
|
||||
|
||||
---
|
||||
@@ -1051,7 +1051,7 @@ staying `PENDING` until something else notices.
|
||||
**On.** The same `health:` block that turns on [health watching](#watch-the-fleets-health). No separate
|
||||
key.
|
||||
|
||||
**Why.** Detection without action just moves the silence. A lead that fires `bridge_send{wait:false}`
|
||||
**Why.** Detection without action just moves the silence. A lead that fires `fleet_send{wait:false}`
|
||||
and polls its ticket gets `pending` forever when the member behind it is already gone — the failure is
|
||||
known inside the daemon and invisible to the only caller who cares. Routing it through CB-568's
|
||||
existing idempotent target-wide failure means the outcome is also *counted*, so a dead delegation stops
|
||||
@@ -1068,7 +1068,7 @@ and has its own test — it is a known edge, not an oversight. The retries also
|
||||
|
||||
## Tell a usage-limit refusal from a real reply
|
||||
|
||||
**What.** A member can end its turn without calling `bridge_reply`. The bridge then scrapes the pane
|
||||
**What.** A member can end its turn without calling `fleet_reply`. The bridge then scrapes the pane
|
||||
and hands that text back as the answer. Sometimes that text is not an answer at all — it is the
|
||||
backend refusing, because the account hit its usage limit. With this on, the bridge matches the scrape
|
||||
against a pattern you configure. On a match it resolves the send as `BACKEND_EXHAUSTED` and carries the
|
||||
@@ -1086,7 +1086,7 @@ one vendor.
|
||||
|
||||
**Gotcha.** The pattern map is built once at startup from the config snapshot, so `exhaustedPattern`
|
||||
is a **deferred** key — adding one to a profile does nothing until the daemon restarts.
|
||||
`bridged.example.yaml` does not say this yet. Also, this stage only *classifies*. Nothing yet stops
|
||||
`fleetd.example.yaml` does not say this yet. Also, this stage only *classifies*. Nothing yet stops
|
||||
the fleet spawning another member onto the same exhausted account, and nothing yet saves the work that
|
||||
member was doing — those are stages B and C of CB-578.
|
||||
|
||||
@@ -1098,7 +1098,7 @@ member was doing — those are stages B and C of CB-578.
|
||||
[the entry above](#tell-a-usage-limit-refusal-from-a-real-reply)), the **credential** behind it is put
|
||||
in quarantine for a cooldown. While it is quarantined, an explicit spawn onto it is refused with a
|
||||
message naming the profile, the credential and roughly how many seconds are left; placement skips it
|
||||
under every policy; and `bridge_profiles` shows it. The quarantine lifts itself — there is no manual
|
||||
under every policy; and `fleet_profiles` shows it. The quarantine lifts itself — there is no manual
|
||||
step.
|
||||
|
||||
**On.** `quarantineCooldownSeconds:` at the top level (default 1800, **deferred** — it is baked into
|
||||
@@ -1123,7 +1123,7 @@ protect a sibling.
|
||||
|
||||
**What.** Every member launch records a `CharterReceipt`: the role, where the charter came from, a
|
||||
sha-256 digest of it, and its size in bytes. It is stored on the session and shown in the roster
|
||||
(`bridge_list` and `GET /members`). The charter text itself is never recorded.
|
||||
(`fleet_list` and `GET /members`). The charter text itself is never recorded.
|
||||
|
||||
**On.** Automatic.
|
||||
|
||||
@@ -1144,10 +1144,10 @@ a `PeerHandle` implementation, this is the line that will not let you skip it.
|
||||
|
||||
## Redeploy the daemon safely
|
||||
|
||||
**What.** `scripts/redeploy-bridged.sh` rebuilds the jar and restarts `bridged` as one command. It
|
||||
**What.** `scripts/redeploy-fleetd.sh` rebuilds the jar and restarts `fleetd` as one command. It
|
||||
builds *before* it stops anything, waits for the old process to actually exit, restarts from a login
|
||||
shell with `cwd = bridged/`, then polls `/healthz` and reports the herdr protocol number, a fresh
|
||||
`bridged listening` line, the config keys accepted or deferred at boot, and any `ERROR` lines since
|
||||
shell with `cwd = fleetd/`, then polls `/healthz` and reports the herdr protocol number, a fresh
|
||||
`fleetd listening` line, the config keys accepted or deferred at boot, and any `ERROR` lines since
|
||||
the restart. `--check` reports state and changes nothing; `--yes` skips the drain prompt;
|
||||
`--no-build` restarts the jar already on disk.
|
||||
|
||||
@@ -1159,7 +1159,7 @@ whole `health:` stack sat merged and doing nothing for two tickets. The script a
|
||||
operator can allow-list **one** auditable command instead of approving a `kill` and a `java -jar`
|
||||
separately every time, which is what a lead would otherwise have to ask for on every deploy.
|
||||
|
||||
**Gotcha.** The check that matters most has no log line anywhere in the daemon: `bridged` inherits
|
||||
**Gotcha.** The check that matters most has no log line anywhere in the daemon: `fleetd` inherits
|
||||
`WORKER_GITEA_TOKEN` from the shell that starts it, via `${SHARED_ENV}/tools/secrets.sh`. Start it
|
||||
from a non-login shell and the variable is empty — the daemon boots normally, `/healthz` is green,
|
||||
and the failure surfaces much later as workers that cannot open a PR. `--check` is the only thing
|
||||
@@ -1177,7 +1177,7 @@ you rather than forcing it.
|
||||
|
||||
## Get told when an async delegation finishes
|
||||
|
||||
**What.** A `bridge_send{wait:false}` ticket that reaches a terminal phase — replied, failed, wedged,
|
||||
**What.** A `fleet_send{wait:false}` ticket that reaches a terminal phase — replied, failed, wedged,
|
||||
or abandoned — now injects a short nudge into the **lead's own pane** telling it to poll. Several
|
||||
tickets finishing at once coalesce into one nudge naming the count. Polling a ticket marks it
|
||||
collected, so a ticket you already read is never nudged about again.
|
||||
@@ -1193,7 +1193,7 @@ budget a ticket needs. Worst case a lead sees up to twice `push_reminders` nudge
|
||||
deliberate price of that isolation.
|
||||
|
||||
**Why.** The charter tells leads to prefer `wait:false` for anything non-trivial, because a blocking
|
||||
`bridge_send` is capped by the caller's own MCP client timeout of about 60 seconds. But until this
|
||||
`fleet_send` is capped by the caller's own MCP client timeout of about 60 seconds. But until this
|
||||
landed, that preferred mode was the one mode with **no notification at all**: `MessageService.reply`
|
||||
returns on the rendezvous fast path before the push loop hears anything, so an async ticket finished
|
||||
in silence and the lead only found out by polling on a hunch.
|
||||
@@ -1210,27 +1210,27 @@ durable artefact; the PR is.** If you orchestrate over REST, poll on a timer.
|
||||
|
||||
## Get told when a worker is waiting on your answer
|
||||
|
||||
**What.** A worker on an async (`wait:false`) delegation that pauses mid-turn in `bridge_ask` now
|
||||
**What.** A worker on an async (`wait:false`) delegation that pauses mid-turn in `fleet_ask` now
|
||||
nudges the lead's pane by itself, naming the exact call that resumes it:
|
||||
|
||||
```
|
||||
Worker term_a asked a question (ticket task-3) — answer it with
|
||||
bridge_send(turnId="term_a#1", content=...) to resume its turn:
|
||||
fleet_send(turnId="term_a#1", content=...) to resume its turn:
|
||||
which config file?
|
||||
```
|
||||
|
||||
`bridge_status{sessionId}` shows the same open question, and so does REST
|
||||
`fleet_status{sessionId}` shows the same open question, and so does REST
|
||||
`GET /sessions/{id}/status` (as `question`, `turnId`, `ticket`).
|
||||
|
||||
**On.** Automatic, on the same terms as the ticket nudge above — it is a **third source in the same
|
||||
per-lead schedule**, not a new push path, so CB-590's one-schedule-per-lead guarantee still holds and
|
||||
it spends from its own `push_reminders` budget.
|
||||
|
||||
**Why.** `bridge_ask` opens a reverse-rendezvous window of about **55 seconds**. A lead polling on its
|
||||
**Why.** `fleet_ask` opens a reverse-rendezvous window of about **55 seconds**. A lead polling on its
|
||||
normal cadence of minutes never saw it, so the worker timed out and carried on without an answer —
|
||||
the ask was, in practice, unusable on the delegation mode the charter tells leads to prefer. Two
|
||||
smaller holes closed with it: REST `GET /tasks/{ticket}` dropped `turnId` on an `ASKING` phase, so a
|
||||
REST caller could read the question and had no way to answer it, and `bridge_status` said nothing
|
||||
REST caller could read the question and had no way to answer it, and `fleet_status` said nothing
|
||||
about an open question at all.
|
||||
|
||||
**Gotcha.** **This closes the window; it does not remove it.** The worker still gets ~55 seconds, and
|
||||
@@ -1273,22 +1273,22 @@ difference beforehand. **Present-and-useless looks identical to absent.**
|
||||
|
||||
## Supervise the daemon without breaking the fleet
|
||||
|
||||
**What.** A launchd unit that restarts `bridged` if it dies, and that still gets the fleet's secrets.
|
||||
`deploy/dev.ltms.bridged.plist` runs `scripts/bridged-launchd-wrapper.sh`, which execs one login shell
|
||||
**What.** A launchd unit that restarts `fleetd` if it dies, and that still gets the fleet's secrets.
|
||||
`deploy/dev.ltms.fleetd.plist` runs `scripts/fleetd-launchd-wrapper.sh`, which execs one login shell
|
||||
in place (`exec /bin/zsh -lc 'exec "$@"' -- "$@"`) and then execs the real java command. One `exec`
|
||||
chain, so launchd keeps tracking the right PID. `scripts/redeploy-bridged.sh` detects whether the
|
||||
chain, so launchd keeps tracking the right PID. `scripts/redeploy-fleetd.sh` detects whether the
|
||||
agent is loaded and switches stop/start to `launchctl unload -w` / `load -w`, falling back to its
|
||||
original kill + `nohup` when it is not.
|
||||
|
||||
**On.** Not automatic, and deliberately so. Install it yourself:
|
||||
|
||||
```bash
|
||||
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
|
||||
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
|
||||
launchctl list | grep bridged
|
||||
cp deploy/dev.ltms.fleetd.plist ~/Library/LaunchAgents/
|
||||
launchctl load -w ~/Library/LaunchAgents/dev.ltms.fleetd.plist
|
||||
launchctl list | grep fleetd
|
||||
```
|
||||
|
||||
`scripts/redeploy-bridged.sh --check` reports whether the agent is installed and whether it is loaded.
|
||||
`scripts/redeploy-fleetd.sh --check` reports whether the agent is installed and whether it is loaded.
|
||||
It is read-only.
|
||||
|
||||
**Why.** The unit shipped with CB-504 was never installed, and could not have worked if it were.
|
||||
@@ -1311,12 +1311,12 @@ restarts to one per ten seconds, it does not cap how many. Fix **CB-600** before
|
||||
|
||||
## See at startup which secrets the daemon actually got
|
||||
|
||||
**What.** `bridged` logs, at startup, every secret environment variable it needs, and whether each one
|
||||
**What.** `fleetd` logs, at startup, every secret environment variable it needs, and whether each one
|
||||
resolved or is `MISSING`. The required set is derived from the loaded config — each non-subscription
|
||||
profile's `tokenEnv`, plus every profile's `gitTokenEnv` — not hard-coded, so a new profile is covered
|
||||
the day it is added.
|
||||
|
||||
**On.** Automatic. Read the `startup secret …` lines at the top of `bridged/bridged.out`.
|
||||
**On.** Automatic. Read the `startup secret …` lines at the top of `fleetd/fleetd.out`.
|
||||
|
||||
**Why.** An empty token used to be completely invisible. The daemon started, `/healthz` went green,
|
||||
and the first sign of trouble came much later and somewhere else — a worker that could not open a PR,
|
||||
@@ -1328,7 +1328,7 @@ one line at startup.
|
||||
deliberate and must stay that way; the log is not a secret store. Also, `MISSING` is a warning, not a
|
||||
refusal: the daemon starts anyway, because refusing to boot over a credential that half the fleet may
|
||||
not need would be worse. So the line has to actually be read. One known false alarm: a non-subscription
|
||||
profile that never sets `tokenEnv` inherits the default name `BRIDGED_WORKER_TOKEN` and is reported
|
||||
profile that never sets `tokenEnv` inherits the default name `FLEETD_WORKER_TOKEN` and is reported
|
||||
missing — which is nearly always a real misconfiguration rather than a bug in the report.
|
||||
|
||||
---
|
||||
@@ -1385,7 +1385,7 @@ and every orchestration call is refused — a failure that looks nothing like it
|
||||
startup is the kinder outcome.
|
||||
|
||||
**Gotcha.** `weight: 0` excludes a profile from *automatic* selection only. An explicit
|
||||
`bridge_spawn{profile:"..."}` bypasses placement entirely and still resolves it, so `weight: 0` is
|
||||
`fleet_spawn{profile:"..."}` bypasses placement entirely and still resolves it, so `weight: 0` is
|
||||
**not** a way to disable a profile. `maxLoad: 0` is, because that cap now applies to explicit spawns
|
||||
too. And a lead's `tab` must match a real herdr tab label exactly (case-insensitively) — the old
|
||||
`tabPrefix` no longer finds a lead, it only guards against a worker's label colliding with the
|
||||
@@ -1395,7 +1395,7 @@ convention.
|
||||
|
||||
## Stop an explicit spawn from busting the cap
|
||||
|
||||
**What.** `maxLoad` is now an unconditional cap. An explicit `bridge_spawn{profile:"X"}` used to skip
|
||||
**What.** `maxLoad` is now an unconditional cap. An explicit `fleet_spawn{profile:"X"}` used to skip
|
||||
the check entirely — the cap applied only to automatic placement — so naming a profile was a way
|
||||
around it. Now an explicit spawn is refused when `live >= cap`, and it is **not** re-routed to
|
||||
another profile.
|
||||
@@ -1453,13 +1453,13 @@ question with a digest instead.
|
||||
|
||||
**Gotcha.** The role charter is combined with the reply charter **only when the profile mounts the
|
||||
bridge MCP**. A non-MCP profile gets the role charter alone — which is correct, since the reply
|
||||
charter tells a member to call `bridge_reply`, and a member with no bridge cannot.
|
||||
charter tells a member to call `fleet_reply`, and a member with no bridge cannot.
|
||||
|
||||
---
|
||||
|
||||
## Tell "busy" from "refusing" in the capacity view
|
||||
|
||||
**What.** In `bridge_list`'s `capacity` view, a quarantined profile's row is forced to `free: 0`
|
||||
**What.** In `fleet_list`'s `capacity` view, a quarantined profile's row is forced to `free: 0`
|
||||
whatever its `maxLoad` and `live` counts say, and gains `credentialId` and `quarantinedForSeconds`.
|
||||
|
||||
**On.** No new knob. The quarantine facts come from the existing `exhaustedPattern`, `credentialId`
|
||||
@@ -1472,17 +1472,17 @@ difference.
|
||||
|
||||
**Gotcha.** The two new fields appear **only** when a profile is actually quarantined, so an ordinary
|
||||
fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares
|
||||
its `QuarantineSource` with `bridge_profiles` so the two surfaces cannot disagree.
|
||||
its `QuarantineSource` with `fleet_profiles` so the two surfaces cannot disagree.
|
||||
|
||||
---
|
||||
|
||||
## Resume a member onto its previous conversation
|
||||
|
||||
**What.** A member's own agent-session id is captured at spawn, survives onto the roster as
|
||||
`agentSessionId` in `bridge_list` and `GET /members`, and `bridge_spawn` accepts `sessionName` and
|
||||
`agentSessionId` in `fleet_list` and `GET /members`, and `fleet_spawn` accepts `sessionName` and
|
||||
`resumeSessionId` to relaunch onto that same conversation.
|
||||
|
||||
**On.** `bridge_spawn{sessionName, resumeSessionId}`.
|
||||
**On.** `fleet_spawn{sessionName, resumeSessionId}`.
|
||||
|
||||
**Why.** The pieces existed but were connected at neither end: nothing persisted the id, and nothing
|
||||
exposed a way to pass it back. So a member that died took its context with it even though the backend
|
||||
@@ -1505,13 +1505,13 @@ not explicitly deny, and the bridge-generated config pins `compaction.auto: true
|
||||
|
||||
**Why.** A worker that stops to ask for permission on a tool call has no one to ask — the lead is not
|
||||
watching its pane, and an unanswered prompt burns the whole turn. Auto-compaction is the same
|
||||
argument: a worker that runs out of context dies mid-turn and loses its `bridge_reply`, so its report
|
||||
argument: a worker that runs out of context dies mid-turn and loses its `fleet_reply`, so its report
|
||||
is gone even though the work was done.
|
||||
|
||||
**Gotcha, and it is a real trade.** opencode's own help calls `--auto` "dangerous!". The blast radius
|
||||
is bounded by everything else about a member — its own worktree, its own branch, off-subscription,
|
||||
and it cannot merge — not by the flag. And `compaction.auto: true` in the generated config **wins
|
||||
over the operator's home config**, so auto-compaction cannot be turned off for bridged workers from
|
||||
over the operator's home config**, so auto-compaction cannot be turned off for fleetd workers from
|
||||
there. That was chosen knowingly: a lost report was judged worse than an unwanted compaction.
|
||||
|
||||
---
|
||||
@@ -1558,7 +1558,7 @@ alongside the worktree, branch and snapshot ref it already reported:
|
||||
… worktree=/wt/x branch=worker/x snapshot=refs/wip/x agentSessionId=abc-123
|
||||
```
|
||||
|
||||
A lead can pass that id back as `bridge_spawn{resumeSessionId: "abc-123"}` on the same profile.
|
||||
A lead can pass that id back as `fleet_spawn{resumeSessionId: "abc-123"}` on the same profile.
|
||||
|
||||
**On.** Automatic, no knob. It appears whenever the released member's backend produced a session id.
|
||||
|
||||
@@ -1581,8 +1581,8 @@ half — it turns *"the work survives"* into *"the work and the thread both surv
|
||||
against `terra: 2` and `sonnet: 1` does not mean "use local, overflow to paid" — it means roughly a
|
||||
quarter of spawns go to a paid profile **while the free box still has a free slot**.
|
||||
|
||||
**On.** `placement: weighted` in `bridged.yaml`, with a `weight:` per profile. `weight: 0` excludes a
|
||||
profile from automatic placement entirely; an explicit `bridge_spawn{profile:...}` bypasses placement
|
||||
**On.** `placement: weighted` in `fleetd.yaml`, with a `weight:` per profile. `weight: 0` excludes a
|
||||
profile from automatic placement entirely; an explicit `fleet_spawn{profile:...}` bypasses placement
|
||||
either way.
|
||||
|
||||
**Why.** The operator's rule is cheapest-first: keep the free boxes busy and pay only for genuine
|
||||
@@ -1597,9 +1597,9 @@ profile sits at `maxLoad` it is filtered out and its score **freezes**, so paid
|
||||
accumulating against it; when the free slot opens the profile returns with a stale score and can
|
||||
*lose* the next pick — a paid spawn while the free box is idle. Second, and worse for the next
|
||||
person: the workaround expresses a **preference order** through a **ratio** knob. Add a profile at
|
||||
weight 150 later and it silently outranks the free box, with nothing to warn you. `bridged.yaml` is
|
||||
weight 150 later and it silently outranks the free box, with nothing to warn you. `fleetd.yaml` is
|
||||
gitignored, so a fresh host starts without this workaround and quietly pays — the reasoning is
|
||||
written into `bridged.example.yaml` next to the key for exactly that reason.
|
||||
written into `fleetd.example.yaml` next to the key for exactly that reason.
|
||||
|
||||
---
|
||||
|
||||
@@ -1633,9 +1633,9 @@ scope. Read this table as the two that were decided **first**, not as the whole
|
||||
|
||||
## A member keeps only the credentials you name
|
||||
|
||||
**What.** A `memberCredentials:` block in `bridged.yaml` lists every credential-shaped variable on
|
||||
**What.** A `memberCredentials:` block in `fleetd.yaml` lists every credential-shaped variable on
|
||||
the host, says which ones a member may keep, and blocks the rest. A blocked name is not unset — it is
|
||||
overwritten with a fixed sentinel string, `blocked-by-bridged-cb596-see-gitea-issue-82`, so a member
|
||||
overwritten with a fixed sentinel string, `blocked-by-fleetd-cb596-see-gitea-issue-82`, so a member
|
||||
that reads it sees *"deliberately blocked"* rather than an empty variable it might quietly work
|
||||
around. Before this, a member pane inherited the operator's **whole** secret store and exactly one
|
||||
name was blocked.
|
||||
@@ -1655,7 +1655,7 @@ memberCredentials:
|
||||
The blocked set is `known` minus `allow`, computed at load. A name in both is an error you cannot
|
||||
make by accident — the intersection is empty by construction, because `allow` wins.
|
||||
|
||||
**On.** Add the block to `bridged.yaml` (there is a commented template in `bridged.example.yaml`) and
|
||||
**On.** Add the block to `fleetd.yaml` (there is a commented template in `fleetd.example.yaml`) and
|
||||
add the matching guarded export to the operator's secret store. **Both halves are needed** — see the
|
||||
gotcha. The block is re-read on every spawn, so editing it takes effect without a restart; only the
|
||||
startup summary line needs one.
|
||||
@@ -1771,32 +1771,32 @@ adding to `known`.
|
||||
only whether herdr answered: `200 {"status":"ok","herdr":{"version":…,"protocol":…}}`, or
|
||||
`503 {"status":"degraded","herdr":"unreachable",…}` on any `HerdrException`. `GET /metrics`
|
||||
renders every registered series as Prometheus text (CB-502): counters
|
||||
`bridged_sends_total`, `bridged_replies_total`, `bridged_push_nudges_total`,
|
||||
`bridged_lead_heartbeat_nudges_total`, `bridged_spawns_total`, `bridged_herdr_calls_total`,
|
||||
`bridged_auth_failures_total`, plus gauges `bridged_sessions` (one series per lifecycle state)
|
||||
and `bridged_inbox_depth` (one series per undrained target).
|
||||
`fleetd_sends_total`, `fleetd_replies_total`, `fleetd_push_nudges_total`,
|
||||
`fleetd_lead_heartbeat_nudges_total`, `fleetd_spawns_total`, `fleetd_herdr_calls_total`,
|
||||
`fleetd_auth_failures_total`, plus gauges `fleetd_sessions` (one series per lifecycle state)
|
||||
and `fleetd_inbox_depth` (one series per undrained target).
|
||||
|
||||
**On.** Both are always registered. `/healthz` carries no authorization check at all — it is
|
||||
reachable by an unauthenticated caller by design. `/metrics` is gated on `Authz.Action.METRICS`:
|
||||
open to `PRIMARY`/`WORKER`/`ARCHITECT`, refused to `ANONYMOUS`.
|
||||
|
||||
**Why.** `BridgedMetrics`'s own class doc states the design intent directly: "each series maps
|
||||
**Why.** `FleetdMetrics`'s own class doc states the design intent directly: "each series maps
|
||||
to a failure mode this project has actually hit, not to whatever was easy to count," and names
|
||||
the two worth watching — a rising `bridged_sends_total{outcome="completion_fallback"}` share
|
||||
(turn detection degrading) and `bridged_push_nudges_total{outcome="exhausted"}` (the primary
|
||||
the two worth watching — a rising `fleetd_sends_total{outcome="completion_fallback"}` share
|
||||
(turn detection degrading) and `fleetd_push_nudges_total{outcome="exhausted"}` (the primary
|
||||
stopped draining its inbox). `/healthz`'s narrow scope traces to CB-504: under supervision the
|
||||
daemon must serve before herdr's socket even exists, so "degraded but alive" needed one cheap,
|
||||
reliable signal.
|
||||
|
||||
**Gotcha.** Neither endpoint proves the fleet actually works. `/healthz` echoes back whatever
|
||||
`protocol` number herdr reports, but nothing in the codebase compares that number against what
|
||||
bridged's own herdr calls need — and the two have already drifted apart in the source itself:
|
||||
fleetd's own herdr calls need — and the two have already drifted apart in the source itself:
|
||||
`AgentControl`'s class doc says it was "ported to herdr protocol 19 (herdr 0.8.0, CB-521)," while
|
||||
`HerdrClient`'s class doc still says "protocol 14, herdr 0.7.0." This is the exact CB-521
|
||||
incident: herdr answers `ping` correctly and `/healthz` goes green, while `agent.start` and the
|
||||
rest of the protocol-19 surface fail because the adapter and the herdr binary disagree on
|
||||
protocol version. `/metrics` has no counter or gauge for that mismatch either — a resulting spawn
|
||||
failure only shows up as a `bridged_spawns_total{outcome=…}` tick, and only once something
|
||||
failure only shows up as a `fleetd_spawns_total{outcome=…}` tick, and only once something
|
||||
actually tries to spawn.
|
||||
|
||||
---
|
||||
@@ -1812,7 +1812,7 @@ Separately, the daemon refuses to start at all when `bind.host` is non-loopback
|
||||
is still `loopback-trust`.
|
||||
|
||||
**On.** `auth.mode: loopback-trust | token` (default `loopback-trust`); `auth.tokenEnv` names the
|
||||
host env var holding the token (default `BRIDGED_API_TOKEN`, read only under `token` mode). The
|
||||
host env var holding the token (default `FLEETD_API_TOKEN`, read only under `token` mode). The
|
||||
bind fail-fast has no separate switch — it always runs in `main()`.
|
||||
|
||||
**Why.** Stated directly in the code: loopback-trust's safety depends entirely on the OS refusing
|
||||
@@ -1837,13 +1837,13 @@ table stating, for each of eight actions (`SPAWN`, `STOP`, `SEND`, `REPLY`, `ASK
|
||||
`READ`, `METRICS`), which role may call it: `SPAWN`/`STOP`/`DRAIN` are the primary alone; `SEND`
|
||||
is primary or architect; `REPLY`/`ASK` require the caller to own the target session (its own
|
||||
pane, checked structurally, never by argument); `READ`/`METRICS` are open to any authenticated
|
||||
role. It is the single gate behind both entry paths (REST and MCP) — `BridgedApp.allow()` and
|
||||
role. It is the single gate behind both entry paths (REST and MCP) — `FleetdApp.allow()` and
|
||||
`BridgeMcp`'s own check both call into it, so the rule can't drift between the two surfaces.
|
||||
Refusals are recorded by `AuditLog`, an append-only JSON-lines trail written by a dedicated
|
||||
`audit` logger to `logs/audit.log` (daily rolling, 30-day retention, 100MB cap), independent of
|
||||
the daemon's normal app log.
|
||||
|
||||
**On.** Always on; not configurable. Every request through `BridgedApp` or `BridgeMcp` passes
|
||||
**On.** Always on; not configurable. Every request through `FleetdApp` or `BridgeMcp` passes
|
||||
through `Authz.permits()`.
|
||||
|
||||
**Why.** The class doc states this plainly: most of the rule was already true de facto — a
|
||||
@@ -1907,15 +1907,15 @@ Checking for the same shape elsewhere found three more fields, all fixed the sam
|
||||
|
||||
## Supervise the daemon on Linux (systemd)
|
||||
|
||||
**What.** `deploy/bridged.service` is a systemd **user** unit (not system-level — "bridged drives
|
||||
the user's herdr, not a system daemon") that runs `java -jar target/bridged.jar bridged.yaml`,
|
||||
**What.** `deploy/fleetd.service` is a systemd **user** unit (not system-level — "fleetd drives
|
||||
the user's herdr, not a system daemon") that runs `java -jar target/fleetd.jar fleetd.yaml`,
|
||||
restarts on failure (`Restart=on-failure`, `RestartSec=10s`, capped at 5 restarts per 120s via
|
||||
`StartLimitBurst`/`StartLimitIntervalSec`), waits on `herdr.service` only advisorially
|
||||
(`Wants=`, not `Requires=`, so a herdr restart never takes bridged down with it), and applies a
|
||||
(`Wants=`, not `Requires=`, so a herdr restart never takes fleetd down with it), and applies a
|
||||
sandboxing profile (`NoNewPrivileges`, `ProtectSystem=strict`, `ProtectHome=read-write`, etc.).
|
||||
|
||||
**On.** Manual install: copy to `~/.config/systemd/user/`, edit `ExecStart`/`WorkingDirectory`/
|
||||
`Environment`, then `systemctl --user daemon-reload && systemctl --user enable --now bridged`. Not
|
||||
`Environment`, then `systemctl --user daemon-reload && systemctl --user enable --now fleetd`. Not
|
||||
currently the live supervision target — the unit file's own header comment says the dogfooded
|
||||
daemon runs on macOS under launchd; this unit is for the Linux gateways CB-308 introduces.
|
||||
|
||||
@@ -1926,21 +1926,21 @@ elsewhere in the wiki) needs Linux hosts, and those need systemd rather than lau
|
||||
"systemd does not source a login shell, so without it the daemon — and every worker — gets a bare
|
||||
default with no JDK/Maven" (same defect class CB-511/CB-594 already fixed for `PATH` on
|
||||
launchd). For secrets, the unit's own comment says plainly "Secrets are NOT set here" and points
|
||||
the operator at a `systemctl --user edit bridged` drop-in or an `EnvironmentFile=`. **It answers
|
||||
the operator at a `systemctl --user edit fleetd` drop-in or an `EnvironmentFile=`. **It answers
|
||||
the login-shell defect only for `PATH`, not for `WORKER_GITEA_TOKEN`/`AI_GATEWAY_TOKEN`.**
|
||||
|
||||
**Systemd answer — YES, it has the underlying defect, undocumented for those two variables
|
||||
specifically.** Comparing to the launchd side: launchd had the identical problem
|
||||
(`deploy/dev.ltms.bridged.plist`'s own comment: "launchd does NOT source .zprofile/.zshrc") and it
|
||||
was fixed by `scripts/bridged-launchd-wrapper.sh`, which execs `zsh -l` so
|
||||
(`deploy/dev.ltms.fleetd.plist`'s own comment: "launchd does NOT source .zprofile/.zshrc") and it
|
||||
was fixed by `scripts/fleetd-launchd-wrapper.sh`, which execs `zsh -l` so
|
||||
`${SHARED_ENV}/tools/secrets.sh` gets sourced — its own header comment names exactly
|
||||
`WORKER_GITEA_TOKEN` and `AI_GATEWAY_TOKEN` as the two secrets this closes the gap for. The
|
||||
systemd unit has **no equivalent wrapper** and does not source that file at all. Its own comment
|
||||
mentions only `BRIDGED_API_TOKEN` as the secret to add via drop-in — it never names
|
||||
mentions only `FLEETD_API_TOKEN` as the secret to add via drop-in — it never names
|
||||
`WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following the unit file's own guidance
|
||||
verbatim would set `BRIDGED_API_TOKEN` and stop there: the daemon boots fine, and the failure
|
||||
verbatim would set `FLEETD_API_TOKEN` and stop there: the daemon boots fine, and the failure
|
||||
surfaces only later as a worker that cannot open a PR or a gateway profile returning 401 — the
|
||||
exact failure mode CB-594 fixed for launchd, left open here. `Bridged.reportRequiredSecrets`
|
||||
exact failure mode CB-594 fixed for launchd, left open here. `Fleetd.reportRequiredSecrets`
|
||||
(called at the top of `main()`) *does* log which secret env-var names resolved on either
|
||||
platform — but that log line can only tell the truth about the names it names; it doesn't fix
|
||||
the sourcing gap, and the unit file gives the operator no prompt to look for it.
|
||||
@@ -2055,7 +2055,7 @@ launch. For an `opencode` member it becomes a per-model `provider.<p>.models.<m>
|
||||
the generated config. Unset ⇒ the backend's own default (Claude Code's built-in auto-compact,
|
||||
opencode's `compaction.auto`).
|
||||
|
||||
**On.** Per profile in `bridged.yaml`, e.g. `autoCompactWindow: 250000`. The value must be in
|
||||
**On.** Per profile in `fleetd.yaml`, e.g. `autoCompactWindow: 250000`. The value must be in
|
||||
`[100000, 1000000]` — the band Claude Code's flag accepts — and the daemon refuses to start with a
|
||||
profile outside it, naming the profile. For opencode the key only takes effect when `model:` is in
|
||||
`provider/model` form (a warning is logged otherwise), because the limit is written under that exact
|
||||
@@ -2136,7 +2136,7 @@ Still to catalogue:
|
||||
repo, so it arrives with its own contract. At spawn the launcher passes `--agent <role>` when the
|
||||
matching file exists.
|
||||
|
||||
**The knob.** Nothing to turn on. The role you pass to `bridge_spawn` selects the file. Absent file ⇒
|
||||
**The knob.** Nothing to turn on. The role you pass to `fleet_spawn` selects the file. Absent file ⇒
|
||||
no `--agent` flag and the member still spawns, so this degrades rather than breaks.
|
||||
|
||||
**Why it exists.** The contract used to travel as an inline argv flag,
|
||||
@@ -2152,15 +2152,15 @@ its charter to a file. Files fix that for both backends and make the contract re
|
||||
`opencode agent list`. The two copies of a role must carry identical bodies or the backends work
|
||||
from different contracts.
|
||||
- **Never put `model:` in an agent file.** Both CLIs honour it *only* when no launch flag is passed,
|
||||
and bridged always passes one (`--model` for claude-code, `-m` for opencode). Measured both ways:
|
||||
and fleetd always passes one (`--model` for claude-code, `-m` for opencode). Measured both ways:
|
||||
an agent pinned to `openai/gpt-5.6-terra` run with `-m opencode/x-preview-f-free` reported
|
||||
`> pin · x-preview-f-free`. A `model:` here is silently overridden on every spawn. The model stays
|
||||
in `bridged.yaml`, which also keeps role and backend as the separate axes the role pools need.
|
||||
in `fleetd.yaml`, which also keeps role and backend as the separate axes the role pools need.
|
||||
- **Both charters travel in one file, or Claude Code will not start** (CB-618). The CLI refuses the
|
||||
two prompt flags together: `Error: Cannot use both --append-system-prompt and
|
||||
--append-system-prompt-file. Please use only one.` The first cut of this feature put the role
|
||||
charter on the file flag and left the reply charter inline, and every claude-code spawn with a role
|
||||
charter died at launch — reported by bridged as `spawn_timeout`, which hides the cause completely.
|
||||
charter died at launch — reported by fleetd as `spawn_timeout`, which hides the cause completely.
|
||||
So when a role charter is present, both go in the one file with the reply charter **last**: last is
|
||||
where the reply rule must sit, because it is the rule that must survive. A member with no role
|
||||
charter keeps the inline flag, which is also the only form that reaches a member with no repo
|
||||
@@ -2171,7 +2171,7 @@ its charter to a file. Files fix that for both backends and make the contract re
|
||||
such key, so the two formats are not interchangeable even though the bodies are identical.
|
||||
- **Measured together, on the real binary.** `claude --agent architect --append-system-prompt-file
|
||||
<both charters>` answered `ROLEOK, REPLYOK, yes` — agent body, role charter and reply charter all
|
||||
in force at once. Both defects above shipped green because every test read the argv bridged builds
|
||||
in force at once. Both defects above shipped green because every test read the argv fleetd builds
|
||||
and none ran the binary that has to accept it (see #113).
|
||||
|
||||
### Found while cataloguing, not by looking for bugs
|
||||
@@ -2183,15 +2183,15 @@ Two of these are filed as their own tickets. They are recorded here because both
|
||||
matching `"opencode"` — routed to the **claude-code** adapter. If `argv:` is unset, the launch
|
||||
command defaults to the misspelled string itself rather than `claude`. No config-load check catches
|
||||
it. See *Multi-profile routing* above.
|
||||
- **The systemd unit has the login-shell secret defect that launchd's had.** `deploy/bridged.service`
|
||||
fixes `PATH` explicitly and names only `BRIDGED_API_TOKEN` as a secret to add. It never sources the
|
||||
- **The systemd unit has the login-shell secret defect that launchd's had.** `deploy/fleetd.service`
|
||||
fixes `PATH` explicitly and names only `FLEETD_API_TOKEN` as a secret to add. It never sources the
|
||||
secret store, and never mentions `WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following
|
||||
the unit's own guidance gets a daemon that boots cleanly and members that cannot open a PR. launchd
|
||||
got `scripts/bridged-launchd-wrapper.sh` for exactly this; the Linux unit has no equivalent.
|
||||
got `scripts/fleetd-launchd-wrapper.sh` for exactly this; the Linux unit has no equivalent.
|
||||
|
||||
One more piece of drift worth knowing while reading these entries: `AgentControl`'s class doc says
|
||||
herdr protocol 19 (herdr 0.8.0), while `HerdrClient`'s still says protocol 14 (0.7.0). Nothing
|
||||
compares the protocol number herdr reports against what bridged actually needs — which is precisely
|
||||
compares the protocol number herdr reports against what fleetd actually needs — which is precisely
|
||||
how `/healthz` once went green while every spawn failed.
|
||||
|
||||
### A note for anyone briefing a worker to read this page
|
||||
|
||||
+1
-1
@@ -64,7 +64,7 @@ belong in the global file and shared ones in the project file.
|
||||
"$schema": "https://opencode.ai/config.json",
|
||||
"instructions": ["CLAUDE.md"],
|
||||
"mcp": {
|
||||
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true }
|
||||
"fleetd": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
+61
-61
@@ -11,15 +11,15 @@
|
||||
|
||||
## 1. What this is, and what it is not
|
||||
|
||||
`bridged` is a **message bus between AI agent sessions**. One session orchestrates (the **lead**),
|
||||
`fleetd` is a **message bus between AI agent sessions**. One session orchestrates (the **lead**),
|
||||
and it delegates work to **members** running in other terminal panes, often on other models and
|
||||
other vendors. All traffic goes through `bridged`'s MCP tools. No session talks to another session,
|
||||
other vendors. All traffic goes through `fleetd`'s MCP tools. No session talks to another session,
|
||||
to a broker, or to the network directly.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
LEAD["lead session<br/>(Claude Code, subscription)"]
|
||||
subgraph BD["bridged — a plain Java daemon"]
|
||||
subgraph BD["fleetd — a plain Java daemon"]
|
||||
SRV["MCP + REST<br/>policy, authz, lifecycle"]
|
||||
INJ["injector<br/>status-gated"]
|
||||
SRV --> INJ
|
||||
@@ -29,9 +29,9 @@ flowchart LR
|
||||
M2["member pane<br/>opencode"]
|
||||
GW["llm.ltms.dev<br/>the one gateway"]
|
||||
|
||||
LEAD -->|"bridge_send"| SRV
|
||||
M1 -.->|"bridge_reply"| SRV
|
||||
M2 -.->|"bridge_reply"| SRV
|
||||
LEAD -->|"fleet_send"| SRV
|
||||
M1 -.->|"fleet_reply"| SRV
|
||||
M2 -.->|"fleet_reply"| SRV
|
||||
INJ -->|"unix socket"| HERDR
|
||||
HERDR --> M1
|
||||
HERDR --> M2
|
||||
@@ -61,7 +61,7 @@ flowchart LR
|
||||
1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the
|
||||
subscription. Only the daemon moves a member off it, at spawn.
|
||||
2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
|
||||
in a `bridge_*` call is discarded silently.
|
||||
in a `fleet_*` call is discarded silently.
|
||||
|
||||
---
|
||||
|
||||
@@ -71,7 +71,7 @@ Four things must be true before anything works. In order, because each one depen
|
||||
|
||||
### 2.1 herdr
|
||||
|
||||
`bridged` does not own terminals. `herdr` does. `bridged` drives it over a Unix socket.
|
||||
`fleetd` does not own terminals. `herdr` does. `fleetd` drives it over a Unix socket.
|
||||
|
||||
```bash
|
||||
herdr --version # 0.8.0 on this host
|
||||
@@ -79,7 +79,7 @@ curl -s http://127.0.0.1:8765/healthz
|
||||
# {"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
|
||||
```
|
||||
|
||||
**The protocol number is the thing to check, not the version.** `bridged`'s adapter speaks one
|
||||
**The protocol number is the thing to check, not the version.** `fleetd`'s adapter speaks one
|
||||
herdr wire protocol. If herdr is upgraded and the protocol moves, `/healthz` still says `ok` —
|
||||
because the socket connects — and **every spawn fails**. Green health with a broken fleet is the
|
||||
normal way this breaks. See trap 2 in §6.
|
||||
@@ -89,8 +89,8 @@ normal way this breaks. See trap 2 in §6.
|
||||
Build and run from this repo. There is no POM at the repo root:
|
||||
|
||||
```bash
|
||||
mvn -f bridged/pom.xml clean install # never pipe this — a pipe hides BUILD FAILURE
|
||||
java -jar bridged/target/bridged.jar bridged/bridged.yaml
|
||||
mvn -f fleetd/pom.xml clean install # never pipe this — a pipe hides BUILD FAILURE
|
||||
java -jar fleetd/target/fleetd.jar fleetd/fleetd.yaml
|
||||
```
|
||||
|
||||
In practice you never run those two by hand. Use the script (§4).
|
||||
@@ -107,7 +107,7 @@ shell**. This one fact causes more lost hours than anything else in the system.
|
||||
request.
|
||||
|
||||
`launchd` does not run a login shell either. That is the only reason
|
||||
`scripts/bridged-launchd-wrapper.sh` exists — it `exec`s `zsh -lc` so the store gets sourced, while
|
||||
`scripts/fleetd-launchd-wrapper.sh` exists — it `exec`s `zsh -lc` so the store gets sourced, while
|
||||
keeping one process so launchd's PID tracking still works. Read its header; it explains the trap
|
||||
better than this paragraph.
|
||||
|
||||
@@ -115,7 +115,7 @@ better than this paragraph.
|
||||
> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing.
|
||||
|
||||
The daemon logs which required secret names resolved at startup
|
||||
(`Bridged.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
|
||||
(`Fleetd.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
|
||||
`subscription: true`**, on purpose, because they need no token. So a green secret report says
|
||||
nothing about your subscription profiles.
|
||||
|
||||
@@ -142,7 +142,7 @@ lead was worse than a daemon that will not boot.
|
||||
|
||||
## 3. Configure
|
||||
|
||||
`bridged/bridged.yaml` is the live config. It is **gitignored**. `bridged.example.yaml` is the
|
||||
`fleetd/fleetd.yaml` is the live config. It is **gitignored**. `fleetd.example.yaml` is the
|
||||
tracked, documented copy. Two consequences you will meet:
|
||||
|
||||
- Members cannot see the live config. They work in worktrees of the tracked repo. So a change to
|
||||
@@ -220,10 +220,10 @@ not assumed.
|
||||
Use the script. Do not hand-roll the steps.
|
||||
|
||||
```bash
|
||||
scripts/redeploy-bridged.sh --check # read-only: reports state, changes nothing
|
||||
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
|
||||
scripts/redeploy-bridged.sh --no-build # restart the jar you already have
|
||||
scripts/redeploy-fleetd.sh --check # read-only: reports state, changes nothing
|
||||
scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
|
||||
scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
|
||||
scripts/redeploy-fleetd.sh --no-build # restart the jar you already have
|
||||
```
|
||||
|
||||
Run `--check` first, always. It is the only thing that reports whether the forge token resolves,
|
||||
@@ -237,8 +237,8 @@ checks to a marker taken before the restart, so old errors cannot be misread as
|
||||
`main` does nothing until you rebuild and restart. Saying "shipped" about code the live daemon has
|
||||
never loaded is a false report.
|
||||
|
||||
**Drain first.** `bridge_list`, collect anything you still want with `bridge_poll`, then
|
||||
`bridge_stop` each member. A restart drops in-flight tickets and rendezvous. A member's report is
|
||||
**Drain first.** `fleet_list`, collect anything you still want with `fleet_poll`, then
|
||||
`fleet_stop` each member. A restart drops in-flight tickets and rendezvous. A member's report is
|
||||
not recoverable once its ticket is gone.
|
||||
|
||||
### Verify — `/healthz` is not enough
|
||||
@@ -246,18 +246,18 @@ not recoverable once its ticket is gone.
|
||||
`/healthz` proves the socket connects. It does not prove a spawn works, that identity resolves, or
|
||||
that the new jar is the one running. Four checks, in order:
|
||||
|
||||
1. **A fresh boot line.** Confirm a new `bridged listening` line at the end of `bridged/bridged.out`,
|
||||
1. **A fresh boot line.** Confirm a new `fleetd listening` line at the end of `fleetd/fleetd.out`,
|
||||
dated after the restart. An old daemon that never died looks identical from outside.
|
||||
2. **Deferred keys.** The startup log names which config keys it accepted and which it deferred. A
|
||||
deferred key needing a restart is usually the whole reason you restarted. Read those lines rather
|
||||
than assuming.
|
||||
3. **Identity.** `bridge_whoami` must still answer `primary`. If the tab label changed, the lead is
|
||||
3. **Identity.** `fleet_whoami` must still answer `primary`. If the tab label changed, the lead is
|
||||
now a worker and every orchestration call is refused.
|
||||
4. **A real spawn.** Spawn one cheap member and stop it. This is the only check that catches a herdr
|
||||
protocol mismatch.
|
||||
|
||||
> Restarting the daemon **cuts your own MCP mount**, and it does not reconnect. So you cannot run
|
||||
> `bridge_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
|
||||
> `fleet_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
|
||||
> This is why the restart is done from the lead but verified after a reconnect.
|
||||
|
||||
### The REST surface
|
||||
@@ -273,20 +273,20 @@ loopback only.
|
||||
| `GET /agents` | agents as herdr sees them |
|
||||
| `GET /members` · `POST /members` · `DELETE /members/{paneId}` | list, spawn, tear down |
|
||||
| `GET /profiles` | the backends configured |
|
||||
| `GET /sessions/{id}/status` | one session — the same view as `bridge_status` |
|
||||
| `GET /sessions/{id}/status` | one session — the same view as `fleet_status` |
|
||||
| `POST /sessions/{id}/message` · `/reply` · `/ask` | the three message kinds |
|
||||
| `GET /sessions/{id}/replies` | drain the reply inbox — **destructive, see trap 6** |
|
||||
| `GET /tasks/{ticket}` | poll a detached ticket |
|
||||
|
||||
### Where it runs, and where the logs are
|
||||
|
||||
- Log file: **`bridged/bridged.out`**, in both supervised and unsupervised modes.
|
||||
- Audit log: `bridged/logs/audit.log`, rotated daily, 30 days kept.
|
||||
- Service units ship in `deploy/`: `bridged.service` for Linux systemd (ordered
|
||||
`After=herdr.service`) and `dev.ltms.bridged.plist` for macOS launchd. `deploy/lavinmq` holds the
|
||||
- Log file: **`fleetd/fleetd.out`**, in both supervised and unsupervised modes.
|
||||
- Audit log: `fleetd/logs/audit.log`, rotated daily, 30 days kept.
|
||||
- Service units ship in `deploy/`: `fleetd.service` for Linux systemd (ordered
|
||||
`After=herdr.service`) and `dev.ltms.fleetd.plist` for macOS launchd. `deploy/lavinmq` holds the
|
||||
optional broker.
|
||||
- **On this host neither is loaded.** The daemon runs as a plain `java -jar` started by
|
||||
`scripts/redeploy-bridged.sh` from a login shell. Verified with `launchctl list | grep bridg`,
|
||||
`scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep bridg`,
|
||||
which returns nothing. If you expected launchd here, that expectation is the bug.
|
||||
|
||||
---
|
||||
@@ -297,36 +297,36 @@ Eleven tools. This is the whole surface.
|
||||
|
||||
| Intent | Tool |
|
||||
|---|---|
|
||||
| Confirm your own role | `bridge_whoami` |
|
||||
| See backends available | `bridge_profiles` |
|
||||
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` |
|
||||
| See the fleet | `bridge_list` → `leads` + `members` |
|
||||
| One member's state | `bridge_status{sessionId}` |
|
||||
| Delegate, blocking | `bridge_send{sessionId, content}` |
|
||||
| Delegate, long task | `bridge_send{sessionId, content, wait:false}` → ticket |
|
||||
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||
| Collect a reply | `bridge_poll{ticket}` or `bridge_poll{target}` |
|
||||
| Clear a reply from the inbox | `bridge_ack{target, msgId}` — **`target`, not `ticket`** |
|
||||
| Member ends its turn | `bridge_reply{content}` |
|
||||
| Member asks the lead | `bridge_ask{question}` |
|
||||
| Tear down | `bridge_stop{paneId}` |
|
||||
| Confirm your own role | `fleet_whoami` |
|
||||
| See backends available | `fleet_profiles` |
|
||||
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` |
|
||||
| See the fleet | `fleet_list` → `leads` + `members` |
|
||||
| One member's state | `fleet_status{sessionId}` |
|
||||
| Delegate, blocking | `fleet_send{sessionId, content}` |
|
||||
| Delegate, long task | `fleet_send{sessionId, content, wait:false}` → ticket |
|
||||
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
|
||||
| Collect a reply | `fleet_poll{ticket}` or `fleet_poll{target}` |
|
||||
| Clear a reply from the inbox | `fleet_ack{target, msgId}` — **`target`, not `ticket`** |
|
||||
| Member ends its turn | `fleet_reply{content}` |
|
||||
| Member asks the lead | `fleet_ask{question}` |
|
||||
| Tear down | `fleet_stop{paneId}` |
|
||||
|
||||
### The loop
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant L as lead
|
||||
participant B as bridged
|
||||
participant B as fleetd
|
||||
participant M as member
|
||||
L->>B: bridge_spawn — all units first
|
||||
L->>B: bridge_send with wait false — then all sends
|
||||
L->>B: fleet_spawn — all units first
|
||||
L->>B: fleet_send with wait false — then all sends
|
||||
B-->>L: ticket
|
||||
B->>M: injected when the member is idle
|
||||
M->>B: bridge_reply
|
||||
M->>B: fleet_reply
|
||||
B-->>L: nudge into the lead's own pane
|
||||
L->>B: bridge_poll by ticket
|
||||
L->>B: bridge_ack by target and msgId
|
||||
L->>B: bridge_stop by paneId
|
||||
L->>B: fleet_poll by ticket
|
||||
L->>B: fleet_ack by target and msgId
|
||||
L->>B: fleet_stop by paneId
|
||||
```
|
||||
|
||||
*Figure 2 — the delegation loop. Spawn is separate from send on purpose.*
|
||||
@@ -334,7 +334,7 @@ sequenceDiagram
|
||||
**Spawn every unit first, then send them all.** Spawning and sending in one loop is how parallel
|
||||
work silently becomes serial. It is the most common way this whole layer is wasted.
|
||||
|
||||
**Prefer `wait:false`.** A blocking `bridge_send` is capped by **your own MCP client timeout**,
|
||||
**Prefer `wait:false`.** A blocking `fleet_send` is capped by **your own MCP client timeout**,
|
||||
about 60 seconds — far below any real task's runtime. The cap is in your client, not in the daemon,
|
||||
so no server setting fixes it.
|
||||
|
||||
@@ -379,7 +379,7 @@ Twelve traps, all hit for real this year. Grouped by where they bite.
|
||||
|
||||
**1. The daemon started from a non-login shell.**
|
||||
Symptom: everything green for hours, then a member cannot open a pull request, or a gateway profile
|
||||
gets a 401. Nothing logs it at the time. Fix: `scripts/redeploy-bridged.sh --check` is the only
|
||||
gets a 401. Nothing logs it at the time. Fix: `scripts/redeploy-fleetd.sh --check` is the only
|
||||
thing that reports it. Restart from a login shell, or via the launchd wrapper.
|
||||
|
||||
**2. `/healthz` is green and every spawn fails.**
|
||||
@@ -394,15 +394,15 @@ removed; a current daemon refuses to start if it is still present.
|
||||
|
||||
**4. A merge is not a deployment.**
|
||||
The running daemon holds its original jar. Rebuild and restart, then prove it with a **fresh**
|
||||
`bridged listening` line. Also: the restart cuts your own MCP mount and it never reconnects, so you
|
||||
`fleetd listening` line. Also: the restart cuts your own MCP mount and it never reconnects, so you
|
||||
cannot verify identity from that session — ask for `/mcp`.
|
||||
|
||||
### Losing a member's work
|
||||
|
||||
**5. The ticket expired.**
|
||||
Ticket time-to-live is about 10 minutes. After that `bridge_poll{ticket}` returns
|
||||
Ticket time-to-live is about 10 minutes. After that `fleet_poll{ticket}` returns
|
||||
`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's
|
||||
**inbox**. Drain it with `bridge_poll{target}`, then `bridge_ack{target, msgId}`.
|
||||
**inbox**. Drain it with `fleet_poll{target}`, then `fleet_ack{target, msgId}`.
|
||||
|
||||
**6. Reading the reply inbox is destructive.**
|
||||
`GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that
|
||||
@@ -419,24 +419,24 @@ queued and restarts the member when it next goes idle. Back up the member's comm
|
||||
|
||||
**9. `liveStatus: working` is not progress.**
|
||||
A member can report busy for hours doing nothing. Check file modification times in its worktree and
|
||||
snapshot `git diff` before and after. Do not read its terminal — use `bridge_status`.
|
||||
snapshot `git diff` before and after. Do not read its terminal — use `fleet_status`.
|
||||
|
||||
**10. Never brief a member to "ask me".**
|
||||
`bridge_ask` blocks for about 55 seconds and no nudge extends it. It is also invisible to
|
||||
`bridge_poll`. If you are running async, you will not see the question in time. Decide before you
|
||||
`fleet_ask` blocks for about 55 seconds and no nudge extends it. It is also invisible to
|
||||
`fleet_poll`. If you are running async, you will not see the question in time. Decide before you
|
||||
delegate, or give the member an explicit default.
|
||||
|
||||
### Merging a member's work
|
||||
|
||||
**11. The reported branch is not the branch it committed to.**
|
||||
`bridge_list` reports the branch **spawn provisioned**, not the one the member actually used. Check
|
||||
`fleet_list` reports the branch **spawn provisioned**, not the one the member actually used. Check
|
||||
`git -C <worktree> branch --show-current`, or you push an empty ref and the merge says
|
||||
"Already up to date".
|
||||
|
||||
**12. Members see a months-old wiki, and cannot see the live config.**
|
||||
`wiki/` is a submodule whose pointer is never advanced, so a member's checkout is a stale snapshot.
|
||||
Never brief "read `wiki/…`" — paste the text, and commit the member's wiki entry yourself. Same
|
||||
shape for `bridged.yaml`: it is gitignored, so a member cannot see the file its change may break.
|
||||
shape for `fleetd.yaml`: it is gitignored, so a member cannot see the file its change may break.
|
||||
|
||||
### The general shape behind several of these
|
||||
|
||||
@@ -463,8 +463,8 @@ before believing it**.
|
||||
| Adding a second host or a non-Claude peer | [10 Cross-Host Messaging](10-Cross-Host-Messaging), [12 Claude → OpenCode](12-Claude-to-OpenCode) |
|
||||
| What is left to build | [8 Roadmap](8-Roadmap) |
|
||||
| The rules every session must follow | `CLAUDE.md` in the repo — it loads into every session |
|
||||
| Rendezvous, `bridge_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code |
|
||||
| Every config key, documented | `bridged/bridged.example.yaml` |
|
||||
| Rendezvous, `fleet_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code |
|
||||
| Every config key, documented | `fleetd/fleetd.example.yaml` |
|
||||
|
||||
**The one thing that is easy to forget.** This repo *is* the bridge, so the charter block in
|
||||
`CLAUDE.md` is not documentation about someone else's system — it is the instruction surface this
|
||||
|
||||
+113
-113
@@ -1,14 +1,14 @@
|
||||
# 2. Herdr Message Server (`bridged`)
|
||||
# 2. Herdr Message Server (`fleetd`)
|
||||
|
||||
> **Status:** 🟢 Proposed primary approach (2026-07-11) — supersedes AgentAPI as the
|
||||
> centric transport. AgentAPI is retained only as a *fallback injector* (see [Approaches](3-Approaches)).
|
||||
|
||||
`bridged` is a small, always-on **message server that controls [herdr](https://herdr.dev)**
|
||||
`fleetd` is a small, always-on **message server that controls [herdr](https://herdr.dev)**
|
||||
and exposes a clean 2-way messaging API between a **primary** Claude Code session (Opus 4.8,
|
||||
on Pro/Max) and one or more **secondary worker** sessions running a different/cheaper model.
|
||||
It replaces the hand-rolled terminal emulation of AgentAPI by standing on herdr's structured
|
||||
socket API: herdr owns the PTYs, multiplexing, persistence, and — crucially — **agent-status
|
||||
events**; `bridged` owns the *policy* (subscription boundary, session lifecycle, delivery
|
||||
events**; `fleetd` owns the *policy* (subscription boundary, session lifecycle, delivery
|
||||
gating) and the *client-facing contract*. That contract is an **MCP server that both the
|
||||
primary and the workers mount** (one unified Claude setup — see
|
||||
[The client contract — MCP](#the-client-contract--mcp-unified-for-primary--workers)), with
|
||||
@@ -20,10 +20,10 @@ AgentAPI re-implements, per process, an in-memory terminal emulator and a *scree
|
||||
heuristic* to guess when the agent is done. herdr already provides all of that as a
|
||||
persistent service, and adds three things AgentAPI cannot:
|
||||
|
||||
| Capability | AgentAPI | herdr (via `bridged`) |
|
||||
| Capability | AgentAPI | herdr (via `fleetd`) |
|
||||
|---|---|---|
|
||||
| Inject a turn into a **worker** | ✅ terminal emulation | ✅ `pane.send_text` + `pane.send_keys` |
|
||||
| Inject a turn into the **primary** | ❌ (only wraps worker) | ◐ same primitive — **only when the primary is a herdr pane** (single-host); split-host wakes via the primary's `Stop`-hook polling `bridged` |
|
||||
| Inject a turn into the **primary** | ❌ (only wraps worker) | ◐ same primitive — **only when the primary is a herdr pane** (single-host); split-host wakes via the primary's `Stop`-hook polling `fleetd` |
|
||||
| "Done / blocked" signal | ⚠ screen-stability heuristic | ✅ `events.subscribe(pane.agent_status_changed)` |
|
||||
| Worker self-reports state | ❌ | ✅ `pane.report_agent` (via herdr `SKILL.md`) |
|
||||
| Multiplex a *herd* of workers + attach/observe | ❌ one server per session | ✅ native workspaces/tabs/panes |
|
||||
@@ -36,9 +36,9 @@ model has **two** shapes, and the distinction is *cross-turn busy-polling* (forb
|
||||
*blocking* (fine):
|
||||
|
||||
- **Blocking request/response (default, short/medium tasks).** The primary issues **one**
|
||||
MCP tool call — `bridge_send(target, task)` (blocking by default) — and `bridged` **holds it
|
||||
MCP tool call — `fleet_send(target, task)` (blocking by default) — and `fleetd` **holds it
|
||||
open** until it observes the turn-done edge (`agent_status: working → idle` — herdr has no
|
||||
`done` status) or a worker `bridge_reply`, then returns
|
||||
`done` status) or a worker `fleet_reply`, then returns
|
||||
the collected reply as the tool result. From the primary's view this is a single tool call
|
||||
parked on a result, exactly like any long-running `Bash` command: it consumes **no**
|
||||
Anthropic quota (the primary isn't looping, it's idle-waiting) and doesn't freeze anything
|
||||
@@ -46,20 +46,20 @@ model has **two** shapes, and the distinction is *cross-turn busy-polling* (forb
|
||||
humans/dashboards — the primary never has to hold it.
|
||||
- **Return-and-reinvoke (long/detached/async tasks).** When a task may outrun a sane request
|
||||
timeout, or is fire-and-forget, the primary's call returns immediately and the reply comes
|
||||
back later **through `bridged`** — `bridged` injects it into the primary's idle pane over
|
||||
herdr (same-host), or a split-host primary's `Stop`-hook long-polls **`bridged`** for it. In
|
||||
neither case does the primary touch a broker: if `bridged` needs durability it queues the
|
||||
back later **through `fleetd`** — `fleetd` injects it into the primary's idle pane over
|
||||
herdr (same-host), or a split-host primary's `Stop`-hook long-polls **`fleetd`** for it. In
|
||||
neither case does the primary touch a broker: if `fleetd` needs durability it queues the
|
||||
message internally (below the gateway) and still delivers by the same route. This is the
|
||||
correct shape for work that outlives a connection.
|
||||
|
||||
What `bridged` does **not** offer is a *held-open bidirectional conversation* — each exchange
|
||||
What `fleetd` does **not** offer is a *held-open bidirectional conversation* — each exchange
|
||||
is one request in, one reply out. That is a feature for a subscription-safe bridge, not a
|
||||
limitation. See **Trade-offs** below.
|
||||
|
||||
## The client contract — MCP (unified for primary + workers)
|
||||
|
||||
Both Claude sessions — the primary Opus **and** every worker — reach `bridged` the same way:
|
||||
they **mount `bridged` as an MCP server**. Claude Code speaks MCP natively, so the bridge
|
||||
Both Claude sessions — the primary Opus **and** every worker — reach `fleetd` the same way:
|
||||
they **mount `fleetd` as an MCP server**. Claude Code speaks MCP natively, so the bridge
|
||||
becomes a set of first-class tools instead of a `curl` the model must be told to run. One
|
||||
config line, identical on both sides, wires the whole mesh:
|
||||
|
||||
@@ -77,27 +77,27 @@ for *non-Claude* callers (webhooks, dashboards, a human CLI); Claude ↔ Claude
|
||||
|
||||
| Caller | Tool | Blocks? | Does |
|
||||
|---|---|---|---|
|
||||
| **Primary** | `bridge_send(message, target?, {block, timeout_seconds, auto_spawn, turn_id})` | `block:true` (default) → yes · `block:false` → no | Deliver a turn to a worker. Blocking form returns the outcome as the tool result (`reply` \| `question` \| `turn_done` \| `timeout`); detached form returns a `dispatch_id` and the reply is **injected into the primary's idle pane** when ready. |
|
||||
| **Primary** | `bridge_status(target?)` | no | Worker's live `agent_status` (`idle`\|`working`\|`blocked`\|`unknown`), queue depth, open rendezvous — and for the *calling* session it also **reports/drains pending messages addressed to it**: how a split-host / non-pane primary pulls replies injection can't serve (subsumes the earlier `bridge_poll`). |
|
||||
| **Worker** | `bridge_reply(text, {final})` | no | Emit a **structured** reply/payload to whoever awaits this turn. |
|
||||
| **Worker** | `bridge_ask(question)` | yes | Worker-initiated question up the chain (true 2-way); parks the worker until the primary answers. |
|
||||
| both | `bridge_list()` | no | List workers + status, for orchestration or a human (the earlier `bridge_sessions`). |
|
||||
| **Primary** | `bridge_spawn({profile?})` · `bridge_stop(target)` · `bridge_read(target, source)` | no | Worker lifecycle (guard-checked spawn, idempotent teardown) and peeking at a detached worker's terminal. |
|
||||
| **Primary** | `fleet_send(message, target?, {block, timeout_seconds, auto_spawn, turn_id})` | `block:true` (default) → yes · `block:false` → no | Deliver a turn to a worker. Blocking form returns the outcome as the tool result (`reply` \| `question` \| `turn_done` \| `timeout`); detached form returns a `dispatch_id` and the reply is **injected into the primary's idle pane** when ready. |
|
||||
| **Primary** | `fleet_status(target?)` | no | Worker's live `agent_status` (`idle`\|`working`\|`blocked`\|`unknown`), queue depth, open rendezvous — and for the *calling* session it also **reports/drains pending messages addressed to it**: how a split-host / non-pane primary pulls replies injection can't serve (subsumes the earlier `fleet_poll`). |
|
||||
| **Worker** | `fleet_reply(text, {final})` | no | Emit a **structured** reply/payload to whoever awaits this turn. |
|
||||
| **Worker** | `fleet_ask(question)` | yes | Worker-initiated question up the chain (true 2-way); parks the worker until the primary answers. |
|
||||
| both | `fleet_list()` | no | List workers + status, for orchestration or a human (the earlier `fleet_sessions`). |
|
||||
| **Primary** | `fleet_spawn({profile?})` · `fleet_stop(target)` · `fleet_read(target, source)` | no | Worker lifecycle (guard-checked spawn, idempotent teardown) and peeking at a detached worker's terminal. |
|
||||
|
||||
> The normative tool-by-tool surface (parameters, outcomes, error model) is
|
||||
> **`docs/MCP-Contract.md`** in the main repo (2026-07-14); this table mirrors it.
|
||||
|
||||
### The rendezvous — why worker → primary needs no keystrokes
|
||||
|
||||
`bridged` is the meeting point. When the primary is parked in a blocking `bridge_send`,
|
||||
`bridged` resolves that pending tool call the instant **either** signal arrives:
|
||||
`fleetd` is the meeting point. When the primary is parked in a blocking `fleet_send`,
|
||||
`fleetd` resolves that pending tool call the instant **either** signal arrives:
|
||||
|
||||
- herdr's status stream shows the pane's turn-done edge — `agent_status_changed: working →
|
||||
idle` (**passive** — fires even for an uncooperative worker; herdr has no `done` status),
|
||||
**or**
|
||||
- the worker calls `bridge_reply(result)` (**active** — a structured payload; preferred).
|
||||
- the worker calls `fleet_reply(result)` (**active** — a structured payload; preferred).
|
||||
|
||||
Because the reply flows back through `bridged`'s own state, the old "herdr types the answer
|
||||
Because the reply flows back through `fleetd`'s own state, the old "herdr types the answer
|
||||
into the primary's pane" path is **no longer needed, even single-host** — the primary reads
|
||||
its answer as an ordinary MCP tool result. herdr keystroke-injection into the *primary* pane
|
||||
survives only as a degraded fallback for a herdr-only (non-MCP) primary.
|
||||
@@ -105,26 +105,26 @@ survives only as a degraded fallback for a herdr-only (non-MCP) primary.
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as "Primary (Opus) — MCP client"
|
||||
participant S as "bridged (MCP server + rendezvous)"
|
||||
participant S as "fleetd (MCP server + rendezvous)"
|
||||
participant H as "herdr"
|
||||
participant W as "Worker claude — MCP client"
|
||||
|
||||
P->>S: "bridge_send(worker, task, block) — tool call PARKS"
|
||||
P->>S: "fleet_send(worker, task, block) — tool call PARKS"
|
||||
S->>H: "send_text + send_keys (gated on idle)"
|
||||
H->>W: "inject turn"
|
||||
activate W
|
||||
H-->>S: "event: agent_status_changed = working"
|
||||
par active payload
|
||||
W->>S: "bridge_reply(result) — structured"
|
||||
W->>S: "fleet_reply(result) — structured"
|
||||
and passive signal
|
||||
H-->>S: "event: agent_status_changed → idle (turn done)"
|
||||
end
|
||||
deactivate W
|
||||
S-->>P: "tool result = reply (unparks bridge_send)"
|
||||
Note over P,W: "worker → primary rode bridged's state,<br/>not a keystroke into the primary pane"
|
||||
S-->>P: "tool result = reply (unparks fleet_send)"
|
||||
Note over P,W: "worker → primary rode fleetd's state,<br/>not a keystroke into the primary pane"
|
||||
```
|
||||
|
||||
*Figure: the reply resolves on whichever of the two signals lands first; `bridge_reply`
|
||||
*Figure: the reply resolves on whichever of the two signals lands first; `fleet_reply`
|
||||
carries the structured payload, the herdr event guarantees the timing even if the worker
|
||||
never cooperates.*
|
||||
|
||||
@@ -135,8 +135,8 @@ it degrades cleanly if a worker is left unmodified:
|
||||
|
||||
| Tier | Worker setup | Reply channel | Worker can ask back? |
|
||||
|---|---|---|---|
|
||||
| **Unified (recommended)** | mounts `bridge` MCP (same one line) | structured `bridge_reply` | yes — `bridge_ask` |
|
||||
| **Hooked (no MCP)** | a `Stop`-hook installed | structured envelope POSTed to `bridged` (see [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply)) | no |
|
||||
| **Unified (recommended)** | mounts `bridge` MCP (same one line) | structured `fleet_reply` | yes — `fleet_ask` |
|
||||
| **Hooked (no MCP)** | a `Stop`-hook installed | structured envelope POSTed to `fleetd` (see [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply)) | no |
|
||||
| **Unmodified (last resort)** | stock `claude` | `pane.read` scrape on the turn-done `working→idle` edge (lossy) | no |
|
||||
|
||||
The **primary-side contract is identical** in both tiers; only the worker's reply fidelity
|
||||
@@ -147,7 +147,7 @@ bridge on either side is subscription-safe by construction (see
|
||||
|
||||
## Architecture
|
||||
|
||||
`bridged` is a **standalone daemon** — one component, two faces:
|
||||
`fleetd` is a **standalone daemon** — one component, two faces:
|
||||
|
||||
- a **SERVER face** — an **MCP server** the Claude Code sessions mount, plus REST/SSE for
|
||||
non-Claude clients, sitting over the session tracker, subscription guard, and reply
|
||||
@@ -156,8 +156,8 @@ bridge on either side is subscription-safe by construction (see
|
||||
to agent-status.
|
||||
|
||||
The Claude sessions themselves live **as panes inside herdr**. Each pane reaches *up* to
|
||||
`bridged`'s MCP server (to send/reply); `bridged`'s herdr client reaches *down* through
|
||||
herdr's socket to drive those same panes and read their status. herdr, `bridged`, the panes,
|
||||
`fleetd`'s MCP server (to send/reply); `fleetd`'s herdr client reaches *down* through
|
||||
herdr's socket to drive those same panes and read their status. herdr, `fleetd`, the panes,
|
||||
and the model are colocated on one host (same-host scenario); a broker is optional for async
|
||||
duplex.
|
||||
|
||||
@@ -168,9 +168,9 @@ flowchart TB
|
||||
WP["worker pane(s) · claude<br/>ANTHROPIC_BASE_URL set · MCP client"]
|
||||
end
|
||||
|
||||
subgraph bridged["bridged — standalone daemon (NOT a claude process)"]
|
||||
subgraph fleetd["fleetd — standalone daemon (NOT a claude process)"]
|
||||
subgraph srv["SERVER face"]
|
||||
MCP["MCP server<br/>bridge_send · reply · ask · status"]
|
||||
MCP["MCP server<br/>fleet_send · reply · ask · status"]
|
||||
REST["REST / SSE<br/>(non-Claude clients)"]
|
||||
POL["policy brain<br/>session tracker · subscription guard<br/>· reply rendezvous"]
|
||||
end
|
||||
@@ -201,13 +201,13 @@ flowchart TB
|
||||
class MODEL,BROKER warn
|
||||
```
|
||||
|
||||
*Figure: `bridged` is one standalone daemon with a **SERVER** face (the MCP endpoint the
|
||||
*Figure: `fleetd` is one standalone daemon with a **SERVER** face (the MCP endpoint the
|
||||
Claude panes mount, over the policy brain) and a **CLIENT** face (the herdr socket client).
|
||||
The Claude sessions are herdr **panes**: they call *up* into the MCP server, while `bridged`'s
|
||||
The Claude sessions are herdr **panes**: they call *up* into the MCP server, while `fleetd`'s
|
||||
client drives them *down* through herdr's socket and gates every injection on live
|
||||
agent-status. Only worker panes carry `ANTHROPIC_BASE_URL`; `bridged` holds no quota, so it
|
||||
subscribes freely. The broker is **`bridged`-internal** (durability / cross-host) — no pane
|
||||
ever addresses it; async delivery is `bridged` injecting an idle pane. (Split-host moves the
|
||||
agent-status. Only worker panes carry `ANTHROPIC_BASE_URL`; `fleetd` holds no quota, so it
|
||||
subscribes freely. The broker is **`fleetd`-internal** (durability / cross-host) — no pane
|
||||
ever addresses it; async delivery is `fleetd` injecting an idle pane. (Split-host moves the
|
||||
primary out of herdr — see [Deployment model](#deployment-model).)*
|
||||
|
||||
### Components
|
||||
@@ -217,12 +217,12 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
|
||||
| **herdr socket client** | NDJSON over `~/.config/herdr/herdr.sock`; correlates responses by `id`; maintains a long-lived `events.subscribe` stream. |
|
||||
| **Session manager** | Maps a logical session → herdr `workspace/tab/pane` id. Spawns the worker `claude` (env-prefixed launch line into a fresh pane's shell), health-checks, and **recycles on context ceiling** (Ralph loop, see below). |
|
||||
| **Injector** | Per-pane FIFO queue. Delivers `send_text` + `send_keys "enter"` **only when** that pane's `agent_status ∈ {idle, blocked}` — never mid-run. |
|
||||
| **Reply rendezvous** | Resolves an awaiting `bridge_send` on whichever lands first: a worker `bridge_reply` (structured, **preferred**), the turn-done `working → idle` status edge (timing guarantee), or — worker-side hook path — a `Stop`-hook envelope; last-resort `pane.read {source:"recent-unwrapped"}` scrape. See [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply). |
|
||||
| **Reply rendezvous** | Resolves an awaiting `fleet_send` on whichever lands first: a worker `fleet_reply` (structured, **preferred**), the turn-done `working → idle` status edge (timing guarantee), or — worker-side hook path — a `Stop`-hook envelope; last-resort `pane.read {source:"recent-unwrapped"}` scrape. See [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply). |
|
||||
| **Subscription guard** | Refuses to spawn a *worker* pane without `ANTHROPIC_BASE_URL`; refuses to *ever* set it on a pane designated *primary*; can assert egress host via `pane.process_info`. |
|
||||
| **SERVER API** | **MCP server** — the Claude-facing contract both primary and workers mount (`bridge_send`/`reply`/`ask`/`status`/`poll`/`sessions`). Plus **REST + SSE** (OpenAPI, AgentAPI-shaped) for non-Claude clients — webhooks, dashboards, a human CLI. |
|
||||
| **Broker connector** *(optional, internal)* | `bridged`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `bridged` will later inject into an idle pane. No Claude session ever connects to it. |
|
||||
| **SERVER API** | **MCP server** — the Claude-facing contract both primary and workers mount (`fleet_send`/`reply`/`ask`/`status`/`poll`/`sessions`). Plus **REST + SSE** (OpenAPI, AgentAPI-shaped) for non-Claude clients — webhooks, dashboards, a human CLI. |
|
||||
| **Broker connector** *(optional, internal)* | `fleetd`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `fleetd` will later inject into an idle pane. No Claude session ever connects to it. |
|
||||
|
||||
## The herdr control contract (what `bridged` drives)
|
||||
## The herdr control contract (what `fleetd` drives)
|
||||
|
||||
> **Verified against the running herdr 0.7.0 (protocol 14), not just docs.** A spike hit the
|
||||
> live socket. `ping` returns `{version:"0.7.0", protocol:14, capabilities:{live_handoff:true}}`;
|
||||
@@ -233,11 +233,11 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
|
||||
> `agent.get`, `agent.focus`, plus `server.agent_manifests` and `pane.report_agent` — so herdr
|
||||
> already models "agents", not just panes. (Also available: `worktree.*` for git-isolated workers.)
|
||||
>
|
||||
> **CB-102 — decided: `bridged` drives the native `agent.*` path** (the pane + `send_text`
|
||||
> **CB-102 — decided: `fleetd` drives the native `agent.*` path** (the pane + `send_text`
|
||||
> workaround below is the documented fallback only). The spike proved against the live daemon that
|
||||
> `agent.start` accepts a **first-class `env` map** that reaches the process environment — so a
|
||||
> worker's `ANTHROPIC_BASE_URL` is injected cleanly and guard-checked, with no shell-prefix
|
||||
> parsing and `bridged`'s own env untouched. herdr also tracks each worker's **Claude session
|
||||
> parsing and `fleetd`'s own env untouched. herdr also tracks each worker's **Claude session
|
||||
> UUID** (`agent_session.value`), grounding the ID contract in herdr's own identity. Observed
|
||||
> `agent.*` schema:
|
||||
>
|
||||
@@ -255,8 +255,8 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
|
||||
>
|
||||
> **CB-108 — worker placement: a tab per worker in a dedicated "worker space."** By default
|
||||
> `agent.start` (no `tab_id`) *splits the currently-focused tab*, so it would clutter — and could
|
||||
> co-tenant — the user's real work tabs. Instead `bridged` puts every worker in its **own tab**
|
||||
> inside a dedicated **workspace** (`worker.workspace`, default `bridged-workers`), found-or-created
|
||||
> co-tenant — the user's real work tabs. Instead `fleetd` puts every worker in its **own tab**
|
||||
> inside a dedicated **workspace** (`worker.workspace`, default `fleetd-workers`), found-or-created
|
||||
> once and shared (a future per-session layout is just a distinct label). The recipe, all verified
|
||||
> live: `workspace.create {label}` (idempotent find-first) → `tab.create {workspace_id}` →
|
||||
> `agent.start {…, tab_id}` → `pane.close {root_pane}` (drop herdr's seed shell so the tab holds
|
||||
@@ -266,7 +266,7 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
|
||||
> the legacy split behaviour.
|
||||
>
|
||||
> **Two more `agent.*` facts the CB-108 build pinned:** (1) an agent's **`name` must be unique**
|
||||
> among running agents — a 2nd `agent.start {name:"claude"}` fails `agent_name_taken`, so `bridged`
|
||||
> among running agents — a 2nd `agent.start {name:"claude"}` fails `agent_name_taken`, so `fleetd`
|
||||
> names workers `claude-<profile>-<nonce>-<seq>` (the per-process `nonce` survives a daemon restart
|
||||
> where old workers linger). (2) herdr detects an agent's **kind and status from terminal output**
|
||||
> (braille spinner / `❯` prompt patterns in the remote `*.toml` manifests), **not** from `name` — so
|
||||
@@ -306,7 +306,7 @@ reference and for herdr builds without the `agent.*` namespace:
|
||||
```
|
||||
|
||||
> **Implementation note (verify in the CLI reference):** herdr's socket may or may not expose
|
||||
> a direct "spawn command + env" primitive. `bridged` uses the robust path — create pane →
|
||||
> a direct "spawn command + env" primitive. `fleetd` uses the robust path — create pane →
|
||||
> `send_text` the env-prefixed launch line — which guarantees `ANTHROPIC_BASE_URL` lands in
|
||||
> the **worker pane's shell only**. If a native spawn call exists, prefer it and pass env
|
||||
> explicitly; the guard invariant is unchanged.
|
||||
@@ -314,8 +314,8 @@ reference and for herdr builds without the `agent.*` namespace:
|
||||
## Reply envelope (how a worker emits a structured reply)
|
||||
|
||||
With the worker **mounting the bridge MCP** (the unified setup), the cleanest path is direct:
|
||||
the worker simply **calls `bridge_reply(result)`** before finishing — a first-class tool call
|
||||
that hands `bridged` a structured payload with no transcript scraping. Prefer this whenever
|
||||
the worker simply **calls `fleet_reply(result)`** before finishing — a first-class tool call
|
||||
that hands `fleetd` a structured payload with no transcript scraping. Prefer this whenever
|
||||
the worker is MCP-mounted.
|
||||
|
||||
The hook path below remains the **fallback** for a herdr-only worker (no MCP) — a Claude Code
|
||||
@@ -325,12 +325,12 @@ envelope" resolves to **a worker-side hook that runs our code at turn end**:
|
||||
1. A **`Stop`-hook** on the worker fires when its turn ends. The hook reads the **last
|
||||
assistant message** from the session transcript (`~/.claude/projects/<proj>/<session>.jsonl`,
|
||||
the path Claude Code exposes to hooks) and POSTs `{session_id, turn_id, status, text,
|
||||
artifacts}` to **`bridged`'s callback endpoint** — the gateway, not a broker (the worker-side
|
||||
hook never writes the broker directly; `bridged` queues internally if it must).
|
||||
2. `bridged` correlates that envelope to the open blocking request by `session_id`/`turn_id`
|
||||
artifacts}` to **`fleetd`'s callback endpoint** — the gateway, not a broker (the worker-side
|
||||
hook never writes the broker directly; `fleetd` queues internally if it must).
|
||||
2. `fleetd` correlates that envelope to the open blocking request by `session_id`/`turn_id`
|
||||
and returns it as the response body. The turn-done `working → idle` status edge is the
|
||||
*timing* signal; the hook payload is the *content*.
|
||||
3. If no hook is installed, `bridged` falls back to `pane.read {source:"recent-unwrapped"}`
|
||||
3. If no hook is installed, `fleetd` falls back to `pane.read {source:"recent-unwrapped"}`
|
||||
and best-effort parses the last assistant block (reuse AgentAPI's `msgfmt`). This is
|
||||
lossy and is the reason the envelope path is preferred.
|
||||
|
||||
@@ -338,7 +338,7 @@ envelope" resolves to **a worker-side hook that runs our code at turn end**:
|
||||
> `turn`/`corr`/`kind`/`body`, and the `kind` verb vocabulary that tells a recipient what to do
|
||||
> next) is defined in [Use Cases → The ID contract](7-Use-Cases#mechanism-4--the-id-contract-envelope).
|
||||
> Landing it is **Stage 2** ([Roadmap](8-Roadmap), ticket `CB-201`). The MCP-mounted worker
|
||||
> emits its reply via `bridge_reply` (structured); the `Stop`-hook envelope remains the fallback
|
||||
> emits its reply via `fleet_reply` (structured); the `Stop`-hook envelope remains the fallback
|
||||
> for a non-MCP worker. Still open: exactly how tool/diff artifacts attach to `review.reply`.
|
||||
|
||||
## Delivery gating & races (the injector is a single writer)
|
||||
@@ -347,7 +347,7 @@ The injector gates on cached `agent_status` (updated by the `events.subscribe` s
|
||||
then calls `send_text`. That read-then-send is a **TOCTOU window**: the pane could leave
|
||||
`idle` between the status read and the keystrokes landing. Mitigations, and the residual gap:
|
||||
|
||||
- **Single writer per pane.** `bridged` is the *only* automated injector into a worker pane;
|
||||
- **Single writer per pane.** `fleetd` is the *only* automated injector into a worker pane;
|
||||
the per-pane FIFO queue serializes deliveries so two turns never interleave. This removes
|
||||
injector-vs-injector races, not injector-vs-agent ones.
|
||||
- **Serialize send within the event loop.** Do the status check and the `send_text`/`send_keys`
|
||||
@@ -360,7 +360,7 @@ then calls `send_text`. That read-then-send is a **TOCTOU window**: the pane cou
|
||||
- **Residual race (accepted).** A human typing into the same worker pane, or the agent
|
||||
self-transitioning to `working` in the millisecond after the gate, can still collide. The
|
||||
cost is a corrupted turn, not a subscription breach; recovery is `ctrl+c` + re-inject.
|
||||
Treat a worker pane as **bridged-owned** (don't hand-drive it) to avoid this.
|
||||
Treat a worker pane as **fleetd-owned** (don't hand-drive it) to avoid this.
|
||||
|
||||
## Message flow
|
||||
|
||||
@@ -369,40 +369,40 @@ then calls `send_text`. That read-then-send is a **TOCTOU window**: the pane cou
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant P as "Primary (Opus)"
|
||||
participant S as "bridged"
|
||||
participant S as "fleetd"
|
||||
participant H as "herdr"
|
||||
participant W as "Worker claude"
|
||||
|
||||
P->>S: "bridge_send(id, task, block) — tool call PARKS"
|
||||
P->>S: "fleet_send(id, task, block) — tool call PARKS"
|
||||
S->>S: "await pane status = idle"
|
||||
S->>H: "pane.send_text + send_keys enter"
|
||||
H->>W: "inject turn"
|
||||
activate W
|
||||
H-->>S: "event: agent_status_changed = working"
|
||||
S-->>P: "SSE: status working (observers only)"
|
||||
W->>S: "bridge_reply(result) — or Stop-hook envelope (fallback)"
|
||||
W->>S: "fleet_reply(result) — or Stop-hook envelope (fallback)"
|
||||
H-->>S: "event: agent_status_changed → idle (turn done)"
|
||||
deactivate W
|
||||
S->>S: "resolve reply (bridge_reply, event, or pane.read fallback)"
|
||||
S->>S: "resolve reply (fleet_reply, event, or pane.read fallback)"
|
||||
S-->>P: "tool result = assistant reply (unparks call)"
|
||||
Note over P: "review diff / result, merge"
|
||||
```
|
||||
|
||||
*Figure: the primary's request blocks; `bridged` gates injection on `idle`, streams status
|
||||
*Figure: the primary's request blocks; `fleetd` gates injection on `idle`, streams status
|
||||
transitions over SSE for observers, and returns the reply — from the worker's `Stop`-hook
|
||||
envelope, falling back to a `recent-unwrapped` scrape — as the blocking call's response body.*
|
||||
|
||||
### Async duplex — bridged-mediated (event bus / worker → recipient)
|
||||
### Async duplex — fleetd-mediated (event bus / worker → recipient)
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant SRC as "Source: webhook/bus (REST) · worker bridge_reply/ask (MCP)"
|
||||
participant S as "bridged (gateway)"
|
||||
participant SRC as "Source: webhook/bus (REST) · worker fleet_reply/ask (MCP)"
|
||||
participant S as "fleetd (gateway)"
|
||||
participant Q as "broker / queue (internal)"
|
||||
participant H as "herdr"
|
||||
participant R as "Recipient pane (idle Claude)"
|
||||
|
||||
SRC->>S: "REST ingress · or MCP bridge_reply / bridge_ask"
|
||||
SRC->>S: "REST ingress · or MCP fleet_reply / fleet_ask"
|
||||
opt durability / cross-host
|
||||
S->>Q: "enqueue (group + XACK)"
|
||||
Q-->>S: "dequeue when ready"
|
||||
@@ -411,18 +411,18 @@ sequenceDiagram
|
||||
H-->>S: "event: idle"
|
||||
S->>H: "pane.send_text + send_keys (inject)"
|
||||
H->>R: "new turn = the message"
|
||||
Note over S,R: "recipient polled nothing — bridged pushed on the idle edge.<br/>split-host primary: its Stop-hook polls bridged, never the broker"
|
||||
Note over S,R: "recipient polled nothing — fleetd pushed on the idle edge.<br/>split-host primary: its Stop-hook polls fleetd, never the broker"
|
||||
```
|
||||
|
||||
*Figure: `bridged` mediates async in both directions. Every source reaches it over the gateway
|
||||
*Figure: `fleetd` mediates async in both directions. Every source reaches it over the gateway
|
||||
(MCP for Claude, REST for external), it optionally parks the message on its **internal** queue,
|
||||
waits for the recipient's idle event, and injects. The queue is a `bridged` implementation
|
||||
waits for the recipient's idle event, and injects. The queue is a `fleetd` implementation
|
||||
detail — no Claude session touches it, and herdr scrollback is never the source of truth.*
|
||||
|
||||
## Worker session lifecycle
|
||||
|
||||
`bridged` treats a worker as a **recyclable** resource, not one immortal session — a
|
||||
long-lived pane fills its context window. When idle-cycle or token caps trip, `bridged`
|
||||
`fleetd` treats a worker as a **recyclable** resource, not one immortal session — a
|
||||
long-lived pane fills its context window. When idle-cycle or token caps trip, `fleetd`
|
||||
kills the pane and respawns fresh (**Ralph loop**).
|
||||
|
||||
**What "state on disk" actually means (be precise — this is easy to hand-wave).** Recycling
|
||||
@@ -438,7 +438,7 @@ the worker externalizes as it works**:
|
||||
continue from it,"* not with the old messages.
|
||||
|
||||
This only works if the worker is disciplined about writing that state **before** a recycle
|
||||
boundary. `bridged` can enforce a checkpoint (inject "commit and update STATE.md" before it
|
||||
boundary. `fleetd` can enforce a checkpoint (inject "commit and update STATE.md" before it
|
||||
kills the pane), but a worker that ignores it loses in-flight context. **Open question:**
|
||||
whether to also snapshot the raw transcript (`--resume` the *same* session on crash-restart,
|
||||
vs. a clean context on a planned recycle) — the two restart reasons may want different
|
||||
@@ -450,7 +450,7 @@ stateDiagram-v2
|
||||
Spawning --> Ready: "claude prompt detected"
|
||||
Ready --> Working: "turn injected"
|
||||
Working --> Blocked: "permission / question"
|
||||
Blocked --> Working: "bridged answers (send_input)"
|
||||
Blocked --> Working: "fleetd answers (send_input)"
|
||||
Working --> Ready: "agent_status → idle (turn done)"
|
||||
Ready --> Recycling: "context / idle cap hit"
|
||||
Recycling --> Spawning: "state persisted to disk"
|
||||
@@ -459,7 +459,7 @@ stateDiagram-v2
|
||||
Ready --> [*]: "drain / shutdown"
|
||||
```
|
||||
|
||||
*Figure: the state machine `bridged` drives per worker. `blocked`, the turn-done
|
||||
*Figure: the state machine `fleetd` drives per worker. `blocked`, the turn-done
|
||||
`working → idle` edge, and `pane.exited` are real herdr signals, not heuristics — the reason
|
||||
herdr-centric beats screen scraping.*
|
||||
|
||||
@@ -468,7 +468,7 @@ herdr-centric beats screen scraping.*
|
||||
The invariant is unchanged from [Architecture](1-Architecture) — **anything that sets
|
||||
`ANTHROPIC_BASE_URL` is, by definition, the worker** — but here it is *enforced in code*:
|
||||
|
||||
- `bridged` is **not** a `claude` process. It consumes zero Anthropic quota, so it may
|
||||
- `fleetd` is **not** a `claude` process. It consumes zero Anthropic quota, so it may
|
||||
busy-poll the broker and hold a permanent herdr event subscription with no policy concern.
|
||||
- The **subscription guard** blocks any spawn of a *worker* pane whose launch does not
|
||||
resolve an **off-subscription** `ANTHROPIC_BASE_URL`, and blocks any attempt to set that
|
||||
@@ -481,14 +481,14 @@ The invariant is unchanged from [Architecture](1-Architecture) — **anything th
|
||||
would pass a naive check while burning subscription-adjacent auth). The guard must
|
||||
validate the **resolved value's host** against an allowlist and, after spawn, confirm
|
||||
egress via `pane.process_info` / a health call to the worker model — not trust the string.
|
||||
- The guard only constrains panes **`bridged` spawns**. A split-host primary on your Mac is
|
||||
a process `bridged` never sees; it cannot inspect that env. There, subscription safety
|
||||
- The guard only constrains panes **`fleetd` spawns**. A split-host primary on your Mac is
|
||||
a process `fleetd` never sees; it cannot inspect that env. There, subscription safety
|
||||
rests on the operator (the Mac `claude` simply is never given the var) plus the fact that
|
||||
the *only* thing crossing to the worker host is **MCP/HTTP traffic to `bridged`**, never an
|
||||
the *only* thing crossing to the worker host is **MCP/HTTP traffic to `fleetd`**, never an
|
||||
endpoint swap and never a broker connection.
|
||||
- Injecting keystrokes into the **primary** pane is subscription-safe: it is simulated
|
||||
typing, identical to the human at the keyboard — the primary still talks to
|
||||
`api.anthropic.com` on Pro/Max. `bridged` never re-points the primary's endpoint. (This
|
||||
`api.anthropic.com` on Pro/Max. `fleetd` never re-points the primary's endpoint. (This
|
||||
path exists only single-host, where the primary is a herdr pane.)
|
||||
- A startup self-check asserts any **locally-hosted** primary pane's env has **no**
|
||||
`ANTHROPIC_BASE_URL` and logs each worker's resolved egress host. It cannot self-check a
|
||||
@@ -497,8 +497,8 @@ The invariant is unchanged from [Architecture](1-Architecture) — **anything th
|
||||
## API surface (SERVER face)
|
||||
|
||||
**Two faces over one core — and REST is the testability surface.** Every feature is implemented
|
||||
as a **REST route**; the **MCP tools are thin adapters over those routes** — e.g. `bridge_send`
|
||||
→ `POST /sessions/{id}/message`, `bridge_status` → `GET /sessions/{id}/status`, `bridge_reply` →
|
||||
as a **REST route**; the **MCP tools are thin adapters over those routes** — e.g. `fleet_send`
|
||||
→ `POST /sessions/{id}/message`, `fleet_status` → `GET /sessions/{id}/status`, `fleet_reply` →
|
||||
`POST /sessions/{id}/reply` (rendezvous). The REST routes are also the surface non-Claude clients
|
||||
use (webhooks, dashboards, a human CLI), kept AgentAPI-shaped for drop-in migration. Because the
|
||||
logic lives in REST, **each feature is acceptance-tested by an HTTP call with no Claude/MCP in
|
||||
@@ -522,15 +522,15 @@ the loop**, and MCP is verified by a **parity test** (tool result == REST result
|
||||
1. **Subscription-safe delegation (the core case).** Primary Opus offloads bulk/routine
|
||||
work — codegen, refactors, test writing, log triage — to a worker on a cheap/local model,
|
||||
keeping Opus's context clean and its quota for review/merge decisions.
|
||||
2. **A herd of specialized workers.** One `bridged` + herdr multiplexes several workers
|
||||
2. **A herd of specialized workers.** One `fleetd` + herdr multiplexes several workers
|
||||
(e.g. a DeepSeek coder, a fast summarizer, a long-context reader), each its own pane,
|
||||
each addressable by `session_id`. Rolls up to a single status sidebar.
|
||||
3. **Async event-bus automation.** A webhook/CI/NATS event hits `bridged`'s **REST ingress**;
|
||||
`bridged` wakes an idle worker by injection, and the result flows back to the primary (or a
|
||||
Slack/Telegram bridge) — all mediated by `bridged`, no human in the loop. The source never
|
||||
3. **Async event-bus automation.** A webhook/CI/NATS event hits `fleetd`'s **REST ingress**;
|
||||
`fleetd` wakes an idle worker by injection, and the result flows back to the primary (or a
|
||||
Slack/Telegram bridge) — all mediated by `fleetd`, no human in the loop. The source never
|
||||
addresses a worker or a broker directly.
|
||||
4. **Human co-pilot from anywhere.** Because herdr persists and detaches, the same worker is
|
||||
reachable from a phone/chat bridge **posting to `bridged`** while you're away, and from the
|
||||
reachable from a phone/chat bridge **posting to `fleetd`** while you're away, and from the
|
||||
attached TUI when you're back.
|
||||
5. **Long-running "perpetual" workers.** The Ralph-loop lifecycle lets a worker run for hours
|
||||
across many context recycles without a human respawning it, state carried on disk.
|
||||
@@ -539,16 +539,16 @@ the loop**, and MCP is verified by a **parity test** (tool result == REST result
|
||||
|
||||
### Single-host (default — simplest, recommended to start)
|
||||
|
||||
Primary, `bridged`, herdr, and workers all on one off-subscription box. `bridged` can inject
|
||||
Primary, `fleetd`, herdr, and workers all on one off-subscription box. `fleetd` can inject
|
||||
into **both** the primary and the worker panes (both local to one herdr), so the broker is
|
||||
optional. Best for a workstation or a single dev box.
|
||||
|
||||
### Split-host (primary local, workers remote near the model)
|
||||
|
||||
Primary Opus runs on your Mac; `bridged` + herdr + workers run on the GPU host next to
|
||||
Primary Opus runs on your Mac; `fleetd` + herdr + workers run on the GPU host next to
|
||||
`ollama.ltms.dev` / GX10 vLLM. herdr's socket is **local-only**, so the Mac reaches the worker
|
||||
host **only over `bridged`'s MCP/HTTP endpoint** — never a remote herdr socket and never the
|
||||
broker (the broker, if any, stays `bridged`-internal on the worker host).
|
||||
host **only over `fleetd`'s MCP/HTTP endpoint** — never a remote herdr socket and never the
|
||||
broker (the broker, if any, stays `fleetd`-internal on the worker host).
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
@@ -557,10 +557,10 @@ flowchart LR
|
||||
end
|
||||
subgraph host["Worker host (off-subscription, near model)"]
|
||||
direction TB
|
||||
BD["bridged<br/>:8080 MCP · REST/SSE"]
|
||||
BD["fleetd<br/>:8080 MCP · REST/SSE"]
|
||||
HS["herdr server"]
|
||||
W2["worker claude pane(s)"]
|
||||
BR["broker / queue<br/>(bridged-owned, internal)"]
|
||||
BR["broker / queue<br/>(fleetd-owned, internal)"]
|
||||
BD -->|"Unix socket"| HS --> W2
|
||||
BD -.->|"durability / cross-host"| BR
|
||||
end
|
||||
@@ -577,22 +577,22 @@ flowchart LR
|
||||
class BR warn
|
||||
```
|
||||
|
||||
*Figure: the Mac's **only** link to the worker host is `bridged`'s MCP/HTTP endpoint — it
|
||||
*Figure: the Mac's **only** link to the worker host is `fleetd`'s MCP/HTTP endpoint — it
|
||||
carries sync replies, and (since the Mac primary isn't a herdr pane) its `Stop`-hook polls that
|
||||
same endpoint for async wake-ups. herdr's socket and the broker stay local and `bridged`-owned.
|
||||
same endpoint for async wake-ups. herdr's socket and the broker stay local and `fleetd`-owned.
|
||||
The subscription boundary tracks the host boundary — nothing on the Mac ever sets
|
||||
`ANTHROPIC_BASE_URL`.*
|
||||
|
||||
**Security:** bind `bridged`'s HTTP to `localhost` and reach it over an SSH tunnel, or front
|
||||
**Security:** bind `fleetd`'s HTTP to `localhost` and reach it over an SSH tunnel, or front
|
||||
it with a bearer token + TLS. Never expose the port unauthenticated — it is an agent-control
|
||||
surface: `POST /message` runs arbitrary prompts, and `POST /keys` sends raw keystrokes
|
||||
(including `ctrl+c`) into a live agent. (This closes the gap left by AgentAPI's open `:3284`.)
|
||||
|
||||
**Multi-tenancy is an open item.** A single `bridged` fronting a *herd* of workers today has
|
||||
**Multi-tenancy is an open item.** A single `fleetd` fronting a *herd* of workers today has
|
||||
**one shared token = full control of every session**; there is no per-session authorization.
|
||||
That is acceptable for a single-operator box but not for shared/multi-user use. Before that,
|
||||
add per-session scoping (a capability token per `session_id`) and an audit log of injected
|
||||
turns. Until then, treat one `bridged` as one trust domain.
|
||||
turns. Until then, treat one `fleetd` as one trust domain.
|
||||
|
||||
## Proposed tech stack
|
||||
|
||||
@@ -603,10 +603,10 @@ turns. Until then, treat one `bridged` as one trust domain.
|
||||
| **herdr transport** | JDK **`UnixDomainSocketAddress` + `SocketChannel`** (native UDS, no dep), **NDJSON** via Jackson, `id`-correlated + a persistent events stream | Native herdr contract; a dedicated **virtual thread** blocks on the event stream. | — |
|
||||
| **SERVER API — Claude** | **Official MCP Java SDK** (streamable-HTTP transport) on an embedded server (Jetty / Spring Boot) | The unified contract both primary and workers mount; native to Claude Code, no shell/`curl`, subscription-safe by construction | stdio MCP adapter (per-session subprocess) if a long-lived HTTP endpoint is undesirable |
|
||||
| **SERVER API — others** | **REST + SSE** via **Javalin** (light) or Spring MVC | Drop-in for AgentAPI-shaped/non-Claude clients; SSE streams status cheaply | JAX-RS (Helidon/Quarkus); gRPC if callers are all code |
|
||||
| **Internal queue** *(optional)* | **Redis Streams** via **Lettuce** (consumer groups, `XACK`, visibility timeout) — `bridged`-owned, below the gateway | Durability + cross-host for async; satisfies the guardrails in [Architecture](1-Architecture). Same-host can start with an in-JVM queue and add this only when durability/cross-host is needed | NATS JetStream for multi-host scale; embedded H2/SQLite for a single host |
|
||||
| **Internal queue** *(optional)* | **Redis Streams** via **Lettuce** (consumer groups, `XACK`, visibility timeout) — `fleetd`-owned, below the gateway | Durability + cross-host for async; satisfies the guardrails in [Architecture](1-Architecture). Same-host can start with an in-JVM queue and add this only when durability/cross-host is needed | NATS JetStream for multi-host scale; embedded H2/SQLite for a single host |
|
||||
| **Config** | **YAML via Jackson** (`jackson-dataformat-yaml`) + env overrides | 12-factor; secrets via env only | MicroProfile Config; Spring config if on Spring Boot |
|
||||
| **Observability** | **SLF4J + Logback**; **Micrometer** → Prometheus `/metrics`; `/healthz` | Ops from day one | OpenTelemetry traces |
|
||||
| **Process supervision** | **systemd** unit (`java -jar` or the native-image binary), ordered after herdr | Restart-on-crash; ordered start (herdr before `bridged`) | Docker Compose colocating herdr + `bridged`; k8s (overkill for one host) |
|
||||
| **Process supervision** | **systemd** unit (`java -jar` or the native-image binary), ordered after herdr | Restart-on-crash; ordered start (herdr before `fleetd`) | Docker Compose colocating herdr + `fleetd`; k8s (overkill for one host) |
|
||||
| **Testing** | **JUnit 5** + a **mock UDS socket** server + golden transcripts; fake `ccs`/`claude` stubs | Deterministic CI without a real TTY | Testcontainers (Redis) + a real herdr for e2e |
|
||||
|
||||
**Recommendation: Java 21+ with virtual threads.** The daemon is almost entirely
|
||||
@@ -652,8 +652,8 @@ final class Guard {
|
||||
throw new GuardException("worker base_url host %s is not an approved off-subscription host".formatted(host));
|
||||
}
|
||||
|
||||
// Only meaningful for a primary bridged itself hosts (single-host). A remote/Mac primary is a
|
||||
// process bridged never sees — its cleanliness is the operator's.
|
||||
// Only meaningful for a primary fleetd itself hosts (single-host). A remote/Mac primary is a
|
||||
// process fleetd never sees — its cleanliness is the operator's.
|
||||
void assertLocalPrimaryClean(Map<String,String> env) {
|
||||
if (env.containsKey("ANTHROPIC_BASE_URL"))
|
||||
throw new GuardException("primary env is tainted — this is the subscription line");
|
||||
@@ -666,25 +666,25 @@ final class Guard {
|
||||
| Milestone | Deliverable | Proves |
|
||||
|---|---|---|
|
||||
| **M0 — Spike** | herdr client + spawn one worker pane + one `send_text`/`send_keys` round-trip | herdr socket drives a real `claude` |
|
||||
| **M1 — Status gate + MCP** | `events.subscribe` → Injector delivers only on `idle`/`blocked`; **MCP `bridge_send`/`status` mounted on the primary**, blocking reply via the rendezvous; SSE status out | No mid-run corruption; real completion signal; primary drives over MCP |
|
||||
| **M2 — Boundary + lifecycle + worker MCP** | Subscription guard + Ralph-loop recycle + worker `bridge_reply`/`bridge_ask` (unified mount) + envelope fallback | Subscription-safe; survives context ceiling; symmetric 2-way |
|
||||
| **M3 — Async + durability + split-host** | Idle-injection async delivery; `bridged`-internal queue (Redis Streams) for durability/cross-host; split-host `Stop`-hook adapter that polls `bridged` | Detached/long work, cross-host, gateway-only (no Claude↔broker) |
|
||||
| **M1 — Status gate + MCP** | `events.subscribe` → Injector delivers only on `idle`/`blocked`; **MCP `fleet_send`/`status` mounted on the primary**, blocking reply via the rendezvous; SSE status out | No mid-run corruption; real completion signal; primary drives over MCP |
|
||||
| **M2 — Boundary + lifecycle + worker MCP** | Subscription guard + Ralph-loop recycle + worker `fleet_reply`/`fleet_ask` (unified mount) + envelope fallback | Subscription-safe; survives context ceiling; symmetric 2-way |
|
||||
| **M3 — Async + durability + split-host** | Idle-injection async delivery; `fleetd`-internal queue (Redis Streams) for durability/cross-host; split-host `Stop`-hook adapter that polls `fleetd` | Detached/long work, cross-host, gateway-only (no Claude↔broker) |
|
||||
| **M4 — Harden** | Auth/TLS, metrics, mock-socket CI, systemd unit | Production shape |
|
||||
|
||||
## Trade-offs & risks
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| **No held-open conversation** — each exchange is one request in, one reply out | By design. Short/medium tasks use a **single blocking call** (fine — no quota burn); long/detached tasks use **return-and-reinvoke**, the reply delivered later **through `bridged`** (idle-pane injection, or a split-host `Stop`-hook polling `bridged`). What's excluded is a persistent bidirectional stream the primary must babysit. See *How the primary actually consumes a reply*. |
|
||||
| **Blocking call can outlive its timeout** on a very long task | Set a request deadline; on timeout `bridged` returns "still working, await async" and the reply lands via `bridged`'s async path (idle-injection) instead of erroring the delegation. Pick async delivery (Mode 2) up front for known-long work. |
|
||||
| **herdr is young / single-dev** — betting transport on it | Durability lives in `bridged`'s **internal queue** (mature Redis/NATS), not herdr — herdr carries only ephemeral delivery + status. The injector is a **pluggable interface** — fall back to `tmux send-keys` or AgentAPI without touching the queue or the gateway contract. |
|
||||
| **herdr socket is local-only** | `bridged`'s **MCP/HTTP** is the sole cross-host link; herdr and the queue stay per-host and `bridged`-owned. |
|
||||
| **No held-open conversation** — each exchange is one request in, one reply out | By design. Short/medium tasks use a **single blocking call** (fine — no quota burn); long/detached tasks use **return-and-reinvoke**, the reply delivered later **through `fleetd`** (idle-pane injection, or a split-host `Stop`-hook polling `fleetd`). What's excluded is a persistent bidirectional stream the primary must babysit. See *How the primary actually consumes a reply*. |
|
||||
| **Blocking call can outlive its timeout** on a very long task | Set a request deadline; on timeout `fleetd` returns "still working, await async" and the reply lands via `fleetd`'s async path (idle-injection) instead of erroring the delegation. Pick async delivery (Mode 2) up front for known-long work. |
|
||||
| **herdr is young / single-dev** — betting transport on it | Durability lives in `fleetd`'s **internal queue** (mature Redis/NATS), not herdr — herdr carries only ephemeral delivery + status. The injector is a **pluggable interface** — fall back to `tmux send-keys` or AgentAPI without touching the queue or the gateway contract. |
|
||||
| **herdr socket is local-only** | `fleetd`'s **MCP/HTTP** is the sole cross-host link; herdr and the queue stay per-host and `fleetd`-owned. |
|
||||
| **Mid-run interrupt still unsolved** | Same as AgentAPI. Injection gates on status; `ctrl+c` via `pane.send_input` is the only (disruptive) interrupt. |
|
||||
| **Spawn-with-env uncertainty in socket API** | Launch via `send_text` of the env-prefixed command → env is provably worker-only; verify native spawn in the CLI reference and prefer it if present. |
|
||||
| **Reply-scrape fragility (fallback path)** | Prefer the structured **envelope** path; scrape `recent-unwrapped` only as a last resort. |
|
||||
| **herdr socket API is unversioned + single-dev churn** | Pin the herdr version in the systemd/Compose unit; keep the socket client behind the `Herdr` interface; probe `ping` (assert `protocol: 14`) on connect and enumerate via `workspace.list`/`pane.list`, failing fast on an unexpected schema. Don't build against `UNCERTAIN` primitives (e.g. native spawn-with-env) until confirmed in the running CLI. |
|
||||
| **SPOF (bridged / herdr / queue)** | Documented in [Architecture](1-Architecture) → *Failure modes*. Key property: the **primary is never downstream** of a bridge component, so a total outage costs workers only, never the subscription session. |
|
||||
| **Injection TOCTOU / shared pane** | Single-writer injector + serialized send; worker panes are bridged-owned. Residual collision corrupts a turn (recoverable), never the subscription boundary. See *Delivery gating & races*. |
|
||||
| **SPOF (fleetd / herdr / queue)** | Documented in [Architecture](1-Architecture) → *Failure modes*. Key property: the **primary is never downstream** of a bridge component, so a total outage costs workers only, never the subscription session. |
|
||||
| **Injection TOCTOU / shared pane** | Single-writer injector + serialized send; worker panes are fleetd-owned. Residual collision corrupts a turn (recoverable), never the subscription boundary. See *Delivery gating & races*. |
|
||||
|
||||
## Related pages
|
||||
|
||||
|
||||
+31
-31
@@ -13,14 +13,14 @@ that is solved, on *how reliably you know when the worker is done or blocked*.
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Q{"How is the message<br/>delivered into a LIVE session?"}
|
||||
Q -->|"structured socket → terminal, + status events"| HD["herdr socket<br/>(via bridged)"]
|
||||
Q -->|"structured socket → terminal, + status events"| HD["herdr socket<br/>(via fleetd)"]
|
||||
Q -->|"HTTP request → terminal emulation"| AA["AgentAPI<br/>(coder/agentapi)"]
|
||||
Q -->|"streaming input generator (in-process)"| SDK["Agent SDK<br/>streaming query()"]
|
||||
Q -->|"worker PULLS on idle via a hook"| BUS["Message-queue<br/>+ Stop-hook long-poll"]
|
||||
Q -->|"raw keystrokes into the tmux pane"| TMUX["tmux send-keys<br/>/ PTY paste"]
|
||||
|
||||
HD --> V0["✅ our leading choice<br/>(status events + multiplex;<br/>symmetric single-host)"]
|
||||
AA --> V1["◐ fallback injector<br/>(swappable behind bridged)"]
|
||||
AA --> V1["◐ fallback injector<br/>(swappable behind fleetd)"]
|
||||
SDK --> V2["✅ if driver is our own code"]
|
||||
BUS --> V3["⚠ async bus events only<br/>(lands at turn boundary)"]
|
||||
TMUX --> V4["⚠ fragile — the raw primitive<br/>herdr/AgentAPI productize"]
|
||||
@@ -44,34 +44,34 @@ explicitly-requested, still-open feature:
|
||||
and [#24947 — `claude inject`](https://github.com/anthropics/claude-code/issues/24947).
|
||||
Every approach below is a way around that gap. **herdr wins because it productizes the
|
||||
sturdiest workaround (terminal automation) as a structured socket API that *also* streams
|
||||
agent-status events** — so `bridged` gets reliable completion/blocked signals instead of
|
||||
agent-status events** — so `fleetd` gets reliable completion/blocked signals instead of
|
||||
scraping a screen.
|
||||
|
||||
## 1. herdr socket API via `bridged` — structured injection + status events *(leading)*
|
||||
## 1. herdr socket API via `fleetd` — structured injection + status events *(leading)*
|
||||
|
||||
[herdr](https://herdr.dev) is a persistent agent multiplexer (a "tmux for agents") with a
|
||||
Unix-socket JSON API. `bridged` (see [Message Server](2-Message-Server)) drives it: `pane.send_text` +
|
||||
Unix-socket JSON API. `fleetd` (see [Message Server](2-Message-Server)) drives it: `pane.send_text` +
|
||||
`pane.send_keys` deliver a turn into the *running* pane, and `events.subscribe`
|
||||
(`pane.agent_status_changed`) reports **working / blocked / idle** as real events (turn-done
|
||||
= the `working → idle` edge; herdr has no `done` status).
|
||||
|
||||
- **Injects into a live session:** yes into the worker. **Worker → primary** rides `bridged`'s
|
||||
**MCP rendezvous** — the worker's `bridge_reply` (or the turn-done idle edge) resolves the primary's
|
||||
blocking `bridge_send` tool call, so no keystroke into the primary pane is needed, even
|
||||
- **Injects into a live session:** yes into the worker. **Worker → primary** rides `fleetd`'s
|
||||
**MCP rendezvous** — the worker's `fleet_reply` (or the turn-done idle edge) resolves the primary's
|
||||
blocking `fleet_send` tool call, so no keystroke into the primary pane is needed, even
|
||||
single-host. Fallbacks: herdr can type into a single-host non-MCP primary (subscription-safe
|
||||
keystrokes); a split-host primary wakes via its own `Stop`-hook polling `bridged` (the async
|
||||
keystrokes); a split-host primary wakes via its own `Stop`-hook polling `fleetd` (the async
|
||||
path — Mode 2 in [Architecture](1-Architecture)).
|
||||
- **North-face contract is MCP — and the sole gateway.** Both primary and workers mount
|
||||
`bridged` as an MCP server (one unified Claude setup); no Claude session ever addresses a
|
||||
`fleetd` as an MCP server (one unified Claude setup); no Claude session ever addresses a
|
||||
broker or peer directly. The herdr injection here is the *south* side, orthogonal to it.
|
||||
- **Completion signal:** structured events — not the screen-stability *guess* AgentAPI makes.
|
||||
The worker can even `pane.report_agent` its own state via herdr's `SKILL.md`.
|
||||
- **Multiplex + persist:** a herd of workers as addressable panes; headless server survives
|
||||
detach/reattach over SSH.
|
||||
- **Subscription-safe:** only the worker pane launches with `ANTHROPIC_BASE_URL`; `bridged`
|
||||
- **Subscription-safe:** only the worker pane launches with `ANTHROPIC_BASE_URL`; `fleetd`
|
||||
is a plain daemon (no quota) that enforces the boundary in code.
|
||||
- **Trade-off:** herdr's socket is **local-only** (`bridged`'s MCP/HTTP spans hosts, not
|
||||
herdr), and it is a young, single-dev project — so `bridged` keeps the injector **pluggable**
|
||||
- **Trade-off:** herdr's socket is **local-only** (`fleetd`'s MCP/HTTP spans hosts, not
|
||||
herdr), and it is a young, single-dev project — so `fleetd` keeps the injector **pluggable**
|
||||
and its durability in an **internal** queue behind the gateway. Replies are best carried as a
|
||||
structured envelope, not scraped.
|
||||
|
||||
@@ -81,13 +81,13 @@ Unix-socket JSON API. `bridged` (see [Message Server](2-Message-Server)) drives
|
||||
HTTP server and drives the CLI's terminal underneath (essentially a hardened, stateful
|
||||
`tmux send-keys` with parsing): `POST /message`, `GET /events` (SSE), `GET /status`. It was
|
||||
the original leading choice; herdr now supersedes it, but it remains a **swappable fallback
|
||||
injector** behind `bridged`'s interface.
|
||||
injector** behind `fleetd`'s interface.
|
||||
|
||||
- **Injects into a live session:** yes — but only the *worker* (it wraps one CLI); the
|
||||
primary direction still needs `bridged`'s async path (idle-injection, or a split-host
|
||||
`Stop`-hook polling `bridged`).
|
||||
primary direction still needs `fleetd`'s async path (idle-injection, or a split-host
|
||||
`Stop`-hook polling `fleetd`).
|
||||
- **Completion signal:** a **screen-stability heuristic**, not structured events.
|
||||
- **Cross-host:** native HTTP — its one edge over herdr, but `bridged` already provides the
|
||||
- **Cross-host:** native HTTP — its one edge over herdr, but `fleetd` already provides the
|
||||
HTTP layer on top of herdr, so that edge is neutralized.
|
||||
- **When to reach for it:** if herdr can't run, or as the second injector implementation to
|
||||
de-risk herdr's immaturity. Its `msgfmt` reply parser is worth reusing regardless.
|
||||
@@ -107,11 +107,11 @@ events and permission callbacks instead of scraping a terminal.
|
||||
|
||||
The **only pure-hooks** way to pull an external message into the **same** session. In
|
||||
`claude-bridge` this is **not** how a Claude session normally receives async work — under the
|
||||
sole-gateway rule `bridged` delivers async by **injecting an idle pane**, and no Claude session
|
||||
sole-gateway rule `fleetd` delivers async by **injecting an idle pane**, and no Claude session
|
||||
polls a queue. The Stop-hook survives in exactly one place: a **split-host primary** that isn't
|
||||
a herdr pane, where the hook long-polls **`bridged`** (not the queue) for wake-ups. The raw
|
||||
a herdr pane, where the hook long-polls **`fleetd`** (not the queue) for wake-ups. The raw
|
||||
mechanism below is shown for the comparison; note the poll target is the gateway, and the queue
|
||||
itself sits *behind* `bridged`.
|
||||
itself sits *behind* `fleetd`.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
@@ -142,7 +142,7 @@ sequenceDiagram
|
||||
[DIY pattern](https://claudefa.st/blog/tools/hooks/stop-hook-task-enforcement)).
|
||||
- **Decisive limitation:** **pull-on-idle only** — the message lands at a **turn boundary**,
|
||||
never mid-turn. Fine for "bus event wakes an idle worker"; wrong for "interrupt a busy
|
||||
worker." (`bridged` injecting into an idle pane has the same idle-boundary property, but
|
||||
worker." (`fleetd` injecting into an idle pane has the same idle-boundary property, but
|
||||
with a real status gate.)
|
||||
- **Lighter cousin:** a `UserPromptSubmit` hook returning `{"additionalContext":"…"}` —
|
||||
fires only *when a prompt is submitted*, so it can't deliver an async push.
|
||||
@@ -167,7 +167,7 @@ status events, and multiplexing — which is why we build on herdr rather than h
|
||||
|
||||
| Approach | Transport | Inject into running session? | Completion signal | Symmetric (both panes)? | Cross-host | Subscription-safe | Fragility |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **herdr via `bridged`** ✅ | socket → terminal + events | ✅ (idle-gated) | ✅ status events² | ◐ single-host¹ | via `bridged` MCP/HTTP (sole gateway) | ✅ (guard in code) | Low–Med (herdr young) |
|
||||
| **herdr via `fleetd`** ✅ | socket → terminal + events | ✅ (idle-gated) | ✅ status events² | ◐ single-host¹ | via `fleetd` MCP/HTTP (sole gateway) | ✅ (guard in code) | Low–Med (herdr young) |
|
||||
| **AgentAPI** ◐ (fallback) | HTTP → terminal emulation | ✅ (worker only) | ⚠ screen-stability | ❌ | ✅ native HTTP | ✅ (worker-only env) | Low |
|
||||
| **Agent SDK streaming** | in-process generator | ✅ | ✅ typed events | n/a | ✅ | ✅ | Low (driver = code) |
|
||||
| **Queue + Stop-hook** | hook long-poll | ✅ (turn boundary) | via turn end | ✅ (symmetric) | ✅ | ✅ | Medium |
|
||||
@@ -176,25 +176,25 @@ status events, and multiplexing — which is why we build on herdr rather than h
|
||||
|
||||
¹ **Symmetric only single-host.** herdr can type into *either* pane, but the primary is a
|
||||
herdr pane only when it runs on the herdr host. In the split-host target (primary on a Mac),
|
||||
worker→primary goes through `bridged` all the same — the primary's `Stop`-hook long-polls
|
||||
`bridged` (never a broker) for the wake-up. The "one mechanism, both directions via herdr"
|
||||
story holds for a single-box setup; across hosts the gateway is still `bridged`, just not by
|
||||
worker→primary goes through `fleetd` all the same — the primary's `Stop`-hook long-polls
|
||||
`fleetd` (never a broker) for the wake-up. The "one mechanism, both directions via herdr"
|
||||
story holds for a single-box setup; across hosts the gateway is still `fleetd`, just not by
|
||||
injection.
|
||||
|
||||
² **"Completion signal" for `bridged` is the timing signal; reply *content* rides the worker's
|
||||
structured `bridge_reply` (preferred), with a `Stop`-hook envelope as fallback** (see
|
||||
² **"Completion signal" for `fleetd` is the timing signal; reply *content* rides the worker's
|
||||
structured `fleet_reply` (preferred), with a `Stop`-hook envelope as fallback** (see
|
||||
[Message Server](2-Message-Server)) — not the status event itself.
|
||||
|
||||
## Recommendation
|
||||
|
||||
- **Primary Opus → worker (the bridge's main path):** **herdr via `bridged`** —
|
||||
- **Primary Opus → worker (the bridge's main path):** **herdr via `fleetd`** —
|
||||
status-gated injection, structured completion/blocked events, symmetric (single-host), multiplexed,
|
||||
persistent, with the subscription boundary enforced in code. Selected. See
|
||||
[Message Server](2-Message-Server) / [Architecture](1-Architecture).
|
||||
- **Keep AgentAPI as a swappable fallback injector** behind `bridged`'s interface, so
|
||||
- **Keep AgentAPI as a swappable fallback injector** behind `fleetd`'s interface, so
|
||||
herdr's immaturity is a de-riskable risk rather than a load-bearing one.
|
||||
- **External event bus → worker (async wake-ups):** the bus hits **`bridged`'s REST ingress**;
|
||||
`bridged` enqueues internally if needed and **injects the idle worker** — the worker runs no
|
||||
- **External event bus → worker (async wake-ups):** the bus hits **`fleetd`'s REST ingress**;
|
||||
`fleetd` enqueues internally if needed and **injects the idle worker** — the worker runs no
|
||||
queue-polling hook. Complementary to the sync path, not a replacement — different trigger
|
||||
shape, same single gateway.
|
||||
- **Avoid hand-rolled `tmux send-keys`** unless neither herdr nor AgentAPI can run; it's the
|
||||
|
||||
+1
-1
@@ -11,7 +11,7 @@
|
||||
- herdr, and why you check the **protocol number** rather than the version;
|
||||
- building and starting the daemon;
|
||||
- the **login shell** rule for `${SHARED_ENV}/tools/secrets.sh` — the single most expensive trap in
|
||||
bring-up, and why `scripts/bridged-launchd-wrapper.sh` exists;
|
||||
bring-up, and why `scripts/fleetd-launchd-wrapper.sh` exists;
|
||||
- the lead's **tab label**, which is how the daemon resolves who the lead is.
|
||||
|
||||
## The one rule that was already right on this page
|
||||
|
||||
+1
-1
@@ -7,7 +7,7 @@
|
||||
|
||||
**Go to [13 User Guide](13-User-Guide):**
|
||||
|
||||
- **§4 Run, and prove it runs** — `scripts/redeploy-bridged.sh`, why a merge is not a deployment,
|
||||
- **§4 Run, and prove it runs** — `scripts/redeploy-fleetd.sh`, why a merge is not a deployment,
|
||||
draining before a restart, and the four checks that go beyond `/healthz`.
|
||||
- **§6 When it breaks** — twelve traps hit for real this year, grouped by bring-up, losing a
|
||||
member's work, and merging a member's work.
|
||||
|
||||
+35
-35
@@ -1,23 +1,23 @@
|
||||
# 6. Team
|
||||
|
||||
The [Message Server](2-Message-Server) (`bridged`) delivers **one turn into one worker**. A
|
||||
The [Message Server](2-Message-Server) (`fleetd`) delivers **one turn into one worker**. A
|
||||
**team** is the layer above it: a **Claude team-lead** that fans a job out across a **mixed
|
||||
fleet** of workers — some on Claude, some on the remote local LLM — and reduces their replies.
|
||||
Same `bridged` delivery, same subscription boundary; this page is only about **orchestration**
|
||||
Same `fleetd` delivery, same subscription boundary; this page is only about **orchestration**
|
||||
— who the workers are, how the lead picks one, and how it runs many at once.
|
||||
|
||||
> Delivery mechanics (blocking `bridge_send` MCP call, status-gated reply) live in
|
||||
> Delivery mechanics (blocking `fleet_send` MCP call, status-gated reply) live in
|
||||
> [Message Server](2-Message-Server). Transport rationale is in [Approaches](3-Approaches).
|
||||
> This page assumes both.
|
||||
|
||||
## The team
|
||||
|
||||
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
|
||||
worker; an **MCP client of `bridged`** (it mounts the bridge like everyone else). It plans,
|
||||
routes, dispatches via `bridge_send`, and integrates — and never sets `ANTHROPIC_BASE_URL`.
|
||||
- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
|
||||
worker; an **MCP client of `fleetd`** (it mounts the bridge like everyone else). It plans,
|
||||
routes, dispatches via `fleet_send`, and integrates — and never sets `ANTHROPIC_BASE_URL`.
|
||||
- **Workers** — a herd of `claude` panes in herdr, each an addressable `fleetd` session
|
||||
with its **own model/env**, each also **mounting the bridge MCP** (unified setup — they
|
||||
reply via `bridge_reply`):
|
||||
reply via `fleet_reply`):
|
||||
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
|
||||
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
|
||||
embarrassingly parallel subtasks.
|
||||
@@ -31,7 +31,7 @@ alike. Scale each kind horizontally by adding panes.
|
||||
```mermaid
|
||||
flowchart TB
|
||||
LEAD["lead — Opus<br/>(Claude Code, env CLEAN)<br/>MCP client"]
|
||||
subgraph BD["bridged — standalone daemon"]
|
||||
subgraph BD["fleetd — standalone daemon"]
|
||||
SRV["SERVER face<br/>MCP · REST/SSE · policy brain"]
|
||||
CLI["CLIENT face<br/>herdr socket"]
|
||||
SRV --> CLI
|
||||
@@ -44,11 +44,11 @@ flowchart TB
|
||||
ANT["api.anthropic.com<br/>(Pro/Max)"]
|
||||
OLL["ollama.ltms.dev<br/>(local model)"]
|
||||
|
||||
LEAD -->|"MCP bridge_send (target role)"| SRV
|
||||
LEAD -->|"MCP fleet_send (target role)"| SRV
|
||||
CLI -->|"Unix socket · send_text · events.subscribe"| HERDR
|
||||
HERDR --> WC1 & WC2 & WL1 & WL2
|
||||
WC1 -.->|"MCP bridge_reply"| SRV
|
||||
WL1 -.->|"MCP bridge_reply"| SRV
|
||||
WC1 -.->|"MCP fleet_reply"| SRV
|
||||
WL1 -.->|"MCP fleet_reply"| SRV
|
||||
WC1 -->|"inference"| ANT
|
||||
WC2 -->|"inference"| ANT
|
||||
WL1 -->|"inference"| OLL
|
||||
@@ -71,14 +71,14 @@ flowchart TB
|
||||
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
|
||||
|
||||
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
|
||||
selection is **policy in the lead**, not a `bridged` concern — `bridged` just delivers to
|
||||
selection is **policy in the lead**, not a `fleetd` concern — `fleetd` just delivers to
|
||||
the session the lead names.
|
||||
|
||||
## Subscription boundary in a team
|
||||
|
||||
Unchanged from [Architecture](1-Architecture), and it scales with the fleet: **only local-worker panes**
|
||||
launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on the
|
||||
subscription. `bridged` enforces which panes may carry the off-subscription env, so adding
|
||||
subscription. `fleetd` enforces which panes may carry the off-subscription env, so adding
|
||||
workers never widens the boundary.
|
||||
|
||||
## Parallel fan-out (map / reduce)
|
||||
@@ -89,48 +89,48 @@ different workers at once, then results are gathered.
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant L as "lead (Opus)"
|
||||
participant B as "bridged"
|
||||
participant B as "fleetd"
|
||||
participant WC as "w-claude-1"
|
||||
participant WL as "w-local-1"
|
||||
|
||||
Note over L: "split job → subtask A (reasoning), subtask B (bulk)"
|
||||
par A → Claude worker
|
||||
L->>B: "bridge_send {role: w-claude, prompt: A}"
|
||||
L->>B: "fleet_send {role: w-claude, prompt: A}"
|
||||
B->>WC: "send_text into running pane"
|
||||
WC-->>B: "bridge_reply (or idle edge)"
|
||||
WC-->>B: "fleet_reply (or idle edge)"
|
||||
B-->>L: "tool result reply A"
|
||||
and B → local worker
|
||||
L->>B: "bridge_send {role: w-local, prompt: B}"
|
||||
L->>B: "fleet_send {role: w-local, prompt: B}"
|
||||
B->>WL: "send_text into running pane"
|
||||
WL-->>B: "bridge_reply (or idle edge)"
|
||||
WL-->>B: "fleet_reply (or idle edge)"
|
||||
B-->>L: "tool result reply B"
|
||||
end
|
||||
Note over L: "reduce → integrate A + B into final answer"
|
||||
```
|
||||
|
||||
- **Map:** the lead issues N concurrent blocking `bridge_send` tool calls (one per subtask →
|
||||
its chosen worker). Each call blocks only *that* request; `bridged` holds it open until the
|
||||
- **Map:** the lead issues N concurrent blocking `fleet_send` tool calls (one per subtask →
|
||||
its chosen worker). Each call blocks only *that* request; `fleetd` holds it open until the
|
||||
worker's turn completes (status-gated) and returns the reply as the tool result.
|
||||
- **Reduce:** the lead collects the N replies and integrates. A slow local worker never
|
||||
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
|
||||
- **Detached / long jobs** use `bridged`'s async path instead of a held request — the result
|
||||
is delivered when ready by `bridged` injecting the lead's idle pane (Mode 2 in
|
||||
[Architecture](1-Architecture)). The lead talks only to `bridged`, never a broker, and never
|
||||
- **Detached / long jobs** use `fleetd`'s async path instead of a held request — the result
|
||||
is delivered when ready by `fleetd` injecting the lead's idle pane (Mode 2 in
|
||||
[Architecture](1-Architecture)). The lead talks only to `fleetd`, never a broker, and never
|
||||
busy-polls across turns.
|
||||
|
||||
Fan-out is bounded by the herd size (pane count) and `bridged`'s concurrency policy, not by
|
||||
Fan-out is bounded by the herd size (pane count) and `fleetd`'s concurrency policy, not by
|
||||
the lead.
|
||||
|
||||
## Knowing the roster
|
||||
|
||||
The lead discovers its team from `bridged` (session list / roles) rather than hard-coding
|
||||
The lead discovers its team from `fleetd` (session list / roles) rather than hard-coding
|
||||
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
|
||||
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
|
||||
|
||||
```markdown
|
||||
## Your team (via bridged)
|
||||
## Your team (via fleetd)
|
||||
You are the team-lead. Delegate through the bridge MCP tools — never launch workers yourself.
|
||||
Roster: call bridge_list for current sessions/roles.
|
||||
Roster: call fleet_list for current sessions/roles.
|
||||
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
|
||||
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
|
||||
|
||||
@@ -139,27 +139,27 @@ of them (concurrent blocking sends), THEN gather — never serialize independent
|
||||
Integrate the reply envelopes; you own the final answer.
|
||||
```
|
||||
|
||||
`bridge_send` **is** the one tool call — no HTTP to hand-roll. Optionally wrap it in a Claude
|
||||
`fleet_send` **is** the one tool call — no HTTP to hand-roll. Optionally wrap it in a Claude
|
||||
Code skill (`/delegate <role> "<task>"`) for ergonomics.
|
||||
|
||||
## What this layer does NOT change
|
||||
|
||||
- **Delivery** is still `bridged` → herdr `pane.send_text` + status events ([Message Server](2-Message-Server)).
|
||||
- **Delivery** is still `fleetd` → herdr `pane.send_text` + status events ([Message Server](2-Message-Server)).
|
||||
- **Completion timing** is still the worker status event; **reply content** rides the worker's
|
||||
`bridge_reply` (or a `Stop`-hook envelope for a herdr-only worker).
|
||||
`fleet_reply` (or a `Stop`-hook envelope for a herdr-only worker).
|
||||
- **Single-host** still applies: herdr's socket is local, so the whole herd lives on the
|
||||
`bridged` host. The lead may be remote — it only needs to reach `bridged`'s MCP endpoint.
|
||||
`fleetd` host. The lead may be remote — it only needs to reach `fleetd`'s MCP endpoint.
|
||||
|
||||
## Open questions
|
||||
|
||||
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `bridged` role-router
|
||||
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `fleetd` role-router
|
||||
(label-based). Start with the former; promote to the latter if routing logic grows.
|
||||
- **Backpressure:** per-role concurrency caps in `bridged` so a fan-out can't exhaust the
|
||||
- **Backpressure:** per-role concurrency caps in `fleetd` so a fan-out can't exhaust the
|
||||
local gateway.
|
||||
- **Result schema:** whether `bridge_reply` payloads should carry structured metadata (worker,
|
||||
- **Result schema:** whether `fleet_reply` payloads should carry structured metadata (worker,
|
||||
model, tokens) to help the lead's reduce step.
|
||||
|
||||
## Status
|
||||
|
||||
🟡 Design (2026-07-11). Orchestration layer over the selected `bridged` server; inherits
|
||||
🟡 Design (2026-07-11). Orchestration layer over the selected `fleetd` server; inherits
|
||||
herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see [Message Server](2-Message-Server).
|
||||
|
||||
+46
-46
@@ -7,7 +7,7 @@ exercises every mechanism the system needs (trigger, discovery, ccs spawn, the I
|
||||
and worker lifecycle), so we design it in full, then catalogue the rest.
|
||||
|
||||
Everything below holds the invariants: the primary stays env-CLEAN on Pro/Max, the worker's
|
||||
model/account come from a **ccs profile**, and both talk only to `bridged` over MCP.
|
||||
model/account come from a **ccs profile**, and both talk only to `fleetd` over MCP.
|
||||
|
||||
## Flagship — a review conversation (Opus ↔ gx00 reviewer)
|
||||
|
||||
@@ -18,42 +18,42 @@ vLLM model. Opus never leaves its subscription; it just calls MCP tools.
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant O as "Opus (primary, env CLEAN)"
|
||||
participant B as "bridged (gateway)"
|
||||
participant B as "fleetd (gateway)"
|
||||
participant C as "ccs + herdr"
|
||||
participant W as "worker · gx00-vllm"
|
||||
|
||||
O->>B: "bridge_list() — what workers can I use?"
|
||||
O->>B: "fleet_list() — what workers can I use?"
|
||||
B-->>O: "profiles:[gx00-vllm=DeepSeek, …] · sessions:[]"
|
||||
O->>B: "bridge_send(to: reviewer@gx00-vllm, {kind: review.request, diff, focus})"
|
||||
O->>B: "fleet_send(to: reviewer@gx00-vllm, {kind: review.request, diff, focus})"
|
||||
Note over B: "guard: ccs env gx00-vllm → base_url host on allowlist ✓"
|
||||
B->>C: "spawn: ccs gx00-vllm claude (new herdr pane)"
|
||||
C->>W: "worker Ready → inject review.request"
|
||||
activate W
|
||||
W->>B: "bridge_reply({kind: review.reply, findings:[…]})"
|
||||
W->>B: "fleet_reply({kind: review.reply, findings:[…]})"
|
||||
deactivate W
|
||||
B-->>O: "tool result = findings"
|
||||
Note over O: "reads findings, wants to dig into one"
|
||||
O->>B: "bridge_send(session s_7f3a, {kind: question — why finding 3 high-sev})"
|
||||
O->>B: "fleet_send(session s_7f3a, {kind: question — why finding 3 high-sev})"
|
||||
B->>W: "inject into the SAME reviewer pane (context still warm)"
|
||||
activate W
|
||||
W->>B: "bridge_reply({kind: answer, …})"
|
||||
W->>B: "fleet_reply({kind: answer, …})"
|
||||
deactivate W
|
||||
B-->>O: "tool result = answer"
|
||||
Note over O,W: "same reviewer session reused across turns → the diff stays in its context"
|
||||
```
|
||||
|
||||
*Figure: discovery → trigger → ccs-spawn (guarded) → structured reply → follow-up on the same
|
||||
warm session. Every arrow from Opus is an MCP tool call to `bridged`; the worker's provider
|
||||
warm session. Every arrow from Opus is an MCP tool call to `fleetd`; the worker's provider
|
||||
routing lives entirely in its ccs profile.*
|
||||
|
||||
The rest of this page is the five mechanisms this scenario needs.
|
||||
|
||||
## Mechanism 1 — triggering a subtask (`bridge_send`)
|
||||
## Mechanism 1 — triggering a subtask (`fleet_send`)
|
||||
|
||||
Opus delegates with **one** tool call:
|
||||
|
||||
```jsonc
|
||||
bridge_send({
|
||||
fleet_send({
|
||||
"to": "reviewer@gx00-vllm", // role@profile, or a live session id
|
||||
"kind": "review.request",
|
||||
"body": { "workspace": "/repo", "base": "main", "head": "HEAD",
|
||||
@@ -62,20 +62,20 @@ bridge_send({
|
||||
})
|
||||
```
|
||||
|
||||
`bridged` resolves the target (spawn-or-reuse, below), injects the turn into the worker's
|
||||
`fleetd` resolves the target (spawn-or-reuse, below), injects the turn into the worker's
|
||||
herdr pane gated on `agent_status`, and — with `block:true` (the default) — holds the call
|
||||
open until the reply lands (worker `bridge_reply` or the turn-done `working→idle` edge),
|
||||
open until the reply lands (worker `fleet_reply` or the turn-done `working→idle` edge),
|
||||
returning it as the tool result.
|
||||
Review turns are short, so they block; a long/detached job would use `block:false` and come
|
||||
back via idle-pane injection ([Mode 2](1-Architecture#traffic-two-modes-across-the-gateway)).
|
||||
|
||||
## Mechanism 2 — worker discovery (knowing your choices)
|
||||
|
||||
Opus shouldn't hard-code pane ids or guess what's available. `bridge_list()` returns both
|
||||
Opus shouldn't hard-code pane ids or guess what's available. `fleet_list()` returns both
|
||||
what's **runnable** and what's **live**:
|
||||
|
||||
```jsonc
|
||||
bridge_list() → {
|
||||
fleet_list() → {
|
||||
"profiles": [ // spawnable = the ccs worker roster (Mechanism 3)
|
||||
{ "name": "gx00-vllm", "model": "DeepSeek-V3", "host": "gx00.ltms.dev", "status": "available" },
|
||||
{ "name": "ollama-local","model": "llama3.1", "host": "ollama.ltms.dev","status": "available" }
|
||||
@@ -86,32 +86,32 @@ bridge_list() → {
|
||||
}
|
||||
```
|
||||
|
||||
This is the "let Opus know its worker choices" surface: the catalogue is `bridged`'s configured
|
||||
This is the "let Opus know its worker choices" surface: the catalogue is `fleetd`'s configured
|
||||
roster of ccs worker profiles, and the live list is what it's already running.
|
||||
|
||||
## Mechanism 3 — spawning via ccs profiles (the flexibility)
|
||||
|
||||
**A worker's identity *is* a ccs profile.** `ccs <profile> [claude-args…]` launches `claude`
|
||||
with that profile's account and provider routing, so `bridged` never hand-assembles env — it
|
||||
with that profile's account and provider routing, so `fleetd` never hand-assembles env — it
|
||||
just picks a profile:
|
||||
|
||||
- **Config.** `bridged.yaml` lists worker profiles by name; each maps to a ccs profile
|
||||
- **Config.** `fleetd.yaml` lists worker profiles by name; each maps to a ccs profile
|
||||
(`ccs api` profile pointing at GX10 vLLM / Ollama), an expected model, and an allowlisted
|
||||
`base_url` host.
|
||||
- **Spawn.** `bridged` tells herdr to open a pane and `send_text`: `ccs gx00-vllm claude`
|
||||
- **Spawn.** `fleetd` tells herdr to open a pane and `send_text`: `ccs gx00-vllm claude`
|
||||
(plus flags — workspace dir, an injected reviewer system prompt). No `ANTHROPIC_BASE_URL=…`
|
||||
prefix; the profile carries it.
|
||||
- **Guard (subscription boundary, in ccs terms).** Before spawning, `bridged` runs
|
||||
- **Guard (subscription boundary, in ccs terms).** Before spawning, `fleetd` runs
|
||||
`ccs env <profile>` and validates the **resolved** `ANTHROPIC_BASE_URL` host against the
|
||||
off-subscription allowlist. A profile that resolves to `api.anthropic.com` (a subscription
|
||||
profile) is **refused as a worker** — that would burn your quota. The primary Opus is *your*
|
||||
session on *your* subscription profile; `bridged` never spawns it.
|
||||
session on *your* subscription profile; `fleetd` never spawns it.
|
||||
- **Swap = repoint.** Changing the reviewer's model is choosing a different ccs profile — no
|
||||
`bridged` code change. Add a profile → it appears in `bridge_list().profiles`.
|
||||
`fleetd` code change. Add a profile → it appears in `fleet_list().profiles`.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
CFG["bridged.yaml<br/>worker profiles"] --> PICK["pick profile<br/>gx00-vllm"]
|
||||
CFG["fleetd.yaml<br/>worker profiles"] --> PICK["pick profile<br/>gx00-vllm"]
|
||||
PICK --> GUARD{"ccs env host<br/>on allowlist?"}
|
||||
GUARD -->|"no (api.anthropic.com)"| REJ["refuse — would burn subscription"]
|
||||
GUARD -->|"yes (gx00.ltms.dev)"| SPAWN["herdr: ccs gx00-vllm claude"]
|
||||
@@ -135,9 +135,9 @@ Every message across the gateway is one envelope. It answers two questions each
|
||||
| `v` | envelope version (`1`) |
|
||||
| `from` | sender identity — `primary:opus`, or `role@profile` / session id for a worker |
|
||||
| `to` | recipient — `role@profile` or a live session id |
|
||||
| `session` | `bridged` session id (the worker session), e.g. `s_7f3a` |
|
||||
| `session` | `fleetd` session id (the worker session), e.g. `s_7f3a` |
|
||||
| `turn` | monotonic counter within the session |
|
||||
| `corr` | correlation id (`session#turn`) — `bridged`'s rendezvous matches a reply to its request |
|
||||
| `corr` | correlation id (`session#turn`) — `fleetd`'s rendezvous matches a reply to its request |
|
||||
| `kind` | **verb.noun** telling the recipient what to do (vocabulary below) |
|
||||
| `body` | kind-specific payload |
|
||||
|
||||
@@ -148,11 +148,11 @@ Every message across the gateway is one envelope. It answers two questions each
|
||||
| `review.request` | primary → worker | `{workspace, base, head\|diff, files?, focus[], instructions}` |
|
||||
| `review.reply` | worker → primary | `{summary, findings:[{file,line,severity,issue,suggestion}], verdict}` |
|
||||
| `question` / `answer` | either way | free-form follow-up tied to the same `session` |
|
||||
| `ask` | worker → primary | worker-initiated blocker (needs a decision) — via `bridge_ask` |
|
||||
| `ask` | worker → primary | worker-initiated blocker (needs a decision) — via `fleet_ask` |
|
||||
| `ack` / `status` | control | delivery/liveness, no new turn |
|
||||
|
||||
The worker learns this contract from a **reviewer skill / `CLAUDE.md` snippet** injected at
|
||||
spawn ("you are a reviewer; requests arrive as `review.request`; reply with `bridge_reply`
|
||||
spawn ("you are a reviewer; requests arrive as `review.request`; reply with `fleet_reply`
|
||||
`kind: review.reply`"). So both ends know the sender and the required next action without a
|
||||
held-open conversation — one request in, one structured reply out.
|
||||
|
||||
@@ -164,30 +164,30 @@ stays warm across follow-ups, and recyclable so it never outgrows its context wi
|
||||
- **Spawn-on-demand** — first `review.request` for a workspace spawns the profile's worker.
|
||||
- **Reuse** — subsequent turns (`question`, next-file review) target the same `session`; its
|
||||
context carries the code under review.
|
||||
- **Recycle (Ralph loop)** — on a context/idle cap `bridged` checkpoints (commit + `STATE.md`)
|
||||
- **Recycle (Ralph loop)** — on a context/idle cap `fleetd` checkpoints (commit + `STATE.md`)
|
||||
and respawns fresh — **not** `claude --resume`. See
|
||||
[Architecture → Worker lifecycle](1-Architecture#worker-lifecycle--the-ralph-loop).
|
||||
- **Drain** — an `idle_ttl` (e.g. 20 min idle) or workspace close tears the pane down.
|
||||
|
||||
Policy knobs in `bridged.yaml`: `idle_ttl`, `context_cap`, `max_workers_per_profile`.
|
||||
Policy knobs in `fleetd.yaml`: `idle_ttl`, `context_cap`, `max_workers_per_profile`.
|
||||
|
||||
## Primary-side directive — *when* to delegate (a `CLAUDE.md` snippet)
|
||||
|
||||
Mechanism 4 gives the **worker** a `CLAUDE.md` snippet so it knows how to answer. The **primary**
|
||||
needs the mirror image: a standing reminder to *reach for the bridge in the first place* instead of
|
||||
spending subscription tokens on work a cheaper worker could do. The trigger is an environment signal
|
||||
— the `bridged` MCP tools being connected (e.g. a `BRIDGED_MCP_URL` marker in the primary's env).
|
||||
— the `fleetd` MCP tools being connected (e.g. a `FLEETD_MCP_URL` marker in the primary's env).
|
||||
When that property is set, this instance is a **bridge primary** and should delegate by default.
|
||||
|
||||
The decision the directive encodes:
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
T["a task arrives"] --> Q{"bridge available?<br/>(BRIDGED_MCP_URL set /<br/>bridged MCP connected)"}
|
||||
T["a task arrives"] --> Q{"bridge available?<br/>(FLEETD_MCP_URL set /<br/>fleetd MCP connected)"}
|
||||
Q -->|"no"| SELF["do it on the primary"]
|
||||
Q -->|"yes"| J{"needs YOUR judgment,<br/>or bulk / mechanical / parallel?"}
|
||||
J -->|"judgment / interactive"| SELF
|
||||
J -->|"bulk / mechanical / parallel"| DEL["bridge_send → worker"]
|
||||
J -->|"bulk / mechanical / parallel"| DEL["fleet_send → worker"]
|
||||
classDef self fill:#2f855a,stroke:#22543d,color:#ffffff;
|
||||
classDef del fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
|
||||
class SELF self
|
||||
@@ -197,7 +197,7 @@ flowchart TD
|
||||
*Figure: an env property (bridge available) flips the default from "do it myself" to "delegate unless
|
||||
it needs my judgment."*
|
||||
|
||||
This is a *suggestion*, not wiring: `bridged` never edits an agent's `CLAUDE.md` (that would cross
|
||||
This is a *suggestion*, not wiring: `fleetd` never edits an agent's `CLAUDE.md` (that would cross
|
||||
the subscription boundary in the wrong direction). The operator writes it.
|
||||
|
||||
### As-built — the `CLAUDE.md` **Bridge communication** section
|
||||
@@ -212,7 +212,7 @@ Why in `CLAUDE.md` and not somewhere else — the four layers, each with a diffe
|
||||
|
||||
| Layer | Carries | Reaches | Cost to the reader |
|
||||
|---|---|---|---|
|
||||
| `REPLY_CHARTER` (`ClaudeCodeLauncher` / `OpenCodeLauncher`) | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, both peer kinds | always in the system prompt |
|
||||
| `REPLY_CHARTER` (`ClaudeCodeLauncher` / `OpenCodeLauncher`) | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every worker, at launch, both peer kinds | always in the system prompt |
|
||||
| **`CLAUDE.md` → Bridge communication** | protocol invariants + orchestration policy | primary **and** every claude-code worker — tracked in git, so worktrees get it free | always in context |
|
||||
| `.claude/skills/{implementer,reviewer}` | per-job procedure: commit/push/PR recipe, finding format | a worker told to load it | on demand |
|
||||
| `docs/MCP-Contract.md` | design detail, flows, error model | anyone who goes looking | on demand |
|
||||
@@ -222,11 +222,11 @@ must obey it.** Duplicating a rule into a skill is how the skill and the charter
|
||||
|
||||
What each part of the section pins down:
|
||||
|
||||
- **Role identification** — **`bridge_whoami`** (below) answers it authoritatively; the section
|
||||
- **Role identification** — **`fleet_whoami`** (below) answers it authoritatively; the section
|
||||
tells the reader to call it rather than infer. A fallback ladder remains for when it is
|
||||
unreachable: the charter in the system prompt (reliable — the launcher appends it in the same
|
||||
branch that mounts the MCP, so bridge tools without a charter is not a reachable state); the mount
|
||||
name (`mcp__bridged__*` for the primary's `.mcp.json` vs `mcp__bridge__*` for a worker's inline
|
||||
name (`mcp__fleetd__*` for the primary's `.mcp.json` vs `mcp__bridge__*` for a worker's inline
|
||||
config); `ANTHROPIC_BASE_URL` (one-way — Claude-model workers run clean, so absence proves
|
||||
nothing); then **fail toward worker**. The two errors are asymmetric: a primary acting as a worker
|
||||
gets refused by the authz gate — loud and self-correcting — while a worker acting as the primary
|
||||
@@ -241,27 +241,27 @@ What each part of the section pins down:
|
||||
turn costs a worker turn, while doing it yourself costs the primary's context and subscription.
|
||||
Independent units fan out (one worktree worker each, all dispatched `wait:false`, then poll)
|
||||
rather than serializing. Plus the policy the tool descriptions can't carry: pass `profile:`
|
||||
explicitly; prefer `wait:false` + `bridge_poll`, since a blocking `bridge_send` is capped by the
|
||||
explicitly; prefer `wait:false` + `fleet_poll`, since a blocking `fleet_send` is capped by the
|
||||
*caller's own* MCP client timeout (~60s) long before a real task finishes; make every delegation
|
||||
self-contained; **name the worker's skill in the first line of `content`** — that instruction is
|
||||
what turns an opt-in skill into a reliable one; you are the merge gate; verify what a worker
|
||||
claims rather than trusting a "clean" report. Delegating work never delegates responsibility.
|
||||
- **Worker** — the turn contract: load the named skill, stay in scope, `bridge_ask` only for a
|
||||
decision that is genuinely the lead's, end with exactly one `bridge_reply`, report only what you
|
||||
- **Worker** — the turn contract: load the named skill, stay in scope, `fleet_ask` only for a
|
||||
decision that is genuinely the lead's, end with exactly one `fleet_reply`, report only what you
|
||||
actually ran, never merge, never commit `.mcp.json` or `wiki/`.
|
||||
|
||||
**Known gap:** *opencode* workers never read `CLAUDE.md` — they receive `REPLY_CHARTER` as an
|
||||
instructions file and nothing else. Any rule a non-Claude peer must obey belongs in the charter, not
|
||||
in this section. The charter currently carries only the reply rule.
|
||||
|
||||
### `bridge_whoami` — asking instead of guessing
|
||||
### `fleet_whoami` — asking instead of guessing
|
||||
|
||||
**Shipped.** The daemon always knew the answer: `ConnectionIdentity` maps a call's loopback peer PID
|
||||
to a herdr pane, and every tool call is already gated on the `Principal` it yields. What was missing
|
||||
was any way for an agent to *ask* — so an agent's own role had to be inferred from side channels the
|
||||
daemon does not control, with a silent failure mode when the inference went the wrong way.
|
||||
|
||||
`bridge_whoami` (no params, `READ` in the [authz table](9-Implementation)) returns that same
|
||||
`fleet_whoami` (no params, `READ` in the [authz table](9-Implementation)) returns that same
|
||||
resolved identity as data:
|
||||
|
||||
```json
|
||||
@@ -270,7 +270,7 @@ resolved identity as data:
|
||||
```
|
||||
|
||||
The primary gets `{"role":"primary"}` and nothing more — deliberately: handing it a `sessionId` it
|
||||
does not own would invite exactly the forged `bridge_reply` that `Authz` refuses. A worker the
|
||||
does not own would invite exactly the forged `fleet_reply` that `Authz` refuses. A worker the
|
||||
session registry has no record of — one that outlived a daemon restart — still gets `role` and
|
||||
`sessionId`, which is the load-bearing part; the registry fields are simply absent rather than
|
||||
invented.
|
||||
@@ -298,7 +298,7 @@ notes after the block). Improvements land *here* first, then propagate to each p
|
||||
|
||||
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
|
||||
|
||||
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
|
||||
`fleetd` is the **sole communication gateway** between agents here. The orchestrating session (the
|
||||
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
|
||||
through its `fleet_*` tools. No session addresses a peer, a broker, or the network directly.
|
||||
|
||||
@@ -479,7 +479,7 @@ Two rules keep this portable, and both were learned by getting them wrong first:
|
||||
meaningless in another project. Each moved to a **Project addendum** section that sits *below*
|
||||
the block and never interleaves with it, so the block can be replaced wholesale without reading it.
|
||||
2. **Every fallback signal must be one-way.** The role ladder originally read the MCP mount name in
|
||||
both directions — `mcp__bridged__*` ⇒ primary, `mcp__bridge__*` ⇒ worker. Only the second half is
|
||||
both directions — `mcp__fleetd__*` ⇒ primary, `mcp__bridge__*` ⇒ worker. Only the second half is
|
||||
real: the launcher hard-codes `bridge` for a worker's inline config, while a primary's mount is
|
||||
named by whoever wrote that project's `.mcp.json`. A two-way reading of a one-way signal is a
|
||||
confident wrong answer, so the block states only the direction that holds.
|
||||
@@ -498,15 +498,15 @@ Same machinery, different `kind`/lifecycle:
|
||||
|
||||
| Use case | Shape |
|
||||
|---|---|
|
||||
| **Delegated refactor / codegen** | `bridge_send(kind: task.request)`, blocking; worker edits + commits on a branch; reply = diff summary. Opus reviews/merges. |
|
||||
| **Delegated refactor / codegen** | `fleet_send(kind: task.request)`, blocking; worker edits + commits on a branch; reply = diff summary. Opus reviews/merges. |
|
||||
| **Test writing / log triage** | Bulk, cheap, parallelizable — a local worker profile; fire several concurrently and reduce. |
|
||||
| **Parallel multi-file review** | A **fleet** of workers, one per area, fan-out/gather — see [Team](6-Team). |
|
||||
| **Perpetual worker** | Long-running across many Ralph recycles, state on disk; woken by async injection ([Mode 2](1-Architecture)). |
|
||||
| **Event-bus automation** | A webhook hits `bridged`'s REST ingress → injects an idle worker → result back to Opus or a chat bridge. No human in the loop. |
|
||||
| **Event-bus automation** | A webhook hits `fleetd`'s REST ingress → injects an idle worker → result back to Opus or a chat bridge. No human in the loop. |
|
||||
|
||||
## Related
|
||||
|
||||
- **[Architecture](1-Architecture)** — invariants, two modes, lifecycle state machine.
|
||||
- **[Message Server](2-Message-Server)** — the `bridged` design these mechanisms live in.
|
||||
- **[Message Server](2-Message-Server)** — the `fleetd` design these mechanisms live in.
|
||||
- **[Roadmap](8-Roadmap)** — the staged plan + tickets that build this review scenario first.
|
||||
- **[Team](6-Team)** — the fleet/orchestration layer above a single review.
|
||||
|
||||
+63
-63
@@ -28,9 +28,9 @@ gantt
|
||||
|
||||
| Stage | Goal | Delivers (usable outcome) |
|
||||
|---|---|---|
|
||||
| **1 — Walking skeleton** | One review, happy path | Opus mounts `bridged` (MCP), calls `bridge_send` with a diff, gets a review back from a real `ccs gx00-vllm claude` worker. Single hardcoded profile, same host, no guard/lifecycle. |
|
||||
| **2 — Contract + guard** | Trust the reply, trust the boundary | Structured [envelope](7-Use-Cases#mechanism-4--the-id-contract-envelope) + worker `bridge_reply`; reply rendezvous; subscription guard via `ccs env`; reviewer skill. |
|
||||
| **3 — Lifecycle + discovery** | Reuse, recycle, choose | Session manager (spawn/reuse/recycle Ralph loop, `idle_ttl`); `bridge_list` roster+live; multiple profiles. (`bridge_ask` landed early in Stage 2.) |
|
||||
| **1 — Walking skeleton** | One review, happy path | Opus mounts `fleetd` (MCP), calls `fleet_send` with a diff, gets a review back from a real `ccs gx00-vllm claude` worker. Single hardcoded profile, same host, no guard/lifecycle. |
|
||||
| **2 — Contract + guard** | Trust the reply, trust the boundary | Structured [envelope](7-Use-Cases#mechanism-4--the-id-contract-envelope) + worker `fleet_reply`; reply rendezvous; subscription guard via `ccs env`; reviewer skill. |
|
||||
| **3 — Lifecycle + discovery** | Reuse, recycle, choose | Session manager (spawn/reuse/recycle Ralph loop, `idle_ttl`); `fleet_list` roster+live; multiple profiles. (`fleet_ask` landed early in Stage 2.) |
|
||||
| **4 — Pluggable peers** | PeerLauncher SPI + two adapters | ✅ Extract `PeerLauncher` SPI in-tree; `ClaudeCodeLauncher` (Stage A) **and** `OpenCodeLauncher` (Stage B, CB-402) routed by a `kind:` discriminator through `CompositePeerLauncher`. Stage C (dynamic external plugin loading) remains future work, gated by a trust/capability model. |
|
||||
| **5 — Harden** | Production shape | ✅ Bearer auth + fail-fast exposure guard, `/metrics` + `/healthz`, mock-socket CI on the Gitea runner, launchd + systemd units, per-session authz + audit log. TLS deliberately terminates at a reverse proxy, not in the daemon. |
|
||||
|
||||
@@ -43,10 +43,10 @@ Consolidated from [Message Server](2-Message-Server#proposed-tech-stack); the cc
|
||||
| **Language** | **Java 21+ (virtual threads)** | Loom fits the blocking socket + MCP + queue + SSE fan-in; **GraalVM `native-image`** recovers the `scp`+systemd single-binary deploy. Kotlin OK (same JVM). |
|
||||
| **REST API core** | **Javalin** (or Spring MVC) + SSE | **The implementation & testability surface** — every feature is a REST endpoint (see [Testability](#testability--the-rest-api-is-the-contract-surface)). |
|
||||
| **MCP server** | **Official MCP Java SDK**, streamable-HTTP (Jetty/Spring) | A **thin adapter over the REST core** — the Claude-facing face of the same features. |
|
||||
| **herdr client** | JDK `UnixDomainSocketAddress` + `SocketChannel`, NDJSON (Jackson) | Native UDS, no dep; event stream on a virtual thread. Pinned to herdr **0.7.0 / protocol 14** ([verified API](2-Message-Server#the-herdr-control-contract-what-bridged-drives)). |
|
||||
| **herdr client** | JDK `UnixDomainSocketAddress` + `SocketChannel`, NDJSON (Jackson) | Native UDS, no dep; event stream on a virtual thread. Pinned to herdr **0.7.0 / protocol 14** ([verified API](2-Message-Server#the-herdr-control-contract-what-fleetd-drives)). |
|
||||
| **Worker spawn** | **`ccs <profile> claude`** into a herdr pane | Profile = worker identity; no env-prefix. |
|
||||
| **Subscription guard** | **`ccs env <profile>`** → resolved `base_url` host allowlist | Boundary check in ccs terms. |
|
||||
| **Config** | **YAML via Jackson** + env | `bridged.yaml`: worker profiles, allowlist, bind addr, lifecycle knobs. |
|
||||
| **Config** | **YAML via Jackson** + env | `fleetd.yaml`: worker profiles, allowlist, bind addr, lifecycle knobs. |
|
||||
| **Internal queue** *(future)* | **Redis Streams via Lettuce** (ack + visibility) | Below the gateway; in-JVM queue OK single-host. |
|
||||
| **Observability** | **SLF4J+Logback**; **Micrometer**→Prometheus `/metrics`; `/healthz` | Stage 5. |
|
||||
| **Supervision** | systemd unit (`java -jar` or native-image), ordered after herdr | Stage 5. |
|
||||
@@ -97,12 +97,12 @@ flowchart TB
|
||||
|
||||
| Feature | REST endpoint | MCP tool | Acceptance test |
|
||||
|---|---|---|---|
|
||||
| Deliver a turn (blocking) | `POST /sessions/{id}/message` | `bridge_send` | reply returned; `working→idle` unblocks; timeout → 202 |
|
||||
| Worker status | `GET /sessions/{id}/status` | `bridge_status` | matches herdr `agent_status` |
|
||||
| Detached dispatch + drain | `POST …/message?block=false` · `GET /sessions/{id}/status` (pending drain) | `bridge_send(block:false)` · `bridge_status` | `dispatch_id` issued; reply retrievable |
|
||||
| Worker reply | `POST /sessions/{id}/reply` | `bridge_reply` | resolves the awaiting request by `corr` |
|
||||
| Worker question | `POST /sessions/{id}/ask` | `bridge_ask` | surfaces to primary; parks worker |
|
||||
| Discovery | `GET /sessions` | `bridge_list` | roster + live match config/herdr |
|
||||
| Deliver a turn (blocking) | `POST /sessions/{id}/message` | `fleet_send` | reply returned; `working→idle` unblocks; timeout → 202 |
|
||||
| Worker status | `GET /sessions/{id}/status` | `fleet_status` | matches herdr `agent_status` |
|
||||
| Detached dispatch + drain | `POST …/message?block=false` · `GET /sessions/{id}/status` (pending drain) | `fleet_send(block:false)` · `fleet_status` | `dispatch_id` issued; reply retrievable |
|
||||
| Worker reply | `POST /sessions/{id}/reply` | `fleet_reply` | resolves the awaiting request by `corr` |
|
||||
| Worker question | `POST /sessions/{id}/ask` | `fleet_ask` | surfaces to primary; parks worker |
|
||||
| Discovery | `GET /sessions` | `fleet_list` | roster + live match config/herdr |
|
||||
| Spawn (guarded) | `POST /sessions` | *(internal)* | rejects on-subscription profile; accepts allowlisted |
|
||||
| Health | `GET /healthz` · `GET /metrics` | — | liveness + Prometheus |
|
||||
|
||||
@@ -114,16 +114,16 @@ Compact scope; expand into detailed tickets when a stage starts (as Stage 1 is b
|
||||
|
||||
| Stage | Tickets |
|
||||
|---|---|
|
||||
| **2** | `CB-201` envelope schema + codec · `CB-202` worker `bridge_reply` tool + reviewer skill · `CB-203` reply rendezvous (corr match; resolve on reply *or* the `working→idle` edge) · `CB-204` subscription guard via `ccs env` + allowlist · `CB-205` blocked-worker path (`bridge_ask`) |
|
||||
| **3** | `CB-301` ✅ session manager (spawn/reuse/recycle) · `CB-301-ext` ✅ per-worker git worktree + config-parity overlay (`97ecc71`) · `CB-302` ✅ worker checkpoint — **shipped as commit→push→**_**worker-opened PR**_ (`64e70ef`: repo-scoped forge-token injection + the implementer skill), which **supersedes** this row's original `STATE.md`-file framing; see `docs/Worker-Git-Workflow.md` · `CB-303` ✅ `idle_ttl`/`context_cap`/drain · `CB-304` ✅ `bridge_list` roster+live (`9fe04bf`) · `CB-305` ✅ multi-profile routing |
|
||||
| **4** | `CB-401` ✅ PeerLauncher SPI Stage A — extracted in-tree, one adapter (`ClaudeCodeLauncher`), core uses the `PeerLauncher` interface, main @ `3aa69a9`. `CB-402` ✅ **Stage B complete** (`ded226a`, dogfooded 2026-07-29) — `OpenCodeLauncher` as the SPI-proving second adapter: shares none of Claude's private seams (no `ANTHROPIC_BASE_URL`, no `SubscriptionGuard`), mounts the bridge MCP via a generated `OPENCODE_CONFIG`, and is routed by `kind:` through `CompositePeerLauncher`. Live spawn→send→`bridge_reply`→teardown verified against opencode 1.18.5. Stage C — dynamic external plugin loading, future, gated by trust/capability model. |
|
||||
| **2** | `CB-201` envelope schema + codec · `CB-202` worker `fleet_reply` tool + reviewer skill · `CB-203` reply rendezvous (corr match; resolve on reply *or* the `working→idle` edge) · `CB-204` subscription guard via `ccs env` + allowlist · `CB-205` blocked-worker path (`fleet_ask`) |
|
||||
| **3** | `CB-301` ✅ session manager (spawn/reuse/recycle) · `CB-301-ext` ✅ per-worker git worktree + config-parity overlay (`97ecc71`) · `CB-302` ✅ worker checkpoint — **shipped as commit→push→**_**worker-opened PR**_ (`64e70ef`: repo-scoped forge-token injection + the implementer skill), which **supersedes** this row's original `STATE.md`-file framing; see `docs/Worker-Git-Workflow.md` · `CB-303` ✅ `idle_ttl`/`context_cap`/drain · `CB-304` ✅ `fleet_list` roster+live (`9fe04bf`) · `CB-305` ✅ multi-profile routing |
|
||||
| **4** | `CB-401` ✅ PeerLauncher SPI Stage A — extracted in-tree, one adapter (`ClaudeCodeLauncher`), core uses the `PeerLauncher` interface, main @ `3aa69a9`. `CB-402` ✅ **Stage B complete** (`ded226a`, dogfooded 2026-07-29) — `OpenCodeLauncher` as the SPI-proving second adapter: shares none of Claude's private seams (no `ANTHROPIC_BASE_URL`, no `SubscriptionGuard`), mounts the bridge MCP via a generated `OPENCODE_CONFIG`, and is routed by `kind:` through `CompositePeerLauncher`. Live spawn→send→`fleet_reply`→teardown verified against opencode 1.18.5. Stage C — dynamic external plugin loading, future, gated by trust/capability model. |
|
||||
| **5** | ✅ **CB-501–505 landed** (the stage as originally scoped): `CB-501` bearer auth + non-loopback-bind fail-fast (TLS at a proxy, not in-JVM — see `docs/CB-5xx-Hardening.md` D3) · `CB-502` `/metrics` (zero-dependency Prometheus renderer, D4) + `/healthz` · `CB-503` mock-socket CI (`.gitea/workflows/ci.yml`) · `CB-504` launchd agent + systemd unit + herdr-socket startup wait · `CB-505` per-session authz table + audit log. **The 5xx line did not stop there** — `CB-506`…`CB-525` shipped after this row was written; see [Stage 5 continued](#cb-506525--stage-5-continued-as-built) below and [Features](11-Features) for the operator-facing ones. |
|
||||
|
||||
## CB-401 — Peer Launcher SPI (Stage 4)
|
||||
|
||||
Stage A has landed on main:
|
||||
|
||||
- ✅ **SPI extracted in-tree** (`dev.ltms.bridged.peer`): `PeerLauncher`, `PeerHandle`, `SpawnRequest`, `Capability`.
|
||||
- ✅ **SPI extracted in-tree** (`dev.ltms.fleetd.peer`): `PeerLauncher`, `PeerHandle`, `SpawnRequest`, `Capability`.
|
||||
- ✅ **One adapter** — `worker.ClaudeCodeLauncher` implements `PeerLauncher`; it is the renamed/adapted `WorkerService` and still performs guard-checked spawn, orphan reap, and teardown.
|
||||
- ✅ **Core decoupled** — `session.SessionManager` now depends on the `PeerLauncher` interface and keys its registry on `PeerHandle.id()` (equal to herdr `paneId` for the Claude adapter, so no value change).
|
||||
- ✅ **Behaviour-preserving** — full green gate on main @ `3aa69a9`, **183 tests**.
|
||||
@@ -133,8 +133,8 @@ Stage A has landed on main:
|
||||
- ✅ `HerdrPeerLauncher` base extracted; `OpenCodeLauncher` implements the three divergent hooks.
|
||||
- ✅ `kind:` discriminator on worker profiles; `CompositePeerLauncher` routes spawn/stop/reap/list by kind.
|
||||
- ✅ The SPI is proven provider-neutral: opencode uses **none** of Claude Code's private launch seams.
|
||||
- ✅ **Live dogfood complete.** Spawn → CB-306 readiness gate → `bridge_send` → structured
|
||||
`bridge_reply` (`replySource: "reply"`, not the completion fallback) → teardown, all through the
|
||||
- ✅ **Live dogfood complete.** Spawn → CB-306 readiness gate → `fleet_send` → structured
|
||||
`fleet_reply` (`replySource: "reply"`, not the completion fallback) → teardown, all through the
|
||||
REST surface against opencode **1.18.5**. The schema-drift risk did not materialise: the adapter
|
||||
was designed against 1.1.31 and its generated `OPENCODE_CONFIG` still validates unchanged.
|
||||
Provider question resolved — opencode's gateway serves **free-tier models with zero credentials**,
|
||||
@@ -148,7 +148,7 @@ The single-host close-out, done **before** cross-host rather than after, because
|
||||
gating concern is the trust model and it inherits whatever identity shape lands here. Full design
|
||||
and decision record: **`docs/CB-5xx-Hardening.md`**.
|
||||
|
||||
The finding that shaped the stage: `bridged` had **exactly one security control — the loopback
|
||||
The finding that shaped the stage: `fleetd` had **exactly one security control — the loopback
|
||||
bind**. `ConnectionIdentity` resolves a worker from its connection (unforgeable), but *any* caller
|
||||
that was not a recognised worker pane — including, had the bind ever widened, an arbitrary remote
|
||||
client — was treated as **the primary**, the most privileged role on the bus. CB-501 inverts that
|
||||
@@ -156,7 +156,7 @@ default: `ANONYMOUS` is now the fallback and `PRIMARY` must be established.
|
||||
|
||||
- **CB-501 ✅ auth.** `auth.mode: loopback-trust` (default, the historical behaviour named honestly)
|
||||
or `token` (bearer required of every non-worker caller). Worker identity is *never* token-gated,
|
||||
so enabling auth cannot lock the fleet out of `bridge_reply`. Constant-time token comparison.
|
||||
so enabling auth cannot lock the fleet out of `fleet_reply`. Constant-time token comparison.
|
||||
**The highest-value line in the stage is a startup check:** a non-loopback `bind.host` under
|
||||
`loopback-trust` now *refuses to start* rather than silently promoting every reachable client to
|
||||
primary. TLS terminates at a reverse proxy by design (D3), not in the JVM.
|
||||
@@ -169,7 +169,7 @@ default: `ANONYMOUS` is now the fallback and `PRIMARY` must be established.
|
||||
so a plain `mvn -B clean install` *is* the mock-socket surface.
|
||||
- **CB-504 ✅ supervision.** launchd agent (the real target — this host is macOS, there is no
|
||||
systemd) **and** a systemd unit for the Linux gateways CB-308 adds. Ordering directives are
|
||||
advisory; the actual fix is that bridged now **waits up to 30s for the herdr socket and then
|
||||
advisory; the actual fix is that fleetd now **waits up to 30s for the herdr socket and then
|
||||
serves degraded** instead of crashing into a restart loop on a boot-order race.
|
||||
- **CB-505 ✅ authz + audit.** The role table enforced on **both** entry paths — and that plural is
|
||||
the point. The wiki has long described MCP as "a thin adapter over the REST core"; at code level
|
||||
@@ -194,7 +194,7 @@ turned "it works when I try it" into guarded behaviour. Operator-facing entries
|
||||
|---|---|
|
||||
| `CB-508` | an opencode profile can pin its own OpenAI-compatible endpoint |
|
||||
| `CB-511` | workers get a real toolchain — the daemon's `PATH` is propagated, plus a per-profile `env:` map. Before this a worker inherited whatever `PATH` the herdr *server* was started with, which on a long-lived herdr can predate your toolchain entirely and leave workers unable to run `mvn` at all |
|
||||
| `CB-517` | `bridge_whoami` — a session asks the daemon for its own role instead of inferring it; and the `CLAUDE.md` bridge block became a portable charter copied verbatim into every project that mounts the bridge. LavinMQ pinned as a durable, self-restarting broker |
|
||||
| `CB-517` | `fleet_whoami` — a session asks the daemon for its own role instead of inferring it; and the `CLAUDE.md` bridge block became a portable charter copied verbatim into every project that mounts the bridge. LavinMQ pinned as a durable, self-restarting broker |
|
||||
| `CB-518` | weighted placement (`placement: weighted`, per-profile `weight` / `maxLoad`), smooth weighted round-robin with failover to the next candidate |
|
||||
| `CB-521` | herdr adapter ported to **protocol 19** (herdr 0.8.0) — a hard version coupling, see the note below |
|
||||
| `CB-522` | `primary.terminal` — the primary may run *inside* a herdr pane. Without the pin the pane lookup reads it as a worker and refuses every orchestration verb, and the failure is self-locking: the learned terminal is populated by the very calls being refused |
|
||||
@@ -214,16 +214,16 @@ turned "it works when I try it" into guarded behaviour. Operator-facing entries
|
||||
|
||||
`CB-505 fix` audit lines are valid JSON · `CB-506` the test suite stays out of the production audit
|
||||
log · `CB-509` JaCoCo coverage reporting · `CB-510` SessionReaper 0 → 86.7% · `CB-512`
|
||||
`bridged_push_nudges_total` increments wired · `CB-513` the MCP-side authorization gate, 27.4 →
|
||||
`fleetd_push_nudges_total` increments wired · `CB-513` the MCP-side authorization gate, 27.4 →
|
||||
57.4% · `CB-514` MessageService timeout/answer/poll/lock edges · `CB-515` the turn-attribution
|
||||
guards regression-protected · `CB-521` the AMQP contract test runnable both locally and in CI.
|
||||
|
||||
> **Version coupling (CB-521).** `bridged` speaks one herdr wire protocol; the adapter is ported
|
||||
> **Version coupling (CB-521).** `fleetd` speaks one herdr wire protocol; the adapter is ported
|
||||
> wholesale on a bump, with no negotiation or compat shim. The daemon can therefore be perfectly
|
||||
> healthy — `/healthz` ok, `bridge_whoami` resolving, profiles listed — while **every spawn fails**,
|
||||
> healthy — `/healthz` ok, `fleet_whoami` resolving, profiles listed — while **every spawn fails**,
|
||||
> because health only pings herdr and never checks that the adapter and the binary agree. Symptom:
|
||||
> `invalid_request: missing field <x>`. After any restart onto a jar carrying an adapter change,
|
||||
> verify with a real `bridge_spawn`, not with `/healthz`.
|
||||
> verify with a real `fleet_spawn`, not with `/healthz`.
|
||||
|
||||
## Release 1.1 — the single-host close-out (open)
|
||||
|
||||
@@ -248,14 +248,14 @@ is a workaround; a config key that is accepted and silently does nothing is brok
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph r1["release 1.1 — one bridged per host"]
|
||||
subgraph r1["release 1.1 — one fleetd per host"]
|
||||
sup["supervision<br/>CB-594 · CB-600"]
|
||||
dur["durability claim<br/>CB-527 · CB-528"]
|
||||
sec["credential scope<br/>CB-593"]
|
||||
bugs["nudge scheduling<br/>CB-590 · CB-598"]
|
||||
cfg["config accuracy<br/>CB-597 · CB-599 · CB-602"]
|
||||
flaky["green build<br/>CB-601 · CB-603"]
|
||||
ask["bridge_ask reach<br/>CB-582"]
|
||||
ask["fleet_ask reach<br/>CB-582"]
|
||||
docs["catalogue debt<br/>CB-595"]
|
||||
val["config validation<br/>CB-604 · CB-606"]
|
||||
op["needs the operator<br/>CB-596"]
|
||||
@@ -284,9 +284,9 @@ two: CB-596's own gap detector found both of them on its first live spawn.*
|
||||
| **Supervision** | CB-594 (#80) · CB-600 (#91) | The launchd unit shipped in `deploy/` and had never been installed, and launchd does not source a login shell — so a supervised daemon got no `WORKER_GITEA_TOKEN` and no `AI_GATEWAY_TOKEN`. The operator had to pick supervision *or* a working fleet. CB-600 then closed the gaps that only bite once the agent is loaded: a log path the script and the plist could silently disagree about, and a failed `load` leaving the agent stopped **and** persistently disabled. |
|
||||
| **Durability** | CB-527 (#10) · CB-528 (#11) | No `basicQos`, so the backlog lived in JVM heap rather than on the broker; no publisher confirms, so a publish to a missing queue was a silent black hole. The wiki promised more durability than the code delivered. Both closed, with contract tests running against a real broker in CI. |
|
||||
| **Credential scope** | CB-593 (#79) | The decision about which inherited credentials a member may keep is now recorded, so a deliberate choice no longer looks like an oversight. The canonical block also stopped claiming a member mounts only the bridge. |
|
||||
| **Nudge scheduling** | CB-590 (#75) · CB-598 (#87) · CB-582 (#61) | Two schedules could inject into one lead pane at once. Then the reminder budget turned out to be a counter carried forward with no memory of *which item* it counted, so work arriving during the backoff window inherited an already-capped count and was never nudged once. CB-582 added the third source: a worker paused in `bridge_ask` now nudges the lead itself, and `bridge_status` and REST both show the open question. **The ~55s window is closed, not removed** — still do not brief a worker to "ask me". |
|
||||
| **Nudge scheduling** | CB-590 (#75) · CB-598 (#87) · CB-582 (#61) | Two schedules could inject into one lead pane at once. Then the reminder budget turned out to be a counter carried forward with no memory of *which item* it counted, so work arriving during the backoff window inherited an already-capped count and was never nudged once. CB-582 added the third source: a worker paused in `fleet_ask` now nudges the lead itself, and `fleet_status` and REST both show the open question. **The ~55s window is closed, not removed** — still do not brief a worker to "ask me". |
|
||||
| **Config accuracy** | CB-597 (#85) · CB-599 (#89) · CB-602 (#96) · CB-604 (#102) | Two documented knobs that are read by nothing; a capacity refusal that surfaced as a bare HTTP 500; no test at all in the code→example direction, so a brand-new key could ship undocumented; and an unknown `kind:` accepted silently and routed to the wrong adapter, which now refuses at config load. |
|
||||
| **Catalogue debt** | CB-595 (#81) | `wiki/11-Features.md` had fallen about fourteen entries behind, worst on the entries that changed what a config key *means*. `bridged.yaml` is gitignored, so that page is the only place an operator could learn them. Cleared — and it turned up two defects on the way (CB-604, CB-606). |
|
||||
| **Catalogue debt** | CB-595 (#81) | `wiki/11-Features.md` had fallen about fourteen entries behind, worst on the entries that changed what a config key *means*. `fleetd.yaml` is gitignored, so that page is the only place an operator could learn them. Cleared — and it turned up two defects on the way (CB-604, CB-606). |
|
||||
| **Green build** | CB-601 (#95) · CB-603 (#100) | Two flaky tests, same root: a test that observes an asynchronous loop must be safe against that loop's thread, and neither the compiler nor a green build will say it is not. |
|
||||
| **Config validation** | CB-604 (#102) · CB-606 (#106) | Four config fields shared one shape: lower-cased in a compact constructor, then compared against exactly **one** string, so a typo fell through to the other branch in silence. The worst was `auth.mode` — a typo of `token` behaved as `loopback-trust`, and `validateAuthExposure()` only fires on a **non-loopback** bind, so the common loopback bind hid it end to end and the daemon authenticated nobody while the config said otherwise. All four now refuse at load, naming the field, the value, the accepted set, and what would have happened. |
|
||||
| **Snapshot pruning** | CB-586 (#67) | Nothing pruned `refs/wip/*`, so CB-578 stage C's snapshots pinned their whole trees forever. The rule that landed needs **both** conditions: the tree is already reachable from `main`, and the ref is older than 24h. Reachability is the floor — a snapshot exists because the work was committed nowhere else, so a plain TTL would delete the only copy. `/members` now reports `wipRefs{count,costBytes}`. The sweep shipped **dead**: a `Long.MIN_VALUE` "never yet" sentinel overflowed the interval gate, which returned before the assignment that would have fixed it, so it never ran once — and every unit test passed, because they all called the seam directly and walked around the gate. |
|
||||
@@ -306,7 +306,7 @@ whole reason CB-592 was built the way it was; CB-596 inherits it.
|
||||
**What the measurement actually said.** Two results, and they point in opposite directions.
|
||||
|
||||
The reassuring one: **CB-592 works.** The member holds the sentinel
|
||||
`blocked-by-bridged-cb592-see-gitea-issue-77`, not the admin token — confirmed by hashing the
|
||||
`blocked-by-fleetd-cb592-see-gitea-issue-77`, not the admin token — confirmed by hashing the
|
||||
sentinel, which is a hardcoded non-secret string, and matching it against what the member reported.
|
||||
The guarded-`export` mechanism does beat the login shell, which is the whole reason it was built that
|
||||
way.
|
||||
@@ -361,7 +361,7 @@ Made 2026-08-16, by reading the admission rule strictly.
|
||||
| Ticket | Why it moved |
|
||||
|---|---|
|
||||
| CB-308 (#6) | Federation. It is what 2.0 *is*. |
|
||||
| CB-589 (#74) | Cost-first placement. `weighted` spreads by ratio with no idea which profile costs money, so paid spawns happen while the free box sits idle — but that is working behaviour that costs too much, with a decisive workaround (`local.weight: 100`), not something broken on one host. A missing **explanation** of the workaround *was* broken, because `bridged.yaml` is gitignored and a fresh host starts without it. **That half shipped in 1.1**; the policy did not. |
|
||||
| CB-589 (#74) | Cost-first placement. `weighted` spreads by ratio with no idea which profile costs money, so paid spawns happen while the free box sits idle — but that is working behaviour that costs too much, with a decisive workaround (`local.weight: 100`), not something broken on one host. A missing **explanation** of the workaround *was* broken, because `fleetd.yaml` is gitignored and a fresh host starts without it. **That half shipped in 1.1**; the policy did not. |
|
||||
| CB-548 (#16) | Architect slots. A new capability, and genuinely blocked: `MemberRegistry.bind` is never called. |
|
||||
| CB-605 (#103) | The systemd unit carries launchd's login-shell secret gap. Bites only when the first Linux gateway is stood up. |
|
||||
| CB-607 (#110) | A member holds `SSH_AUTH_SOCK`, so it can sign with every key the operator's agent holds — broader than the repo-scoped token CB-302 built. Allowed **on purpose** today, because worktree remotes are `ssh://` and blocking it stops members pushing. A documented, deliberate scope reduction, not a break. |
|
||||
@@ -375,7 +375,7 @@ Made 2026-08-16, by reading the admission rule strictly.
|
||||
`e11160695fbe`, and the startup line reads
|
||||
`memberCredentials: 34 known name(s), 5 allowed — blocking 29 on every spawn`. A merge is not a
|
||||
deployment, and this change alters what a member's environment contains.
|
||||
4. ~~Confirm `bridge_whoami` still answers `primary`.~~ Done: `primary`. Note the trap — the redeploy
|
||||
4. ~~Confirm `fleet_whoami` still answers `primary`.~~ Done: `primary`. Note the trap — the redeploy
|
||||
cuts the lead's own bridge MCP mount. It reconnected by itself the second time and did not the
|
||||
first, so do not rely on either; if the tools are gone, ask the operator to run `/mcp`.
|
||||
5. ~~Apply the secret-store half of CB-596.~~ Done by the operator on 2026-08-17: `secrets.sh` backed
|
||||
@@ -406,7 +406,7 @@ installing it is the operator's decision. Read the Stage 5 table as *built*, not
|
||||
|
||||
A cross-cutting track that came out of a **"communication break" review** of the reverse
|
||||
(worker → primary) path. The bridge *pushes* to workers (the status-gated injector) but only
|
||||
*pulls* to the primary — a worker reply resolves only an **already-open** blocking `bridge_send`;
|
||||
*pulls* to the primary — a worker reply resolves only an **already-open** blocking `fleet_send`;
|
||||
MCP is client-initiated (bridge = server, primary = client), so the server can't call into the
|
||||
primary. That asymmetry is the root of all three tickets.
|
||||
|
||||
@@ -435,9 +435,9 @@ spawn fail fast instead of stalling.*
|
||||
- **CB-307 — reliable worker → primary delivery.** Behind one `ReplyInbox` port so the adapter is
|
||||
swappable (see [Implementation → `msg`](9-Implementation#msg--the-service-core-4-classes)).
|
||||
- **Stage 1 ✅ (main `ba6b4a5`).** `ReplyInbox` port + `InMemoryReplyInbox` (soft-state, dedup by
|
||||
`msgId`). A `bridge_reply` with no open send is now **held** instead of dropped; the primary
|
||||
drains it by target via `bridge_poll(target)` / `GET /sessions/{id}/replies`. No broker, no new
|
||||
dependency. **Only terminal replies are queued** — `bridge_ask` (interactive) and the
|
||||
`msgId`). A `fleet_reply` with no open send is now **held** instead of dropped; the primary
|
||||
drains it by target via `fleet_poll(target)` / `GET /sessions/{id}/replies`. No broker, no new
|
||||
dependency. **Only terminal replies are queued** — `fleet_ask` (interactive) and the
|
||||
completion/failure fallbacks are deliberately *not* (would risk double-delivery). Live on the
|
||||
running daemon.
|
||||
- **Stage 2 ✅ (main `2bc5f3a`).** `AmqpReplyInbox` behind the *same* port for cross-restart
|
||||
@@ -448,7 +448,7 @@ spawn fail fast instead of stalling.*
|
||||
the primary drains, so a `java -jar` bounce leaves them on the broker for redelivery. `broker:`
|
||||
config absent → in-memory, present → AMQP. Contract test `@Tag("contract")` runs against a RabbitMQ
|
||||
container. Dogfooded live (reply survived a daemon bounce, redelivered + acked exactly once).
|
||||
**`bridged` stays soft-state — the broker owns message durability, not the bus.**
|
||||
**`fleetd` stays soft-state — the broker owns message durability, not the bus.**
|
||||
- **Stage 3 — active push-to-primary + reminder ✅ (main `d4c9704`, gitea #5 closed).** The durable
|
||||
inbox is a *landing zone* but delivery was still **pull** (the primary had to poll). Stage 3 makes
|
||||
it **active**: a `ReplyPushLoop` nudges the primary the moment a reply lands with no open send, and
|
||||
@@ -459,11 +459,11 @@ spawn fail fast instead of stalling.*
|
||||
`push_backoff_ms`=15000). **Ack = drain**: stop when `inbox.peek(target).isEmpty()`. A single-slot
|
||||
`PrimaryRegistry` learns the primary from orchestration-side tools; an off-host / non-herdr primary
|
||||
leaves it empty → the loop is a no-op and delivery degrades to pull (reply never lost). Optional
|
||||
per-`msgId` `bridge_ack` tool for finer control than drain-all. Dogfooded live end-to-end.
|
||||
per-`msgId` `fleet_ack` tool for finer control than drain-all. Dogfooded live end-to-end.
|
||||
|
||||
- **CB-308 — multi-host federation ⏳ (design note, gitea #6; depends on CB-307 Stage 2).** A primary
|
||||
on host A delegating to workers on hosts B, C… with no host learning another's terminals. Built on
|
||||
CB-307's broker fabric: a **per-host gateway** (evolved `bridged` owning its local herdr +
|
||||
CB-307's broker fabric: a **per-host gateway** (evolved `fleetd` owning its local herdr +
|
||||
registry), **per-agent broker channels** `agent.<globalId>.inbox` (owning gateway = sole consumer),
|
||||
a **federated roster** (soft-state presence on a `roster.*` topic = [CB-304](9-Implementation)'s
|
||||
`rosterView`, federated), and a one-line **routing fork** (`local ? inject : publish`). Five
|
||||
@@ -475,14 +475,14 @@ spawn fail fast instead of stalling.*
|
||||
|
||||
## Stage 1 — detailed tickets
|
||||
|
||||
**Goal:** Opus, from its own subscription session, mounts `bridged` and gets a code review
|
||||
**Goal:** Opus, from its own subscription session, mounts `fleetd` and gets a code review
|
||||
back from a real worker spawned under the `ccs ltms-local` profile (routing to the gx00 vLLM at
|
||||
`http://gx00.gw:8000`) — same host, one hardcoded profile, reply via a pane/agent read (envelope
|
||||
comes in Stage 2). This is the thinnest end-to-end vertical slice.
|
||||
|
||||
**Definition of done for the stage:** `CB-107` demo passes.
|
||||
|
||||
> **Build status — Stage 1 COMPLETE** (in `bridged/`, Maven · Java 25 · **307 unit/acceptance tests
|
||||
> **Build status — Stage 1 COMPLETE** (in `fleetd/`, Maven · Java 25 · **307 unit/acceptance tests
|
||||
> green** as of the CB-5xx close-out; the live-herdr and broker contract tests run separately via
|
||||
> `mvn test -Pcontract`). The `CB-107` end-to-end demo gate passes and the
|
||||
> bridge is **dogfooded**: an Opus primary delegates real tasks to off-subscription workers that
|
||||
@@ -490,8 +490,8 @@ comes in Stage 2). This is the thinnest end-to-end vertical slice.
|
||||
> ✅ **CB-101** herdr client — connection-per-call, contract-tested vs live 0.7.0.
|
||||
> ✅ **CB-102** worker spawn — native `agent.*` with env injection, guard-checked, live-verified.
|
||||
> ✅ **CB-103** status-gated injector — per-target FIFO, one message per turn, TOCTOU closed.
|
||||
> ✅ **CB-104** blocking `bridge_send` — `POST /sessions/{id}/message` + reply rendezvous.
|
||||
> ✅ **CB-105** MCP adapter — `bridge_send`/`bridge_reply`/`bridge_status` over the REST core, with
|
||||
> ✅ **CB-104** blocking `fleet_send` — `POST /sessions/{id}/message` + reply rendezvous.
|
||||
> ✅ **CB-105** MCP adapter — `fleet_send`/`fleet_reply`/`fleet_status` over the REST core, with
|
||||
> connection-based caller identity (loopback peer PID → herdr pane).
|
||||
> ✅ **CB-106** config + logging — Jackson YAML + Logback.
|
||||
> ✅ **CB-107** e2e review demo (stage gate) — runs green end-to-end; the worker is verifiably the
|
||||
@@ -507,8 +507,8 @@ comes in Stage 2). This is the thinnest end-to-end vertical slice.
|
||||
> numbers (CB-106…CB-118 in git) reuse the CB-106/107/108 slots this plan assigned above to
|
||||
> config/e2e-demo/placement — identify the items below by name, not number.** Shipped:
|
||||
> completion fallback (a confirmed `working→idle` turn resolves a send), async fire-and-poll
|
||||
> (`wait:false` + `bridge_poll`, beats the caller's MCP call timeout), fleet MCP tools
|
||||
> (`bridge_spawn`/`bridge_list`/`bridge_stop`/`bridge_profiles`/`bridge_poll`), failure detection
|
||||
> (`wait:false` + `fleet_poll`, beats the caller's MCP call timeout), fleet MCP tools
|
||||
> (`fleet_spawn`/`fleet_list`/`fleet_stop`/`fleet_profiles`/`fleet_poll`), failure detection
|
||||
> for wedged (`unknown`), vanished, and never-ready workers, multi-profile workers with
|
||||
> per-profile base_url guards, worker cwd inheritance (never `$HOME`), and a readiness gate that
|
||||
> holds delivery until the worker's Claude has connected the bridge MCP.
|
||||
@@ -520,19 +520,19 @@ comes in Stage 2). This is the thinnest end-to-end vertical slice.
|
||||
> completion-baseline clip fix so the CB-115 misattribution guard holds for assistant blocks over
|
||||
> the scrape cap (CB-118). Verified by a single-worker conversation harness, a **1-primary /
|
||||
> N-worker fan-out issue-hunt**, and a **sustained 5-minute stateful back-and-forth** (30 turns,
|
||||
> every one a clean `bridge_reply`, running total held) — all under `e2e/`.
|
||||
> every one a clean `fleet_reply`, running total held) — all under `e2e/`.
|
||||
>
|
||||
> **Stage 2 — rich message semantics COMPLETE.** The contract-and-guard stage has landed (the
|
||||
> +13 tests above are its coverage):
|
||||
> ✅ **CB-205** `bridge_ask` reverse rendezvous — a worker pauses its delegated turn to ask the
|
||||
> ✅ **CB-205** `fleet_ask` reverse rendezvous — a worker pauses its delegated turn to ask the
|
||||
> primary and resumes the **same** turn with the answer. The question surfaces on the primary's own
|
||||
> blocked `bridge_send` carrying a `turn_id`, and the primary answers by sending on that `turn_id`.
|
||||
> Live e2e (`e2e/bridge_ask_test.py`): the worker asked in ~9s and, after the primary answered,
|
||||
> resumed and replied via a clean `bridge_reply`.
|
||||
> blocked `fleet_send` carrying a `turn_id`, and the primary answers by sending on that `turn_id`.
|
||||
> Live e2e (`e2e/fleet_ask_test.py`): the worker asked in ~9s and, after the primary answered,
|
||||
> resumed and replied via a clean `fleet_reply`.
|
||||
> ✅ **CB-202** reviewer-role skill (`.claude/skills/reviewer/SKILL.md`) — the playbook a worker
|
||||
> loads to review a scoped assignment, ask the lead via `bridge_ask` when the call is genuinely
|
||||
> theirs, and report exactly one structured finding via `bridge_reply`. Pairs the already-shipped
|
||||
> `bridge_reply`/`bridge_ask` tools with the role guidance for using them.
|
||||
> loads to review a scoped assignment, ask the lead via `fleet_ask` when the call is genuinely
|
||||
> theirs, and report exactly one structured finding via `fleet_reply`. Pairs the already-shipped
|
||||
> `fleet_reply`/`fleet_ask` tools with the role guidance for using them.
|
||||
> ⚠️ **CB-201 descoped** — the planned `{from,to,corr}` envelope schema + codec was **not** built.
|
||||
> Connection identity (loopback peer PID → herdr pane) already routes every reply and question to
|
||||
> the right session robustly, so CB-201 shipped as a *lightweight* addition only where the reverse
|
||||
@@ -549,7 +549,7 @@ comes in Stage 2). This is the thinnest end-to-end vertical slice.
|
||||
and `subscribe(subs) → BlockingQueue<Event>` (fed by an event-loop virtual thread) for
|
||||
`pane.agent_status_changed`. Probe `ping` on connect (assert `protocol: 14`); enumerate via
|
||||
`workspace.list`/`pane.list`; fail fast on schema mismatch. **Build against the
|
||||
[verified 0.7.0 API](2-Message-Server#the-herdr-control-contract-what-bridged-drives), not the docs.**
|
||||
[verified 0.7.0 API](2-Message-Server#the-herdr-control-contract-what-fleetd-drives), not the docs.**
|
||||
**Acceptance.** JUnit test against a **mock UDS socket** replays a `workspace.create` round-trip
|
||||
and one event; a **contract test vs the real herdr** asserts `ping.protocol == 14` and the
|
||||
`workspace.list`/`pane.list` shapes.
|
||||
@@ -557,7 +557,7 @@ and one event; a **contract test vs the real herdr** asserts `ping.protocol == 1
|
||||
|
||||
### CB-102 — spawn a worker (native `agent.*`) ✅
|
||||
**Decided (spike done).** herdr's `agent.start {name, argv, env}` carries the worker command and
|
||||
a **first-class `env` map that reaches the process** (proven via `agent.read`), so `bridged`
|
||||
a **first-class `env` map that reaches the process** (proven via `agent.read`), so `fleetd`
|
||||
injects `ANTHROPIC_BASE_URL` there — guard-checked before the call — with no shell prefix and its
|
||||
own env untouched. Chosen over the pane + `send_text` fallback. herdr tracks each worker's Claude
|
||||
**session UUID** (`agent.list`/`agent.get`), which grounds the ID contract. Teardown is
|
||||
@@ -579,7 +579,7 @@ pane; status-check + send serialized on the event-loop virtual thread (close the
|
||||
two rapid deliveries never interleave (golden transcript).
|
||||
**Deps.** CB-101.
|
||||
|
||||
### CB-104 — blocking `bridge_send` as a **REST endpoint** + reply capture
|
||||
### CB-104 — blocking `fleet_send` as a **REST endpoint** + reply capture
|
||||
**Scope.** Implement the feature as **`POST /sessions/{id}/message`** (`{content}`, blocking) →
|
||||
resolve to the (single, hardcoded) worker pane → inject → wait for the turn-done `working→idle` edge → return
|
||||
`pane.read {source:"recent-unwrapped"}` of the last assistant block in the response body. (No
|
||||
@@ -591,17 +591,17 @@ timeout returns a typed "still working" (HTTP 202-style) response.
|
||||
**Deps.** CB-102, CB-103.
|
||||
|
||||
### CB-105 — MCP adapter over the REST core (SERVER face)
|
||||
**Scope.** Streamable-HTTP MCP server exposing `bridge_send`/`bridge_status` as **thin adapters
|
||||
**Scope.** Streamable-HTTP MCP server exposing `fleet_send`/`fleet_status` as **thin adapters
|
||||
over the CB-104 REST routes** (`POST /sessions/{id}/message`, `GET /sessions/{id}/status`); bind
|
||||
`127.0.0.1:8080`. One-line mount: `claude mcp add --transport http bridge
|
||||
http://127.0.0.1:8080/mcp`.
|
||||
**Acceptance.** **Parity test** — `bridge_send` via MCP and `POST …/message` via REST produce
|
||||
identical results for the same input; `claude mcp list` shows `bridge` connected; `bridge_status`
|
||||
**Acceptance.** **Parity test** — `fleet_send` via MCP and `POST …/message` via REST produce
|
||||
identical results for the same input; `claude mcp list` shows `bridge` connected; `fleet_status`
|
||||
returns the live `agent_status`.
|
||||
**Deps.** CB-104.
|
||||
|
||||
### CB-106 — config + wiring
|
||||
**Scope.** `bridged.yaml` load (**Jackson YAML**): one worker profile (`gx00-vllm` → ccs profile,
|
||||
**Scope.** `fleetd.yaml` load (**Jackson YAML**): one worker profile (`gx00-vllm` → ccs profile,
|
||||
model, `base_url` host), bind address, workspace root. **SLF4J/Logback** startup line logging the
|
||||
resolved worker command (secrets redacted). Run as a foreground process (systemd deferred to
|
||||
Stage 5).
|
||||
@@ -610,8 +610,8 @@ logs the resolved (redacted) spawn command.
|
||||
**Deps.** none (parallel with CB-101).
|
||||
|
||||
### CB-107 — end-to-end review demo (stage gate)
|
||||
**Scope.** Scripted demo: start herdr → start `bridged` → mount MCP on a primary `claude` →
|
||||
from the primary, `bridge_send` a real diff with `kind:"review.request"` (body carried as text
|
||||
**Scope.** Scripted demo: start herdr → start `fleetd` → mount MCP on a primary `claude` →
|
||||
from the primary, `fleet_send` a real diff with `kind:"review.request"` (body carried as text
|
||||
for now) → assert a review comes back as the tool result. Document the exact steps in
|
||||
[Setup](4-Setup).
|
||||
**Acceptance.** The demo runs green on one host end-to-end; the worker is verifiably the
|
||||
@@ -622,7 +622,7 @@ for now) → assert a review comes back as the tool result. Document the exact s
|
||||
flowchart LR
|
||||
CB101["CB-101<br/>herdr client"] --> CB102["CB-102<br/>ccs spawn"]
|
||||
CB102 --> CB103["CB-103<br/>injector"]
|
||||
CB103 --> CB104["CB-104<br/>bridge_send"]
|
||||
CB103 --> CB104["CB-104<br/>fleet_send"]
|
||||
CB104 --> CB105["CB-105<br/>MCP server"]
|
||||
CB106["CB-106<br/>config"] --> CB107
|
||||
CB105 --> CB107["CB-107<br/>e2e demo (gate)"]
|
||||
|
||||
+46
-46
@@ -1,13 +1,13 @@
|
||||
# 9. Implementation Architecture (as-built)
|
||||
|
||||
> **Scope.** This is the *as-built* code map of the `bridged` module — the actual packages,
|
||||
> **Scope.** This is the *as-built* code map of the `fleetd` module — the actual packages,
|
||||
> classes, flows, and state machines in the source tree, as a companion to the design-level
|
||||
> [1. Architecture](1-Architecture) and [2. Message Server](2-Message-Server). Every enum,
|
||||
> constant, and route below was verified against source at main `3aa69a9`; the `msg`-layer
|
||||
> **reply-inbox** (CB-307 Stage 1 `ba6b4a5`), its **AMQP durable adapter** (Stage 2 `2bc5f3a`), and
|
||||
> the **active push-to-primary loop** (Stage 3 `d4c9704`) are folded in below.
|
||||
|
||||
`bridged` is a single-host Java 25 / Maven daemon: the sole gateway between an on-subscription
|
||||
`fleetd` is a single-host Java 25 / Maven daemon: the sole gateway between an on-subscription
|
||||
**primary** (Opus) and off-subscription **workers**, speaking to the **herdr** PTY manager
|
||||
(protocol 19, herdr 0.8.0 — ported in CB-521) over a Unix-domain socket. It presents two equivalent faces — a REST
|
||||
server (the testability seam) and an MCP server — over one shared service core.
|
||||
@@ -16,17 +16,17 @@ server (the testability seam) and an MCP server — over one shared service core
|
||||
|
||||
The daemon is layered. The two north faces (REST + MCP) are thin adapters over one service core
|
||||
(`msg`); the core drives delivery through `inject`, which speaks to the outside world only through
|
||||
`herdr`. `guard` sits on the spawn path; `config` and the `Bridged` entry point wire it all.
|
||||
`herdr`. `guard` sits on the spawn path; `config` and the `Fleetd` entry point wire it all.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
primary["Primary (Opus)<br/>on-subscription"]
|
||||
worker["Worker (Claude Code)<br/>off-subscription"]
|
||||
|
||||
subgraph bridged["bridged daemon"]
|
||||
subgraph fleetd["fleetd daemon"]
|
||||
direction TB
|
||||
subgraph faces["North faces (thin adapters)"]
|
||||
rest["rest.BridgedApp<br/>REST · Javalin"]
|
||||
rest["rest.FleetdApp<br/>REST · Javalin"]
|
||||
mcp["mcp.BridgeMcp<br/>MCP · /mcp servlet"]
|
||||
end
|
||||
msg["msg.MessageService + Rendezvous + ReplyInbox<br/>service core · rendezvous + held-reply inbox"]
|
||||
@@ -38,9 +38,9 @@ flowchart TB
|
||||
|
||||
herdrd["herdr daemon<br/>(PTY manager)"]
|
||||
|
||||
primary -->|"bridge_send / answer"| rest
|
||||
primary -->|"fleet_send / answer"| rest
|
||||
primary -->|"MCP tool calls"| mcp
|
||||
worker -->|"bridge_reply / bridge_ask"| mcp
|
||||
worker -->|"fleet_reply / fleet_ask"| mcp
|
||||
rest --> msg
|
||||
mcp --> msg
|
||||
mcp --> ccl
|
||||
@@ -64,7 +64,7 @@ the adapter and gates spawn.*
|
||||
|
||||
## Package & class reference
|
||||
|
||||
Fourteen packages under `dev.ltms.bridged`: `auth`, `config`, `guard`, `herdr`, `inject`, `lead`,
|
||||
Fourteen packages under `dev.ltms.fleetd`: `auth`, `config`, `guard`, `herdr`, `inject`, `lead`,
|
||||
`mcp`, `member`, `metrics`, `msg`, `peer`, `placement`, `rest`, `session`. Below, per layer: the
|
||||
classes, their kind, and their role. Method signatures are abbreviated; see source for full
|
||||
contracts. The sections below do not yet cover every package — `auth`, `herdr`, `inject`, `msg`,
|
||||
@@ -78,7 +78,7 @@ Java records.
|
||||
|
||||
| Class | Kind | Role |
|
||||
|---|---|---|
|
||||
| `HerdrClient` | interface | The client face — the only herdr speaker in `bridged` (`call(method, params)`). |
|
||||
| `HerdrClient` | interface | The client face — the only herdr speaker in `fleetd` (`call(method, params)`). |
|
||||
| `UnixSocketHerdrClient` | class | JDK Unix-socket transport; fresh connection per `call`, one frame, one line, close. |
|
||||
| `HerdrCodec` | class | Encodes/decodes JSON-RPC frames (`encode`, `decodeResult`). |
|
||||
| `HerdrException` | class | Runtime failure carrying an optional herdr protocol `code()`. |
|
||||
@@ -97,8 +97,8 @@ text (`AgentControl.SUBMIT_KEY`) — never a shell prefix. `closeTab`/`PaneLocat
|
||||
### `inject` — status-gated delivery + turn lifecycle (6 classes)
|
||||
|
||||
The single-writer delivery layer. Polls worker status, queues per target, injects only when safe,
|
||||
and **synthesizes turn boundaries** so a blocked `bridge_send` resolves even when a worker never
|
||||
calls `bridge_reply`. This is where the turn state machine lives.
|
||||
and **synthesizes turn boundaries** so a blocked `fleet_send` resolves even when a worker never
|
||||
calls `fleet_reply`. This is where the turn state machine lives.
|
||||
|
||||
| Class | Kind | Role |
|
||||
|---|---|---|
|
||||
@@ -111,17 +111,17 @@ calls `bridge_reply`. This is where the turn state machine lives.
|
||||
|
||||
### `msg` — the service core (6 classes)
|
||||
|
||||
Owns the forward rendezvous (`bridge_send` → `bridge_reply`) and the reverse rendezvous
|
||||
(`bridge_ask` → answer), plus async fire-and-poll — and, since CB-307, the **reply inbox** that
|
||||
Owns the forward rendezvous (`fleet_send` → `fleet_reply`) and the reverse rendezvous
|
||||
(`fleet_ask` → answer), plus async fire-and-poll — and, since CB-307, the **reply inbox** that
|
||||
holds a worker's terminal reply when *no* forward send is open (instead of dropping it) and the
|
||||
**push loop** that actively nudges the primary to drain it.
|
||||
|
||||
| Class | Kind | Role |
|
||||
|---|---|---|
|
||||
| `MessageService` | class | Orchestrates send/reply/ask/answer + async dispatch (`send`, `answer`, `ask`, `sendAsync`, `poll`), and routes `bridge_reply` through **`reply`** (resolve an open send, else publish to the inbox) + **`drainReplies`** (peek-then-ack a target's held replies). |
|
||||
| `MessageService` | class | Orchestrates send/reply/ask/answer + async dispatch (`send`, `answer`, `ask`, `sendAsync`, `poll`), and routes `fleet_reply` through **`reply`** (resolve an open send, else publish to the inbox) + **`drainReplies`** (peek-then-ack a target's held replies). |
|
||||
| `Rendezvous` | class | Low-level registry of forward waiters + reverse-ask futures (`open`, `resolve`, `resolveQuestion`, `openAsk`, `answerAsk`, `closeAsk`, `resolveCompletion`, `resolveFailure`). **Untouched by CB-307** — it stays a pure synchronization primitive; a `false` from `resolve` (no live waiter) is what triggers the inbox publish, one layer up in `MessageService`. |
|
||||
| `ReplyInbox` | interface | The port (CB-307): `publish(target, msgId, content)` (idempotent, dedup by `msgId`), `peek(target)` (non-destructive FIFO snapshot), `ack(target, msgId)`. Nested `InboxMessage(msgId, target, content)` record. Two adapters implement the *same* port — soft-state in-memory and durable AMQP. |
|
||||
| `InMemoryReplyInbox` | class | Stage-1 default adapter — per-target FIFO in a `ConcurrentHashMap<String, LinkedHashMap<msgId, InboxMessage>>`, thread-safe, dedup by `msgId`. **Soft-state, not persistence** — undrained replies are lost on a `java -jar` bounce, consistent with "bridged stays soft-state; the broker owns durability." Selected when `broker:` config is absent. |
|
||||
| `InMemoryReplyInbox` | class | Stage-1 default adapter — per-target FIFO in a `ConcurrentHashMap<String, LinkedHashMap<msgId, InboxMessage>>`, thread-safe, dedup by `msgId`. **Soft-state, not persistence** — undrained replies are lost on a `java -jar` bounce, consistent with "fleetd stays soft-state; the broker owns durability." Selected when `broker:` config is absent. |
|
||||
| `AmqpReplyInbox` | class | Stage-2 durable adapter (CB-307 `2bc5f3a`) — one durable queue `agent.<target>.inbox` per target; **consume-and-hold with deferred manual ack** (a manual-ack consumer pulls persistent messages into an in-memory held map but doesn't ack until the primary drains, so a bounce leaves them on the broker for redelivery). Dedup keys the held map by `msgId`; automatic connection + topology recovery re-declares queues and clears stale delivery-tags. LavinMQ default, RabbitMQ by URI swap. Selected when `broker:` config is present. |
|
||||
| `ReplyPushLoop` | class | Stage-3 active push (CB-307 `d4c9704`) — a dedicated status-gated scheduled loop (mechanism (b), **not** the worker `Injector`, so it stays decoupled from `WorkerPresence`). `onReplyQueued(target)` (called on the no-waiter branch of `reply`) starts a bounded reminder loop: at each tick `decide(target, count)` returns `INJECT` / `WAIT_BUSY` / `STOP`, injecting a *drain nudge* into the primary's own pane via `AgentControl.send` only when `status(primary).injectable()`. Stops when `peek(target)` is empty (ack = drain) or the reminder cap is reached; idempotent per target. |
|
||||
|
||||
@@ -157,7 +157,7 @@ cannot claim to be someone else.
|
||||
|
||||
**Two consequences worth stating.** First, a member's role reaches the principal only through
|
||||
`MemberRegistry`, and only architects bind; a developer and a reviewer are both `Role.WORKER` at
|
||||
this layer, and the difference between them lives in the roster (`bridge_list`), not in the
|
||||
this layer, and the difference between them lives in the roster (`fleet_list`), not in the
|
||||
principal. Second, `SEND` is the one action an architect gains over a worker, which is what lets
|
||||
two architects talk to each other without the lead relaying every message.
|
||||
|
||||
@@ -176,40 +176,40 @@ is `member.ClaudeCodeLauncher`; future adapters (e.g. Codex) implement the same
|
||||
|
||||
### `mcp` — the MCP north face (7 classes)
|
||||
|
||||
Exposes `bridged` as a Streamable-HTTP MCP endpoint and resolves caller identity from the
|
||||
Exposes `fleetd` as a Streamable-HTTP MCP endpoint and resolves caller identity from the
|
||||
**connection**, not from tool arguments (unspoofable).
|
||||
|
||||
| Class | Kind | Role |
|
||||
|---|---|---|
|
||||
| `BridgeMcp` | class | Builds the MCP server, registers `bridge_*` tools, holds thin tool adapters (`servlet`, `send`, `answer`, `ask`, `reply`, `spawn`). |
|
||||
| `BridgeMcp` | class | Builds the MCP server, registers `fleet_*` tools, holds thin tool adapters (`servlet`, `send`, `answer`, `ask`, `reply`, `spawn`). |
|
||||
| `ConnectionIdentity` | class | Resolves *who is calling*: peer PID → herdr pane → `terminal_id` (`resolve`, `callerTerminal`, `cwdForPid`). |
|
||||
| `PrimaryRegistry` | class | Single-slot thread-safe holder of the primary's `terminal_id` (CB-307 Stage 3). `record(id)` learns it from orchestration-side tools (`bridge_send`/`bridge_spawn`) when the caller resolves to a non-null terminal that is *not* a registered worker; `primaryTerminal()` / `isKnown()` feed the push loop. Optional constructor-pin (`primary.terminal` config) for an operator override; stays empty (→ push degrades to pull) when the primary is off-host or non-herdr. |
|
||||
| `PrimaryRegistry` | class | Single-slot thread-safe holder of the primary's `terminal_id` (CB-307 Stage 3). `record(id)` learns it from orchestration-side tools (`fleet_send`/`fleet_spawn`) when the caller resolves to a non-null terminal that is *not* a registered worker; `primaryTerminal()` / `isKnown()` feed the push loop. Optional constructor-pin (`primary.terminal` config) for an operator override; stays empty (→ push degrades to pull) when the primary is off-host or non-herdr. |
|
||||
| `PeerPidLookup` | interface | Abstracts OS peer-PID lookup for a loopback source port. |
|
||||
| `LsofPeerPidLookup` | class | `lsof` impl; excludes bridged's own PID. |
|
||||
| `LsofPeerPidLookup` | class | `lsof` impl; excludes fleetd's own PID. |
|
||||
| `ProcessCwdLookup` | interface | Abstracts PID → cwd lookup. |
|
||||
| `LsofProcessCwdLookup` | class | `lsof`-based cwd lookup for spawn inheritance. |
|
||||
|
||||
**Tool surface:** `bridge_send` (delegate + block; with `turnId`, answers an ask) · `bridge_reply`
|
||||
**Tool surface:** `fleet_send` (delegate + block; with `turnId`, answers an ask) · `fleet_reply`
|
||||
(worker → structured answer; if no send is open it is now **held in the reply inbox**, not
|
||||
errored) · `bridge_ask` (worker pauses to ask primary) · `bridge_status` · `bridge_poll` (async
|
||||
errored) · `fleet_ask` (worker pauses to ask primary) · `fleet_status` · `fleet_poll` (async
|
||||
ticket; with an optional `target`, **drains that worker's held replies** from the inbox) ·
|
||||
`bridge_ack` (ack one held reply by `msgId` — finer than drain-all) · `bridge_spawn` ·
|
||||
`bridge_list` · `bridge_stop` · `bridge_profiles`.
|
||||
**Invariant:** `bridge_reply`/`bridge_ask` accept *no* identity argument; it comes only from the
|
||||
transport context. **CB-307 scope:** only a *terminal* `bridge_reply` with no open send is queued —
|
||||
`bridge_ask` (interactive; the worker blocks and can't consume a late answer) and the injector's
|
||||
`fleet_ack` (ack one held reply by `msgId` — finer than drain-all) · `fleet_spawn` ·
|
||||
`fleet_list` · `fleet_stop` · `fleet_profiles`.
|
||||
**Invariant:** `fleet_reply`/`fleet_ask` accept *no* identity argument; it comes only from the
|
||||
transport context. **CB-307 scope:** only a *terminal* `fleet_reply` with no open send is queued —
|
||||
`fleet_ask` (interactive; the worker blocks and can't consume a late answer) and the injector's
|
||||
completion/failure fallbacks (they target a *captured* waiter, CB-116) are **never** queued.
|
||||
|
||||
### `rest` · `member` · `lead` · `guard` · `config` — the edge
|
||||
|
||||
| Class | Kind | Role |
|
||||
|---|---|---|
|
||||
| `rest.BridgedApp` | class | Javalin routes; validates bodies, maps `Outcome` → HTTP status. |
|
||||
| `rest.FleetdApp` | class | Javalin routes; validates bodies, maps `Outcome` → HTTP status. |
|
||||
| `member.ClaudeCodeLauncher` | class | Guard-checked spawn, orphan-pane reaping at boot, teardown. The first-class `PeerLauncher` adapter for Claude Code over herdr (`spawn`, `reapOrphanWorkers`, `stop`, `list`, `profiles`, `capabilities`). |
|
||||
| `guard.SubscriptionGuard` | class | Host-allowlist + primary-cleanliness enforcement (`assertWorker`, `assertPrimaryClean`). |
|
||||
| `guard.GuardException` | class | Thrown on any subscription-boundary violation. |
|
||||
| `config.BridgedConfig` | record | YAML config with defaults; single legacy worker or named `workers` map (`load`, `workerProfiles`, `defaultProfile`). |
|
||||
| `Bridged` | class | Static `main` — wires real collaborators and starts the server. |
|
||||
| `config.FleetdConfig` | record | YAML config with defaults; single legacy worker or named `workers` map (`load`, `workerProfiles`, `defaultProfile`). |
|
||||
| `Fleetd` | class | Static `main` — wires real collaborators and starts the server. |
|
||||
|
||||
**REST routes** (the acceptance surface — every capability is reachable here without MCP):
|
||||
|
||||
@@ -221,16 +221,16 @@ completion/failure fallbacks (they target a *captured* waiter, CB-116) are **nev
|
||||
| `GET /profiles` | Configured worker profiles + default |
|
||||
| `POST /workers` | Spawn a worker (`?profile=`, `?cwd=`, or JSON body) |
|
||||
| `DELETE /workers/{paneId}` | Stop a worker |
|
||||
| `POST /sessions/{id}/message` | `bridge_send` — blocking, `wait:false`, or answer via `turnId` |
|
||||
| `POST /sessions/{id}/reply` | `bridge_reply` (worker) |
|
||||
| `POST /sessions/{id}/ask` | `bridge_ask` (worker → primary) |
|
||||
| `POST /sessions/{id}/message` | `fleet_send` — blocking, `wait:false`, or answer via `turnId` |
|
||||
| `POST /sessions/{id}/reply` | `fleet_reply` (worker) |
|
||||
| `POST /sessions/{id}/ask` | `fleet_ask` (worker → primary) |
|
||||
| `GET /sessions/{id}/status` | Worker lifecycle + MCP readiness |
|
||||
| `GET /sessions/{id}/replies` | Drain a worker's **held replies** from the inbox (CB-307): peek-then-ack, second call returns `[]` |
|
||||
| `GET /tasks/{ticket}` | Poll an async (`wait:false`) send |
|
||||
|
||||
## Bootstrap wiring
|
||||
|
||||
`Bridged.main` assembles the object graph in dependency order, then starts both faces.
|
||||
`Fleetd.main` assembles the object graph in dependency order, then starts both faces.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
@@ -240,19 +240,19 @@ flowchart LR
|
||||
ccl --> rv["Rendezvous → CompletionResolver<br/>→ Injector + StatusPoller"]
|
||||
rv --> ms["MessageService"]
|
||||
ms --> mcp["BridgeMcp<br/>(connection identity)"]
|
||||
ms --> app["BridgedApp<br/>(start Javalin)"]
|
||||
ms --> app["FleetdApp<br/>(start Javalin)"]
|
||||
mcp --> app
|
||||
```
|
||||
|
||||
*Figure 2 — startup wiring in `Bridged.main`. The guard asserts the primary env is clean before
|
||||
*Figure 2 — startup wiring in `Fleetd.main`. The guard asserts the primary env is clean before
|
||||
anything else; `PeerLauncher` is wired as an interface, with `ClaudeCodeLauncher` as the first
|
||||
adapter; `reapOrphanWorkers()` clears stale panes from a prior daemon restart.*
|
||||
|
||||
## Flow: forward rendezvous (`bridge_send` → `bridge_reply`)
|
||||
## Flow: forward rendezvous (`fleet_send` → `fleet_reply`)
|
||||
|
||||
The main path. A primary's send blocks on a per-session waiter that resolves on the worker's
|
||||
explicit reply — **or**, as a fallback, on a confirmed `working → idle` turn boundary the injector
|
||||
observes (so a worker that finishes without calling `bridge_reply` still unblocks the caller).
|
||||
observes (so a worker that finishes without calling `fleet_reply` still unblocks the caller).
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
@@ -270,7 +270,7 @@ sequenceDiagram
|
||||
IJ->>W: inject when injectable (one msg/turn)
|
||||
IJ-->>MS: onDelivered (capture waiter + baseline)
|
||||
alt worker replies explicitly
|
||||
W->>RV: bridge_reply → resolve(REPLY)
|
||||
W->>RV: fleet_reply → resolve(REPLY)
|
||||
else turn completes without reply
|
||||
IJ-->>RV: onTurnComplete → resolveCompletion(COMPLETION)
|
||||
else worker fails / wedges
|
||||
@@ -284,11 +284,11 @@ sequenceDiagram
|
||||
(delivered) or `TIMED_OUT_QUEUED` (never delivered).*
|
||||
|
||||
**The stranded-reply case (CB-307).** The figure is the happy path — a live waiter exists. The
|
||||
failure that motivated CB-307 is the *reverse*: a worker finishes and calls `bridge_reply` **after**
|
||||
failure that motivated CB-307 is the *reverse*: a worker finishes and calls `fleet_reply` **after**
|
||||
its send has already timed out (the ~60s client window) or was never opened, so `Rendezvous.resolve`
|
||||
finds no waiter. Before CB-307 that reply was silently discarded (the worker got an error / REST
|
||||
`409`). Now `MessageService.reply` publishes it to the `ReplyInbox` under a fresh `msgId`, and the
|
||||
primary collects it later keyed by target — `bridge_poll(target)` or `GET /sessions/{id}/replies`
|
||||
primary collects it later keyed by target — `fleet_poll(target)` or `GET /sessions/{id}/replies`
|
||||
(peek → deliver → ack, so an in-flight failure re-surfaces it). With `broker:` configured the
|
||||
`AmqpReplyInbox` gives this **cross-restart durability** (Stage 2) — the held reply survives a
|
||||
`java -jar` bounce and is redelivered; without it the in-memory adapter is soft-state (undrained
|
||||
@@ -296,7 +296,7 @@ replies clear on restart).
|
||||
|
||||
Delivery is no longer purely pull. When the reply lands with no waiter, `reply` also calls
|
||||
`ReplyPushLoop.onReplyQueued(target)`, which **actively nudges the primary to drain** (Stage 3): it
|
||||
injects a *"run `bridge_poll(target=…)`"* turn into the primary's own herdr pane — the same
|
||||
injects a *"run `fleet_poll(target=…)`"* turn into the primary's own herdr pane — the same
|
||||
`AgentControl.send` primitive that delivers to workers, pointed at the primary — but only when the
|
||||
primary is `injectable()` (never mid-turn), bounded to `primary.push_reminders` nudges on a
|
||||
`push_backoff_ms` schedule, and stopping the instant a drain empties the inbox. The nudge carries
|
||||
@@ -304,7 +304,7 @@ primary is `injectable()` (never mid-turn), bounded to `primary.push_reminders`
|
||||
push fails the durable inbox is still the backstop. An off-host / non-herdr primary leaves
|
||||
`PrimaryRegistry` empty, so the loop is a no-op and delivery cleanly degrades to pull.
|
||||
|
||||
## Flow: reverse rendezvous (`bridge_ask` → answer)
|
||||
## Flow: reverse rendezvous (`fleet_ask` → answer)
|
||||
|
||||
A worker pauses its own turn to ask the primary a question; the question surfaces on the primary's
|
||||
open forward send, and the answer resumes the *same* worker turn. Duplicate asks from one session
|
||||
@@ -328,7 +328,7 @@ sequenceDiagram
|
||||
P->>MS: answer(turnId, content)
|
||||
MS->>RV: answerAsk(turnId, content) → unblock worker
|
||||
MS->>RV: open(workerSession) — new forward waiter
|
||||
W-->>RV: resumes turn → bridge_reply
|
||||
W-->>RV: resumes turn → fleet_reply
|
||||
RV-->>P: Reply
|
||||
else no forward send open
|
||||
RV-->>W: AskOutcome.NO_WAITER (closeAsk)
|
||||
@@ -413,7 +413,7 @@ in `ClaudeCodeLauncher.spawn()` **before any herdr call**: the worker's `ANTHROP
|
||||
be on the allowlist (Stage-1: `gx00.gw`, `ollama.ltms.dev`). `assertPrimaryClean` (called at
|
||||
startup) hard-stops if the primary env carries any `ANTHROPIC_BASE_URL`. Spawn injects
|
||||
`ANTHROPIC_BASE_URL`/`ANTHROPIC_MODEL`/`CLAUDE_CONFIG_DIR`/`ANTHROPIC_AUTH_TOKEN` into the
|
||||
**worker's** env only — `bridged`'s own env is never mutated. CWD resolves
|
||||
**worker's** env only — `fleetd`'s own env is never mutated. CWD resolves
|
||||
`requestedCwd → profile cwd → caller cwd → daemon user.dir → "."`.
|
||||
|
||||
## Concurrency model (at a glance)
|
||||
@@ -435,7 +435,7 @@ startup) hard-stops if the primary env carries any `ANTHROPIC_BASE_URL`. Spawn i
|
||||
## Related pages
|
||||
|
||||
- [1. Architecture](1-Architecture) — system, invariants, the two modes (design level).
|
||||
- [2. Message Server](2-Message-Server) — the `bridged` design rationale, herdr control contract.
|
||||
- [2. Message Server](2-Message-Server) — the `fleetd` design rationale, herdr control contract.
|
||||
- [8. Roadmap](8-Roadmap) — stages, tickets, and the feature ⇄ endpoint ⇄ test map.
|
||||
|
||||
---
|
||||
|
||||
+17
-17
@@ -12,23 +12,23 @@ A **subscription-safe bridge** that lets a lead **Claude Code** session on Pro/M
|
||||
> inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper model. It also runs
|
||||
> non-Claude members (`kind: opencode`) as first-class peers.
|
||||
|
||||
## Leading approach — herdr-centric message server (`bridged`)
|
||||
## Leading approach — herdr-centric message server (`fleetd`)
|
||||
|
||||
A small always-on message server, **`bridged`**, controls
|
||||
A small always-on message server, **`fleetd`**, controls
|
||||
[herdr](https://herdr.dev) (an agent multiplexer, "tmux for agents") over its Unix-socket
|
||||
API and exposes a clean 2-way messaging API as an **MCP server that both the primary and the
|
||||
workers mount** — one unified Claude setup and the **sole communication gateway** (REST/SSE
|
||||
stays for non-Claude clients; any broker is `bridged`-internal, below the gateway). herdr owns
|
||||
stays for non-Claude clients; any broker is `fleetd`-internal, below the gateway). herdr owns
|
||||
the PTYs, multiplexing, persistence, and **agent-status
|
||||
events**; `bridged` owns policy (subscription boundary, session lifecycle, status-gated
|
||||
events**; `fleetd` owns policy (subscription boundary, session lifecycle, status-gated
|
||||
delivery) and the client contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at
|
||||
the gateway, `https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and
|
||||
calls `bridged`'s MCP tools.
|
||||
calls `fleetd`'s MCP tools.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
|
||||
subgraph BD["bridged — standalone daemon (not a claude process)"]
|
||||
subgraph BD["fleetd — standalone daemon (not a claude process)"]
|
||||
SRV["SERVER face<br/>MCP · REST/SSE · policy"]
|
||||
CLI["CLIENT face<br/>status-gated injector · herdr socket"]
|
||||
SRV --> CLI
|
||||
@@ -37,8 +37,8 @@ flowchart LR
|
||||
W["member pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
|
||||
M["llm.ltms.dev<br/>(the one gateway)"]
|
||||
|
||||
OPUS -->|"MCP bridge_send (blocks)"| SRV
|
||||
W -.->|"MCP bridge_reply"| SRV
|
||||
OPUS -->|"MCP fleet_send (blocks)"| SRV
|
||||
W -.->|"MCP fleet_reply"| SRV
|
||||
CLI -->|"Unix socket<br/>send_text · events.subscribe"| HERDR
|
||||
HERDR -->|"drives PTY"| W
|
||||
W -->|"inference"| M
|
||||
@@ -50,28 +50,28 @@ flowchart LR
|
||||
```
|
||||
|
||||
- **Subscription boundary:** the *lead* never sets `ANTHROPIC_BASE_URL` (stays on Pro/Max). Only a
|
||||
member process is moved off-subscription, and only by the daemon, at spawn; `bridged` is a plain
|
||||
member process is moved off-subscription, and only by the daemon, at spawn; `fleetd` is a plain
|
||||
daemon with no Anthropic quota, so it may poll and subscribe freely. **One exception, and it
|
||||
costs money:** a profile marked `subscription: true` runs its members on *your* plan on purpose —
|
||||
see [13 User Guide](13-User-Guide) → *The knobs that cost money*. An `opencode` member sits
|
||||
outside this boundary entirely and uses its own provider credential.
|
||||
- **One gateway:** `bridged` is the **sole communication path** for every session. The lead and
|
||||
- **One gateway:** `fleetd` is the **sole communication path** for every session. The lead and
|
||||
every member mount it as an MCP server and talk *only* to it — **no session ever addresses a
|
||||
broker, a peer, or the network directly.** How the mount happens differs by backend: a
|
||||
`claude-code` member gets one `claude mcp add` line or a shared `.mcp.json`, while an `opencode`
|
||||
member gets a generated `opencode.json` and no `ANTHROPIC_*` variables at all. Any broker or
|
||||
queue is `bridged`-internal, below the gateway — and it is optional. With `broker:` commented
|
||||
queue is `fleetd`-internal, below the gateway — and it is optional. With `broker:` commented
|
||||
out the daemon uses an in-memory inbox; when it is on, it is **LavinMQ** over AMQP.
|
||||
- **How the lead gets a reply:** in the simple case, **one blocking MCP tool call** (`bridge_send`)
|
||||
that `bridged` holds open until the member replies (`bridge_reply`) or its turn completes, then
|
||||
- **How the lead gets a reply:** in the simple case, **one blocking MCP tool call** (`fleet_send`)
|
||||
that `fleetd` holds open until the member replies (`fleet_reply`) or its turn completes, then
|
||||
returns the reply as the tool result. In practice that call is capped by the lead's own MCP client
|
||||
timeout (about 60 seconds), so real work uses `wait:false` and a ticket instead. Either way the
|
||||
lead never busy-polls and never touches a broker: when a detached ticket goes terminal, `bridged`
|
||||
lead never busy-polls and never touches a broker: when a detached ticket goes terminal, `fleetd`
|
||||
**injects the lead's own idle pane** to wake it.
|
||||
- **Different model** per member process sidesteps Claude Code's lack of per-subagent
|
||||
provider routing — the worker isn't a subagent, it's its own configured process.
|
||||
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) was kept on paper as a
|
||||
swappable *fallback injector*. It was **never built** — `grep -ri agentapi bridged/src/main`
|
||||
swappable *fallback injector*. It was **never built** — `grep -ri agentapi fleetd/src/main`
|
||||
returns nothing, and the only injection path in the shipped code is the herdr one. Treat it as a
|
||||
discarded option, not a fallback you can switch to. See [Approaches](3-Approaches) for why herdr
|
||||
won and [Message Server](2-Message-Server) for the full design.
|
||||
@@ -81,7 +81,7 @@ flowchart LR
|
||||
Read in order (the sidebar mirrors this):
|
||||
|
||||
1. **[Architecture](1-Architecture)** — process model, the two invariants, two traffic modes
|
||||
2. **[Message Server](2-Message-Server)** — 🟢 **`bridged`**, the herdr-centric message server (primary approach)
|
||||
2. **[Message Server](2-Message-Server)** — 🟢 **`fleetd`**, the herdr-centric message server (primary approach)
|
||||
3. **[Approaches](3-Approaches)** — herdr-centric vs AgentAPI vs Agent SDK vs bus/tmux (research matrix)
|
||||
4. **[Setup](4-Setup)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §2** instead.
|
||||
5. **[Operations](5-Operations)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §4 and §6** instead.
|
||||
@@ -100,7 +100,7 @@ Read in order (the sidebar mirrors this):
|
||||
the daemon live on this host. Chapters 4 and 5 were never written past their scope note; chapter 13
|
||||
replaced them.
|
||||
|
||||
The **herdr-centric `bridged` message server** was selected on 2026-07-11, superseding the AgentAPI
|
||||
The **herdr-centric `fleetd` message server** was selected on 2026-07-11, superseding the AgentAPI
|
||||
plan of 2026-07-08. AgentAPI is retained as a fallback injector and has not been needed. See
|
||||
**[Message Server](2-Message-Server)** for the design and **[13 User Guide](13-User-Guide)** for how
|
||||
to run it.
|
||||
|
||||
+2
-2
@@ -5,7 +5,7 @@
|
||||
**Chapters**
|
||||
|
||||
1. [Architecture](1-Architecture) — system · 2 invariants · 2 modes
|
||||
2. [Message Server](2-Message-Server) — the `bridged` design
|
||||
2. [Message Server](2-Message-Server) — the `fleetd` design
|
||||
3. [Approaches](3-Approaches) — transports compared, why herdr
|
||||
4. [Setup](4-Setup) — ⚫ superseded by 13
|
||||
5. [Operations](5-Operations) — ⚫ superseded by 13
|
||||
@@ -19,4 +19,4 @@
|
||||
13. **[User Guide](13-User-Guide)** — 🟢 install · configure · run · delegate · the traps
|
||||
|
||||
---
|
||||
🟢 herdr-centric `bridged` · AgentAPI = fallback
|
||||
🟢 herdr-centric `fleetd` · AgentAPI = fallback
|
||||
|
||||
Reference in New Issue
Block a user