.models.
(Claude Code, subscription)"]
- subgraph BD["bridged — a plain Java daemon"]
+ subgraph BD["fleetd — a plain Java daemon"]
SRV["MCP + REST
policy, authz, lifecycle"]
INJ["injector
status-gated"]
SRV --> INJ
@@ -29,9 +29,9 @@ flowchart LR
M2["member pane
opencode"]
GW["llm.ltms.dev
the one gateway"]
- LEAD -->|"bridge_send"| SRV
- M1 -.->|"bridge_reply"| SRV
- M2 -.->|"bridge_reply"| SRV
+ LEAD -->|"fleet_send"| SRV
+ M1 -.->|"fleet_reply"| SRV
+ M2 -.->|"fleet_reply"| SRV
INJ -->|"unix socket"| HERDR
HERDR --> M1
HERDR --> M2
@@ -61,7 +61,7 @@ flowchart LR
1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the
subscription. Only the daemon moves a member off it, at spawn.
2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
- in a `bridge_*` call is discarded silently.
+ in a `fleet_*` call is discarded silently.
---
@@ -71,7 +71,7 @@ Four things must be true before anything works. In order, because each one depen
### 2.1 herdr
-`bridged` does not own terminals. `herdr` does. `bridged` drives it over a Unix socket.
+`fleetd` does not own terminals. `herdr` does. `fleetd` drives it over a Unix socket.
```bash
herdr --version # 0.8.0 on this host
@@ -79,7 +79,7 @@ curl -s http://127.0.0.1:8765/healthz
# {"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
```
-**The protocol number is the thing to check, not the version.** `bridged`'s adapter speaks one
+**The protocol number is the thing to check, not the version.** `fleetd`'s adapter speaks one
herdr wire protocol. If herdr is upgraded and the protocol moves, `/healthz` still says `ok` —
because the socket connects — and **every spawn fails**. Green health with a broken fleet is the
normal way this breaks. See trap 2 in §6.
@@ -89,8 +89,8 @@ normal way this breaks. See trap 2 in §6.
Build and run from this repo. There is no POM at the repo root:
```bash
-mvn -f bridged/pom.xml clean install # never pipe this — a pipe hides BUILD FAILURE
-java -jar bridged/target/bridged.jar bridged/bridged.yaml
+mvn -f fleetd/pom.xml clean install # never pipe this — a pipe hides BUILD FAILURE
+java -jar fleetd/target/fleetd.jar fleetd/fleetd.yaml
```
In practice you never run those two by hand. Use the script (§4).
@@ -107,7 +107,7 @@ shell**. This one fact causes more lost hours than anything else in the system.
request.
`launchd` does not run a login shell either. That is the only reason
-`scripts/bridged-launchd-wrapper.sh` exists — it `exec`s `zsh -lc` so the store gets sourced, while
+`scripts/fleetd-launchd-wrapper.sh` exists — it `exec`s `zsh -lc` so the store gets sourced, while
keeping one process so launchd's PID tracking still works. Read its header; it explains the trap
better than this paragraph.
@@ -115,7 +115,7 @@ better than this paragraph.
> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing.
The daemon logs which required secret names resolved at startup
-(`Bridged.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
+(`Fleetd.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
`subscription: true`**, on purpose, because they need no token. So a green secret report says
nothing about your subscription profiles.
@@ -142,7 +142,7 @@ lead was worse than a daemon that will not boot.
## 3. Configure
-`bridged/bridged.yaml` is the live config. It is **gitignored**. `bridged.example.yaml` is the
+`fleetd/fleetd.yaml` is the live config. It is **gitignored**. `fleetd.example.yaml` is the
tracked, documented copy. Two consequences you will meet:
- Members cannot see the live config. They work in worktrees of the tracked repo. So a change to
@@ -220,10 +220,10 @@ not assumed.
Use the script. Do not hand-roll the steps.
```bash
-scripts/redeploy-bridged.sh --check # read-only: reports state, changes nothing
-scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
-scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
-scripts/redeploy-bridged.sh --no-build # restart the jar you already have
+scripts/redeploy-fleetd.sh --check # read-only: reports state, changes nothing
+scripts/redeploy-fleetd.sh # build, confirm drain, restart, verify
+scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already checked)
+scripts/redeploy-fleetd.sh --no-build # restart the jar you already have
```
Run `--check` first, always. It is the only thing that reports whether the forge token resolves,
@@ -237,8 +237,8 @@ checks to a marker taken before the restart, so old errors cannot be misread as
`main` does nothing until you rebuild and restart. Saying "shipped" about code the live daemon has
never loaded is a false report.
-**Drain first.** `bridge_list`, collect anything you still want with `bridge_poll`, then
-`bridge_stop` each member. A restart drops in-flight tickets and rendezvous. A member's report is
+**Drain first.** `fleet_list`, collect anything you still want with `fleet_poll`, then
+`fleet_stop` each member. A restart drops in-flight tickets and rendezvous. A member's report is
not recoverable once its ticket is gone.
### Verify — `/healthz` is not enough
@@ -246,18 +246,18 @@ not recoverable once its ticket is gone.
`/healthz` proves the socket connects. It does not prove a spawn works, that identity resolves, or
that the new jar is the one running. Four checks, in order:
-1. **A fresh boot line.** Confirm a new `bridged listening` line at the end of `bridged/bridged.out`,
+1. **A fresh boot line.** Confirm a new `fleetd listening` line at the end of `fleetd/fleetd.out`,
dated after the restart. An old daemon that never died looks identical from outside.
2. **Deferred keys.** The startup log names which config keys it accepted and which it deferred. A
deferred key needing a restart is usually the whole reason you restarted. Read those lines rather
than assuming.
-3. **Identity.** `bridge_whoami` must still answer `primary`. If the tab label changed, the lead is
+3. **Identity.** `fleet_whoami` must still answer `primary`. If the tab label changed, the lead is
now a worker and every orchestration call is refused.
4. **A real spawn.** Spawn one cheap member and stop it. This is the only check that catches a herdr
protocol mismatch.
> Restarting the daemon **cuts your own MCP mount**, and it does not reconnect. So you cannot run
-> `bridge_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
+> `fleet_whoami` from the session that restarted it. Ask the operator to run `/mcp` to reconnect.
> This is why the restart is done from the lead but verified after a reconnect.
### The REST surface
@@ -273,20 +273,20 @@ loopback only.
| `GET /agents` | agents as herdr sees them |
| `GET /members` · `POST /members` · `DELETE /members/{paneId}` | list, spawn, tear down |
| `GET /profiles` | the backends configured |
-| `GET /sessions/{id}/status` | one session — the same view as `bridge_status` |
+| `GET /sessions/{id}/status` | one session — the same view as `fleet_status` |
| `POST /sessions/{id}/message` · `/reply` · `/ask` | the three message kinds |
| `GET /sessions/{id}/replies` | drain the reply inbox — **destructive, see trap 6** |
| `GET /tasks/{ticket}` | poll a detached ticket |
### Where it runs, and where the logs are
-- Log file: **`bridged/bridged.out`**, in both supervised and unsupervised modes.
-- Audit log: `bridged/logs/audit.log`, rotated daily, 30 days kept.
-- Service units ship in `deploy/`: `bridged.service` for Linux systemd (ordered
- `After=herdr.service`) and `dev.ltms.bridged.plist` for macOS launchd. `deploy/lavinmq` holds the
+- Log file: **`fleetd/fleetd.out`**, in both supervised and unsupervised modes.
+- Audit log: `fleetd/logs/audit.log`, rotated daily, 30 days kept.
+- Service units ship in `deploy/`: `fleetd.service` for Linux systemd (ordered
+ `After=herdr.service`) and `dev.ltms.fleetd.plist` for macOS launchd. `deploy/lavinmq` holds the
optional broker.
- **On this host neither is loaded.** The daemon runs as a plain `java -jar` started by
- `scripts/redeploy-bridged.sh` from a login shell. Verified with `launchctl list | grep bridg`,
+ `scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep bridg`,
which returns nothing. If you expected launchd here, that expectation is the bug.
---
@@ -297,36 +297,36 @@ Eleven tools. This is the whole surface.
| Intent | Tool |
|---|---|
-| Confirm your own role | `bridge_whoami` |
-| See backends available | `bridge_profiles` |
-| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` |
-| See the fleet | `bridge_list` → `leads` + `members` |
-| One member's state | `bridge_status{sessionId}` |
-| Delegate, blocking | `bridge_send{sessionId, content}` |
-| Delegate, long task | `bridge_send{sessionId, content, wait:false}` → ticket |
-| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
-| Collect a reply | `bridge_poll{ticket}` or `bridge_poll{target}` |
-| Clear a reply from the inbox | `bridge_ack{target, msgId}` — **`target`, not `ticket`** |
-| Member ends its turn | `bridge_reply{content}` |
-| Member asks the lead | `bridge_ask{question}` |
-| Tear down | `bridge_stop{paneId}` |
+| Confirm your own role | `fleet_whoami` |
+| See backends available | `fleet_profiles` |
+| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` |
+| See the fleet | `fleet_list` → `leads` + `members` |
+| One member's state | `fleet_status{sessionId}` |
+| Delegate, blocking | `fleet_send{sessionId, content}` |
+| Delegate, long task | `fleet_send{sessionId, content, wait:false}` → ticket |
+| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
+| Collect a reply | `fleet_poll{ticket}` or `fleet_poll{target}` |
+| Clear a reply from the inbox | `fleet_ack{target, msgId}` — **`target`, not `ticket`** |
+| Member ends its turn | `fleet_reply{content}` |
+| Member asks the lead | `fleet_ask{question}` |
+| Tear down | `fleet_stop{paneId}` |
### The loop
```mermaid
sequenceDiagram
participant L as lead
- participant B as bridged
+ participant B as fleetd
participant M as member
- L->>B: bridge_spawn — all units first
- L->>B: bridge_send with wait false — then all sends
+ L->>B: fleet_spawn — all units first
+ L->>B: fleet_send with wait false — then all sends
B-->>L: ticket
B->>M: injected when the member is idle
- M->>B: bridge_reply
+ M->>B: fleet_reply
B-->>L: nudge into the lead's own pane
- L->>B: bridge_poll by ticket
- L->>B: bridge_ack by target and msgId
- L->>B: bridge_stop by paneId
+ L->>B: fleet_poll by ticket
+ L->>B: fleet_ack by target and msgId
+ L->>B: fleet_stop by paneId
```
*Figure 2 — the delegation loop. Spawn is separate from send on purpose.*
@@ -334,7 +334,7 @@ sequenceDiagram
**Spawn every unit first, then send them all.** Spawning and sending in one loop is how parallel
work silently becomes serial. It is the most common way this whole layer is wasted.
-**Prefer `wait:false`.** A blocking `bridge_send` is capped by **your own MCP client timeout**,
+**Prefer `wait:false`.** A blocking `fleet_send` is capped by **your own MCP client timeout**,
about 60 seconds — far below any real task's runtime. The cap is in your client, not in the daemon,
so no server setting fixes it.
@@ -379,7 +379,7 @@ Twelve traps, all hit for real this year. Grouped by where they bite.
**1. The daemon started from a non-login shell.**
Symptom: everything green for hours, then a member cannot open a pull request, or a gateway profile
-gets a 401. Nothing logs it at the time. Fix: `scripts/redeploy-bridged.sh --check` is the only
+gets a 401. Nothing logs it at the time. Fix: `scripts/redeploy-fleetd.sh --check` is the only
thing that reports it. Restart from a login shell, or via the launchd wrapper.
**2. `/healthz` is green and every spawn fails.**
@@ -394,15 +394,15 @@ removed; a current daemon refuses to start if it is still present.
**4. A merge is not a deployment.**
The running daemon holds its original jar. Rebuild and restart, then prove it with a **fresh**
-`bridged listening` line. Also: the restart cuts your own MCP mount and it never reconnects, so you
+`fleetd listening` line. Also: the restart cuts your own MCP mount and it never reconnects, so you
cannot verify identity from that session — ask for `/mcp`.
### Losing a member's work
**5. The ticket expired.**
-Ticket time-to-live is about 10 minutes. After that `bridge_poll{ticket}` returns
+Ticket time-to-live is about 10 minutes. After that `fleet_poll{ticket}` returns
`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's
-**inbox**. Drain it with `bridge_poll{target}`, then `bridge_ack{target, msgId}`.
+**inbox**. Drain it with `fleet_poll{target}`, then `fleet_ack{target, msgId}`.
**6. Reading the reply inbox is destructive.**
`GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that
@@ -419,24 +419,24 @@ queued and restarts the member when it next goes idle. Back up the member's comm
**9. `liveStatus: working` is not progress.**
A member can report busy for hours doing nothing. Check file modification times in its worktree and
-snapshot `git diff` before and after. Do not read its terminal — use `bridge_status`.
+snapshot `git diff` before and after. Do not read its terminal — use `fleet_status`.
**10. Never brief a member to "ask me".**
-`bridge_ask` blocks for about 55 seconds and no nudge extends it. It is also invisible to
-`bridge_poll`. If you are running async, you will not see the question in time. Decide before you
+`fleet_ask` blocks for about 55 seconds and no nudge extends it. It is also invisible to
+`fleet_poll`. If you are running async, you will not see the question in time. Decide before you
delegate, or give the member an explicit default.
### Merging a member's work
**11. The reported branch is not the branch it committed to.**
-`bridge_list` reports the branch **spawn provisioned**, not the one the member actually used. Check
+`fleet_list` reports the branch **spawn provisioned**, not the one the member actually used. Check
`git -C
not a keystroke into the primary pane"
+ S-->>P: "tool result = reply (unparks fleet_send)"
+ Note over P,W: "worker → primary rode fleetd's state,
not a keystroke into the primary pane"
```
-*Figure: the reply resolves on whichever of the two signals lands first; `bridge_reply`
+*Figure: the reply resolves on whichever of the two signals lands first; `fleet_reply`
carries the structured payload, the herdr event guarantees the timing even if the worker
never cooperates.*
@@ -135,8 +135,8 @@ it degrades cleanly if a worker is left unmodified:
| Tier | Worker setup | Reply channel | Worker can ask back? |
|---|---|---|---|
-| **Unified (recommended)** | mounts `bridge` MCP (same one line) | structured `bridge_reply` | yes — `bridge_ask` |
-| **Hooked (no MCP)** | a `Stop`-hook installed | structured envelope POSTed to `bridged` (see [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply)) | no |
+| **Unified (recommended)** | mounts `bridge` MCP (same one line) | structured `fleet_reply` | yes — `fleet_ask` |
+| **Hooked (no MCP)** | a `Stop`-hook installed | structured envelope POSTed to `fleetd` (see [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply)) | no |
| **Unmodified (last resort)** | stock `claude` | `pane.read` scrape on the turn-done `working→idle` edge (lossy) | no |
The **primary-side contract is identical** in both tiers; only the worker's reply fidelity
@@ -147,7 +147,7 @@ bridge on either side is subscription-safe by construction (see
## Architecture
-`bridged` is a **standalone daemon** — one component, two faces:
+`fleetd` is a **standalone daemon** — one component, two faces:
- a **SERVER face** — an **MCP server** the Claude Code sessions mount, plus REST/SSE for
non-Claude clients, sitting over the session tracker, subscription guard, and reply
@@ -156,8 +156,8 @@ bridge on either side is subscription-safe by construction (see
to agent-status.
The Claude sessions themselves live **as panes inside herdr**. Each pane reaches *up* to
-`bridged`'s MCP server (to send/reply); `bridged`'s herdr client reaches *down* through
-herdr's socket to drive those same panes and read their status. herdr, `bridged`, the panes,
+`fleetd`'s MCP server (to send/reply); `fleetd`'s herdr client reaches *down* through
+herdr's socket to drive those same panes and read their status. herdr, `fleetd`, the panes,
and the model are colocated on one host (same-host scenario); a broker is optional for async
duplex.
@@ -168,9 +168,9 @@ flowchart TB
WP["worker pane(s) · claude
ANTHROPIC_BASE_URL set · MCP client"]
end
- subgraph bridged["bridged — standalone daemon (NOT a claude process)"]
+ subgraph fleetd["fleetd — standalone daemon (NOT a claude process)"]
subgraph srv["SERVER face"]
- MCP["MCP server
bridge_send · reply · ask · status"]
+ MCP["MCP server
fleet_send · reply · ask · status"]
REST["REST / SSE
(non-Claude clients)"]
POL["policy brain
session tracker · subscription guard
· reply rendezvous"]
end
@@ -201,13 +201,13 @@ flowchart TB
class MODEL,BROKER warn
```
-*Figure: `bridged` is one standalone daemon with a **SERVER** face (the MCP endpoint the
+*Figure: `fleetd` is one standalone daemon with a **SERVER** face (the MCP endpoint the
Claude panes mount, over the policy brain) and a **CLIENT** face (the herdr socket client).
-The Claude sessions are herdr **panes**: they call *up* into the MCP server, while `bridged`'s
+The Claude sessions are herdr **panes**: they call *up* into the MCP server, while `fleetd`'s
client drives them *down* through herdr's socket and gates every injection on live
-agent-status. Only worker panes carry `ANTHROPIC_BASE_URL`; `bridged` holds no quota, so it
-subscribes freely. The broker is **`bridged`-internal** (durability / cross-host) — no pane
-ever addresses it; async delivery is `bridged` injecting an idle pane. (Split-host moves the
+agent-status. Only worker panes carry `ANTHROPIC_BASE_URL`; `fleetd` holds no quota, so it
+subscribes freely. The broker is **`fleetd`-internal** (durability / cross-host) — no pane
+ever addresses it; async delivery is `fleetd` injecting an idle pane. (Split-host moves the
primary out of herdr — see [Deployment model](#deployment-model).)*
### Components
@@ -217,12 +217,12 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
| **herdr socket client** | NDJSON over `~/.config/herdr/herdr.sock`; correlates responses by `id`; maintains a long-lived `events.subscribe` stream. |
| **Session manager** | Maps a logical session → herdr `workspace/tab/pane` id. Spawns the worker `claude` (env-prefixed launch line into a fresh pane's shell), health-checks, and **recycles on context ceiling** (Ralph loop, see below). |
| **Injector** | Per-pane FIFO queue. Delivers `send_text` + `send_keys "enter"` **only when** that pane's `agent_status ∈ {idle, blocked}` — never mid-run. |
-| **Reply rendezvous** | Resolves an awaiting `bridge_send` on whichever lands first: a worker `bridge_reply` (structured, **preferred**), the turn-done `working → idle` status edge (timing guarantee), or — worker-side hook path — a `Stop`-hook envelope; last-resort `pane.read {source:"recent-unwrapped"}` scrape. See [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply). |
+| **Reply rendezvous** | Resolves an awaiting `fleet_send` on whichever lands first: a worker `fleet_reply` (structured, **preferred**), the turn-done `working → idle` status edge (timing guarantee), or — worker-side hook path — a `Stop`-hook envelope; last-resort `pane.read {source:"recent-unwrapped"}` scrape. See [Reply envelope](#reply-envelope-how-a-worker-emits-a-structured-reply). |
| **Subscription guard** | Refuses to spawn a *worker* pane without `ANTHROPIC_BASE_URL`; refuses to *ever* set it on a pane designated *primary*; can assert egress host via `pane.process_info`. |
-| **SERVER API** | **MCP server** — the Claude-facing contract both primary and workers mount (`bridge_send`/`reply`/`ask`/`status`/`poll`/`sessions`). Plus **REST + SSE** (OpenAPI, AgentAPI-shaped) for non-Claude clients — webhooks, dashboards, a human CLI. |
-| **Broker connector** *(optional, internal)* | `bridged`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `bridged` will later inject into an idle pane. No Claude session ever connects to it. |
+| **SERVER API** | **MCP server** — the Claude-facing contract both primary and workers mount (`fleet_send`/`reply`/`ask`/`status`/`poll`/`sessions`). Plus **REST + SSE** (OpenAPI, AgentAPI-shaped) for non-Claude clients — webhooks, dashboards, a human CLI. |
+| **Broker connector** *(optional, internal)* | `fleetd`-owned durability + cross-host transport, **below the gateway**. Enqueues async messages `fleetd` will later inject into an idle pane. No Claude session ever connects to it. |
-## The herdr control contract (what `bridged` drives)
+## The herdr control contract (what `fleetd` drives)
> **Verified against the running herdr 0.7.0 (protocol 14), not just docs.** A spike hit the
> live socket. `ping` returns `{version:"0.7.0", protocol:14, capabilities:{live_handoff:true}}`;
@@ -233,11 +233,11 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
> `agent.get`, `agent.focus`, plus `server.agent_manifests` and `pane.report_agent` — so herdr
> already models "agents", not just panes. (Also available: `worktree.*` for git-isolated workers.)
>
-> **CB-102 — decided: `bridged` drives the native `agent.*` path** (the pane + `send_text`
+> **CB-102 — decided: `fleetd` drives the native `agent.*` path** (the pane + `send_text`
> workaround below is the documented fallback only). The spike proved against the live daemon that
> `agent.start` accepts a **first-class `env` map** that reaches the process environment — so a
> worker's `ANTHROPIC_BASE_URL` is injected cleanly and guard-checked, with no shell-prefix
-> parsing and `bridged`'s own env untouched. herdr also tracks each worker's **Claude session
+> parsing and `fleetd`'s own env untouched. herdr also tracks each worker's **Claude session
> UUID** (`agent_session.value`), grounding the ID contract in herdr's own identity. Observed
> `agent.*` schema:
>
@@ -255,8 +255,8 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
>
> **CB-108 — worker placement: a tab per worker in a dedicated "worker space."** By default
> `agent.start` (no `tab_id`) *splits the currently-focused tab*, so it would clutter — and could
-> co-tenant — the user's real work tabs. Instead `bridged` puts every worker in its **own tab**
-> inside a dedicated **workspace** (`worker.workspace`, default `bridged-workers`), found-or-created
+> co-tenant — the user's real work tabs. Instead `fleetd` puts every worker in its **own tab**
+> inside a dedicated **workspace** (`worker.workspace`, default `fleetd-workers`), found-or-created
> once and shared (a future per-session layout is just a distinct label). The recipe, all verified
> live: `workspace.create {label}` (idempotent find-first) → `tab.create {workspace_id}` →
> `agent.start {…, tab_id}` → `pane.close {root_pane}` (drop herdr's seed shell so the tab holds
@@ -266,7 +266,7 @@ primary out of herdr — see [Deployment model](#deployment-model).)*
> the legacy split behaviour.
>
> **Two more `agent.*` facts the CB-108 build pinned:** (1) an agent's **`name` must be unique**
-> among running agents — a 2nd `agent.start {name:"claude"}` fails `agent_name_taken`, so `bridged`
+> among running agents — a 2nd `agent.start {name:"claude"}` fails `agent_name_taken`, so `fleetd`
> names workers `claude-
split-host primary: its Stop-hook polls bridged, never the broker"
+ Note over S,R: "recipient polled nothing — fleetd pushed on the idle edge.
split-host primary: its Stop-hook polls fleetd, never the broker"
```
-*Figure: `bridged` mediates async in both directions. Every source reaches it over the gateway
+*Figure: `fleetd` mediates async in both directions. Every source reaches it over the gateway
(MCP for Claude, REST for external), it optionally parks the message on its **internal** queue,
-waits for the recipient's idle event, and injects. The queue is a `bridged` implementation
+waits for the recipient's idle event, and injects. The queue is a `fleetd` implementation
detail — no Claude session touches it, and herdr scrollback is never the source of truth.*
## Worker session lifecycle
-`bridged` treats a worker as a **recyclable** resource, not one immortal session — a
-long-lived pane fills its context window. When idle-cycle or token caps trip, `bridged`
+`fleetd` treats a worker as a **recyclable** resource, not one immortal session — a
+long-lived pane fills its context window. When idle-cycle or token caps trip, `fleetd`
kills the pane and respawns fresh (**Ralph loop**).
**What "state on disk" actually means (be precise — this is easy to hand-wave).** Recycling
@@ -438,7 +438,7 @@ the worker externalizes as it works**:
continue from it,"* not with the old messages.
This only works if the worker is disciplined about writing that state **before** a recycle
-boundary. `bridged` can enforce a checkpoint (inject "commit and update STATE.md" before it
+boundary. `fleetd` can enforce a checkpoint (inject "commit and update STATE.md" before it
kills the pane), but a worker that ignores it loses in-flight context. **Open question:**
whether to also snapshot the raw transcript (`--resume` the *same* session on crash-restart,
vs. a clean context on a planned recycle) — the two restart reasons may want different
@@ -450,7 +450,7 @@ stateDiagram-v2
Spawning --> Ready: "claude prompt detected"
Ready --> Working: "turn injected"
Working --> Blocked: "permission / question"
- Blocked --> Working: "bridged answers (send_input)"
+ Blocked --> Working: "fleetd answers (send_input)"
Working --> Ready: "agent_status → idle (turn done)"
Ready --> Recycling: "context / idle cap hit"
Recycling --> Spawning: "state persisted to disk"
@@ -459,7 +459,7 @@ stateDiagram-v2
Ready --> [*]: "drain / shutdown"
```
-*Figure: the state machine `bridged` drives per worker. `blocked`, the turn-done
+*Figure: the state machine `fleetd` drives per worker. `blocked`, the turn-done
`working → idle` edge, and `pane.exited` are real herdr signals, not heuristics — the reason
herdr-centric beats screen scraping.*
@@ -468,7 +468,7 @@ herdr-centric beats screen scraping.*
The invariant is unchanged from [Architecture](1-Architecture) — **anything that sets
`ANTHROPIC_BASE_URL` is, by definition, the worker** — but here it is *enforced in code*:
-- `bridged` is **not** a `claude` process. It consumes zero Anthropic quota, so it may
+- `fleetd` is **not** a `claude` process. It consumes zero Anthropic quota, so it may
busy-poll the broker and hold a permanent herdr event subscription with no policy concern.
- The **subscription guard** blocks any spawn of a *worker* pane whose launch does not
resolve an **off-subscription** `ANTHROPIC_BASE_URL`, and blocks any attempt to set that
@@ -481,14 +481,14 @@ The invariant is unchanged from [Architecture](1-Architecture) — **anything th
would pass a naive check while burning subscription-adjacent auth). The guard must
validate the **resolved value's host** against an allowlist and, after spawn, confirm
egress via `pane.process_info` / a health call to the worker model — not trust the string.
- - The guard only constrains panes **`bridged` spawns**. A split-host primary on your Mac is
- a process `bridged` never sees; it cannot inspect that env. There, subscription safety
+ - The guard only constrains panes **`fleetd` spawns**. A split-host primary on your Mac is
+ a process `fleetd` never sees; it cannot inspect that env. There, subscription safety
rests on the operator (the Mac `claude` simply is never given the var) plus the fact that
- the *only* thing crossing to the worker host is **MCP/HTTP traffic to `bridged`**, never an
+ the *only* thing crossing to the worker host is **MCP/HTTP traffic to `fleetd`**, never an
endpoint swap and never a broker connection.
- Injecting keystrokes into the **primary** pane is subscription-safe: it is simulated
typing, identical to the human at the keyboard — the primary still talks to
- `api.anthropic.com` on Pro/Max. `bridged` never re-points the primary's endpoint. (This
+ `api.anthropic.com` on Pro/Max. `fleetd` never re-points the primary's endpoint. (This
path exists only single-host, where the primary is a herdr pane.)
- A startup self-check asserts any **locally-hosted** primary pane's env has **no**
`ANTHROPIC_BASE_URL` and logs each worker's resolved egress host. It cannot self-check a
@@ -497,8 +497,8 @@ The invariant is unchanged from [Architecture](1-Architecture) — **anything th
## API surface (SERVER face)
**Two faces over one core — and REST is the testability surface.** Every feature is implemented
-as a **REST route**; the **MCP tools are thin adapters over those routes** — e.g. `bridge_send`
-→ `POST /sessions/{id}/message`, `bridge_status` → `GET /sessions/{id}/status`, `bridge_reply` →
+as a **REST route**; the **MCP tools are thin adapters over those routes** — e.g. `fleet_send`
+→ `POST /sessions/{id}/message`, `fleet_status` → `GET /sessions/{id}/status`, `fleet_reply` →
`POST /sessions/{id}/reply` (rendezvous). The REST routes are also the surface non-Claude clients
use (webhooks, dashboards, a human CLI), kept AgentAPI-shaped for drop-in migration. Because the
logic lives in REST, **each feature is acceptance-tested by an HTTP call with no Claude/MCP in
@@ -522,15 +522,15 @@ the loop**, and MCP is verified by a **parity test** (tool result == REST result
1. **Subscription-safe delegation (the core case).** Primary Opus offloads bulk/routine
work — codegen, refactors, test writing, log triage — to a worker on a cheap/local model,
keeping Opus's context clean and its quota for review/merge decisions.
-2. **A herd of specialized workers.** One `bridged` + herdr multiplexes several workers
+2. **A herd of specialized workers.** One `fleetd` + herdr multiplexes several workers
(e.g. a DeepSeek coder, a fast summarizer, a long-context reader), each its own pane,
each addressable by `session_id`. Rolls up to a single status sidebar.
-3. **Async event-bus automation.** A webhook/CI/NATS event hits `bridged`'s **REST ingress**;
- `bridged` wakes an idle worker by injection, and the result flows back to the primary (or a
- Slack/Telegram bridge) — all mediated by `bridged`, no human in the loop. The source never
+3. **Async event-bus automation.** A webhook/CI/NATS event hits `fleetd`'s **REST ingress**;
+ `fleetd` wakes an idle worker by injection, and the result flows back to the primary (or a
+ Slack/Telegram bridge) — all mediated by `fleetd`, no human in the loop. The source never
addresses a worker or a broker directly.
4. **Human co-pilot from anywhere.** Because herdr persists and detaches, the same worker is
- reachable from a phone/chat bridge **posting to `bridged`** while you're away, and from the
+ reachable from a phone/chat bridge **posting to `fleetd`** while you're away, and from the
attached TUI when you're back.
5. **Long-running "perpetual" workers.** The Ralph-loop lifecycle lets a worker run for hours
across many context recycles without a human respawning it, state carried on disk.
@@ -539,16 +539,16 @@ the loop**, and MCP is verified by a **parity test** (tool result == REST result
### Single-host (default — simplest, recommended to start)
-Primary, `bridged`, herdr, and workers all on one off-subscription box. `bridged` can inject
+Primary, `fleetd`, herdr, and workers all on one off-subscription box. `fleetd` can inject
into **both** the primary and the worker panes (both local to one herdr), so the broker is
optional. Best for a workstation or a single dev box.
### Split-host (primary local, workers remote near the model)
-Primary Opus runs on your Mac; `bridged` + herdr + workers run on the GPU host next to
+Primary Opus runs on your Mac; `fleetd` + herdr + workers run on the GPU host next to
`ollama.ltms.dev` / GX10 vLLM. herdr's socket is **local-only**, so the Mac reaches the worker
-host **only over `bridged`'s MCP/HTTP endpoint** — never a remote herdr socket and never the
-broker (the broker, if any, stays `bridged`-internal on the worker host).
+host **only over `fleetd`'s MCP/HTTP endpoint** — never a remote herdr socket and never the
+broker (the broker, if any, stays `fleetd`-internal on the worker host).
```mermaid
flowchart LR
@@ -557,10 +557,10 @@ flowchart LR
end
subgraph host["Worker host (off-subscription, near model)"]
direction TB
- BD["bridged
:8080 MCP · REST/SSE"]
+ BD["fleetd
:8080 MCP · REST/SSE"]
HS["herdr server"]
W2["worker claude pane(s)"]
- BR["broker / queue
(bridged-owned, internal)"]
+ BR["broker / queue
(fleetd-owned, internal)"]
BD -->|"Unix socket"| HS --> W2
BD -.->|"durability / cross-host"| BR
end
@@ -577,22 +577,22 @@ flowchart LR
class BR warn
```
-*Figure: the Mac's **only** link to the worker host is `bridged`'s MCP/HTTP endpoint — it
+*Figure: the Mac's **only** link to the worker host is `fleetd`'s MCP/HTTP endpoint — it
carries sync replies, and (since the Mac primary isn't a herdr pane) its `Stop`-hook polls that
-same endpoint for async wake-ups. herdr's socket and the broker stay local and `bridged`-owned.
+same endpoint for async wake-ups. herdr's socket and the broker stay local and `fleetd`-owned.
The subscription boundary tracks the host boundary — nothing on the Mac ever sets
`ANTHROPIC_BASE_URL`.*
-**Security:** bind `bridged`'s HTTP to `localhost` and reach it over an SSH tunnel, or front
+**Security:** bind `fleetd`'s HTTP to `localhost` and reach it over an SSH tunnel, or front
it with a bearer token + TLS. Never expose the port unauthenticated — it is an agent-control
surface: `POST /message` runs arbitrary prompts, and `POST /keys` sends raw keystrokes
(including `ctrl+c`) into a live agent. (This closes the gap left by AgentAPI's open `:3284`.)
-**Multi-tenancy is an open item.** A single `bridged` fronting a *herd* of workers today has
+**Multi-tenancy is an open item.** A single `fleetd` fronting a *herd* of workers today has
**one shared token = full control of every session**; there is no per-session authorization.
That is acceptable for a single-operator box but not for shared/multi-user use. Before that,
add per-session scoping (a capability token per `session_id`) and an audit log of injected
-turns. Until then, treat one `bridged` as one trust domain.
+turns. Until then, treat one `fleetd` as one trust domain.
## Proposed tech stack
@@ -603,10 +603,10 @@ turns. Until then, treat one `bridged` as one trust domain.
| **herdr transport** | JDK **`UnixDomainSocketAddress` + `SocketChannel`** (native UDS, no dep), **NDJSON** via Jackson, `id`-correlated + a persistent events stream | Native herdr contract; a dedicated **virtual thread** blocks on the event stream. | — |
| **SERVER API — Claude** | **Official MCP Java SDK** (streamable-HTTP transport) on an embedded server (Jetty / Spring Boot) | The unified contract both primary and workers mount; native to Claude Code, no shell/`curl`, subscription-safe by construction | stdio MCP adapter (per-session subprocess) if a long-lived HTTP endpoint is undesirable |
| **SERVER API — others** | **REST + SSE** via **Javalin** (light) or Spring MVC | Drop-in for AgentAPI-shaped/non-Claude clients; SSE streams status cheaply | JAX-RS (Helidon/Quarkus); gRPC if callers are all code |
-| **Internal queue** *(optional)* | **Redis Streams** via **Lettuce** (consumer groups, `XACK`, visibility timeout) — `bridged`-owned, below the gateway | Durability + cross-host for async; satisfies the guardrails in [Architecture](1-Architecture). Same-host can start with an in-JVM queue and add this only when durability/cross-host is needed | NATS JetStream for multi-host scale; embedded H2/SQLite for a single host |
+| **Internal queue** *(optional)* | **Redis Streams** via **Lettuce** (consumer groups, `XACK`, visibility timeout) — `fleetd`-owned, below the gateway | Durability + cross-host for async; satisfies the guardrails in [Architecture](1-Architecture). Same-host can start with an in-JVM queue and add this only when durability/cross-host is needed | NATS JetStream for multi-host scale; embedded H2/SQLite for a single host |
| **Config** | **YAML via Jackson** (`jackson-dataformat-yaml`) + env overrides | 12-factor; secrets via env only | MicroProfile Config; Spring config if on Spring Boot |
| **Observability** | **SLF4J + Logback**; **Micrometer** → Prometheus `/metrics`; `/healthz` | Ops from day one | OpenTelemetry traces |
-| **Process supervision** | **systemd** unit (`java -jar` or the native-image binary), ordered after herdr | Restart-on-crash; ordered start (herdr before `bridged`) | Docker Compose colocating herdr + `bridged`; k8s (overkill for one host) |
+| **Process supervision** | **systemd** unit (`java -jar` or the native-image binary), ordered after herdr | Restart-on-crash; ordered start (herdr before `fleetd`) | Docker Compose colocating herdr + `fleetd`; k8s (overkill for one host) |
| **Testing** | **JUnit 5** + a **mock UDS socket** server + golden transcripts; fake `ccs`/`claude` stubs | Deterministic CI without a real TTY | Testcontainers (Redis) + a real herdr for e2e |
**Recommendation: Java 21+ with virtual threads.** The daemon is almost entirely
@@ -652,8 +652,8 @@ final class Guard {
throw new GuardException("worker base_url host %s is not an approved off-subscription host".formatted(host));
}
- // Only meaningful for a primary bridged itself hosts (single-host). A remote/Mac primary is a
- // process bridged never sees — its cleanliness is the operator's.
+ // Only meaningful for a primary fleetd itself hosts (single-host). A remote/Mac primary is a
+ // process fleetd never sees — its cleanliness is the operator's.
void assertLocalPrimaryClean(Map
delivered into a LIVE session?"}
- Q -->|"structured socket → terminal, + status events"| HD["herdr socket
(via bridged)"]
+ Q -->|"structured socket → terminal, + status events"| HD["herdr socket
(via fleetd)"]
Q -->|"HTTP request → terminal emulation"| AA["AgentAPI
(coder/agentapi)"]
Q -->|"streaming input generator (in-process)"| SDK["Agent SDK
streaming query()"]
Q -->|"worker PULLS on idle via a hook"| BUS["Message-queue
+ Stop-hook long-poll"]
Q -->|"raw keystrokes into the tmux pane"| TMUX["tmux send-keys
/ PTY paste"]
HD --> V0["✅ our leading choice
(status events + multiplex;
symmetric single-host)"]
- AA --> V1["◐ fallback injector
(swappable behind bridged)"]
+ AA --> V1["◐ fallback injector
(swappable behind fleetd)"]
SDK --> V2["✅ if driver is our own code"]
BUS --> V3["⚠ async bus events only
(lands at turn boundary)"]
TMUX --> V4["⚠ fragile — the raw primitive
herdr/AgentAPI productize"]
@@ -44,34 +44,34 @@ explicitly-requested, still-open feature:
and [#24947 — `claude inject`](https://github.com/anthropics/claude-code/issues/24947).
Every approach below is a way around that gap. **herdr wins because it productizes the
sturdiest workaround (terminal automation) as a structured socket API that *also* streams
-agent-status events** — so `bridged` gets reliable completion/blocked signals instead of
+agent-status events** — so `fleetd` gets reliable completion/blocked signals instead of
scraping a screen.
-## 1. herdr socket API via `bridged` — structured injection + status events *(leading)*
+## 1. herdr socket API via `fleetd` — structured injection + status events *(leading)*
[herdr](https://herdr.dev) is a persistent agent multiplexer (a "tmux for agents") with a
-Unix-socket JSON API. `bridged` (see [Message Server](2-Message-Server)) drives it: `pane.send_text` +
+Unix-socket JSON API. `fleetd` (see [Message Server](2-Message-Server)) drives it: `pane.send_text` +
`pane.send_keys` deliver a turn into the *running* pane, and `events.subscribe`
(`pane.agent_status_changed`) reports **working / blocked / idle** as real events (turn-done
= the `working → idle` edge; herdr has no `done` status).
-- **Injects into a live session:** yes into the worker. **Worker → primary** rides `bridged`'s
- **MCP rendezvous** — the worker's `bridge_reply` (or the turn-done idle edge) resolves the primary's
- blocking `bridge_send` tool call, so no keystroke into the primary pane is needed, even
+- **Injects into a live session:** yes into the worker. **Worker → primary** rides `fleetd`'s
+ **MCP rendezvous** — the worker's `fleet_reply` (or the turn-done idle edge) resolves the primary's
+ blocking `fleet_send` tool call, so no keystroke into the primary pane is needed, even
single-host. Fallbacks: herdr can type into a single-host non-MCP primary (subscription-safe
- keystrokes); a split-host primary wakes via its own `Stop`-hook polling `bridged` (the async
+ keystrokes); a split-host primary wakes via its own `Stop`-hook polling `fleetd` (the async
path — Mode 2 in [Architecture](1-Architecture)).
- **North-face contract is MCP — and the sole gateway.** Both primary and workers mount
- `bridged` as an MCP server (one unified Claude setup); no Claude session ever addresses a
+ `fleetd` as an MCP server (one unified Claude setup); no Claude session ever addresses a
broker or peer directly. The herdr injection here is the *south* side, orthogonal to it.
- **Completion signal:** structured events — not the screen-stability *guess* AgentAPI makes.
The worker can even `pane.report_agent` its own state via herdr's `SKILL.md`.
- **Multiplex + persist:** a herd of workers as addressable panes; headless server survives
detach/reattach over SSH.
-- **Subscription-safe:** only the worker pane launches with `ANTHROPIC_BASE_URL`; `bridged`
+- **Subscription-safe:** only the worker pane launches with `ANTHROPIC_BASE_URL`; `fleetd`
is a plain daemon (no quota) that enforces the boundary in code.
-- **Trade-off:** herdr's socket is **local-only** (`bridged`'s MCP/HTTP spans hosts, not
- herdr), and it is a young, single-dev project — so `bridged` keeps the injector **pluggable**
+- **Trade-off:** herdr's socket is **local-only** (`fleetd`'s MCP/HTTP spans hosts, not
+ herdr), and it is a young, single-dev project — so `fleetd` keeps the injector **pluggable**
and its durability in an **internal** queue behind the gateway. Replies are best carried as a
structured envelope, not scraped.
@@ -81,13 +81,13 @@ Unix-socket JSON API. `bridged` (see [Message Server](2-Message-Server)) drives
HTTP server and drives the CLI's terminal underneath (essentially a hardened, stateful
`tmux send-keys` with parsing): `POST /message`, `GET /events` (SSE), `GET /status`. It was
the original leading choice; herdr now supersedes it, but it remains a **swappable fallback
-injector** behind `bridged`'s interface.
+injector** behind `fleetd`'s interface.
- **Injects into a live session:** yes — but only the *worker* (it wraps one CLI); the
- primary direction still needs `bridged`'s async path (idle-injection, or a split-host
- `Stop`-hook polling `bridged`).
+ primary direction still needs `fleetd`'s async path (idle-injection, or a split-host
+ `Stop`-hook polling `fleetd`).
- **Completion signal:** a **screen-stability heuristic**, not structured events.
-- **Cross-host:** native HTTP — its one edge over herdr, but `bridged` already provides the
+- **Cross-host:** native HTTP — its one edge over herdr, but `fleetd` already provides the
HTTP layer on top of herdr, so that edge is neutralized.
- **When to reach for it:** if herdr can't run, or as the second injector implementation to
de-risk herdr's immaturity. Its `msgfmt` reply parser is worth reusing regardless.
@@ -107,11 +107,11 @@ events and permission callbacks instead of scraping a terminal.
The **only pure-hooks** way to pull an external message into the **same** session. In
`claude-bridge` this is **not** how a Claude session normally receives async work — under the
-sole-gateway rule `bridged` delivers async by **injecting an idle pane**, and no Claude session
+sole-gateway rule `fleetd` delivers async by **injecting an idle pane**, and no Claude session
polls a queue. The Stop-hook survives in exactly one place: a **split-host primary** that isn't
-a herdr pane, where the hook long-polls **`bridged`** (not the queue) for wake-ups. The raw
+a herdr pane, where the hook long-polls **`fleetd`** (not the queue) for wake-ups. The raw
mechanism below is shown for the comparison; note the poll target is the gateway, and the queue
-itself sits *behind* `bridged`.
+itself sits *behind* `fleetd`.
```mermaid
sequenceDiagram
@@ -142,7 +142,7 @@ sequenceDiagram
[DIY pattern](https://claudefa.st/blog/tools/hooks/stop-hook-task-enforcement)).
- **Decisive limitation:** **pull-on-idle only** — the message lands at a **turn boundary**,
never mid-turn. Fine for "bus event wakes an idle worker"; wrong for "interrupt a busy
- worker." (`bridged` injecting into an idle pane has the same idle-boundary property, but
+ worker." (`fleetd` injecting into an idle pane has the same idle-boundary property, but
with a real status gate.)
- **Lighter cousin:** a `UserPromptSubmit` hook returning `{"additionalContext":"…"}` —
fires only *when a prompt is submitted*, so it can't deliver an async push.
@@ -167,7 +167,7 @@ status events, and multiplexing — which is why we build on herdr rather than h
| Approach | Transport | Inject into running session? | Completion signal | Symmetric (both panes)? | Cross-host | Subscription-safe | Fragility |
|---|---|---|---|---|---|---|---|
-| **herdr via `bridged`** ✅ | socket → terminal + events | ✅ (idle-gated) | ✅ status events² | ◐ single-host¹ | via `bridged` MCP/HTTP (sole gateway) | ✅ (guard in code) | Low–Med (herdr young) |
+| **herdr via `fleetd`** ✅ | socket → terminal + events | ✅ (idle-gated) | ✅ status events² | ◐ single-host¹ | via `fleetd` MCP/HTTP (sole gateway) | ✅ (guard in code) | Low–Med (herdr young) |
| **AgentAPI** ◐ (fallback) | HTTP → terminal emulation | ✅ (worker only) | ⚠ screen-stability | ❌ | ✅ native HTTP | ✅ (worker-only env) | Low |
| **Agent SDK streaming** | in-process generator | ✅ | ✅ typed events | n/a | ✅ | ✅ | Low (driver = code) |
| **Queue + Stop-hook** | hook long-poll | ✅ (turn boundary) | via turn end | ✅ (symmetric) | ✅ | ✅ | Medium |
@@ -176,25 +176,25 @@ status events, and multiplexing — which is why we build on herdr rather than h
¹ **Symmetric only single-host.** herdr can type into *either* pane, but the primary is a
herdr pane only when it runs on the herdr host. In the split-host target (primary on a Mac),
-worker→primary goes through `bridged` all the same — the primary's `Stop`-hook long-polls
-`bridged` (never a broker) for the wake-up. The "one mechanism, both directions via herdr"
-story holds for a single-box setup; across hosts the gateway is still `bridged`, just not by
+worker→primary goes through `fleetd` all the same — the primary's `Stop`-hook long-polls
+`fleetd` (never a broker) for the wake-up. The "one mechanism, both directions via herdr"
+story holds for a single-box setup; across hosts the gateway is still `fleetd`, just not by
injection.
-² **"Completion signal" for `bridged` is the timing signal; reply *content* rides the worker's
-structured `bridge_reply` (preferred), with a `Stop`-hook envelope as fallback** (see
+² **"Completion signal" for `fleetd` is the timing signal; reply *content* rides the worker's
+structured `fleet_reply` (preferred), with a `Stop`-hook envelope as fallback** (see
[Message Server](2-Message-Server)) — not the status event itself.
## Recommendation
-- **Primary Opus → worker (the bridge's main path):** **herdr via `bridged`** —
+- **Primary Opus → worker (the bridge's main path):** **herdr via `fleetd`** —
status-gated injection, structured completion/blocked events, symmetric (single-host), multiplexed,
persistent, with the subscription boundary enforced in code. Selected. See
[Message Server](2-Message-Server) / [Architecture](1-Architecture).
-- **Keep AgentAPI as a swappable fallback injector** behind `bridged`'s interface, so
+- **Keep AgentAPI as a swappable fallback injector** behind `fleetd`'s interface, so
herdr's immaturity is a de-riskable risk rather than a load-bearing one.
-- **External event bus → worker (async wake-ups):** the bus hits **`bridged`'s REST ingress**;
- `bridged` enqueues internally if needed and **injects the idle worker** — the worker runs no
+- **External event bus → worker (async wake-ups):** the bus hits **`fleetd`'s REST ingress**;
+ `fleetd` enqueues internally if needed and **injects the idle worker** — the worker runs no
queue-polling hook. Complementary to the sync path, not a replacement — different trigger
shape, same single gateway.
- **Avoid hand-rolled `tmux send-keys`** unless neither herdr nor AgentAPI can run; it's the
diff --git a/4-Setup.md b/4-Setup.md
index e8e78f8..08508fb 100644
--- a/4-Setup.md
+++ b/4-Setup.md
@@ -11,7 +11,7 @@
- herdr, and why you check the **protocol number** rather than the version;
- building and starting the daemon;
- the **login shell** rule for `${SHARED_ENV}/tools/secrets.sh` — the single most expensive trap in
- bring-up, and why `scripts/bridged-launchd-wrapper.sh` exists;
+ bring-up, and why `scripts/fleetd-launchd-wrapper.sh` exists;
- the lead's **tab label**, which is how the daemon resolves who the lead is.
## The one rule that was already right on this page
diff --git a/5-Operations.md b/5-Operations.md
index 405899c..abb69b3 100644
--- a/5-Operations.md
+++ b/5-Operations.md
@@ -7,7 +7,7 @@
**Go to [13 User Guide](13-User-Guide):**
-- **§4 Run, and prove it runs** — `scripts/redeploy-bridged.sh`, why a merge is not a deployment,
+- **§4 Run, and prove it runs** — `scripts/redeploy-fleetd.sh`, why a merge is not a deployment,
draining before a restart, and the four checks that go beyond `/healthz`.
- **§6 When it breaks** — twelve traps hit for real this year, grouped by bring-up, losing a
member's work, and merging a member's work.
diff --git a/6-Team.md b/6-Team.md
index b9944f2..ea9b457 100644
--- a/6-Team.md
+++ b/6-Team.md
@@ -1,23 +1,23 @@
# 6. Team
-The [Message Server](2-Message-Server) (`bridged`) delivers **one turn into one worker**. A
+The [Message Server](2-Message-Server) (`fleetd`) delivers **one turn into one worker**. A
**team** is the layer above it: a **Claude team-lead** that fans a job out across a **mixed
fleet** of workers — some on Claude, some on the remote local LLM — and reduces their replies.
-Same `bridged` delivery, same subscription boundary; this page is only about **orchestration**
+Same `fleetd` delivery, same subscription boundary; this page is only about **orchestration**
— who the workers are, how the lead picks one, and how it runs many at once.
-> Delivery mechanics (blocking `bridge_send` MCP call, status-gated reply) live in
+> Delivery mechanics (blocking `fleet_send` MCP call, status-gated reply) live in
> [Message Server](2-Message-Server). Transport rationale is in [Approaches](3-Approaches).
> This page assumes both.
## The team
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
- worker; an **MCP client of `bridged`** (it mounts the bridge like everyone else). It plans,
- routes, dispatches via `bridge_send`, and integrates — and never sets `ANTHROPIC_BASE_URL`.
-- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session
+ worker; an **MCP client of `fleetd`** (it mounts the bridge like everyone else). It plans,
+ routes, dispatches via `fleet_send`, and integrates — and never sets `ANTHROPIC_BASE_URL`.
+- **Workers** — a herd of `claude` panes in herdr, each an addressable `fleetd` session
with its **own model/env**, each also **mounting the bridge MCP** (unified setup — they
- reply via `bridge_reply`):
+ reply via `fleet_reply`):
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
embarrassingly parallel subtasks.
@@ -31,7 +31,7 @@ alike. Scale each kind horizontally by adding panes.
```mermaid
flowchart TB
LEAD["lead — Opus
(Claude Code, env CLEAN)
MCP client"]
- subgraph BD["bridged — standalone daemon"]
+ subgraph BD["fleetd — standalone daemon"]
SRV["SERVER face
MCP · REST/SSE · policy brain"]
CLI["CLIENT face
herdr socket"]
SRV --> CLI
@@ -44,11 +44,11 @@ flowchart TB
ANT["api.anthropic.com
(Pro/Max)"]
OLL["ollama.ltms.dev
(local model)"]
- LEAD -->|"MCP bridge_send (target role)"| SRV
+ LEAD -->|"MCP fleet_send (target role)"| SRV
CLI -->|"Unix socket · send_text · events.subscribe"| HERDR
HERDR --> WC1 & WC2 & WL1 & WL2
- WC1 -.->|"MCP bridge_reply"| SRV
- WL1 -.->|"MCP bridge_reply"| SRV
+ WC1 -.->|"MCP fleet_reply"| SRV
+ WL1 -.->|"MCP fleet_reply"| SRV
WC1 -->|"inference"| ANT
WC2 -->|"inference"| ANT
WL1 -->|"inference"| OLL
@@ -71,14 +71,14 @@ flowchart TB
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
-selection is **policy in the lead**, not a `bridged` concern — `bridged` just delivers to
+selection is **policy in the lead**, not a `fleetd` concern — `fleetd` just delivers to
the session the lead names.
## Subscription boundary in a team
Unchanged from [Architecture](1-Architecture), and it scales with the fleet: **only local-worker panes**
launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on the
-subscription. `bridged` enforces which panes may carry the off-subscription env, so adding
+subscription. `fleetd` enforces which panes may carry the off-subscription env, so adding
workers never widens the boundary.
## Parallel fan-out (map / reduce)
@@ -89,48 +89,48 @@ different workers at once, then results are gathered.
```mermaid
sequenceDiagram
participant L as "lead (Opus)"
- participant B as "bridged"
+ participant B as "fleetd"
participant WC as "w-claude-1"
participant WL as "w-local-1"
Note over L: "split job → subtask A (reasoning), subtask B (bulk)"
par A → Claude worker
- L->>B: "bridge_send {role: w-claude, prompt: A}"
+ L->>B: "fleet_send {role: w-claude, prompt: A}"
B->>WC: "send_text into running pane"
- WC-->>B: "bridge_reply (or idle edge)"
+ WC-->>B: "fleet_reply (or idle edge)"
B-->>L: "tool result reply A"
and B → local worker
- L->>B: "bridge_send {role: w-local, prompt: B}"
+ L->>B: "fleet_send {role: w-local, prompt: B}"
B->>WL: "send_text into running pane"
- WL-->>B: "bridge_reply (or idle edge)"
+ WL-->>B: "fleet_reply (or idle edge)"
B-->>L: "tool result reply B"
end
Note over L: "reduce → integrate A + B into final answer"
```
-- **Map:** the lead issues N concurrent blocking `bridge_send` tool calls (one per subtask →
- its chosen worker). Each call blocks only *that* request; `bridged` holds it open until the
+- **Map:** the lead issues N concurrent blocking `fleet_send` tool calls (one per subtask →
+ its chosen worker). Each call blocks only *that* request; `fleetd` holds it open until the
worker's turn completes (status-gated) and returns the reply as the tool result.
- **Reduce:** the lead collects the N replies and integrates. A slow local worker never
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
-- **Detached / long jobs** use `bridged`'s async path instead of a held request — the result
- is delivered when ready by `bridged` injecting the lead's idle pane (Mode 2 in
- [Architecture](1-Architecture)). The lead talks only to `bridged`, never a broker, and never
+- **Detached / long jobs** use `fleetd`'s async path instead of a held request — the result
+ is delivered when ready by `fleetd` injecting the lead's idle pane (Mode 2 in
+ [Architecture](1-Architecture)). The lead talks only to `fleetd`, never a broker, and never
busy-polls across turns.
-Fan-out is bounded by the herd size (pane count) and `bridged`'s concurrency policy, not by
+Fan-out is bounded by the herd size (pane count) and `fleetd`'s concurrency policy, not by
the lead.
## Knowing the roster
-The lead discovers its team from `bridged` (session list / roles) rather than hard-coding
+The lead discovers its team from `fleetd` (session list / roles) rather than hard-coding
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
```markdown
-## Your team (via bridged)
+## Your team (via fleetd)
You are the team-lead. Delegate through the bridge MCP tools — never launch workers yourself.
-Roster: call bridge_list for current sessions/roles.
+Roster: call fleet_list for current sessions/roles.
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
@@ -139,27 +139,27 @@ of them (concurrent blocking sends), THEN gather — never serialize independent
Integrate the reply envelopes; you own the final answer.
```
-`bridge_send` **is** the one tool call — no HTTP to hand-roll. Optionally wrap it in a Claude
+`fleet_send` **is** the one tool call — no HTTP to hand-roll. Optionally wrap it in a Claude
Code skill (`/delegate
worker profiles"] --> PICK["pick profile
gx00-vllm"]
+ CFG["fleetd.yaml
worker profiles"] --> PICK["pick profile
gx00-vllm"]
PICK --> GUARD{"ccs env host
on allowlist?"}
GUARD -->|"no (api.anthropic.com)"| REJ["refuse — would burn subscription"]
GUARD -->|"yes (gx00.ltms.dev)"| SPAWN["herdr: ccs gx00-vllm claude"]
@@ -135,9 +135,9 @@ Every message across the gateway is one envelope. It answers two questions each
| `v` | envelope version (`1`) |
| `from` | sender identity — `primary:opus`, or `role@profile` / session id for a worker |
| `to` | recipient — `role@profile` or a live session id |
-| `session` | `bridged` session id (the worker session), e.g. `s_7f3a` |
+| `session` | `fleetd` session id (the worker session), e.g. `s_7f3a` |
| `turn` | monotonic counter within the session |
-| `corr` | correlation id (`session#turn`) — `bridged`'s rendezvous matches a reply to its request |
+| `corr` | correlation id (`session#turn`) — `fleetd`'s rendezvous matches a reply to its request |
| `kind` | **verb.noun** telling the recipient what to do (vocabulary below) |
| `body` | kind-specific payload |
@@ -148,11 +148,11 @@ Every message across the gateway is one envelope. It answers two questions each
| `review.request` | primary → worker | `{workspace, base, head\|diff, files?, focus[], instructions}` |
| `review.reply` | worker → primary | `{summary, findings:[{file,line,severity,issue,suggestion}], verdict}` |
| `question` / `answer` | either way | free-form follow-up tied to the same `session` |
-| `ask` | worker → primary | worker-initiated blocker (needs a decision) — via `bridge_ask` |
+| `ask` | worker → primary | worker-initiated blocker (needs a decision) — via `fleet_ask` |
| `ack` / `status` | control | delivery/liveness, no new turn |
The worker learns this contract from a **reviewer skill / `CLAUDE.md` snippet** injected at
-spawn ("you are a reviewer; requests arrive as `review.request`; reply with `bridge_reply`
+spawn ("you are a reviewer; requests arrive as `review.request`; reply with `fleet_reply`
`kind: review.reply`"). So both ends know the sender and the required next action without a
held-open conversation — one request in, one structured reply out.
@@ -164,30 +164,30 @@ stays warm across follow-ups, and recyclable so it never outgrows its context wi
- **Spawn-on-demand** — first `review.request` for a workspace spawns the profile's worker.
- **Reuse** — subsequent turns (`question`, next-file review) target the same `session`; its
context carries the code under review.
-- **Recycle (Ralph loop)** — on a context/idle cap `bridged` checkpoints (commit + `STATE.md`)
+- **Recycle (Ralph loop)** — on a context/idle cap `fleetd` checkpoints (commit + `STATE.md`)
and respawns fresh — **not** `claude --resume`. See
[Architecture → Worker lifecycle](1-Architecture#worker-lifecycle--the-ralph-loop).
- **Drain** — an `idle_ttl` (e.g. 20 min idle) or workspace close tears the pane down.
-Policy knobs in `bridged.yaml`: `idle_ttl`, `context_cap`, `max_workers_per_profile`.
+Policy knobs in `fleetd.yaml`: `idle_ttl`, `context_cap`, `max_workers_per_profile`.
## Primary-side directive — *when* to delegate (a `CLAUDE.md` snippet)
Mechanism 4 gives the **worker** a `CLAUDE.md` snippet so it knows how to answer. The **primary**
needs the mirror image: a standing reminder to *reach for the bridge in the first place* instead of
spending subscription tokens on work a cheaper worker could do. The trigger is an environment signal
-— the `bridged` MCP tools being connected (e.g. a `BRIDGED_MCP_URL` marker in the primary's env).
+— the `fleetd` MCP tools being connected (e.g. a `FLEETD_MCP_URL` marker in the primary's env).
When that property is set, this instance is a **bridge primary** and should delegate by default.
The decision the directive encodes:
```mermaid
flowchart TD
- T["a task arrives"] --> Q{"bridge available?
(BRIDGED_MCP_URL set /
bridged MCP connected)"}
+ T["a task arrives"] --> Q{"bridge available?
(FLEETD_MCP_URL set /
fleetd MCP connected)"}
Q -->|"no"| SELF["do it on the primary"]
Q -->|"yes"| J{"needs YOUR judgment,
or bulk / mechanical / parallel?"}
J -->|"judgment / interactive"| SELF
- J -->|"bulk / mechanical / parallel"| DEL["bridge_send → worker"]
+ J -->|"bulk / mechanical / parallel"| DEL["fleet_send → worker"]
classDef self fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef del fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
class SELF self
@@ -197,7 +197,7 @@ flowchart TD
*Figure: an env property (bridge available) flips the default from "do it myself" to "delegate unless
it needs my judgment."*
-This is a *suggestion*, not wiring: `bridged` never edits an agent's `CLAUDE.md` (that would cross
+This is a *suggestion*, not wiring: `fleetd` never edits an agent's `CLAUDE.md` (that would cross
the subscription boundary in the wrong direction). The operator writes it.
### As-built — the `CLAUDE.md` **Bridge communication** section
@@ -212,7 +212,7 @@ Why in `CLAUDE.md` and not somewhere else — the four layers, each with a diffe
| Layer | Carries | Reaches | Cost to the reader |
|---|---|---|---|
-| `REPLY_CHARTER` (`ClaudeCodeLauncher` / `OpenCodeLauncher`) | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, both peer kinds | always in the system prompt |
+| `REPLY_CHARTER` (`ClaudeCodeLauncher` / `OpenCodeLauncher`) | the one rule that must survive with no repo: *end every turn with `fleet_reply`* | every worker, at launch, both peer kinds | always in the system prompt |
| **`CLAUDE.md` → Bridge communication** | protocol invariants + orchestration policy | primary **and** every claude-code worker — tracked in git, so worktrees get it free | always in context |
| `.claude/skills/{implementer,reviewer}` | per-job procedure: commit/push/PR recipe, finding format | a worker told to load it | on demand |
| `docs/MCP-Contract.md` | design detail, flows, error model | anyone who goes looking | on demand |
@@ -222,11 +222,11 @@ must obey it.** Duplicating a rule into a skill is how the skill and the charter
What each part of the section pins down:
-- **Role identification** — **`bridge_whoami`** (below) answers it authoritatively; the section
+- **Role identification** — **`fleet_whoami`** (below) answers it authoritatively; the section
tells the reader to call it rather than infer. A fallback ladder remains for when it is
unreachable: the charter in the system prompt (reliable — the launcher appends it in the same
branch that mounts the MCP, so bridge tools without a charter is not a reachable state); the mount
- name (`mcp__bridged__*` for the primary's `.mcp.json` vs `mcp__bridge__*` for a worker's inline
+ name (`mcp__fleetd__*` for the primary's `.mcp.json` vs `mcp__bridge__*` for a worker's inline
config); `ANTHROPIC_BASE_URL` (one-way — Claude-model workers run clean, so absence proves
nothing); then **fail toward worker**. The two errors are asymmetric: a primary acting as a worker
gets refused by the authz gate — loud and self-correcting — while a worker acting as the primary
@@ -241,27 +241,27 @@ What each part of the section pins down:
turn costs a worker turn, while doing it yourself costs the primary's context and subscription.
Independent units fan out (one worktree worker each, all dispatched `wait:false`, then poll)
rather than serializing. Plus the policy the tool descriptions can't carry: pass `profile:`
- explicitly; prefer `wait:false` + `bridge_poll`, since a blocking `bridge_send` is capped by the
+ explicitly; prefer `wait:false` + `fleet_poll`, since a blocking `fleet_send` is capped by the
*caller's own* MCP client timeout (~60s) long before a real task finishes; make every delegation
self-contained; **name the worker's skill in the first line of `content`** — that instruction is
what turns an opt-in skill into a reliable one; you are the merge gate; verify what a worker
claims rather than trusting a "clean" report. Delegating work never delegates responsibility.
-- **Worker** — the turn contract: load the named skill, stay in scope, `bridge_ask` only for a
- decision that is genuinely the lead's, end with exactly one `bridge_reply`, report only what you
+- **Worker** — the turn contract: load the named skill, stay in scope, `fleet_ask` only for a
+ decision that is genuinely the lead's, end with exactly one `fleet_reply`, report only what you
actually ran, never merge, never commit `.mcp.json` or `wiki/`.
**Known gap:** *opencode* workers never read `CLAUDE.md` — they receive `REPLY_CHARTER` as an
instructions file and nothing else. Any rule a non-Claude peer must obey belongs in the charter, not
in this section. The charter currently carries only the reply rule.
-### `bridge_whoami` — asking instead of guessing
+### `fleet_whoami` — asking instead of guessing
**Shipped.** The daemon always knew the answer: `ConnectionIdentity` maps a call's loopback peer PID
to a herdr pane, and every tool call is already gated on the `Principal` it yields. What was missing
was any way for an agent to *ask* — so an agent's own role had to be inferred from side channels the
daemon does not control, with a silent failure mode when the inference went the wrong way.
-`bridge_whoami` (no params, `READ` in the [authz table](9-Implementation)) returns that same
+`fleet_whoami` (no params, `READ` in the [authz table](9-Implementation)) returns that same
resolved identity as data:
```json
@@ -270,7 +270,7 @@ resolved identity as data:
```
The primary gets `{"role":"primary"}` and nothing more — deliberately: handing it a `sessionId` it
-does not own would invite exactly the forged `bridge_reply` that `Authz` refuses. A worker the
+does not own would invite exactly the forged `fleet_reply` that `Authz` refuses. A worker the
session registry has no record of — one that outlived a daemon restart — still gets `role` and
`sessionId`, which is the load-bearing part; the registry fields are simply absent rather than
invented.
@@ -298,7 +298,7 @@ notes after the block). Improvements land *here* first, then propagate to each p
If no `fleet_*` MCP tools are mounted in this session, this section does not apply — skip it.
-`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
+`fleetd` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `fleet_*` tools. No session addresses a peer, a broker, or the network directly.
@@ -479,7 +479,7 @@ Two rules keep this portable, and both were learned by getting them wrong first:
meaningless in another project. Each moved to a **Project addendum** section that sits *below*
the block and never interleaves with it, so the block can be replaced wholesale without reading it.
2. **Every fallback signal must be one-way.** The role ladder originally read the MCP mount name in
- both directions — `mcp__bridged__*` ⇒ primary, `mcp__bridge__*` ⇒ worker. Only the second half is
+ both directions — `mcp__fleetd__*` ⇒ primary, `mcp__bridge__*` ⇒ worker. Only the second half is
real: the launcher hard-codes `bridge` for a worker's inline config, while a primary's mount is
named by whoever wrote that project's `.mcp.json`. A two-way reading of a one-way signal is a
confident wrong answer, so the block states only the direction that holds.
@@ -498,15 +498,15 @@ Same machinery, different `kind`/lifecycle:
| Use case | Shape |
|---|---|
-| **Delegated refactor / codegen** | `bridge_send(kind: task.request)`, blocking; worker edits + commits on a branch; reply = diff summary. Opus reviews/merges. |
+| **Delegated refactor / codegen** | `fleet_send(kind: task.request)`, blocking; worker edits + commits on a branch; reply = diff summary. Opus reviews/merges. |
| **Test writing / log triage** | Bulk, cheap, parallelizable — a local worker profile; fire several concurrently and reduce. |
| **Parallel multi-file review** | A **fleet** of workers, one per area, fan-out/gather — see [Team](6-Team). |
| **Perpetual worker** | Long-running across many Ralph recycles, state on disk; woken by async injection ([Mode 2](1-Architecture)). |
-| **Event-bus automation** | A webhook hits `bridged`'s REST ingress → injects an idle worker → result back to Opus or a chat bridge. No human in the loop. |
+| **Event-bus automation** | A webhook hits `fleetd`'s REST ingress → injects an idle worker → result back to Opus or a chat bridge. No human in the loop. |
## Related
- **[Architecture](1-Architecture)** — invariants, two modes, lifecycle state machine.
-- **[Message Server](2-Message-Server)** — the `bridged` design these mechanisms live in.
+- **[Message Server](2-Message-Server)** — the `fleetd` design these mechanisms live in.
- **[Roadmap](8-Roadmap)** — the staged plan + tickets that build this review scenario first.
- **[Team](6-Team)** — the fleet/orchestration layer above a single review.
diff --git a/8-Roadmap.md b/8-Roadmap.md
index 5f4c962..6a3fe9a 100644
--- a/8-Roadmap.md
+++ b/8-Roadmap.md
@@ -28,9 +28,9 @@ gantt
| Stage | Goal | Delivers (usable outcome) |
|---|---|---|
-| **1 — Walking skeleton** | One review, happy path | Opus mounts `bridged` (MCP), calls `bridge_send` with a diff, gets a review back from a real `ccs gx00-vllm claude` worker. Single hardcoded profile, same host, no guard/lifecycle. |
-| **2 — Contract + guard** | Trust the reply, trust the boundary | Structured [envelope](7-Use-Cases#mechanism-4--the-id-contract-envelope) + worker `bridge_reply`; reply rendezvous; subscription guard via `ccs env`; reviewer skill. |
-| **3 — Lifecycle + discovery** | Reuse, recycle, choose | Session manager (spawn/reuse/recycle Ralph loop, `idle_ttl`); `bridge_list` roster+live; multiple profiles. (`bridge_ask` landed early in Stage 2.) |
+| **1 — Walking skeleton** | One review, happy path | Opus mounts `fleetd` (MCP), calls `fleet_send` with a diff, gets a review back from a real `ccs gx00-vllm claude` worker. Single hardcoded profile, same host, no guard/lifecycle. |
+| **2 — Contract + guard** | Trust the reply, trust the boundary | Structured [envelope](7-Use-Cases#mechanism-4--the-id-contract-envelope) + worker `fleet_reply`; reply rendezvous; subscription guard via `ccs env`; reviewer skill. |
+| **3 — Lifecycle + discovery** | Reuse, recycle, choose | Session manager (spawn/reuse/recycle Ralph loop, `idle_ttl`); `fleet_list` roster+live; multiple profiles. (`fleet_ask` landed early in Stage 2.) |
| **4 — Pluggable peers** | PeerLauncher SPI + two adapters | ✅ Extract `PeerLauncher` SPI in-tree; `ClaudeCodeLauncher` (Stage A) **and** `OpenCodeLauncher` (Stage B, CB-402) routed by a `kind:` discriminator through `CompositePeerLauncher`. Stage C (dynamic external plugin loading) remains future work, gated by a trust/capability model. |
| **5 — Harden** | Production shape | ✅ Bearer auth + fail-fast exposure guard, `/metrics` + `/healthz`, mock-socket CI on the Gitea runner, launchd + systemd units, per-session authz + audit log. TLS deliberately terminates at a reverse proxy, not in the daemon. |
@@ -43,10 +43,10 @@ Consolidated from [Message Server](2-Message-Server#proposed-tech-stack); the cc
| **Language** | **Java 21+ (virtual threads)** | Loom fits the blocking socket + MCP + queue + SSE fan-in; **GraalVM `native-image`** recovers the `scp`+systemd single-binary deploy. Kotlin OK (same JVM). |
| **REST API core** | **Javalin** (or Spring MVC) + SSE | **The implementation & testability surface** — every feature is a REST endpoint (see [Testability](#testability--the-rest-api-is-the-contract-surface)). |
| **MCP server** | **Official MCP Java SDK**, streamable-HTTP (Jetty/Spring) | A **thin adapter over the REST core** — the Claude-facing face of the same features. |
-| **herdr client** | JDK `UnixDomainSocketAddress` + `SocketChannel`, NDJSON (Jackson) | Native UDS, no dep; event stream on a virtual thread. Pinned to herdr **0.7.0 / protocol 14** ([verified API](2-Message-Server#the-herdr-control-contract-what-bridged-drives)). |
+| **herdr client** | JDK `UnixDomainSocketAddress` + `SocketChannel`, NDJSON (Jackson) | Native UDS, no dep; event stream on a virtual thread. Pinned to herdr **0.7.0 / protocol 14** ([verified API](2-Message-Server#the-herdr-control-contract-what-fleetd-drives)). |
| **Worker spawn** | **`ccs
CB-594 · CB-600"]
dur["durability claim
CB-527 · CB-528"]
sec["credential scope
CB-593"]
bugs["nudge scheduling
CB-590 · CB-598"]
cfg["config accuracy
CB-597 · CB-599 · CB-602"]
flaky["green build
CB-601 · CB-603"]
- ask["bridge_ask reach
CB-582"]
+ ask["fleet_ask reach
CB-582"]
docs["catalogue debt
CB-595"]
val["config validation
CB-604 · CB-606"]
op["needs the operator
CB-596"]
@@ -284,9 +284,9 @@ two: CB-596's own gap detector found both of them on its first live spawn.*
| **Supervision** | CB-594 (#80) · CB-600 (#91) | The launchd unit shipped in `deploy/` and had never been installed, and launchd does not source a login shell — so a supervised daemon got no `WORKER_GITEA_TOKEN` and no `AI_GATEWAY_TOKEN`. The operator had to pick supervision *or* a working fleet. CB-600 then closed the gaps that only bite once the agent is loaded: a log path the script and the plist could silently disagree about, and a failed `load` leaving the agent stopped **and** persistently disabled. |
| **Durability** | CB-527 (#10) · CB-528 (#11) | No `basicQos`, so the backlog lived in JVM heap rather than on the broker; no publisher confirms, so a publish to a missing queue was a silent black hole. The wiki promised more durability than the code delivered. Both closed, with contract tests running against a real broker in CI. |
| **Credential scope** | CB-593 (#79) | The decision about which inherited credentials a member may keep is now recorded, so a deliberate choice no longer looks like an oversight. The canonical block also stopped claiming a member mounts only the bridge. |
-| **Nudge scheduling** | CB-590 (#75) · CB-598 (#87) · CB-582 (#61) | Two schedules could inject into one lead pane at once. Then the reminder budget turned out to be a counter carried forward with no memory of *which item* it counted, so work arriving during the backoff window inherited an already-capped count and was never nudged once. CB-582 added the third source: a worker paused in `bridge_ask` now nudges the lead itself, and `bridge_status` and REST both show the open question. **The ~55s window is closed, not removed** — still do not brief a worker to "ask me". |
+| **Nudge scheduling** | CB-590 (#75) · CB-598 (#87) · CB-582 (#61) | Two schedules could inject into one lead pane at once. Then the reminder budget turned out to be a counter carried forward with no memory of *which item* it counted, so work arriving during the backoff window inherited an already-capped count and was never nudged once. CB-582 added the third source: a worker paused in `fleet_ask` now nudges the lead itself, and `fleet_status` and REST both show the open question. **The ~55s window is closed, not removed** — still do not brief a worker to "ask me". |
| **Config accuracy** | CB-597 (#85) · CB-599 (#89) · CB-602 (#96) · CB-604 (#102) | Two documented knobs that are read by nothing; a capacity refusal that surfaced as a bare HTTP 500; no test at all in the code→example direction, so a brand-new key could ship undocumented; and an unknown `kind:` accepted silently and routed to the wrong adapter, which now refuses at config load. |
-| **Catalogue debt** | CB-595 (#81) | `wiki/11-Features.md` had fallen about fourteen entries behind, worst on the entries that changed what a config key *means*. `bridged.yaml` is gitignored, so that page is the only place an operator could learn them. Cleared — and it turned up two defects on the way (CB-604, CB-606). |
+| **Catalogue debt** | CB-595 (#81) | `wiki/11-Features.md` had fallen about fourteen entries behind, worst on the entries that changed what a config key *means*. `fleetd.yaml` is gitignored, so that page is the only place an operator could learn them. Cleared — and it turned up two defects on the way (CB-604, CB-606). |
| **Green build** | CB-601 (#95) · CB-603 (#100) | Two flaky tests, same root: a test that observes an asynchronous loop must be safe against that loop's thread, and neither the compiler nor a green build will say it is not. |
| **Config validation** | CB-604 (#102) · CB-606 (#106) | Four config fields shared one shape: lower-cased in a compact constructor, then compared against exactly **one** string, so a typo fell through to the other branch in silence. The worst was `auth.mode` — a typo of `token` behaved as `loopback-trust`, and `validateAuthExposure()` only fires on a **non-loopback** bind, so the common loopback bind hid it end to end and the daemon authenticated nobody while the config said otherwise. All four now refuse at load, naming the field, the value, the accepted set, and what would have happened. |
| **Snapshot pruning** | CB-586 (#67) | Nothing pruned `refs/wip/*`, so CB-578 stage C's snapshots pinned their whole trees forever. The rule that landed needs **both** conditions: the tree is already reachable from `main`, and the ref is older than 24h. Reachability is the floor — a snapshot exists because the work was committed nowhere else, so a plain TTL would delete the only copy. `/members` now reports `wipRefs{count,costBytes}`. The sweep shipped **dead**: a `Long.MIN_VALUE` "never yet" sentinel overflowed the interval gate, which returned before the assignment that would have fixed it, so it never ran once — and every unit test passed, because they all called the seam directly and walked around the gate. |
@@ -306,7 +306,7 @@ whole reason CB-592 was built the way it was; CB-596 inherits it.
**What the measurement actually said.** Two results, and they point in opposite directions.
The reassuring one: **CB-592 works.** The member holds the sentinel
-`blocked-by-bridged-cb592-see-gitea-issue-77`, not the admin token — confirmed by hashing the
+`blocked-by-fleetd-cb592-see-gitea-issue-77`, not the admin token — confirmed by hashing the
sentinel, which is a hardcoded non-secret string, and matching it against what the member reported.
The guarded-`export` mechanism does beat the login shell, which is the whole reason it was built that
way.
@@ -361,7 +361,7 @@ Made 2026-08-16, by reading the admission rule strictly.
| Ticket | Why it moved |
|---|---|
| CB-308 (#6) | Federation. It is what 2.0 *is*. |
-| CB-589 (#74) | Cost-first placement. `weighted` spreads by ratio with no idea which profile costs money, so paid spawns happen while the free box sits idle — but that is working behaviour that costs too much, with a decisive workaround (`local.weight: 100`), not something broken on one host. A missing **explanation** of the workaround *was* broken, because `bridged.yaml` is gitignored and a fresh host starts without it. **That half shipped in 1.1**; the policy did not. |
+| CB-589 (#74) | Cost-first placement. `weighted` spreads by ratio with no idea which profile costs money, so paid spawns happen while the free box sits idle — but that is working behaviour that costs too much, with a decisive workaround (`local.weight: 100`), not something broken on one host. A missing **explanation** of the workaround *was* broken, because `fleetd.yaml` is gitignored and a fresh host starts without it. **That half shipped in 1.1**; the policy did not. |
| CB-548 (#16) | Architect slots. A new capability, and genuinely blocked: `MemberRegistry.bind` is never called. |
| CB-605 (#103) | The systemd unit carries launchd's login-shell secret gap. Bites only when the first Linux gateway is stood up. |
| CB-607 (#110) | A member holds `SSH_AUTH_SOCK`, so it can sign with every key the operator's agent holds — broader than the repo-scoped token CB-302 built. Allowed **on purpose** today, because worktree remotes are `ssh://` and blocking it stops members pushing. A documented, deliberate scope reduction, not a break. |
@@ -375,7 +375,7 @@ Made 2026-08-16, by reading the admission rule strictly.
`e11160695fbe`, and the startup line reads
`memberCredentials: 34 known name(s), 5 allowed — blocking 29 on every spawn`. A merge is not a
deployment, and this change alters what a member's environment contains.
-4. ~~Confirm `bridge_whoami` still answers `primary`.~~ Done: `primary`. Note the trap — the redeploy
+4. ~~Confirm `fleet_whoami` still answers `primary`.~~ Done: `primary`. Note the trap — the redeploy
cuts the lead's own bridge MCP mount. It reconnected by itself the second time and did not the
first, so do not rely on either; if the tools are gone, ask the operator to run `/mcp`.
5. ~~Apply the secret-store half of CB-596.~~ Done by the operator on 2026-08-17: `secrets.sh` backed
@@ -406,7 +406,7 @@ installing it is the operator's decision. Read the Stage 5 table as *built*, not
A cross-cutting track that came out of a **"communication break" review** of the reverse
(worker → primary) path. The bridge *pushes* to workers (the status-gated injector) but only
-*pulls* to the primary — a worker reply resolves only an **already-open** blocking `bridge_send`;
+*pulls* to the primary — a worker reply resolves only an **already-open** blocking `fleet_send`;
MCP is client-initiated (bridge = server, primary = client), so the server can't call into the
primary. That asymmetry is the root of all three tickets.
@@ -435,9 +435,9 @@ spawn fail fast instead of stalling.*
- **CB-307 — reliable worker → primary delivery.** Behind one `ReplyInbox` port so the adapter is
swappable (see [Implementation → `msg`](9-Implementation#msg--the-service-core-4-classes)).
- **Stage 1 ✅ (main `ba6b4a5`).** `ReplyInbox` port + `InMemoryReplyInbox` (soft-state, dedup by
- `msgId`). A `bridge_reply` with no open send is now **held** instead of dropped; the primary
- drains it by target via `bridge_poll(target)` / `GET /sessions/{id}/replies`. No broker, no new
- dependency. **Only terminal replies are queued** — `bridge_ask` (interactive) and the
+ `msgId`). A `fleet_reply` with no open send is now **held** instead of dropped; the primary
+ drains it by target via `fleet_poll(target)` / `GET /sessions/{id}/replies`. No broker, no new
+ dependency. **Only terminal replies are queued** — `fleet_ask` (interactive) and the
completion/failure fallbacks are deliberately *not* (would risk double-delivery). Live on the
running daemon.
- **Stage 2 ✅ (main `2bc5f3a`).** `AmqpReplyInbox` behind the *same* port for cross-restart
@@ -448,7 +448,7 @@ spawn fail fast instead of stalling.*
the primary drains, so a `java -jar` bounce leaves them on the broker for redelivery. `broker:`
config absent → in-memory, present → AMQP. Contract test `@Tag("contract")` runs against a RabbitMQ
container. Dogfooded live (reply survived a daemon bounce, redelivered + acked exactly once).
- **`bridged` stays soft-state — the broker owns message durability, not the bus.**
+ **`fleetd` stays soft-state — the broker owns message durability, not the bus.**
- **Stage 3 — active push-to-primary + reminder ✅ (main `d4c9704`, gitea #5 closed).** The durable
inbox is a *landing zone* but delivery was still **pull** (the primary had to poll). Stage 3 makes
it **active**: a `ReplyPushLoop` nudges the primary the moment a reply lands with no open send, and
@@ -459,11 +459,11 @@ spawn fail fast instead of stalling.*
`push_backoff_ms`=15000). **Ack = drain**: stop when `inbox.peek(target).isEmpty()`. A single-slot
`PrimaryRegistry` learns the primary from orchestration-side tools; an off-host / non-herdr primary
leaves it empty → the loop is a no-op and delivery degrades to pull (reply never lost). Optional
- per-`msgId` `bridge_ack` tool for finer control than drain-all. Dogfooded live end-to-end.
+ per-`msgId` `fleet_ack` tool for finer control than drain-all. Dogfooded live end-to-end.
- **CB-308 — multi-host federation ⏳ (design note, gitea #6; depends on CB-307 Stage 2).** A primary
on host A delegating to workers on hosts B, C… with no host learning another's terminals. Built on
- CB-307's broker fabric: a **per-host gateway** (evolved `bridged` owning its local herdr +
+ CB-307's broker fabric: a **per-host gateway** (evolved `fleetd` owning its local herdr +
registry), **per-agent broker channels** `agent.
herdr client"] --> CB102["CB-102
ccs spawn"]
CB102 --> CB103["CB-103
injector"]
- CB103 --> CB104["CB-104
bridge_send"]
+ CB103 --> CB104["CB-104
fleet_send"]
CB104 --> CB105["CB-105
MCP server"]
CB106["CB-106
config"] --> CB107
CB105 --> CB107["CB-107
e2e demo (gate)"]
diff --git a/9-Implementation.md b/9-Implementation.md
index 9d74506..00b6711 100644
--- a/9-Implementation.md
+++ b/9-Implementation.md
@@ -1,13 +1,13 @@
# 9. Implementation Architecture (as-built)
-> **Scope.** This is the *as-built* code map of the `bridged` module — the actual packages,
+> **Scope.** This is the *as-built* code map of the `fleetd` module — the actual packages,
> classes, flows, and state machines in the source tree, as a companion to the design-level
> [1. Architecture](1-Architecture) and [2. Message Server](2-Message-Server). Every enum,
> constant, and route below was verified against source at main `3aa69a9`; the `msg`-layer
> **reply-inbox** (CB-307 Stage 1 `ba6b4a5`), its **AMQP durable adapter** (Stage 2 `2bc5f3a`), and
> the **active push-to-primary loop** (Stage 3 `d4c9704`) are folded in below.
-`bridged` is a single-host Java 25 / Maven daemon: the sole gateway between an on-subscription
+`fleetd` is a single-host Java 25 / Maven daemon: the sole gateway between an on-subscription
**primary** (Opus) and off-subscription **workers**, speaking to the **herdr** PTY manager
(protocol 19, herdr 0.8.0 — ported in CB-521) over a Unix-domain socket. It presents two equivalent faces — a REST
server (the testability seam) and an MCP server — over one shared service core.
@@ -16,17 +16,17 @@ server (the testability seam) and an MCP server — over one shared service core
The daemon is layered. The two north faces (REST + MCP) are thin adapters over one service core
(`msg`); the core drives delivery through `inject`, which speaks to the outside world only through
-`herdr`. `guard` sits on the spawn path; `config` and the `Bridged` entry point wire it all.
+`herdr`. `guard` sits on the spawn path; `config` and the `Fleetd` entry point wire it all.
```mermaid
flowchart TB
primary["Primary (Opus)
on-subscription"]
worker["Worker (Claude Code)
off-subscription"]
- subgraph bridged["bridged daemon"]
+ subgraph fleetd["fleetd daemon"]
direction TB
subgraph faces["North faces (thin adapters)"]
- rest["rest.BridgedApp
REST · Javalin"]
+ rest["rest.FleetdApp
REST · Javalin"]
mcp["mcp.BridgeMcp
MCP · /mcp servlet"]
end
msg["msg.MessageService + Rendezvous + ReplyInbox
service core · rendezvous + held-reply inbox"]
@@ -38,9 +38,9 @@ flowchart TB
herdrd["herdr daemon
(PTY manager)"]
- primary -->|"bridge_send / answer"| rest
+ primary -->|"fleet_send / answer"| rest
primary -->|"MCP tool calls"| mcp
- worker -->|"bridge_reply / bridge_ask"| mcp
+ worker -->|"fleet_reply / fleet_ask"| mcp
rest --> msg
mcp --> msg
mcp --> ccl
@@ -64,7 +64,7 @@ the adapter and gates spawn.*
## Package & class reference
-Fourteen packages under `dev.ltms.bridged`: `auth`, `config`, `guard`, `herdr`, `inject`, `lead`,
+Fourteen packages under `dev.ltms.fleetd`: `auth`, `config`, `guard`, `herdr`, `inject`, `lead`,
`mcp`, `member`, `metrics`, `msg`, `peer`, `placement`, `rest`, `session`. Below, per layer: the
classes, their kind, and their role. Method signatures are abbreviated; see source for full
contracts. The sections below do not yet cover every package — `auth`, `herdr`, `inject`, `msg`,
@@ -78,7 +78,7 @@ Java records.
| Class | Kind | Role |
|---|---|---|
-| `HerdrClient` | interface | The client face — the only herdr speaker in `bridged` (`call(method, params)`). |
+| `HerdrClient` | interface | The client face — the only herdr speaker in `fleetd` (`call(method, params)`). |
| `UnixSocketHerdrClient` | class | JDK Unix-socket transport; fresh connection per `call`, one frame, one line, close. |
| `HerdrCodec` | class | Encodes/decodes JSON-RPC frames (`encode`, `decodeResult`). |
| `HerdrException` | class | Runtime failure carrying an optional herdr protocol `code()`. |
@@ -97,8 +97,8 @@ text (`AgentControl.SUBMIT_KEY`) — never a shell prefix. `closeTab`/`PaneLocat
### `inject` — status-gated delivery + turn lifecycle (6 classes)
The single-writer delivery layer. Polls worker status, queues per target, injects only when safe,
-and **synthesizes turn boundaries** so a blocked `bridge_send` resolves even when a worker never
-calls `bridge_reply`. This is where the turn state machine lives.
+and **synthesizes turn boundaries** so a blocked `fleet_send` resolves even when a worker never
+calls `fleet_reply`. This is where the turn state machine lives.
| Class | Kind | Role |
|---|---|---|
@@ -111,17 +111,17 @@ calls `bridge_reply`. This is where the turn state machine lives.
### `msg` — the service core (6 classes)
-Owns the forward rendezvous (`bridge_send` → `bridge_reply`) and the reverse rendezvous
-(`bridge_ask` → answer), plus async fire-and-poll — and, since CB-307, the **reply inbox** that
+Owns the forward rendezvous (`fleet_send` → `fleet_reply`) and the reverse rendezvous
+(`fleet_ask` → answer), plus async fire-and-poll — and, since CB-307, the **reply inbox** that
holds a worker's terminal reply when *no* forward send is open (instead of dropping it) and the
**push loop** that actively nudges the primary to drain it.
| Class | Kind | Role |
|---|---|---|
-| `MessageService` | class | Orchestrates send/reply/ask/answer + async dispatch (`send`, `answer`, `ask`, `sendAsync`, `poll`), and routes `bridge_reply` through **`reply`** (resolve an open send, else publish to the inbox) + **`drainReplies`** (peek-then-ack a target's held replies). |
+| `MessageService` | class | Orchestrates send/reply/ask/answer + async dispatch (`send`, `answer`, `ask`, `sendAsync`, `poll`), and routes `fleet_reply` through **`reply`** (resolve an open send, else publish to the inbox) + **`drainReplies`** (peek-then-ack a target's held replies). |
| `Rendezvous` | class | Low-level registry of forward waiters + reverse-ask futures (`open`, `resolve`, `resolveQuestion`, `openAsk`, `answerAsk`, `closeAsk`, `resolveCompletion`, `resolveFailure`). **Untouched by CB-307** — it stays a pure synchronization primitive; a `false` from `resolve` (no live waiter) is what triggers the inbox publish, one layer up in `MessageService`. |
| `ReplyInbox` | interface | The port (CB-307): `publish(target, msgId, content)` (idempotent, dedup by `msgId`), `peek(target)` (non-destructive FIFO snapshot), `ack(target, msgId)`. Nested `InboxMessage(msgId, target, content)` record. Two adapters implement the *same* port — soft-state in-memory and durable AMQP. |
-| `InMemoryReplyInbox` | class | Stage-1 default adapter — per-target FIFO in a `ConcurrentHashMap
→ Injector + StatusPoller"]
rv --> ms["MessageService"]
ms --> mcp["BridgeMcp
(connection identity)"]
- ms --> app["BridgedApp
(start Javalin)"]
+ ms --> app["FleetdApp
(start Javalin)"]
mcp --> app
```
-*Figure 2 — startup wiring in `Bridged.main`. The guard asserts the primary env is clean before
+*Figure 2 — startup wiring in `Fleetd.main`. The guard asserts the primary env is clean before
anything else; `PeerLauncher` is wired as an interface, with `ClaudeCodeLauncher` as the first
adapter; `reapOrphanWorkers()` clears stale panes from a prior daemon restart.*
-## Flow: forward rendezvous (`bridge_send` → `bridge_reply`)
+## Flow: forward rendezvous (`fleet_send` → `fleet_reply`)
The main path. A primary's send blocks on a per-session waiter that resolves on the worker's
explicit reply — **or**, as a fallback, on a confirmed `working → idle` turn boundary the injector
-observes (so a worker that finishes without calling `bridge_reply` still unblocks the caller).
+observes (so a worker that finishes without calling `fleet_reply` still unblocks the caller).
```mermaid
sequenceDiagram
@@ -270,7 +270,7 @@ sequenceDiagram
IJ->>W: inject when injectable (one msg/turn)
IJ-->>MS: onDelivered (capture waiter + baseline)
alt worker replies explicitly
- W->>RV: bridge_reply → resolve(REPLY)
+ W->>RV: fleet_reply → resolve(REPLY)
else turn completes without reply
IJ-->>RV: onTurnComplete → resolveCompletion(COMPLETION)
else worker fails / wedges
@@ -284,11 +284,11 @@ sequenceDiagram
(delivered) or `TIMED_OUT_QUEUED` (never delivered).*
**The stranded-reply case (CB-307).** The figure is the happy path — a live waiter exists. The
-failure that motivated CB-307 is the *reverse*: a worker finishes and calls `bridge_reply` **after**
+failure that motivated CB-307 is the *reverse*: a worker finishes and calls `fleet_reply` **after**
its send has already timed out (the ~60s client window) or was never opened, so `Rendezvous.resolve`
finds no waiter. Before CB-307 that reply was silently discarded (the worker got an error / REST
`409`). Now `MessageService.reply` publishes it to the `ReplyInbox` under a fresh `msgId`, and the
-primary collects it later keyed by target — `bridge_poll(target)` or `GET /sessions/{id}/replies`
+primary collects it later keyed by target — `fleet_poll(target)` or `GET /sessions/{id}/replies`
(peek → deliver → ack, so an in-flight failure re-surfaces it). With `broker:` configured the
`AmqpReplyInbox` gives this **cross-restart durability** (Stage 2) — the held reply survives a
`java -jar` bounce and is redelivered; without it the in-memory adapter is soft-state (undrained
@@ -296,7 +296,7 @@ replies clear on restart).
Delivery is no longer purely pull. When the reply lands with no waiter, `reply` also calls
`ReplyPushLoop.onReplyQueued(target)`, which **actively nudges the primary to drain** (Stage 3): it
-injects a *"run `bridge_poll(target=…)`"* turn into the primary's own herdr pane — the same
+injects a *"run `fleet_poll(target=…)`"* turn into the primary's own herdr pane — the same
`AgentControl.send` primitive that delivers to workers, pointed at the primary — but only when the
primary is `injectable()` (never mid-turn), bounded to `primary.push_reminders` nudges on a
`push_backoff_ms` schedule, and stopping the instant a drain empties the inbox. The nudge carries
@@ -304,7 +304,7 @@ primary is `injectable()` (never mid-turn), bounded to `primary.push_reminders`
push fails the durable inbox is still the backstop. An off-host / non-herdr primary leaves
`PrimaryRegistry` empty, so the loop is a no-op and delivery cleanly degrades to pull.
-## Flow: reverse rendezvous (`bridge_ask` → answer)
+## Flow: reverse rendezvous (`fleet_ask` → answer)
A worker pauses its own turn to ask the primary a question; the question surfaces on the primary's
open forward send, and the answer resumes the *same* worker turn. Duplicate asks from one session
@@ -328,7 +328,7 @@ sequenceDiagram
P->>MS: answer(turnId, content)
MS->>RV: answerAsk(turnId, content) → unblock worker
MS->>RV: open(workerSession) — new forward waiter
- W-->>RV: resumes turn → bridge_reply
+ W-->>RV: resumes turn → fleet_reply
RV-->>P: Reply
else no forward send open
RV-->>W: AskOutcome.NO_WAITER (closeAsk)
@@ -413,7 +413,7 @@ in `ClaudeCodeLauncher.spawn()` **before any herdr call**: the worker's `ANTHROP
be on the allowlist (Stage-1: `gx00.gw`, `ollama.ltms.dev`). `assertPrimaryClean` (called at
startup) hard-stops if the primary env carries any `ANTHROPIC_BASE_URL`. Spawn injects
`ANTHROPIC_BASE_URL`/`ANTHROPIC_MODEL`/`CLAUDE_CONFIG_DIR`/`ANTHROPIC_AUTH_TOKEN` into the
-**worker's** env only — `bridged`'s own env is never mutated. CWD resolves
+**worker's** env only — `fleetd`'s own env is never mutated. CWD resolves
`requestedCwd → profile cwd → caller cwd → daemon user.dir → "."`.
## Concurrency model (at a glance)
@@ -435,7 +435,7 @@ startup) hard-stops if the primary env carries any `ANTHROPIC_BASE_URL`. Spawn i
## Related pages
- [1. Architecture](1-Architecture) — system, invariants, the two modes (design level).
-- [2. Message Server](2-Message-Server) — the `bridged` design rationale, herdr control contract.
+- [2. Message Server](2-Message-Server) — the `fleetd` design rationale, herdr control contract.
- [8. Roadmap](8-Roadmap) — stages, tickets, and the feature ⇄ endpoint ⇄ test map.
---
diff --git a/Home.md b/Home.md
index 3e6c0e2..8a7ce0e 100644
--- a/Home.md
+++ b/Home.md
@@ -12,23 +12,23 @@ A **subscription-safe bridge** that lets a lead **Claude Code** session on Pro/M
> inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper model. It also runs
> non-Claude members (`kind: opencode`) as first-class peers.
-## Leading approach — herdr-centric message server (`bridged`)
+## Leading approach — herdr-centric message server (`fleetd`)
-A small always-on message server, **`bridged`**, controls
+A small always-on message server, **`fleetd`**, controls
[herdr](https://herdr.dev) (an agent multiplexer, "tmux for agents") over its Unix-socket
API and exposes a clean 2-way messaging API as an **MCP server that both the primary and the
workers mount** — one unified Claude setup and the **sole communication gateway** (REST/SSE
-stays for non-Claude clients; any broker is `bridged`-internal, below the gateway). herdr owns
+stays for non-Claude clients; any broker is `fleetd`-internal, below the gateway). herdr owns
the PTYs, multiplexing, persistence, and **agent-status
-events**; `bridged` owns policy (subscription boundary, session lifecycle, status-gated
+events**; `fleetd` owns policy (subscription boundary, session lifecycle, status-gated
delivery) and the client contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at
the gateway, `https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and
-calls `bridged`'s MCP tools.
+calls `fleetd`'s MCP tools.
```mermaid
flowchart LR
OPUS["Opus — primary
(Claude Code, env CLEAN)
MCP client"]
- subgraph BD["bridged — standalone daemon (not a claude process)"]
+ subgraph BD["fleetd — standalone daemon (not a claude process)"]
SRV["SERVER face
MCP · REST/SSE · policy"]
CLI["CLIENT face
status-gated injector · herdr socket"]
SRV --> CLI
@@ -37,8 +37,8 @@ flowchart LR
W["member pane
ANTHROPIC_BASE_URL set
MCP client"]
M["llm.ltms.dev
(the one gateway)"]
- OPUS -->|"MCP bridge_send (blocks)"| SRV
- W -.->|"MCP bridge_reply"| SRV
+ OPUS -->|"MCP fleet_send (blocks)"| SRV
+ W -.->|"MCP fleet_reply"| SRV
CLI -->|"Unix socket
send_text · events.subscribe"| HERDR
HERDR -->|"drives PTY"| W
W -->|"inference"| M
@@ -50,28 +50,28 @@ flowchart LR
```
- **Subscription boundary:** the *lead* never sets `ANTHROPIC_BASE_URL` (stays on Pro/Max). Only a
- member process is moved off-subscription, and only by the daemon, at spawn; `bridged` is a plain
+ member process is moved off-subscription, and only by the daemon, at spawn; `fleetd` is a plain
daemon with no Anthropic quota, so it may poll and subscribe freely. **One exception, and it
costs money:** a profile marked `subscription: true` runs its members on *your* plan on purpose —
see [13 User Guide](13-User-Guide) → *The knobs that cost money*. An `opencode` member sits
outside this boundary entirely and uses its own provider credential.
-- **One gateway:** `bridged` is the **sole communication path** for every session. The lead and
+- **One gateway:** `fleetd` is the **sole communication path** for every session. The lead and
every member mount it as an MCP server and talk *only* to it — **no session ever addresses a
broker, a peer, or the network directly.** How the mount happens differs by backend: a
`claude-code` member gets one `claude mcp add` line or a shared `.mcp.json`, while an `opencode`
member gets a generated `opencode.json` and no `ANTHROPIC_*` variables at all. Any broker or
- queue is `bridged`-internal, below the gateway — and it is optional. With `broker:` commented
+ queue is `fleetd`-internal, below the gateway — and it is optional. With `broker:` commented
out the daemon uses an in-memory inbox; when it is on, it is **LavinMQ** over AMQP.
-- **How the lead gets a reply:** in the simple case, **one blocking MCP tool call** (`bridge_send`)
- that `bridged` holds open until the member replies (`bridge_reply`) or its turn completes, then
+- **How the lead gets a reply:** in the simple case, **one blocking MCP tool call** (`fleet_send`)
+ that `fleetd` holds open until the member replies (`fleet_reply`) or its turn completes, then
returns the reply as the tool result. In practice that call is capped by the lead's own MCP client
timeout (about 60 seconds), so real work uses `wait:false` and a ticket instead. Either way the
- lead never busy-polls and never touches a broker: when a detached ticket goes terminal, `bridged`
+ lead never busy-polls and never touches a broker: when a detached ticket goes terminal, `fleetd`
**injects the lead's own idle pane** to wake it.
- **Different model** per member process sidesteps Claude Code's lack of per-subagent
provider routing — the worker isn't a subagent, it's its own configured process.
- **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) was kept on paper as a
- swappable *fallback injector*. It was **never built** — `grep -ri agentapi bridged/src/main`
+ swappable *fallback injector*. It was **never built** — `grep -ri agentapi fleetd/src/main`
returns nothing, and the only injection path in the shipped code is the herdr one. Treat it as a
discarded option, not a fallback you can switch to. See [Approaches](3-Approaches) for why herdr
won and [Message Server](2-Message-Server) for the full design.
@@ -81,7 +81,7 @@ flowchart LR
Read in order (the sidebar mirrors this):
1. **[Architecture](1-Architecture)** — process model, the two invariants, two traffic modes
-2. **[Message Server](2-Message-Server)** — 🟢 **`bridged`**, the herdr-centric message server (primary approach)
+2. **[Message Server](2-Message-Server)** — 🟢 **`fleetd`**, the herdr-centric message server (primary approach)
3. **[Approaches](3-Approaches)** — herdr-centric vs AgentAPI vs Agent SDK vs bus/tmux (research matrix)
4. **[Setup](4-Setup)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §2** instead.
5. **[Operations](5-Operations)** — 🟠 stub, never written. Use **[13 User Guide](13-User-Guide) §4 and §6** instead.
@@ -100,7 +100,7 @@ Read in order (the sidebar mirrors this):
the daemon live on this host. Chapters 4 and 5 were never written past their scope note; chapter 13
replaced them.
-The **herdr-centric `bridged` message server** was selected on 2026-07-11, superseding the AgentAPI
+The **herdr-centric `fleetd` message server** was selected on 2026-07-11, superseding the AgentAPI
plan of 2026-07-08. AgentAPI is retained as a fallback injector and has not been needed. See
**[Message Server](2-Message-Server)** for the design and **[13 User Guide](13-User-Guide)** for how
to run it.
diff --git a/_Sidebar.md b/_Sidebar.md
index 860c35a..6f955f6 100644
--- a/_Sidebar.md
+++ b/_Sidebar.md
@@ -5,7 +5,7 @@
**Chapters**
1. [Architecture](1-Architecture) — system · 2 invariants · 2 modes
-2. [Message Server](2-Message-Server) — the `bridged` design
+2. [Message Server](2-Message-Server) — the `fleetd` design
3. [Approaches](3-Approaches) — transports compared, why herdr
4. [Setup](4-Setup) — ⚫ superseded by 13
5. [Operations](5-Operations) — ⚫ superseded by 13
@@ -19,4 +19,4 @@
13. **[User Guide](13-User-Guide)** — 🟢 install · configure · run · delegate · the traps
---
-🟢 herdr-centric `bridged` · AgentAPI = fallback
+🟢 herdr-centric `fleetd` · AgentAPI = fallback