diff --git a/README.md b/README.md index e7baf11..4154fee 100644 --- a/README.md +++ b/README.md @@ -9,23 +9,23 @@ Sibling of [`crush-bridge`](https://git.ltms.dev/systems/vms) (which drives a he process*, so it inherits `CLAUDE.md`, hooks, skills, and MCP — just pointed at a cheaper/local model. -## Leading approach — herdr-centric message server (`bridged`) +## Leading approach — herdr-centric message server (`fleetd`) -A small always-on message server, **`bridged`**, controls +A small always-on message server, **`fleetd`**, controls [herdr](https://herdr.dev) (an agent multiplexer) over its Unix-socket API and exposes a clean 2-way messaging API as an **MCP server that both the primary and the workers mount** — one unified Claude setup and the **sole communication gateway** (REST/SSE stays for non-Claude -clients; any broker is `bridged`-internal, below the gateway). -herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `bridged` owns +clients; any broker is `fleetd`-internal, below the gateway). +herdr owns the PTYs, multiplexing, persistence, and **agent-status events**; `fleetd` owns policy (subscription boundary, session lifecycle, status-gated delivery) and the client contract. A Claude member launches with `ANTHROPIC_BASE_URL` pointed at the gateway, `https://llm.ltms.dev/anthropic`, plus a bearer token; the lead stays env-clean and calls -`bridged`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it. +`fleetd`'s MCP tools. See the wiki's **[13 User Guide](wiki/13-User-Guide.md)** to run it. ```mermaid flowchart LR OPUS["Opus — primary
(Claude Code, env CLEAN)
MCP client"] - subgraph BD["bridged — standalone daemon (not a claude process)"] + subgraph BD["fleetd — standalone daemon (not a claude process)"] SRV["SERVER face
MCP · REST/SSE · policy"] CLI["CLIENT face
status-gated injector · herdr socket"] SRV --> CLI @@ -47,27 +47,27 @@ flowchart LR ``` - **Subscription boundary:** the *primary* never sets `ANTHROPIC_BASE_URL` (stays on - Pro/Max). Only the *secondary* process is off-subscription — and `bridged` itself is a + Pro/Max). Only the *secondary* process is off-subscription — and `fleetd` itself is a plain daemon (no Anthropic quota), so it may poll/subscribe freely. -- **One gateway (unified MCP setup):** `bridged` is the **sole communication path** for every +- **One gateway (unified MCP setup):** `fleetd` is the **sole communication path** for every Claude session. Primary and workers each mount it as an MCP server (one `claude mcp add` line, same on both) and talk over MCP tools — `fleet_send` / `fleet_reply` / `fleet_status` (with `fleet_ask` planned for the blocked-worker path). **No Claude session ever addresses a broker, a peer, or the network - directly**; any queue is `bridged`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so + directly**; any queue is `fleetd`-internal. MCP tool I/O never sets `ANTHROPIC_BASE_URL`, so mounting the bridge is subscription-safe by construction. **Tool naming:** the tools were renamed from `bridge_*` to `fleet_*` (CB-622). The daemon still answers the old `bridge_*` names for one release, but they are deprecated — use the `fleet_*` names. - **How the primary consumes a reply:** a single **blocking MCP call** (`fleet_send`); - `bridged` holds it open until the worker calls `fleet_reply` or its turn hits + `fleetd` holds it open until the worker calls `fleet_reply` or its turn hits `agent_status=done`, then returns the reply as the tool result. No cross-turn busy-poll, so no quota burn. SSE is an optional side-channel for humans/dashboards watching status. -- **Worker → primary** rides `bridged`'s **MCP rendezvous** — the reply resolves the primary's - blocking call (or, for detached work, `bridged` **injects the primary's idle pane** when it's +- **Worker → primary** rides `fleetd`'s **MCP rendezvous** — the reply resolves the primary's + blocking call (or, for detached work, `fleetd` **injects the primary's idle pane** when it's ready), so *no keystroke-into-primary and no broker are involved, even single-host*. The one exception: a split-host primary that isn't a herdr pane wakes via its own `Stop`-hook, which - polls **`bridged`** (never a broker). See the wiki for the two topologies. + polls **`fleetd`** (never a broker). See the wiki for the two topologies. - **Different model per process** sidesteps Claude Code's lack of per-subagent provider routing — the worker isn't a subagent, it's its own configured process. - **AgentAPI** ([`coder/agentapi`](https://github.com/coder/agentapi)) is retained only as a @@ -90,7 +90,7 @@ Gitea wiki. ## Status -🟢 **Implemented & dogfooded** — the herdr-centric **`bridged`** message server is built and in +🟢 **Implemented & dogfooded** — the herdr-centric **`fleetd`** message server is built and in real use: an Opus primary delegates tasks to off-subscription workers that reply through the bridge (code reviews delegated this way have produced committed bug fixes). Selected as the primary approach 2026-07-11, superseding the AgentAPI plan (2026-07-08); AgentAPI retained as a diff --git a/bridged/src/main/java/dev/ltms/fleet/member/OpenCodeLauncher.java b/bridged/src/main/java/dev/ltms/fleet/member/OpenCodeLauncher.java index 9b760c7..f018f9b 100644 --- a/bridged/src/main/java/dev/ltms/fleet/member/OpenCodeLauncher.java +++ b/bridged/src/main/java/dev/ltms/fleet/member/OpenCodeLauncher.java @@ -362,7 +362,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher { options.put("baseURL", openAiBaseUrl(cfg.baseUrl())); // vLLM and friends usually ignore the key, but the AI SDK still requires a non-empty one. String token = resolveEnv(cfg.tokenEnv()); - options.put("apiKey", (token == null || token.isBlank()) ? "bridged-local-noauth" : token); + options.put("apiKey", (token == null || token.isBlank()) ? "fleetd-local-noauth" : token); provider.putObject("models").putObject(modelId).put("name", modelId); } diff --git a/docs/CB-301-Session-Manager.md b/docs/CB-301-Session-Manager.md index 57ad7fd..a4305f4 100644 --- a/docs/CB-301-Session-Manager.md +++ b/docs/CB-301-Session-Manager.md @@ -2,7 +2,7 @@ **Status:** design spec for review → delegate implementation. **Grounded in:** `WorkerService`, `Injector`/`StatusPoller`/`TurnListener`, `MessageService`, -`BridgeMcp`, `BridgedApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)). +`FleetMcp`, `FleetApp` (see [wiki 9. Implementation](../wiki/9-Implementation.md)). ## Problem @@ -43,7 +43,7 @@ Under no-reuse, a released session is terminal. A new `acquire` always creates a subscription-guarded spawn/teardown mechanics; `SessionManager` adds the registry, lifecycle, and ownership on top. -**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from +**Package:** new `dev.ltms.fleet.session` — keeps the registry/lifecycle concern separate from the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`. ### `WorkerSession` (record or small mutable holder) @@ -96,10 +96,10 @@ final class SessionManager { ### Integration points -- **`Bridged.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the +- **`Fleetd.main`** — construct `SessionManager(workerService, ...)`; wire it as/decorating the `TurnListener` alongside `CompletionResolver` so it sees turn boundaries, and give it the `WorkerPresence` signal for `READY`. -- **`BridgeMcp.spawn` / `BridgedApp.spawnWorker`** — route spawn through `SessionManager.acquire` +- **`FleetMcp.spawn` / `FleetApp.spawnWorker`** — route spawn through `SessionManager.acquire` (carry `callerTerminal` as `ownerTerminal`). **`fleet_stop` / `DELETE /workers/{paneId}`** → `SessionManager.release`. - **`fleet_list` / `GET /sessions` (CB-304 later)** — read `SessionManager.roster()`. diff --git a/docs/CB-301-ext-Worktree-Provisioning.md b/docs/CB-301-ext-Worktree-Provisioning.md index 6f75678..eac46c6 100644 --- a/docs/CB-301-ext-Worktree-Provisioning.md +++ b/docs/CB-301-ext-Worktree-Provisioning.md @@ -2,12 +2,12 @@ **Status:** ✅ shipped — implemented at commit `97ecc71` (per-worker git worktree + config-parity overlay). As-built: `session/GitWorktrees.java` behind the `Worktrees` port, wired in -`Bridged.main` and configurable via `worktreeRoot` / per-profile `parityOverlay` +`Fleetd.main` and configurable via `worktreeRoot` / per-profile `parityOverlay` (see `bridged.example.yaml`). Branch/worktree surface in `fleet_list` landed with CB-304 (`9fe04bf`); the worker-opened-PR checkpoint landed as CB-302 (`64e70ef`). **Extends:** [CB-301 Session Manager](CB-301-Session-Manager.md) (shipped, commit `54d907c`). **Realizes:** the config-parity requirement in [Worker Git Workflow](Worker-Git-Workflow.md). -**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `BridgedConfig.Worker`, +**Grounded in:** `SessionManager`, `WorkerService.spawn/effectiveCwd`, `FleetConfig.Worker`, `inject/…LsofPeerPidLookup` (the `ProcessBuilder` exec pattern). ## Problem @@ -43,7 +43,7 @@ provider. ### `WorktreeRequest` (new, nullable = "no worktree") ```java -package dev.ltms.bridged.session; +package dev.ltms.fleet.session; /** Ask acquire() to provision an isolated worktree. null ⇒ run in the shared primary tree. */ public record WorktreeRequest(String ticketSlug, String baseRef) { // ticketSlug seeds the branch name; baseRef null/blank ⇒ current HEAD of the repo. @@ -63,7 +63,7 @@ untouched. ### `Worktrees` seam (new) ```java -package dev.ltms.bridged.session; +package dev.ltms.fleet.session; public interface Worktrees { /** git -C worktree add -b . Returns the worktree path. */ String add(String repoRoot, String branch, String baseRef); @@ -112,7 +112,7 @@ public void release(String paneId) { } ``` -### Config — `BridgedConfig.Worker.parityOverlay` + a `worktreeRoot` +### Config — `FleetConfig.Worker.parityOverlay` + a `worktreeRoot` - Add `List parityOverlay` to the `Worker` record (12th field). Compact-constructor default when null/empty: `[".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]` (missing paths are diff --git a/docs/CB-306-Spawn-Readiness-Gate.md b/docs/CB-306-Spawn-Readiness-Gate.md index 9a97bf0..e104d26 100644 --- a/docs/CB-306-Spawn-Readiness-Gate.md +++ b/docs/CB-306-Spawn-Readiness-Gate.md @@ -55,7 +55,7 @@ a crash, and a slow start all present as "never becomes injectable" and all corr or `spawnReadyTimeoutMs` elapses. 3. **Ready** → return the `WorkerHandle(paneId, terminalId)` as today. 4. **Timeout** → the launcher **closes the pane it started** (and its tab, via the same path - `release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.bridged.peer`). + `release`/`stop` uses) and throws **`PeerUnreachableException`** (new, in `dev.ltms.fleet.peer`). No orphan pane is left behind — the launcher cleans up its own failed birth. `spawnReadyTimeoutMs == 0` (or unset) **disables** the gate = legacy non-blocking behaviour, so the @@ -80,7 +80,7 @@ Unit tests (add to the existing `ClaudeCodeLauncher` test): ## 5. Config Add to the launcher-level config (a bridged-level knob, not per-profile) in `bridged.yaml` + -`BridgedConfig`: +`FleetConfig`: ```yaml spawn_ready_timeout_ms: 20000 # 0 disables the gate (legacy non-blocking spawn) @@ -88,7 +88,7 @@ spawn_ready_poll_ms: 300 ``` Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sane defaults in code -(`20000` / `300`). Keep the names consistent with existing config field style in `BridgedConfig`. +(`20000` / `300`). Keep the names consistent with existing config field style in `FleetConfig`. ## 6. Core / MCP propagation @@ -100,7 +100,7 @@ Jackson ignores unknown keys, so omitting them in existing YAML is safe; pick sa - The **non-worktree** path registers the session only *after* `spawn` returns, so a throw means no half-live `SPAWNING` session is ever registered — confirm this and add a test. - `fleet_spawn` (MCP verb) must return an **error result** carrying the exception message, not a - success with a dead session. Trace `BridgeMcp`/`BridgedApp` spawn handlers and make sure the + success with a dead session. Trace `FleetMcp`/`FleetApp` spawn handlers and make sure the exception becomes a clean tool error, not an uncaught 500 with a stack trace. **Out of scope (do NOT do here):** gating `fleet_send` on session `READY` (existing status-gate + @@ -111,7 +111,7 @@ any config `kind:` discriminator, any second adapter. - `ClaudeCodeLauncher.spawn` blocks until injectable or throws `PeerUnreachableException` + self-reaps the pane; gate disabled when timeout is 0. -- New `PeerUnreachableException` in `dev.ltms.bridged.peer`. +- New `PeerUnreachableException` in `dev.ltms.fleet.peer`. - Config knobs wired (`spawn_ready_timeout_ms`, `spawn_ready_poll_ms`) with safe defaults. - Existing `SPAWNING→READY` MCP-contact transition untouched. - New unit tests (ready / timeout+reap / disabled) green; **all existing tests still pass unchanged**. diff --git a/docs/CB-307-Reliable-Delivery.md b/docs/CB-307-Reliable-Delivery.md index 5de1d1e..e827e09 100644 --- a/docs/CB-307-Reliable-Delivery.md +++ b/docs/CB-307-Reliable-Delivery.md @@ -19,8 +19,8 @@ is currently open** for that worker: - `Rendezvous.resolve(session, content)` → `complete(...)` → `waiters.get(session) == null` → returns `false` (`msg/Rendezvous.java:212-215`). - The `content` string is **never retained** — it is dropped. The worker is told it failed: - `BridgeMcp.reply` returns `error("no send is awaiting a reply for this worker")` (`mcp/BridgeMcp.java:270-272`); - REST returns `409 no_pending_send` (`rest/BridgedApp.java:339-345`). + `FleetMcp.reply` returns `error("no send is awaiting a reply for this worker")` (`mcp/FleetMcp.java:270-272`); + REST returns `409 no_pending_send` (`rest/FleetApp.java:339-345`). This is the observed "communication break": a worker that finishes just after its `fleet_send` timed out (the ~60s sync window) replies into the void. There is **no message-id, dedup, or ack** anywhere in @@ -28,13 +28,13 @@ the message path today. ## 2. What to build -### 2.1 The port — `dev.ltms.bridged.msg.ReplyInbox` +### 2.1 The port — `dev.ltms.fleet.msg.ReplyInbox` A thin interface owned by the `msg` layer. The in-memory adapter is Stage 1; the AMQP adapter (Stage 2) implements the **same** interface, so keep it broker-agnostic. ```java -package dev.ltms.bridged.msg; +package dev.ltms.fleet.msg; import java.util.List; @@ -70,7 +70,7 @@ public interface ReplyInbox { seen-set — your call; preserve insertion order). - `peek` returns an immutable copy; `ack` removes by `msgId`. Thread-safe (concurrent publish vs. drain). - **This is soft-state, NOT persistence.** Lost on a `java -jar` bounce — that is correct and consistent - with "bridged stays soft-state." Do **not** add any file/DB backing. + with "fleetd stays soft-state." Do **not** add any file/DB backing. ### 2.3 Publish seam — route reply through the service layer @@ -89,9 +89,9 @@ in `MessageService`, which already owns the `Rendezvous` and will own the `Reply } ``` - Repoint the two callers off the bare `rendezvous.resolve(...)` onto `messages.reply(...)`: - - `BridgeMcp.reply` (`mcp/BridgeMcp.java:262-273`) — on success return a normal ack; **remove** the + - `FleetMcp.reply` (`mcp/FleetMcp.java:262-273`) — on success return a normal ack; **remove** the `error("no send is awaiting a reply…")` branch (that case is now a successful queue). - - `BridgedApp.replyMessage` (`rest/BridgedApp.java:330-346`) — return `200` (queued) instead of + - `FleetApp.replyMessage` (`rest/FleetApp.java:330-346`) — return `200` (queued) instead of `409 no_pending_send`. **DO NOT touch the QUESTION path.** `fleet_ask` / `rendezvous.resolveQuestion` must keep today's @@ -120,12 +120,12 @@ The primary re-checks a worker it delegated to. Expose a drain keyed by **worker ## 3. Config **None for Stage 1.** The in-memory adapter is the unconditional default — wire `new InMemoryReplyInbox()` -into `MessageService` in `Bridged.main`. Do **not** add a `broker:` config block (that arrives with the +into `MessageService` in `Fleetd.main`. Do **not** add a `broker:` config block (that arrives with the Stage-2 AMQP adapter: absent → in-memory, present → AMQP). ## 4. Acceptance criteria (what the primary will verify) -1. New `ReplyInbox` + `InboxMessage` + `InMemoryReplyInbox` in `dev.ltms.bridged.msg`. +1. New `ReplyInbox` + `InboxMessage` + `InMemoryReplyInbox` in `dev.ltms.fleet.msg`. 2. `fleet_reply` with **no open send** now **succeeds and queues** (no more `error` / `409`); the reply is later retrievable and identical. 3. The queued reply is drainable by the primary keyed by target; draining **acks** it (a second drain @@ -152,7 +152,7 @@ Stage-2 AMQP adapter: absent → in-memory, present → AMQP). - **`.mcp.json` is `--skip-worktree` in your worktree — never edit, `git add`, or commit it.** - **`wiki/` is a submodule — never run git in it; never touch it.** - Commit only your feature changes (the new port/adapter, the `msg`/`mcp`/`rest` wiring, tests, and if - you add config wiring in `Bridged.java`). Nothing else. + you add config wiring in `Fleetd.java`). Nothing else. - Work only inside your assigned worktree on your feature branch. The primary fast-forwards `main` after re-gating — do not touch `main`. - Java 25 idioms are welcome (unnamed `_` params, records). Keep the diff minimal and match surrounding style. @@ -171,7 +171,7 @@ removal, msgId dedup, cross-restart redelivery). It is `@Tag("contract")`, so th untouched. Run it explicitly when Docker (or a broker) is available: ```bash -cd bridged +cd fleetd mvn -Pcontract test -Dtest=AmqpReplyInboxContractTest # local: spins a RabbitMQ Testcontainers fixture ``` diff --git a/docs/CB-308-Multi-Host-Federation.md b/docs/CB-308-Multi-Host-Federation.md index 6e0730f..67c5b48 100644 --- a/docs/CB-308-Multi-Host-Federation.md +++ b/docs/CB-308-Multi-Host-Federation.md @@ -18,7 +18,7 @@ The design rests on three pieces (the shape this ticket proposes): 1. **Dedicated per-agent channels** — every agent has its own addressable inbox on the broker. 2. **A federated agent directory** — a global "who/where/status" lookup, assembled from per-host presence, not a central database. -3. **A per-host gateway** — each host runs a `bridged` that owns its local herdr, registers/manages +3. **A per-host gateway** — each host runs a `fleetd` that owns its local herdr, registers/manages its own sessions, and proxies messages to/from other hosts over the broker. ## 2. What is single-host today (the assumptions to break) @@ -27,7 +27,7 @@ The design rests on three pieces (the shape this ticket proposes): flowchart TB subgraph host["Single host (today)"] primary["primary
(MCP client)"] - daemon["bridged daemon
127.0.0.1:8765"] + daemon["fleetd daemon
127.0.0.1:8765"] reg["in-process registry
keyed by PeerHandle.id() == paneId"] herdr["herdr
(local unix-socket PTY mux)"] w1["worker pane wQ:p1"] @@ -48,21 +48,21 @@ Three concrete bake-ins assume one host: |---|---|---| | **herdr is local** | `herdr/` unix socket `~/.config/herdr/herdr.sock` | You cannot drive another host's PTYs → each host **must** own its herdr. This is why a per-host gateway is mandatory. | | **registry is in-process, keyed by `paneId`** | `session/SessionManager` | `paneId` (e.g. `wQ:p2B`) is a herdr-local coordinate — meaningless off-host. Routing needs a host-unique id. | -| **loopback, no authn** | `rest/BridgedApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. | +| **loopback, no authn** | `rest/FleetApp` binds `127.0.0.1:8765` | Fine on one host; the moment a second host can talk to a gateway, that link is a trust boundary. | ## 3. Target architecture ```mermaid flowchart TB subgraph hostA["HOST A"] - gA["gateway = bridged A"] + gA["gateway = fleetd A"] regA["local registry + herdr"] primary["primary (MCP client)"] gA --- regA primary --- gA end subgraph hostB["HOST B"] - gB["gateway = bridged B"] + gB["gateway = fleetd B"] regB["local registry + herdr"] wb["worker panes"] gB --- regB @@ -96,13 +96,13 @@ host's terminals.* delayed-message exchange (the remind/backoff loop for free) — the same reasons CB-307 picked it. - **Federated agent directory** = a **soft-state, bridge-owned** roster, *not* a broker-stored - database. Per the persistence-boundary decision (bridged is soft-state; the broker owns *message* + database. Per the persistence-boundary decision (fleetd is soft-state; the broker owns *message* durability, not *who/where/status*), each gateway announces its local agents `(globalId, host, status, capabilities)` on a `roster.*` presence topic with periodic heartbeats. Every gateway builds an eventually-consistent **union view** — literally CB-304's `rosterView`, federated. A stale entry expires by missed heartbeat (reuses CB-303's idle/TTL thinking). -- **Per-host gateway** = today's `bridged` daemon, evolved. It already registers/manages sessions +- **Per-host gateway** = today's `fleetd` daemon, evolved. It already registers/manages sessions and controls its local herdr; multi-host adds exactly two responsibilities: (a) a broker client that consumes its agents' inboxes and injects into local herdr, and (b) presence announce + union-roster assembly. Evolution, not rewrite. @@ -123,7 +123,7 @@ broker. A sender is oblivious to which branch it took.* ## 4. What CB-307 already provides vs. what is net-new **CB-307 delivers the transport half** and is independently valuable on a single host: the AMQP -broker fabric, the `bridged → broker` client/adapter, at-least-once + idempotent (dedup-by-id) +broker fabric, the `fleetd → broker` client/adapter, at-least-once + idempotent (dedup-by-id) delivery, DLQ, and delayed-retry (remind). That *is* the "proxy cross-host message" backbone; extending the same broker from "worker→primary reliability" to "gateway↔gateway" is incremental. diff --git a/docs/CB-401-Peer-Launcher-SPI.md b/docs/CB-401-Peer-Launcher-SPI.md index 2e334bd..b72aa22 100644 --- a/docs/CB-401-Peer-Launcher-SPI.md +++ b/docs/CB-401-Peer-Launcher-SPI.md @@ -60,7 +60,7 @@ Where Claude/herdr specifics actually live today: | `claude---` naming, `WORKER_NAME` regex, orphan reap (CB-117) | `WorkerService` | **→ adapter** (naming is a herdr-label detail) | | `GITEA_TOKEN` / `GITEA_HOST` injection (CB-302 checkpoint) | `WorkerService.spawn` | **→ adapter** + a **capability** (§6) | | tab/pane placement, worker space, tab labels | `WorkerService.spawnInTab/spawnAsPane` via herdr `WorkspaceControl` | **→ adapter** (herdr transport detail) | -| `BridgedConfig.Worker` profile shape (`baseUrl`, `model`, `configDir`, …) | `config` | **mostly adapter-shaped** — see §5 | +| `FleetConfig.Worker` profile shape (`baseUrl`, `model`, `configDir`, …) | `config` | **mostly adapter-shaped** — see §5 | | FSM, registry, roster, `reapIdle`/`drainAll`/`contextCap`, `rosterView` | `SessionManager` | **stays core** | | turn/completion detection (`TurnListener`, `CompletionResolver`, `StatusPoller`, `WorkerPresence`) | `inject/` | **stays core**, but reads herdr terminal output → transport-coupled (§4b) | | message store & routing | `msg/` | **stays core** | @@ -118,7 +118,7 @@ hard-wire "turns come from herdr". ## 5. Config shape -`BridgedConfig.Worker` is Claude-shaped (`baseUrl`, `model`, `configDir`, `tokenEnv`). Rather than +`FleetConfig.Worker` is Claude-shaped (`baseUrl`, `model`, `configDir`, `tokenEnv`). Rather than break existing YAML, CB-401 keeps `workers:` exactly as-is and treats those fields as the **ClaudeCodeLauncher's** profile schema. A future peer kind adds a `kind:` discriminator (default `"claude-code"`) selecting the launcher; unknown-kind → clear config error. No migration of @@ -185,13 +185,13 @@ messages, not verified facts. Deliverable for CB-401 Stage A — mechanical, behaviour-preserving: 1. `PeerLauncher` interface + `PeerHandle` (opaque id) + `SpawnRequest` (profile, requestedCwd, - callerCwd) + `Capability` enum, new package `dev.ltms.bridged.peer`. + callerCwd) + `Capability` enum, new package `dev.ltms.fleet.peer`. 2. `ClaudeCodeLauncher implements PeerLauncher` = today's `WorkerService`, adapted: `spawn(...)` returns a `PeerHandle` (id = paneId), `capabilities()` declares `MID_TURN_ASK, SELF_PR(when token), WORKTREE, ORPHAN_REAP`. 3. `SessionManager` depends on `PeerLauncher`, not `WorkerService` concretely; routing keys on `PeerHandle.id()` (== paneId today, so zero value change). -4. `Bridged.main` wires the concrete `ClaudeCodeLauncher` behind the interface. +4. `Fleetd.main` wires the concrete `ClaudeCodeLauncher` behind the interface. 5. **No behaviour change, no config change.** Full green gate: `ide_sync` → `ide_diagnostics` (0 errors/0 warnings) → `mvn clean install` with `MVN_EXIT` captured (no masking pipe). All existing tests pass unchanged; add tests only for the new `PeerHandle` indirection. diff --git a/docs/CB-402-OpenCode-Adapter.md b/docs/CB-402-OpenCode-Adapter.md index 282fd03..c1de564 100644 --- a/docs/CB-402-OpenCode-Adapter.md +++ b/docs/CB-402-OpenCode-Adapter.md @@ -11,7 +11,7 @@ All five increments of §4 are done, including increment 5 (the §5 live checkli ## 1. Goal -Prove the [`PeerLauncher`](../bridged/src/main/java/dev/ltms/bridged/peer/PeerLauncher.java) SPI +Prove the [`PeerLauncher`](../bridged/src/main/java/dev/ltms/fleet/peer/PeerLauncher.java) SPI actually holds for a **non-Claude** coding agent by shipping a second, first-class in-tree adapter: **opencode** (`opencode` 1.1.31, a provider-agnostic terminal coding agent). @@ -56,8 +56,8 @@ flowchart TB *Figure 1 — the two concerns tangled inside today's single launcher; CB-402 splits them.* -There is also a **Stage-A deferral** to finish: `Bridged.main` still casts -`(ClaudeCodeLauncher) workers` at the `BridgeMcp` and `BridgedApp` constructors. Those two +There is also a **Stage-A deferral** to finish: `Fleetd.main` still casts +`(ClaudeCodeLauncher) workers` at the `FleetMcp` and `FleetApp` constructors. Those two callers only invoke `profiles()`, `defaultProfile()`, and `list()` — **all already on the `PeerLauncher` interface**. The cast survives for one reason only: `PeerLauncher.list()` returns `List` (element type erased) while the callers use `Agent` element methods in their @@ -68,7 +68,7 @@ roster join. Finishing the migration is therefore small and contained (§4.D). ## 3. Target design Template-Method base + two thin adapters + a routing composite that keeps the Stage-A seam -(one `PeerLauncher` reference held by `SessionManager` / `BridgeMcp` / `BridgedApp`) intact. +(one `PeerLauncher` reference held by `SessionManager` / `FleetMcp` / `FleetApp`) intact. ```mermaid flowchart TB @@ -114,7 +114,7 @@ Two adapter **hooks** (abstract): protected abstract String namePrefix(); // "claude" | "opencode" /** Build the peer-specific launch: env map + argv. Runs any pre-spawn guard here. */ -protected abstract Launch buildLaunch(BridgedConfig.Worker cfg, SpawnRequest req); +protected abstract Launch buildLaunch(FleetConfig.Worker cfg, SpawnRequest req); record Launch(Map env, List argv) {} ``` @@ -126,7 +126,7 @@ so the opencode adapter never reaps a `claude-*` pane and vice-versa. The compos ### B. `kind:` config discriminator -Add one field to `BridgedConfig.Worker`: +Add one field to `FleetConfig.Worker`: ```java String kind // "claude-code" (default) | "opencode" @@ -164,13 +164,13 @@ String kind // "claude-code" (default) | "opencode" ### D. `CompositePeerLauncher` + finish the Stage-A migration -- `Bridged.main` groups configured profiles by `kind`, instantiates one launcher per kind +- `Fleetd.main` groups configured profiles by `kind`, instantiates one launcher per kind present, and wraps them in `CompositePeerLauncher implements PeerLauncher`. - Routing methods (`spawn(req)`, `effectiveCwd(req)`, `parityOverlay(name)`) dispatch by the profile's kind. Fan-out methods (`list()`, `reapOrphanWorkers()`, `capabilities()`, `profiles()`, `defaultProfile()`) merge across sub-launchers. `stop(id)` tries each (teardown only knows the pane id) — already best-effort/idempotent. -- **Migrate `BridgeMcp` + `BridgedApp` to the `PeerLauncher` interface**, dropping both +- **Migrate `FleetMcp` + `FleetApp` to the `PeerLauncher` interface**, dropping both `(ClaudeCodeLauncher)` casts. Only friction is `list()`'s `List`; resolve by giving the SPI a typed roster element (small neutral `PeerAgent` view exposing `id()`/`name()`/status) that the CB-304 roster join consumes — or, minimally, narrow at the callsite. Prefer the typed view. @@ -179,7 +179,7 @@ String kind // "claude-code" (default) | "opencode" sequenceDiagram autonumber participant P as Primary - participant M as BridgeMcp / REST + participant M as FleetMcp / REST participant C as CompositePeerLauncher participant O as OpenCodeLauncher participant B as HerdrPeerLauncher (base) @@ -244,7 +244,7 @@ until it is "major" (Stage-B whole), per the CB-401 bar. - **Unit (hermetic):** base-extraction regression (existing `ClaudeCodeLauncher` tests pass unchanged); `OpenCodeLauncher.buildLaunch` env/argv/config assertions; `kind` normalization - in `BridgedConfigTest`; `CompositePeerLauncher` routing + fan-out (merge of `profiles()`, + in `FleetConfigTest`; `CompositePeerLauncher` routing + fan-out (merge of `profiles()`, summed `reapOrphanWorkers()`, per-kind reap isolation) with fake sub-launchers. - **Live (dogfood, manual):** the §5 checklist on the running daemon. - **Gate (primary):** IDE diagnostics 0/0 on every changed file, `mvn clean install` green with @@ -269,7 +269,7 @@ until it is "major" (Stage-B whole), per the CB-401 bar. ## 8. As-built — live dogfood (2026-07-29) -Run against `bridged` on `127.0.0.1:8766` at main `19cdf8d`, with opencode **1.18.5** installed +Run against `fleetd` on `127.0.0.1:8766` at main `19cdf8d`, with opencode **1.18.5** installed via Homebrew. Every §5 risk is now a verified fact rather than an assumption. **The version-drift risk was the real one, and it did not bite.** This adapter was designed against diff --git a/docs/CB-500-Multi-Tier-Coordination.md b/docs/CB-500-Multi-Tier-Coordination.md index 9bc3d59..964c32c 100644 --- a/docs/CB-500-Multi-Tier-Coordination.md +++ b/docs/CB-500-Multi-Tier-Coordination.md @@ -38,7 +38,7 @@ never toolchain ownership (§7). flowchart TB human["human (types)"] primary["PRIMARY (Opus)
MCP client — pull-only"] - daemon["bridged daemon
127.0.0.1:8765 (single host)"] + daemon["fleetd daemon
127.0.0.1:8765 (single host)"] comp["CompositePeerLauncher
routes by kind"] cc["ClaudeCodeLauncher"] oc["OpenCodeLauncher"] @@ -80,7 +80,7 @@ flowchart TB m1["main A: Opus
MCP client"] m2["main B: cloud module
MCP client"] end - subgraph bus["bridged fabric (CB-307/308 substrate)"] + subgraph bus["fleetd fabric (CB-307/308 substrate)"] chan["per-agent inbox channels
agent.<globalId>.inbox"] roster["federated roster (union view)"] end @@ -146,7 +146,7 @@ env-manager" rule (§7).* ```mermaid sequenceDiagram participant M as main (delegator) - participant D as bridged + participant D as fleetd participant SL as SandboxLauncher participant SB as sandbox (peer-owned) participant A as agent in sandbox @@ -189,7 +189,7 @@ flowchart TB m2["main B: cloud module
MCP client (pull-only)"] end reg["PrimaryRegistry → multi-slot
(terminal per main)"] - subgraph fabric["bridged"] + subgraph fabric["fleetd"] ca["agent.A.inbox"] cb["agent.B.inbox"] push["ReplyPushLoop → N terminals"] @@ -211,7 +211,7 @@ transport is already peer-neutral — what was missing is N pull endpoints).* ```mermaid sequenceDiagram participant MA as main A (Opus) - participant BR as bridged / broker + participant BR as fleetd / broker participant MB as main B (cloud) MA->>BR: fleet_send(to = main B, msg) BR->>BR: publish agent.B.inbox (durable, msg id) @@ -233,7 +233,7 @@ push-loop fan-out; relax "orchestration tools only the primary calls" to "any re ## 6. Development C — Orchestrator tier > **SUPERSEDED — do not implement this model.** The operator rejected a supervisor above the lead. -> The human continues to drive the pre-existing lead directly; bridged neither spawns nor resumes that +> The human continues to drive the pre-existing lead directly; fleetd neither spawns nor resumes that > lead. What replaced this proposal is **one lead, two short-lived advisory architects, and N workers**: > the lead engages architects sideways for a strong-model assessment, then discards them. Architect > slots are declared in `architects:` (see Gitea issue #16), rather than making leads managed sessions. @@ -243,7 +243,7 @@ push-loop fan-out; relax "orchestration tools only the primary calls" to "any re > **Historical alternative retained.** The text and figures below record the considered model and why it > was rejected: it re-rooted the human-facing session above the lead, violating the still-true premise -> that configured leaders pre-exist, are recognised, and cannot be resumed by bridged. +> that configured leaders pre-exist, are recognised, and cannot be resumed by fleetd. The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker* sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**. @@ -385,7 +385,7 @@ ticket is an extension of an existing pattern (CB-402 for A, CB-308 for B/C). The follow-up question — *"clarify the architecture when we have distributed agents in sandboxes"* — resolves the fork left open in §4 and §9. **Decision: Development A (sandbox launcher) and CB-308 -(per-host federation) *compose*, not compete — each host runs a `bridged` gateway whose launcher +(per-host federation) *compose*, not compete — each host runs a `fleetd` gateway whose launcher spawns agents into that host's *local* sandboxes.** A sandbox is never reached across the network; it is reached by the gateway sitting next to it. @@ -393,7 +393,7 @@ is reached by the gateway sitting next to it. The bus delivers a turn by **herdr keystroke-injection** — `Injector → AgentControl.send` writes into a PTY that its **local** herdr owns. The broker moves *messages and presence*, **never keystrokes**. -So an agent's PTY must live in a herdr that *some* `bridged` instance drives locally: a remote +So an agent's PTY must live in a herdr that *some* `fleetd` instance drives locally: a remote container with no local herdr **cannot be injected into**. That rules out a central daemon reaching remote PTYs, and collapses the design to a single identity: @@ -403,7 +403,7 @@ remote PTYs, and collapses the design to a single identity: flowchart TB subgraph hostA["HOST A — gateway"] mA["main / orchestrator
MCP client → LOCAL gateway"] - gA["bridged A
herdr + CompositePeerLauncher
(incl. SandboxLauncher)"] + gA["fleetd A
herdr + CompositePeerLauncher
(incl. SandboxLauncher)"] cBEa["sandbox: backend
(local container)"] cFEa["sandbox: frontend
(local container)"] mA --- gA @@ -415,7 +415,7 @@ flowchart TB roster["roster.* (federated presence)"] end subgraph hostB["HOST B — gateway"] - gB["bridged B
herdr + SandboxLauncher"] + gB["fleetd B
herdr + SandboxLauncher"] cBEb["sandbox: backend
(local container)"] gB -->|"spawn → PTY in B's herdr"| cBEb end @@ -470,7 +470,7 @@ unchanged — the sandbox is transparent to it.* | Concern | Provided by | |---|---| -| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `bridged`, evolved) | +| Per-host gateway (owns local herdr + sessions) | **CB-308** (today's `fleetd`, evolved) | | Spawn into a local sandbox / role→image | **Development A** `SandboxLauncher` (§4), routed by `CompositePeerLauncher` | | Addressing a remote sandboxed agent | **CB-308** global id + federated roster (host + role as metadata) | | Orphan reap after a gateway restart | **CB-117** per-gateway, summed by the composite — each reaps only its **local** herdr | diff --git a/docs/CB-591-Gateway-Migration.md b/docs/CB-591-Gateway-Migration.md index f690ed9..e18d40f 100644 --- a/docs/CB-591-Gateway-Migration.md +++ b/docs/CB-591-Gateway-Migration.md @@ -51,7 +51,7 @@ Three facts about our side that decide the shape of this work: `ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config). `local` sets no `tokenEnv` today, because a direct vLLM needs no token. 2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is - `guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the + `guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Fleetd.java:93` and handed to the launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing this makes every `local` spawn throw. 3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is @@ -132,7 +132,7 @@ own provider credentials. So 3b needs **no allowlist change**; only 3a does. **Both still need a restart, for a different reason.** `tokenEnv` is resolved by `HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process -environment**. The running `bridged` inherited its environment when it started, so a variable added to +environment**. The running `fleetd` inherited its environment when it started, so a variable added to `secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh` (`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use @@ -207,7 +207,7 @@ worker's honesty rule leans on it ("never claim the result of a check you had no - **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow. - **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful - for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a + for library work. But `mcpUrl` in `FleetConfig.Profile` is a **single `String`**, so a claude-code member can mount exactly one MCP — this needs a code change, not a config edit. Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7 @@ -383,7 +383,7 @@ one turn; this one costs the whole task and is indistinguishable from a slow wor - `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear. - **Reasoning survives both surfaces** — see §3b above. - The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-` - key rather than the `bridged-local-noauth` placeholder. + key rather than the `fleetd-local-noauth` placeholder. - `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local` spawned without throwing, which is the check that catches a missed restart. diff --git a/docs/CB-5xx-Hardening.md b/docs/CB-5xx-Hardening.md index 48d2d27..ada5f87 100644 --- a/docs/CB-5xx-Hardening.md +++ b/docs/CB-5xx-Hardening.md @@ -12,12 +12,12 @@ not after it. ## 1. Why this stage is not optional bookkeeping -`bridged` today has **exactly one security control: the loopback bind**. Every other guarantee +`fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee rests on it. The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone — the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another -worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `bridged` host); the +worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the token path is the split-host fallback."* The token path does not exist yet. That leaves a seam that is **latent today and load-bearing the moment the bind moves**: @@ -112,7 +112,7 @@ The roadmap says "systemd unit". **This host is macOS — there is no systemd on not found), and the daemon that has been dogfooded for weeks runs as a bare foreground `java -jar`. Ship **both**: -- `deploy/dev.ltms.bridged.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and +- `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and ordered start after herdr. - `deploy/bridged.service` — systemd unit for the Linux gateways CB-308 introduces. @@ -158,10 +158,10 @@ configured. It already leaks nothing but herdr's version and up/down. ### 3.1 There are TWO entry paths, and only one of them has identity today The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not -literally true, and the difference is security-relevant.** `BridgeMcp` calls `MessageService` / +literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` / `SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is mounted as a raw servlet on Jetty's `ServletContextHandler` -(`BridgedApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through +(`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through Javalin's `before` filters at all. The current split is the mirror image of what you'd expect: @@ -226,7 +226,7 @@ was answered from the forge itself. Kept here as the decision record.* 1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon; TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an - `amqps://` URI. No keystore handling in `bridged`. + `amqps://` URI. No keystore handling in `fleetd`. 2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the diff --git a/docs/M4-Fleet-Health.md b/docs/M4-Fleet-Health.md index 99c74d7..dd9da54 100644 --- a/docs/M4-Fleet-Health.md +++ b/docs/M4-Fleet-Health.md @@ -72,7 +72,7 @@ post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`. ### 2.2 Corrections made during design -The first state table missed `BUSY` in bridged plus `IDLE` or `DONE` in herdr. It would have found +The first state table missed `BUSY` in fleetd plus `IDLE` or `DONE` in herdr. It would have found the real trace only through a late, weak stall timer. The final model adds `TURN_BOUNDARY_LOST` as a strong disagreement state. @@ -161,7 +161,7 @@ request or merge state. If committed work appears with `REPLY_STRANDED` or | State | Exact evidence | Certainty and action | |---|---|---| -| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the bridged-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. | +| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the fleetd-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. | A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global link is a target fault, not a control-link fault. @@ -298,7 +298,7 @@ Compiled brakes apply even if config asks for more: - at most two pane reads occur in one fleet tick; - targets rotate fairly; - only a normalised digest and optional clipped local excerpt are stored; -- no pane excerpt leaves bridged in a human webhook. +- no pane excerpt leaves fleetd in a human webhook. Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress comparison. @@ -330,7 +330,7 @@ The reader uses AMQP `content_type`, never body sniffing: ```text Legacy v0: text/plain -Typed family: application/vnd.ltms.bridged.inbox-message+json +Typed family: application/vnd.ltms.fleet.inbox-message+json ``` A legacy reply may begin with `{`. It remains plain text because its media type is `text/plain`. @@ -520,7 +520,7 @@ A repaired result uses distinct `RECONCILED_COMPLETION` values in `Rendezvous`, task poll source, and metrics. The lead sees: ```text -[repaired completion - bridged detected a lost turn boundary. The member did not call +[repaired completion - fleetd detected a lost turn boundary. The member did not call fleet_reply; pane-derived text follows and may be partial] ``` diff --git a/docs/MCP-Contract.md b/docs/MCP-Contract.md index 4fac394..2e79765 100644 --- a/docs/MCP-Contract.md +++ b/docs/MCP-Contract.md @@ -1,4 +1,4 @@ -# MCP Contract — `bridged`'s unified gateway +# MCP Contract — `fleetd`'s unified gateway > **Status: 🔴 HISTORICAL DESIGN — do NOT use as the tool reference.** Written 2026-07-14, before > any MCP code existed. The system shipped and this page never caught up, so **its tool names, @@ -14,13 +14,13 @@ > > **The authoritative tool surface is the live MCP schema** (each tool's own description and > parameters, as mounted), with the intent→tool table in `CLAUDE.md` as the short form. Both were -> checked against `mcp/BridgeMcp.java` on 2026-08-17 and are accurate. +> checked against `mcp/FleetMcp.java` on 2026-08-17 and are accurate. > > What is still worth reading here is **§6 — the flows and the error model** (rendezvous, > `fleet_ask`, detached delivery, the turn-done fallback). The shapes it describes are the ones > that shipped; only the names around them drifted. Rewriting this page is tracked as **CB-609**. -`bridged` is the **sole communication gateway** for every Claude session in the bridge. Both +`fleetd` is the **sole communication gateway** for every Claude session in the bridge. Both the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code) mount the *same* MCP server with a single `claude mcp add` line, and talk only through its tools. No Claude session ever addresses a broker, a peer, or the network directly. @@ -35,18 +35,18 @@ semantics, and how it maps onto the code already in the tree. These come from the project's core invariants and bound every decision below. 1. **One server, both roles.** The primary and all workers mount an identical server. The - catalog must serve both, and `bridged` must decide *who is calling* from the connection — + catalog must serve both, and `fleetd` must decide *who is calling* from the connection — never from a caller-supplied argument that could be spoofed. 2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards `ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription. Enforced today by [`SubscriptionGuard`](1-Architecture). 3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a - *single* MCP call that `bridged` holds open — never a cross-turn poll loop that would burn + *single* MCP call that `fleetd` holds open — never a cross-turn poll loop that would burn subscription quota. 4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing [`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most one message per turn. -5. **`bridged` owns policy; herdr owns PTYs.** MCP tools express *intent*; `bridged` +5. **`fleetd` owns policy; herdr owns PTYs.** MCP tools express *intent*; `fleetd` translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls. --- @@ -59,7 +59,7 @@ face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients an ```mermaid flowchart LR OPUS["Opus — primary
(Claude Code, env CLEAN)
MCP client"] - subgraph BD["bridged — standalone daemon"] + subgraph BD["fleetd — standalone daemon"] MCP["MCP server (north face)
bridge_send · bridge_reply
bridge_ask · bridge_status · lifecycle"] RDV["rendezvous registry
(blocking-call waiters)"] INJ["Injector + StatusPoller
(status-gated writer)"] @@ -87,10 +87,10 @@ flowchart LR ## 3. Identity & addressing -Because the same server is mounted by everyone, `bridged` resolves the caller's role on every +Because the same server is mounted by everyone, `fleetd` resolves the caller's role on every request — this is the linchpin of the whole contract and has no code yet. -- **Workers are known.** `bridged` spawns every worker +- **Workers are known.** `fleetd` spawns every worker ([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on the returned [`Agent`]. When a call arrives on a connection that maps to a known worker, the caller is *that* worker — so **workers never pass a target**; routing is implicit. @@ -106,13 +106,13 @@ request — this is the linchpin of the whole contract and has no code yet. ## 4. Transport -`bridged` is a long-lived daemon serving **multiple** concurrent clients (one primary + N +`fleetd` is a long-lived daemon serving **multiple** concurrent clients (one primary + N workers), so a per-client stdio child is the wrong shape. The recommended transport is **streamable-HTTP / SSE** on the same bind as the REST face: ```bash # identical on primary and every worker -claude mcp add --transport http bridged http://127.0.0.1:8080/mcp +claude mcp add --transport http fleetd http://127.0.0.1:8080/mcp ``` This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions). @@ -157,7 +157,7 @@ This adds an MCP-server dependency the pom does not yet carry. See [Open decisio - **Params:** `text` (required); `final` (default `true`). - **Behavior:** resolve the primary waiter registered against this worker with `text`. If no - waiter exists (detached delegation), `bridged` **injects the primary's idle pane** instead. + waiter exists (detached delegation), `fleetd` **injects the primary's idle pane** instead. Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit. #### `bridge_ask` @@ -216,7 +216,7 @@ One blocking call, zero polls. ```mermaid sequenceDiagram participant P as Primary (Opus) - participant B as bridged (MCP + Injector) + participant B as fleetd (MCP + Injector) participant H as herdr participant W as Worker (Claude) @@ -237,7 +237,7 @@ The worker pauses mid-turn to ask; the primary answers; the worker resumes in th ```mermaid sequenceDiagram participant P as Primary - participant B as bridged + participant B as fleetd participant W as Worker P->>B: fleet_send("do X", target=w) — blocks @@ -258,7 +258,7 @@ The primary does not block; the reply arrives later in its idle pane. ```mermaid sequenceDiagram participant P as Primary - participant B as bridged + participant B as fleetd participant W as Worker P->>B: fleet_send("do X", target=w, block=false) @@ -272,13 +272,13 @@ sequenceDiagram ### 6.4 Uncooperative worker — turn-done fallback -A worker that never calls `fleet_reply` still returns a result: `bridged` reads its terminal +A worker that never calls `fleet_reply` still returns a result: `fleetd` reads its terminal tail when the turn completes. ```mermaid sequenceDiagram participant P as Primary - participant B as bridged + participant B as fleetd participant W as Worker P->>B: fleet_send("do X", target=w) — blocks @@ -358,7 +358,7 @@ seam. Only the **rendezvous registry** and the **caller-identity resolver** are | `bridge_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only | | `bridge_read` | `AgentControl.read` | MCP adapter only | -Because the REST routes in `BridgedApp` already exercise the collaborators, MCP tools are +Because the REST routes in `FleetApp` already exercise the collaborators, MCP tools are validated by **parity** against those routes, not by re-testing behavior. --- diff --git a/docs/RELEASE-NOTES-1.0.0.md b/docs/RELEASE-NOTES-1.0.0.md index 82c96c5..d902ab4 100644 --- a/docs/RELEASE-NOTES-1.0.0.md +++ b/docs/RELEASE-NOTES-1.0.0.md @@ -1,8 +1,8 @@ # v1.0.0 — One leader, one host, complete -This is the first release of **`bridged`**. +This is the first release of **`fleetd`**. -`bridged` lets one main Claude Code session (the **leader**, on your Pro/Max subscription) +`fleetd` lets one main Claude Code session (the **leader**, on your Pro/Max subscription) run a team of **workers** — extra Claude Code sessions on a cheaper or local model, and non-Claude agents too. The leader's own session is never touched: it stays on subscription, with a clean environment. diff --git a/docs/Team.md b/docs/Team.md index d6ccb11..aeffc84 100644 --- a/docs/Team.md +++ b/docs/Team.md @@ -1,9 +1,9 @@ # Team — lead orchestrating a mixed Claude + local-LLM fleet -The message server (`bridged`) delivers **one turn into one worker**. A **team** is the +The message server (`fleetd`) delivers **one turn into one worker**. A **team** is the layer above it: a **Claude team-lead** that fans a job out across a **mixed fleet** of workers — some on Claude, some on the remote local LLM — and reduces their replies. Same -`bridged` delivery, same subscription boundary; this doc is only about **orchestration** — +`fleetd` delivery, same subscription boundary; this doc is only about **orchestration** — who the workers are, how the lead picks one, and how it runs many at once. > Delivery mechanics (blocking `POST /message`, status-gated reply envelope) live in the @@ -12,9 +12,9 @@ who the workers are, how the lead picks one, and how it runs many at once. ## The team - **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a - worker; a **thin client of `bridged`**. It plans, routes, dispatches, and integrates, and + worker; a **thin client of `fleetd`**. It plans, routes, dispatches, and integrates, and never sets `ANTHROPIC_BASE_URL`. -- **Workers** — a herd of `claude` panes in herdr, each an addressable `bridged` session +- **Workers** — a herd of `claude` panes in herdr, each an addressable `fleetd` session with its **own model/env**: - **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks. - **Local workers** (`ANTHROPIC_BASE_URL=https://llm.ltms.dev/anthropic`) — bulk, cheap, or @@ -28,7 +28,7 @@ MCP) — only its model differs. Scale each kind horizontally by adding panes. ```mermaid flowchart TB LEAD["lead — Opus
(Claude Code, env CLEAN)"] - BD["bridged
message server + router"] + BD["fleetd
message server + router"] HERDR["herdr
panes · agent-status"] WC1["w-claude-1
Sonnet · CLEAN"] WC2["w-claude-2
Sonnet · CLEAN"] @@ -62,14 +62,14 @@ flowchart TB | `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel | The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker -selection is **policy in the lead**, not a `bridged` concern — `bridged` just delivers to +selection is **policy in the lead**, not a `fleetd` concern — `fleetd` just delivers to the session the lead names. ## Subscription boundary in a team Unchanged from the base architecture, and it scales with the fleet: **only local-worker panes** launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on -the subscription. `bridged` enforces which panes may carry the off-subscription env, so +the subscription. `fleetd` enforces which panes may carry the off-subscription env, so adding workers never widens the boundary. ## Parallel fan-out (map / reduce) @@ -80,7 +80,7 @@ different workers at once, then results are gathered. ```mermaid sequenceDiagram participant L as lead (Opus) - participant B as bridged + participant B as fleetd participant WC as w-claude-1 participant WL as w-local-1 @@ -100,26 +100,26 @@ sequenceDiagram ``` - **Map:** the lead issues N concurrent blocking `POST /message` calls (one per subtask → its - chosen worker). Each call blocks only *that* request; `bridged` holds it open until the + chosen worker). Each call blocks only *that* request; `fleetd` holds it open until the worker's turn completes (status-gated) and returns the reply envelope. - **Reduce:** the lead collects the N envelopes and integrates. A slow local worker never blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum. - **Detached / long jobs** use the async broker path instead of a held request (Channel 2 in the base architecture), so the lead never busy-polls across turns. -Fan-out is bounded by the herd size (pane count) and `bridged`'s concurrency policy, not by +Fan-out is bounded by the herd size (pane count) and `fleetd`'s concurrency policy, not by the lead. ## Knowing the roster -The lead discovers its team from `bridged` (session list / roles) rather than hard-coding +The lead discovers its team from `fleetd` (session list / roles) rather than hard-coding pane ids, so workers can be added or restarted without editing the lead. A minimal charter in the lead's `CLAUDE.md` turns Opus into the orchestrator: ```markdown -## Your team (via bridged) -You are the team-lead. Delegate through the bridged client — never launch workers yourself. -Roster: ask bridged for current sessions/roles. +## Your team (via fleetd) +You are the team-lead. Delegate through the fleetd client — never launch workers yourself. +Roster: ask fleetd for current sessions/roles. - w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks. - w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks. @@ -133,22 +133,22 @@ tool instead of hand-rolling the HTTP request. ## What this layer does NOT change -- **Delivery** is still `bridged` → herdr `pane.send_text` + status events (Message-Server). +- **Delivery** is still `fleetd` → herdr `pane.send_text` + status events (Message-Server). - **Completion timing** is still the worker status event; **reply content** still rides the worker `Stop`-hook envelope. - **Single-host** still applies: herdr's socket is local, so the whole herd lives on the - `bridged` host. The lead may be remote — it only needs HTTP to `bridged`. + `fleetd` host. The lead may be remote — it only needs HTTP to `fleetd`. ## Open questions -- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `bridged` role-router +- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `fleetd` role-router (label-based). Start with the former; promote to the latter if routing logic grows. -- **Backpressure:** per-role concurrency caps in `bridged` so a fan-out can't exhaust the +- **Backpressure:** per-role concurrency caps in `fleetd` so a fan-out can't exhaust the local gateway. - **Result schema:** whether reply envelopes should carry structured metadata (worker, model, tokens) to help the lead's reduce step. ## Status -🟡 Design (2026-07-11). Orchestration layer over the selected `bridged` server; inherits +🟡 Design (2026-07-11). Orchestration layer over the selected `fleetd` server; inherits herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see the Message-Server design. diff --git a/docs/Worker-Git-Workflow.md b/docs/Worker-Git-Workflow.md index 4c45939..d6af7a5 100644 --- a/docs/Worker-Git-Workflow.md +++ b/docs/Worker-Git-Workflow.md @@ -76,7 +76,7 @@ worktree checks out anyway. Amber is the real gap — untracked local config the 3. **Never overlay the git plumbing** — the worktree's own `.git` file/branch is what gives isolation; that's the *one* thing that must differ from the main tree. -The overlay set lives in config (`BridgedConfig.Worker.parityOverlay` — a list of repo-relative +The overlay set lives in config (`FleetConfig.Worker.parityOverlay` — a list of repo-relative paths, with sane defaults) so it's auditable and per-repo tunable. > **Trust note (deliberate).** Hydrating local config means the primary's local secrets/tokens @@ -117,7 +117,7 @@ earlier `STATE.md` idea — a PR is reviewable, mergeable, and self-describing.* - **Remote:** `ssh://git@git.ltms.dev:2224/fleet/fleetd.git` (gitea). Push is over **SSH** — a worker running as the same user with the same keys can `git push` **with no extra credential**. -- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `bridged`). The +- **gitea is NOT in the project `.mcp.json`** (only `jetbrains`, `intellij-index`, `fleetd`). The primary's gitea MCP comes from a global/user config, so **workers do not inherit it**. A worker gets only the `bridge` MCP mounted (via `--mcp-config` launch flag). - **No gitea CLI** (`tea`) installed; `glab` is present but is the GitLab CLI (wrong backend). @@ -153,11 +153,11 @@ and the token is a single scoped secret the daemon injects like it already injec |---|---|---| | Worktree provision/teardown | **CB-301 ext** — `SessionManager.acquire`/`release`; `WorkerSession` gains `worktree`, `branch` | daemon shells out to `git worktree add/remove` | | **Config-parity overlay** | **CB-301 ext** — `SessionManager.acquire`, after `git worktree add` | symlink/copy the `parityOverlay` set into the worktree so the worker is a full peer; **this is what makes worktrees viable, not a dead-end** | -| Overlay config | `BridgedConfig.Worker.parityOverlay` — repo-relative paths, sane defaults | auditable, per-repo tunable; keep explicit + minimal (trust) | +| Overlay config | `FleetConfig.Worker.parityOverlay` — repo-relative paths, sane defaults | auditable, per-repo tunable; keep explicit + minimal (trust) | | Branch naming | `worker/-` off `main` (or a configured base) | one branch per session | | Commit + push + PR handoff | **CB-302** — worker-driven, guided by the skill | push = SSH; PR = option A | | Implementer skill | `.claude/skills/implementer/SKILL.md` | worktree-aware playbook (see below); mounts automatically since workers inherit repo cwd | -| gitea token injection | `WorkerService` env + `BridgedConfig` | repo-scoped, minimal perms | +| gitea token injection | `WorkerService` env + `FleetConfig` | repo-scoped, minimal perms | | PR review + merge | Primary (has gitea MCP + judgment) | merge on green; the human/primary gate stays | ## Implementer skill (outline) diff --git a/docs/Worker-Startup-and-Trust.md b/docs/Worker-Startup-and-Trust.md index 121a865..a9edda8 100644 --- a/docs/Worker-Startup-and-Trust.md +++ b/docs/Worker-Startup-and-Trust.md @@ -1,6 +1,6 @@ # Worker startup: working directory & the folder-trust prompt -When `bridged` spawns a worker, the worker CLI may show an **interactive startup prompt** before it +When `fleetd` spawns a worker, the worker CLI may show an **interactive startup prompt** before it is ready to accept a task — most importantly a *"Do you trust the files in this folder?"* dialog. An unattended worker parked on that prompt never becomes injectable: the status-gated injector waits for `idle`/`blocked`, the task is never delivered, and (worst case) a stray Enter answers the dialog @@ -18,7 +18,7 @@ flowchart TD B -->|"yes — told otherwise"| C["use that cwd"] B -->|"no"| D{"caller PID resolvable?
(MCP peer PID)"} D -->|"yes"| E["cwd = the primary's cwd
lsof -a -p PID -d cwd"] - D -->|"no (REST / off-host)"| F["cwd = bridged daemon cwd
(never $HOME by assumption)"] + D -->|"no (REST / off-host)"| F["cwd = fleetd daemon cwd
(never $HOME by assumption)"] C --> G["ensureWorkspace → tab.create → agent.start {cwd}"] E --> G F --> G @@ -52,9 +52,9 @@ only affect the seed shell, which the bridge closes). |---|--------|------| | 1 | Explicit `cwd` — a per-profile `cwd:` in config, or a spawn argument | "told otherwise" — pin a fixed workdir | | 2 | The **primary's cwd**, auto-detected from the `fleet_spawn` caller | normal MCP spawn from the primary | -| 3 | The `bridged` daemon's own cwd | REST spawn / off-host caller — **never `$HOME`** | +| 3 | The `fleetd` daemon's own cwd | REST spawn / off-host caller — **never `$HOME`** | -The primary's cwd (source 2) is discoverable with no new plumbing: `bridged` already resolves the MCP +The primary's cwd (source 2) is discoverable with no new plumbing: `fleetd` already resolves the MCP caller's loopback **peer PID** for connection identity (`ConnectionIdentity` → `LsofPeerPidLookup`); the same PID yields its cwd via `lsof -a -p -d cwd -Fn` (the `n…` line). The primary maps to no worker pane (it is not a worker), but its PID and cwd are still readable. @@ -62,7 +62,7 @@ worker pane (it is not a worker), but its PID and cwd are still readable. ```mermaid sequenceDiagram participant P as "Primary (main)" - participant B as "bridged" + participant B as "fleetd" participant O as "OS (lsof)" participant H as "herdr" P->>B: "fleet_spawn {profile} (no cwd)" @@ -77,7 +77,7 @@ sequenceDiagram *Figure 2 — a no-cwd spawn inherits the primary's directory from the caller's PID.* -> **Status:** implemented (CB-112). `bridged` threads the resolved `cwd` onto **`agent.start {cwd}`** +> **Status:** implemented (CB-112). `fleetd` threads the resolved `cwd` onto **`agent.start {cwd}`** > (verified: the worker process is rooted there), keeping the single shared worker space. On an MCP > `fleet_spawn` the primary's cwd is auto-detected from the caller's PID; over REST (no MCP caller) > it is the explicit `cwd` param else the daemon's cwd. Both placements (`tab` and legacy `pane`)