diff --git a/14-Fleet-Manager.md b/14-Fleet-Manager.md new file mode 100644 index 0000000..16f609e --- /dev/null +++ b/14-Fleet-Manager.md @@ -0,0 +1,111 @@ +# Fleet Manager + +`fleet-manager` gives one host a status view of several `fleetd` fleets. A fleet entry has a name, a REST address, and an optional project path. The manager is a separate Python package. It uses only REST requests and does not import or run inside `fleetd` (`README.md:3-12`, `fleet_manager/probe.py:10-13`, `pyproject.toml:1-9`). + +This is not multi-host federation. Entries marked `remote` are not probed. The manager reports them as remote instead (`fleet_manager/probe.py:117-121`, `tests/test_probe.py:31-34`). + +## Install and run + +The package needs Python 3.11 or later. Its runtime dependencies are Python standard-library modules only (`pyproject.toml:3-6`). + +For development, create a virtual environment and install the package with its test and lint tools: + +```sh +python -m venv .venv && .venv/bin/pip install -e '.[dev]' +``` + +This command is the documented development setup (`README.md:54-59`). + +The package registers the `fleet-status` command (`pyproject.toml:8-9`). It also runs as a Python module: + +```sh +fleet-status +python -m fleet_manager.cli +python -m fleet_manager.cli path/to/fleets.json +``` + +The module command accepts at most the first command-line argument as the registry path. Without an argument, it reads `fleets.json` beside the package source, not the current directory (`fleet_manager/cli.py:54-57`). If that file does not exist, it prints setup help to standard error and exits with status 2 (`fleet_manager/cli.py:58-62`). + +## `fleets.json` registry + +The registry is JSON with a required top-level `fleets` array. The command reads `reg["fleets"]`, so a missing `fleets` key stops the command with a key lookup error (`fleet_manager/cli.py:49-51`, `fleet_manager/cli.py:63-65`). + +Each array item is one fleet. The following example uses only the sample values: + +```json +{ + "fleets": [ + { + "name": "fleetd", + "rest": "http://127.0.0.1:8765", + "project": "/path/to/the/project/this/fleet/works/on" + }, + { + "name": "other-host-fleet", + "remote": true, + "project": "/path/on/that/host", + "note": "REST is loopback-only, not reachable from here yet" + } + ] +} +``` + +The sample shape is from `fleets.example.json:1-16`. + +| Field | Required | Meaning and missing-field result | +|---|---|---| +| `name` | Yes | The displayed fleet name. A missing value stops probing that entry with a key lookup error (`fleet_manager/probe.py:114-115`). | +| `rest` | Yes for a local fleet | The base REST URL. A missing value stops a local probe with a key lookup error. A remote entry does not read it (`fleet_manager/probe.py:114-115`, `fleet_manager/probe.py:117-124`). | +| `project` | No | An optional project path shown in the status output. When absent, no project line is shown (`fleet_manager/probe.py:114-115`, `fleet_manager/cli.py:32-33`). | +| `remote` | No | When true, the manager does not call the fleet and reports `remote`. When absent or false, it treats the entry as local (`fleet_manager/probe.py:117-123`). | +| `note` | No | An optional message for a remote fleet. If absent, the manager uses `on another host; REST is loopback-only, not probed from here` (`fleet_manager/probe.py:117-121`). | + +## `fleet-status` + +`fleet-status` is the installed status command. It loads every entry in `fleets`, probes each one, and prints the result (`pyproject.toml:8-9`, `fleet_manager/cli.py:63-66`). + +The output has a time-stamped `FLEET STATUS` heading. For each fleet, it shows the name, status, REST URL, optional project and note, a verdict when the fleet is reachable, and each returned member's role, profile, state, live status, branch, and progress result (`fleet_manager/cli.py:27-46`). + +## `python -m fleet_manager.cli` + +This module command runs the same `main` function as `fleet-status` (`fleet_manager/cli.py:54-70`, `pyproject.toml:8-9`). It uses the package-side default registry with no path, or the first supplied path with one (`fleet_manager/cli.py:54-57`). + +## How a probe works + +For a local fleet, the manager requests `/healthz` first. If the request fails, it reports `down`, adds an `/healthz unreachable` note, and does not request members (`fleet_manager/probe.py:123-129`, `tests/test_probe.py:37-42`). A successful health response supplies the displayed status and optional herdr version (`fleet_manager/probe.py:130-131`). + +It then requests `/members`. It reads the `workers` array when present. Each worker can supply `profile`, `role`, `state`, `liveStatus`, `worktree`, and `branch` (`fleet_manager/probe.py:133-145`). + +The manager does not trust a busy `liveStatus` by itself. For members whose live status is `working` or `busy`, it records the newest modification time in the worktree, waits four seconds by default, then records it again (`fleet_manager/probe.py:36-55`, `fleet_manager/probe.py:81`, `fleet_manager/probe.py:136-150`). A later time means `working`. No change means `stalled`, with the newest file age. An unavailable worktree gives `unknown`, not `stalled` (`fleet_manager/probe.py:64-78`, `tests/test_probe.py:86-92`). + +A fleet is `WORKING` when any watched member changed files. It is `STALLED` when a watched member did not change files. It is `busy (no worktree progress signal)` when a busy member has no usable worktree. A fleet with no members is `idle (no members)`; otherwise it can be `idle` (`fleet_manager/probe.py:156-166`, `tests/test_probe.py:45-57`). + +## Limits of the probe + +The manager handles several fleet entries on one host. It does not probe a remote entry, so it is not a cross-host status or federation tool (`README.md:3-7`, `fleet_manager/probe.py:117-121`, `tests/test_probe.py:31-34`). + +The probe cannot prove that a busy member is making useful progress. It only compares worktree modification times. A member can be busy without a usable worktree, and the result is `unknown` rather than a progress decision (`fleet_manager/probe.py:64-78`, `fleet_manager/probe.py:143-150`, `tests/test_probe.py:86-92`). + +## Two limits that come from `fleetd`, not from this tool + +Neither limit is visible in the `fleet-manager` source, because neither belongs to it. Both were +measured against running daemons, and both shape how the manager can be used. + +**Two `fleetd` daemons must not share one herdr session.** When they do, each daemon's reaper +treats the other's panes as orphans and tears them down, so the two fleets kill each other's +members. Give each daemon its own herdr session. + +**The manager must not run inside a herdr pane.** `fleetd` resolves a caller's role from the +connection, and it identifies its lead by a tab label. A Claude session sitting in a pane that some +*other* daemon does not name as its lead is resolved as a **worker** by that daemon, and every +orchestration call is then refused. A process outside any pane is resolved as the primary, which is +why the manager sits outside and speaks REST. + +Together these are the reason the manager is a separate REST client rather than a second daemon or +an in-pane session. + +## What this page could not confirm + +The `fleet-manager` source alone does not describe either limit above, and it was the only source +read for the rest of this page. Issues [#160](https://git.ltms.dev/fleet/fleetd/issues/160) and +[#156](https://git.ltms.dev/fleet/fleetd/issues/156) carry the measurements behind them. diff --git a/15-REST-API-Reference.md b/15-REST-API-Reference.md new file mode 100644 index 0000000..1352fb7 --- /dev/null +++ b/15-REST-API-Reference.md @@ -0,0 +1,231 @@ +# REST API Reference + +`fleetd` exposes 14 HTTP routes, built once in `FleetApp.build()` +(`fleetd/src/main/java/dev/ltms/fleet/rest/FleetApp.java:126-160`). No other page collects them in +one place — each route is currently documented only where it happens to matter for some other +topic, spread across seven pages. + +## Why this page exists + +The MCP tools (their contracts are defined in `FleetMcp.java`, not yet collected on their own +wiki page) and the REST routes on this page are **siblings that both sit on top of the same +service objects** — they are not a wrapper around each other in either direction. `fleetd`'s startup code builds one `MessageService` and one `SessionManager` and hands +the *same instances* to both faces: the MCP server (`Fleetd.java:524`) and the REST app +(`Fleetd.java:607`), from a `MessageService` built once at `Fleetd.java:449-450`. The MCP server's +own file header says this directly: its tools are "thin adapters over the same `MessageService`/ +`Rendezvous` the REST routes use — so the two are validated by parity, not by re-implementing +behaviour" (`FleetMcp.java:54-57`); `FleetApp.java:38-41` makes the same claim from the REST side. + +This matters for a reader in a concrete way: **you cannot infer REST behaviour from the MCP tool +contract.** The two surfaces read from the same source of truth, but each has its own request +shape, its own field names, and — as the table below shows — its own coverage. Two MCP tools have +no REST route at all, and several REST routes have no MCP tool. + +## All routes + +Every route below is wired in `FleetApp.build()` (`FleetApp.java:126-160`). The "Role" column is +the `Authz.Action` each handler checks (`Authz.java:44-71`) — "open" means the handler makes no +`allow()` call at all, so it needs no `Authorization` header. + +| Method | Path | What it does | Role required | Matching MCP tool | +|---|---|---|---|---| +| GET | `/healthz` | Liveness + herdr reachability | open (no auth check) | none | +| GET | `/metrics` | Prometheus scrape (only when metrics are configured) | any authenticated caller (`METRICS`) | none | +| GET | `/sessions` | Raw herdr workspace list, one row per workspace | any authenticated caller (`READ`) | none | +| GET | `/agents` | Raw herdr agent list (every agent herdr tracks) | any authenticated caller (`READ`) | none | +| GET | `/members` | The fleetd-owned worker roster, joined with live herdr status | any authenticated caller (`READ`) | `fleet_list` (its `members` section only — `fleet_list` also adds peer `leads` and capacity data this route does not have) | +| GET | `/profiles` | Configured worker profiles and the default one | any authenticated caller (`READ`) | `fleet_profiles` | +| POST | `/members` | Spawn a worker | primary only (`SPAWN`) | `fleet_spawn` | +| DELETE | `/members/{paneId}` | Tear a worker down by pane id | primary only (`STOP`) | `fleet_stop` | +| POST | `/sessions/{id}/message` | Deliver a turn to a session, or answer a worker's `fleet_ask` | primary or architect (`SEND`) | `fleet_send` | +| POST | `/sessions/{id}/reply` | A worker's (or peer lead's) structured reply for its own turn | caller must own the session (`REPLY`) | `fleet_reply` | +| GET | `/sessions/{id}/replies` | Drain a session's held-reply inbox | primary only (`DRAIN`) | `fleet_poll` (with `target` set) | +| POST | `/sessions/{id}/ask` | A worker's mid-turn question to the primary | caller must own the session (`ASK`) | `fleet_ask` | +| GET | `/sessions/{id}/status` | A session's live lifecycle status and readiness | any authenticated caller (`READ`) | `fleet_status` | +| GET | `/tasks/{ticket}` | Poll an async (`wait:false`) send by its ticket | any authenticated caller (`READ`) | `fleet_poll` (with `ticket` set) | + +**MCP tools with no REST route:** `fleet_ack` (`FleetMcp.java:1137-1146`, calls +`messages.ackReply(target, msgId)` directly — no `FleetApp` handler exists for it) and +`fleet_whoami` (`FleetMcp.java:1233-1246` — identity is resolved per-connection by +`ConnectionIdentity`/`CallerResolver`, so there is nothing for a REST caller to query; each REST +request already carries the resolved `Principal` under the `fleetd.caller` context attribute, +`FleetApp.java:56,138-142`). + +## Per-route detail + +### `POST /sessions/{id}/message` + +Handler: `sendMessage`, `FleetApp.java:431-474`. This is `fleet_send`'s REST face, and it also +carries `fleet_send`'s "answer a worker's `fleet_ask`" mode. + +Body fields (`FleetApp.java:441-445`): + +| Field | Type | Default | Meaning | +|---|---|---|---| +| `content` | string | required | The message to deliver, or your answer when `turnId` is set. A blank value is a `400 bad_request`. | +| `turnId` | string | none | When set, this call answers a worker's open `fleet_ask` instead of starting a new delegation, and it always blocks. | +| `timeoutMs` | integer | `25000` | Max time to wait for a reply. Clamped to `[1, 120000]` (`FleetApp.java:49-50,454`). | +| `wait` | boolean | `true` | `false` returns a ticket immediately (fire-and-poll); the caller then polls `GET /tasks/{ticket}`. | + +Response shape depends on the outcome (`writeReply`, `FleetApp.java:481-511`): + +| Outcome | Status | Body | +|---|---|---| +| Worker asked a question | 202 | `{sessionId, status:"question", question, turnId}` | +| Answering a stale/expired turn | 409 | `{sessionId, error:"stale_turn", detail}` | +| Worker replied normally | 200 | `{sessionId, reply, replySource:"reply"}` | +| Worker finished without calling `fleet_reply` | 200 | `{sessionId, reply, replySource:"transcript"}` — a scrape of the worker's terminal tail, not a structured reply | +| Timed out, worker still working | 202 | `{sessionId, status:"working", detail}` | +| Timed out, message still queued | 202 | `{sessionId, status:"queued", detail}` | +| Worker busy on another turn | 202 | `{sessionId, status:"busy", detail}` | +| Worker failed | 202 | `{sessionId, status:"failed", detail}` | +| Backend credential exhausted | 202 | `{sessionId, status:"backend_exhausted", detail}` | + +A herdr-level failure during the blocking send (not the `turnId` path) maps through +`herdrError` (`FleetApp.java:650-656`): a `..._not_found` herdr error code becomes +`404 session_not_found`, anything else becomes `502 herdr_error`. + +### `POST /sessions/{id}/reply` + +Handler: `replyMessage`, `FleetApp.java:554-571`. This is `fleet_reply`'s REST face. + +Body: `{"content": "..."}`. Unlike `sendMessage`, there is no blank check on `content` — a missing +or empty value is delivered as `""`. Response: `200 {sessionId, delivered: true}`. + +The path's session id is the caller's own identity claim, and this is the one place that claim is +checked over REST — over MCP a worker's identity already comes from its connection, never an +argument, so it could never reply as another worker. Over REST the id in the URL used to be +trusted outright; `Authz.Action.REPLY` (`caller.ownsSession(targetSession)`, `Authz.java:66`) is +what closes that (`FleetApp.java:556-561`). + +`messages.reply()` either resolves an open blocking `POST /sessions/{id}/message`, or — when no +send is currently open for that session — queues the reply into the session's inbox +(`FleetApp.java:550-553`). See the warning below about draining that inbox. + +### `POST /sessions/{id}/ask` + +Handler: `askMessage`, `FleetApp.java:518-548`. This is `fleet_ask`'s REST face: a worker's +mid-turn question to whichever primary has a blocking send open on it. + +Body fields (`FleetApp.java:526-528`): `question` (string, required — blank is `400 bad_request`), +`timeoutMs` (default `55000`, clamped to `[1, 115000]`, `FleetApp.java:52-53,537`). + +Response: + +| Outcome | Status | Body | +|---|---|---| +| Primary answered | 200 | `{sessionId, answered:true, answer}` | +| No primary is waiting on this session | 409 | `{sessionId, error:"no_pending_send", detail}` | +| Primary did not answer in time | 202 | `{sessionId, status:"no_answer", detail}` | + +### `GET /sessions/{id}/replies` + +Handler: `drainReplies`, `FleetApp.java:578-588`. Response: `200 {sessionId, replies:[{msgId, +content}, ...]}`. + +**Warning — this call drains the inbox on its first read.** The javadoc calls the semantics +"at-least-once": reading removes the messages, so a second call returns nothing even if the first +caller never processed the result (`FleetApp.java:573-577`). If you need to keep the response, +write it to a file the first time you call this route — a second call will not give it back to +you. + +### `GET /tasks/{ticket}` + +Handler: `taskStatus`, `FleetApp.java:622-647`. Polls an async (`wait:false`) send by the ticket +`POST /sessions/{id}/message` returned. + +Response: `200 {ticket, phase, reply?, replySource?, detail?, turnId?}` — `reply`/`replySource` are +present once the phase is terminal with a result, `turnId` is present while the phase is +`asking` (the worker opened a `fleet_ask` mid-turn and this ticket now needs an answer, +`FleetApp.java:641-645`). An unknown or expired ticket returns `404 {error:"unknown_ticket", +detail}`. + +### `GET /members` + +Handler: `listMembers`, `FleetApp.java:309-329`. The fleetd-owned roster, joined against live herdr +status. + +**The response key is `workers`, not `members`** (`FleetApp.java:322`). The route was renamed from +`/workers` to `/members` in CB-557, but the body's key was left alone. A caller that reads +`body["members"]` gets nothing and sees an empty fleet rather than an error, so read `body["workers"]`. +`fleet-manager` already does (`fleet_manager/probe.py:134`). + +The join is on the **terminal id**, not the pane coordinate: the registry key is a host-unique id +while a pane coordinate is a per-daemon counter, so joining on the pane would mismatch as soon as a +second herdr daemon is configured (`FleetApp.java:314-317`). + +Each row is a `SessionManager.rosterView`, carrying at least `sessionId`, `paneId`, `profile`, +`role`, `state`, `worktree`, `branch`, `owner` and `liveStatus`. + +The body also carries an optional `wipRefs` object (`count`, `costBytes`) describing the +`refs/wip` snapshot store. It is **absent** until a worktree session has established the repo, so a +fleet that has never snapshotted reports nothing rather than zero (`FleetApp.java:324-327`). + +### `POST /members` + +Handler: `spawnMember`, `FleetApp.java:346-396`. Spawns a worker. + +Accepts `role`, `profile`, `cwd`, `worktree`, `ticket` either as query parameters or as a JSON +body — the body is only parsed when `profile`, `cwd` or `worktree` are missing from the query +string (`FleetApp.java:355-370`). `role` defaults to `dev` (`MemberRole.parse`, values are +`architect`, `dev`, `reviewer`); an unrecognised role is `400 {error:"unknown_role", detail}`. +`worktree` is either `"true"` (requires `ticket`) or a ticket-slug string +(`worktreeRequest`, `FleetApp.java:398-409`). + +On success: `201` with the new `MemberSession` view — `{terminalId, paneId, profile, cwd, +ownerTerminal, state, worktree?, branch?}` (`FleetApp.java:672-687`). + +Error mapping (`FleetApp.java:379-396`): + +| Condition | Status | Body | +|---|---|---| +| Subscription-boundary guard tripped | 403 | `{error:"subscription_boundary", detail}` | +| No profile has spawn capacity right now | 503 | `{error:"no_capacity", detail}` | +| Unknown profile name | 400 | `{error:"unknown_profile", detail}` | +| The peer never became reachable | 502 | `{error:"spawn_timeout", detail}` | + +One case is not covered by this table: `worktreeRequest` (`FleetApp.java:398-409`) is called +**before** the `try` block that catches `IllegalArgumentException`, so `worktree=true` with no +`ticket` throws past every one of these handlers rather than becoming a `400`. This is a gap +found while writing this page — it has not been reported as a bug, and no other caller of this +route has been checked for the same pattern. + +### `DELETE /members/{paneId}` + +Handler: `stopMember`, `FleetApp.java:416-423`. No body. Calls `sessions.release(paneId)` and +always answers `204`, with no path that surfaces an unknown `paneId` back to the caller. Whether +`sessions.release` itself does anything for an id it does not recognise was not checked for this +page. + +### `GET /healthz` + +Handler: `healthz`, `FleetApp.java:229-266`. The only route with no `allow()` call at all — it +takes no `Authorization` header and answers any caller that can reach the port. + +It pings the lead herdr daemon and, when a member daemon is separately configured +(`memberHerdrSocket`, `FleetApp.java:210-227`), pings that one too: + +| Configuration | Success body | Failure | +|---|---|---| +| One daemon (the common case) | `200 {status:"ok", herdr:{version, protocol}}` | `503 {status:"degraded", herdr:"unreachable", detail}` | +| Two daemons configured | adds `"member":{version, protocol}`, and `"protocolMismatch": true` when the two `protocol` numbers differ | a member-daemon failure is its own `503 {status:"degraded", herdr:"member unreachable", detail}`, reported separately from a lead-daemon failure so it is never masked by a healthy lead | + +**Warning — a green `/healthz` does not mean spawning works.** The socket can connect while the +herdr wire protocol on the other end has moved, so `ping` succeeds and every spawn still fails +(`FleetApp.java:218-227`). Check the `protocol` number in the response, and — where a member +daemon is configured — check the `member` key and the `protocolMismatch` flag specifically, since +every spawn goes through the member daemon, not the lead one. + +## What the REST surface is for + +Every capability behind an MCP tool is also reachable here, so the system can be driven and +tested with plain HTTP — no Claude session and no MCP client in the loop. This is what makes the +routes on this page an acceptance-test surface in their own right, not only a debugging aid: a +script can spawn a worker, send it a turn, and drain its replies using nothing but `curl` and the +field names on this page. + +This page does not cover the rendezvous flows themselves (the blocking-send / `fleet_ask` / +detached-delivery / turn-done sequences) — those are diagrammed in `docs/MCP-Contract.md` §6. That +section's diagrams describe the flow shapes correctly, but its field names (`turn_id`, +`block=false`, `outcome`) predate the code and do not match the request and response fields +documented on this page — use the names given here, not the ones in that document. diff --git a/16-Security-and-Trust-Boundary.md b/16-Security-and-Trust-Boundary.md new file mode 100644 index 0000000..e48eef8 --- /dev/null +++ b/16-Security-and-Trust-Boundary.md @@ -0,0 +1,201 @@ +# 16. Security and Trust Boundary + +> **Scope.** This page pulls the security model into one place. The model is real, but the code +> that carries it is spread across three areas — the guard package, the auth package, and the +> member package — so no single page showed it whole before this one. Every claim below was +> checked against the source at `main` `a814d1e` (2026-08-31). Where this page and the code ever +> disagree again, trust the code, not this page. + +`fleetd` draws four separate lines of defense: a **subscription boundary** that stops a worker +from billing the operator's paid Claude plan, an **identity and authorization** model that stops +one member from acting as another, an **environment control** on what a spawned member inherits +from the operator's own shell, and a **token scope** limit on what a member can do to the git +forge. The sections below cover each in turn, using the class that enforces it. + +```mermaid +flowchart TB + op["Operator's shell
own secret store"] + + subgraph primary_side["Primary — on subscription"] + primary["Primary (lead)
ANTHROPIC_BASE_URL must be absent"] + end + + subgraph member_side["Member — off subscription"] + member["Worker / architect pane
login shell"] + claude["claude / opencode process"] + end + + guard["guard.SubscriptionGuard
assertWorker / assertPrimaryClean"] + authz["auth.Authz + mcp.ConnectionIdentity
role table, identity from the connection"] + scrub["member.MemberEnvAllowList
+ EnvAllowListScrub (ZDOTDIR)"] + token["Repo-scoped forge token
write:repository only"] + + op -- "re-sources on login" --> member + guard -- "checked before spawn" --> member + guard -- "checked at daemon start" --> primary + authz -- "gates every fleet_* call" --> primary + authz -- "gates every fleet_* call" --> member + scrub -- "blanks names after the shell chain runs" --> member + member --> claude + token -- "injected into the worker env" --> claude + + classDef boundary fill:#6b46c1,stroke:#3d2a75,color:#ffffff; + class guard,authz,scrub,token boundary +``` +*Figure: the four controls (purple) that separate the operator's subscription and secrets from a +spawned member. The environment controls stop at the process environment — the argv gap in the +last section of this page is not one of the boxes above, because nothing here defends it.* + +## 1. The subscription boundary + +`SubscriptionGuard` (`fleetd/src/main/java/dev/ltms/fleet/guard/SubscriptionGuard.java`) enforces +one rule in two directions: a worker must run off the operator's paid subscription, and the +primary must never run on a base URL at all. + +- **`assertWorker(baseUrl)`** (`SubscriptionGuard.java:32-46`) refuses to spawn a worker whose + `ANTHROPIC_BASE_URL` is blank, not a valid URL, or resolves to a host outside a configured + allowlist. The allowlist is `FleetConfig.Guard.offSubscriptionHosts` + (`FleetConfig.java:1156-1159`) and it **defaults to an empty list**. Because the list is a config + value that differs by host, this page names no hostname. +- **`assertPrimaryClean(env)`** (`SubscriptionGuard.java:49-55`) refuses to let the daemon start if + its own environment carries `ANTHROPIC_BASE_URL`. `Fleetd.java:131-132` calls this against + `System.getenv()` at startup, before the daemon does anything else. If the primary's shell were + ever pointed at a base URL, the daemon would not come up. +- **`ClaudeCodeLauncher.buildLaunch`** (`fleetd/src/main/java/dev/ltms/fleet/member/ClaudeCodeLauncher.java:191-213`) + calls `assertWorker` right before every non-subscription spawn. + +**The opt-out: `subscription: true`.** A profile can deliberately choose to run a worker on the +operator's own subscription instead of an off-subscription endpoint (CB-539). This is meant for a +profile with no off-subscription endpoint to point at. It is enforced at two separate points, so a +worker cannot slip past the guard on the subscription path: + +- **At spawn time**, `ClaudeCodeLauncher.buildLaunch` (`ClaudeCodeLauncher.java:196-213, 227-233`) + refuses a profile that sets both `subscription: true` and a `baseUrl` — the two say opposite + things, so the launcher throws rather than picking one. On the subscription path it also strips + `ANTHROPIC_BASE_URL` and `ANTHROPIC_AUTH_TOKEN` from the worker's environment as a + belt-and-braces step, even though `subscription: true` should never carry them. +- **At config-load time**, `FleetConfig.validateSubscriptionProfiles()` (`FleetConfig.java:2058-2075`) + refuses to start the daemon at all if any `subscription: true` profile's `env:` block names + `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. The reason this must be a fatal, loud refusal + and not a silent strip: on the subscription path, `SubscriptionGuard` is skipped entirely for + that profile, so nothing else would catch a repoint (`FleetConfig.java:2044-2052`). + +## 2. Identity and authorization + +A member's identity comes from **the connection it calls on**, never from an argument in the +request. `ConnectionIdentity.resolve` (`fleetd/src/main/java/dev/ltms/fleet/mcp/ConnectionIdentity.java:42-48`) +takes only the remote address and port of the MCP connection. It maps the connecting process's PID +to a herdr pane through `PaneLocator`, and that pane maps to a worker's `terminal_id`. Both steps +use sources the caller cannot influence: the OS reports the real PID for a loopback connection, and +herdr owns the PID-to-pane mapping. A caller that maps to no worker pane — an off-host client, or +the primary itself — resolves to `null` (`ConnectionIdentity.java:9-14`). + +The role table lives in `Authz.permits` (`fleetd/src/main/java/dev/ltms/fleet/auth/Authz.java:44-72`): + +| Action | Who may perform it | Why (from the code's own comments) | +|---|---|---| +| `SPAWN`, `STOP`, `DRAIN` | primary only | fleet lifecycle is the primary's alone; an architect deliberately does not get these, even though it coordinates workers, so it cannot tear a fleet down | +| `SEND` | primary or architect | an architect delegates to workers, which is the point of the role, but still has no lifecycle rights; a worker sending would be it escalating into the orchestrator role | +| `REPLY`, `ASK` | the caller must own the target session | the load-bearing rule: a caller acts only as the pane it occupies | +| `READ`, `METRICS` | primary, worker, or architect | observation carries no secrets, so every authenticated role may read | + +The `REPLY`/`ASK` row is what stops one member from forging a reply for another member's +rendezvous. Because identity is read from the connection and not from a session id typed into the +request, a member cannot claim a session id that is not its own pane's. An unauthenticated caller +is refused for every action (`Authz.java:44-47`). + +## 3. What a member inherits, and what is blocked + +This is the part of the model that most needs care, because two different controls exist and only +one of them actually works once a login shell has run. + +**Why a spawn-time overlay is not enough.** A herdr pane runs a **login shell**, and that shell +re-sources the operator's own secret store. A control applied at pane-creation time — an +environment overlay handed to herdr before the shell starts — is therefore overwritten the moment +the operator's own `~/.zshrc` and secret files run (`FleetConfig.java:1182-1183`, +`EnvAllowListScrub.java:17-25`). This is the `deny-by-default` policy +(`FleetConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT`, `FleetConfig.java:1223`): it shadows each +known-but-not-allowed name in the pane-creation overlay, but a login shell that re-exports the same +name defeats it. + +**What actually holds is the `allow-list` policy** (`FleetConfig.MemberCredentials.POLICY_ALLOW_LIST`, +`FleetConfig.java:1235`), implemented in `EnvAllowListScrub` +(`fleetd/src/main/java/dev/ltms/fleet/member/EnvAllowListScrub.java`). It generates a per-spawn +`ZDOTDIR` whose startup files run the scrub **after** the whole operator chain has sourced, +blanking every exported variable that is not on a derived allow-list. Because the scrub runs last, +nothing sourced later can undo it (`EnvAllowListScrub.java:20-25`). The scrub is sourced from both +the generated `.zshrc` and `.zlogin`, because herdr opens a login zsh on macOS (where `.zlogin` +runs) and a plain interactive zsh on Linux (where it does not) — a scrub placed in only one of +those files would silently do nothing on the other platform (`EnvAllowListScrub.java:27-34`). This +control is **zsh-only**: it depends on `$ZDOTDIR` being a zsh mechanism, so a member whose pane +runs a different shell is not covered by it. + +**The effective allow-list is wider than what an operator writes.** `MemberEnvAllowList.derive` +(`fleetd/src/main/java/dev/ltms/fleet/member/MemberEnvAllowList.java:108-128`) builds the kept-name +set as a **union** of: + +- `INFRASTRUCTURE_PASSTHROUGH` — locations and shell settings such as `PATH`, `HOME`, `ZDOTDIR`, + and the `XDG_*` roots (`MemberEnvAllowList.java:78-84`), never credentials; +- every configured profile's `gitTokenEnv`, `gitHostEnv`, `tokenEnv`, and every key of that + profile's `env:` map (`MemberEnvAllowList.java:110-118`) — for **every** profile, not only the + one being spawned, so adding a profile can only widen the list, never narrow another spawn's + scrub; +- the operator's own `memberCredentials.allow:` list (`MemberEnvAllowList.java:120-126`). + +So the list a spawn actually keeps is a strict superset of what an operator types under `allow:`. +An operator who reads only their own `allow:` entry will underestimate what a member can keep. + +**`SSH_AUTH_SOCK` is excluded from that union** (`MemberEnvAllowList.java:122`). It is the +operator's ssh-agent socket, and a member holding it can sign with the operator's own keys, so it +is governed only by the separate `memberCredentials.sshAuthSock` setting +(`FleetConfig.java:1208-1216`), never by appearing in `allow:`. + +**The limit of an environment control.** Blocking the agent socket in the environment does not +stop a member reaching a passphrase-free private key **file** sitting on disk outside `~/.ssh`. An +environment control covers what a process can read from its environment variables; it says +nothing about the filesystem. A control at one layer does not close a gap at another. + +## 4. Tokens + +A member is given a repo-scoped forge token with **`write:repository`** scope only — enough to +create a branch and open a pull request — never the operator's own admin-scoped +`GITEA_ACCESS_TOKEN` (`fleetd/docs/Worker-Git-Workflow.md:139-144`). `GitWorktrees`'s credential +helper reads this scoped token from `WORKER_GITEA_TOKEN` +(`fleetd/src/main/java/dev/ltms/fleet/session/GitWorktrees.java:70-72`) and warns, in its own +comment, that a silent fallback to the primary's admin token would turn a loud crash into a quiet +privilege leak (`GitWorktrees.java:52-58`). `ClaudeCodeLauncher.buildLaunch` injects the scoped +token into the worker's environment as part of every spawn (`ClaudeCodeLauncher.java:240`). A lead +gets no git token at all — it reviews and merges through the operator's own credentials, and +`LeadLauncher.leadEnv` never grants it the scoped token a member gets +(`fleetd/src/main/java/dev/ltms/fleet/lead/LeadLauncher.java:287-288`). + +The reason for the narrow scope is direct: a member opens its own pull request and must never be +able to merge it. A leaked or misused member token can open pull requests. It cannot merge, delete +a branch it does not own, or touch the org. Merge authority stays with the operator or the primary. + +## The one honest limitation: argv is world-readable + +A process's command-line arguments (`argv`) are visible to any other process on the same host — +`ps -axww` prints them for every user session. If a program passes a credential as +`NAME=value` on its command line, that value is exposed the same way a value printed to a shared +log would be. Nothing in `fleetd` checks for this, and nothing in this page's other three sections +defends against it. + +`fleetd`, herdr, and `claude` are clean on this specific point: `fleetd` passes a member's +environment to herdr as a JSON field in the spawn request, not as command-line arguments — +`WorkspaceControl.createTab` and `WorkspaceControl.splitPane` both carry `env` as a separate +`params` entry alongside `cwd`, never appended to argv +(`fleetd/src/main/java/dev/ltms/fleet/herdr/WorkspaceControl.java:84-93, 100-109`). + +But this guarantee is narrow. It covers `fleetd`, herdr, and `claude` — nothing else. **A +different program started on the same host can still leak a credential through its own argv**, and +every control in this page — the subscription guard, the identity check, the environment scrub, +the token scope — is bypassed the moment that happens, because none of them touch argv. This is a +gap in the host, not in `fleetd`'s own code, and it is recorded here rather than left undocumented, +because a security page that hides its own limits is worse than one with none. + +--- + +This page does not cover herdr's own process isolation, the LavinMQ broker's per-vhost isolation +between two fleets, or session lifecycle and reaping. See [1. Architecture](1-Architecture) and +[9. Implementation Architecture](9-Implementation) for those. diff --git a/Home.md b/Home.md index 4ff8527..3819763 100644 --- a/Home.md +++ b/Home.md @@ -94,6 +94,9 @@ Read in order (the sidebar mirrors this): 11. **[Features](11-Features)** — what it can do, the knob that turns it on, why it exists, the gotcha 12. **[Claude → OpenCode](12-Claude-to-OpenCode)** — porting a workspace to a second host 13. **[User Guide](13-User-Guide)** — 🟢 **the operator page.** Install, configure, run, delegate, and the traps. +14. **[Fleet Manager](14-Fleet-Manager)** — the separate tool that shows several fleets on one host +15. **[REST API Reference](15-REST-API-Reference)** — every route, the role it needs, and its real body +16. **[Security & Trust Boundary](16-Security-and-Trust-Boundary)** — the subscription guard, the role table, and what a member's pane inherits ## Status diff --git a/_Sidebar.md b/_Sidebar.md index 10dd9d2..fcd651d 100644 --- a/_Sidebar.md +++ b/_Sidebar.md @@ -17,6 +17,9 @@ 11. [Features](11-Features) — what it can do · the knob that turns it on · why · the gotcha 12. [Claude → OpenCode](12-Claude-to-OpenCode) — porting a workspace to a second host 13. **[User Guide](13-User-Guide)** — 🟢 install · configure · run · delegate · the traps +14. [Fleet Manager](14-Fleet-Manager) — many fleets on one host, over REST +15. [REST API Reference](15-REST-API-Reference) — all 14 routes, roles, and bodies +16. [Security & Trust Boundary](16-Security-and-Trust-Boundary) — the guard · authz · what a member inherits --- 🟢 herdr-centric `fleetd` · AgentAPI = research, never built