#168 §B: add chapters 14, 15 and 16

Three pages the audit asked for that did not exist.

14 Fleet Manager — fleet-manager had zero coverage in this wiki, so an
operator had no way to learn it exists. Written from its own source: the
fleets.json shape from the parser rather than the example file, what each
command does, and how the probe decides working vs stalled by comparing
worktree modification times twice. It also records the two limits that come
from fleetd rather than from the tool: two daemons sharing one herdr session
tear down each other's members, and a session inside a pane is resolved as a
worker, which is why the manager sits outside and speaks REST.

15 REST API Reference — the 14 routes were documented nowhere as a set. The
page leads with why that matters: MCP and REST are sibling adapters over one
shared MessageService, so REST behaviour cannot be inferred from the MCP
contract. fleet_ack and fleet_whoami have no route at all.

  It also records a trap found while writing it: GET /members returns its
  rows under a "workers" key (FleetApp.java:322). The route was renamed from
  /workers in CB-557 and the body key was left behind, so a caller reading
  body["members"] sees an empty fleet instead of an error.

16 Security & Trust Boundary — the guard, the role table, the member
credential scrub and token scope were spread across three pages. Collected,
with the limits stated rather than glossed: the effective allow-list is a
union and so a strict superset of what an operator writes under allow:; the
ZDOTDIR scrub is zsh-only; blocking SSH_AUTH_SOCK does not stop a member
reaching a passphrase-free key file; and argv is world-readable, which
bypasses every environment control described on the page.

Home and _Sidebar link all three.
Dai Ha
2026-08-31 10:49:14 +07:00
parent cd3a12ffa2
commit 4912b7acaf
5 changed files with 549 additions and 0 deletions
+111
@@ -0,0 +1,111 @@
# Fleet Manager
`fleet-manager` gives one host a status view of several `fleetd` fleets. A fleet entry has a name, a REST address, and an optional project path. The manager is a separate Python package. It uses only REST requests and does not import or run inside `fleetd` (`README.md:3-12`, `fleet_manager/probe.py:10-13`, `pyproject.toml:1-9`).
This is not multi-host federation. Entries marked `remote` are not probed. The manager reports them as remote instead (`fleet_manager/probe.py:117-121`, `tests/test_probe.py:31-34`).
## Install and run
The package needs Python 3.11 or later. Its runtime dependencies are Python standard-library modules only (`pyproject.toml:3-6`).
For development, create a virtual environment and install the package with its test and lint tools:
```sh
python -m venv .venv && .venv/bin/pip install -e '.[dev]'
```
This command is the documented development setup (`README.md:54-59`).
The package registers the `fleet-status` command (`pyproject.toml:8-9`). It also runs as a Python module:
```sh
fleet-status
python -m fleet_manager.cli
python -m fleet_manager.cli path/to/fleets.json
```
The module command accepts at most the first command-line argument as the registry path. Without an argument, it reads `fleets.json` beside the package source, not the current directory (`fleet_manager/cli.py:54-57`). If that file does not exist, it prints setup help to standard error and exits with status 2 (`fleet_manager/cli.py:58-62`).
## `fleets.json` registry
The registry is JSON with a required top-level `fleets` array. The command reads `reg["fleets"]`, so a missing `fleets` key stops the command with a key lookup error (`fleet_manager/cli.py:49-51`, `fleet_manager/cli.py:63-65`).
Each array item is one fleet. The following example uses only the sample values:
```json
{
"fleets": [
{
"name": "fleetd",
"rest": "http://127.0.0.1:8765",
"project": "/path/to/the/project/this/fleet/works/on"
},
{
"name": "other-host-fleet",
"remote": true,
"project": "/path/on/that/host",
"note": "REST is loopback-only, not reachable from here yet"
}
]
}
```
The sample shape is from `fleets.example.json:1-16`.
| Field | Required | Meaning and missing-field result |
|---|---|---|
| `name` | Yes | The displayed fleet name. A missing value stops probing that entry with a key lookup error (`fleet_manager/probe.py:114-115`). |
| `rest` | Yes for a local fleet | The base REST URL. A missing value stops a local probe with a key lookup error. A remote entry does not read it (`fleet_manager/probe.py:114-115`, `fleet_manager/probe.py:117-124`). |
| `project` | No | An optional project path shown in the status output. When absent, no project line is shown (`fleet_manager/probe.py:114-115`, `fleet_manager/cli.py:32-33`). |
| `remote` | No | When true, the manager does not call the fleet and reports `remote`. When absent or false, it treats the entry as local (`fleet_manager/probe.py:117-123`). |
| `note` | No | An optional message for a remote fleet. If absent, the manager uses `on another host; REST is loopback-only, not probed from here` (`fleet_manager/probe.py:117-121`). |
## `fleet-status`
`fleet-status` is the installed status command. It loads every entry in `fleets`, probes each one, and prints the result (`pyproject.toml:8-9`, `fleet_manager/cli.py:63-66`).
The output has a time-stamped `FLEET STATUS` heading. For each fleet, it shows the name, status, REST URL, optional project and note, a verdict when the fleet is reachable, and each returned member's role, profile, state, live status, branch, and progress result (`fleet_manager/cli.py:27-46`).
## `python -m fleet_manager.cli`
This module command runs the same `main` function as `fleet-status` (`fleet_manager/cli.py:54-70`, `pyproject.toml:8-9`). It uses the package-side default registry with no path, or the first supplied path with one (`fleet_manager/cli.py:54-57`).
## How a probe works
For a local fleet, the manager requests `/healthz` first. If the request fails, it reports `down`, adds an `/healthz unreachable` note, and does not request members (`fleet_manager/probe.py:123-129`, `tests/test_probe.py:37-42`). A successful health response supplies the displayed status and optional herdr version (`fleet_manager/probe.py:130-131`).
It then requests `/members`. It reads the `workers` array when present. Each worker can supply `profile`, `role`, `state`, `liveStatus`, `worktree`, and `branch` (`fleet_manager/probe.py:133-145`).
The manager does not trust a busy `liveStatus` by itself. For members whose live status is `working` or `busy`, it records the newest modification time in the worktree, waits four seconds by default, then records it again (`fleet_manager/probe.py:36-55`, `fleet_manager/probe.py:81`, `fleet_manager/probe.py:136-150`). A later time means `working`. No change means `stalled`, with the newest file age. An unavailable worktree gives `unknown`, not `stalled` (`fleet_manager/probe.py:64-78`, `tests/test_probe.py:86-92`).
A fleet is `WORKING` when any watched member changed files. It is `STALLED` when a watched member did not change files. It is `busy (no worktree progress signal)` when a busy member has no usable worktree. A fleet with no members is `idle (no members)`; otherwise it can be `idle` (`fleet_manager/probe.py:156-166`, `tests/test_probe.py:45-57`).
## Limits of the probe
The manager handles several fleet entries on one host. It does not probe a remote entry, so it is not a cross-host status or federation tool (`README.md:3-7`, `fleet_manager/probe.py:117-121`, `tests/test_probe.py:31-34`).
The probe cannot prove that a busy member is making useful progress. It only compares worktree modification times. A member can be busy without a usable worktree, and the result is `unknown` rather than a progress decision (`fleet_manager/probe.py:64-78`, `fleet_manager/probe.py:143-150`, `tests/test_probe.py:86-92`).
## Two limits that come from `fleetd`, not from this tool
Neither limit is visible in the `fleet-manager` source, because neither belongs to it. Both were
measured against running daemons, and both shape how the manager can be used.
**Two `fleetd` daemons must not share one herdr session.** When they do, each daemon's reaper
treats the other's panes as orphans and tears them down, so the two fleets kill each other's
members. Give each daemon its own herdr session.
**The manager must not run inside a herdr pane.** `fleetd` resolves a caller's role from the
connection, and it identifies its lead by a tab label. A Claude session sitting in a pane that some
*other* daemon does not name as its lead is resolved as a **worker** by that daemon, and every
orchestration call is then refused. A process outside any pane is resolved as the primary, which is
why the manager sits outside and speaks REST.
Together these are the reason the manager is a separate REST client rather than a second daemon or
an in-pane session.
## What this page could not confirm
The `fleet-manager` source alone does not describe either limit above, and it was the only source
read for the rest of this page. Issues [#160](https://git.ltms.dev/fleet/fleetd/issues/160) and
[#156](https://git.ltms.dev/fleet/fleetd/issues/156) carry the measurements behind them.
+231
@@ -0,0 +1,231 @@
# REST API Reference
`fleetd` exposes 14 HTTP routes, built once in `FleetApp.build()`
(`fleetd/src/main/java/dev/ltms/fleet/rest/FleetApp.java:126-160`). No other page collects them in
one place — each route is currently documented only where it happens to matter for some other
topic, spread across seven pages.
## Why this page exists
The MCP tools (their contracts are defined in `FleetMcp.java`, not yet collected on their own
wiki page) and the REST routes on this page are **siblings that both sit on top of the same
service objects** — they are not a wrapper around each other in either direction. `fleetd`'s startup code builds one `MessageService` and one `SessionManager` and hands
the *same instances* to both faces: the MCP server (`Fleetd.java:524`) and the REST app
(`Fleetd.java:607`), from a `MessageService` built once at `Fleetd.java:449-450`. The MCP server's
own file header says this directly: its tools are "thin adapters over the same `MessageService`/
`Rendezvous` the REST routes use — so the two are validated by parity, not by re-implementing
behaviour" (`FleetMcp.java:54-57`); `FleetApp.java:38-41` makes the same claim from the REST side.
This matters for a reader in a concrete way: **you cannot infer REST behaviour from the MCP tool
contract.** The two surfaces read from the same source of truth, but each has its own request
shape, its own field names, and — as the table below shows — its own coverage. Two MCP tools have
no REST route at all, and several REST routes have no MCP tool.
## All routes
Every route below is wired in `FleetApp.build()` (`FleetApp.java:126-160`). The "Role" column is
the `Authz.Action` each handler checks (`Authz.java:44-71`) — "open" means the handler makes no
`allow()` call at all, so it needs no `Authorization` header.
| Method | Path | What it does | Role required | Matching MCP tool |
|---|---|---|---|---|
| GET | `/healthz` | Liveness + herdr reachability | open (no auth check) | none |
| GET | `/metrics` | Prometheus scrape (only when metrics are configured) | any authenticated caller (`METRICS`) | none |
| GET | `/sessions` | Raw herdr workspace list, one row per workspace | any authenticated caller (`READ`) | none |
| GET | `/agents` | Raw herdr agent list (every agent herdr tracks) | any authenticated caller (`READ`) | none |
| GET | `/members` | The fleetd-owned worker roster, joined with live herdr status | any authenticated caller (`READ`) | `fleet_list` (its `members` section only — `fleet_list` also adds peer `leads` and capacity data this route does not have) |
| GET | `/profiles` | Configured worker profiles and the default one | any authenticated caller (`READ`) | `fleet_profiles` |
| POST | `/members` | Spawn a worker | primary only (`SPAWN`) | `fleet_spawn` |
| DELETE | `/members/{paneId}` | Tear a worker down by pane id | primary only (`STOP`) | `fleet_stop` |
| POST | `/sessions/{id}/message` | Deliver a turn to a session, or answer a worker's `fleet_ask` | primary or architect (`SEND`) | `fleet_send` |
| POST | `/sessions/{id}/reply` | A worker's (or peer lead's) structured reply for its own turn | caller must own the session (`REPLY`) | `fleet_reply` |
| GET | `/sessions/{id}/replies` | Drain a session's held-reply inbox | primary only (`DRAIN`) | `fleet_poll` (with `target` set) |
| POST | `/sessions/{id}/ask` | A worker's mid-turn question to the primary | caller must own the session (`ASK`) | `fleet_ask` |
| GET | `/sessions/{id}/status` | A session's live lifecycle status and readiness | any authenticated caller (`READ`) | `fleet_status` |
| GET | `/tasks/{ticket}` | Poll an async (`wait:false`) send by its ticket | any authenticated caller (`READ`) | `fleet_poll` (with `ticket` set) |
**MCP tools with no REST route:** `fleet_ack` (`FleetMcp.java:1137-1146`, calls
`messages.ackReply(target, msgId)` directly — no `FleetApp` handler exists for it) and
`fleet_whoami` (`FleetMcp.java:1233-1246` — identity is resolved per-connection by
`ConnectionIdentity`/`CallerResolver`, so there is nothing for a REST caller to query; each REST
request already carries the resolved `Principal` under the `fleetd.caller` context attribute,
`FleetApp.java:56,138-142`).
## Per-route detail
### `POST /sessions/{id}/message`
Handler: `sendMessage`, `FleetApp.java:431-474`. This is `fleet_send`'s REST face, and it also
carries `fleet_send`'s "answer a worker's `fleet_ask`" mode.
Body fields (`FleetApp.java:441-445`):
| Field | Type | Default | Meaning |
|---|---|---|---|
| `content` | string | required | The message to deliver, or your answer when `turnId` is set. A blank value is a `400 bad_request`. |
| `turnId` | string | none | When set, this call answers a worker's open `fleet_ask` instead of starting a new delegation, and it always blocks. |
| `timeoutMs` | integer | `25000` | Max time to wait for a reply. Clamped to `[1, 120000]` (`FleetApp.java:49-50,454`). |
| `wait` | boolean | `true` | `false` returns a ticket immediately (fire-and-poll); the caller then polls `GET /tasks/{ticket}`. |
Response shape depends on the outcome (`writeReply`, `FleetApp.java:481-511`):
| Outcome | Status | Body |
|---|---|---|
| Worker asked a question | 202 | `{sessionId, status:"question", question, turnId}` |
| Answering a stale/expired turn | 409 | `{sessionId, error:"stale_turn", detail}` |
| Worker replied normally | 200 | `{sessionId, reply, replySource:"reply"}` |
| Worker finished without calling `fleet_reply` | 200 | `{sessionId, reply, replySource:"transcript"}` — a scrape of the worker's terminal tail, not a structured reply |
| Timed out, worker still working | 202 | `{sessionId, status:"working", detail}` |
| Timed out, message still queued | 202 | `{sessionId, status:"queued", detail}` |
| Worker busy on another turn | 202 | `{sessionId, status:"busy", detail}` |
| Worker failed | 202 | `{sessionId, status:"failed", detail}` |
| Backend credential exhausted | 202 | `{sessionId, status:"backend_exhausted", detail}` |
A herdr-level failure during the blocking send (not the `turnId` path) maps through
`herdrError` (`FleetApp.java:650-656`): a `..._not_found` herdr error code becomes
`404 session_not_found`, anything else becomes `502 herdr_error`.
### `POST /sessions/{id}/reply`
Handler: `replyMessage`, `FleetApp.java:554-571`. This is `fleet_reply`'s REST face.
Body: `{"content": "..."}`. Unlike `sendMessage`, there is no blank check on `content` — a missing
or empty value is delivered as `""`. Response: `200 {sessionId, delivered: true}`.
The path's session id is the caller's own identity claim, and this is the one place that claim is
checked over REST — over MCP a worker's identity already comes from its connection, never an
argument, so it could never reply as another worker. Over REST the id in the URL used to be
trusted outright; `Authz.Action.REPLY` (`caller.ownsSession(targetSession)`, `Authz.java:66`) is
what closes that (`FleetApp.java:556-561`).
`messages.reply()` either resolves an open blocking `POST /sessions/{id}/message`, or — when no
send is currently open for that session — queues the reply into the session's inbox
(`FleetApp.java:550-553`). See the warning below about draining that inbox.
### `POST /sessions/{id}/ask`
Handler: `askMessage`, `FleetApp.java:518-548`. This is `fleet_ask`'s REST face: a worker's
mid-turn question to whichever primary has a blocking send open on it.
Body fields (`FleetApp.java:526-528`): `question` (string, required — blank is `400 bad_request`),
`timeoutMs` (default `55000`, clamped to `[1, 115000]`, `FleetApp.java:52-53,537`).
Response:
| Outcome | Status | Body |
|---|---|---|
| Primary answered | 200 | `{sessionId, answered:true, answer}` |
| No primary is waiting on this session | 409 | `{sessionId, error:"no_pending_send", detail}` |
| Primary did not answer in time | 202 | `{sessionId, status:"no_answer", detail}` |
### `GET /sessions/{id}/replies`
Handler: `drainReplies`, `FleetApp.java:578-588`. Response: `200 {sessionId, replies:[{msgId,
content}, ...]}`.
**Warning — this call drains the inbox on its first read.** The javadoc calls the semantics
"at-least-once": reading removes the messages, so a second call returns nothing even if the first
caller never processed the result (`FleetApp.java:573-577`). If you need to keep the response,
write it to a file the first time you call this route — a second call will not give it back to
you.
### `GET /tasks/{ticket}`
Handler: `taskStatus`, `FleetApp.java:622-647`. Polls an async (`wait:false`) send by the ticket
`POST /sessions/{id}/message` returned.
Response: `200 {ticket, phase, reply?, replySource?, detail?, turnId?}` — `reply`/`replySource` are
present once the phase is terminal with a result, `turnId` is present while the phase is
`asking` (the worker opened a `fleet_ask` mid-turn and this ticket now needs an answer,
`FleetApp.java:641-645`). An unknown or expired ticket returns `404 {error:"unknown_ticket",
detail}`.
### `GET /members`
Handler: `listMembers`, `FleetApp.java:309-329`. The fleetd-owned roster, joined against live herdr
status.
**The response key is `workers`, not `members`** (`FleetApp.java:322`). The route was renamed from
`/workers` to `/members` in CB-557, but the body's key was left alone. A caller that reads
`body["members"]` gets nothing and sees an empty fleet rather than an error, so read `body["workers"]`.
`fleet-manager` already does (`fleet_manager/probe.py:134`).
The join is on the **terminal id**, not the pane coordinate: the registry key is a host-unique id
while a pane coordinate is a per-daemon counter, so joining on the pane would mismatch as soon as a
second herdr daemon is configured (`FleetApp.java:314-317`).
Each row is a `SessionManager.rosterView`, carrying at least `sessionId`, `paneId`, `profile`,
`role`, `state`, `worktree`, `branch`, `owner` and `liveStatus`.
The body also carries an optional `wipRefs` object (`count`, `costBytes`) describing the
`refs/wip` snapshot store. It is **absent** until a worktree session has established the repo, so a
fleet that has never snapshotted reports nothing rather than zero (`FleetApp.java:324-327`).
### `POST /members`
Handler: `spawnMember`, `FleetApp.java:346-396`. Spawns a worker.
Accepts `role`, `profile`, `cwd`, `worktree`, `ticket` either as query parameters or as a JSON
body — the body is only parsed when `profile`, `cwd` or `worktree` are missing from the query
string (`FleetApp.java:355-370`). `role` defaults to `dev` (`MemberRole.parse`, values are
`architect`, `dev`, `reviewer`); an unrecognised role is `400 {error:"unknown_role", detail}`.
`worktree` is either `"true"` (requires `ticket`) or a ticket-slug string
(`worktreeRequest`, `FleetApp.java:398-409`).
On success: `201` with the new `MemberSession` view — `{terminalId, paneId, profile, cwd,
ownerTerminal, state, worktree?, branch?}` (`FleetApp.java:672-687`).
Error mapping (`FleetApp.java:379-396`):
| Condition | Status | Body |
|---|---|---|
| Subscription-boundary guard tripped | 403 | `{error:"subscription_boundary", detail}` |
| No profile has spawn capacity right now | 503 | `{error:"no_capacity", detail}` |
| Unknown profile name | 400 | `{error:"unknown_profile", detail}` |
| The peer never became reachable | 502 | `{error:"spawn_timeout", detail}` |
One case is not covered by this table: `worktreeRequest` (`FleetApp.java:398-409`) is called
**before** the `try` block that catches `IllegalArgumentException`, so `worktree=true` with no
`ticket` throws past every one of these handlers rather than becoming a `400`. This is a gap
found while writing this page — it has not been reported as a bug, and no other caller of this
route has been checked for the same pattern.
### `DELETE /members/{paneId}`
Handler: `stopMember`, `FleetApp.java:416-423`. No body. Calls `sessions.release(paneId)` and
always answers `204`, with no path that surfaces an unknown `paneId` back to the caller. Whether
`sessions.release` itself does anything for an id it does not recognise was not checked for this
page.
### `GET /healthz`
Handler: `healthz`, `FleetApp.java:229-266`. The only route with no `allow()` call at all — it
takes no `Authorization` header and answers any caller that can reach the port.
It pings the lead herdr daemon and, when a member daemon is separately configured
(`memberHerdrSocket`, `FleetApp.java:210-227`), pings that one too:
| Configuration | Success body | Failure |
|---|---|---|
| One daemon (the common case) | `200 {status:"ok", herdr:{version, protocol}}` | `503 {status:"degraded", herdr:"unreachable", detail}` |
| Two daemons configured | adds `"member":{version, protocol}`, and `"protocolMismatch": true` when the two `protocol` numbers differ | a member-daemon failure is its own `503 {status:"degraded", herdr:"member unreachable", detail}`, reported separately from a lead-daemon failure so it is never masked by a healthy lead |
**Warning — a green `/healthz` does not mean spawning works.** The socket can connect while the
herdr wire protocol on the other end has moved, so `ping` succeeds and every spawn still fails
(`FleetApp.java:218-227`). Check the `protocol` number in the response, and — where a member
daemon is configured — check the `member` key and the `protocolMismatch` flag specifically, since
every spawn goes through the member daemon, not the lead one.
## What the REST surface is for
Every capability behind an MCP tool is also reachable here, so the system can be driven and
tested with plain HTTP — no Claude session and no MCP client in the loop. This is what makes the
routes on this page an acceptance-test surface in their own right, not only a debugging aid: a
script can spawn a worker, send it a turn, and drain its replies using nothing but `curl` and the
field names on this page.
This page does not cover the rendezvous flows themselves (the blocking-send / `fleet_ask` /
detached-delivery / turn-done sequences) — those are diagrammed in `docs/MCP-Contract.md` §6. That
section's diagrams describe the flow shapes correctly, but its field names (`turn_id`,
`block=false`, `outcome`) predate the code and do not match the request and response fields
documented on this page — use the names given here, not the ones in that document.
+201
@@ -0,0 +1,201 @@
# 16. Security and Trust Boundary
> **Scope.** This page pulls the security model into one place. The model is real, but the code
> that carries it is spread across three areas — the guard package, the auth package, and the
> member package — so no single page showed it whole before this one. Every claim below was
> checked against the source at `main` `a814d1e` (2026-08-31). Where this page and the code ever
> disagree again, trust the code, not this page.
`fleetd` draws four separate lines of defense: a **subscription boundary** that stops a worker
from billing the operator's paid Claude plan, an **identity and authorization** model that stops
one member from acting as another, an **environment control** on what a spawned member inherits
from the operator's own shell, and a **token scope** limit on what a member can do to the git
forge. The sections below cover each in turn, using the class that enforces it.
```mermaid
flowchart TB
op["Operator's shell<br/>own secret store"]
subgraph primary_side["Primary — on subscription"]
primary["Primary (lead)<br/>ANTHROPIC_BASE_URL must be absent"]
end
subgraph member_side["Member — off subscription"]
member["Worker / architect pane<br/>login shell"]
claude["claude / opencode process"]
end
guard["guard.SubscriptionGuard<br/>assertWorker / assertPrimaryClean"]
authz["auth.Authz + mcp.ConnectionIdentity<br/>role table, identity from the connection"]
scrub["member.MemberEnvAllowList<br/>+ EnvAllowListScrub (ZDOTDIR)"]
token["Repo-scoped forge token<br/>write:repository only"]
op -- "re-sources on login" --> member
guard -- "checked before spawn" --> member
guard -- "checked at daemon start" --> primary
authz -- "gates every fleet_* call" --> primary
authz -- "gates every fleet_* call" --> member
scrub -- "blanks names after the shell chain runs" --> member
member --> claude
token -- "injected into the worker env" --> claude
classDef boundary fill:#6b46c1,stroke:#3d2a75,color:#ffffff;
class guard,authz,scrub,token boundary
```
*Figure: the four controls (purple) that separate the operator's subscription and secrets from a
spawned member. The environment controls stop at the process environment — the argv gap in the
last section of this page is not one of the boxes above, because nothing here defends it.*
## 1. The subscription boundary
`SubscriptionGuard` (`fleetd/src/main/java/dev/ltms/fleet/guard/SubscriptionGuard.java`) enforces
one rule in two directions: a worker must run off the operator's paid subscription, and the
primary must never run on a base URL at all.
- **`assertWorker(baseUrl)`** (`SubscriptionGuard.java:32-46`) refuses to spawn a worker whose
`ANTHROPIC_BASE_URL` is blank, not a valid URL, or resolves to a host outside a configured
allowlist. The allowlist is `FleetConfig.Guard.offSubscriptionHosts`
(`FleetConfig.java:1156-1159`) and it **defaults to an empty list**. Because the list is a config
value that differs by host, this page names no hostname.
- **`assertPrimaryClean(env)`** (`SubscriptionGuard.java:49-55`) refuses to let the daemon start if
its own environment carries `ANTHROPIC_BASE_URL`. `Fleetd.java:131-132` calls this against
`System.getenv()` at startup, before the daemon does anything else. If the primary's shell were
ever pointed at a base URL, the daemon would not come up.
- **`ClaudeCodeLauncher.buildLaunch`** (`fleetd/src/main/java/dev/ltms/fleet/member/ClaudeCodeLauncher.java:191-213`)
calls `assertWorker` right before every non-subscription spawn.
**The opt-out: `subscription: true`.** A profile can deliberately choose to run a worker on the
operator's own subscription instead of an off-subscription endpoint (CB-539). This is meant for a
profile with no off-subscription endpoint to point at. It is enforced at two separate points, so a
worker cannot slip past the guard on the subscription path:
- **At spawn time**, `ClaudeCodeLauncher.buildLaunch` (`ClaudeCodeLauncher.java:196-213, 227-233`)
refuses a profile that sets both `subscription: true` and a `baseUrl` — the two say opposite
things, so the launcher throws rather than picking one. On the subscription path it also strips
`ANTHROPIC_BASE_URL` and `ANTHROPIC_AUTH_TOKEN` from the worker's environment as a
belt-and-braces step, even though `subscription: true` should never carry them.
- **At config-load time**, `FleetConfig.validateSubscriptionProfiles()` (`FleetConfig.java:2058-2075`)
refuses to start the daemon at all if any `subscription: true` profile's `env:` block names
`ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. The reason this must be a fatal, loud refusal
and not a silent strip: on the subscription path, `SubscriptionGuard` is skipped entirely for
that profile, so nothing else would catch a repoint (`FleetConfig.java:2044-2052`).
## 2. Identity and authorization
A member's identity comes from **the connection it calls on**, never from an argument in the
request. `ConnectionIdentity.resolve` (`fleetd/src/main/java/dev/ltms/fleet/mcp/ConnectionIdentity.java:42-48`)
takes only the remote address and port of the MCP connection. It maps the connecting process's PID
to a herdr pane through `PaneLocator`, and that pane maps to a worker's `terminal_id`. Both steps
use sources the caller cannot influence: the OS reports the real PID for a loopback connection, and
herdr owns the PID-to-pane mapping. A caller that maps to no worker pane — an off-host client, or
the primary itself — resolves to `null` (`ConnectionIdentity.java:9-14`).
The role table lives in `Authz.permits` (`fleetd/src/main/java/dev/ltms/fleet/auth/Authz.java:44-72`):
| Action | Who may perform it | Why (from the code's own comments) |
|---|---|---|
| `SPAWN`, `STOP`, `DRAIN` | primary only | fleet lifecycle is the primary's alone; an architect deliberately does not get these, even though it coordinates workers, so it cannot tear a fleet down |
| `SEND` | primary or architect | an architect delegates to workers, which is the point of the role, but still has no lifecycle rights; a worker sending would be it escalating into the orchestrator role |
| `REPLY`, `ASK` | the caller must own the target session | the load-bearing rule: a caller acts only as the pane it occupies |
| `READ`, `METRICS` | primary, worker, or architect | observation carries no secrets, so every authenticated role may read |
The `REPLY`/`ASK` row is what stops one member from forging a reply for another member's
rendezvous. Because identity is read from the connection and not from a session id typed into the
request, a member cannot claim a session id that is not its own pane's. An unauthenticated caller
is refused for every action (`Authz.java:44-47`).
## 3. What a member inherits, and what is blocked
This is the part of the model that most needs care, because two different controls exist and only
one of them actually works once a login shell has run.
**Why a spawn-time overlay is not enough.** A herdr pane runs a **login shell**, and that shell
re-sources the operator's own secret store. A control applied at pane-creation time — an
environment overlay handed to herdr before the shell starts — is therefore overwritten the moment
the operator's own `~/.zshrc` and secret files run (`FleetConfig.java:1182-1183`,
`EnvAllowListScrub.java:17-25`). This is the `deny-by-default` policy
(`FleetConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT`, `FleetConfig.java:1223`): it shadows each
known-but-not-allowed name in the pane-creation overlay, but a login shell that re-exports the same
name defeats it.
**What actually holds is the `allow-list` policy** (`FleetConfig.MemberCredentials.POLICY_ALLOW_LIST`,
`FleetConfig.java:1235`), implemented in `EnvAllowListScrub`
(`fleetd/src/main/java/dev/ltms/fleet/member/EnvAllowListScrub.java`). It generates a per-spawn
`ZDOTDIR` whose startup files run the scrub **after** the whole operator chain has sourced,
blanking every exported variable that is not on a derived allow-list. Because the scrub runs last,
nothing sourced later can undo it (`EnvAllowListScrub.java:20-25`). The scrub is sourced from both
the generated `.zshrc` and `.zlogin`, because herdr opens a login zsh on macOS (where `.zlogin`
runs) and a plain interactive zsh on Linux (where it does not) — a scrub placed in only one of
those files would silently do nothing on the other platform (`EnvAllowListScrub.java:27-34`). This
control is **zsh-only**: it depends on `$ZDOTDIR` being a zsh mechanism, so a member whose pane
runs a different shell is not covered by it.
**The effective allow-list is wider than what an operator writes.** `MemberEnvAllowList.derive`
(`fleetd/src/main/java/dev/ltms/fleet/member/MemberEnvAllowList.java:108-128`) builds the kept-name
set as a **union** of:
- `INFRASTRUCTURE_PASSTHROUGH` — locations and shell settings such as `PATH`, `HOME`, `ZDOTDIR`,
and the `XDG_*` roots (`MemberEnvAllowList.java:78-84`), never credentials;
- every configured profile's `gitTokenEnv`, `gitHostEnv`, `tokenEnv`, and every key of that
profile's `env:` map (`MemberEnvAllowList.java:110-118`) — for **every** profile, not only the
one being spawned, so adding a profile can only widen the list, never narrow another spawn's
scrub;
- the operator's own `memberCredentials.allow:` list (`MemberEnvAllowList.java:120-126`).
So the list a spawn actually keeps is a strict superset of what an operator types under `allow:`.
An operator who reads only their own `allow:` entry will underestimate what a member can keep.
**`SSH_AUTH_SOCK` is excluded from that union** (`MemberEnvAllowList.java:122`). It is the
operator's ssh-agent socket, and a member holding it can sign with the operator's own keys, so it
is governed only by the separate `memberCredentials.sshAuthSock` setting
(`FleetConfig.java:1208-1216`), never by appearing in `allow:`.
**The limit of an environment control.** Blocking the agent socket in the environment does not
stop a member reaching a passphrase-free private key **file** sitting on disk outside `~/.ssh`. An
environment control covers what a process can read from its environment variables; it says
nothing about the filesystem. A control at one layer does not close a gap at another.
## 4. Tokens
A member is given a repo-scoped forge token with **`write:repository`** scope only — enough to
create a branch and open a pull request — never the operator's own admin-scoped
`GITEA_ACCESS_TOKEN` (`fleetd/docs/Worker-Git-Workflow.md:139-144`). `GitWorktrees`'s credential
helper reads this scoped token from `WORKER_GITEA_TOKEN`
(`fleetd/src/main/java/dev/ltms/fleet/session/GitWorktrees.java:70-72`) and warns, in its own
comment, that a silent fallback to the primary's admin token would turn a loud crash into a quiet
privilege leak (`GitWorktrees.java:52-58`). `ClaudeCodeLauncher.buildLaunch` injects the scoped
token into the worker's environment as part of every spawn (`ClaudeCodeLauncher.java:240`). A lead
gets no git token at all — it reviews and merges through the operator's own credentials, and
`LeadLauncher.leadEnv` never grants it the scoped token a member gets
(`fleetd/src/main/java/dev/ltms/fleet/lead/LeadLauncher.java:287-288`).
The reason for the narrow scope is direct: a member opens its own pull request and must never be
able to merge it. A leaked or misused member token can open pull requests. It cannot merge, delete
a branch it does not own, or touch the org. Merge authority stays with the operator or the primary.
## The one honest limitation: argv is world-readable
A process's command-line arguments (`argv`) are visible to any other process on the same host —
`ps -axww` prints them for every user session. If a program passes a credential as
`NAME=value` on its command line, that value is exposed the same way a value printed to a shared
log would be. Nothing in `fleetd` checks for this, and nothing in this page's other three sections
defends against it.
`fleetd`, herdr, and `claude` are clean on this specific point: `fleetd` passes a member's
environment to herdr as a JSON field in the spawn request, not as command-line arguments —
`WorkspaceControl.createTab` and `WorkspaceControl.splitPane` both carry `env` as a separate
`params` entry alongside `cwd`, never appended to argv
(`fleetd/src/main/java/dev/ltms/fleet/herdr/WorkspaceControl.java:84-93, 100-109`).
But this guarantee is narrow. It covers `fleetd`, herdr, and `claude` — nothing else. **A
different program started on the same host can still leak a credential through its own argv**, and
every control in this page — the subscription guard, the identity check, the environment scrub,
the token scope — is bypassed the moment that happens, because none of them touch argv. This is a
gap in the host, not in `fleetd`'s own code, and it is recorded here rather than left undocumented,
because a security page that hides its own limits is worse than one with none.
---
This page does not cover herdr's own process isolation, the LavinMQ broker's per-vhost isolation
between two fleets, or session lifecycle and reaping. See [1. Architecture](1-Architecture) and
[9. Implementation Architecture](9-Implementation) for those.
+3
@@ -94,6 +94,9 @@ Read in order (the sidebar mirrors this):
11. **[Features](11-Features)** — what it can do, the knob that turns it on, why it exists, the gotcha
12. **[Claude → OpenCode](12-Claude-to-OpenCode)** — porting a workspace to a second host
13. **[User Guide](13-User-Guide)** — 🟢 **the operator page.** Install, configure, run, delegate, and the traps.
14. **[Fleet Manager](14-Fleet-Manager)** — the separate tool that shows several fleets on one host
15. **[REST API Reference](15-REST-API-Reference)** — every route, the role it needs, and its real body
16. **[Security & Trust Boundary](16-Security-and-Trust-Boundary)** — the subscription guard, the role table, and what a member's pane inherits
## Status
+3
@@ -17,6 +17,9 @@
11. [Features](11-Features) — what it can do · the knob that turns it on · why · the gotcha
12. [Claude → OpenCode](12-Claude-to-OpenCode) — porting a workspace to a second host
13. **[User Guide](13-User-Guide)** — 🟢 install · configure · run · delegate · the traps
14. [Fleet Manager](14-Fleet-Manager) — many fleets on one host, over REST
15. [REST API Reference](15-REST-API-Reference) — all 14 routes, roles, and bodies
16. [Security & Trust Boundary](16-Security-and-Trust-Boundary) — the guard · authz · what a member inherits
---
🟢 herdr-centric `fleetd` · AgentAPI = research, never built