#168: rebuild chapter 6 against the source

The page taught a role-addressed send that does not exist. fleet_send takes
sessionId, content, timeoutMs, wait, turnId and coordId — never role or prompt
(FleetMcp.java:1096-1108).

Other claims replaced:

- It said every worker is a Claude Code process and the only other kind is a
  "remote local LLM". opencode is a first-class launcher kind. The page now says
  what really differs: claude-code mounts MCP through an inline --mcp-config,
  opencode through a generated opencode.json plus OPENCODE_CONFIG, and opencode
  reads provider credentials from its own auth.json.
- It described "fleetd's concurrency policy" and role routing. The real controls
  are per-profile maxLoad and weight, plus a placement policy. weighted uses
  smooth weighted round-robin and follows weight ratios; it is not cheapest-first
  (WeightedRoundRobinPolicy.java:7-48).
- A named profile bypasses placement but not maxLoad, and a full profile fails
  the spawn rather than silently moving it elsewhere.
- It named a backend host. The guard's allowlist defaults to an empty list, so
  the page names none.

Also corrects the two member roles against Authz: a worker may reply, ask and
read; an architect may also send; only a primary may spawn, stop or drain.
Dai Ha
2026-08-31 10:36:10 +07:00
parent e9c4f96b6f
commit 426f5a2ff4
+130 -132
@@ -1,165 +1,163 @@
# 6. Team
The [Message Server](2-Message-Server) (`fleetd`) delivers **one turn into one worker**. A
**team** is the layer above it: a **Claude team-lead** that fans a job out across a **mixed
fleet** of workers — some on Claude, some on the remote local LLM — and reduces their replies.
Same `fleetd` delivery, same subscription boundary; this page is only about **orchestration**
— who the workers are, how the lead picks one, and how it runs many at once.
`fleetd` delivers one turn to one session. This page describes the layer above
that delivery: how a lead splits work, starts members, and collects their work.
The delivery rules belong to [Message Server](2-Message-Server).
> Delivery mechanics (blocking `fleet_send` MCP call, status-gated reply) live in
> [Message Server](2-Message-Server). Transport rationale is in [Approaches](3-Approaches).
> This page assumes both.
## The team shape
## The team
A primary lead orchestrates the work. It splits a job into units, starts members
for the units, sends each member its own brief, and decides the final result.
Only a primary can start, stop, or drain members. An architect may send work, but
a worker may not (`Authz.java:49-70`).
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
worker; an **MCP client of `fleetd`** (it mounts the bridge like everyone else). It plans,
routes, dispatches via `fleet_send`, and integrates — and never sets `ANTHROPIC_BASE_URL`.
- **Workers** — a herd of `claude` panes in herdr, each an addressable `fleetd` session
with its **own model/env**, each also **mounting the bridge MCP** (unified setup — they
reply via `fleet_reply`):
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
- **Local workers** (`ANTHROPIC_BASE_URL=https://ollama.ltms.dev`) — bulk, cheap, or
embarrassingly parallel subtasks.
Leads may coordinate with peer leads. They do not assign tasks to each other.
A lead addresses a member by its `sessionId`, not by a role
(`FleetMcp.java:1096-1108`). `fleet_list` reports peer leads and members, including
each member's `sessionId`, role, and profile (`FleetMcp.java:1185-1202`).
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
MCP) — only its model differs. The **same one MCP-mount line** wires lead and every worker
alike. Scale each kind horizontally by adding panes.
### Topology
The diagram shows the normal direction of work. A member reply goes back through
`fleetd` after the member finishes its turn.
```mermaid
flowchart TB
LEAD["lead — Opus<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["fleetd — standalone daemon"]
SRV["SERVER face<br/>MCP · REST/SSE · policy brain"]
CLI["CLIENT face<br/>herdr socket"]
SRV --> CLI
end
HERDR["herdr<br/>panes · agent-status"]
WC1["w-claude-1<br/>Sonnet · CLEAN"]
WC2["w-claude-2<br/>Sonnet · CLEAN"]
WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
ANT["api.anthropic.com<br/>(Pro/Max)"]
OLL["ollama.ltms.dev<br/>(local model)"]
Lead["Lead"]
Fleet["fleetd"]
Peer["Peer lead"]
MemberA["Member A"]
MemberB["Member B"]
LEAD -->|"MCP fleet_send (target role)"| SRV
CLI -->|"Unix socket · send_text · events.subscribe"| HERDR
HERDR --> WC1 & WC2 & WL1 & WL2
WC1 -.->|"MCP fleet_reply"| SRV
WL1 -.->|"MCP fleet_reply"| SRV
WC1 -->|"inference"| ANT
WC2 -->|"inference"| ANT
WL1 -->|"inference"| OLL
WL2 -->|"inference"| OLL
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef local fill:#6b46c1,stroke:#44337a,color:#ffffff;
class LEAD,WC1,WC2 ext
class SRV,CLI,HERDR core
class WL1,WL2 local
Lead -->|"spawn and send"| Fleet
Fleet -->|"deliver work"| MemberA
Fleet -->|"deliver work"| MemberB
MemberA -->|"reply"| Fleet
MemberB -->|"reply"| Fleet
Fleet -->|"return or hold reply"| Lead
Lead <-->|"coordinate"| Peer
```
### Roles & routing
*The lead delegates downward to members. Peer leads only coordinate.*
| Role | Env | Model | Route here when… |
|---|---|---|---|
| `lead` | clean | Opus (sub) | always — it does the routing |
| `w-claude-*` | clean | Sonnet (sub) | task needs Claude-grade reasoning / careful edits |
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
## Addressing and roles
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
selection is **policy in the lead**, not a `fleetd` concern — `fleetd` just delivers to
the session the lead names.
`fleet_send` has one required field: `content`. Its optional addressing fields
are `sessionId`, `turnId`, and `coordId` (`FleetMcp.java:1096-1108`).
## Subscription boundary in a team
| Use | Field | Meaning |
|---|---|---|
| Send work to a member | `sessionId` | The member's herdr terminal ID. |
| Answer a member's `fleet_ask` | `turnId` | Routes the answer into that member's same turn. |
| Coordinate with a peer lead on another daemon | `coordId` | Routes a coordination message to that lead's mailbox. |
Unchanged from [Architecture](1-Architecture), and it scales with the fleet: **only local-worker panes**
launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on the
subscription. `fleetd` enforces which panes may carry the off-subscription env, so adding
workers never widens the boundary.
`fleet_send` also accepts `wait` and `timeoutMs`. `wait` defaults to `true`.
With `wait: false`, the call returns a ticket. The lead checks that ticket with
`fleet_poll` (`FleetMcp.java:1088-1108`, `FleetMcp.java:1125-1134`).
## Parallel fan-out (map / reduce)
Do not use `role` or `prompt` with `fleet_send`. They are not fields in its
schema (`FleetMcp.java:1096-1108`).
The lead's advantage over a single bridge is **concurrency**: independent subtasks go to
different workers at once, then results are gathered.
There are two member authorization roles: `worker` and `architect`. A worker can
reply, ask, and read. An architect can also send work. Only a primary can spawn,
stop, or drain members (`Authz.java:44-71`).
```mermaid
sequenceDiagram
participant L as "lead (Opus)"
participant B as "fleetd"
participant WC as "w-claude-1"
participant WL as "w-local-1"
The `fleet_spawn` field named `role` is separate from that authorization table.
It selects a member contract: `dev`, `reviewer`, or `architect`. `profile` selects
the backend. The two fields are independent, so a reviewer can use the same
profile as the developer it reviews (`FleetMcp.java:1148-1173`).
Note over L: "split job → subtask A (reasoning), subtask B (bulk)"
par A → Claude worker
L->>B: "fleet_send {role: w-claude, prompt: A}"
B->>WC: "send_text into running pane"
WC-->>B: "fleet_reply (or idle edge)"
B-->>L: "tool result reply A"
and B → local worker
L->>B: "fleet_send {role: w-local, prompt: B}"
B->>WL: "send_text into running pane"
WL-->>B: "fleet_reply (or idle edge)"
B-->>L: "tool result reply B"
end
Note over L: "reduce → integrate A + B into final answer"
## Claude Code and OpenCode members
A profile chooses one launcher kind. The supported kinds are `claude-code`, which
is the default, and `opencode` (`FleetConfig.java:265-270`). Both use the shared
member transport for pane placement, readiness, teardown, listing, and working
directory handling (`OpenCodeLauncher.java:30-35`).
The launchers differ in important ways:
| Kind | Launch and configuration | MCP servers |
|---|---|---|
| `claude-code` | Uses command-line flags. The model and other profile arguments come from its launch arguments. | When configured, its inline `--mcp-config` includes the `fleet` server. It also includes `intellij` when the IDE MCP URL is configured. |
| `opencode` | Writes an ephemeral `opencode.json`, points `OPENCODE_CONFIG` to it, and passes the provider/model selector with `-m`. It reads provider credentials from its global `auth.json`. | When configured, its file config includes the remote `fleet` server. It also includes the remote `intellij` server when the IDE MCP URL is configured. |
Claude Code mounts MCP through `--mcp-config` (`ClaudeCodeLauncher.java:283-356`).
OpenCode mounts MCP through its generated file and `OPENCODE_CONFIG`
(`OpenCodeLauncher.java:39-49`, `OpenCodeLauncher.java:340-367`). The `fleet` and
`intellij` servers are conditional on profile configuration in both launchers.
The subscription guard applies to Claude Code profiles. OpenCode has no
`ANTHROPIC_BASE_URL` injection and no `SubscriptionGuard` (`OpenCodeLauncher.java:37-47`).
The guard's allowed off-subscription hosts come from configuration. The default
list is empty, so this page names no backend host (`FleetConfig.java:1149-1163`).
## Start members, then send work
For independent units, the lead starts all needed members before it sends briefs.
It then sends all briefs without waiting for each one. This avoids turning
independent work into serial work.
```text
1. fleet_spawn for each unit
2. fleet_send with each returned sessionId and wait: false
3. fleet_poll each returned ticket
4. fleet_ack each reply after the lead processes it
5. review the change, then merge if the lead accepts it
```
- **Map:** the lead issues N concurrent blocking `fleet_send` tool calls (one per subtask →
its chosen worker). Each call blocks only *that* request; `fleetd` holds it open until the
worker's turn completes (status-gated) and returns the reply as the tool result.
- **Reduce:** the lead collects the N replies and integrates. A slow local worker never
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
- **Detached / long jobs** use `fleetd`'s async path instead of a held request — the result
is delivered when ready by `fleetd` injecting the lead's idle pane (Mode 2 in
[Architecture](1-Architecture)). The lead talks only to `fleetd`, never a broker, and never
busy-polls across turns.
Example calls use `sessionId` and `content`, not a role address:
Fan-out is bounded by the herd size (pane count) and `fleetd`'s concurrency policy, not by
the lead.
```text
fleet_spawn {
role: "dev",
profile: "chosen-profile",
worktree: true,
ticket: "task-123"
}
## Knowing the roster
fleet_send {
sessionId: "returned-session-id",
content: "Implement the assigned unit.",
wait: false
}
The lead discovers its team from `fleetd` (session list / roles) rather than hard-coding
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
```markdown
## Your team (via fleetd)
You are the team-lead. Delegate through the bridge MCP tools — never launch workers yourself.
Roster: call fleet_list for current sessions/roles.
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
Route each subtask by the rubric in the Team page. For independent subtasks, DISPATCH ALL
of them (concurrent blocking sends), THEN gather — never serialize independent work.
Integrate the reply envelopes; you own the final answer.
fleet_poll { ticket: "returned-ticket" }
fleet_ack { target: "returned-session-id", msgId: "returned-message-id" }
```
`fleet_send` **is** the one tool call — no HTTP to hand-roll. Optionally wrap it in a Claude
Code skill (`/delegate <role> "<task>"`) for ergonomics.
`fleet_spawn` returns a `sessionId` for `fleet_send` and a `paneId` for
`fleet_stop` (`FleetMcp.java:1156-1173`). `fleet_ack` needs the member session ID
and the reply message ID (`FleetMcp.java:1137-1145`).
## What this layer does NOT change
## Profile placement and capacity
- **Delivery** is still `fleetd` → herdr `pane.send_text` + status events ([Message Server](2-Message-Server)).
- **Completion timing** is still the worker status event; **reply content** rides the worker's
`fleet_reply` (or a `Stop`-hook envelope for a herdr-only worker).
- **Single-host** still applies: herdr's socket is local, so the whole herd lives on the
`fleetd` host. The lead may be remote — it only needs to reach `fleetd`'s MCP endpoint.
The lead may name a `profile` in `fleet_spawn`. That bypasses automatic placement,
but `maxLoad` still applies. If that named profile is at its cap, the spawn fails;
it does not silently move to another profile (`CompositePeerLauncher.java:260-273`,
`CompositePeerLauncher.java:326-397`).
## Open questions
When a spawn omits `profile`, `fleetd` builds candidates from the requested
contract's profile pool. An absent or empty pool uses all configured profiles
(`CompositePeerLauncher.java:276-286`, `CompositePeerLauncher.java:401-431`).
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `fleetd` role-router
(label-based). Start with the former; promote to the latter if routing logic grows.
- **Backpressure:** per-role concurrency caps in `fleetd` so a fan-out can't exhaust the
local gateway.
- **Result schema:** whether `fleet_reply` payloads should carry structured metadata (worker,
model, tokens) to help the lead's reduce step.
Each profile can set:
## Status
- `maxLoad`, the maximum number of live members. An absent value means unlimited.
A value of `0` allows no live members (`FleetConfig.java:251-264`).
- `weight`, the relative value for automatic placement. A value of zero or less
excludes the profile from automatic placement, but an explicit profile can still
select it if it has capacity (`FleetConfig.java:242-260`).
🟡 Design (2026-07-11). Orchestration layer over the selected `fleetd` server; inherits
herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see [Message Server](2-Message-Server).
Available candidates exclude weight-disabled, unreachable, quarantined, and
at-cap profiles (`PlacementPolicyUtil.java:14-36`). The configured placement policy
then chooses from those candidates. The default is `fixed`; `round-robin` and
`weighted` are also supported (`PlacementPolicies.java:12-40`).
The `weighted` policy uses smooth weighted round-robin. It adds weights, picks
the highest score, then subtracts the total weight from that winner. Over time,
the selection follows weight ratios. It is not a cheapest-first policy
(`WeightedRoundRobinPolicy.java:7-48`).
## Limits of this page
This page does not describe the exact worker turn lifecycle or the transport
between `fleetd` and herdr. See [Message Server](2-Message-Server) for delivery
details.