# Team — lead orchestrating a mixed Claude + local-LLM fleet
The message server (`fleetd`) delivers **one turn into one worker**. A **team** is the
layer above it: a **Claude team-lead** that fans a job out across a **mixed fleet** of
workers — some on Claude, some on the remote local LLM — and reduces their replies. Same
`fleetd` delivery, same subscription boundary; this doc is only about **orchestration** —
who the workers are, how the lead picks one, and how it runs many at once.
> Delivery mechanics (blocking `POST /message`, status-gated reply envelope) live in the
> Message-Server design. Transport rationale is in Approaches. This doc assumes both.
## The team
- **Team-lead** — the primary **Opus** (Claude Code, env **CLEAN**, on Pro/Max). Not a
worker; a **thin client of `fleetd`**. It plans, routes, dispatches, and integrates, and
never sets `ANTHROPIC_BASE_URL`.
- **Workers** — a herd of `claude` panes in herdr, each an addressable `fleetd` session
with its **own model/env**:
- **Claude workers** (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
- **Local workers** (`ANTHROPIC_BASE_URL=https://llm.ltms.dev/anthropic`) — bulk, cheap, or
embarrassingly parallel subtasks.
Every worker is still a *real Claude Code process* (inherits `CLAUDE.md`, hooks, skills,
MCP) — only its model differs. Scale each kind horizontally by adding panes.
### Topology
```mermaid
flowchart TB
LEAD["lead — Opus
(Claude Code, env CLEAN)"]
BD["fleetd
message server + router"]
HERDR["herdr
panes · agent-status"]
WC1["w-claude-1
Sonnet · CLEAN"]
WC2["w-claude-2
Sonnet · CLEAN"]
WL1["w-local-1
ANTHROPIC_BASE_URL set"]
WL2["w-local-2
ANTHROPIC_BASE_URL set"]
ANT["api.anthropic.com
(Pro/Max)"]
OLL["llm.ltms.dev
(gateway to the local model)"]
LEAD -->|"blocking POST /message (target role)"| BD
BD -->|"Unix socket · send_text · events.subscribe"| HERDR
HERDR --> WC1 & WC2 & WL1 & WL2
WC1 --> ANT
WC2 --> ANT
WL1 --> OLL
WL2 --> OLL
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
classDef local fill:#6b46c1,stroke:#44337a,color:#ffffff;
class LEAD,WC1,WC2 ext
class BD,HERDR core
class WL1,WL2 local
```
### Roles & routing
| Role | Env | Model | Route here when… |
|---|---|---|---|
| `lead` | clean | Opus (sub) | always — it does the routing |
| `w-claude-*` | clean | Sonnet (sub) | task needs Claude-grade reasoning / careful edits |
| `w-local-*` | `ANTHROPIC_BASE_URL` set | local LLM | task is bulk / cheap / embarrassingly parallel |
The lead applies this rubric itself, guided by its `CLAUDE.md` team charter (below). Worker
selection is **policy in the lead**, not a `fleetd` concern — `fleetd` just delivers to
the session the lead names.
## Subscription boundary in a team
Unchanged from the base architecture, and it scales with the fleet: **only local-worker
panes** launch with `ANTHROPIC_BASE_URL`. The lead and every Claude worker stay env-clean on
the subscription. `fleetd` enforces which panes may carry the off-subscription env, so
adding workers never widens the boundary.
## Parallel fan-out (map / reduce)
The lead's advantage over a single bridge is **concurrency**: independent subtasks go to
different workers at once, then results are gathered.
```mermaid
sequenceDiagram
participant L as lead (Opus)
participant B as fleetd
participant WC as w-claude-1
participant WL as w-local-1
Note over L: split job → subtask A (reasoning), subtask B (bulk)
par A → Claude worker
L->>B: POST /message {role: w-claude, prompt: A}
B->>WC: send_text into running pane
WC-->>B: status working → idle + Stop-hook envelope
B-->>L: 200 reply A
and B → local worker
L->>B: POST /message {role: w-local, prompt: B}
B->>WL: send_text into running pane
WL-->>B: status working → idle + Stop-hook envelope
B-->>L: 200 reply B
end
Note over L: reduce → integrate A + B into final answer
```
- **Map:** the lead issues N concurrent blocking `POST /message` calls (one per subtask → its
chosen worker). Each call blocks only *that* request; `fleetd` holds it open until the
worker's turn completes (status-gated) and returns the reply envelope.
- **Reduce:** the lead collects the N envelopes and integrates. A slow local worker never
blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
- **Detached / long jobs** use the async broker path instead of a held request (Channel 2 in
the base architecture), so the lead never busy-polls across turns.
Fan-out is bounded by the herd size (pane count) and `fleetd`'s concurrency policy, not by
the lead.
## Knowing the roster
The lead discovers its team from `fleetd` (session list / roles) rather than hard-coding
pane ids, so workers can be added or restarted without editing the lead. A minimal charter
in the lead's `CLAUDE.md` turns Opus into the orchestrator:
```markdown
## Your team (via fleetd)
You are the team-lead. Delegate through the fleetd client — never launch workers yourself.
Roster: ask fleetd for current sessions/roles.
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
- w-local-* — remote local LLM. Bulk, cheap, or parallelizable subtasks.
Route each subtask by the rubric in the Team design. For independent subtasks, DISPATCH ALL
of them (concurrent blocking sends), THEN gather — never serialize independent work.
Integrate the reply envelopes; you own the final answer.
```
Wrap the send as a Claude Code skill (`/delegate ""`) so the lead calls one
tool instead of hand-rolling the HTTP request.
## What this layer does NOT change
- **Delivery** is still `fleetd` → herdr `pane.send_text` + status events (Message-Server).
- **Completion timing** is still the worker status event; **reply content** still rides the
worker `Stop`-hook envelope.
- **Single-host** still applies: herdr's socket is local, so the whole herd lives on the
`fleetd` host. The lead may be remote — it only needs HTTP to `fleetd`.
## Open questions
- **Routing intelligence:** rubric-in-`CLAUDE.md` (lead decides) vs. a `fleetd` role-router
(label-based). Start with the former; promote to the latter if routing logic grows.
- **Backpressure:** per-role concurrency caps in `fleetd` so a fan-out can't exhaust the
local gateway.
- **Result schema:** whether reply envelopes should carry structured metadata (worker, model,
tokens) to help the lead's reduce step.
## Status
🟡 Design (2026-07-11). Orchestration layer over the selected `fleetd` server; inherits
herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see the Message-Server design.