Files
fleetd/docs/Team.md
T
Dai Ha 7d41ccccee
CI / build (push) Failing after 1m4s
CI / contract (push) Successful in 1m31s
docs: replace the decommissioned ollama.ltms.dev host
A page audit found the dead host still named in four places outside the wiki.
ollama.ltms.dev no longer exists; the one front door is llm.ltms.dev, and the
path differs by backend - /anthropic for claude-code, /v1 for opencode.
Anyone following README.md today points a member at nothing.

README.md also now links the new wiki chapter 13, the operator guide, since
the README's own next step used to be the never-written Setup page.

Not touched: SubscriptionGuard.java:14 still names ollama.ltms.dev, but only
in a javadoc line describing the Stage-1 example allowlist. It is a comment
about history, not a live default, so it stays until that class is next
edited for its own reasons.
2026-08-17 16:19:19 +02:00

6.8 KiB

Team — lead orchestrating a mixed Claude + local-LLM fleet

The message server (bridged) delivers one turn into one worker. A team is the layer above it: a Claude team-lead that fans a job out across a mixed fleet of workers — some on Claude, some on the remote local LLM — and reduces their replies. Same bridged delivery, same subscription boundary; this doc is only about orchestration — who the workers are, how the lead picks one, and how it runs many at once.

Delivery mechanics (blocking POST /message, status-gated reply envelope) live in the Message-Server design. Transport rationale is in Approaches. This doc assumes both.

The team

  • Team-lead — the primary Opus (Claude Code, env CLEAN, on Pro/Max). Not a worker; a thin client of bridged. It plans, routes, dispatches, and integrates, and never sets ANTHROPIC_BASE_URL.
  • Workers — a herd of claude panes in herdr, each an addressable bridged session with its own model/env:
    • Claude workers (clean env, e.g. Sonnet) — reasoning-heavy or high-accuracy subtasks.
    • Local workers (ANTHROPIC_BASE_URL=https://llm.ltms.dev/anthropic) — bulk, cheap, or embarrassingly parallel subtasks.

Every worker is still a real Claude Code process (inherits CLAUDE.md, hooks, skills, MCP) — only its model differs. Scale each kind horizontally by adding panes.

Topology

flowchart TB
    LEAD["lead — Opus<br/>(Claude Code, env CLEAN)"]
    BD["bridged<br/>message server + router"]
    HERDR["herdr<br/>panes · agent-status"]
    WC1["w-claude-1<br/>Sonnet · CLEAN"]
    WC2["w-claude-2<br/>Sonnet · CLEAN"]
    WL1["w-local-1<br/>ANTHROPIC_BASE_URL set"]
    WL2["w-local-2<br/>ANTHROPIC_BASE_URL set"]
    ANT["api.anthropic.com<br/>(Pro/Max)"]
    OLL["llm.ltms.dev<br/>(gateway to the local model)"]

    LEAD -->|"blocking POST /message (target role)"| BD
    BD -->|"Unix socket · send_text · events.subscribe"| HERDR
    HERDR --> WC1 & WC2 & WL1 & WL2
    WC1 --> ANT
    WC2 --> ANT
    WL1 --> OLL
    WL2 --> OLL

    classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
    classDef local fill:#6b46c1,stroke:#44337a,color:#ffffff;
    class LEAD,WC1,WC2 ext
    class BD,HERDR core
    class WL1,WL2 local

Roles & routing

Role Env Model Route here when…
lead clean Opus (sub) always — it does the routing
w-claude-* clean Sonnet (sub) task needs Claude-grade reasoning / careful edits
w-local-* ANTHROPIC_BASE_URL set local LLM task is bulk / cheap / embarrassingly parallel

The lead applies this rubric itself, guided by its CLAUDE.md team charter (below). Worker selection is policy in the lead, not a bridged concern — bridged just delivers to the session the lead names.

Subscription boundary in a team

Unchanged from the base architecture, and it scales with the fleet: only local-worker panes launch with ANTHROPIC_BASE_URL. The lead and every Claude worker stay env-clean on the subscription. bridged enforces which panes may carry the off-subscription env, so adding workers never widens the boundary.

Parallel fan-out (map / reduce)

The lead's advantage over a single bridge is concurrency: independent subtasks go to different workers at once, then results are gathered.

sequenceDiagram
    participant L as lead (Opus)
    participant B as bridged
    participant WC as w-claude-1
    participant WL as w-local-1

    Note over L: split job → subtask A (reasoning), subtask B (bulk)
    par A → Claude worker
        L->>B: POST /message {role: w-claude, prompt: A}
        B->>WC: send_text into running pane
        WC-->>B: status working → idle + Stop-hook envelope
        B-->>L: 200 reply A
    and B → local worker
        L->>B: POST /message {role: w-local, prompt: B}
        B->>WL: send_text into running pane
        WL-->>B: status working → idle + Stop-hook envelope
        B-->>L: 200 reply B
    end
    Note over L: reduce → integrate A + B into final answer
  • Map: the lead issues N concurrent blocking POST /message calls (one per subtask → its chosen worker). Each call blocks only that request; bridged holds it open until the worker's turn completes (status-gated) and returns the reply envelope.
  • Reduce: the lead collects the N envelopes and integrates. A slow local worker never blocks a fast Claude worker — wall-clock ≈ the slowest single subtask, not the sum.
  • Detached / long jobs use the async broker path instead of a held request (Channel 2 in the base architecture), so the lead never busy-polls across turns.

Fan-out is bounded by the herd size (pane count) and bridged's concurrency policy, not by the lead.

Knowing the roster

The lead discovers its team from bridged (session list / roles) rather than hard-coding pane ids, so workers can be added or restarted without editing the lead. A minimal charter in the lead's CLAUDE.md turns Opus into the orchestrator:

## Your team (via bridged)
You are the team-lead. Delegate through the bridged client — never launch workers yourself.
Roster: ask bridged for current sessions/roles.
- w-claude-* — Claude Sonnet. Reasoning-heavy / high-accuracy subtasks.
- w-local-*  — remote local LLM. Bulk, cheap, or parallelizable subtasks.

Route each subtask by the rubric in the Team design. For independent subtasks, DISPATCH ALL
of them (concurrent blocking sends), THEN gather — never serialize independent work.
Integrate the reply envelopes; you own the final answer.

Wrap the send as a Claude Code skill (/delegate <role> "<task>") so the lead calls one tool instead of hand-rolling the HTTP request.

What this layer does NOT change

  • Delivery is still bridged → herdr pane.send_text + status events (Message-Server).
  • Completion timing is still the worker status event; reply content still rides the worker Stop-hook envelope.
  • Single-host still applies: herdr's socket is local, so the whole herd lives on the bridged host. The lead may be remote — it only needs HTTP to bridged.

Open questions

  • Routing intelligence: rubric-in-CLAUDE.md (lead decides) vs. a bridged role-router (label-based). Start with the former; promote to the latter if routing logic grows.
  • Backpressure: per-role concurrency caps in bridged so a fan-out can't exhaust the local gateway.
  • Result schema: whether reply envelopes should carry structured metadata (worker, model, tokens) to help the lead's reduce step.

Status

🟡 Design (2026-07-11). Orchestration layer over the selected bridged server; inherits herdr (chosen) + AgentAPI (fallback). Delivery unchanged — see the Message-Server design.