13
1 Architecture
Dai Ha edited this page 2026-08-31 10:29:05 +07:00

1. Architecture

fleet lets a primary Claude Code session — the one on your own subscription — drive one or more members: separate Claude Code or OpenCode sessions that can run a different, cheaper or local model. A standalone daemon, fleetd, sits between them. It drives herdr, a terminal multiplexer, and exposes an MCP server that every session mounts. MCP (Model Context Protocol) is the client contract every session uses to talk to fleetd; see Message Server for the full tool schema.

A member is either a worker (does assigned work, replies, never delegates) or an architect (can delegate to workers but never spawns or tears one down). fleetd resolves its own name as fleet to every MCP client (fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:314).

One invariant and two facts explain almost everything else in this page. They are easy to miss because the old design docs got them wrong.

The subscription boundary — why fleet exists at all

This is the whole reason fleetd exists. The primary stays on the operator's subscription. A worker can run on a different, often cheaper, backend. There is no proxy on the primary, so there is no accidental bill. fleetd enforces this with one class, SubscriptionGuard (fleetd/src/main/java/dev/ltms/fleet/guard/SubscriptionGuard.java:9-21, its class-level javadoc):

  • A worker must egress to an off-subscription host. Before fleetd spawns a worker, it calls guard.assertWorker(baseUrl). This checks the profile's ANTHROPIC_BASE_URL against a configured allowlist of hostnames. It throws if the URL is missing, cannot be parsed, or is not on that list (SubscriptionGuard.java:32-45). The allowlist is a plain config value — an empty list unless the operator sets one — read from guard.offSubscriptionHosts in fleetd.yaml (fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:1153-1163, the Guard record). The configured hostnames differ per host, so this page names none.
  • The primary must never carry ANTHROPIC_BASE_URL. fleetd checks its own startup environment against this rule once, at boot, and refuses to start if the variable is set (SubscriptionGuard.java:49-54, called from fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:131-132). This guards fleetd's own process, not a worker's. The primary Claude session's own environment is a separate concern, upstream of fleetd.
  • A profile may opt out on purpose, with subscription: true. Some backends have no off-subscription endpoint to point at. For those, a profile can be marked subscription: true to run deliberately on the operator's subscription instead. When set, the launcher writes no ANTHROPIC_BASE_URL or ANTHROPIC_AUTH_TOKEN for that worker at all, and skips the guard's assertWorker check for that profile only. Every other profile keeps the hard refusal (FleetConfig.java:271-280, the subscription field's javadoc; enforced at fleetd/src/main/java/dev/ltms/fleet/member/ClaudeCodeLauncher.java:189-234). Setting both subscription: true and a baseUrl on the same profile is a contradiction, and fleetd refuses it loudly at two different points: once at spawn time, as an IllegalStateException that names the profile (ClaudeCodeLauncher.java:198-205), and once at config load time, if the profile's env: block tries to smuggle ANTHROPIC_BASE_URL or ANTHROPIC_AUTH_TOKEN back in past the skipped guard (FleetConfig.java:2040-2070, validateSubscriptionProfiles()).

Fact 1 — fleetd never forks a member process

fleetd does not run claude or opencode itself, and it does not fork a child process for a member. Spawning a member is two herdr calls over a Unix domain socket:

  1. tab.create (or pane.split), carrying the member's working directory and its environment map — this is where a worker's ANTHROPIC_BASE_URL and token are set (fleetd/src/main/java/dev/ltms/fleet/herdr/WorkspaceControl.java:84-94).
  2. agent.start, carrying the agent's name, kind (claude or opencode), and its argument list, targeting the pane just created (fleetd/src/main/java/dev/ltms/fleet/herdr/AgentControl.java:101-108).

herdr is the one that forks the pseudo-terminal (PTY) and starts the process, and it does so as its own OS user — the user herdr itself runs as, not fleetd's user (fleetd/src/main/java/dev/ltms/fleet/herdr/UnixSocketHerdrClient.java:1-30 — the client connects over a Unix domain socket, one connection per call).

This single fact explains three things that would otherwise look arbitrary:

  • The credential model. A member's environment comes from the env map fleetd hands to herdr at tab.create, not from a process fleetd controls directly. Whatever the shell that seeds that pane already has access to, the member can reach too — this is why the credential scrub work in the project's history exists at all.
  • Why fleetd and its members share one filesystem. fleetd, herdr, and every member pane must sit on the same machine and the same filesystem, because the Unix socket and the worktree paths fleetd hands to herdr are plain local paths, not a network protocol.
  • Why running a member as a different OS user needs a second herdr daemon. herdr forks as its own user, so the only way to change a member's user is to run a second herdr process under that other user and point fleetd at its socket. fleetd supports this today as memberHerdrSocket — a second herdr socket for members, separate from the lead's socket (fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java:36,84; wired up in fleetd/src/main/java/dev/ltms/fleet/Fleetd.java:153-154). With no memberHerdrSocket configured, the lead and its members share the one herdr daemon.

Fact 2 — identity comes from the connection, never an argument

When a session calls an MCP tool, fleetd does not read who is calling from a parameter in the call. It resolves the caller from the network connection itself: the loopback peer PID (from the operating system) mapped to a herdr pane (from herdr), or the primary if no pane matches (fleetd/src/main/java/dev/ltms/fleet/mcp/ConnectionIdentity.java:16-48). fleetd builds this resolved identity into every MCP call's context before the tool handler ever runs (FleetMcp.java:157-172).

This is why a worker can never reply as another worker, or as the primary: fleet_reply's caller is always the connection it arrived on, never a field in the request (FleetMcp.java:215 — "fleet_reply's identity is the CONNECTION, never an argument"). The resolved identity is one of three roles — primary, worker, or architect — reported by fleet_whoami and readable in code as Principal.Role (fleetd/src/main/java/dev/ltms/fleet/auth/Principal.java:16-95).

Authz is the table that decides what each role may do, checked on every MCP call and every REST call (fleetd/src/main/java/dev/ltms/fleet/auth/Authz.java:44-72):

Action Who may do it Why
SPAWN, STOP, DRAIN primary only fleet lifecycle is the primary's alone; an architect coordinates members but never stands them up or tears them down
SEND primary or architect delegating a turn; a worker sending would be it escalating its own role
REPLY, ASK only the caller that owns that session a caller may act only as the pane it occupies — this is what stops one member forging a reply for another
READ, METRICS primary, worker, or architect read-only status carries no secret and is open to every authenticated role

System overview

flowchart TB
    subgraph herd["herdr — owns the PTYs and panes"]
        PP["primary pane<br/>MCP client"]
        WP["worker pane(s)<br/>MCP client"]
        AP["architect pane(s)<br/>MCP client"]
    end

    subgraph FD["fleetd — the daemon, no Anthropic quota of its own"]
        MCP["MCP server (FleetMcp)<br/>+ REST surface (FleetApp)"]
        ID["ConnectionIdentity + Authz<br/>who is calling, what they may do"]
        MSG["MessageService<br/>rendezvous + status-gated injector"]
        MCP --> ID
        ID --> MSG
    end

    BR["broker (optional)<br/>AMQP — LavinMQ in production"]

    PP -->|"fleet_send / fleet_spawn / fleet_status"| MCP
    WP -->|"fleet_reply / fleet_ask"| MCP
    AP -->|"fleet_send / fleet_reply"| MCP
    MSG -->|"Unix domain socket<br/>tab.create, agent.start, agent.prompt"| herd
    MSG -.->|"durable reply inbox<br/>lead-to-lead coordination"| BR

Figure: every session — primary, worker, or architect — reaches fleetd only through its MCP tools. fleetd resolves the caller's identity from the connection, checks it against Authz, then either answers directly or drives herdr over the Unix socket described in Fact 1. The broker is optional infrastructure fleetd owns; no session ever talks to it directly.

Components

Component Role Notes
fleetd The daemon: an MCP server (FleetMcp) plus a REST surface (FleetApp) over one shared MessageService, so both are validated by parity rather than by re-implementing behaviour. Holds no Anthropic subscription quota itself. Registered MCP server name is fleet (FleetMcp.java:314). Listens on 127.0.0.1:8765 by default (FleetConfig.java:183-186).
MCP tools The full registered set, all on FleetMcp: fleet_send, fleet_reply, fleet_ask, fleet_status, fleet_poll, fleet_ack, fleet_spawn, fleet_list, fleet_stop, fleet_profiles, fleet_whoami. FleetMcp.java:301-327. There is no fleet_read tool — an earlier design doc named one, but it was never built.
REST routes /healthz, /metrics, /sessions, /agents, /members (GET/POST), /members/{paneId} (DELETE), /profiles, /sessions/{id}/message, /sessions/{id}/reply, /sessions/{id}/replies, /sessions/{id}/ask, /sessions/{id}/status, /tasks/{ticket}, plus the /mcp servlet. FleetApp.java:143-158, build(), read in full. There is no /events route — an earlier design doc described an SSE (GET /events) status stream; it was never built.
herdr An external terminal multiplexer. Owns the PTYs, the panes, and the live agent_status events. Forks every member process as its own OS user, over its own Unix socket. Socket is local-only per Fact 1 above.
Worker A member that does the assigned work and replies. Runs as a real claude or opencode process, inheriting the project's CLAUDE.md, hooks, skills and MCP mounts — never orchestrates or spawns. Launcher kind is claude-code (default) or opencode, set per profile (FleetConfig.java:265-266, 333, 335).
Architect A member that can delegate to workers (fleet_send) but has no lifecycle rights — it cannot fleet_spawn or fleet_stop a peer. Authz.java:53-58.
Broker (optional) An external AMQP broker fleetd owns for durable reply delivery across a restart, and for lead-to-lead coordination across daemons. Its presence swaps the in-memory reply inbox for the AMQP-backed one; absent, fleetd stays soft-state. FleetConfig.java:655-673. Production default is LavinMQ; a stock RabbitMQ works too, since both speak AMQP 0-9-1. Not Redis Streams or NATS JetStream — an earlier design doc named those; the shipped adapter is AMQP only.

How a delegation travels

sequenceDiagram
    autonumber
    participant P as "Primary — MCP client"
    participant F as "fleetd"
    participant H as "herdr"
    participant M as "Member pane (worker or architect)"

    P->>F: "fleet_send(sessionId, content)"
    F->>F: "resolve caller from the connection, check Authz.SEND"
    F->>F: "MessageService opens a rendezvous, injector waits for the pane to go idle"
    F->>H: "agent.prompt (inject the turn)"
    H->>M: "new turn"
    activate M
    Note over M: "member works the turn"
    M->>F: "fleet_reply(content)"
    Note over F: "caller resolved from the connection —<br/>a member can only reply as itself"
    H-->>F: "agent_status: working -> idle (turn done)"
    deactivate M
    F-->>P: "tool result = the reply (unparks the fleet_send call)"

Figure: one blocking delegation. fleet_send parks until the rendezvous resolves — on the member's fleet_reply, or on the pane going idle again, whichever comes first. A member's fleet_spawn step (the two herdr calls from Fact 1) already happened before this exchange; this diagram covers only the turn itself. fleetd never asks the primary to poll — the reply comes back as the tool result of the original call.

Two traffic modes: blocking, or fire-and-poll

fleet_send has two modes, both handled by the same class, MessageService (fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java, class-level javadoc read in full).

  • Blocking (the default). MessageService.send delivers the content and blocks the caller until the member's fleet_reply resolves the rendezvous, or the call times out (MessageService.java:498-519, the send overloads). This is the diagram above.
  • Fire-and-poll (wait:false). The MCP tool schema for fleet_send documents wait as "Block for the reply (default true); false returns a ticket to poll" (fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java:1100, the sendTool() schema). MessageService.sendAsync runs that same blocking send on a background virtual thread and returns a ticket immediately (MessageService.java:710-753); the caller collects the result later with fleet_poll(ticket), which reads back PENDING, DONE (with the reply), or FAILED (MessageService.java:765-793, the poll method).

Why the async path exists at all — stated directly in the code, not something I inferred: "A caller's MCP client caps a blocking call at ~60s, but a real delegated task runs for minutes" (MessageService.java:40-41, the class javadoc). A blocking fleet_send is capped by the calling client's own MCP timeout, not by anything fleetd controls — so any task expected to run longer than about a minute should use wait:false and fleet_poll, never a long blocking call.

What this page does not cover

This page covers the daemon, the MCP and REST surface, the identity model, the subscription boundary, the broker, and the two traffic modes. Every claim above carries a file:line reference into the source.

It does not cover the worker lifecycle, the idle reaper, or how a reply is recycled — those are in Message Server. It also names no hostname on the subscription-boundary allowlist, because that list is a plain config value and differs per host.

  • Message Server — the fleetd design in depth: MCP tool schema, herdr control, reply rendezvous, worker lifecycle, REST API.
  • Approaches — why herdr, and the transport alternatives considered and rejected (including the AgentAPI adapter, which was never built).
  • Team — a primary orchestrating a mixed fleet of Claude and OpenCode members.
  • 13 User Guide — the written procedure: install, run, and what to do when it breaks. Setup and Operations are older pages that now point here.

Sources