13
2 Message Server
Dai Ha edited this page 2026-08-31 10:29:05 +07:00

2. Message Server (fleetd)

fleetd is the always-on daemon that runs the fleet. It controls herdr, an agent multiplexer, over herdr's Unix-socket API. It gives every Claude Code session — the lead (the orchestrating session, called "primary" in the code) and every member (a spawned worker) — one clean way to send tasks and get replies back.

fleetd has two faces:

  • a SERVER face — an MCP server at /mcp, plus a REST API, sitting over the policy layer (session tracking, the subscription guard, and the reply rendezvous); and
  • a CLIENT face — a herdr socket client that starts members, sends text into their panes, and reads their live status.

The two faces are separate in the code. The MCP server is built and mounted in dev.ltms.fleet.mcp.FleetMcp (FleetMcp.java:150-174, FleetMcp.java:461-463), the REST routes are built in dev.ltms.fleet.rest.FleetApp.build() (FleetApp.java:126-160), and both sit on top of the same dev.ltms.fleet.msg.MessageService (FleetMcp.java:127, FleetApp.java:62).

A note on this page's sources. Everything below with a file:line reference was read directly from the code in this worktree. docs/MCP-Contract.md is not a normative reference for tool or route names — only its §6 (the rendezvous flows) is current; the rest is a pre-build design document whose names never caught up with the shipped code.

Why herdr, not a hand-rolled terminal reader

The project's earlier approach, AgentAPI, would have re-implemented a terminal emulator and guessed when an agent was done from screen stability. herdr already solves that as a running service: it owns the panes, tells fleetd the instant an agent's status changes (idle/working/blocked), and survives a detach/reattach over SSH. fleetd's job is the policy on top: which panes may be written to, when, and how a reply gets back to whoever sent the task.

flowchart TB
    subgraph herd["herdr — agent multiplexer"]
        PP["lead pane<br/>MCP client"]
        WP["member pane(s)<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
    end

    subgraph fleetd["fleetd — standalone daemon"]
        subgraph srv["SERVER face"]
            MCP["MCP server (/mcp)<br/>fleet_send · fleet_reply · fleet_ask · ..."]
            REST["REST API"]
            POL["policy: SessionManager<br/>SubscriptionGuard · MessageService/Rendezvous"]
        end
        subgraph cli["CLIENT face"]
            INJ["Injector<br/>status-gated FIFO per pane"]
            HCL["herdr socket client"]
        end
        MCP --> POL
        REST --> POL
        POL --> INJ --> HCL
    end

    PP -->|"MCP tools"| MCP
    WP -->|"MCP tools"| MCP
    HCL -->|"Unix socket"| herd

    classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
    classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
    class MCP,REST,POL,INJ,HCL core
    class PP,WP ext

Figure: fleetd is one daemon with a SERVER face (MCP + REST, over the policy layer) and a CLIENT face (the herdr socket client). Verified: FleetMcp.java:150-174 (MCP build), FleetApp.java:126-160 (REST routes), FleetMcp.java:127 / FleetApp.java:62 (both share one MessageService).

Mounting the MCP server

Both the lead and every member reach fleetd the same way: they mount it as an MCP server. The daemon listens by default on 127.0.0.1:8765 and serves the MCP endpoint at /mcp. The port comes from FleetConfig.Bind, which defaults to 8765 when the config does not set one (FleetConfig.java:183-187); the endpoint path is set at FleetMcp.java:152 (.mcpEndpoint("/mcp")). The server reports its own name as fleet at MCP initialize time (FleetMcp.java:313-315: .serverInfo("fleet", "0.1.0")).

claude mcp add --transport http fleet http://127.0.0.1:8765/mcp

The port in the command above is the default. A real deployment's actual port comes from its bind: block in fleetd.yaml (or bridged.yaml, depending on the host) — check that file before assuming 8765.

Tool reference

The registered tool set is built once, in FleetMcp's constructor (FleetMcp.java:301-327). The parameters listed below come from each tool's own schema method (FleetMcp.java:1088-1246), not from any older document.

Tool Who calls it Parameters Returns
fleet_send lead content (required), sessionId, timeoutMs, wait (default true), turnId, coordId Delegates content to the member named by sessionId and, by default, blocks for its reply. wait:false returns a ticket to poll with fleet_poll instead of blocking. Passing turnId (instead of sessionId) answers a member's open fleet_ask question. Passing coordId (instead of sessionId/turnId) sends to a peer lead's mailbox on another daemon — this is coordination between leads, not a task. sessionId, turnId and coordId are mutually exclusive.
fleet_reply member content (required) Ends a delegated turn with a structured answer. The member's identity comes from its connection, never an argument, so a member can only ever reply as itself.
fleet_ask member question (required), timeoutMs Pauses the member's current turn to ask the lead a question, and blocks until the lead answers (the lead answers with fleet_send{turnId, content}). Default timeout is 55 seconds and the cap is 115 seconds — kept under a typical MCP client's own ~60s call cap so the tool returns a clean timeout instead of the client just severing the connection.
fleet_status lead sessionId (required) The member's live status: idle, working, blocked, or unknown. If the member is paused mid-turn in an async fleet_ask, the result also carries the open question and its turnId.
fleet_poll lead ticket, target With ticket: the state of a fleet_send{wait:false} delegation — pending, done (with the reply), asking, or failed. With target instead: drains that member's reply inbox (replies that arrived when no fleet_send was open waiting for them).
fleet_ack lead target (required), msgId (required) Removes one specific reply from a member's inbox, leaving any others queued.
fleet_spawn lead role, profile, cwd, worktree, ticket, sessionName, resumeSessionId Starts a new member. role picks the contract (dev, reviewer, or architect; default dev); profile picks the backend (default is the configured default profile). worktree:true (with ticket) or worktree:<slug> provisions an isolated git worktree. Returns the member's sessionId (for fleet_send) and paneId (for fleet_stop).
fleet_list lead none The whole fleet: leads (peer orchestrators, each with sessionId, name, live status, and self:true on the caller's own row) and members (each with sessionId, paneId, role, profile, state, and — when set — worktree/branch/owner). When capacity facts are configured, also a capacity row per profile.
fleet_stop lead paneId (required) Tears a member down by its pane id.
fleet_profiles lead none The configured backend profiles, the default one, and (when CB-578's stage B quarantine has tripped) which profiles are currently refusing new spawns and for how long.
fleet_whoami any none The caller's own resolved role (primary, architect, or worker) and identity — the caller never has to guess its own role from a side channel.

fleet_read does not exist. The full registration list is at FleetMcp.java:301-327, and it has no tool by that name. Older pages named one; it was never built.

The rendezvous — a blocking call, not polling

When the lead calls fleet_send and blocks, fleetd does not make the lead poll. It parks the MCP call and resolves it the instant one of two things happens: the member calls fleet_reply, or (a fallback, CB-106) the member's turn ends without ever calling fleet_reply, in which case fleetd returns the scraped transcript tail instead, clearly flagged as such. The outcome list is MessageService.Outcome (MessageService.java:65-103), and FleetMcp.formatReply renders it back to the tool caller (FleetMcp.java:541-564).

sequenceDiagram
    autonumber
    participant L as Lead — MCP client
    participant F as "fleetd (MessageService + Injector)"
    participant M as Member — MCP client

    L->>F: fleet_send(sessionId, content) — call blocks
    F->>F: Injector delivers content when the pane is idle
    activate M
    M->>M: works the turn
    M->>F: fleet_reply(content)
    deactivate M
    F-->>L: tool result = the member's reply
    Note over L,M: if the member's turn ends with no fleet_reply,<br/>fleetd returns the scraped transcript tail instead

Figure: the send/reply round trip. Verified against MessageService.java:65-103 (the outcome enum) and FleetMcp.java:486-501 / 541-564 (fleet_send's handler and reply rendering).

fleet_ask — a member asking the lead back

fleet_ask is the reverse direction: a member pauses its own turn to ask the lead a question, and the lead answers by calling fleet_send again with turnId set instead of sessionId. The member then resumes the same turn. See FleetMcp.ask (FleetMcp.java:521-538), and the turnId branch in fleet_send (FleetMcp.java:198-203, calling FleetMcp.answer at FleetMcp.java:508-514).

sequenceDiagram
    autonumber
    participant L as Lead
    participant F as fleetd
    participant M as Member

    L->>F: fleet_send(sessionId, content) — call blocks
    F->>M: deliver content
    activate M
    M->>F: fleet_ask(question) — member's own call blocks
    F-->>L: tool result = QUESTION(text=question, turnId)
    L->>F: fleet_send(turnId, content=answer) — call blocks again
    F-->>M: unblocks fleet_ask with the answer
    M->>M: resumes the same turn
    M->>F: fleet_reply(content)
    deactivate M
    F-->>L: tool result = the member's reply

Figure: the reverse rendezvous. QUESTION is one of MessageService.Outcome's values (MessageService.java:91); FleetMcp.formatReply turns it into a tool result that names the turnId and tells the lead how to answer (FleetMcp.java:556-558).

Detached delegation (wait:false)

A blocking fleet_send call is capped by the caller's own MCP client timeout — usually around 60 seconds — but a real task can run for minutes. fleet_send{wait:false} runs the same delegation on a background thread and returns a ticket right away; the lead checks on it later with fleet_poll{ticket}. See the class-level doc comment on MessageService (MessageService.java:40-44, describing sendAsync) and the fleet_poll tool handler (FleetMcp.java:645-670).

A finished ticket is kept for 10 minutes after it completes (MessageService.java:62, TICKET_TTL_NANOS), then pruned. So the ten minutes are the window to collect the report, whatever the task's own runtime was.

This was a real bug until 2026-08-31 (#197): the TTL was measured from when the ticket was created, so the window was ten minutes minus however long the task ran. Any delegation lasting more than ten minutes had its report destroyed the moment it arrived, and fleet_poll{target} returned nothing rather than holding it. It is fixed, but a daemon that has not been redeployed since still behaves the old way. Separately, if a member calls fleet_reply while no fleet_send is open waiting for it, fleetd does not drop the reply — it queues it in that member's inbox, and a background push loop (ReplyPushLoop) nudges the lead's own pane to go check when this happens (MessageService.java:370-383, MessageService.reply). The mechanism that carries the nudge into the pane lives in ReplyPushLoop; MessageService.reply calls pushLoop.onReplyQueued(session) when a reply arrives with no waiter.

REST API

Every feature is also reachable over plain HTTP, so it can be tested and driven without an MCP client. FleetApp.build() (FleetApp.java:126-160) registers every route the daemon has; the table below is that complete list.

Method + path Purpose
GET /healthz Liveness. Also confirms herdr is reachable (and, if a second memberHerdrSocket is configured, that the member daemon is too).
GET /metrics Prometheus scrape (only registered when a metrics registry is configured).
GET /sessions Herdr workspaces, one row per workspace, with live agent status.
GET /agents Every agent herdr tracks, keyed by its Claude session id.
GET /members The fleet's own session roster, merged with live herdr status — this is what fleet_list is built from.
GET /profiles Configured backend profiles and the default one.
POST /members Spawn a member (fleet_spawn's REST equivalent).
DELETE /members/{paneId} Tear a member down (fleet_stop's REST equivalent).
POST /sessions/{id}/message Deliver a turn, blocking by default (fleet_send's REST equivalent); {"wait": false} in the body returns a ticket instead.
POST /sessions/{id}/reply A member's structured reply (fleet_reply's REST equivalent).
GET /sessions/{id}/replies Drain a member's reply inbox.
POST /sessions/{id}/ask A member's mid-turn question (fleet_ask's REST equivalent).
GET /sessions/{id}/status Live status, plus ready (whether the injector can currently deliver to this target) and any open question.
GET /tasks/{ticket} Poll an async (wait:false) delegation.

There is no GET /events route and no Server-Sent Events route of any kind. The table above lists every app.get, app.post and app.delete call in FleetApp.build(), and that is the whole set. An older design doc described an SSE status stream; it was never built.

The subscription boundary

The one invariant the whole design protects: whatever sets ANTHROPIC_BASE_URL is a member, never the lead. This is enforced in code, not only documented. SubscriptionGuard holds both checks (SubscriptionGuard.java):

  • assertWorker(baseUrl) refuses to spawn a member whose ANTHROPIC_BASE_URL is missing, or whose host is not on the configured off-subscription allowlist.
  • assertPrimaryClean(env) refuses to let the lead's own environment carry ANTHROPIC_BASE_URL at all.

Both throw GuardException on violation, which fleet_spawn turns into a subscription boundary: … tool error (FleetMcp.java:825-826) and the REST route into an HTTP 403 (FleetApp.java:384-385).

  • Architecture — the two-invariant model this page refines.
  • Approaches — why herdr was chosen; AgentAPI as discarded research.
  • Home — project overview.

Sources

  • herdr — socket API · GitHub
  • docs/MCP-Contract.md §6 only — the rest of that file is pre-build design and does not match the shipped tool names or routes.