docs/MCP-Contract.md was written 2026-07-14, before any MCP code existed,
and never caught up. CLAUDE.md points every session in the fleet at it.
Audited against the code today. The drift was not confined to the tool
table the ticket reported:
section 3 still described the OLD identity rule - "any connection that
does not map to a known worker is treated as a primary".
That was a real privilege bug, fixed since by the ancestry
walk in #161. The page still taught it.
section 4 names port 8080 (the mount is 8765) and says the pom does
not yet carry an MCP dependency.
section 5 named fleet_read and fleet_cancel, which do not exist, and
omitted fleet_poll, fleet_ack, fleet_profiles, fleet_whoami
and fleet_list, which do.
section 8 says turn_id where the code says turnId, and has no row for
the exhausted outcome CB-578 added.
sections
9, 10, 11 pre-build planning: "new work" columns, open decisions long
since decided, CB-1xx placeholders.
Every one of those is the same defect: a hand-maintained second copy of
something the code already states. So the copy is deleted rather than
corrected - correcting it just restarts the clock.
What survives is the flows and the status gating, because a flow is a
shape rather than a name, and shapes are what this page was ever good
for. They are rewritten with the names checked against the code, and
extended with what has been learned since: the ~60s cap on a blocking
send, the ~55s ask window, and the three ways the turn-done fallback
loses a report (clipped, echoed brief, slow member).
389 lines -> 188.
The names that remain are guarded. McpContractDocTest fails if the page
names a fleet_* tool FleetMcp does not register, and - because an empty
set is a subset of everything - a second test pins that both sides
actually found names, so the check cannot pass by checking nothing. A
third pins the "this is not the tool reference" sentence, which is the
fix itself: without it someone helpfully re-adds a tool table.
Mutation-tested both ways, 0 compile errors each: adding `fleet_read` to
the doc fails theDocNamesNoToolThatDoesNotExist ("names [fleet_read] ...
Checked 6 name(s)"); removing the disclaimer fails
theDocStillDisclaimsBeingTheToolReference.
All 5 mermaid diagrams render under mermaid-cli.
CLAUDE.md's pointer said "section 6 only" and now names the guard
instead. It is in the project addendum, so the canonical block is
untouched - verified still byte-identical with the wiki template.
REST is split out to #252: 14 routes, documented nowhere, and it IS a
supported operator surface - one of them drains on read.
1232 tests, 0 failures.
7.6 KiB
MCP flows and error model — fleetd
What this page is. The flows: how a delegation, a clarification, a detached task and a silent member each travel through
fleetd. These shapes are what shipped, and they are hard to read off the code because they span the MCP face, the rendezvous registry, theInjectorand herdr.What this page is NOT: a tool reference. It deliberately holds no tool catalogue, no parameter tables and no REST paths. The live MCP schema is the authority — each tool's own description and parameters, as mounted — with the intent→tool table in
CLAUDE.mdas the short form.That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped; the page did not follow. By August it named two tools that do not exist, omitted five that do, had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve, and — worst — still described an identity model ("any connection that does not map to a known worker is treated as a primary") that was a real privilege bug, fixed since by the ancestry walk in fleetd #161. Every one of those errors is the same error: a second, hand-maintained copy of something the code already states. So the second copy is gone rather than corrected. Only the flows remain, because a flow is a shape rather than a name, and shapes are what this page was ever good for.
The names that do appear below are checked by
McpContractDocTest, which fails if this page names afleet_*tool the server does not register. That test is the whole reason it is safe to write a tool name here at all.
1. Rendezvous flows
1.1 Delegation — happy path
One blocking call, zero polls. The lead's call is held open by fleetd until the member answers.
sequenceDiagram
participant P as "Lead (primary)"
participant B as "fleetd (MCP + Injector)"
participant H as herdr
participant W as "Member"
P->>B: "fleet_send{sessionId, content} — blocks"
B->>B: "register waiter(sessionId)"
B->>H: "agent.send — only in an injectable window"
H-->>W: "prompt injected"
W->>W: "works the turn"
W->>B: "fleet_reply{content}"
B->>B: "resolve waiter"
B-->>P: "{ outcome: reply }"
The cap that matters: a blocking fleet_send is bounded by the caller's own MCP client
timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow
in §1.3, or the lead's call returns while the member is still working.
1.2 Clarification — reverse rendezvous
The member pauses mid-turn to ask, the lead answers, and the member resumes the same turn with its context intact.
sequenceDiagram
participant P as "Lead"
participant B as fleetd
participant W as "Member"
P->>B: "fleet_send{sessionId, content} — blocks"
B-->>W: "content injected"
W->>B: "fleet_ask{question} — member blocks"
B-->>P: "{ outcome: question, turnId }"
P->>B: "fleet_send{turnId, content} — answers THIS turn"
B-->>W: "fleet_ask returns the answer"
W->>W: "resumes the same turn"
W->>B: "fleet_reply{content}"
B-->>P: "{ outcome: reply }"
Answer with turnId, never sessionId. A sessionId send starts a new turn; it does not
resolve the waiting fleet_ask.
The window is about 55 seconds and no nudge extends it. So never brief a member to "ask me": decide before delegating, or give the member an explicit default to fall back on.
1.3 Detached delegation — the lead does not block
The lead gets a ticket immediately and collects the answer later. This is the flow for any real task, because of the ~60s cap in §1.1.
sequenceDiagram
participant P as "Lead"
participant B as fleetd
participant W as "Member"
P->>B: "fleet_send{sessionId, content, wait:false}"
B-->>P: "accepted — ticket"
P->>P: "continues its own work"
W->>B: "fleet_reply{content}"
Note over B: "no waiter is blocked — the reply is held"
B->>B: "nudge the lead's own pane (status-gated)"
P->>B: "fleet_poll{ticket}"
B-->>P: "the member's report"
P->>B: "fleet_ack{target, msgId}"
A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.
1.4 The member never replies — turn-done fallback
A member that ends its turn without fleet_reply still produces something: fleetd reads its
pane tail. This is a fallback, not a channel — it is lossy in three separate ways, and every
one of them has produced a wrong answer in practice.
sequenceDiagram
participant P as "Lead"
participant B as fleetd
participant W as "Member"
P->>B: "fleet_send — blocks or detaches"
B-->>W: "content injected"
W->>W: "works, never calls fleet_reply"
B->>B: "StatusPoller sees the turn end"
B->>B: "read the pane tail"
B->>B: "classify: exhausted? echoed brief? real report?"
B-->>P: "{ outcome: turn_done } or a named failure"
The three ways it goes wrong, and what each looks like now:
| What happened | What the lead used to get | What it gets today |
|---|---|---|
| The report is longer than the scrape window | The end silently cut off | Still clipped, but marked partial |
| The member never started — spent credential | The lead's own brief echoed back as a report | A named failure: backend exhausted |
| The member is simply slow | A tail of work in progress | Unchanged — read it as a hint, not a result |
The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in
it from the member. It is suppressed now, but the general rule stands — check the member's
worktree with git log before believing a report you did not watch arrive.
2. Status gating
Delivery only happens in a safe window. fleet_send is a producer for the Injector, which
already enforces this through AgentStatus.injectable().
stateDiagram-v2
[*] --> IDLE
IDLE --> WORKING: "message delivered, picked up"
WORKING --> IDLE: "turn done"
WORKING --> BLOCKED: "awaits input"
BLOCKED --> WORKING: "input delivered"
IDLE --> UNKNOWN: "detection glitch"
BLOCKED --> UNKNOWN: "detection glitch"
UNKNOWN --> IDLE: "re-detected"
note right of IDLE
injectable — deliver head of FIFO
end note
note right of BLOCKED
injectable — deliver head of FIFO
end note
note right of WORKING
NOT injectable — counts as pickup
end note
note right of UNKNOWN
NOT injectable, NOT a pickup — wait
end note
At most one message per turn. After a send, the Injector waits for a WORKING pickup before
delivering the next, with a grace-poll fallback for turns that finish faster than the poll
interval.
Two consequences a lead feels directly:
- A second send to a busy member never lands. It reports as queued and times out. The member is fine; the message simply waits, and then restarts the member when it next goes idle.
- A spawned member is not deliverable until it has mounted the MCP. Until then a send waits on that gate for about 60 seconds and then fails without ever reaching the pane.
UNKNOWN is deliberately neither injectable nor a pickup. A pane whose status cannot be read is
not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an
unreadable pane as a ready one.