Files
fleetd/docs/MCP-Contract.md
Dai Ha e897e5257b
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
fleetd #114: delete the drifted tool catalogue, keep the flows, guard the names
docs/MCP-Contract.md was written 2026-07-14, before any MCP code existed,
and never caught up. CLAUDE.md points every session in the fleet at it.

Audited against the code today. The drift was not confined to the tool
table the ticket reported:

  section 3  still described the OLD identity rule - "any connection that
             does not map to a known worker is treated as a primary".
             That was a real privilege bug, fixed since by the ancestry
             walk in #161. The page still taught it.
  section 4  names port 8080 (the mount is 8765) and says the pom does
             not yet carry an MCP dependency.
  section 5  named fleet_read and fleet_cancel, which do not exist, and
             omitted fleet_poll, fleet_ack, fleet_profiles, fleet_whoami
             and fleet_list, which do.
  section 8  says turn_id where the code says turnId, and has no row for
             the exhausted outcome CB-578 added.
  sections
  9, 10, 11  pre-build planning: "new work" columns, open decisions long
             since decided, CB-1xx placeholders.

Every one of those is the same defect: a hand-maintained second copy of
something the code already states. So the copy is deleted rather than
corrected - correcting it just restarts the clock.

What survives is the flows and the status gating, because a flow is a
shape rather than a name, and shapes are what this page was ever good
for. They are rewritten with the names checked against the code, and
extended with what has been learned since: the ~60s cap on a blocking
send, the ~55s ask window, and the three ways the turn-done fallback
loses a report (clipped, echoed brief, slow member).

389 lines -> 188.

The names that remain are guarded. McpContractDocTest fails if the page
names a fleet_* tool FleetMcp does not register, and - because an empty
set is a subset of everything - a second test pins that both sides
actually found names, so the check cannot pass by checking nothing. A
third pins the "this is not the tool reference" sentence, which is the
fix itself: without it someone helpfully re-adds a tool table.

Mutation-tested both ways, 0 compile errors each: adding `fleet_read` to
the doc fails theDocNamesNoToolThatDoesNotExist ("names [fleet_read] ...
Checked 6 name(s)"); removing the disclaimer fails
theDocStillDisclaimsBeingTheToolReference.

All 5 mermaid diagrams render under mermaid-cli.

CLAUDE.md's pointer said "section 6 only" and now names the guard
instead. It is in the project addendum, so the canonical block is
untouched - verified still byte-identical with the wiki template.

REST is split out to #252: 14 routes, documented nowhere, and it IS a
supported operator surface - one of them drains on read.

1232 tests, 0 failures.
2026-09-03 12:52:12 +07:00

7.6 KiB

MCP flows and error model — fleetd

What this page is. The flows: how a delegation, a clarification, a detached task and a silent member each travel through fleetd. These shapes are what shipped, and they are hard to read off the code because they span the MCP face, the rendezvous registry, the Injector and herdr.

What this page is NOT: a tool reference. It deliberately holds no tool catalogue, no parameter tables and no REST paths. The live MCP schema is the authority — each tool's own description and parameters, as mounted — with the intent→tool table in CLAUDE.md as the short form.

That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped; the page did not follow. By August it named two tools that do not exist, omitted five that do, had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve, and — worst — still described an identity model ("any connection that does not map to a known worker is treated as a primary") that was a real privilege bug, fixed since by the ancestry walk in fleetd #161. Every one of those errors is the same error: a second, hand-maintained copy of something the code already states. So the second copy is gone rather than corrected. Only the flows remain, because a flow is a shape rather than a name, and shapes are what this page was ever good for.

The names that do appear below are checked by McpContractDocTest, which fails if this page names a fleet_* tool the server does not register. That test is the whole reason it is safe to write a tool name here at all.


1. Rendezvous flows

1.1 Delegation — happy path

One blocking call, zero polls. The lead's call is held open by fleetd until the member answers.

sequenceDiagram
    participant P as "Lead (primary)"
    participant B as "fleetd (MCP + Injector)"
    participant H as herdr
    participant W as "Member"

    P->>B: "fleet_send{sessionId, content} — blocks"
    B->>B: "register waiter(sessionId)"
    B->>H: "agent.send — only in an injectable window"
    H-->>W: "prompt injected"
    W->>W: "works the turn"
    W->>B: "fleet_reply{content}"
    B->>B: "resolve waiter"
    B-->>P: "{ outcome: reply }"

The cap that matters: a blocking fleet_send is bounded by the caller's own MCP client timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow in §1.3, or the lead's call returns while the member is still working.

1.2 Clarification — reverse rendezvous

The member pauses mid-turn to ask, the lead answers, and the member resumes the same turn with its context intact.

sequenceDiagram
    participant P as "Lead"
    participant B as fleetd
    participant W as "Member"

    P->>B: "fleet_send{sessionId, content} — blocks"
    B-->>W: "content injected"
    W->>B: "fleet_ask{question} — member blocks"
    B-->>P: "{ outcome: question, turnId }"
    P->>B: "fleet_send{turnId, content} — answers THIS turn"
    B-->>W: "fleet_ask returns the answer"
    W->>W: "resumes the same turn"
    W->>B: "fleet_reply{content}"
    B-->>P: "{ outcome: reply }"

Answer with turnId, never sessionId. A sessionId send starts a new turn; it does not resolve the waiting fleet_ask.

The window is about 55 seconds and no nudge extends it. So never brief a member to "ask me": decide before delegating, or give the member an explicit default to fall back on.

1.3 Detached delegation — the lead does not block

The lead gets a ticket immediately and collects the answer later. This is the flow for any real task, because of the ~60s cap in §1.1.

sequenceDiagram
    participant P as "Lead"
    participant B as fleetd
    participant W as "Member"

    P->>B: "fleet_send{sessionId, content, wait:false}"
    B-->>P: "accepted — ticket"
    P->>P: "continues its own work"
    W->>B: "fleet_reply{content}"
    Note over B: "no waiter is blocked — the reply is held"
    B->>B: "nudge the lead's own pane (status-gated)"
    P->>B: "fleet_poll{ticket}"
    B-->>P: "the member's report"
    P->>B: "fleet_ack{target, msgId}"

A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.

1.4 The member never replies — turn-done fallback

A member that ends its turn without fleet_reply still produces something: fleetd reads its pane tail. This is a fallback, not a channel — it is lossy in three separate ways, and every one of them has produced a wrong answer in practice.

sequenceDiagram
    participant P as "Lead"
    participant B as fleetd
    participant W as "Member"

    P->>B: "fleet_send — blocks or detaches"
    B-->>W: "content injected"
    W->>W: "works, never calls fleet_reply"
    B->>B: "StatusPoller sees the turn end"
    B->>B: "read the pane tail"
    B->>B: "classify: exhausted? echoed brief? real report?"
    B-->>P: "{ outcome: turn_done } or a named failure"

The three ways it goes wrong, and what each looks like now:

What happened What the lead used to get What it gets today
The report is longer than the scrape window The end silently cut off Still clipped, but marked partial
The member never started — spent credential The lead's own brief echoed back as a report A named failure: backend exhausted
The member is simply slow A tail of work in progress Unchanged — read it as a hint, not a result

The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in it from the member. It is suppressed now, but the general rule stands — check the member's worktree with git log before believing a report you did not watch arrive.


2. Status gating

Delivery only happens in a safe window. fleet_send is a producer for the Injector, which already enforces this through AgentStatus.injectable().

stateDiagram-v2
    [*] --> IDLE
    IDLE --> WORKING: "message delivered, picked up"
    WORKING --> IDLE: "turn done"
    WORKING --> BLOCKED: "awaits input"
    BLOCKED --> WORKING: "input delivered"
    IDLE --> UNKNOWN: "detection glitch"
    BLOCKED --> UNKNOWN: "detection glitch"
    UNKNOWN --> IDLE: "re-detected"

    note right of IDLE
        injectable — deliver head of FIFO
    end note
    note right of BLOCKED
        injectable — deliver head of FIFO
    end note
    note right of WORKING
        NOT injectable — counts as pickup
    end note
    note right of UNKNOWN
        NOT injectable, NOT a pickup — wait
    end note

At most one message per turn. After a send, the Injector waits for a WORKING pickup before delivering the next, with a grace-poll fallback for turns that finish faster than the poll interval.

Two consequences a lead feels directly:

  • A second send to a busy member never lands. It reports as queued and times out. The member is fine; the message simply waits, and then restarts the member when it next goes idle.
  • A spawned member is not deliverable until it has mounted the MCP. Until then a send waits on that gate for about 60 seconds and then fails without ever reaching the pane.

UNKNOWN is deliberately neither injectable nor a pickup. A pane whose status cannot be read is not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an unreadable pane as a ready one.