Files
fleetd/docs/MCP-Contract.md
T
Dai Ha 291dc02c77 fleetd #743: document that a hand-opened pane is reachable with no config
The lead's intent->tool table had no row for messaging an unconfigured pane,
and the user-scope instruction file said outright that the fleet has no route
to an interactive session unless an operator registers it as a collaborator.
That claim is false and it is load-bearing: a session reading it concludes the
exchange is impossible and stops, which is what happened here.

Delivery is gated on presence, not on SEND. contextExtractor runs on every MCP
request including initialize, markTrackedCallerPresent enrols an observer into
MemberPresence, and deliverableTo tests presence before the lead and
collaborator maps. So connecting the server is the enrolment, and fleet_reply
is gated on owning your own pane, which every pane does.

Add the table row, and add the enrolment side of the deliverability gate to the
flows page next to the existing "a spawned member is not deliverable until it
has mounted the MCP" bullet, which is the same gate read the other way.

Measured on a real pane, not a fake: trinotes answered with no fleet config, no
restart, and its fleet_* tools still deferred and unloaded.
2026-10-05 08:56:01 +02:00

8.1 KiB

MCP flows and error model — fleetd

What this page is. The flows: how a delegation, a clarification, a detached task and a silent member each travel through fleetd. These shapes are what shipped, and they are hard to read off the code because they span the MCP face, the rendezvous registry, the Injector and herdr.

What this page is NOT: a tool reference. It deliberately holds no tool catalogue, no parameter tables and no REST paths. The live MCP schema is the authority — each tool's own description and parameters, as mounted — with the intent→tool table in CLAUDE.md as the short form.

That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped; the page did not follow. By August it named two tools that do not exist, omitted five that do, had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve, and — worst — still described an identity model ("any connection that does not map to a known worker is treated as a primary") that was a real privilege bug, fixed since by the ancestry walk in fleetd #161. Every one of those errors is the same error: a second, hand-maintained copy of something the code already states. So the second copy is gone rather than corrected. Only the flows remain, because a flow is a shape rather than a name, and shapes are what this page was ever good for.

The names that do appear below are checked by McpContractDocTest, which fails if this page names a fleet_* tool the server does not register. That test is the whole reason it is safe to write a tool name here at all.


1. Rendezvous flows

1.1 Delegation — happy path

One blocking call, zero polls. The lead's call is held open by fleetd until the member answers.

sequenceDiagram
    participant P as "Lead (primary)"
    participant B as "fleetd (MCP + Injector)"
    participant H as herdr
    participant W as "Member"

    P->>B: "fleet_send{sessionId, content} — blocks"
    B->>B: "register waiter(sessionId)"
    B->>H: "agent.send — only in an injectable window"
    H-->>W: "prompt injected"
    W->>W: "works the turn"
    W->>B: "fleet_reply{content}"
    B->>B: "resolve waiter"
    B-->>P: "{ outcome: reply }"

The cap that matters: a blocking fleet_send is bounded by the caller's own MCP client timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow in §1.3, or the lead's call returns while the member is still working.

1.2 Clarification — reverse rendezvous

The member pauses mid-turn to ask, the lead answers, and the member resumes the same turn with its context intact.

sequenceDiagram
    participant P as "Lead"
    participant B as fleetd
    participant W as "Member"

    P->>B: "fleet_send{sessionId, content} — blocks"
    B-->>W: "content injected"
    W->>B: "fleet_ask{question} — member blocks"
    B-->>P: "{ outcome: question, turnId }"
    P->>B: "fleet_send{turnId, content} — answers THIS turn"
    B-->>W: "fleet_ask returns the answer"
    W->>W: "resumes the same turn"
    W->>B: "fleet_reply{content}"
    B-->>P: "{ outcome: reply }"

Answer with turnId, never sessionId. A sessionId send starts a new turn; it does not resolve the waiting fleet_ask.

The window is about 55 seconds and no nudge extends it. So never brief a member to "ask me": decide before delegating, or give the member an explicit default to fall back on.

1.3 Detached delegation — the lead does not block

The lead gets a ticket immediately and collects the answer later. This is the flow for any real task, because of the ~60s cap in §1.1.

sequenceDiagram
    participant P as "Lead"
    participant B as fleetd
    participant W as "Member"

    P->>B: "fleet_send{sessionId, content, wait:false}"
    B-->>P: "accepted — ticket"
    P->>P: "continues its own work"
    W->>B: "fleet_reply{content}"
    Note over B: "no waiter is blocked — the reply is held"
    B->>B: "nudge the lead's own pane (status-gated)"
    P->>B: "fleet_poll{ticket}"
    B-->>P: "the member's report"
    P->>B: "fleet_ack{target, msgId}"

A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.

1.4 The member never replies — turn-done fallback

A member that ends its turn without fleet_reply still produces something: fleetd reads its pane tail. This is a fallback, not a channel — it is lossy in three separate ways, and every one of them has produced a wrong answer in practice.

sequenceDiagram
    participant P as "Lead"
    participant B as fleetd
    participant W as "Member"

    P->>B: "fleet_send — blocks or detaches"
    B-->>W: "content injected"
    W->>W: "works, never calls fleet_reply"
    B->>B: "StatusPoller sees the turn end"
    B->>B: "read the pane tail"
    B->>B: "classify: exhausted? echoed brief? real report?"
    B-->>P: "{ outcome: turn_done } or a named failure"

The three ways it goes wrong, and what each looks like now:

What happened What the lead used to get What it gets today
The report is longer than the scrape window The end silently cut off Still clipped, but marked partial
The member never started — spent credential The lead's own brief echoed back as a report A named failure: backend exhausted
The member is simply slow A tail of work in progress Unchanged — read it as a hint, not a result

The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in it from the member. It is suppressed now, but the general rule stands — check the member's worktree with git log before believing a report you did not watch arrive.


2. Status gating

Delivery only happens in a safe window. fleet_send is a producer for the Injector, which already enforces this through AgentStatus.injectable().

stateDiagram-v2
    [*] --> IDLE
    IDLE --> WORKING: "message delivered, picked up"
    WORKING --> IDLE: "turn done"
    WORKING --> BLOCKED: "awaits input"
    BLOCKED --> WORKING: "input delivered"
    IDLE --> UNKNOWN: "detection glitch"
    BLOCKED --> UNKNOWN: "detection glitch"
    UNKNOWN --> IDLE: "re-detected"

    note right of IDLE
        injectable — deliver head of FIFO
    end note
    note right of BLOCKED
        injectable — deliver head of FIFO
    end note
    note right of WORKING
        NOT injectable — counts as pickup
    end note
    note right of UNKNOWN
        NOT injectable, NOT a pickup — wait
    end note

At most one message per turn. After a send, the Injector waits for a WORKING pickup before delivering the next, with a grace-poll fallback for turns that finish faster than the poll interval.

Two consequences a lead feels directly:

  • A second send to a busy member never lands. It reports as queued and times out. The member is fine; the message simply waits, and then restarts the member when it next goes idle.
  • A spawned member is not deliverable until it has mounted the MCP. Until then a send waits on that gate for about 60 seconds and then fails without ever reaching the pane.
  • The same gate is what makes an unconfigured pane deliverable. contextExtractor runs on every MCP request, initialize included, and markTrackedCallerPresent enrols a spawned member or an observer into MemberPresence; deliverableTo then tests presence before the lead and collaborator maps. So mounting the server is the enrolment, and a tab a person opened by hand can be sent to with no config and no restart. It answers with fleet_reply — it cannot fleet_send, because Authz keeps SEND to a primary, an architect or a collaborator.

UNKNOWN is deliberately neither injectable nor a pickup. A pane whose status cannot be read is not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an unreadable pane as a ready one.