Lead-to-lead coordination has a send half with tools and a receive half without #361

Closed
opened 2026-09-05 07:17:27 +02:00 by ltms · 1 comment
Owner

Built the first real cross-host link on 2026-09-05 (Mac fleet coordId: mac ↔ fleet01
coordId: fleet01, shared /coord vhost). It works. Using it exposed that the MCP surface only
covers sending.

Today a lead has fleet_send{coordId}, fleet_reply, and fleet_list (which reports its own
coord-id). That is the whole surface.

1. No peer discovery

fleet_list returns:

"leads":[{"sessionId":"term_65a4c36d11af27","name":"opus","status":"working","self":true}],
"coordinator":{"selfId":"mac","configured":true}

fleet01 appears nowhere. Those leads rows come from CallerResolver.leads(), which resolves
local panes only — FleetMcp.java:1093 says as much.

So a lead cannot learn that a peer exists, or what coord-id addresses it. I only knew to write
coordId: "fleet01" because I had configured both daemons myself an hour earlier. Any other
session would have to be told out of band, which is exactly what this channel is supposed to
replace.

2. No way to read your own mailbox

Receiving is push-only: LeadCoordLoop consumes and types into the pane. There is no fleet_poll
equivalent for coordination messages, and nothing reports how many are waiting.

To find out whether a message I sent was delivered or still queued I had to ssh to the other
host and run docker exec fleet-lavinmq lavinmqctl list_queues -p coord, then read journald. None
of that is reachable through MCP, and a lead on a host with no broker shell access is blind to its
own inbox.

3. fleet_send{coordId} reports success for a message the peer may never see

FleetMcp.java:721-727:

leadChannel.publish(coordId, msg);
...
return text("delivered to peer lead " + coordId + " (msgId " + msg.msgId() + ")");

"Delivered" means published to the queue. It does not mean the peer's pane received it. Those
two came apart on the first real message: fleet_send reported success and the message then sat
undelivered through three daemon restarts, because of the duplicate-lead-tab bug (#359). With #359
present the gap is permanent — the peer never sees the message and the sender is told "delivered"
every time.

This is the same shape as #359 and #360: a green result for a control that did not take effect.

Suggested direction (candidate, not a decided design)

One mechanism may cover all three. An AMQP passive queue declare on lead.<coordId>.inbox
returns both message_count and consumer_count. I confirmed this from a client during the
build:

lead.fleet01.inbox visible: messages=0 consumers=1

consumers=1 means a daemon is attached and consuming; messages=N is the backlog. That is real
information about a peer, obtained with no new infrastructure and no presence protocol.

Sketch:

  • a coordinator.peers: [fleet01, ...] config list, so the operator declares who exists rather
    than the daemon guessing;
  • fleet_list.coordinator gains its own mailbox state (pending, consumers) and a peers array
    carrying the same per peer, so "is my peer up and is it reading" becomes answerable;
  • fleet_send{coordId} reworded to say what it did — published to the mailbox — and warning when
    the target queue has zero consumers, which is the observable form of "this will not arrive".

Implementation hazard worth stating up front: in AMQP 0-9-1 a passive declare of a queue that
does not exist closes the channel. Any implementation needs a channel it can afford to lose, or
must reopen after a miss. Getting this wrong turns a peer-status call into a broken publish path.

Not in scope here

#359 is the reason a delivered-but-never-injected message can sit forever; it is filed separately
and should be fixed on its own. This issue is about the missing tools, not that bug.

Built the first real cross-host link on 2026-09-05 (Mac fleet `coordId: mac` ↔ fleet01 `coordId: fleet01`, shared `/coord` vhost). It works. Using it exposed that the MCP surface only covers sending. Today a lead has `fleet_send{coordId}`, `fleet_reply`, and `fleet_list` (which reports its own coord-id). That is the whole surface. ## 1. No peer discovery `fleet_list` returns: ```json "leads":[{"sessionId":"term_65a4c36d11af27","name":"opus","status":"working","self":true}], "coordinator":{"selfId":"mac","configured":true} ``` `fleet01` appears nowhere. Those `leads` rows come from `CallerResolver.leads()`, which resolves **local panes only** — `FleetMcp.java:1093` says as much. So a lead cannot learn that a peer exists, or what coord-id addresses it. I only knew to write `coordId: "fleet01"` because I had configured both daemons myself an hour earlier. Any other session would have to be told out of band, which is exactly what this channel is supposed to replace. ## 2. No way to read your own mailbox Receiving is push-only: `LeadCoordLoop` consumes and types into the pane. There is no `fleet_poll` equivalent for coordination messages, and nothing reports how many are waiting. To find out whether a message I sent was delivered or still queued I had to `ssh` to the other host and run `docker exec fleet-lavinmq lavinmqctl list_queues -p coord`, then read journald. None of that is reachable through MCP, and a lead on a host with no broker shell access is blind to its own inbox. ## 3. `fleet_send{coordId}` reports success for a message the peer may never see `FleetMcp.java:721-727`: ```java leadChannel.publish(coordId, msg); ... return text("delivered to peer lead " + coordId + " (msgId " + msg.msgId() + ")"); ``` "Delivered" means **published to the queue**. It does not mean the peer's pane received it. Those two came apart on the first real message: `fleet_send` reported success and the message then sat undelivered through three daemon restarts, because of the duplicate-lead-tab bug (#359). With #359 present the gap is permanent — the peer never sees the message and the sender is told "delivered" every time. This is the same shape as #359 and #360: a green result for a control that did not take effect. ## Suggested direction (candidate, not a decided design) One mechanism may cover all three. An AMQP **passive queue declare** on `lead.<coordId>.inbox` returns both `message_count` and `consumer_count`. I confirmed this from a client during the build: ``` lead.fleet01.inbox visible: messages=0 consumers=1 ``` `consumers=1` means a daemon is attached and consuming; `messages=N` is the backlog. That is real information about a peer, obtained with no new infrastructure and no presence protocol. Sketch: - a `coordinator.peers: [fleet01, ...]` config list, so the operator declares who exists rather than the daemon guessing; - `fleet_list.coordinator` gains its own mailbox state (`pending`, `consumers`) and a `peers` array carrying the same per peer, so "is my peer up and is it reading" becomes answerable; - `fleet_send{coordId}` reworded to say what it did — published to the mailbox — and warning when the target queue has **zero consumers**, which is the observable form of "this will not arrive". **Implementation hazard worth stating up front:** in AMQP 0-9-1 a passive declare of a queue that does not exist **closes the channel**. Any implementation needs a channel it can afford to lose, or must reopen after a miss. Getting this wrong turns a peer-status call into a broken publish path. ## Not in scope here #359 is the reason a delivered-but-never-injected message can sit forever; it is filed separately and should be fixed on its own. This issue is about the missing tools, not that bug.
Author
Owner

Correction to gap 3 in the description. I wrote that with #359 present "the gap is permanent — the peer never sees the message". That is wrong, and an architect review caught it.

LeadCoordLoop.resolveLocalLead() (LeadCoordLoop.java:174-197) returns null and logs a warning that names the fix:

lead coordination: {} leads are known and none is named "{}" — cannot tell which pane a peer
message is for; name one lead after coordinator.selfId to fix this

tick() then leaves the message unacked, so the broker keeps holding it. Once a lead is named after coordinator.selfId — or the extra lead tabs go away — it is delivered. That matches what actually happened here: the first message sat through three daemon restarts and was delivered after the stale herdr tabs were archived.

So the accurate statement is: it stalls loudly and then recovers. It is not silent, and it is not data loss.

The gap this issue is about is unchanged and still real: fleet_send{coordId} says "delivered to peer lead" when it has only published, so the sender cannot tell a stall from a delivery. The fix is the honest wording plus the consumer count, not a claim about permanence.

**Correction to gap 3 in the description.** I wrote that with #359 present "the gap is permanent — the peer never sees the message". That is wrong, and an architect review caught it. `LeadCoordLoop.resolveLocalLead()` (`LeadCoordLoop.java:174-197`) returns `null` and logs a warning that names the fix: ``` lead coordination: {} leads are known and none is named "{}" — cannot tell which pane a peer message is for; name one lead after coordinator.selfId to fix this ``` `tick()` then leaves the message **unacked**, so the broker keeps holding it. Once a lead is named after `coordinator.selfId` — or the extra lead tabs go away — it is delivered. That matches what actually happened here: the first message sat through three daemon restarts and *was* delivered after the stale herdr tabs were archived. So the accurate statement is: it **stalls loudly and then recovers**. It is not silent, and it is not data loss. The gap this issue is about is unchanged and still real: `fleet_send{coordId}` says "delivered to peer lead" when it has only published, so the sender cannot tell a stall from a delivery. The fix is the honest wording plus the consumer count, not a claim about permanence.
ltms closed this issue 2026-09-05 08:21:33 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#361