fleetd #361: fleet_list reports lead coordination state

Dai Ha
2026-09-05 13:22:05 +07:00
parent 26a7ea41f7
commit 3a4bd751e7
+56
@@ -4473,3 +4473,59 @@ now does, and `aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket`
`MessageService` now carries six test-only hooks, one per race of this kind. Each exists because its
window is unreachable through the public API — which is also why each bug was invisible — but six is
enough that the next one needs a harder look than "the file already does this".
## `fleet_list` reports lead coordination state instead of guessing at it
**What it does.** A lead's `fleet_list` now answers three questions it could not answer before: is
my own coordination mailbox there and is anyone reading it, what is waiting in it, and what is the
state of each peer daemon I know about. Each mailbox row carries a `status` of `exists`, `absent`
or `unknown`. `pending` and `consumers` appear **only** when `status` is `exists`.
The reading is a passive AMQP queue declare, not a presence protocol. `consumers: 1` means a daemon
is attached and consuming; `pending: N` is the backlog.
`fleet_send{coordId}` also stopped saying "delivered". It now says the message was published and
durably confirmed by the broker, which is what a publisher confirm actually proves. When the target
mailbox exists but has zero consumers, the (still successful) result adds a warning that nobody is
reading it right now.
**On.** Add `peers:` to the `coordinator:` block — the operator declares who exists, because the
daemon never guesses:
```yaml
coordinator:
uriEnv: LEAD_COORD_URI
selfId: mac
peers: [fleet01]
```
Omit it and you still get your own mailbox row. An undeclared peer can still reach you and be
reached by `fleet_send`; it simply does not get a row. Like the rest of the `coordinator:` block,
`peers` is read once at boot, so a change needs a daemon restart.
**Why it exists.** The first real cross-host link (Mac ↔ fleet01, 2026-09-05) worked, and using it
showed the MCP surface only covered sending. A lead could not learn that a peer existed, could not
read its own inbox, and was told "delivered" for a message the peer never saw — that one sat
undelivered through three daemon restarts because of the duplicate-lead-tab bug (#359). Finding
that out needed an `ssh` to the other host and `lavinmqctl list_queues`. None of it was reachable
through MCP, and a lead on a host with no broker shell was blind to its own inbox. fleetd #361.
**One thing to know for maintenance.** The three-state `status` is the whole point, and it is easy
to collapse back into a boolean. The first cut of this feature did exactly that: one `absent()`
value stood for both "the broker said there is no such queue" and "I could not check", so a
self-probe timeout rendered as `pending: 0, consumers: 0` — indistinguishable from a mailbox that
is genuinely empty and genuinely unread. That is the same overstatement the ticket exists to fix,
one level down.
`LeadMailbox.isMissingQueue` is the discriminator that keeps them apart, and it is narrow on
purpose: only an `IOException` whose cause is a `ShutdownSignalException` carrying an
`AMQP.Channel.Close` with reply code 404 counts as a confirmed absence. Measured on merge: making
it return `true` unconditionally restored the original defect and left 1389 tests green. It now has
five tests of its own, and the same mutation gives 4 failures.
Two other things this feature depends on, both easy to undo by accident. `inspect` runs its passive
declare on a **throwaway** channel, because in AMQP 0-9-1 a passive declare of a missing queue
closes the channel it ran on — reusing the publish channel would let one miss break every later
publish on that instance. And `FleetMcp.probe` cancels a timed-out probe rather than abandoning it;
without that, a hung (not down) broker would orphan one channel per `fleet_list` call until the
connection's channel-max ran out, breaking publish by a different route.