diff --git a/11-Features.md b/11-Features.md index 525e42e..1d9bffe 100644 --- a/11-Features.md +++ b/11-Features.md @@ -4473,3 +4473,59 @@ now does, and `aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket` `MessageService` now carries six test-only hooks, one per race of this kind. Each exists because its window is unreachable through the public API — which is also why each bug was invisible — but six is enough that the next one needs a harder look than "the file already does this". + +## `fleet_list` reports lead coordination state instead of guessing at it + +**What it does.** A lead's `fleet_list` now answers three questions it could not answer before: is +my own coordination mailbox there and is anyone reading it, what is waiting in it, and what is the +state of each peer daemon I know about. Each mailbox row carries a `status` of `exists`, `absent` +or `unknown`. `pending` and `consumers` appear **only** when `status` is `exists`. + +The reading is a passive AMQP queue declare, not a presence protocol. `consumers: 1` means a daemon +is attached and consuming; `pending: N` is the backlog. + +`fleet_send{coordId}` also stopped saying "delivered". It now says the message was published and +durably confirmed by the broker, which is what a publisher confirm actually proves. When the target +mailbox exists but has zero consumers, the (still successful) result adds a warning that nobody is +reading it right now. + +**On.** Add `peers:` to the `coordinator:` block — the operator declares who exists, because the +daemon never guesses: + +```yaml +coordinator: + uriEnv: LEAD_COORD_URI + selfId: mac + peers: [fleet01] +``` + +Omit it and you still get your own mailbox row. An undeclared peer can still reach you and be +reached by `fleet_send`; it simply does not get a row. Like the rest of the `coordinator:` block, +`peers` is read once at boot, so a change needs a daemon restart. + +**Why it exists.** The first real cross-host link (Mac ↔ fleet01, 2026-09-05) worked, and using it +showed the MCP surface only covered sending. A lead could not learn that a peer existed, could not +read its own inbox, and was told "delivered" for a message the peer never saw — that one sat +undelivered through three daemon restarts because of the duplicate-lead-tab bug (#359). Finding +that out needed an `ssh` to the other host and `lavinmqctl list_queues`. None of it was reachable +through MCP, and a lead on a host with no broker shell was blind to its own inbox. fleetd #361. + +**One thing to know for maintenance.** The three-state `status` is the whole point, and it is easy +to collapse back into a boolean. The first cut of this feature did exactly that: one `absent()` +value stood for both "the broker said there is no such queue" and "I could not check", so a +self-probe timeout rendered as `pending: 0, consumers: 0` — indistinguishable from a mailbox that +is genuinely empty and genuinely unread. That is the same overstatement the ticket exists to fix, +one level down. + +`LeadMailbox.isMissingQueue` is the discriminator that keeps them apart, and it is narrow on +purpose: only an `IOException` whose cause is a `ShutdownSignalException` carrying an +`AMQP.Channel.Close` with reply code 404 counts as a confirmed absence. Measured on merge: making +it return `true` unconditionally restored the original defect and left 1389 tests green. It now has +five tests of its own, and the same mutation gives 4 failures. + +Two other things this feature depends on, both easy to undo by accident. `inspect` runs its passive +declare on a **throwaway** channel, because in AMQP 0-9-1 a passive declare of a missing queue +closes the channel it ran on — reusing the publish channel would let one miss break every later +publish on that instance. And `FleetMcp.probe` cancels a timed-out probe rather than abandoning it; +without that, a hung (not down) broker would orphan one channel per `fleet_list` call until the +connection's channel-max ran out, breaking publish by a different route.