From 226794a4fb03fcf44fb6d610758f47b14a2939d1 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Thu, 27 Aug 2026 22:16:44 +0700 Subject: [PATCH] Features: fleet health reaches its fault states (CB-640/641/643) + the /fleets-status skill (CB-642) --- 11-Features.md | 67 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 67 insertions(+) diff --git a/11-Features.md b/11-Features.md index a6fb2b0..a35de78 100644 --- a/11-Features.md +++ b/11-Features.md @@ -2106,6 +2106,73 @@ coord-id it was already told. --- +## Fleet health can finally reach its fault states + +**What.** `FleetHealthMonitor` ticks on its own timer, well away from the 250ms delivery poller. Each +tick makes one `agent.list` call plus one roster snapshot, then classifies every member through the +pure `FleetHealth.decide`. There are 14 health states and 9 of them are faults. A fault is logged +**when the state changes**, never once per tick. Two terminal states, `GONE` and `NEVER_READY`, also +call `failTarget` (`MessageService.abandon`), which fails every ticket still waiting on that member. + +**On.** The `health:` block in `fleetd.yaml`: `enabled: true`, `intervalSeconds: 30` (floor 15), +`workingSuspectAfterSeconds: 600` (floor 300 — how long a `BUSY` member with no activity must sit +before it is `STALL_SUSPECTED`), and `notifications.mode`. `fleet_list` reports `healthCoverage`: +`off`, `detection-only`, or `full`. It is `full` only for `mode: webhook`. + +**Why.** The monitor shipped half-wired and nothing said so. `tick` passed a hardcoded +`NOT_YET_OBSERVED = false` for 7 of the 12 `HealthSnapshot` fields, so **8 of the 9 fault states +could never be reached** — only `TURN_BOUNDARY_LOST` could fire. `GONE` and `NEVER_READY` were among +the dead ones, which means CB-580's repair never ran: a ticket waiting on a member that had vanished +still hung to the 30-minute async timeout. Everything compiled and every test passed, because each +test called the classifier directly and walked around the caller. Three units fixed it: CB-640 added +the message-layer facts (`hasQueuedDelivery`, `hasStrandedReply`, `hasOrphanedDelegation`), CB-641 +added the herdr and clock facts (`present`, `targetNotFound`, `controlLinkDown`, +`readinessGraceElapsed`, `stalled`), and CB-643 joined them and deleted the placeholder. The rule +that came out of it: **a snapshot field with no publisher is a dead state, and nothing else will tell +you.** + +**Gotcha.** Three things to know before you read the log. + +- With `notifications.mode: disabled` the coverage is `detection-only`. Every fault the monitor finds + goes to the log and nowhere else. The daemon prints a warning about this at startup. +- A member that is already faulty when the daemon starts logs its fault **once**, on the first tick, + and then stays quiet. Grepping recent lines can therefore show nothing while the fault is live. +- `DELEGATION_ORPHANED` needs the fact to hold for **two consecutive ticks**. An async ticket exists + for a moment before its virtual thread opens the rendezvous waiter, so for that instant nothing is + accepted or queued behind it and a single tick would report a fault that clears immediately. The + gate costs one interval of delay on a real orphan. + +--- + +## `/fleets-status` — every fleet that shares one broker + +**What.** A primary-side skill in `.claude/skills/fleets-status/`. It reports the state of every +fleet that shares one LavinMQ instance, in three tiers: (1) the local daemon — process, deployed jar +against `HEAD`, `/healthz`, `fleet_list` with capacity and `healthCoverage`, and the WARN/ERROR lines +since the last `fleetd listening` marker; (2) the broker itself over its management API — each +vhost's queues, depths and consumers; (3) cross-fleet lead coordination, read from the +`lead..inbox` queues. Every tier ends with what it **cannot** see, and a fleet it cannot +reach appears in the report as `unknown` with the reason, never as an omission. + +**On.** Type `/fleets-status`. Tier 1 always runs. Tier 2 runs only when +`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` resolve in a login shell. + +**Why.** There was no way to ask one question about the whole fleet estate. A lead could see its own +daemon and nothing else, and the remote daemon's REST port is not reachable from here. The broker is +the one thing both fleets touch, so its management API answers for a host you cannot log in to. The +sharpest fact it gives you: **a vhost with queues but zero consumers means that fleet's daemon is +down while its durable state survives.** + +**Gotcha.** Tier 2 is blocked today. The AMQP user in `LAVINMQ_URI` connects fine on 5672 but gets +HTTP 401 from the management API on 15672 — an AMQP connection is not monitoring access. An operator +must create a separate read-only user with the `monitoring` tag, allowed on both vhosts, and store it +under the two names above in `${SHARED_ENV}/tools/secrets.sh`. Also: the skill forbids printing +`LAVINMQ_URI`, because the password is inside it. Output that could carry it is redacted with +`sed -E 's#://[^@]*@#://@#g'`, and the `g` is not optional — without it a line holding two +URIs leaks the second one, which `scripts/redeploy-fleetd.sh --check` prints. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from