Features: fleet health reaches its fault states (CB-640/641/643) + the /fleets-status skill (CB-642)
+67
@@ -2106,6 +2106,73 @@ coord-id it was already told.
|
||||
|
||||
---
|
||||
|
||||
## Fleet health can finally reach its fault states
|
||||
|
||||
**What.** `FleetHealthMonitor` ticks on its own timer, well away from the 250ms delivery poller. Each
|
||||
tick makes one `agent.list` call plus one roster snapshot, then classifies every member through the
|
||||
pure `FleetHealth.decide`. There are 14 health states and 9 of them are faults. A fault is logged
|
||||
**when the state changes**, never once per tick. Two terminal states, `GONE` and `NEVER_READY`, also
|
||||
call `failTarget` (`MessageService.abandon`), which fails every ticket still waiting on that member.
|
||||
|
||||
**On.** The `health:` block in `fleetd.yaml`: `enabled: true`, `intervalSeconds: 30` (floor 15),
|
||||
`workingSuspectAfterSeconds: 600` (floor 300 — how long a `BUSY` member with no activity must sit
|
||||
before it is `STALL_SUSPECTED`), and `notifications.mode`. `fleet_list` reports `healthCoverage`:
|
||||
`off`, `detection-only`, or `full`. It is `full` only for `mode: webhook`.
|
||||
|
||||
**Why.** The monitor shipped half-wired and nothing said so. `tick` passed a hardcoded
|
||||
`NOT_YET_OBSERVED = false` for 7 of the 12 `HealthSnapshot` fields, so **8 of the 9 fault states
|
||||
could never be reached** — only `TURN_BOUNDARY_LOST` could fire. `GONE` and `NEVER_READY` were among
|
||||
the dead ones, which means CB-580's repair never ran: a ticket waiting on a member that had vanished
|
||||
still hung to the 30-minute async timeout. Everything compiled and every test passed, because each
|
||||
test called the classifier directly and walked around the caller. Three units fixed it: CB-640 added
|
||||
the message-layer facts (`hasQueuedDelivery`, `hasStrandedReply`, `hasOrphanedDelegation`), CB-641
|
||||
added the herdr and clock facts (`present`, `targetNotFound`, `controlLinkDown`,
|
||||
`readinessGraceElapsed`, `stalled`), and CB-643 joined them and deleted the placeholder. The rule
|
||||
that came out of it: **a snapshot field with no publisher is a dead state, and nothing else will tell
|
||||
you.**
|
||||
|
||||
**Gotcha.** Three things to know before you read the log.
|
||||
|
||||
- With `notifications.mode: disabled` the coverage is `detection-only`. Every fault the monitor finds
|
||||
goes to the log and nowhere else. The daemon prints a warning about this at startup.
|
||||
- A member that is already faulty when the daemon starts logs its fault **once**, on the first tick,
|
||||
and then stays quiet. Grepping recent lines can therefore show nothing while the fault is live.
|
||||
- `DELEGATION_ORPHANED` needs the fact to hold for **two consecutive ticks**. An async ticket exists
|
||||
for a moment before its virtual thread opens the rendezvous waiter, so for that instant nothing is
|
||||
accepted or queued behind it and a single tick would report a fault that clears immediately. The
|
||||
gate costs one interval of delay on a real orphan.
|
||||
|
||||
---
|
||||
|
||||
## `/fleets-status` — every fleet that shares one broker
|
||||
|
||||
**What.** A primary-side skill in `.claude/skills/fleets-status/`. It reports the state of every
|
||||
fleet that shares one LavinMQ instance, in three tiers: (1) the local daemon — process, deployed jar
|
||||
against `HEAD`, `/healthz`, `fleet_list` with capacity and `healthCoverage`, and the WARN/ERROR lines
|
||||
since the last `fleetd listening` marker; (2) the broker itself over its management API — each
|
||||
vhost's queues, depths and consumers; (3) cross-fleet lead coordination, read from the
|
||||
`lead.<coordId>.inbox` queues. Every tier ends with what it **cannot** see, and a fleet it cannot
|
||||
reach appears in the report as `unknown` with the reason, never as an omission.
|
||||
|
||||
**On.** Type `/fleets-status`. Tier 1 always runs. Tier 2 runs only when
|
||||
`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` resolve in a login shell.
|
||||
|
||||
**Why.** There was no way to ask one question about the whole fleet estate. A lead could see its own
|
||||
daemon and nothing else, and the remote daemon's REST port is not reachable from here. The broker is
|
||||
the one thing both fleets touch, so its management API answers for a host you cannot log in to. The
|
||||
sharpest fact it gives you: **a vhost with queues but zero consumers means that fleet's daemon is
|
||||
down while its durable state survives.**
|
||||
|
||||
**Gotcha.** Tier 2 is blocked today. The AMQP user in `LAVINMQ_URI` connects fine on 5672 but gets
|
||||
HTTP 401 from the management API on 15672 — an AMQP connection is not monitoring access. An operator
|
||||
must create a separate read-only user with the `monitoring` tag, allowed on both vhosts, and store it
|
||||
under the two names above in `${SHARED_ENV}/tools/secrets.sh`. Also: the skill forbids printing
|
||||
`LAVINMQ_URI`, because the password is inside it. Output that could carry it is redacted with
|
||||
`sed -E 's#://[^@]*@#://<redacted>@#g'`, and the `g` is not optional — without it a line holding two
|
||||
URIs leaks the second one, which `scripts/redeploy-fleetd.sh --check` prints.
|
||||
|
||||
---
|
||||
|
||||
## Backfill status
|
||||
|
||||
This page was started after the fact, so it is **not yet complete**. Entries above are written from
|
||||
|
||||
Reference in New Issue
Block a user