Features: fleet health reaches its fault states (CB-640/641/643) + the /fleets-status skill (CB-642)

Dai Ha
2026-08-27 22:16:44 +07:00
parent 3357960dd1
commit 226794a4fb
+67
@@ -2106,6 +2106,73 @@ coord-id it was already told.
---
## Fleet health can finally reach its fault states
**What.** `FleetHealthMonitor` ticks on its own timer, well away from the 250ms delivery poller. Each
tick makes one `agent.list` call plus one roster snapshot, then classifies every member through the
pure `FleetHealth.decide`. There are 14 health states and 9 of them are faults. A fault is logged
**when the state changes**, never once per tick. Two terminal states, `GONE` and `NEVER_READY`, also
call `failTarget` (`MessageService.abandon`), which fails every ticket still waiting on that member.
**On.** The `health:` block in `fleetd.yaml`: `enabled: true`, `intervalSeconds: 30` (floor 15),
`workingSuspectAfterSeconds: 600` (floor 300 — how long a `BUSY` member with no activity must sit
before it is `STALL_SUSPECTED`), and `notifications.mode`. `fleet_list` reports `healthCoverage`:
`off`, `detection-only`, or `full`. It is `full` only for `mode: webhook`.
**Why.** The monitor shipped half-wired and nothing said so. `tick` passed a hardcoded
`NOT_YET_OBSERVED = false` for 7 of the 12 `HealthSnapshot` fields, so **8 of the 9 fault states
could never be reached** — only `TURN_BOUNDARY_LOST` could fire. `GONE` and `NEVER_READY` were among
the dead ones, which means CB-580's repair never ran: a ticket waiting on a member that had vanished
still hung to the 30-minute async timeout. Everything compiled and every test passed, because each
test called the classifier directly and walked around the caller. Three units fixed it: CB-640 added
the message-layer facts (`hasQueuedDelivery`, `hasStrandedReply`, `hasOrphanedDelegation`), CB-641
added the herdr and clock facts (`present`, `targetNotFound`, `controlLinkDown`,
`readinessGraceElapsed`, `stalled`), and CB-643 joined them and deleted the placeholder. The rule
that came out of it: **a snapshot field with no publisher is a dead state, and nothing else will tell
you.**
**Gotcha.** Three things to know before you read the log.
- With `notifications.mode: disabled` the coverage is `detection-only`. Every fault the monitor finds
goes to the log and nowhere else. The daemon prints a warning about this at startup.
- A member that is already faulty when the daemon starts logs its fault **once**, on the first tick,
and then stays quiet. Grepping recent lines can therefore show nothing while the fault is live.
- `DELEGATION_ORPHANED` needs the fact to hold for **two consecutive ticks**. An async ticket exists
for a moment before its virtual thread opens the rendezvous waiter, so for that instant nothing is
accepted or queued behind it and a single tick would report a fault that clears immediately. The
gate costs one interval of delay on a real orphan.
---
## `/fleets-status` — every fleet that shares one broker
**What.** A primary-side skill in `.claude/skills/fleets-status/`. It reports the state of every
fleet that shares one LavinMQ instance, in three tiers: (1) the local daemon — process, deployed jar
against `HEAD`, `/healthz`, `fleet_list` with capacity and `healthCoverage`, and the WARN/ERROR lines
since the last `fleetd listening` marker; (2) the broker itself over its management API — each
vhost's queues, depths and consumers; (3) cross-fleet lead coordination, read from the
`lead.<coordId>.inbox` queues. Every tier ends with what it **cannot** see, and a fleet it cannot
reach appears in the report as `unknown` with the reason, never as an omission.
**On.** Type `/fleets-status`. Tier 1 always runs. Tier 2 runs only when
`LAVINMQ_MANAGEMENT_USER` and `LAVINMQ_MANAGEMENT_PASSWORD` resolve in a login shell.
**Why.** There was no way to ask one question about the whole fleet estate. A lead could see its own
daemon and nothing else, and the remote daemon's REST port is not reachable from here. The broker is
the one thing both fleets touch, so its management API answers for a host you cannot log in to. The
sharpest fact it gives you: **a vhost with queues but zero consumers means that fleet's daemon is
down while its durable state survives.**
**Gotcha.** Tier 2 is blocked today. The AMQP user in `LAVINMQ_URI` connects fine on 5672 but gets
HTTP 401 from the management API on 15672 — an AMQP connection is not monitoring access. An operator
must create a separate read-only user with the `monitoring` tag, allowed on both vhosts, and store it
under the two names above in `${SHARED_ENV}/tools/secrets.sh`. Also: the skill forbids printing
`LAVINMQ_URI`, because the password is inside it. Output that could carry it is redacted with
`sed -E 's#://[^@]*@#://<redacted>@#g'`, and the `g` is not optional — without it a line holding two
URIs leaks the second one, which `scripts/redeploy-fleetd.sh --check` prints.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from