Features: watch the fleet's health (CB-573)

Dai Ha
2026-08-15 06:21:08 +02:00
parent 6eaa19a868
commit 891bd9fc0b
+31
@@ -54,6 +54,7 @@ six weeks, and the table alone will not carry it.
| [See free fleet capacity](#see-free-fleet-capacity) | automatic | CB-573 | `mcp/BridgeMcp` |
| [Answer a question on an async delegation](#answer-a-question-on-an-async-delegation) | automatic | CB-574 | `msg/MessageService` |
| [Learn why a delegation died](#learn-why-a-delegation-died) | automatic | CB-568 | `inject/CompletionResolver` |
| [Watch the fleet's health](#watch-the-fleets-health) | `health:` | CB-573 | `health/FleetHealthMonitor` |
Nearly every knob above lives in one file, on one profile:
@@ -933,6 +934,36 @@ was queued but not yet delivered sat until its timeout, which is 30 minutes on a
---
## Watch the fleet's health
**What.** An opt-in background observer. Each tick it takes **one** `AgentControl.list()` for the whole
fleet and one in-memory roster snapshot, joins them, and classifies every member. A member entering a
fault state logs one WARN; recovering logs one INFO. `bridge_list` reports `healthCoverage`, which is
`off`, `detection-only`, or `full`.
**On.** A `health:` block in `bridged.yaml` with `enabled: true`. `intervalSeconds` defaults to 30 and
is floored at 15. With no block at all nothing is constructed and no herdr call is ever made.
**Why.** A fault is usually a *disagreement between two views*, not a value you can read from one of
them. A member that says `BUSY` in the session FSM while herdr says `DONE` has lost its turn boundary —
a real trace sat in that state for eighteen minutes with its ticket still `PENDING` and no fallback
firing. That is why both views must come from the same instant: one list call per tick, never one per
member. Reading panes is the exception, not the method.
**Gotcha.** Detection and notification are **separate keys** on purpose. An earlier design required a
webhook before `health.enabled` could be turned on, which would have removed real local detection to
avoid a narrower human-notification gap. So health runs with no sink configured and reports
`detection-only` — treat that value as "nobody will be paged", not as "health is off".
Two more. `tick()` catches `Throwable` and reschedules in a `finally`, because a
`ScheduledExecutorService` never re-runs a task that threw: the earlier version rescheduled as its last
statement, so the first `agents.list()` failure would have stopped health permanently and silently —
exactly when the control link is down, the highest-priority state in the model. And the snapshot fields
this unit cannot yet supply are the named constant `NOT_YET_OBSERVED`, not bare `false`, because to this
classifier `false` means "no fault" rather than "not known yet".
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from