Features: the charter tool check now covers reload, and the quarantine row reports the attempt

fleetd #474 removed the limit this page warned about. The charter
tool-surface check ran at startup only, so a reload could install a
charter naming a tool the server does not register. ConfigRef.reload()
now runs the same check and refuses the whole reload. Replaced the
"until #474 lands" warning, which was about to tell a reader they could
not rely on something they now can.

fleetd #473 added quarantineAttempt beside quarantinedForSeconds in both
fleet_profiles and fleet_list capacity rows. Recorded what a large count
with a flat cooldown means (normal -- the ceiling caps the wait, not the
streak) and that both values come from one map read.
Dai Ha
2026-09-10 20:43:35 +07:00
parent 14798b39e7
commit 363116648b
+31 -7
@@ -1001,13 +1001,21 @@ out of each charter and asks `FleetTool`, the one enum the MCP server derives it
from, whether that tool exists. If one does not, the daemon refuses to start and the message names
both the charter key and the unknown tool.
**Know the limit of that check: it runs at startup only (fleetd #474).** The paragraph above says
validation runs on both the startup and the reload path, and that is true of the shape checks. It
is **not** yet true of the tool-name check, which has one call site, in `Fleetd.main`. Charters are
hot and are read live at each spawn, so editing `fleet.charters:` on a running daemon to name a tool
that does not exist is accepted, applied, and delivered to the next member. A restart would refuse
the same file. Until #474 lands, treat a charter edit made without a restart as unchecked, and read
the startup log after the next restart to find out whether it was valid.
**A reload is checked too, not only startup (fleetd #474).** For a short while this check had one
call site, in `Fleetd.main`, so editing `fleet.charters:` on a running daemon to name a tool that
does not exist was accepted, applied and delivered to the next member — while a restart would have
refused the very same file. Charters are hot and are read live at each spawn, so that gap was
reachable. `ConfigRef.reload()` now runs the same check, and a reload that fails it is refused whole
and keeps the running config, exactly like any other validation failure. The daemon logs
`config reload from <path> refused, keeping the running config: <message>`, and the message names
both the charter key and the unknown tool.
*Why it was built this way:* the check cannot live inside `FleetConfig`, because the canonical tool
set is in the `mcp` package and config is loaded before the MCP server exists. So `ConfigRef` takes
it as a `Consumer<FleetConfig>` and `Fleetd.main` supplies it — `Fleetd` is the one seam that
already holds both a loaded config and the `mcp` package. *The gotcha:* nothing yet stops someone
reverting `Fleetd.main` to the two-argument `ConfigRef` constructor. That one edit turns this gate
off with the whole test suite green, which is why the wiring is worth reading before you trust it.
**Do not put secrets in charter text.** There is deliberately no `${ENV}` interpolation. The
OpenCode adapter writes the composed charter to a temp file so its CLI can read it, and that file is
@@ -1674,6 +1682,22 @@ difference.
fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares
its `QuarantineSource` with `fleet_profiles` so the two surfaces cannot disagree.
**The row also says which attempt this is, not only how long is left (fleetd #466, #473).** Both
`fleet_profiles` and `fleet_list`'s capacity rows carry `quarantineAttempt` beside
`quarantinedForSeconds`: `1` for a first refusal, `2` for the second in a row, and so on. The
cooldown escalates with that count, so `quarantinedForSeconds: 1800, quarantineAttempt: 1` and
`quarantinedForSeconds: 7200, quarantineAttempt: 3` mean very different things about the credential.
*Why it exists:* the seconds alone cannot tell a weekly subscription limit apart from a one-off
capacity blip. A lead seeing attempt 4 knows retrying is pointless and should move the work to
another credential, or tell the operator. *Gotcha:* the count keeps growing past the cooldown
ceiling, so a large `quarantineAttempt` with a flat `quarantinedForSeconds` is normal, not a bug —
the ceiling caps the wait, not the streak. A quiet gap long enough to clear the quarantine resets
the count to 1. The flat two-argument `BackendQuarantine` constructor still reports a real growing
count even though its own cooldown never escalates, because the streak is a fact either way. Both
values come from **one** `status()` call that does a single map read, so the reported count can
never disagree with the cooldown it describes.
---
## Resume a member onto its previous conversation