Features: a usage limit names the model to turn off; the detector is hot

fleetd #446, merged as 1fb6176. One new entry:

- exhaustedPattern is now a hot config key, so detection can be armed without a
  daemon restart. It used to be compiled once into a startup map.
- the quarantine warning names the fix: which model that profile runs, and the
  exact edit to make.
- fleet_profiles reports model and reason on a quarantined row.

Also corrects the last gotcha on the fleetd #439 entry. It said the compat
overloads default to showing the coordinator row and pointed at #463 as open.
#463 shipped: the default is now false, so a forgotten argument means a missing
row rather than a silent disclosure. The count is corrected too -- listFleet has
seven declarations and six inherit the default; the old text said six overloads
without saying one of them writes it.

Both entries name what is still unproven: nothing tests that main() wires the
quarantine sink at all (fleetd #460).
Dai Ha
2026-09-10 19:18:56 +07:00
parent fa326b42de
commit 8155fd601e
+58 -4
@@ -5262,9 +5262,63 @@ client that tests `if (coordinator)` behave differently from one that tests
- **A worker cannot tell "no coordination configured" from "not for you".** Both look like a
missing key. That is the intended trade: the alternative leaks the fact that peers exist.
- **The lead sees no change.** Same key, same fields.
- **The compat overloads still default to showing the row.** `listFleet` has six overloads that do
not take the caller's role, and they inherit `true` from one line. Nothing is exposed today —
the one production call site passes the real answer — but a new call site that forgets the
argument would disclose the row silently. Tracked in fleetd #463.
- **The compat overloads used to default to showing the row, and no longer do.** `listFleet` has
seven declarations. Six of them do not take the caller's role, and they all inherited the
default from one line. That default was `true`, so a call site that forgot the argument would
have disclosed the row silently. It is now `false`: a missing identity means a missing row,
which is a visible bug rather than a quiet disclosure. Fixed in fleetd #463. A test calls a
compat overload with no boolean and asserts the `coordinator` key is absent.
fleetd #439.
## A usage limit now says which model to turn off, and detection can be armed without a restart
**What.** Three related changes to how fleetd reacts when a backend reports that a subscription
limit is reached.
- `exhaustedPattern` is now a **hot** config key. Before, it was compiled once into a map at
daemon startup, so arming detection for a profile needed a restart. Now the pattern is read
live per profile, and compiled patterns are cached by pattern string so a check does not
recompile a regex every time.
- The warning logged when a credential is quarantined **names the fix**, not just the fact:
```
usage-limit fix: profile 'X' runs model 'Y' — set `enabled: false` on that model's entry
under models.allow in fleetd.yaml to stop new spawns landing on it (models: is hot, no
restart needed); remove the line again once the subscription window resets
```
- `fleet_profiles` reports two new fields on a quarantined row: `model` (the model that profile
runs) and `reason` (the backend text that triggered the most recent quarantine of that
credential).
**The knob.** `exhaustedPattern:` on a profile, which is opt-in and off by default. It is now hot,
so a reload arms it. Turning a model off is `enabled: false` on its `models.allow` entry, which
was already hot.
**Why it exists.** The two halves of model gating were reloadable in opposite directions, and it
was the wrong way round. An operator could already turn a model off at runtime, but could not
arm the detector that tells them to — that needed a restart, on a feature whose whole point is
reacting to a limit while the fleet is running. The `reason` field exists because "this
credential is quarantined" does not say whether the cause was a usage limit or something else,
and the two need different responses. The warning names the fix because an operator reading a
quarantine line should not have to work out which of several configured model names to edit.
**Gotchas.**
- **`reason` is only written where a quarantine actually happens**, keyed by credential id — the
same key the cooldown itself uses. So a reason can never be reported for a quarantine that did
not occur.
- **A pattern has to match real text from that specific backend.** Do not guess one. A guessed
regex gives you a profile that reports `exhaustionDetectionArmed: true` and silently never
fires, which is worse than an honest `false`.
- **There is no `off:` config key.** The flip is `enabled: false`. `off` is the name of the
*reported* set (`modelsOff`), which is an output.
- **Nothing re-enables a model automatically.** "On again when the limit lifts" is a manual edit,
which is cheap because `models:` is hot and needs no restart. An automatic backoff probe was
considered and rejected: a probe spends quota to discover quota, so on a metered subscription
it burns the first tokens of every new window on discovery instead of work.
- **Nothing proves the daemon wires this sink.** The quarantine sink was extracted into a factory
and what it logs is pinned by a test, but replacing `main()`'s call to that factory with an
inert lambda leaves the whole suite green. Tracked in fleetd #460.
fleetd #446.