Features: the exhaustion quarantine now escalates (fleetd #466)

The 'Stop spawning onto an exhausted account' entry described a flat
cooldown, which stopped being true with fleetd #466. Updated rather than
duplicated: the cooldown now doubles on each consecutive exhaustion of the
same credential, capped at 12x the base, so a chronically exhausted
credential is retried about a dozen times a week instead of ~336.

quarantineCooldownSeconds is now the BASE of the backoff, and the entry
says so where an operator reads what the key does.

Four gotchas added, all of them things an operator can get wrong:

- The reset is a TIME PROXY. Nothing reports a successful spawn back to the
  tracker, so what clears the streak is a base cooldown of quiet, not
  evidence the account recovered. Read it as 'we have not been told it is
  still broken'.
- The multiplier and ceiling are constants, not YAML. No new knob, no new
  hot/cold question.
- The ceiling is deliberate: an unbounded backoff is a permanent outage
  only a restart clears, which is worse than the flat retrying it replaced.
- Cooling-off is a SEPARATE mechanism and is not escalated, because
  escalating a 5xx storm would turn it into a multi-hour outage.

Entry count unchanged at 157 - this updates an existing entry.
Dai Ha
2026-09-10 19:54:54 +07:00
parent 8155fd601e
commit 3dde75ebca
+23 -1
@@ -1262,8 +1262,13 @@ message naming the profile, the credential and roughly how many seconds are left
under every policy; and `fleet_profiles` shows it. The quarantine lifts itself — there is no manual
step.
Since fleetd #466 the cooldown **escalates**. Each consecutive exhaustion of the same credential
doubles the wait, capped at 12x the base — about 6 hours at the 1800s default. A credential that
keeps reporting exhausted is therefore retried roughly a dozen times a week instead of about 336
times.
**On.** `quarantineCooldownSeconds:` at the top level (default 1800, **deferred** — it is baked into
the tracker at startup). Per profile, `credentialId:` (**hot**) says which credential this profile
the tracker at startup). Since fleetd #466 it is the **base** of the backoff, not the whole of it. Per profile, `credentialId:` (**hot**) says which credential this profile
spends. Quarantine itself only ever fires for a profile that has an `exhaustedPattern`, so a fleet
with no patterns configured behaves exactly as before.
@@ -1278,6 +1283,23 @@ straight onto the same dead account under the other name. Profiles that share an
`credentialId`. A profile that sets none quarantines alone, under its own name — safe, but it will not
protect a sibling.
**More gotchas, from the escalation (fleetd #466).**
- **The reset is a time proxy, not a success signal.** Nothing in the daemon reports a *successful*
spawn back to the quarantine tracker, so "the account started working again" cannot be observed
there. What clears the streak is a base cooldown's worth of quiet — no further exhaustion report
for that credential. That is the best available evidence, not proof. Read it as "we have not been
told it is still broken", never as "it is fixed".
- **The multiplier and the ceiling are constants, not config.** Doubling (2.0) and the 12x cap live
in `BackendQuarantine`, so there is no YAML knob for either and no new hot/cold question.
`quarantineCooldownSeconds` stays deferred.
- **There is a ceiling on purpose.** An unbounded backoff becomes a permanent outage that only a
restart clears, which would be worse than the flat-rate retrying it replaced.
- **Cooling-off is a different mechanism and is NOT escalated.** A credential that throws repeated
*non-exhaustion* errors (an HTTP 5xx storm) gets `BackendOutagePolicy`'s flat 60s, with no repeat
tracking. Escalating that would turn a transient storm into a multi-hour outage. The two states
are reported separately and a profile can be in both at once.
---
## See which charter a member got