diff --git a/11-Features.md b/11-Features.md index 65dbe4a..5d46580 100644 --- a/11-Features.md +++ b/11-Features.md @@ -1262,8 +1262,13 @@ message naming the profile, the credential and roughly how many seconds are left under every policy; and `fleet_profiles` shows it. The quarantine lifts itself — there is no manual step. +Since fleetd #466 the cooldown **escalates**. Each consecutive exhaustion of the same credential +doubles the wait, capped at 12x the base — about 6 hours at the 1800s default. A credential that +keeps reporting exhausted is therefore retried roughly a dozen times a week instead of about 336 +times. + **On.** `quarantineCooldownSeconds:` at the top level (default 1800, **deferred** — it is baked into -the tracker at startup). Per profile, `credentialId:` (**hot**) says which credential this profile +the tracker at startup). Since fleetd #466 it is the **base** of the backoff, not the whole of it. Per profile, `credentialId:` (**hot**) says which credential this profile spends. Quarantine itself only ever fires for a profile that has an `exhaustedPattern`, so a fleet with no patterns configured behaves exactly as before. @@ -1278,6 +1283,23 @@ straight onto the same dead account under the other name. Profiles that share an `credentialId`. A profile that sets none quarantines alone, under its own name — safe, but it will not protect a sibling. +**More gotchas, from the escalation (fleetd #466).** + +- **The reset is a time proxy, not a success signal.** Nothing in the daemon reports a *successful* + spawn back to the quarantine tracker, so "the account started working again" cannot be observed + there. What clears the streak is a base cooldown's worth of quiet — no further exhaustion report + for that credential. That is the best available evidence, not proof. Read it as "we have not been + told it is still broken", never as "it is fixed". +- **The multiplier and the ceiling are constants, not config.** Doubling (2.0) and the 12x cap live + in `BackendQuarantine`, so there is no YAML knob for either and no new hot/cold question. + `quarantineCooldownSeconds` stays deferred. +- **There is a ceiling on purpose.** An unbounded backoff becomes a permanent outage that only a + restart clears, which would be worse than the flat-rate retrying it replaced. +- **Cooling-off is a different mechanism and is NOT escalated.** A credential that throws repeated + *non-exhaustion* errors (an HTTP 5xx storm) gets `BackendOutagePolicy`'s flat 60s, with no repeat + tracking. Escalating that would turn a transient storm into a multi-hour outage. The two states + are reported separately and a profile can be in both at once. + --- ## See which charter a member got