Features: stop spawning onto an exhausted account (CB-578 stage B)

Dai Ha
2026-08-15 10:34:40 +02:00
parent a4fc43f465
commit f4c70b7786
+28
@@ -58,6 +58,7 @@ six weeks, and the table alone will not carry it.
| [Learn why a delegation died](#learn-why-a-delegation-died) | automatic | CB-568 | `inject/CompletionResolver` |
| [Watch the fleet's health](#watch-the-fleets-health) | `health:` | CB-573 | `health/FleetHealthMonitor` |
| [Tell a usage-limit refusal from a real reply](#tell-a-usage-limit-refusal-from-a-real-reply) | profile `exhaustedPattern:` | CB-578 | `inject/CompletionResolver` |
| [Stop spawning onto an exhausted account](#stop-spawning-onto-an-exhausted-account) | `quarantineCooldownSeconds:` + profile `credentialId:` | CB-578 | `placement/BackendQuarantine` |
| [See which charter a member got](#see-which-charter-a-member-got) | automatic | CB-571 | `peer/CharterReceipt` |
Nearly every knob above lives in one file, on one profile:
@@ -1048,6 +1049,33 @@ member was doing — those are stages B and C of CB-578.
---
## Stop spawning onto an exhausted account
**What.** When a member's turn is classified as a usage-limit refusal (see
[the entry above](#tell-a-usage-limit-refusal-from-a-real-reply)), the **credential** behind it is put
in quarantine for a cooldown. While it is quarantined, an explicit spawn onto it is refused with a
message naming the profile, the credential and roughly how many seconds are left; placement skips it
under every policy; and `bridge_profiles` shows it. The quarantine lifts itself — there is no manual
step.
**On.** `quarantineCooldownSeconds:` at the top level (default 1800, **deferred** — it is baked into
the tracker at startup). Per profile, `credentialId:` (**hot**) says which credential this profile
spends. Quarantine itself only ever fires for a profile that has an `exhaustedPattern`, so a fleet
with no patterns configured behaves exactly as before.
**Why.** Detecting the refusal was only half the problem. Without this, the fleet answers an exhausted
account by spawning another member onto it, which fails the same way, and the operator sees a run of
dead workers rather than one clear cause.
**Gotcha.** It quarantines the **credential, not the profile name**, and that distinction is the whole
point. In this fleet `sol` and `terra` are two different models billing **one** OpenAI account. Locking
only the profile that happened to report the refusal leaves its sibling live, and the next spawn walks
straight onto the same dead account under the other name. Profiles that share an account must share a
`credentialId`. A profile that sets none quarantines alone, under its own name — safe, but it will not
protect a sibling.
---
## See which charter a member got
**What.** Every member launch records a `CharterReceipt`: the role, where the charter came from, a