Features: the two limits that bound any usage-limit recovery

Dai Ha
2026-09-10 08:46:21 +07:00
parent 2a570c5c06
commit ad3732a908
+19
@@ -4826,4 +4826,23 @@ subscription is the thing that actually runs out.
own behaviour is tested; that it is still *called* is not. Same shape as the six
`FleetConfig.validateXxx()` startup calls — fleetd #398 owns closing it.
**Two limits of the detection this reports on.** Both bound what any recovery feature can do, so
read them before designing one.
- **Detection needs a worker turn that finished.** `CompletionResolver` classifies a usage limit by
matching the member's own terminal output when its turn completes. There is no HTTP status or
header path into this. So fleetd cannot learn a limit has been hit until a member has run and
ended, and it cannot cheaply ask "is the limit lifted yet?" — any auto-resume has to spend real
backend work to find out.
- **No reset time is delivered, anywhere.** The captured live refusal is `The usage limit has been
reached. Try again later.` — there is no time in it, and nothing in the refusal path reads a
`Retry-After`. `BackendQuarantine` stores `now + cooldownNanos` and nothing else. Recovery can
therefore only ever be a fixed timer or a backoff probe, never a resume scheduled for the real
reset moment.
- **Every `subscription: true` profile shares one quarantine key.** With no explicit `credentialId`,
`effectiveCredentialId()` returns the `<subscription>` sentinel, which is right for a single
Claude plan. The consequence is easy to miss: on this fleet the lead's own profile (`opus`) shares
that key with the worker fan-out profile (`sonnet`). Arm `exhaustedPattern` on `sonnet` alone and
worker exhaustion starts quarantining the lead's seat too. Arm them together, or not at all.
fleetd #395.