Features: the two limits that bound any usage-limit recovery
+19
@@ -4826,4 +4826,23 @@ subscription is the thing that actually runs out.
|
||||
own behaviour is tested; that it is still *called* is not. Same shape as the six
|
||||
`FleetConfig.validateXxx()` startup calls — fleetd #398 owns closing it.
|
||||
|
||||
**Two limits of the detection this reports on.** Both bound what any recovery feature can do, so
|
||||
read them before designing one.
|
||||
|
||||
- **Detection needs a worker turn that finished.** `CompletionResolver` classifies a usage limit by
|
||||
matching the member's own terminal output when its turn completes. There is no HTTP status or
|
||||
header path into this. So fleetd cannot learn a limit has been hit until a member has run and
|
||||
ended, and it cannot cheaply ask "is the limit lifted yet?" — any auto-resume has to spend real
|
||||
backend work to find out.
|
||||
- **No reset time is delivered, anywhere.** The captured live refusal is `The usage limit has been
|
||||
reached. Try again later.` — there is no time in it, and nothing in the refusal path reads a
|
||||
`Retry-After`. `BackendQuarantine` stores `now + cooldownNanos` and nothing else. Recovery can
|
||||
therefore only ever be a fixed timer or a backoff probe, never a resume scheduled for the real
|
||||
reset moment.
|
||||
- **Every `subscription: true` profile shares one quarantine key.** With no explicit `credentialId`,
|
||||
`effectiveCredentialId()` returns the `<subscription>` sentinel, which is right for a single
|
||||
Claude plan. The consequence is easy to miss: on this fleet the lead's own profile (`opus`) shares
|
||||
that key with the worker fan-out profile (`sonnet`). Arm `exhaustedPattern` on `sonnet` alone and
|
||||
worker exhaustion starts quarantining the lead's seat too. Arm them together, or not at all.
|
||||
|
||||
fleetd #395.
|
||||
|
||||
Reference in New Issue
Block a user