A profile with no exhaustedPattern has usage-limit detection silently OFF — 6 of 8 live profiles, including both Claude subscription ones #395

Closed
opened 2026-09-10 03:00:11 +02:00 by ltms · 2 comments
Owner

The defect

exhaustedPattern is opt-in with no default, and the opt-out is silent:

// FleetConfig.java:504-506
// exhaustedPattern stays null when unset/blank (opt-in) — no defaulting, no vendor
// wording: an unconfigured profile keeps today's completion-fallback behaviour exactly.
exhaustedPattern = (exhaustedPattern == null || exhaustedPattern.isBlank()) ? null : exhaustedPattern;

null means CompletionResolver has no pattern to match, so a usage-limit refusal is never classified as BACKEND_EXHAUSTED, ExhaustionSink is never called, and BackendQuarantine.quarantine() never runs. That profile can never be quarantined for hitting a usage limit.

Nothing warns. Nothing in fleet_profiles or fleet_list distinguishes "not quarantined because it is healthy" from "not quarantined because nothing can ever quarantine it".

The comment above is honest about the mechanism and reads as a deliberate, safe default. It is neither safe nor visible: the feature it disables is the one that protects a metered subscription.

Measured on the live Mac fleet

profile        exhaustedPattern                          subscription
local          ABSENT -> detection OFF                   false
local-direct   ABSENT -> detection OFF                   false
gx             ABSENT -> detection OFF                   false
opus           ABSENT -> detection OFF                   TRUE
sonnet         ABSENT -> detection OFF                   TRUE
sol            "The usage limit has been reached"         false
terra          "The usage limit has been reached"         false
xf             ABSENT -> detection OFF                   false

2 of 8 profiles have limit detection. 6 do not, and that includes both subscription: true profiles — the ones running on the operator's own metered Claude plan, where an unnoticed usage limit is most expensive.

The coupling that makes it sharper

effectiveCredentialId() (FleetConfig.java:739-743) returns the <subscription> sentinel for any subscription: true profile with no explicit credentialId. So opus and sonnet share one quarantine key — documented as intentional, and correct for a single Claude plan.

The consequence is that the lead's own profile (opus) shares a quarantine with the worker fan-out profile (sonnet). Today that never fires, because neither has a pattern. Add a pattern to sonnet alone and worker exhaustion starts taking out the lead's seat. Whoever configures this needs to be told that, and nothing tells them.

Detection also requires a completed worker turn

CompletionResolver classifies by matching the worker's terminal output at turn completion (CompletionResolver.java:320-331, 370-385), with a raw-screen fallback at :471-506. There is no HTTP status or header path. So:

  • fleetd cannot learn about a limit without a worker turn having run and finished.
  • Any automatic recovery probe must therefore perform real backend work; it cannot cheaply ask "is the limit lifted yet?".

Worth stating in the docs, because it bounds what any recovery feature can do.

Related finding: no reset time exists anywhere

Searched the production and test sources for Retry-After and every reset/duration variant: no match in the refusal path. The classifier keeps only the first matching line (CompletionResolver.java:625-635), and the one captured live refusal is The usage limit has been reached. Try again later. — no time in it.

BackendQuarantine stores now + cooldownNanos and nothing else (BackendQuarantine.java:61-77). So recovery today is a fixed timer, and it can only ever become a backoff probe, never a scheduled resume. A reset time cannot be honoured because none is delivered.

Fix

  1. Warn at config load for any profile with no exhaustedPattern, naming the profile and saying limit detection is off for it. Cheapest change, and it converts a silent gap into a visible one.
  2. Report it. fleet_profiles should show per-profile whether limit detection is armed, so free: N can be read correctly.
  3. Do not add a vendor-text default. The existing comment is right that guessing wording is worse than nothing; a default that silently stops matching after a vendor reword is the same defect with extra confidence.
  4. Document that detection needs a completed worker turn, and that <subscription> couples every subscription profile into one quarantine.

Acceptance

  • Starting the daemon with a profile that has no exhaustedPattern produces one clear warning naming that profile.
  • fleet_profiles distinguishes "detection armed" from "detection off".
  • A test asserts both. Today no test covers the unset-pattern path's effect on quarantine reachability.

Found by a spike I ran while designing operator-facing model gating (allow-list, runtime on/off, limit-driven off/on). The gating design assumed the "off at the limit" half already worked; it works for 2 profiles of 8.

## The defect `exhaustedPattern` is opt-in with no default, and the opt-out is silent: ```java // FleetConfig.java:504-506 // exhaustedPattern stays null when unset/blank (opt-in) — no defaulting, no vendor // wording: an unconfigured profile keeps today's completion-fallback behaviour exactly. exhaustedPattern = (exhaustedPattern == null || exhaustedPattern.isBlank()) ? null : exhaustedPattern; ``` `null` means `CompletionResolver` has no pattern to match, so a usage-limit refusal is never classified as `BACKEND_EXHAUSTED`, `ExhaustionSink` is never called, and `BackendQuarantine.quarantine()` never runs. **That profile can never be quarantined for hitting a usage limit.** Nothing warns. Nothing in `fleet_profiles` or `fleet_list` distinguishes "not quarantined because it is healthy" from "not quarantined because nothing can ever quarantine it". The comment above is honest about the mechanism and reads as a deliberate, safe default. It is neither safe nor visible: the feature it disables is the one that protects a metered subscription. ## Measured on the live Mac fleet ``` profile exhaustedPattern subscription local ABSENT -> detection OFF false local-direct ABSENT -> detection OFF false gx ABSENT -> detection OFF false opus ABSENT -> detection OFF TRUE sonnet ABSENT -> detection OFF TRUE sol "The usage limit has been reached" false terra "The usage limit has been reached" false xf ABSENT -> detection OFF false ``` **2 of 8 profiles have limit detection. 6 do not, and that includes both `subscription: true` profiles** — the ones running on the operator's own metered Claude plan, where an unnoticed usage limit is most expensive. ## The coupling that makes it sharper `effectiveCredentialId()` (`FleetConfig.java:739-743`) returns the `<subscription>` sentinel for any `subscription: true` profile with no explicit `credentialId`. So `opus` and `sonnet` share one quarantine key — documented as intentional, and correct for a single Claude plan. The consequence is that the *lead's own* profile (`opus`) shares a quarantine with the worker fan-out profile (`sonnet`). Today that never fires, because neither has a pattern. Add a pattern to `sonnet` alone and worker exhaustion starts taking out the lead's seat. Whoever configures this needs to be told that, and nothing tells them. ## Detection also requires a completed worker turn `CompletionResolver` classifies by matching the worker's **terminal output** at turn completion (`CompletionResolver.java:320-331`, `370-385`), with a raw-screen fallback at `:471-506`. There is no HTTP status or header path. So: - fleetd cannot learn about a limit without a worker turn having run and finished. - Any automatic recovery probe must therefore perform real backend work; it cannot cheaply ask "is the limit lifted yet?". Worth stating in the docs, because it bounds what any recovery feature can do. ## Related finding: no reset time exists anywhere Searched the production and test sources for `Retry-After` and every reset/duration variant: **no match in the refusal path.** The classifier keeps only the first matching line (`CompletionResolver.java:625-635`), and the one captured live refusal is `The usage limit has been reached. Try again later.` — no time in it. `BackendQuarantine` stores `now + cooldownNanos` and nothing else (`BackendQuarantine.java:61-77`). So recovery today is a fixed timer, and it can only ever become a backoff probe, never a scheduled resume. A reset time cannot be honoured because none is delivered. ## Fix 1. **Warn at config load** for any profile with no `exhaustedPattern`, naming the profile and saying limit detection is off for it. Cheapest change, and it converts a silent gap into a visible one. 2. **Report it.** `fleet_profiles` should show per-profile whether limit detection is armed, so `free: N` can be read correctly. 3. **Do not add a vendor-text default.** The existing comment is right that guessing wording is worse than nothing; a default that silently stops matching after a vendor reword is the same defect with extra confidence. 4. Document that detection needs a completed worker turn, and that `<subscription>` couples every subscription profile into one quarantine. ## Acceptance - Starting the daemon with a profile that has no `exhaustedPattern` produces one clear warning naming that profile. - `fleet_profiles` distinguishes "detection armed" from "detection off". - A test asserts both. Today no test covers the unset-pattern path's effect on quarantine reachability. Found by a spike I ran while designing operator-facing model gating (allow-list, runtime on/off, limit-driven off/on). The gating design assumed the "off at the limit" half already worked; it works for 2 profiles of 8.
Author
Owner

Fixed and merged as 7180b1a

PR #401, merged onto main with a --no-ff merge. Merged build measured by me: Tests run: 1474, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS, 0 compile errors.

Against the acceptance list

Acceptance State
Starting the daemon with an unarmed profile produces one clear warning naming it Done — one aggregate WARN naming every unarmed profile, plus a second louder WARN if any of them is subscription: true. All-armed configs log INFO, so the healthy case is also on the record
fleet_profiles distinguishes armed from off Done — exhaustionDetectionArmed, one boolean per profile, on both fleet_profiles and GET /profiles. Always emitted, not conditional like quarantined
A test asserts both Done — ExhaustedPatternGapReportTest (a @TempDir config reproducing the live 8-profile shape) and FleetProfilesArmedFieldTest
Fix item 3: no vendor-text default Honoured — nothing was defaulted. exhaustedPattern stays opt-in
Fix item 4: document the two constraints Done — Features wiki, "Startup says which profiles have usage-limit detection turned off" (ad3732a): detection needs a completed worker turn, no reset time is ever delivered, and every subscription: true profile shares the <subscription> quarantine key, so arming sonnet alone would start quarantining the lead's opus seat

One gap left open on purpose

Deleting the reportExhaustedPatternGap(cfg) call from Fleetd.java:140 leaves the suite green (measured: 1472 tests, 0 failures in the worker's tree). Both new tests call the method directly, and no test starts the main() sequence, so the report's behaviour is pinned but its invocation is not.

That is a pre-existing shape, not something this PR introduced — the six cfg.validateXxx() calls beside it have the same hole. It is folded into #398's scope, recorded in the merge commit, and written into the wiki entry as a known gap. Not tracked here.

What this issue does not close

The "off when the limit is reached, on when it lifts" half of the model-gating design is still open, and this ticket bounded it usefully: because no reset time is delivered anywhere, the "on again" side can only be a backoff probe that spends real backend work, never a scheduled resume. That decision is with the operator and is not blocking anything else.

Closing.

## Fixed and merged as `7180b1a` PR #401, merged onto `main` with a `--no-ff` merge. Merged build measured by me: `Tests run: 1474, Failures: 0, Errors: 0, Skipped: 0`, **BUILD SUCCESS**, 0 compile errors. ### Against the acceptance list | Acceptance | State | |---|---| | Starting the daemon with an unarmed profile produces one clear warning naming it | **Done** — one aggregate `WARN` naming every unarmed profile, plus a second louder `WARN` if any of them is `subscription: true`. All-armed configs log `INFO`, so the healthy case is also on the record | | `fleet_profiles` distinguishes armed from off | **Done** — `exhaustionDetectionArmed`, one boolean per profile, on both `fleet_profiles` and `GET /profiles`. Always emitted, not conditional like `quarantined` | | A test asserts both | **Done** — `ExhaustedPatternGapReportTest` (a `@TempDir` config reproducing the live 8-profile shape) and `FleetProfilesArmedFieldTest` | | Fix item 3: no vendor-text default | **Honoured** — nothing was defaulted. `exhaustedPattern` stays opt-in | | Fix item 4: document the two constraints | **Done** — Features wiki, "Startup says which profiles have usage-limit detection turned off" (`ad3732a`): detection needs a completed worker turn, no reset time is ever delivered, and every `subscription: true` profile shares the `<subscription>` quarantine key, so arming `sonnet` alone would start quarantining the lead's `opus` seat | ### One gap left open on purpose Deleting the `reportExhaustedPatternGap(cfg)` call from `Fleetd.java:140` leaves the suite green (measured: 1472 tests, 0 failures in the worker's tree). Both new tests call the method directly, and no test starts the `main()` sequence, so the report's behaviour is pinned but its invocation is not. That is a pre-existing shape, not something this PR introduced — the six `cfg.validateXxx()` calls beside it have the same hole. It is folded into #398's scope, recorded in the merge commit, and written into the wiki entry as a known gap. Not tracked here. ### What this issue does not close The "off when the limit is reached, on when it lifts" half of the model-gating design is still open, and this ticket bounded it usefully: because no reset time is delivered anywhere, the "on again" side can only be a backoff probe that spends real backend work, never a scheduled resume. That decision is with the operator and is not blocking anything else. Closing.
ltms closed this issue 2026-09-10 03:46:47 +02:00
Author
Owner

Correction to this ticket's premise — the gap was never silence

I wrote that exhaustedPattern unset means detection is off "silently, with no warning at config
load
". That is false, and the fleet01 lead measured it on a real startup:

02:11:19.607 INFO  dev.ltms.fleet.Fleetd - backend-exhausted classification (CB-578 stage A):
    off (no profile has an exhaustedPattern configured; profiles: [gx, local, opus, xf])
02:11:19.609 INFO  dev.ltms.fleet.Fleetd - backend-error classification (fleetd #201 Unit 5):
    off (no profile has an errorPattern configured; profiles: [gx, local, opus, xf])

Fleetd.java:382 and :403, both via CompletionResolver.coverage(). It warns at every startup
and it names the unconfigured profiles. That code predates this ticket.

So the accurate statement is: the fact is recorded, at INFO, where nobody looks. It is not on a
queryable surface, which is a different problem with a different fix.

Why the wording matters rather than being pedantry

"No warning exists" invites the next person to add the warning that is already there — a second
reporter saying the same thing in the same place, and now two lines to keep in step. The reason to
correct a closed ticket is that the ticket is what someone reads before touching this code.

The fix that shipped was right for the real gap. exhaustionDetectionArmed puts the fact on
fleet_profiles and GET /profiles, where a lead already looks before reading free: N. That is
the queryable surface the INFO line is not. The implementation needed one correction (#404, the
lookup read the live config while detection read the startup map) and one more during that review (a
mutation showed the field could go permanently false with the suite green). Both are merged.

A second, worse instance of the same confusion — filed as #415

While checking this, the same coverage() helper turned out to produce a line that is not merely
buried but false. backend-error classification: off is wrong: every profile without an
errorPattern falls back to CompletionResolver's built-in (?i)\bAPI Error\s*: pattern
(CompletionResolver.java:84), so classification is running. For exhaustedPattern the same word is
true, because no fallback exists.

One generic helper, two keys that differ in what unset means, and the false answer is the reassuring
one. Confirmed in this tree, not taken on report.

Credit where it belongs

The fleet01 lead also caught themselves nearly making the mirror error: their first journal grep
returned only the errorPattern line, and their pattern had excluded the exhaustedPattern one — so
they were one sentence from reporting that it has no coverage log at all. A grep that returns
nothing is evidence about the grep first.
Worth recording next to this correction, because the same
zero is what my original wording was built on.

## Correction to this ticket's premise — the gap was never silence I wrote that `exhaustedPattern` unset means detection is off "**silently, with no warning at config load**". That is false, and the fleet01 lead measured it on a real startup: ``` 02:11:19.607 INFO dev.ltms.fleet.Fleetd - backend-exhausted classification (CB-578 stage A): off (no profile has an exhaustedPattern configured; profiles: [gx, local, opus, xf]) 02:11:19.609 INFO dev.ltms.fleet.Fleetd - backend-error classification (fleetd #201 Unit 5): off (no profile has an errorPattern configured; profiles: [gx, local, opus, xf]) ``` `Fleetd.java:382` and `:403`, both via `CompletionResolver.coverage()`. It warns at **every** startup and it **names the unconfigured profiles**. That code predates this ticket. So the accurate statement is: **the fact is recorded, at INFO, where nobody looks.** It is not on a queryable surface, which is a different problem with a different fix. ## Why the wording matters rather than being pedantry "No warning exists" invites the next person to add the warning that is already there — a second reporter saying the same thing in the same place, and now two lines to keep in step. The reason to correct a closed ticket is that the ticket is what someone reads before touching this code. **The fix that shipped was right for the real gap.** `exhaustionDetectionArmed` puts the fact on `fleet_profiles` and `GET /profiles`, where a lead already looks before reading `free: N`. That is the queryable surface the INFO line is not. The implementation needed one correction (#404, the lookup read the live config while detection read the startup map) and one more during that review (a mutation showed the field could go permanently false with the suite green). Both are merged. ## A second, worse instance of the same confusion — filed as #415 While checking this, the same `coverage()` helper turned out to produce a line that is not merely buried but **false**. `backend-error classification: off` is wrong: every profile without an `errorPattern` falls back to `CompletionResolver`'s built-in `(?i)\bAPI Error\s*:` pattern (`CompletionResolver.java:84`), so classification is running. For `exhaustedPattern` the same word is true, because no fallback exists. One generic helper, two keys that differ in what unset means, and the false answer is the reassuring one. Confirmed in this tree, not taken on report. ## Credit where it belongs The fleet01 lead also caught themselves nearly making the mirror error: their first journal grep returned only the `errorPattern` line, and their pattern had excluded the `exhaustedPattern` one — so they were one sentence from reporting that it has no coverage log at all. **A grep that returns nothing is evidence about the grep first.** Worth recording next to this correction, because the same zero is what my original wording was built on.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#395