The model gate can be turned off at runtime but its usage-limit detector cannot be turned on at runtime #446

Closed
opened 2026-09-10 11:40:19 +02:00 by ltms · 1 comment
Owner

The operator asked for three things from model gating:

  1. a central allow-list of usable models
  2. runtime on/off without editing profiles
  3. runtime model limit monitoring — off when a subscription limit is reached, on again when the limit lifts

1 and 2 are shipped and live. I redeployed the Mac's daemon today; fleet_profiles now returns modelGateArmed: true (that key was absent before the restart, which is how I knew the running jar predated the gate). The live config allows 7 models covering all 8 profiles, none turned off, so modelsOff is correctly absent. models: is a hot key, so an on/off flip takes effect without a restart.

3 is not shipped, and this ticket is about a structural reason it cannot be, not just missing code.

The correct config key, since I got it wrong once already

Turning a model off is enabled: false on its models.allow entry — FleetConfig.Models.ModelEntry is record ModelEntry(String model, Boolean enabled) at FleetConfig.java:1419, and an absent enabled: means the same as enabled: true:

models:
  allow:
    - model: openai/gpt-5.6-terra
      enabled: false      # <- this is the flip. NOT "off: true"

The first version of this ticket said off: true. There is no off: config key — I grepped for it and it appears nowhere in main or test sources. off is only the name of the reported set (modelsOff, ModelGateState.off()), which is an output, not an input. Anyone implementing this should use enabled:.

The asymmetry

Measured against main today:

Half of the feature Config key Reload behaviour Where
the gate (turn a model off) models: hot — a reload takes effect on the next spawn, or for status the next report ConfigRef.java:160-166
the detector (notice a usage limit) exhaustedPattern deferred — compiled once at startup, a reload never re-reads it Fleetd.java:381-386, ConfigRef.java:628-634
// Fleetd.java:381-386 — compiled once at startup, keyed by profile name
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
...
        exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));

So today an operator can turn a model off at runtime, but cannot arm the thing that would tell them to. Arming detection for a profile needs a daemon restart. That is the wrong way round for a feature whose whole purpose is to react to a limit while the fleet is running.

Current coverage, measured live

fleet_profiles right now:

exhaustionDetectionArmed: sonnet:false, local-direct:false, local:false, opus:false,
                          xf:false, terra:true, gx:false, sol:true

and in the config file: 2 exhaustedPattern: rows, 0 errorPattern: rows, across 8 profile blocks.

So six of eight profiles cannot detect their own usage limit at all, and no profile can detect a plain backend error. I have a live datapoint for why the second one matters: terra failed with "Our servers are currently overloaded" and was not quarantined. That is a backend error, not exhaustion, and with no errorPattern anywhere neither quarantine nor cooling-off engaged. The profile stayed in the pool and kept being chosen.

What to build

The direction is decided: warn and let the operator flip. Do not auto-probe.

  1. Make exhaustedPattern hot, so detection can be armed without a restart. This is the load-bearing change. It means moving the compile out of Fleetd.main's startup map and reading the pattern live per profile, the same way the gate reads models: through models0(). Watch the cost: compiling a regex per check is wasteful, so cache by pattern string and invalidate when the string changes — do not recompile on every message.
  2. When exhaustion is detected, log a WARNING that names the fix. Not just "profile X is quarantined". Name the model that profile runs on, and give the exact edit — enabled: false on that model's models.allow entry, removed again when the subscription window resets. An operator reading the log should not have to work out which of 7 model names to touch.
  3. Make it observable in the tool surface, so a lead can see it without reading the daemon log — the reason a limit was hit and which model it points at.

Read CompositePeerLauncher.modelGateState()'s javadoc (around :788) before you start. It records a hard-won rule from fleetd #404: "armed" and "which models are off" must come from one read of the same accessor the gate enforces on, so the report can never disagree with the behaviour. Any new reporting this ticket adds must follow the same rule.

What NOT to build, and why

No automatic backoff probe. The obvious design is to retry the model periodically and turn it back on when a call succeeds. I am refusing that: a probe spends quota to discover quota. On a metered subscription the probe is itself the scarce resource it is trying to measure, and a fleet that probes on a timer will burn the first tokens of every new window on discovery rather than work.

No automatic re-enable at all, for now. So "on when the limit lifts" is, in this ticket, a manual flip of enabled: false back to absent — which is cheap precisely because models: is hot and needs no restart.

The open question that is the operator's, not mine

Whether a later ticket should add an automatic re-enable, and if so on what signal. The candidates I can see are all imperfect: a wall-clock reset time (needs the provider's window, which differs per provider and is not in the config), a successful call made by some other path (free, but it may never happen if every path is gated off), or a probe (rejected above). I am not choosing between these, and I am deliberately not implementing a guess. Ask before building it.

Out of scope

  • Do not add exhaustedPattern or errorPattern values to any profile in fleetd.yaml. A pattern has to match real text from that specific backend, and I have observed real failure text from exactly one of the eight. Guessing a regex would produce a detector that looks armed and silently never fires — worse than an honest false in exhaustionDetectionArmed.
  • Do not change the quarantine duration or the cooling-off window.
  • Do not fold this into the gate's on/off code. The gate is correct and live; this is about the detector that should drive it.
The operator asked for three things from model gating: 1. a central allow-list of usable models 2. runtime on/off without editing profiles 3. runtime model **limit monitoring** — off when a subscription limit is reached, on again when the limit lifts **1 and 2 are shipped and live.** I redeployed the Mac's daemon today; `fleet_profiles` now returns `modelGateArmed: true` (that key was absent before the restart, which is how I knew the running jar predated the gate). The live config allows 7 models covering all 8 profiles, none turned off, so `modelsOff` is correctly absent. `models:` is a hot key, so an on/off flip takes effect without a restart. **3 is not shipped, and this ticket is about a structural reason it cannot be, not just missing code.** ## The correct config key, since I got it wrong once already Turning a model off is `enabled: false` on its `models.allow` entry — `FleetConfig.Models.ModelEntry` is `record ModelEntry(String model, Boolean enabled)` at `FleetConfig.java:1419`, and an absent `enabled:` means the same as `enabled: true`: ```yaml models: allow: - model: openai/gpt-5.6-terra enabled: false # <- this is the flip. NOT "off: true" ``` The first version of this ticket said `off: true`. **There is no `off:` config key** — I grepped for it and it appears nowhere in main or test sources. `off` is only the name of the *reported* set (`modelsOff`, `ModelGateState.off()`), which is an output, not an input. Anyone implementing this should use `enabled:`. ## The asymmetry Measured against `main` today: | Half of the feature | Config key | Reload behaviour | Where | |---|---|---|---| | the **gate** (turn a model off) | `models:` | **hot** — a reload takes effect on the next spawn, or for status the next report | `ConfigRef.java:160-166` | | the **detector** (notice a usage limit) | `exhaustedPattern` | **deferred** — compiled once at startup, a reload never re-reads it | `Fleetd.java:381-386`, `ConfigRef.java:628-634` | ```java // Fleetd.java:381-386 — compiled once at startup, keyed by profile name Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>(); ... exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern())); ``` So today an operator **can** turn a model off at runtime, but **cannot** arm the thing that would tell them to. Arming detection for a profile needs a daemon restart. That is the wrong way round for a feature whose whole purpose is to react to a limit while the fleet is running. ## Current coverage, measured live `fleet_profiles` right now: ``` exhaustionDetectionArmed: sonnet:false, local-direct:false, local:false, opus:false, xf:false, terra:true, gx:false, sol:true ``` and in the config file: **2** `exhaustedPattern:` rows, **0** `errorPattern:` rows, across **8** profile blocks. So six of eight profiles cannot detect their own usage limit at all, and **no** profile can detect a plain backend error. I have a live datapoint for why the second one matters: `terra` failed with "Our servers are currently overloaded" and was **not** quarantined. That is a backend error, not exhaustion, and with no `errorPattern` anywhere neither quarantine nor cooling-off engaged. The profile stayed in the pool and kept being chosen. ## What to build **The direction is decided: warn and let the operator flip. Do not auto-probe.** 1. **Make `exhaustedPattern` hot**, so detection can be armed without a restart. This is the load-bearing change. It means moving the compile out of `Fleetd.main`'s startup map and reading the pattern live per profile, the same way the gate reads `models:` through `models0()`. Watch the cost: compiling a regex per check is wasteful, so cache by pattern string and invalidate when the string changes — do not recompile on every message. 2. **When exhaustion is detected, log a WARNING that names the fix.** Not just "profile X is quarantined". Name the model that profile runs on, and give the exact edit — `enabled: false` on that model's `models.allow` entry, removed again when the subscription window resets. An operator reading the log should not have to work out which of 7 model names to touch. 3. **Make it observable in the tool surface**, so a lead can see it without reading the daemon log — the reason a limit was hit and which model it points at. Read `CompositePeerLauncher.modelGateState()`'s javadoc (around `:788`) before you start. It records a hard-won rule from fleetd #404: "armed" and "which models are off" must come from **one** read of the same accessor the gate enforces on, so the report can never disagree with the behaviour. Any new reporting this ticket adds must follow the same rule. ## What NOT to build, and why **No automatic backoff probe.** The obvious design is to retry the model periodically and turn it back on when a call succeeds. I am refusing that: a probe spends quota to discover quota. On a metered subscription the probe is itself the scarce resource it is trying to measure, and a fleet that probes on a timer will burn the first tokens of every new window on discovery rather than work. **No automatic re-enable at all, for now.** So "on when the limit lifts" is, in this ticket, a manual flip of `enabled: false` back to absent — which is cheap precisely because `models:` is hot and needs no restart. ## The open question that is the operator's, not mine Whether a *later* ticket should add an automatic re-enable, and if so on what signal. The candidates I can see are all imperfect: a wall-clock reset time (needs the provider's window, which differs per provider and is not in the config), a successful call made by some other path (free, but it may never happen if every path is gated off), or a probe (rejected above). I am not choosing between these, and I am deliberately not implementing a guess. Ask before building it. ## Out of scope - Do not add `exhaustedPattern` or `errorPattern` values to any profile in `fleetd.yaml`. A pattern has to match real text from that specific backend, and I have observed real failure text from exactly one of the eight. Guessing a regex would produce a detector that looks armed and silently never fires — worse than an honest `false` in `exhaustionDetectionArmed`. - Do not change the quarantine duration or the cooling-off window. - Do not fold this into the gate's on/off code. The gate is correct and live; this is about the detector that should drive it.
Author
Owner

Merged as 1fb6176 (PR #457), after three rounds.

mvn -B clean test on the merged tree 953ce11: Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS, 0 compile-error blocks.

Where the mutation numbers come from

Being exact, because the two trees are not the same. The battery ran on merge commit 3d2d521, tree 1dcca23 — origin/main at 235644c plus this branch. The merge that landed has tree 953ce11, which is that plus one file: CharterToolSurfaceTest, from #464's merge (49df792) that landed while the battery was running. 1600 + 1 = 1601, which is the arithmetic the two trees predict. The one extra file was verified green on its own merge.

On the battery tree FleetMcp.java auto-merged against #463's change to the same file, so that build was also the gate on the auto-merge — a clean auto-merge is not a compiling merge.

The three cells

Cell What it changes Result
CONTROL nothing 1600 green
M5 the round-2 survivor, verbatim. The call site stops using either pinned method and logs a literal instead KILLED — profileWithModelGetsTheActionableFix, profileWithNoModelGetsTheFallbackNotAFix
M6 (Cell B) the ternary's branches swapped: a profile with a model gets the no-model text and vice versa. Both extracted methods and both their unit tests untouched KILLED — the same two tests, independently
M7 main()'s call to the new factory replaced with an inert lambda. The factory and its test stay perfect; the daemon quarantines nothing and logs nothing SURVIVED — 1600 green

M5 was the point of round 3 and it is closed. The ListAppender on the class's own logger was the right idiom choice: it answers both cells from one mechanism, because it reads what the sink actually logged for each profile shape.

M7, and the pattern it completes

M7 is the same defect shape as round 2, moved one level up. It is worth stating plainly because it has now happened twice in one ticket:

Extracting a thing in order to pin it creates the seam the test then lands on.

  • Round 2 extracted the warning text, pinned the two methods, and left the call site that selects between them unpinned. That was M5.
  • Round 3 extracted the sink, pinned the factory's behaviour, and left main()'s wiring of it unpinned. That is M7.

The question to ask on any fix of this shape is which of three things a test now reaches: the value, the call site, or the selection between values. Extraction only ever answers the first. Each round moved the boundary up by one and pinned what it had just exposed.

M7 is not a reason to hold the merge, and per the boundary I set in the round-3 brief it does not get a round 4. Proving main() wires this factory means driving daemon startup, which nothing here does. Recording it on #460 with a cheaper option named: a source-reading assertion of the kind #439 used would kill M7 without starting the daemon.

Round 3's disclosed caveat, kept on the record

The worker reported that its first Cell A run was corrupted because it started a second concurrent mvn against the same module directory while a background clean test was in flight. It killed that run, checked for stray ForkedBooter processes, and re-ran cleanly. That disclosure was the right call and it cost it nothing. The numbers above are from my own battery, not that run.

I hit the same hazard from the other side in this session: I started the merged-tree build with & inside an already-backgrounded task, so the task reported success immediately while mvn was still running, and the log had no BUILD SUCCESS line in it. pgrep found 7 maven processes still live, so I waited for that build rather than starting a second one against the same module directory. Closing #446.

Merged as `1fb6176` (PR #457), after three rounds. `mvn -B clean test` on the merged tree `953ce11`: **`Tests run: 1601, Failures: 0, Errors: 0, Skipped: 0`**, `BUILD SUCCESS`, 0 compile-error blocks. ## Where the mutation numbers come from Being exact, because the two trees are not the same. The battery ran on merge commit `3d2d521`, tree `1dcca23` — `origin/main` at `235644c` plus this branch. The merge that landed has tree `953ce11`, which is that plus one file: `CharterToolSurfaceTest`, from #464's merge (`49df792`) that landed while the battery was running. 1600 + 1 = 1601, which is the arithmetic the two trees predict. The one extra file was verified green on its own merge. On the battery tree `FleetMcp.java` auto-merged against #463's change to the same file, so that build was also the gate on the auto-merge — a clean auto-merge is not a compiling merge. ## The three cells | Cell | What it changes | Result | |---|---|---| | CONTROL | nothing | 1600 green | | **M5** | **the round-2 survivor, verbatim.** The call site stops using either pinned method and logs a literal instead | **KILLED** — `profileWithModelGetsTheActionableFix`, `profileWithNoModelGetsTheFallbackNotAFix` | | **M6** (Cell B) | the ternary's branches swapped: a profile *with* a model gets the no-model text and vice versa. Both extracted methods and both their unit tests untouched | **KILLED** — the same two tests, independently | | **M7** | `main()`'s call to the new factory replaced with an inert lambda. The factory and its test stay perfect; the daemon quarantines nothing and logs nothing | **SURVIVED** — 1600 green | M5 was the point of round 3 and it is closed. The `ListAppender` on the class's own logger was the right idiom choice: it answers both cells from one mechanism, because it reads what the sink actually logged for each profile shape. ## M7, and the pattern it completes M7 is the same defect shape as round 2, moved one level up. It is worth stating plainly because it has now happened twice in one ticket: **Extracting a thing in order to pin it creates the seam the test then lands on.** - Round 2 extracted the warning **text**, pinned the two methods, and left the call site that *selects between them* unpinned. That was M5. - Round 3 extracted the **sink**, pinned the factory's behaviour, and left `main()`'s *wiring of it* unpinned. That is M7. The question to ask on any fix of this shape is which of three things a test now reaches: **the value**, **the call site**, or **the selection between values**. Extraction only ever answers the first. Each round moved the boundary up by one and pinned what it had just exposed. M7 is not a reason to hold the merge, and per the boundary I set in the round-3 brief it does not get a round 4. Proving `main()` wires this factory means driving daemon startup, which nothing here does. Recording it on #460 with a cheaper option named: a source-reading assertion of the kind #439 used would kill M7 without starting the daemon. ## Round 3's disclosed caveat, kept on the record The worker reported that its first Cell A run was corrupted because it started a second concurrent `mvn` against the same module directory while a background `clean test` was in flight. It killed that run, checked for stray `ForkedBooter` processes, and re-ran cleanly. That disclosure was the right call and it cost it nothing. The numbers above are from my own battery, not that run. I hit the same hazard from the other side in this session: I started the merged-tree build with `&` inside an already-backgrounded task, so the task reported success immediately while `mvn` was still running, and the log had no `BUILD SUCCESS` line in it. `pgrep` found 7 maven processes still live, so I waited for that build rather than starting a second one against the same module directory. Closing #446.
ltms closed this issue 2026-09-10 14:17:10 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#446