Model gating units 2+3: enforce the allow-list at spawn, and turn one model off at runtime without editing profiles #422

Closed
opened 2026-09-10 06:39:58 +02:00 by ltms · 2 comments
Owner

Follow-up to the "central allow-list of usable models" ticket, which shipped config-load validation only. FleetConfig.Models' javadoc says so and names what was left out:

Out of scope here, deliberately: nothing in this block is read at spawn time — enforcing it against a live spawn, an on/off runtime switch, and any interaction with BackendQuarantine are separate units. This block is config-load validation only.

This ticket is the first two of those: enforce at spawn and runtime on/off.

Why these are one unit, not two

Split apart, each half ships inert. A gate with no flag to read always allows; a flag nothing reads always does nothing. Both would pass a full build with every test green. That is the silent-default-disables-features shape, and it has already cost this project a shipped-but-off feature.

So: one worker, one PR.

What the operator actually needs

Real case, today. xf runs opencode/nemotron-3-ultra-free, and the opencode subscription ran out of its weekly allowance. sol and terra run openai/gpt-5.6-sol / openai/gpt-5.6-terra on a subscription that just came back. The operator wants to turn a model off while its allowance is gone and on when it lifts, without editing the profiles: block — because a profile's model:, argv, env and the rest are deferred launch settings, so editing them needs a full daemon restart, and a restart drops every in-flight ticket.

Turning a model off must therefore be a hot change.

The design trap, and the one thing that must not be got wrong

allow and a new on/off flag answer two different questions, and collapsing them makes the feature unusable:

  • allow = which model ids may appear in profiles: at all. Checked at config load by validateModels().
  • on/off = whether a spawn onto that model is permitted right now. Checked at spawn.

If "off" were implemented by removing the entry from allow, then turning a model off would make validateModels() refuse the whole reload — because the profile still names that model, and ConfigRef.reload() runs validateAll() and refuses the config outright rather than applying it partially (ConfigRef.java:449-455). The operator's only way to turn a model off would then be to also edit profiles:, which is exactly what this ticket exists to avoid.

So an entry that is off stays in allow. It stays valid to name; it just stops being spawnable.

Where the flag goes

Models.ModelEntry is already a record for exactly this reason — its javadoc says:

named as a record rather than a bare string on purpose: a later unit needs to hang an on/off state and a load-limit state off each entry, and a bare List<String> cannot grow those fields without changing the YAML shape underneath every operator who already wrote one

So add the field there. Absent ⇒ on. Every config on every host was written before this field existed and must keep working unchanged — the same compatibility reasoning that makes an absent models.allow: mean "off".

Where the gate goes

CompositePeerLauncher already holds this exact family of spawn refusals, and the model gate is a fourth sibling:

CompositePeerLauncher.java:332   enforceNotQuarantined(requestedProfile);
CompositePeerLauncher.java:333   enforceNotCoolingOff(requestedProfile);
CompositePeerLauncher.java:334   enforceMaxLoad(requestedProfile);

Read the live config, not a snapshot. profiles0() at :301 is the pattern to copy — it calls profileConfigs.get() on a supplier every time, which is why maxLoad / weight / credentialId are hot. A field captured at construction would give you #416 in reverse: an operator turns a model off, the reload reports success, and spawns keep landing on it forever.

Both halves of the gate, or it is a one-way gate

There are two paths into a spawn, and each family of refusal above covers both:

  1. Explicit profile — fleet_spawn{profile: "xf"} — refused by the enforceX methods at :332-334.
  2. Weighted placement — an unqualified spawn — filtered by the candidate-set methods, quarantinedProfiles at :445 and coolingOffProfiles at :470, which feed PlacementContext at :352.

Covering only path 1 means an unqualified fleet_spawn still lands on a disabled model. Covering only path 2 means an explicit spawn does. Both, or the unit is not done. This is the a-one-way-gate-is-not-a-gate lesson: ask which states open it, not only which close it.

models must be reclassified, or the reload report goes false

models is in ConfigRef.DEFERRED_KEYS (ConfigRef.java:227-229), and that is correct today — the class javadoc explains why at :46-51: a good edit "has nothing built at startup to rebuild", so it is reported deferred rather than silently swallowed.

This ticket makes that false. Once the spawn gate reads models live, a models: edit takes effect on the very next spawn. A reload that still reports "these changes need a restart to take effect" would be telling the operator to restart for a change that already applied — and the operator would restart, dropping live members, for nothing.

Work out the honest class and move it. Read the hot / deferred / cold / split definitions in the ConfigRef class javadoc and pick from them; do not invent a fourth. Two facts to weigh:

  • The gate reads live ⇒ that half is hot.
  • validateModels() re-runs on reload and refuses a bad edit outright ⇒ already handled by the catch block, and not a reason to call the key cold.

State your reasoning in the javadoc next to the classification, the way every other key there does. If the answer is split, say which half is live and which needs a restart, in the same shape as the existing health: / coordinator: / fleet: bullets.

ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest both read those sets and will hold you to triaging the key. Do not add an exclusion to get them green — that is the escape hatch #323 was filed for.

The receipt must read what the gate reads

fleet_profiles should report which models are currently off, so an operator can see the gate's state without reading the config file. That report must read the same source the gate reads — the live config.

This is #404 exactly: exhaustionDetectionArmed read the live config while the detection it described read the startup snapshot, so the status field promised something the behaviour could not deliver. Do not repeat it in the other direction either. If the gate reads live and the report reads live, they agree by construction; anything else needs a test proving they agree.

Acceptance criteria

  1. A model entry marked off, with its profile unchanged, refuses fleet_spawn{profile: <that profile>}. The message names the model and says the operator turned it off — distinct wording from the quarantine and cool-off refusals, so a lead reading a spawn failure can tell the three apart.
  2. An unqualified fleet_spawn never places onto a profile whose model is off. When every candidate's model is off, the error names that as the cause rather than reporting a generic "no candidates".
  3. An entry with no on/off field set behaves exactly as today. Prove it: a config written before this field existed spawns unchanged.
  4. Turning a model off and on again is a hot change — no restart, and the reload does not tell the operator a restart is needed. Test this with a real ConfigRef.reload(), and assert applied().
  5. models is reclassified in ConfigRef with its reasoning written down, and both top-level coverage tests pass without a new exclusion.
  6. fleet_profiles reports the off models, read from the live config.
  7. A profile whose model is off is still a valid config — validateModels() must not refuse it. This is criterion 3's mirror and the whole point of the ticket; test it directly.
  8. Disabling a model disables every profile that names it. Today deepseek-v4-flash is named by both local and local-direct, so turning it off takes out two profiles at once. That is the correct meaning — the model is what is rate-limited, not the profile — but it must be documented in the field's javadoc, because it will surprise someone.

Prove your own tests

For each of criteria 1, 2 and 4, break the fix on purpose and confirm the matching test goes red, one mutation at a time, restored afterwards:

  • A — make the gate read a snapshot captured at construction instead of the live supplier. Criterion 4's test must fail.
  • B — remove the candidate-set filter, keeping the explicit-profile gate. Criterion 2's test must fail.
  • C — remove the explicit-profile gate, keeping the candidate filter. Criterion 1's test must fail.

Report the exact failing test name and assertion message for each. A mutation that leaves the suite green means that half is not pinned, and the unit is not done. Mutation B and C exist because they are the only thing that proves the two halves are independently load-bearing — a single test that happens to exercise both would hide a missing one.

Out of scope — do not build these

  • Automatic detection of a subscription limit, and automatic re-enable when it lifts. That is a separate unit and it is blocked on an operator decision that has not been made: whether re-enabling is an automatic backoff probe or a warning plus a manual flip. Do not guess, do not build "just the detection half", and do not leave a hook for it. The on/off flag this ticket adds is the seam that unit will use.
  • Any change to BackendQuarantine or BackendOutagePolicy. The model gate is a fourth, independent reason to refuse a spawn — it must not reuse or extend the quarantine machinery, which is for backend-reported outages, not operator intent.
  • Validating that a model id is one the backend actually knows. The operator explicitly declined that.
  • Editing fleetd.yaml. It is gitignored and you cannot see the real one. Reproduce the shape you need in a @TempDir fixture.

Reference — the live shape

models:
  allow:
    - model: deepseek-v4-flash            # local, local-direct
    - model: gx/deepseek-v4-flash         # gx
    - model: claude-opus-5                # opus (subscription)
    - model: claude-sonnet-5              # sonnet (subscription)
    - model: openai/gpt-5.6-sol           # sol
    - model: openai/gpt-5.6-terra         # terra
    - model: opencode/nemotron-3-ultra-free   # xf

Model ids are a single flat opaque-string namespace on purpose — a bare Claude id and a provider-prefixed opencode id both fit unchanged, because the comparison is exact string equality and never parses a provider prefix or branches on a profile's kind:. Keep that property.

Follow-up to the "central allow-list of usable models" ticket, which shipped **config-load validation only**. `FleetConfig.Models`' javadoc says so and names what was left out: > **Out of scope here, deliberately:** nothing in this block is read at spawn time — enforcing it against a live spawn, an on/off runtime switch, and any interaction with `BackendQuarantine` are separate units. This block is config-load validation only. This ticket is the first two of those: **enforce at spawn** and **runtime on/off**. ## Why these are one unit, not two Split apart, each half ships inert. A gate with no flag to read always allows; a flag nothing reads always does nothing. Both would pass a full build with every test green. That is the [[silent-default-disables-features]] shape, and it has already cost this project a shipped-but-off feature. So: one worker, one PR. ## What the operator actually needs Real case, today. `xf` runs `opencode/nemotron-3-ultra-free`, and the opencode subscription ran out of its weekly allowance. `sol` and `terra` run `openai/gpt-5.6-sol` / `openai/gpt-5.6-terra` on a subscription that just came back. The operator wants to turn a model **off** while its allowance is gone and **on** when it lifts, **without editing the `profiles:` block** — because a profile's `model:`, `argv`, `env` and the rest are deferred launch settings, so editing them needs a full daemon restart, and a restart drops every in-flight ticket. Turning a model off must therefore be a **hot** change. ## The design trap, and the one thing that must not be got wrong `allow` and a new on/off flag answer **two different questions**, and collapsing them makes the feature unusable: - **`allow`** = which model ids may appear in `profiles:` at all. Checked at config load by `validateModels()`. - **on/off** = whether a spawn onto that model is permitted **right now**. Checked at spawn. If "off" were implemented by removing the entry from `allow`, then turning a model off would make `validateModels()` **refuse the whole reload** — because the profile still names that model, and `ConfigRef.reload()` runs `validateAll()` and refuses the config outright rather than applying it partially (`ConfigRef.java:449-455`). The operator's only way to turn a model off would then be to *also* edit `profiles:`, which is exactly what this ticket exists to avoid. **So an entry that is off stays in `allow`.** It stays valid to name; it just stops being spawnable. ## Where the flag goes `Models.ModelEntry` is already a record for exactly this reason — its javadoc says: > named as a record rather than a bare string on purpose: a later unit needs to hang an on/off state and a load-limit state off each entry, and a bare `List<String>` cannot grow those fields without changing the YAML shape underneath every operator who already wrote one So add the field there. **Absent ⇒ on.** Every config on every host was written before this field existed and must keep working unchanged — the same compatibility reasoning that makes an absent `models.allow:` mean "off". ## Where the gate goes `CompositePeerLauncher` already holds this exact family of spawn refusals, and the model gate is a fourth sibling: ``` CompositePeerLauncher.java:332 enforceNotQuarantined(requestedProfile); CompositePeerLauncher.java:333 enforceNotCoolingOff(requestedProfile); CompositePeerLauncher.java:334 enforceMaxLoad(requestedProfile); ``` Read the live config, not a snapshot. `profiles0()` at `:301` is the pattern to copy — it calls `profileConfigs.get()` on a supplier every time, which is why `maxLoad` / `weight` / `credentialId` are hot. A field captured at construction would give you #416 in reverse: an operator turns a model off, the reload reports success, and spawns keep landing on it forever. ### Both halves of the gate, or it is a one-way gate There are **two** paths into a spawn, and each family of refusal above covers both: 1. **Explicit profile** — `fleet_spawn{profile: "xf"}` — refused by the `enforceX` methods at `:332-334`. 2. **Weighted placement** — an unqualified spawn — filtered by the candidate-set methods, `quarantinedProfiles` at `:445` and `coolingOffProfiles` at `:470`, which feed `PlacementContext` at `:352`. Covering only path 1 means an unqualified `fleet_spawn` still lands on a disabled model. Covering only path 2 means an explicit spawn does. **Both, or the unit is not done.** This is the [[a-one-way-gate-is-not-a-gate]] lesson: ask which states open it, not only which close it. ## `models` must be reclassified, or the reload report goes false `models` is in `ConfigRef.DEFERRED_KEYS` (`ConfigRef.java:227-229`), and that is **correct today** — the class javadoc explains why at `:46-51`: a good edit "has nothing built at startup to rebuild", so it is reported deferred rather than silently swallowed. **This ticket makes that false.** Once the spawn gate reads `models` live, a `models:` edit takes effect on the very next spawn. A reload that still reports "these changes need a restart to take effect" would be telling the operator to restart for a change that already applied — and the operator would restart, dropping live members, for nothing. Work out the honest class and move it. Read the `hot` / `deferred` / `cold` / `split` definitions in the `ConfigRef` class javadoc and pick from them; do not invent a fourth. Two facts to weigh: - The gate reads live ⇒ that half is hot. - `validateModels()` re-runs on reload and refuses a bad edit outright ⇒ already handled by the catch block, and not a reason to call the key cold. State your reasoning in the javadoc next to the classification, the way every other key there does. If the answer is `split`, say which half is live and which needs a restart, in the same shape as the existing `health:` / `coordinator:` / `fleet:` bullets. `ConfigRefTopLevelCoverageTest` and `ConfigRefTopLevelReportingCoverageTest` both read those sets and will hold you to triaging the key. Do not add an exclusion to get them green — that is the escape hatch #323 was filed for. ## The receipt must read what the gate reads `fleet_profiles` should report which models are currently off, so an operator can see the gate's state without reading the config file. **That report must read the same source the gate reads** — the live config. This is #404 exactly: `exhaustionDetectionArmed` read the live config while the detection it described read the startup snapshot, so the status field promised something the behaviour could not deliver. Do not repeat it in the other direction either. If the gate reads live and the report reads live, they agree by construction; anything else needs a test proving they agree. ## Acceptance criteria 1. A model entry marked off, with its profile unchanged, refuses `fleet_spawn{profile: <that profile>}`. The message names the model and says the operator turned it off — distinct wording from the quarantine and cool-off refusals, so a lead reading a spawn failure can tell the three apart. 2. An unqualified `fleet_spawn` never places onto a profile whose model is off. When **every** candidate's model is off, the error names that as the cause rather than reporting a generic "no candidates". 3. An entry with no on/off field set behaves exactly as today. Prove it: a config written before this field existed spawns unchanged. 4. Turning a model off and on again is a **hot** change — no restart, and the reload does not tell the operator a restart is needed. Test this with a real `ConfigRef.reload()`, and assert `applied()`. 5. `models` is reclassified in `ConfigRef` with its reasoning written down, and both top-level coverage tests pass without a new exclusion. 6. `fleet_profiles` reports the off models, read from the live config. 7. A profile whose model is off is still a **valid** config — `validateModels()` must not refuse it. This is criterion 3's mirror and the whole point of the ticket; test it directly. 8. Disabling a model disables **every** profile that names it. Today `deepseek-v4-flash` is named by both `local` and `local-direct`, so turning it off takes out two profiles at once. That is the correct meaning — the model is what is rate-limited, not the profile — but it must be documented in the field's javadoc, because it will surprise someone. ## Prove your own tests For each of criteria 1, 2 and 4, break the fix on purpose and confirm the matching test goes red, one mutation at a time, restored afterwards: - **A** — make the gate read a snapshot captured at construction instead of the live supplier. Criterion 4's test must fail. - **B** — remove the candidate-set filter, keeping the explicit-profile gate. Criterion 2's test must fail. - **C** — remove the explicit-profile gate, keeping the candidate filter. Criterion 1's test must fail. Report the exact failing test name and assertion message for each. **A mutation that leaves the suite green means that half is not pinned**, and the unit is not done. Mutation B and C exist because they are the only thing that proves the two halves are independently load-bearing — a single test that happens to exercise both would hide a missing one. ## Out of scope — do not build these - **Automatic detection of a subscription limit, and automatic re-enable when it lifts.** That is a separate unit and it is **blocked on an operator decision** that has not been made: whether re-enabling is an automatic backoff probe or a warning plus a manual flip. Do not guess, do not build "just the detection half", and do not leave a hook for it. The on/off flag this ticket adds is the seam that unit will use. - Any change to `BackendQuarantine` or `BackendOutagePolicy`. The model gate is a fourth, independent reason to refuse a spawn — it must not reuse or extend the quarantine machinery, which is for backend-reported outages, not operator intent. - Validating that a model id is one the backend actually knows. The operator explicitly declined that. - Editing `fleetd.yaml`. It is gitignored and you cannot see the real one. Reproduce the shape you need in a `@TempDir` fixture. ## Reference — the live shape ```yaml models: allow: - model: deepseek-v4-flash # local, local-direct - model: gx/deepseek-v4-flash # gx - model: claude-opus-5 # opus (subscription) - model: claude-sonnet-5 # sonnet (subscription) - model: openai/gpt-5.6-sol # sol - model: openai/gpt-5.6-terra # terra - model: opencode/nemotron-3-ultra-free # xf ``` Model ids are a single flat opaque-string namespace on purpose — a bare Claude id and a provider-prefixed opencode id both fit unchanged, because the comparison is exact string equality and never parses a provider prefix or branches on a profile's `kind:`. Keep that property.
Author
Owner

Measured on the second host: the gate ships green and inert there, and nothing says so

Checked fleet01's live config read-only over ssh, because this feature is driven entirely by a gitignored file and no test can see one:

grep -c '^models:' fleetd/fleetd.yaml        -> 0
top-level keys: bind, broker, configReload, coordinator, fleet, guard, health,
                herdrSocket, lifecycle, memberCredentials, memberSkills,
                placement, profiles, worktreeRoot
profiles: gx, local, opus, xf

The Mac has a models: block with 7 entries. fleet01 has none.

This is not a defect in PR #429. models0() normalising a missing block to NO_MODELS_CONFIGURED so the gate never fires is the correct back-compat behaviour, and it is deliberately documented that way. The problem is what the operator is told.

On a host with no models: block, every part of this feature is silently absent: enforceModelEnabled never refuses, modelOffProfiles is always empty, and disabledModels() returns an empty set which fleet_profiles will report as "nothing is off" — indistinguishable from "everything is on and the gate is working". An operator who flips a model off on the Mac, sees it work, and expects the same lever on fleet01 gets nothing, with no warning at any point.

That is the "gitignored config ships inert" shape, and it is also the #415 shape one level up: "off" and "no fallback configured" are different facts and must read differently.

Follow-up unit (not for the worker currently on #429 — do not add scope mid-round)

A startup coverage line plus a status field, in the style already established by CompletionResolver.coverage (#415) and FleetHealthMonitor.coverage:

  • Startup log: one line saying which state the model gate is in — off (no models: block configured), armed (N models allowed, none turned off), or armed (N allowed, M turned off: <ids>). The three must be distinguishable, and the "no block" case must not read as "nothing is off".
  • fleet_profiles/GET /profiles: report the gate's state, not only the off-list. An empty disabledModels() currently conflates "no block" with "block present, nothing off".
  • The #404 rule applies: whatever field reports the state must read models0() — the same accessor enforceModelEnabled and modelOffProfiles read. PR #429 already got that right for disabledModels(); the new field must not introduce a second source.
  • Test the caller's choice of argument, not just the helper: the #415 lesson (instance 26) is that a coverage(...)-style helper with the right signature can still be called with the wrong arguments and leave the whole suite green. Assert what Fleetd passes.

This also needs a decision I am recording rather than leaving open: a missing models: block stays permitted. Requiring one would break every existing config, including fleet01's, and the allow-list was shipped in #398 as opt-in on purpose. The fix is to make the inert state visible, never to make it fatal.

## Measured on the second host: the gate ships green and inert there, and nothing says so Checked fleet01's live config read-only over `ssh`, because this feature is driven entirely by a gitignored file and no test can see one: ``` grep -c '^models:' fleetd/fleetd.yaml -> 0 top-level keys: bind, broker, configReload, coordinator, fleet, guard, health, herdrSocket, lifecycle, memberCredentials, memberSkills, placement, profiles, worktreeRoot profiles: gx, local, opus, xf ``` The Mac has a `models:` block with 7 entries. **fleet01 has none.** **This is not a defect in PR #429.** `models0()` normalising a missing block to `NO_MODELS_CONFIGURED` so the gate never fires is the correct back-compat behaviour, and it is deliberately documented that way. The problem is what the operator is told. On a host with no `models:` block, every part of this feature is silently absent: `enforceModelEnabled` never refuses, `modelOffProfiles` is always empty, and `disabledModels()` returns an empty set which `fleet_profiles` will report as "nothing is off" — indistinguishable from "everything is on and the gate is working". An operator who flips a model off on the Mac, sees it work, and expects the same lever on fleet01 gets nothing, with no warning at any point. That is the "gitignored config ships inert" shape, and it is also the #415 shape one level up: **"off" and "no fallback configured" are different facts and must read differently.** ## Follow-up unit (not for the worker currently on #429 — do not add scope mid-round) A startup coverage line plus a status field, in the style already established by `CompletionResolver.coverage` (#415) and `FleetHealthMonitor.coverage`: - **Startup log**: one line saying which state the model gate is in — `off (no models: block configured)`, `armed (N models allowed, none turned off)`, or `armed (N allowed, M turned off: <ids>)`. The three must be distinguishable, and the "no block" case must not read as "nothing is off". - **`fleet_profiles`/`GET /profiles`**: report the gate's state, not only the off-list. An empty `disabledModels()` currently conflates "no block" with "block present, nothing off". - **The #404 rule applies**: whatever field reports the state must read `models0()` — the same accessor `enforceModelEnabled` and `modelOffProfiles` read. PR #429 already got that right for `disabledModels()`; the new field must not introduce a second source. - **Test the caller's choice of argument**, not just the helper: the #415 lesson (instance 26) is that a `coverage(...)`-style helper with the right signature can still be called with the wrong arguments and leave the whole suite green. Assert what `Fleetd` passes. This also needs a decision I am recording rather than leaving open: **a missing `models:` block stays permitted.** Requiring one would break every existing config, including fleet01's, and the allow-list was shipped in #398 as opt-in on purpose. The fix is to make the inert state visible, never to make it fatal.
Author
Owner

The follow-up unit is shipped and live. Closing, with one deviation from what I asked for.

The follow-up in the comment above asked for three distinguishable gate states, in both the startup log and the tool surface, both reading models0(). Checked against the running daemon and the code, not the PR.

The live daemon's startup line:

16:34:33.505 INFO [main] dev.ltms.fleet.Fleetd - model gate (fleetd #422): armed (models: block present; 0 models currently turned off)

The live tool surface (fleet_profiles right now):

{"profiles":["sonnet","local-direct","local","opus","xf","terra","gx","sol"],
 "default":"sonnet",
 "exhaustionDetectionArmed":{"sonnet":false,"local-direct":false,"local":false,"opus":false,
                             "xf":false,"terra":true,"gx":false,"sol":true},
 "modelGateArmed":true}

All three states are distinguishable, which was the actual defect this ticket named — an empty disabledModels() conflating "no block" with "block armed, nothing off":

state startup log tool surface
no models: block not configured (no models: block — nothing is gated, and nothing can be) modelGateArmed: false
block, nothing off armed (models: block present; 0 models currently turned off) modelGateArmed: true, no modelsOff
block, N off armed (N model(s) turned off: [ids]) modelGateArmed: true + modelsOff

The #404 rule is followed, and the code says why. One read, one source:

FleetMcp.java:1268   PeerLauncher.ModelGateState modelGate = workers.modelGateState();
FleetMcp.java:1269   result.put("modelGateArmed", modelGate.configured());
Fleetd.java:241      log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));

and that accessor is the one the gate enforces on:

CompositePeerLauncher.java:805-806   modelGateState()      -> models0()
CompositePeerLauncher.java:570       enforceModelEnabled   -> models0().offIds()
CompositePeerLauncher.java:578       modelOffProfiles      -> models0().offIds()
PeerLauncher.java:341                disabledModels()      -> modelGateState().off()

PeerLauncher.java:337-338 even carries the instruction that stops this drifting again: "do not override this method separately from that one." That is the right shape.

The deviation, stated plainly

I asked for armed (N models allowed, none turned off). The shipped line does not report N. It says "block present", not how many models the allow-list covers, and fleet_profiles does not carry that count either.

I am not reopening for it. The defect this ticket was filed about was the three states reading identically, and that is fixed in both surfaces. The allowed-model count is a nice-to-have that conflates nothing — an operator can read the count from the config file, which is the same place they would go to change it. Recording it so nobody later reads my original wording as an unmet criterion.

Decision from the comment above, still standing

A missing models: block stays permitted. fleet01 has no models: block and must keep working. The fix was to make the inert state visible, never to make it fatal — and modelGateArmed: false is exactly that. Closing.

Units 1 and 2 of the operator's three-part ask are therefore live. Unit 3 (limit monitoring — turn a model off when a subscription limit is reached) is #446, which is a structural problem rather than missing code: the gate's key is hot and the detector's key is deferred, so a model can be turned off at runtime but detection cannot be armed at runtime. That is delegated.

## The follow-up unit is shipped and live. Closing, with one deviation from what I asked for. The follow-up in the comment above asked for three distinguishable gate states, in both the startup log and the tool surface, both reading `models0()`. Checked against the running daemon and the code, not the PR. **The live daemon's startup line:** ``` 16:34:33.505 INFO [main] dev.ltms.fleet.Fleetd - model gate (fleetd #422): armed (models: block present; 0 models currently turned off) ``` **The live tool surface** (`fleet_profiles` right now): ```json {"profiles":["sonnet","local-direct","local","opus","xf","terra","gx","sol"], "default":"sonnet", "exhaustionDetectionArmed":{"sonnet":false,"local-direct":false,"local":false,"opus":false, "xf":false,"terra":true,"gx":false,"sol":true}, "modelGateArmed":true} ``` **All three states are distinguishable**, which was the actual defect this ticket named — an empty `disabledModels()` conflating "no block" with "block armed, nothing off": | state | startup log | tool surface | |---|---|---| | no `models:` block | `not configured (no models: block — nothing is gated, and nothing can be)` | `modelGateArmed: false` | | block, nothing off | `armed (models: block present; 0 models currently turned off)` | `modelGateArmed: true`, no `modelsOff` | | block, N off | `armed (N model(s) turned off: [ids])` | `modelGateArmed: true` + `modelsOff` | **The #404 rule is followed, and the code says why.** One read, one source: ``` FleetMcp.java:1268 PeerLauncher.ModelGateState modelGate = workers.modelGateState(); FleetMcp.java:1269 result.put("modelGateArmed", modelGate.configured()); Fleetd.java:241 log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState())); ``` and that accessor is the one the gate enforces on: ``` CompositePeerLauncher.java:805-806 modelGateState() -> models0() CompositePeerLauncher.java:570 enforceModelEnabled -> models0().offIds() CompositePeerLauncher.java:578 modelOffProfiles -> models0().offIds() PeerLauncher.java:341 disabledModels() -> modelGateState().off() ``` `PeerLauncher.java:337-338` even carries the instruction that stops this drifting again: *"do not override this method separately from that one."* That is the right shape. ### The deviation, stated plainly I asked for `armed (N models allowed, none turned off)`. **The shipped line does not report N.** It says "block present", not how many models the allow-list covers, and `fleet_profiles` does not carry that count either. I am not reopening for it. The defect this ticket was filed about was the three states reading identically, and that is fixed in both surfaces. The allowed-model count is a nice-to-have that conflates nothing — an operator can read the count from the config file, which is the same place they would go to change it. Recording it so nobody later reads my original wording as an unmet criterion. ### Decision from the comment above, still standing **A missing `models:` block stays permitted.** fleet01 has no `models:` block and must keep working. The fix was to make the inert state visible, never to make it fatal — and `modelGateArmed: false` is exactly that. Closing. Units 1 and 2 of the operator's three-part ask are therefore live. Unit 3 (limit monitoring — turn a model off when a subscription limit is reached) is #446, which is a structural problem rather than missing code: the gate's key is hot and the detector's key is deferred, so a model can be turned off at runtime but detection cannot be armed at runtime. That is delegated.
ltms closed this issue 2026-09-10 12:48:01 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#422