fleetd #422: enforce model allow-list on/off at spawn, hot reload #429

Merged
ltms merged 2 commits from worker/422-model-gate-spawn-c29f48-6 into main 2026-09-10 07:44:36 +02:00
Member

What changed

Ships both halves the earlier allow-list ticket left out, in one PR (a gate with no flag always allows; a flag nothing reads does nothing):

  1. Spawn-time enforcement. CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent spawn-refusal reason (operator intent), never layered onto BackendQuarantine/BackendOutagePolicy (backend-reported outage). Wired into the explicit-profile branch alongside enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad. modelOffProfiles() feeds the same exclusion into PlacementContext (new modelOff set) for unqualified spawns via PlacementPolicyUtil, counted into its own bucket in emptyException so "every candidate's model is off" is named as the cause.
  2. Hot on/off switch. FleetConfig.Models.ModelEntry gains enabled (absent/true = on, false = off). Turning a model off never removes it from allow: — validateModels() checks membership only, so an off model stays valid config and a still-configured profile naming it does not refuse reload (this is the whole point of the ticket). Models.offIds() is the one live accessor both the gate and the status report read.
  3. fleet_profiles/GET /profiles parity. PeerLauncher.disabledModels() (default Set.of()) lets both surfaces report off models by reading the exact same accessor the gate reads — the fleetd #404 lesson: a status field must read the source the behaviour reads.
  4. ConfigRef reclassification. models moved from DEFERRED_KEYS to HOT_EXCLUDED_TOP_LEVEL_KEYS: nothing about it is baked into a startup-built object anymore — membership re-validates in full on every reload() via validateAll(), and the on/off half is read live everywhere. New tally: 5 cold, 13 deferred, 3 split, 4 hot-excluded (25 total, unchanged sum). Both ConfigRefTopLevelCoverageTest and ConfigRefTopLevelReportingCoverageTest pass with no new exclusion added just to force green.

Files changed

  • FleetConfig.java — ModelEntry.enabled/isEnabled(), Models.offIds(), javadoc
  • CompositePeerLauncher.java — enforceModelEnabled, modelOffProfiles, live models supplier threaded through every constructor, disabledModels() override
  • PeerLauncher.java — disabledModels() default method
  • PlacementContext.java — new modelOff record component + back-compat 6-arg ctor
  • PlacementPolicyUtil.java — available()/emptyException() gain the model-off bucket
  • FleetMcp.java — profilesView reports modelsOff
  • ConfigRef.java — models reclassified deferred → hot-excluded, class doc + tally updated
  • ConfigRefTopLevelCoverageTest.java — models added to the hot-excluded set + assertion
  • FleetConfigTest.java, CompositePeerLauncherTest.java — new tests, see below

Tests added

  • FleetConfigTest: an old-style allow entry with no enabled: field stays on; an off model is still valid config for validateModels(); one enabled: false entry disables every profile naming it; explicit enabled: true also stays on.
  • CompositePeerLauncherTest: explicit spawn onto an off-model profile is refused with wording distinct from quarantine/cool-off; an enabled-model profile still spawns while another model is off; one model disables every profile naming it; an old-style entry never refuses a spawn; a model absent from allow: is never gated; unqualified spawn skips an off candidate and lands elsewhere; unqualified spawn names model-off as the cause when every candidate is off; a real ConfigRef.reload() proves the on/off switch is hot (turn off → reload → next spawn refused → turn on → reload → next spawn succeeds, no restart); disabledModels() matches the gate exactly.

Self-proof: three deliberate mutations, each restored

  • Mutation A (gate reads a snapshot captured at construction instead of the live supplier): CompositePeerLauncherTest#modelOnOffIsHotReloadedThroughARealConfigRef failed — AssertionFailedError: the very next spawn must see the reload, with no restart ==> Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.
  • Mutation B (candidate-set filter removed, explicit-profile gate kept): CompositePeerLauncherTest#placementSkipsAnOffModelProfileAndRoutesToAnotherOne failed — expected: <sonnet> but was: <local>; CompositePeerLauncherTest#automaticPlacementNamesModelOffWhenEveryCandidateIsOffModel failed — AssertionFailedError: both candidates share the off model — nothing is available ==> Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.
  • Mutation C (explicit-profile gate removed, candidate filter kept): CompositePeerLauncherTest#explicitSpawnOntoAnOffModelProfileIsRefused failed — AssertionFailedError: Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.

All three restored before the final build.

Out of scope (per the ticket, not built)

Automatic detection/re-enable of subscription-limit exhaustion; any change to BackendQuarantine/BackendOutagePolicy; validating that a model id is one the backend actually knows.

Build

mvn clean install from fleetd/: Tests run: 1519, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

## What changed Ships both halves the earlier allow-list ticket left out, in one PR (a gate with no flag always allows; a flag nothing reads does nothing): 1. **Spawn-time enforcement.** `CompositePeerLauncher.enforceModelEnabled` is a FOURTH, independent spawn-refusal reason (operator intent), never layered onto `BackendQuarantine`/`BackendOutagePolicy` (backend-reported outage). Wired into the explicit-profile branch alongside `enforceNotQuarantined`/`enforceNotCoolingOff`/`enforceMaxLoad`. `modelOffProfiles()` feeds the same exclusion into `PlacementContext` (new `modelOff` set) for unqualified spawns via `PlacementPolicyUtil`, counted into its own bucket in `emptyException` so "every candidate's model is off" is named as the cause. 2. **Hot on/off switch.** `FleetConfig.Models.ModelEntry` gains `enabled` (absent/`true` = on, `false` = off). Turning a model off never removes it from `allow:` — `validateModels()` checks membership only, so an off model stays valid config and a still-configured profile naming it does not refuse reload (this is the whole point of the ticket). `Models.offIds()` is the one live accessor both the gate and the status report read. 3. **`fleet_profiles`/`GET /profiles` parity.** `PeerLauncher.disabledModels()` (default `Set.of()`) lets both surfaces report off models by reading the exact same accessor the gate reads — the fleetd #404 lesson: a status field must read the source the behaviour reads. 4. **`ConfigRef` reclassification.** `models` moved from `DEFERRED_KEYS` to `HOT_EXCLUDED_TOP_LEVEL_KEYS`: nothing about it is baked into a startup-built object anymore — membership re-validates in full on every `reload()` via `validateAll()`, and the on/off half is read live everywhere. New tally: 5 cold, 13 deferred, 3 split, 4 hot-excluded (25 total, unchanged sum). Both `ConfigRefTopLevelCoverageTest` and `ConfigRefTopLevelReportingCoverageTest` pass with no new exclusion added just to force green. ## Files changed - `FleetConfig.java` — `ModelEntry.enabled`/`isEnabled()`, `Models.offIds()`, javadoc - `CompositePeerLauncher.java` — `enforceModelEnabled`, `modelOffProfiles`, live `models` supplier threaded through every constructor, `disabledModels()` override - `PeerLauncher.java` — `disabledModels()` default method - `PlacementContext.java` — new `modelOff` record component + back-compat 6-arg ctor - `PlacementPolicyUtil.java` — `available()`/`emptyException()` gain the model-off bucket - `FleetMcp.java` — `profilesView` reports `modelsOff` - `ConfigRef.java` — `models` reclassified deferred → hot-excluded, class doc + tally updated - `ConfigRefTopLevelCoverageTest.java` — `models` added to the hot-excluded set + assertion - `FleetConfigTest.java`, `CompositePeerLauncherTest.java` — new tests, see below ## Tests added - `FleetConfigTest`: an old-style allow entry with no `enabled:` field stays on; an off model is still valid config for `validateModels()`; one `enabled: false` entry disables every profile naming it; explicit `enabled: true` also stays on. - `CompositePeerLauncherTest`: explicit spawn onto an off-model profile is refused with wording distinct from quarantine/cool-off; an enabled-model profile still spawns while another model is off; one model disables every profile naming it; an old-style entry never refuses a spawn; a model absent from `allow:` is never gated; unqualified spawn skips an off candidate and lands elsewhere; unqualified spawn names model-off as the cause when every candidate is off; **a real `ConfigRef.reload()`** proves the on/off switch is hot (turn off → reload → next spawn refused → turn on → reload → next spawn succeeds, no restart); `disabledModels()` matches the gate exactly. ## Self-proof: three deliberate mutations, each restored - **Mutation A** (gate reads a snapshot captured at construction instead of the live supplier): `CompositePeerLauncherTest#modelOnOffIsHotReloadedThroughARealConfigRef` failed — `AssertionFailedError: the very next spawn must see the reload, with no restart ==> Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.` - **Mutation B** (candidate-set filter removed, explicit-profile gate kept): `CompositePeerLauncherTest#placementSkipsAnOffModelProfileAndRoutesToAnotherOne` failed — `expected: <sonnet> but was: <local>`; `CompositePeerLauncherTest#automaticPlacementNamesModelOffWhenEveryCandidateIsOffModel` failed — `AssertionFailedError: both candidates share the off model — nothing is available ==> Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.` - **Mutation C** (explicit-profile gate removed, candidate filter kept): `CompositePeerLauncherTest#explicitSpawnOntoAnOffModelProfileIsRefused` failed — `AssertionFailedError: Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.` All three restored before the final build. ## Out of scope (per the ticket, not built) Automatic detection/re-enable of subscription-limit exhaustion; any change to `BackendQuarantine`/`BackendOutagePolicy`; validating that a model id is one the backend actually knows. ## Build `mvn clean install` from `fleetd/`: `Tests run: 1519, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`.
agent added 1 commit 2026-09-10 07:08:57 +02:00
fleetd #422: enforce the model allow-list on/off state at spawn, hot
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 1m30s
ea02c7b248
Ships the two halves left out of the earlier allow-list ticket in one PR,
since apart they are inert: a gate with no flag always allows, and a flag
nothing reads does nothing.

- FleetConfig.Models.ModelEntry gains `enabled` (default on; absent/true =
  on, false = off). Turning a model off never removes it from `allow:` —
  validateModels() checks membership only, so an off model stays valid
  config and a still-configured profile naming it does not refuse reload.
  Models.offIds() is the one live accessor both the gate and the status
  report read.
- CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent
  spawn-refusal reason (operator intent) — never layered onto
  BackendQuarantine/BackendOutagePolicy, which are backend-reported outage.
  Wired into the explicit-profile branch. modelOffProfiles() feeds the same
  off-model exclusion into PlacementContext for unqualified spawns via
  PlacementPolicyUtil (a new modelOff set, counted into its own bucket in
  emptyException so "all off" is named as the cause, not generic).
  Both read models0(), a live Supplier<FleetConfig.Models>, so a reload
  reaches the very next spawn — no restart.
- PeerLauncher.disabledModels() (default empty) lets fleet_profiles/
  GET /profiles report off models by reading the exact same accessor the
  gate reads (the fleetd #404 lesson: a status field must read the source
  the behaviour reads).
- ConfigRef: `models` reclassified from deferred to hot-excluded — nothing
  about it is baked into a startup-built object anymore; membership is
  re-validated in full on every reload via validateAll(), and the on/off
  half is read live everywhere. Tally: 5 cold, 13 deferred, 3 split, 4
  hot-excluded (25 total). ConfigRefTopLevelCoverageTest and
  ConfigRefTopLevelReportingCoverageTest updated with no new exclusion
  added just to force green.

Tests: FleetConfigTest (old-style fixture stays on; an off model is still
valid config; one model disables every profile naming it),
CompositePeerLauncherTest (explicit refusal wording distinct from
quarantine/cool-off; unqualified spawn skips an off candidate and names
model-off when every candidate is off; a real ConfigRef.reload() proves
the hot path; disabledModels() matches the gate).
Owner

Lead review — not merging yet. One hole, in the default placement policy.

Verified on ea02c7b myself: mvn -f <abs>/fleetd/pom.xml clean install → Tests run: 1519, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS, 0 compile errors.

What is right, and I am keeping all of it:

  • Both spawn paths are gated — enforceModelEnabled for an explicit profile and modelOffProfiles feeding PlacementContext for automatic placement. Covering only one would have been a one-way gate.
  • allow membership stays separate from the on/off state, so validateModels() never refuses a reload because a still-configured profile names an off model. That was the design constraint and it is honoured, and the ModelEntry#enabled javadoc explains why better than the ticket did.
  • Boolean enabled with absent meaning on, plus the back-compat single-arg constructor and the back-compat PlacementContext constructor.
  • disabledModels() reads models0() — the same accessor the gate reads. That is the #404 lesson applied without being asked.
  • modelOff kept as its own bucket in PlacementPolicyUtil.emptyException rather than merged into quarantine or cooling off, so the refusal message never blames the backend for an operator's decision.

The hole: FixedPlacementPolicy never calls PlacementPolicyUtil.available().

The candidate filter went into PlacementPolicyUtil.available(ctx). RoundRobinPlacementPolicy:17 and WeightedRoundRobinPolicy:21 both call it. FixedPlacementPolicy filters inline instead, at FixedPlacementPolicy.java:44 (the default-profile fast path) and :49-50 (the fallback walk), checking quarantined, coolingOff, unreachable and excluded() — and not modelOff().

fixed is the default: PlacementPolicies.fromName returns it for an absent or blank name. So on any fleet that has not set a placement policy, an unqualified fleet_spawn still lands on a profile whose model the operator turned off. The kill switch does not fire on the path most fleets use.

Both new placement tests use PlacementPolicies.weighted(). Every PlacementPolicies.fixed() call in the new tests is an explicit-spawn test, and those bypass placement entirely — so nothing here covers it.

Measured, with this PR's own test wiring and only the policy swapped:

CompositePeerLauncherTest.fixedPlacementSkipsAnOffModelProfileToo (lead probe)
  expected: <sonnet> but was: <local>
  Tests run: 2, Failures: 1     (the weighted test passed; the fixed one did not)

local names deepseek-v4-flash, which the probe turns off. The spawn landed on it.

This is the shape worth naming for next time: FixedPlacementPolicy's own javadoc opens "Four exceptions walk past the default" and lists quarantine, cooling off, unreachable and weight-0. Every previous independent exclusion source had to be added there by hand. #422 adds a fifth and did not. A filter placed in a shared helper is only shared by the callers that call the helper.

Follow-up briefed to the implementer: consult ctx.modelOff() at both filter sites, add the refusal reason in the same style as the other four, fix the two now-stale "quarantined, cooling off, unreachable, or weight-0" messages, update the javadoc's count and bullet list, and add fixed-policy tests for both sites plus the all-off case — with a mutation proof per direction.

## Lead review — not merging yet. One hole, in the default placement policy. Verified on `ea02c7b` myself: `mvn -f <abs>/fleetd/pom.xml clean install` → `Tests run: 1519, Failures: 0, Errors: 0, Skipped: 0`, `BUILD SUCCESS`, 0 compile errors. **What is right, and I am keeping all of it:** - Both spawn paths are gated — `enforceModelEnabled` for an explicit profile and `modelOffProfiles` feeding `PlacementContext` for automatic placement. Covering only one would have been a one-way gate. - `allow` membership stays separate from the on/off state, so `validateModels()` never refuses a reload because a still-configured profile names an off model. That was the design constraint and it is honoured, and the `ModelEntry#enabled` javadoc explains why better than the ticket did. - `Boolean enabled` with absent meaning on, plus the back-compat single-arg constructor and the back-compat `PlacementContext` constructor. - `disabledModels()` reads `models0()` — the same accessor the gate reads. That is the #404 lesson applied without being asked. - `modelOff` kept as its own bucket in `PlacementPolicyUtil.emptyException` rather than merged into quarantine or cooling off, so the refusal message never blames the backend for an operator's decision. **The hole: `FixedPlacementPolicy` never calls `PlacementPolicyUtil.available()`.** The candidate filter went into `PlacementPolicyUtil.available(ctx)`. `RoundRobinPlacementPolicy:17` and `WeightedRoundRobinPolicy:21` both call it. `FixedPlacementPolicy` filters inline instead, at `FixedPlacementPolicy.java:44` (the default-profile fast path) and `:49-50` (the fallback walk), checking `quarantined`, `coolingOff`, `unreachable` and `excluded()` — and not `modelOff()`. `fixed` is the **default**: `PlacementPolicies.fromName` returns it for an absent or blank name. So on any fleet that has not set a placement policy, an unqualified `fleet_spawn` still lands on a profile whose model the operator turned off. The kill switch does not fire on the path most fleets use. Both new placement tests use `PlacementPolicies.weighted()`. Every `PlacementPolicies.fixed()` call in the new tests is an explicit-spawn test, and those bypass placement entirely — so nothing here covers it. **Measured, with this PR's own test wiring and only the policy swapped:** ``` CompositePeerLauncherTest.fixedPlacementSkipsAnOffModelProfileToo (lead probe) expected: <sonnet> but was: <local> Tests run: 2, Failures: 1 (the weighted test passed; the fixed one did not) ``` `local` names `deepseek-v4-flash`, which the probe turns off. The spawn landed on it. This is the shape worth naming for next time: `FixedPlacementPolicy`'s own javadoc opens "**Four** exceptions walk past the default" and lists quarantine, cooling off, unreachable and weight-0. Every previous independent exclusion source had to be added there by hand. #422 adds a fifth and did not. A filter placed in a shared helper is only shared by the callers that call the helper. Follow-up briefed to the implementer: consult `ctx.modelOff()` at both filter sites, add the refusal reason in the same style as the other four, fix the two now-stale "quarantined, cooling off, unreachable, or weight-0" messages, update the javadoc's count and bullet list, and add fixed-policy tests for both sites plus the all-off case — with a mutation proof per direction.
agent added 1 commit 2026-09-10 07:25:44 +02:00
fleetd #422 follow-up: gate FixedPlacementPolicy on modelOff too
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
e2801b9bbc
FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName
returns it for an absent/blank name) and it built its own inline candidate
filter instead of calling PlacementPolicyUtil.available(). That filter checked
quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an
unqualified fleet_spawn on any fleet without an explicit placement: policy
could still land on a profile whose model the operator turned off.

- Add ctx.modelOff() to both filter sites: the default-profile fast path and
  the fallback walk over ctx.candidates().
- Add a modelOff refusal reason to the default-profile reasons list, worded as
  an operator decision ("turned off in models.allow"), matching
  enforceModelEnabled. Quarantine and cooling off still take priority when a
  profile is also model-off, matching CompositePeerLauncher's explicit-spawn
  check order.
- Update the two stale "excluded from automatic selection" messages to name
  model-off, consistent with PlacementPolicyUtil.emptyException.
- Update the class javadoc: four exceptions -> five, with a new bullet for
  model-off (fleetd #422).

Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path),
fixedFallbackWalkSkipsModelOffCandidate (fallback walk),
fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and
fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest
gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of
the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under
PlacementPolicies.fixed(). The two existing weighted()-based tests are
untouched.
ltms merged commit eab8185d7b into main 2026-09-10 07:44:36 +02:00
Sign in to join this conversation.