fleetd #422: enforce model allow-list on/off at spawn, hot reload #429
Reference in New Issue
Block a user
Delete Branch "worker/422-model-gate-spawn-c29f48-6"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What changed
Ships both halves the earlier allow-list ticket left out, in one PR (a gate with no flag always allows; a flag nothing reads does nothing):
CompositePeerLauncher.enforceModelEnabledis a FOURTH, independent spawn-refusal reason (operator intent), never layered ontoBackendQuarantine/BackendOutagePolicy(backend-reported outage). Wired into the explicit-profile branch alongsideenforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad.modelOffProfiles()feeds the same exclusion intoPlacementContext(newmodelOffset) for unqualified spawns viaPlacementPolicyUtil, counted into its own bucket inemptyExceptionso "every candidate's model is off" is named as the cause.FleetConfig.Models.ModelEntrygainsenabled(absent/true= on,false= off). Turning a model off never removes it fromallow:—validateModels()checks membership only, so an off model stays valid config and a still-configured profile naming it does not refuse reload (this is the whole point of the ticket).Models.offIds()is the one live accessor both the gate and the status report read.fleet_profiles/GET /profilesparity.PeerLauncher.disabledModels()(defaultSet.of()) lets both surfaces report off models by reading the exact same accessor the gate reads — the fleetd #404 lesson: a status field must read the source the behaviour reads.ConfigRefreclassification.modelsmoved fromDEFERRED_KEYStoHOT_EXCLUDED_TOP_LEVEL_KEYS: nothing about it is baked into a startup-built object anymore — membership re-validates in full on everyreload()viavalidateAll(), and the on/off half is read live everywhere. New tally: 5 cold, 13 deferred, 3 split, 4 hot-excluded (25 total, unchanged sum). BothConfigRefTopLevelCoverageTestandConfigRefTopLevelReportingCoverageTestpass with no new exclusion added just to force green.Files changed
FleetConfig.java—ModelEntry.enabled/isEnabled(),Models.offIds(), javadocCompositePeerLauncher.java—enforceModelEnabled,modelOffProfiles, livemodelssupplier threaded through every constructor,disabledModels()overridePeerLauncher.java—disabledModels()default methodPlacementContext.java— newmodelOffrecord component + back-compat 6-arg ctorPlacementPolicyUtil.java—available()/emptyException()gain the model-off bucketFleetMcp.java—profilesViewreportsmodelsOffConfigRef.java—modelsreclassified deferred → hot-excluded, class doc + tally updatedConfigRefTopLevelCoverageTest.java—modelsadded to the hot-excluded set + assertionFleetConfigTest.java,CompositePeerLauncherTest.java— new tests, see belowTests added
FleetConfigTest: an old-style allow entry with noenabled:field stays on; an off model is still valid config forvalidateModels(); oneenabled: falseentry disables every profile naming it; explicitenabled: truealso stays on.CompositePeerLauncherTest: explicit spawn onto an off-model profile is refused with wording distinct from quarantine/cool-off; an enabled-model profile still spawns while another model is off; one model disables every profile naming it; an old-style entry never refuses a spawn; a model absent fromallow:is never gated; unqualified spawn skips an off candidate and lands elsewhere; unqualified spawn names model-off as the cause when every candidate is off; a realConfigRef.reload()proves the on/off switch is hot (turn off → reload → next spawn refused → turn on → reload → next spawn succeeds, no restart);disabledModels()matches the gate exactly.Self-proof: three deliberate mutations, each restored
CompositePeerLauncherTest#modelOnOffIsHotReloadedThroughARealConfigReffailed —AssertionFailedError: the very next spawn must see the reload, with no restart ==> Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.CompositePeerLauncherTest#placementSkipsAnOffModelProfileAndRoutesToAnotherOnefailed —expected: <sonnet> but was: <local>;CompositePeerLauncherTest#automaticPlacementNamesModelOffWhenEveryCandidateIsOffModelfailed —AssertionFailedError: both candidates share the off model — nothing is available ==> Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.CompositePeerLauncherTest#explicitSpawnOntoAnOffModelProfileIsRefusedfailed —AssertionFailedError: Expected dev.ltms.fleet.placement.PlacementException to be thrown, but nothing was thrown.All three restored before the final build.
Out of scope (per the ticket, not built)
Automatic detection/re-enable of subscription-limit exhaustion; any change to
BackendQuarantine/BackendOutagePolicy; validating that a model id is one the backend actually knows.Build
mvn clean installfromfleetd/:Tests run: 1519, Failures: 0, Errors: 0, Skipped: 0—BUILD SUCCESS.Lead review — not merging yet. One hole, in the default placement policy.
Verified on
ea02c7bmyself:mvn -f <abs>/fleetd/pom.xml clean install→Tests run: 1519, Failures: 0, Errors: 0, Skipped: 0,BUILD SUCCESS, 0 compile errors.What is right, and I am keeping all of it:
enforceModelEnabledfor an explicit profile andmodelOffProfilesfeedingPlacementContextfor automatic placement. Covering only one would have been a one-way gate.allowmembership stays separate from the on/off state, sovalidateModels()never refuses a reload because a still-configured profile names an off model. That was the design constraint and it is honoured, and theModelEntry#enabledjavadoc explains why better than the ticket did.Boolean enabledwith absent meaning on, plus the back-compat single-arg constructor and the back-compatPlacementContextconstructor.disabledModels()readsmodels0()— the same accessor the gate reads. That is the #404 lesson applied without being asked.modelOffkept as its own bucket inPlacementPolicyUtil.emptyExceptionrather than merged into quarantine or cooling off, so the refusal message never blames the backend for an operator's decision.The hole:
FixedPlacementPolicynever callsPlacementPolicyUtil.available().The candidate filter went into
PlacementPolicyUtil.available(ctx).RoundRobinPlacementPolicy:17andWeightedRoundRobinPolicy:21both call it.FixedPlacementPolicyfilters inline instead, atFixedPlacementPolicy.java:44(the default-profile fast path) and:49-50(the fallback walk), checkingquarantined,coolingOff,unreachableandexcluded()— and notmodelOff().fixedis the default:PlacementPolicies.fromNamereturns it for an absent or blank name. So on any fleet that has not set a placement policy, an unqualifiedfleet_spawnstill lands on a profile whose model the operator turned off. The kill switch does not fire on the path most fleets use.Both new placement tests use
PlacementPolicies.weighted(). EveryPlacementPolicies.fixed()call in the new tests is an explicit-spawn test, and those bypass placement entirely — so nothing here covers it.Measured, with this PR's own test wiring and only the policy swapped:
localnamesdeepseek-v4-flash, which the probe turns off. The spawn landed on it.This is the shape worth naming for next time:
FixedPlacementPolicy's own javadoc opens "Four exceptions walk past the default" and lists quarantine, cooling off, unreachable and weight-0. Every previous independent exclusion source had to be added there by hand. #422 adds a fifth and did not. A filter placed in a shared helper is only shared by the callers that call the helper.Follow-up briefed to the implementer: consult
ctx.modelOff()at both filter sites, add the refusal reason in the same style as the other four, fix the two now-stale "quarantined, cooling off, unreachable, or weight-0" messages, update the javadoc's count and bullet list, and add fixed-policy tests for both sites plus the all-off case — with a mutation proof per direction.FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName returns it for an absent/blank name) and it built its own inline candidate filter instead of calling PlacementPolicyUtil.available(). That filter checked quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an unqualified fleet_spawn on any fleet without an explicit placement: policy could still land on a profile whose model the operator turned off. - Add ctx.modelOff() to both filter sites: the default-profile fast path and the fallback walk over ctx.candidates(). - Add a modelOff refusal reason to the default-profile reasons list, worded as an operator decision ("turned off in models.allow"), matching enforceModelEnabled. Quarantine and cooling off still take priority when a profile is also model-off, matching CompositePeerLauncher's explicit-spawn check order. - Update the two stale "excluded from automatic selection" messages to name model-off, consistent with PlacementPolicyUtil.emptyException. - Update the class javadoc: four exceptions -> five, with a new bullet for model-off (fleetd #422). Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path), fixedFallbackWalkSkipsModelOffCandidate (fallback walk), fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under PlacementPolicies.fixed(). The two existing weighted()-based tests are untouched.