From 9e3db6ee6a9d7d0115d2c3f5e371194510cc4473 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Thu, 10 Sep 2026 12:59:59 +0700 Subject: [PATCH] Features: runtime model on/off (#422) and real architect-slot revocation (#424) Two new entries, plus a correction to the #398 allow-list entry: its gotcha said `models:` is deferred and needs a restart. Since #422 the block is hot, so that line was telling operators the opposite of what the code does. The #422 entry carries the spawn-path diagram, because the same trap has now been hit twice: an explicit profile is REFUSED on a bad model while a blank one is ROUTED AROUND it, so resolving a profile name before calling spawn silently moves the caller from one path to the other. That is the open regression on PR #430. Also records that fleet01 has no `models:` block, so the gate is inert there, and that automatic limit detection is not built. --- 11-Features.md | 110 +++++++++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 107 insertions(+), 3 deletions(-) diff --git a/11-Features.md b/11-Features.md index 9afb314..b57e438 100644 --- a/11-Features.md +++ b/11-Features.md @@ -4891,9 +4891,10 @@ configurations. The list is the source of truth; nothing else can be. - **Adding a profile means adding its model in the same edit.** Otherwise the daemon will not start. That is the intended trade: the failure is loud and immediate rather than silent and later. -- **`models:` is deferred, not hot.** A reload validates the new block — so a bad edit is refused - and the running config is kept — but a good edit is reported as needing a restart. Edit-then-reload - is a half-change. Restart with `scripts/redeploy-fleetd.sh`. +- **`models:` used to be deferred. Since fleetd #422 it is hot.** A reload validates the new block, + so a bad edit is still refused and the running config kept — and a good edit now takes effect on + the very next spawn, with no restart. If you are reading an older note that says a `models:` edit + needs a restart, that note is stale. - **The list is a name gate, not a capability check.** It says the operator permits this id. It does not say the backend serves it, that the credential may use it, or that the id is spelled the way the provider spells it. A model on the list can still fail at spawn. @@ -5072,3 +5073,106 @@ general rule: when a fix adds a parameter so the caller can supply a missing fac broken. fleetd #415. Found by the fleet01 lead on their own startup log. + +## Turning a model off at runtime, without editing `profiles:` + +**What.** Each entry in `models.allow` now takes an `enabled:` flag. Set it to `false` and the +fleet stops spawning on every profile that names that model — at once, with no restart and no edit +to `profiles:`. Set it back to `true` and those profiles are usable again. + +```yaml +models: + allow: + - model: claude-sonnet-5 + - model: openai/gpt-5.6-terra + enabled: false # every profile naming this model is now unspawnable + - model: gx/deepseek-v4-flash +``` + +`enabled:` is optional and defaults to on, so every existing `models:` block keeps working +untouched. An entry that is off **stays in `allow`** — do not delete it. `allow` answers "may the +fleet use this id at all", and `enabled` answers "may it use it right now". Removing the line +instead of turning it off makes every profile naming that model fail validation, and the whole +reload is refused. + +`fleet_profiles` and `GET /profiles` report the off set, so a lead can see the gate state without +reading the config file. + +**The knob.** `models.allow[].enabled`. Absent means on. + +**Why it exists.** A subscription runs out. When it does, the profiles on it must stop taking work +while the rest of the fleet keeps going. Before this the only ways to do that were to edit every +affected `profiles:` entry, or to let each spawn fail against the dead backend and wait for the +quarantine to catch up. One central switch keyed by **model** is the right shape, because a model +is what a subscription sells — several profiles usually share one. + +**Where the gate sits, and the one thing to know about it.** `fleet_spawn` takes two different +paths, and they treat a bad profile in opposite ways: + +```mermaid +flowchart TD + A["fleet_spawn"] --> B{"did the caller
name a profile?"} + B -->|"yes — an operator override"| C["check quarantine, cooling off,
maxLoad, model-off"] + C --> D["REFUSE the spawn"] + B -->|"no"| E["placement: build the
quarantined / coolingOff / modelOff sets"] + E --> F["the policy SKIPS those profiles"] + F --> G["spawn on the next candidate"] + classDef stop fill:#b7791f,stroke:#7b341e,color:#ffffff; + classDef go fill:#2f855a,stroke:#22543d,color:#ffffff; + class D stop + class G go +``` + +Naming an off-model profile explicitly is refused, and that is deliberate — it is the operator +overriding, so a clear refusal beats a silent redirect. Leaving the profile blank routes around it +instead. Both behaviours are correct; they are just not the same, and code that resolves a profile +name *before* calling `spawn` silently moves itself from the second path to the first. + +**Gotchas.** + +- **A host with no `models:` block is not gated.** The whole feature is inert there, and that is + permitted, never fatal. On this fleet, `fleet01` has no `models:` block, so the gate ships green + and does nothing on that host. Check with `grep -c '^models:' fleetd.yaml` before assuming a + second host is protected. +- **The gate is a name gate, like `allow` itself.** Turning a model off stops fleetd spawning on + it. It does not reach a member that is already running. +- **Off is not quarantine, and the message says so.** A model-off refusal names + `models.allow`; a quarantine names the credential and the seconds left. When a profile is both, + quarantine is reported, because a backend fact outranks an operator preference. +- **`fixed` is the default policy, and it applies this filter in two places** — the default + fast path and the fallback walk. The first round of this feature put the filter in the shared + helper `PlacementPolicyUtil.available()`, which `fixed` never calls, so it shipped green and + inert under the default policy with 1519 tests passing. Both new tests had used `weighted()`. + +**Still missing.** Detecting a subscription limit and flipping the switch automatically is not +built. Today an operator turns the model off and back on by hand. That decision is open. + +fleetd #422. + +## Removing an architect slot now actually revokes it + +**What.** An architect's extra rights come from a slot in `fleet.architects`. Delete that slot from +the config and reload, and the session bound to it drops to worker rights on its very next request. +Before, it kept full architect rights until the daemon restarted. + +**The knob.** None. This is how `fleet.architects` behaves on reload. + +**Why it exists.** An architect can do things a worker cannot, so "revoke" has to mean revoke. The +old code flattened the slot list once when the daemon started and never read it again, so removing +a slot only refused the *next* spawn. The session already holding the privilege kept it — a +revocation that did not revoke, with nothing in the logs to say so. + +**Gotchas.** + +- **The privilege is revoked; the slot stays occupied.** The demoted session keeps its slot key + until it unbinds. That is on purpose: freeing the key at once would let a second terminal claim + the slot the operator was actually trying to shut down, and it would break `unbind`, which needs + the original terminal-to-slot pair intact. +- **The demoted session can still finish its turn.** `fleet_reply` is authorised by terminal + identity, not by role, so a demoted architect ends its turn normally instead of stalling. +- **Do not confuse the binding with the privilege.** They are separate, and wording that treats + them as one thing is what produced the first, wrong version of this fix — the ticket asked for a + test that a bound architect "survives the rebuild", which is true of the binding and false of the + privilege. + +fleetd #424.