Features: runtime model on/off (#422) and real architect-slot revocation (#424)

Two new entries, plus a correction to the #398 allow-list entry: its gotcha
said `models:` is deferred and needs a restart. Since #422 the block is hot,
so that line was telling operators the opposite of what the code does.

The #422 entry carries the spawn-path diagram, because the same trap has now
been hit twice: an explicit profile is REFUSED on a bad model while a blank one
is ROUTED AROUND it, so resolving a profile name before calling spawn silently
moves the caller from one path to the other. That is the open regression on
PR #430.

Also records that fleet01 has no `models:` block, so the gate is inert there,
and that automatic limit detection is not built.
Dai Ha
2026-09-10 12:59:59 +07:00
parent efffebef49
commit 9e3db6ee6a
+107 -3
@@ -4891,9 +4891,10 @@ configurations. The list is the source of truth; nothing else can be.
- **Adding a profile means adding its model in the same edit.** Otherwise the daemon will not
start. That is the intended trade: the failure is loud and immediate rather than silent and later.
- **`models:` is deferred, not hot.** A reload validates the new block — so a bad edit is refused
and the running config is kept — but a good edit is reported as needing a restart. Edit-then-reload
is a half-change. Restart with `scripts/redeploy-fleetd.sh`.
- **`models:` used to be deferred. Since fleetd #422 it is hot.** A reload validates the new block,
so a bad edit is still refused and the running config kept — and a good edit now takes effect on
the very next spawn, with no restart. If you are reading an older note that says a `models:` edit
needs a restart, that note is stale.
- **The list is a name gate, not a capability check.** It says the operator permits this id. It does
not say the backend serves it, that the credential may use it, or that the id is spelled the way
the provider spells it. A model on the list can still fail at spawn.
@@ -5072,3 +5073,106 @@ general rule: when a fix adds a parameter so the caller can supply a missing fac
broken.
fleetd #415. Found by the fleet01 lead on their own startup log.
## Turning a model off at runtime, without editing `profiles:`
**What.** Each entry in `models.allow` now takes an `enabled:` flag. Set it to `false` and the
fleet stops spawning on every profile that names that model — at once, with no restart and no edit
to `profiles:`. Set it back to `true` and those profiles are usable again.
```yaml
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
enabled: false # every profile naming this model is now unspawnable
- model: gx/deepseek-v4-flash
```
`enabled:` is optional and defaults to on, so every existing `models:` block keeps working
untouched. An entry that is off **stays in `allow`** — do not delete it. `allow` answers "may the
fleet use this id at all", and `enabled` answers "may it use it right now". Removing the line
instead of turning it off makes every profile naming that model fail validation, and the whole
reload is refused.
`fleet_profiles` and `GET /profiles` report the off set, so a lead can see the gate state without
reading the config file.
**The knob.** `models.allow[].enabled`. Absent means on.
**Why it exists.** A subscription runs out. When it does, the profiles on it must stop taking work
while the rest of the fleet keeps going. Before this the only ways to do that were to edit every
affected `profiles:` entry, or to let each spawn fail against the dead backend and wait for the
quarantine to catch up. One central switch keyed by **model** is the right shape, because a model
is what a subscription sells — several profiles usually share one.
**Where the gate sits, and the one thing to know about it.** `fleet_spawn` takes two different
paths, and they treat a bad profile in opposite ways:
```mermaid
flowchart TD
A["fleet_spawn"] --> B{"did the caller<br/>name a profile?"}
B -->|"yes — an operator override"| C["check quarantine, cooling off,<br/>maxLoad, model-off"]
C --> D["REFUSE the spawn"]
B -->|"no"| E["placement: build the<br/>quarantined / coolingOff / modelOff sets"]
E --> F["the policy SKIPS those profiles"]
F --> G["spawn on the next candidate"]
classDef stop fill:#b7791f,stroke:#7b341e,color:#ffffff;
classDef go fill:#2f855a,stroke:#22543d,color:#ffffff;
class D stop
class G go
```
Naming an off-model profile explicitly is refused, and that is deliberate — it is the operator
overriding, so a clear refusal beats a silent redirect. Leaving the profile blank routes around it
instead. Both behaviours are correct; they are just not the same, and code that resolves a profile
name *before* calling `spawn` silently moves itself from the second path to the first.
**Gotchas.**
- **A host with no `models:` block is not gated.** The whole feature is inert there, and that is
permitted, never fatal. On this fleet, `fleet01` has no `models:` block, so the gate ships green
and does nothing on that host. Check with `grep -c '^models:' fleetd.yaml` before assuming a
second host is protected.
- **The gate is a name gate, like `allow` itself.** Turning a model off stops fleetd spawning on
it. It does not reach a member that is already running.
- **Off is not quarantine, and the message says so.** A model-off refusal names
`models.allow`; a quarantine names the credential and the seconds left. When a profile is both,
quarantine is reported, because a backend fact outranks an operator preference.
- **`fixed` is the default policy, and it applies this filter in two places** — the default
fast path and the fallback walk. The first round of this feature put the filter in the shared
helper `PlacementPolicyUtil.available()`, which `fixed` never calls, so it shipped green and
inert under the default policy with 1519 tests passing. Both new tests had used `weighted()`.
**Still missing.** Detecting a subscription limit and flipping the switch automatically is not
built. Today an operator turns the model off and back on by hand. That decision is open.
fleetd #422.
## Removing an architect slot now actually revokes it
**What.** An architect's extra rights come from a slot in `fleet.architects`. Delete that slot from
the config and reload, and the session bound to it drops to worker rights on its very next request.
Before, it kept full architect rights until the daemon restarted.
**The knob.** None. This is how `fleet.architects` behaves on reload.
**Why it exists.** An architect can do things a worker cannot, so "revoke" has to mean revoke. The
old code flattened the slot list once when the daemon started and never read it again, so removing
a slot only refused the *next* spawn. The session already holding the privilege kept it — a
revocation that did not revoke, with nothing in the logs to say so.
**Gotchas.**
- **The privilege is revoked; the slot stays occupied.** The demoted session keeps its slot key
until it unbinds. That is on purpose: freeing the key at once would let a second terminal claim
the slot the operator was actually trying to shut down, and it would break `unbind`, which needs
the original terminal-to-slot pair intact.
- **The demoted session can still finish its turn.** `fleet_reply` is authorised by terminal
identity, not by role, so a demoted architect ends its turn normally instead of stalling.
- **Do not confuse the binding with the privilege.** They are separate, and wording that treats
them as one thing is what produced the first, wrong version of this fix — the ticket asked for a
test that a bound architect "survives the rebuild", which is true of the binding and false of the
privilege.
fleetd #424.