Features: opencode model read-back and silent-substitution quarantine (#175)

Dai Ha
2026-09-02 18:06:36 +07:00
parent a6dcfa7789
commit c5033578a5
+37
@@ -79,6 +79,7 @@ six weeks, and the table alone will not carry it.
| [An exhausted backend is quarantined even from a chrome-only pane](#an-exhausted-backend-is-quarantined-even-from-a-chrome-only-pane) | automatic | #211 | `inject/CompletionResolver` |
| [The credential scrub follows the member's own user](#the-credential-scrub-follows-the-members-own-user) | `memberLoginShell:` + `worktreeGroup:` | #213 | `member/HerdrPeerLauncher` |
| [An opencode member's config follows the member's own user](#an-opencode-members-config-follows-the-members-own-user) | `memberHerdrSocket:` + `worktreeGroup:` | #219 | `member/OpenCodeLauncher` |
| [The model opencode actually ran is read back and checked](#the-model-opencode-actually-ran-is-read-back-and-checked) | automatic (opencode profiles) | #175 | `member/OpenCodeSessionDiscovery`, `member/OpenCodeLauncher` |
| [Every file a member must read follows the member's own user](#every-file-a-member-must-read-follows-the-members-own-user) | `memberHerdrSocket:` + `worktreeGroup:` | #222 #224 | `member/ClaudeCodeLauncher`, `session/GitWorktrees` |
| [A role fleetd cannot bind is refused](#a-role-fleetd-cannot-bind-is-refused-not-quietly-downgraded) | automatic | #123 | `auth/MemberRegistry` |
@@ -2813,6 +2814,42 @@ beats an answer read from the wrong directory.
---
## The model opencode actually ran is read back and checked
**What.** After an opencode member's session row appears, fleetd reads the model opencode actually
resolved and compares it to what the profile asked for. On a real mismatch it logs an ERROR naming
both models and quarantines that profile's credential through the existing exhaustion path.
**On.** Automatic, for opencode profiles that configure a `model:`. claude-code profiles are not
affected — they have no opencode session row, and the check is structurally unreachable for them.
**Why it exists.** `opencode` does not fail on an unknown `-m`. It silently falls back to a default.
The `xf` profile named a model that had been withdrawn from the catalogue, so every `xf` member ran
on a **paid** model at the expensive `high` variant while the profile was configured to be free. That
was 97.6% of dev spawns for about a day.
The second harm is worse than the bill. `xf` declared no `credentialId`, because it was supposed to
be free — so five concurrent members could hammer the shared paid credential while `fleet_list`
reported the paid profiles as barely loaded. The accounting built to protect that credential was
looking the other way. Nothing logged any of it; it was found because the operator happened to notice
the member's own UI naming the wrong model.
The general shape is worth naming: **fleetd set a backend option by flag and treated the process
starting as proof the option took effect.** #232 tracks the same assumption for `autoCompactWindow`.
**The gotcha: the check runs late, and "unknown" is never a mismatch.** opencode writes its session
row only when the session is first persisted, so a check inside `spawn()` runs before the evidence
exists and can never fire — an earlier attempt shipped exactly that and was closed. This one runs
from the same late-resolve step that fills in `agentSessionId`, so a resolved `agentSessionId` is
also the sign the model check ran.
The comparison is deliberately narrow, because a false positive stops all work on a credential and is
worse than the bug it prevents. The profile string and the stored JSON are in different formats
(`openai/gpt-5.6-terra` against `{"id":"gpt-5.6-terra","providerID":"openai"}`), and a profile may
carry no provider prefix at all. So the id is always compared, the provider only when **both** sides
have one, and absent, incomplete or unparseable evidence is UNKNOWN — never a mismatch, never a
quarantine.
## Every file a member must read follows the member's own user
**What.** Two more places where fleetd used to write a file into its own process's filesystem and