fleetd asks opencode for a model and never checks it got that model, so a withdrawn name silently bills a paid credential #175

Closed
opened 2026-08-28 00:52:11 +02:00 by ltms · 2 comments
Owner

What happened

The xf profile named opencode/x-preview-f-free. That model was withdrawn from the OpenCode catalogue at some point. opencode does not fail on an unknown -m: it silently falls back to a default.

So every xf member ran on:

{"id":"gpt-5.6-sol","providerID":"openai","variant":"high"}

Read out of opencode's own session database, not from asking a model what it was — a model's self-report is not evidence.

xf carried weight: 80 against terra's 2, so this was 97.6% of dev spawns. Confirmed running that way from at least 2026-08-27 21:48 until the profile was changed on 2026-08-28.

Why it matters

Three separate harms, and the money is the least interesting one.

  1. It billed the paid shared OpenAI credential, at the expensive high variant, for work that was configured to be free.
  2. It escaped the accounting built to prevent exactly this. sol and terra declare credentialId: openai-shared, which is what maxLoad and quarantine use to protect that credential. xf declares no credentialId, because it was supposed to be free. So five concurrent xf members could hammer the shared credential while fleet_list reported sol and terra as barely loaded. The capacity view was wrong in the one direction that costs money.
  3. Nothing anywhere reported it. No log line, no warning, no capacity anomaly. It was found because the operator noticed the member's own UI naming the wrong model.

Root cause

HerdrPeerLauncher passes -m <profile.model> to opencode and treats the process starting as success. There is no read-back. The request is fire-and-forget, and the backend is permitted to substitute silently.

The local catalogue cache at ~/.cache/opencode/models.json made this worse: it listed 29 free models, including a complete and entirely healthy-looking entry for the withdrawn x-preview-f-free. After opencode models --refresh, only 6 free models are actually live. Stale cache is why a name that had stopped working still looked correct to anyone who checked.

Suggested direction

  1. Verify the model after spawn. opencode writes the resolved model into its session database. Read it back once on the first turn and compare it to what the profile asked for. On mismatch: log an ERROR naming both, and quarantine the profile rather than letting it keep spawning.
  2. Fail loudly on an unknown model name at spawn, if opencode offers any way to ask before running. Silent substitution must not be an accepted outcome.
  3. Treat "the backend agreed to my request" as unproven until read back. This is the general shape: the same assumption exists anywhere fleetd sets a backend option by flag and never confirms it took. autoCompactWindow is set the same way and has the same exposure.

Not verified

  • I did not check whether claude-code has the same exposure for its --model flag. Worth checking; the launcher has the same fire-and-forget shape there.
  • I did not establish exactly when x-preview-f-free was withdrawn, only that it was already gone by 2026-08-27.
## What happened The `xf` profile named `opencode/x-preview-f-free`. That model was **withdrawn from the OpenCode catalogue** at some point. `opencode` does not fail on an unknown `-m`: it silently falls back to a default. So every `xf` member ran on: ```json {"id":"gpt-5.6-sol","providerID":"openai","variant":"high"} ``` Read out of opencode's own session database, not from asking a model what it was — a model's self-report is not evidence. `xf` carried `weight: 80` against terra's `2`, so this was **97.6% of dev spawns**. Confirmed running that way from at least 2026-08-27 21:48 until the profile was changed on 2026-08-28. ## Why it matters Three separate harms, and the money is the least interesting one. 1. **It billed the paid shared OpenAI credential**, at the expensive `high` variant, for work that was configured to be free. 2. **It escaped the accounting built to prevent exactly this.** `sol` and `terra` declare `credentialId: openai-shared`, which is what `maxLoad` and quarantine use to protect that credential. `xf` declares no `credentialId`, because it was supposed to be free. So five concurrent `xf` members could hammer the shared credential while `fleet_list` reported `sol` and `terra` as barely loaded. The capacity view was wrong in the one direction that costs money. 3. **Nothing anywhere reported it.** No log line, no warning, no capacity anomaly. It was found because the operator noticed the member's own UI naming the wrong model. ## Root cause `HerdrPeerLauncher` passes `-m <profile.model>` to `opencode` and treats the process starting as success. There is no read-back. The request is fire-and-forget, and the backend is permitted to substitute silently. The local catalogue cache at `~/.cache/opencode/models.json` made this worse: it listed **29** free models, including a complete and entirely healthy-looking entry for the withdrawn `x-preview-f-free`. After `opencode models --refresh`, only **6** free models are actually live. Stale cache is why a name that had stopped working still looked correct to anyone who checked. ## Suggested direction 1. **Verify the model after spawn.** opencode writes the resolved model into its session database. Read it back once on the first turn and compare it to what the profile asked for. On mismatch: log an ERROR naming both, and quarantine the profile rather than letting it keep spawning. 2. **Fail loudly on an unknown model name at spawn**, if opencode offers any way to ask before running. Silent substitution must not be an accepted outcome. 3. **Treat "the backend agreed to my request" as unproven until read back.** This is the general shape: the same assumption exists anywhere fleetd sets a backend option by flag and never confirms it took. `autoCompactWindow` is set the same way and has the same exposure. ## Not verified - I did not check whether `claude-code` has the same exposure for its `--model` flag. Worth checking; the launcher has the same fire-and-forget shape there. - I did not establish exactly when `x-preview-f-free` was withdrawn, only that it was already gone by 2026-08-27.
Author
Owner

Status update. PR #203 is closed as superseded — see the comment there for the detail. Two things changed under it:

  • Its read-back ran inside spawn(), before opencode writes the session row, so the check could never fire.
  • OpenCodeSessionDiscovery has since been rewritten to SQLite (#206 / #207), so the JSON-tree code the PR extended no longer exists.

The corrected requirement for this ticket, once #209 lands:

The model read-back must run from the same late resolve step #209 introduces — a retained PeerHandle plus a re-read driven from roster()/find() while the answer is still unknown — not from spawn(). That step is already opening the session row for agentSessionId, so it can read provider_id/model_id from the same query at no extra cost.

Carry over from #203, both correct:

  • Route a mismatch through the existing ExhaustionSink quarantine path rather than a new mechanism.
  • Absent, incomplete, or unparseable evidence is unknown, not a mismatch, and must never quarantine a working profile.

Add one thing #203 did not have: a test that proves the check actually runs at the point where the row exists. A test that injects the record proves only that the comparison works.

Blocked on #209.

Status update. PR #203 is closed as superseded — see the comment there for the detail. Two things changed under it: - Its read-back ran inside `spawn()`, before opencode writes the session row, so the check could never fire. - `OpenCodeSessionDiscovery` has since been rewritten to SQLite (#206 / #207), so the JSON-tree code the PR extended no longer exists. **The corrected requirement for this ticket**, once #209 lands: The model read-back must run from the same late resolve step #209 introduces — a retained `PeerHandle` plus a re-read driven from `roster()`/`find()` while the answer is still unknown — not from `spawn()`. That step is already opening the `session` row for `agentSessionId`, so it can read `provider_id`/`model_id` from the same query at no extra cost. Carry over from #203, both correct: - Route a mismatch through the existing `ExhaustionSink` quarantine path rather than a new mechanism. - Absent, incomplete, or unparseable evidence is **unknown**, not a mismatch, and must never quarantine a working profile. Add one thing #203 did not have: a test that proves the check actually runs at the point where the row exists. A test that injects the record proves only that the comparison works. Blocked on #209.
Author
Owner

Fixed in PR #231, merged to main (3fbd43f). Live on the daemon since pid 62834, jar 59d8b428b1a2.

What shipped

OpenCodeSessionDiscovery.actualModelForDirectory(directory) reads the session.model column through its own independent query, so a schema without that column can never break id resolution. OpenCodeLauncher.SessionAwareHandle.agentSessionId() then compares what came back against what the profile asked for, and on a real mismatch logs an ERROR naming both and quarantines through the existing ExhaustionSink. No new mechanism.

It runs on #209's late-resolve path, not in spawn(). That is the whole difference from #203, which was closed because its check ran before opencode writes the row and could therefore never fire.

Two corrections to this ticket's own suggested direction

  • The ticket said to read provider_id/model_id from the session row. Those columns do not exist. The real column is model, holding JSON: {"id":"gpt-5.6-terra","providerID":"openai"}.
  • The cost column is 0.0 on every one of the 123 rows, so cost is not a usable signal for this. Anyone tempted to detect paid usage that way should stop there.

The comparison rules, and why they matter more than the check

The profile string and the database JSON are in different formats, and not every profile carries a provider prefix:

profile asks database stores verdict
openai/gpt-5.6-terra {"id":"gpt-5.6-terra","providerID":"openai"} match
gx/deepseek-v4-flash {"id":"deepseek-v4-flash","providerID":"gx"} match
deepseek-v4-flash (no /) any provider match on id alone
claude-sonnet-5 no row exists not applicable

A naive string compare marks all of these as mismatches and quarantines the whole fleet. That is worse than the bug: this one costs money, that one stops all work. claude-code is excluded structurally — SessionAwareHandle is only ever built by OpenCodeLauncher.spawn().

One gap found at review

parseModel accepts JSON with an id and no providerID, returning a record with a null provider. A provider-prefixed profile then read that as a mismatch and would have quarantined on incomplete evidence, which acceptance rule 4 forbids. I checked reachability before calling it a defect — all 123 live rows carry both fields, so it could not fire today — and asked for the fix anyway, because the consequence is asymmetric and this ticket exists precisely because a backend changed behaviour silently.

Round 2 narrowed the comparison to "compare the provider only when both sides have one", keeping the id comparison (the part that actually caught xf) fully intact. The test that proves it narrowed the rule rather than switching the check off is the one where the id genuinely differs and it still quarantines.

Verification

Independent build by the lead: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.

Live probe on the redeployed daemon — a terra member on a fresh worktree, which is the case that must NOT trigger:

profile terra asks : openai/gpt-5.6-terra
database recorded  : {"id":"gpt-5.6-terra","providerID":"openai"}
agentSessionId     : ses_f9e358628ffeNZrbb0VUQTmICN   (so the check ran)
model-mismatch ERR : none
terra quarantine   : none — free: 1
ERROR lines since restart : 0

The resolved agentSessionId is what proves the check actually executed rather than being skipped.

Still open

The two "not verified" items in the original report stand: I did not check whether claude-code has the same exposure for --model, and the withdrawal date of x-preview-f-free was never established. The general shape — setting a backend option by flag and never confirming it took — is now filed as #232 for autoCompactWindow.

Fixed in PR #231, merged to `main` (3fbd43f). Live on the daemon since pid 62834, jar `59d8b428b1a2`. ## What shipped `OpenCodeSessionDiscovery.actualModelForDirectory(directory)` reads the `session.model` column through its own independent query, so a schema without that column can never break id resolution. `OpenCodeLauncher.SessionAwareHandle.agentSessionId()` then compares what came back against what the profile asked for, and on a real mismatch logs an ERROR naming both and quarantines through the existing `ExhaustionSink`. No new mechanism. **It runs on #209's late-resolve path, not in `spawn()`.** That is the whole difference from #203, which was closed because its check ran before opencode writes the row and could therefore never fire. ## Two corrections to this ticket's own suggested direction - The ticket said to read `provider_id`/`model_id` from the session row. **Those columns do not exist.** The real column is `model`, holding JSON: `{"id":"gpt-5.6-terra","providerID":"openai"}`. - The `cost` column is `0.0` on every one of the 123 rows, so cost is not a usable signal for this. Anyone tempted to detect paid usage that way should stop there. ## The comparison rules, and why they matter more than the check The profile string and the database JSON are in different formats, and not every profile carries a provider prefix: | profile asks | database stores | verdict | |---|---|---| | `openai/gpt-5.6-terra` | `{"id":"gpt-5.6-terra","providerID":"openai"}` | match | | `gx/deepseek-v4-flash` | `{"id":"deepseek-v4-flash","providerID":"gx"}` | match | | `deepseek-v4-flash` (no `/`) | any provider | match on id alone | | `claude-sonnet-5` | no row exists | not applicable | A naive string compare marks all of these as mismatches and quarantines the whole fleet. That is worse than the bug: this one costs money, that one stops all work. claude-code is excluded structurally — `SessionAwareHandle` is only ever built by `OpenCodeLauncher.spawn()`. ## One gap found at review `parseModel` accepts JSON with an `id` and no `providerID`, returning a record with a null provider. A provider-prefixed profile then read that as a mismatch and would have quarantined on **incomplete** evidence, which acceptance rule 4 forbids. I checked reachability before calling it a defect — all 123 live rows carry both fields, so it could not fire today — and asked for the fix anyway, because the consequence is asymmetric and this ticket exists precisely because a backend changed behaviour silently. Round 2 narrowed the comparison to "compare the provider only when **both** sides have one", keeping the id comparison (the part that actually caught `xf`) fully intact. The test that proves it narrowed the rule rather than switching the check off is the one where the id genuinely differs and it still quarantines. ## Verification Independent build by the lead: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped. Live probe on the redeployed daemon — a `terra` member on a fresh worktree, which is the case that must NOT trigger: ``` profile terra asks : openai/gpt-5.6-terra database recorded : {"id":"gpt-5.6-terra","providerID":"openai"} agentSessionId : ses_f9e358628ffeNZrbb0VUQTmICN (so the check ran) model-mismatch ERR : none terra quarantine : none — free: 1 ERROR lines since restart : 0 ``` The resolved `agentSessionId` is what proves the check actually executed rather than being skipped. ## Still open The two "not verified" items in the original report stand: I did not check whether `claude-code` has the same exposure for `--model`, and the withdrawal date of `x-preview-f-free` was never established. The general shape — setting a backend option by flag and never confirming it took — is now filed as #232 for `autoCompactWindow`.
ltms closed this issue 2026-09-02 13:06:09 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#175