A fast model trips the 2000ms crash floor: a real answer is reported as a failed turn #376

Closed
opened 2026-09-08 23:06:12 +02:00 by ltms · 2 comments
Owner

What happened

Measured on fleet01 on 2026-09-09, on a daemon built from 127e683.

I spawned a dev member on the xf profile (model opencode/mimo-v2.5-free) and sent it a short question. The turn came back like this:

{"status":"failed","detail":"member term_65aff0bbd321811 went BUSY -> DONE in 1575ms
 (floor 2000ms) — too fast to be real work, most likely a backend error before any
 work started: ...

The member was fine. The scrape inside that same detail field holds its real answer:

I am running on opencode/mimo-v2.5-free.
▣  Dev · MiMo V2.5 Free · 6.0s

It also showed $0.00 spent and fleet Connected. So the model ran, used the bridge, and answered correctly — in 1575ms, which is below the floor.

Why it matters

CompletionResolver.MIN_TURN_NANOS (fleetd#164) treats a BUSY -> DONE inside 2000ms as a crash signature. That was a fair reading when every backend took seconds to produce a first token. It is no longer fair. On fleet01 I timed four free OpenCode models on a trivial prompt today:

model reply "OK"
ling-3.0-flash-fin-free 5s
mimo-v2.5-free 6s
nemotron-3.5-lightning-free 7s
nemotron-3-ultra-free 90s, timed out with no answer

Those are whole-process times for opencode run, including start-up. Inside an already-running pane the model itself answers far faster, and 1575ms is what that looks like.

The consequence is not a lost answer — the scrape still carries it. The consequence is a wrong status. A lead reading status: failed will re-send, or mark the member broken, or pick a different profile. The faster the model, the more often this fires. It punishes exactly the backend you want.

What I am not claiming

I have not measured how often this hits a real worker brief, which takes much longer than 1575ms. My one data point is a short probe question. The failure mode is real; its frequency on normal work is unmeasured.

Possible directions (not a decision)

  • The floor exists to catch a crash. A crash and a fast answer differ in the pane content, not only in the timing. The scrape already happens — the floor could ask whether the pane holds an answer before calling it a crash.
  • Or make the floor per-profile, so a fast model can declare a lower one.
  • Or keep the floor but stop reporting failed when the scrape looks like a real reply.

Measure before fixing. The first question is which real turns actually land inside the floor, and whether any of them is genuinely a crash.

Where

  • fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java:68 — the MIN_TURN_NANOS javadoc
  • :311 — the floor check
  • :545, :564 — the failure message quoted above
## What happened Measured on fleet01 on 2026-09-09, on a daemon built from `127e683`. I spawned a `dev` member on the `xf` profile (model `opencode/mimo-v2.5-free`) and sent it a short question. The turn came back like this: ``` {"status":"failed","detail":"member term_65aff0bbd321811 went BUSY -> DONE in 1575ms (floor 2000ms) — too fast to be real work, most likely a backend error before any work started: ... ``` The member was fine. The scrape inside that same `detail` field holds its real answer: ``` I am running on opencode/mimo-v2.5-free. ▣ Dev · MiMo V2.5 Free · 6.0s ``` It also showed `$0.00 spent` and `fleet Connected`. So the model ran, used the bridge, and answered correctly — in 1575ms, which is below the floor. ## Why it matters `CompletionResolver.MIN_TURN_NANOS` (fleetd#164) treats a `BUSY -> DONE` inside 2000ms as a crash signature. That was a fair reading when every backend took seconds to produce a first token. It is no longer fair. On fleet01 I timed four free OpenCode models on a trivial prompt today: | model | reply "OK" | |---|---| | `ling-3.0-flash-fin-free` | 5s | | `mimo-v2.5-free` | 6s | | `nemotron-3.5-lightning-free` | 7s | | `nemotron-3-ultra-free` | 90s, timed out with no answer | Those are whole-process times for `opencode run`, including start-up. Inside an already-running pane the model itself answers far faster, and 1575ms is what that looks like. The consequence is not a lost answer — the scrape still carries it. The consequence is a **wrong status**. A lead reading `status: failed` will re-send, or mark the member broken, or pick a different profile. The faster the model, the more often this fires. It punishes exactly the backend you want. ## What I am not claiming I have not measured how often this hits a real worker brief, which takes much longer than 1575ms. My one data point is a short probe question. The failure mode is real; its frequency on normal work is unmeasured. ## Possible directions (not a decision) - The floor exists to catch a crash. A crash and a fast answer differ in the **pane content**, not only in the timing. The scrape already happens — the floor could ask whether the pane holds an answer before calling it a crash. - Or make the floor per-profile, so a fast model can declare a lower one. - Or keep the floor but stop reporting `failed` when the scrape looks like a real reply. Measure before fixing. The first question is which real turns actually land inside the floor, and whether any of them is genuinely a crash. ## Where - `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java:68` — the `MIN_TURN_NANOS` javadoc - `:311` — the floor check - `:545`, `:564` — the failure message quoted above
Author
Owner

Correction to the date in the issue body: the measurement was 2026-09-08 UTC, not 2026-09-09.

Both dates came from a real clock, which is why this is worth writing down rather than just editing away. The fleet01 lead flagged the mismatch, so I compared three clocks:

  mac  UTC   2026-09-08T21:07:03+0000    mac local 2026-09-09T04:07 +07
  f01  UTC   2026-09-08T21:06:59+0000    (fleet01 runs UTC)
  gitea      Tue, 08 Sep 2026 21:06:59 GMT

They agree within 4 seconds. No clock is skewed. The Mac runs at +07, so its local date had already rolled over to the 9th while fleet01 and the forge were still on the 8th.

The reason this matters here: every timestamp in this issue is read against journalctl --user -u fleetd on fleet01, and that journal is in UTC. A date stamped in Mac local time does not line up with the log lines it describes. So dates about fleet01 belong in UTC.

Nothing else in the issue changes: the floor is still 2000ms, the measured turn was still 1575ms, and the model times are unaffected.

Correction to the date in the issue body: the measurement was **2026-09-08 UTC**, not 2026-09-09. Both dates came from a real clock, which is why this is worth writing down rather than just editing away. The fleet01 lead flagged the mismatch, so I compared three clocks: ``` mac UTC 2026-09-08T21:07:03+0000 mac local 2026-09-09T04:07 +07 f01 UTC 2026-09-08T21:06:59+0000 (fleet01 runs UTC) gitea Tue, 08 Sep 2026 21:06:59 GMT ``` They agree within 4 seconds. No clock is skewed. The Mac runs at +07, so its local date had already rolled over to the 9th while fleet01 and the forge were still on the 8th. The reason this matters here: every timestamp in this issue is read against `journalctl --user -u fleetd` on fleet01, and that journal is in UTC. A date stamped in Mac local time does not line up with the log lines it describes. So dates about fleet01 belong in UTC. Nothing else in the issue changes: the floor is still 2000ms, the measured turn was still 1575ms, and the model times are unaffected.
Author
Owner

Fixed in 2302b3b, but not the way this ticket suggested. Writing down what was rejected, because the rejected version is the obvious one and someone will try it again.

What shipped

The verdict is unchanged: a turn that settles inside MIN_TURN_NANOS is still FAILED, and the floor value is still 2000ms. Only the claim about why is withdrawn.

The old message asserted "most likely a backend error before any work started". When no errorPattern matches, that cause is a guess — and both hosts run with none configured, which they log at boot as backend-error classification: off. A reader who believes the guess stops looking at the pane, and the pane is where the answer actually was.

What I rejected, and how I know it was wrong

The ticket's first suggested direction — inspect the pane inside the floor, resolve a COMPLETION when it holds a real reply — was implemented. Its test for "a real reply" was: non-blank, plus a ., ! or ? anywhere in the text.

That is unsafe twice over:

  1. lastAssistantBlock falls back to the whole pane when it finds no ⏺ marker (CompletionResolver.java:730). On a crash there is no assistant block, so the candidate "reply" is the entire screen.
  2. A crash pane almost always contains a full stop — in a file path, a version number, or a hostname.

I did not argue this from reading. I ran that implementation against the new guard test:

aPlausibleLookingReplyInsideTheFloorStillFails:380
  expected: <FAILED> but was: <COMPLETION>

pane: Error: connection reset while loading src/main/java/Foo.java v1.2.3

A backend failure reported as a successful completion. That is a loud wrong answer traded for a silent one, which is the one trade this path must never make.

Why the obvious repair does not work either

Requiring the ⏺ marker as positive evidence of a reply would be safe. But ⏺ is Claude Code chrome, and an opencode pane never carries it — and an opencode member on mimo-v2.5-free is what raised this ticket. So the safe rule would fix nothing for the backend that actually needs it.

There is no reliable cross-backend marker for "this is a real reply." That is the finding, and it is now in the failTooFast javadoc so the next reader meets it before rewriting this.

Tests

  • aPlausibleLookingReplyInsideTheFloorStillFails — pins the safety property against exactly the rejected approach. It passes today and fails against that implementation. Verified by running it, not assumed.
  • theTooFastFailureDoesNotAssertACauseItCannotKnow — pins the wording.
  • aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink was changed. It asserted the phrase "too fast to be real work", which carried the withdrawn claim. It now asserts what it was really guarding: the floor alone fails the turn, the reason stays generic, the pane is carried, and the typed sink is never notified.

mvn clean install: 1441 tests, 0 failures, 0 errors.

What is still not fixed

A lead reading status: failed on a genuinely fast real turn still sees a failure. That is deliberate — it is the loud direction — but it is not free, and this ticket does not solve it. Doing so needs a signal that does not exist today: something from the backend adapter saying whether a turn produced output, rather than a guess at terminal text. If that ever exists, this is the first place to use it.

Also still unmeasured, as the ticket admitted: how often real work lands inside the floor. My one data point remains a short probe question.

Fixed in `2302b3b`, but **not** the way this ticket suggested. Writing down what was rejected, because the rejected version is the obvious one and someone will try it again. ## What shipped The verdict is unchanged: a turn that settles inside `MIN_TURN_NANOS` is still `FAILED`, and the floor value is still 2000ms. Only the claim about **why** is withdrawn. The old message asserted "most likely a backend error before any work started". When no `errorPattern` matches, that cause is a guess — and both hosts run with none configured, which they log at boot as `backend-error classification: off`. A reader who believes the guess stops looking at the pane, and the pane is where the answer actually was. ## What I rejected, and how I know it was wrong The ticket's first suggested direction — inspect the pane inside the floor, resolve a `COMPLETION` when it holds a real reply — was implemented. Its test for "a real reply" was: non-blank, plus a `.`, `!` or `?` anywhere in the text. That is unsafe twice over: 1. `lastAssistantBlock` falls back to **the whole pane** when it finds no `⏺` marker (`CompletionResolver.java:730`). On a crash there is no assistant block, so the candidate "reply" is the entire screen. 2. A crash pane almost always contains a full stop — in a file path, a version number, or a hostname. I did not argue this from reading. I ran that implementation against the new guard test: ``` aPlausibleLookingReplyInsideTheFloorStillFails:380 expected: <FAILED> but was: <COMPLETION> pane: Error: connection reset while loading src/main/java/Foo.java v1.2.3 ``` A backend failure reported as a successful completion. That is a loud wrong answer traded for a silent one, which is the one trade this path must never make. ## Why the obvious repair does not work either Requiring the `⏺` marker as *positive* evidence of a reply would be safe. But `⏺` is Claude Code chrome, and an opencode pane never carries it — and an opencode member on `mimo-v2.5-free` is what raised this ticket. So the safe rule would fix nothing for the backend that actually needs it. **There is no reliable cross-backend marker for "this is a real reply."** That is the finding, and it is now in the `failTooFast` javadoc so the next reader meets it before rewriting this. ## Tests - `aPlausibleLookingReplyInsideTheFloorStillFails` — pins the safety property against exactly the rejected approach. It passes today and fails against that implementation. Verified by running it, not assumed. - `theTooFastFailureDoesNotAssertACauseItCannotKnow` — pins the wording. - `aNonMatchInsideTheFloorStaysGenericAndNeverNotifiesTheSink` was changed. It asserted the phrase "too fast to be real work", which carried the withdrawn claim. It now asserts what it was really guarding: the floor alone fails the turn, the reason stays generic, the pane is carried, and the typed sink is never notified. `mvn clean install`: 1441 tests, 0 failures, 0 errors. ## What is still not fixed A lead reading `status: failed` on a genuinely fast real turn still sees a failure. That is deliberate — it is the loud direction — but it is not free, and this ticket does not solve it. Doing so needs a signal that does not exist today: something from the **backend adapter** saying whether a turn produced output, rather than a guess at terminal text. If that ever exists, this is the first place to use it. Also still unmeasured, as the ticket admitted: how often real work lands inside the floor. My one data point remains a short probe question.
ltms closed this issue 2026-09-09 02:28:43 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#376