A backend-exhausted pane with no ⏺ marker is classified as an empty scrape, so the credential is never quarantined #211

Closed
opened 2026-08-31 17:08:58 +02:00 by ltms · 1 comment
Owner

Found while closing PR #200. The worker there flagged something adjacent in its report; I checked it in the code and the real mechanism is different from their description, and the consequence is more serious than "the error is not classified".

The path

CompletionResolver.resolveCompletion(...) classifies in this order:

  1. assistantBlock = lastAssistantBlock(agents.read(...)) (~line 244)
  2. if (scrapeFailed || tail.isEmpty()) { fail(..., emptyScrapeReason(...)); return; } (~line 259)
  3. BACKEND_EXHAUSTED match → fail(...) and notify exhaustionSink (~line 276)
  4. BACKEND_ERROR match → fail(...) (~line 293)

Step 2 returns before 3 and 4 ever run.

Now lastAssistantBlock (~line 441):

int marker = raw.lastIndexOf('⏺');
String block = marker >= 0 ? raw.substring(marker + 1) : raw;
for (String line : block.split("\n", -1)) {
    if (isBoundary(line)) break;   // first TUI boundary ends the assistant message
    ...
}

With no ⏺ on the pane, block is the whole raw screen — not blank, so the worker's "returns blank" framing is not quite right. But the boundary scan then starts at the top of that screen, and isBoundary treats ╭, │, ╰, ┌, └, ❯, ⏵, ⎿, ⚠ and horizontal rules as boundaries. A pane whose visible top line is chrome — a box border, a warning line, the input box — breaks on the first line and returns "".

So the trigger is the conjunction: no ⏺ marker anywhere on the pane, and the first visible line is TUI chrome. Then tail.isEmpty(), step 2 fires, and the pane is reported as "empty scrape" no matter what the rest of the screen said.

Why it matters more than a wrong label

Step 3 is not only a classification — it is the only thing that calls exhaustionSink. That sink is what quarantines an exhausted credential so placement stops spawning onto it.

So a member whose backend refused the turn for usage exhaustion, on a pane shaped as above, is recorded as "the member produced nothing" and the credential is never quarantined. fleetd keeps handing work to a backend that cannot run it, and each attempt looks like a member that silently did nothing. That is the failure mode where a usage limit takes out several members and the roster gives no reason.

BACKEND_ERROR (step 4) loses its label the same way, but has no side effect beyond the message.

What I did and did not check

  • I read the ordering and lastAssistantBlock/isBoundary in main and confirmed the control flow above.
  • I did not reproduce it against a live exhausted backend, and I do not know how often a real Claude Code or opencode pane ends with no ⏺ and leading chrome. The frequency is unmeasured. The code path is real; the rate is not established.

Suggested direction

Do not simply reorder 3 and 4 above 2 — the empty-scrape fail exists for a good reason (#164), and pattern-matching a blank string finds nothing anyway. The useful change is to run the classification against the raw pane, not only the boundary-trimmed assistant block, before falling back to the generic empty-scrape failure:

  • Keep lastAssistantBlock as the source for the reply text.
  • Match BACKEND_EXHAUSTED / BACKEND_ERROR against the raw scrape as well, so a marker-less pane still classifies and still reaches exhaustionSink.
  • Keep the empty-scrape failure as the last resort it is meant to be.

Note the tension to resolve explicitly: matching the raw screen widens what the patterns can hit, including chrome and scrollback that is not this turn's output. BACKEND_EXHAUSTED has a real side effect, so a false positive quarantines a working credential. Whoever takes this should say which way they resolved that and test both directions — a missed exhaustion and a spurious quarantine are both live risks, and they pull opposite ways.

Acceptance

  • A test with a pane containing an exhaustion line, no ⏺, and a leading ╭ chrome line: it must classify as BACKEND_EXHAUSTED and notify the sink, not as an empty scrape.
  • The same for a backend-error line.
  • A test that an ordinary pane with a normal ⏺ block is unchanged.
  • A test that pins whichever false-positive decision is taken.
  • Watch each fail before keeping it.

Related: #164 (the empty-scrape fail this collides with), #201 (a real classification mechanism instead of hard-coded strings — this issue is more evidence for it).

Found while closing PR #200. The worker there flagged something adjacent in its report; I checked it in the code and the real mechanism is different from their description, and the consequence is more serious than "the error is not classified". ## The path `CompletionResolver.resolveCompletion(...)` classifies in this order: 1. `assistantBlock = lastAssistantBlock(agents.read(...))` (~line 244) 2. **`if (scrapeFailed || tail.isEmpty()) { fail(..., emptyScrapeReason(...)); return; }`** (~line 259) 3. `BACKEND_EXHAUSTED` match → `fail(...)` **and notify `exhaustionSink`** (~line 276) 4. `BACKEND_ERROR` match → `fail(...)` (~line 293) Step 2 returns before 3 and 4 ever run. Now `lastAssistantBlock` (~line 441): ```java int marker = raw.lastIndexOf('⏺'); String block = marker >= 0 ? raw.substring(marker + 1) : raw; for (String line : block.split("\n", -1)) { if (isBoundary(line)) break; // first TUI boundary ends the assistant message ... } ``` With no `⏺` on the pane, `block` is the **whole raw screen** — not blank, so the worker's "returns blank" framing is not quite right. But the boundary scan then starts at the top of that screen, and `isBoundary` treats `╭`, `│`, `╰`, `┌`, `└`, `❯`, `⏵`, `⎿`, `⚠` and horizontal rules as boundaries. A pane whose visible top line is chrome — a box border, a warning line, the input box — breaks on the **first** line and returns `""`. So the trigger is the conjunction: **no `⏺` marker anywhere on the pane, and the first visible line is TUI chrome.** Then `tail.isEmpty()`, step 2 fires, and the pane is reported as "empty scrape" no matter what the rest of the screen said. ## Why it matters more than a wrong label Step 3 is not only a classification — it is the **only** thing that calls `exhaustionSink`. That sink is what quarantines an exhausted credential so placement stops spawning onto it. So a member whose backend refused the turn for usage exhaustion, on a pane shaped as above, is recorded as "the member produced nothing" and **the credential is never quarantined**. fleetd keeps handing work to a backend that cannot run it, and each attempt looks like a member that silently did nothing. That is the failure mode where a usage limit takes out several members and the roster gives no reason. `BACKEND_ERROR` (step 4) loses its label the same way, but has no side effect beyond the message. ## What I did and did not check - I read the ordering and `lastAssistantBlock`/`isBoundary` in `main` and confirmed the control flow above. - I did **not** reproduce it against a live exhausted backend, and I do not know how often a real Claude Code or opencode pane ends with no `⏺` and leading chrome. The frequency is unmeasured. The code path is real; the rate is not established. ## Suggested direction Do not simply reorder 3 and 4 above 2 — the empty-scrape fail exists for a good reason (#164), and pattern-matching a blank string finds nothing anyway. The useful change is to run the classification against the **raw** pane, not only the boundary-trimmed assistant block, before falling back to the generic empty-scrape failure: - Keep `lastAssistantBlock` as the source for the reply *text*. - Match `BACKEND_EXHAUSTED` / `BACKEND_ERROR` against the raw scrape as well, so a marker-less pane still classifies and still reaches `exhaustionSink`. - Keep the empty-scrape failure as the last resort it is meant to be. Note the tension to resolve explicitly: matching the raw screen widens what the patterns can hit, including chrome and scrollback that is not this turn's output. `BACKEND_EXHAUSTED` has a real side effect, so a false positive quarantines a working credential. Whoever takes this should say which way they resolved that and test both directions — a missed exhaustion and a spurious quarantine are both live risks, and they pull opposite ways. ## Acceptance - A test with a pane containing an exhaustion line, no `⏺`, and a leading `╭` chrome line: it must classify as `BACKEND_EXHAUSTED` **and** notify the sink, not as an empty scrape. - The same for a backend-error line. - A test that an ordinary pane with a normal `⏺` block is unchanged. - A test that pins whichever false-positive decision is taken. - Watch each fail before keeping it. Related: #164 (the empty-scrape fail this collides with), #201 (a real classification mechanism instead of hard-coded strings — this issue is more evidence for it).
Author
Owner

Merged to main as 049e7d9, with a follow-up at cabcd87.

The follow-up closed a real gap in the merged fix: the raw-scrape fallback's backend-error branch put only the matched line into the failure, not the pane. That is the wrong way round — the fallback runs precisely because the trimmed block was empty, so the raw scrape is the only copy of whatever the member managed to say. It now carries the pane, clipped to the same cap the normal path uses. The two extra assertions were watched failing first.

Full suite green (1081 tests as of c3fa113).

Merged to main as `049e7d9`, with a follow-up at `cabcd87`. The follow-up closed a real gap in the merged fix: the raw-scrape fallback's backend-error branch put only the matched line into the failure, not the pane. That is the wrong way round — the fallback runs precisely because the trimmed block was empty, so the raw scrape is the **only** copy of whatever the member managed to say. It now carries the pane, clipped to the same cap the normal path uses. The two extra assertions were watched failing first. Full suite green (1081 tests as of `c3fa113`).
ltms closed this issue 2026-09-01 09:13:03 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#211