fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback #216

Closed
agent wants to merge 0 commits from worker/cb211-exhaustion-classification-9546e0-2 into main
Member

Fixes fleetd#211.

CompletionResolver.resolve() classified in this order: (1) extract the last assistant block, (2) if the tail is empty or the scrape failed, fail and return, (3) BACKEND_EXHAUSTED match (notifies exhaustionSink), (4) BACKEND_ERROR match. Step 2 could return before 3 and 4 ever ran.

lastAssistantBlock() finds no usable text whenever there is no ⏺ marker AND the pane's first visible line is TUI chrome — the boundary scan starts at the top of the raw screen and breaks on the first chrome line, returning "". Since BACKEND_EXHAUSTED is the only caller of exhaustionSink, this meant an exhausted backend credential was recorded as "produced nothing" instead of being quarantined, so fleetd kept handing it work.

Fix: run the same BACKEND_EXHAUSTED / BACKEND_ERROR classification against the raw (untrimmed) scrape, but only as a fallback inside the empty-scrape failure branch — not on the normal path, and steps 3/4 are not moved above step 2. A pane that already yields a usable assistant block never reaches this branch, so the existing narrow match is byte-for-byte unchanged. lastAssistantBlock stays the sole source of the reply text; only classification ever consults the raw scrape.

Tests added (CompletionResolverTest): watched each fail against the pre-fix code before keeping it.

  • anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink — asserts the sink is actually called, not just the label. Failed pre-fix with expected: <BACKEND_EXHAUSTED> but was: <FAILED>.
  • aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape — failed pre-fix with the failure text being the generic empty-scrape reason instead of carrying the matched backend-error line.
  • anOrdinaryPaneWithANormalAssistantBlockIsUnaffectedByTheRawScrapeFallback — pins that a usable block bypasses the fallback entirely (same outcome, same text, sink not called), even when text before the ⏺ marker would itself match the exhausted pattern.
  • aGenuinelyEmptyScrapeStillFailsAsEmptyAndNeverNotifiesTheSink — false-positive pin: a pane with no exhaustion/backend-error text anywhere still reports the ordinary empty-scrape failure and never calls the sink.

Build: mvn clean install from fleetd/ — full output read end to end (no pipe). Tests run: 1075, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.

Fixes fleetd#211. CompletionResolver.resolve() classified in this order: (1) extract the last assistant block, (2) if the tail is empty or the scrape failed, fail and return, (3) BACKEND_EXHAUSTED match (notifies exhaustionSink), (4) BACKEND_ERROR match. Step 2 could return before 3 and 4 ever ran. lastAssistantBlock() finds no usable text whenever there is no `⏺` marker AND the pane's first visible line is TUI chrome — the boundary scan starts at the top of the raw screen and breaks on the first chrome line, returning `""`. Since BACKEND_EXHAUSTED is the *only* caller of exhaustionSink, this meant an exhausted backend credential was recorded as "produced nothing" instead of being quarantined, so fleetd kept handing it work. **Fix**: run the same BACKEND_EXHAUSTED / BACKEND_ERROR classification against the raw (untrimmed) scrape, but only as a fallback inside the empty-scrape failure branch — not on the normal path, and steps 3/4 are not moved above step 2. A pane that already yields a usable assistant block never reaches this branch, so the existing narrow match is byte-for-byte unchanged. `lastAssistantBlock` stays the sole source of the reply *text*; only classification ever consults the raw scrape. **Tests added** (CompletionResolverTest): watched each fail against the pre-fix code before keeping it. - `anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink` — asserts the sink is actually called, not just the label. Failed pre-fix with `expected: <BACKEND_EXHAUSTED> but was: <FAILED>`. - `aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape` — failed pre-fix with the failure text being the generic empty-scrape reason instead of carrying the matched backend-error line. - `anOrdinaryPaneWithANormalAssistantBlockIsUnaffectedByTheRawScrapeFallback` — pins that a usable block bypasses the fallback entirely (same outcome, same text, sink not called), even when text before the `⏺` marker would itself match the exhausted pattern. - `aGenuinelyEmptyScrapeStillFailsAsEmptyAndNeverNotifiesTheSink` — false-positive pin: a pane with no exhaustion/backend-error text anywhere still reports the ordinary empty-scrape failure and never calls the sink. **Build**: `mvn clean install` from `fleetd/` — full output read end to end (no pipe). `Tests run: 1075, Failures: 0, Errors: 0, Skipped: 0`. `BUILD SUCCESS`.
agent added 1 commit 2026-09-01 05:50:47 +02:00
fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m17s
445a45f6e1
CompletionResolver.resolve() returned an empty-scrape failure before the
BACKEND_EXHAUSTED / BACKEND_ERROR classification ever ran, whenever
lastAssistantBlock() found no usable text — most commonly a pane with no ⏺
marker at all, whose boundary scan starts at the top of the raw screen and
breaks immediately on the first line of TUI chrome. Since BACKEND_EXHAUSTED
is the only caller of exhaustionSink, this meant an exhausted backend was
recorded as "produced nothing" instead of being quarantined.

Fix: run the same two classifications against the raw (untrimmed) scrape as
a fallback, only inside the empty-scrape failure branch. A pane that already
yields a usable assistant block never reaches this branch, so the existing
narrow match is unchanged. lastAssistantBlock stays the sole source of the
reply text; only classification ever consults the raw scrape.
Owner

Merged to main as 049e7d9. Full suite on the merged result: Tests run: 1079, Failures: 0, Errors: 0, Skipped: 0 — mvn clean install, unpiped, run by me.

The fix is right and the placement is exactly what the brief asked for. I checked the fallback against the normal path line by line: same reason format, same resolveExhausted guard, same inFlight.remove, and the sink called only on the resolution that won the race. The third test is the strongest one here — a pane whose text above the ⏺ marker would itself match the exhausted pattern, pinning that a usable block never reaches the fallback at all.

One follow-up, fixed at merge (cabcd87). The fallback's backend-error branch carried only the matched line. The normal path deliberately appends the pane tail, and its own comment says why: the BACKEND_ERROR pattern is a heuristic, so a member that reported about an error while forgetting fleet_reply matches it too, and dropping the rest of the pane destroys the report — which is the defect fleetd#164 exists to fix.

That reasoning is stronger in the fallback, not weaker. The fallback only runs when the trimmed block was empty, so the raw scrape is the only copy of whatever the member said. A lead would have read one matched line and nothing else.

fail(target, turn, "member " + target + " ended on a backend error: " + backendError
        + "\n--- pane tail ---\n" + clip(raw));

Clipped to the same cap the normal path uses, since a raw screen has no boundary trimming to bound its size. I extended your aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape and watched the new assertion fail first:

AssertionFailedError: fleetd#164: the failure must carry the pane, not only the
matched line ... expected: <true> but was: <false>

Nothing else needed changing.

Merged to `main` as `049e7d9`. Full suite on the merged result: **Tests run: 1079, Failures: 0, Errors: 0, Skipped: 0** — `mvn clean install`, unpiped, run by me. The fix is right and the placement is exactly what the brief asked for. I checked the fallback against the normal path line by line: same reason format, same `resolveExhausted` guard, same `inFlight.remove`, and the sink called only on the resolution that won the race. The third test is the strongest one here — a pane whose text *above* the `⏺` marker would itself match the exhausted pattern, pinning that a usable block never reaches the fallback at all. **One follow-up, fixed at merge (`cabcd87`).** The fallback's backend-error branch carried only the matched line. The normal path deliberately appends the pane tail, and its own comment says why: the `BACKEND_ERROR` pattern is a heuristic, so a member that reported *about* an error while forgetting `fleet_reply` matches it too, and dropping the rest of the pane destroys the report — which is the defect fleetd#164 exists to fix. That reasoning is stronger in the fallback, not weaker. The fallback only runs when the trimmed block was empty, so the raw scrape is the **only** copy of whatever the member said. A lead would have read one matched line and nothing else. ```java fail(target, turn, "member " + target + " ended on a backend error: " + backendError + "\n--- pane tail ---\n" + clip(raw)); ``` Clipped to the same cap the normal path uses, since a raw screen has no boundary trimming to bound its size. I extended your `aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape` and watched the new assertion fail first: ``` AssertionFailedError: fleetd#164: the failure must carry the pane, not only the matched line ... expected: <true> but was: <false> ``` Nothing else needed changing.
ltms closed this pull request 2026-09-01 09:13:26 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m17s

Pull request closed

Sign in to join this conversation.