fleetd#164: empty or suspiciously fast scrape must fail, not resolve as a success #180

Merged
ltms merged 1 commits from worker/fleetd-164-empty-scrape-ff768e-2 into main 2026-08-28 01:10:27 +02:00
Member

Fixes fleetd#164.

The defect: CompletionResolver.resolve() treated any confirmed BUSY -> DONE transition as proof a turn finished, and any empty scrape (whether the read failed, or genuinely produced nothing) resolved the caller's send as a successful reply carrying "". A member whose backend rejects the turn (e.g. HTTP 400) crashes in about a second, producing the exact same transition a real completion does, and the caller could not tell a lost turn from a real empty answer.

The fix:

  1. Added MIN_TURN_NANOS (2 seconds) — a named floor below which a BUSY -> DONE transition is treated as a crash signature and resolved as a failure, never a reply. The failure names the member, gives both timings, and carries whatever is on the pane (usually the backend's own error) for context.
  2. An empty scrape (read failure, or a clean read producing zero characters) now also resolves as a failure naming the member, instead of a success carrying "".
  3. Threaded an injectable LongSupplier clock through CompletionResolver, matching the existing nowNanos pattern in SessionManager/MessageService, so the floor is testable without a real sleep.

Out of scope (per the ticket): the scrape-quality problem (terminal UI junk / echoed prompt / previous-turn leftovers in a non-empty scrape) is untouched.

Test encoding the bug: CompletionResolverTest.resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent asserted that a failed read resolved as COMPLETION with an empty string — that was the bug. Renamed to aFailedScrapeResolvesAsAFailureEvenWithABaselinePresent and now asserts FAILED.

Tests: mvn clean install — Tests run: 973, Failures: 0, Errors: 0, Skipped: 0. New tests: anEmptyScrapeResolvesAsAFailureNamingTheMember, aBusyToDoneTransitionInsideTheFloorResolvesAsAFailure, aBusyToDoneTransitionJustOutsideTheFloorResolvesNormally.

Fixes fleetd#164. **The defect:** `CompletionResolver.resolve()` treated any confirmed BUSY -> DONE transition as proof a turn finished, and any empty scrape (whether the read failed, or genuinely produced nothing) resolved the caller's send as a *successful* reply carrying `""`. A member whose backend rejects the turn (e.g. HTTP 400) crashes in about a second, producing the exact same transition a real completion does, and the caller could not tell a lost turn from a real empty answer. **The fix:** 1. Added `MIN_TURN_NANOS` (2 seconds) — a named floor below which a BUSY -> DONE transition is treated as a crash signature and resolved as a **failure**, never a reply. The failure names the member, gives both timings, and carries whatever is on the pane (usually the backend's own error) for context. 2. An empty scrape (read failure, or a clean read producing zero characters) now also resolves as a **failure** naming the member, instead of a success carrying `""`. 3. Threaded an injectable `LongSupplier` clock through `CompletionResolver`, matching the existing `nowNanos` pattern in `SessionManager`/`MessageService`, so the floor is testable without a real sleep. **Out of scope (per the ticket):** the scrape-quality problem (terminal UI junk / echoed prompt / previous-turn leftovers in a *non-empty* scrape) is untouched. **Test encoding the bug:** `CompletionResolverTest.resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresent` asserted that a failed read resolved as `COMPLETION` with an empty string — that was the bug. Renamed to `aFailedScrapeResolvesAsAFailureEvenWithABaselinePresent` and now asserts `FAILED`. **Tests:** `mvn clean install` — `Tests run: 973, Failures: 0, Errors: 0, Skipped: 0`. New tests: `anEmptyScrapeResolvesAsAFailureNamingTheMember`, `aBusyToDoneTransitionInsideTheFloorResolvesAsAFailure`, `aBusyToDoneTransitionJustOutsideTheFloorResolvesNormally`.
agent added 1 commit 2026-08-28 01:01:25 +02:00
fleetd#164: an empty or suspiciously fast scrape must fail, never resolve as a success
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Successful in 1m20s
3bfa82839b
CompletionResolver.resolve() used to hand the caller a successful "" reply whenever a
turn's scrape came back empty (whether the read failed, or genuinely produced nothing),
making a lost turn indistinguishable from a real empty answer. It also had no way to
tell a crashed backend's near-instant BUSY -> DONE transition apart from a genuine
completion.

Add MIN_TURN_NANOS (2s), a named floor below which a completed turn is treated as a
crash signature and failed rather than resolved as a reply. Fail on any empty scrape
(read failure or a clean-but-empty read) instead of resolving with "". Both failures
name the member and carry whatever is on the pane for context.

Thread an injectable LongSupplier clock through CompletionResolver (matching the
SessionManager/MessageService nowNanos pattern) so the floor is testable without a
real sleep.
ltms merged commit 430f5b0dae into main 2026-08-28 01:10:27 +02:00
ltms deleted branch worker/fleetd-164-empty-scrape-ff768e-2 2026-08-28 01:10:27 +02:00
Sign in to join this conversation.