fleetd#164: empty or suspiciously fast scrape must fail, not resolve as a success #180
Reference in New Issue
Block a user
Delete Branch "worker/fleetd-164-empty-scrape-ff768e-2"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Fixes fleetd#164.
The defect:
CompletionResolver.resolve()treated any confirmed BUSY -> DONE transition as proof a turn finished, and any empty scrape (whether the read failed, or genuinely produced nothing) resolved the caller's send as a successful reply carrying"". A member whose backend rejects the turn (e.g. HTTP 400) crashes in about a second, producing the exact same transition a real completion does, and the caller could not tell a lost turn from a real empty answer.The fix:
MIN_TURN_NANOS(2 seconds) — a named floor below which a BUSY -> DONE transition is treated as a crash signature and resolved as a failure, never a reply. The failure names the member, gives both timings, and carries whatever is on the pane (usually the backend's own error) for context."".LongSupplierclock throughCompletionResolver, matching the existingnowNanospattern inSessionManager/MessageService, so the floor is testable without a real sleep.Out of scope (per the ticket): the scrape-quality problem (terminal UI junk / echoed prompt / previous-turn leftovers in a non-empty scrape) is untouched.
Test encoding the bug:
CompletionResolverTest.resolvesWhenTheScrapeItselfFailsEvenWithABaselinePresentasserted that a failed read resolved asCOMPLETIONwith an empty string — that was the bug. Renamed toaFailedScrapeResolvesAsAFailureEvenWithABaselinePresentand now assertsFAILED.Tests:
mvn clean install—Tests run: 973, Failures: 0, Errors: 0, Skipped: 0. New tests:anEmptyScrapeResolvesAsAFailureNamingTheMember,aBusyToDoneTransitionInsideTheFloorResolvesAsAFailure,aBusyToDoneTransitionJustOutsideTheFloorResolvesNormally.