fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS #617

Merged
ltms merged 1 commits from worker/fleetd-615-e05481-5 into main 2026-09-22 05:21:21 +02:00
Member

Fixes fleetd #615.

LeadRollover.runRollover made two unwrapped agents.send calls (/clear and the bootstrap text). HerdrException is unchecked, and the production continuationRunner is a bare virtual thread with no uncaught-exception handler, so a throw from either call killed the continuation silently. confirm() had already written IN_PROGRESS into outcomes before scheduling the continuation, and nothing ever overwrote that entry with a terminal state — so status(token) reported IN_PROGRESS forever.

Fix: wrap the whole continuation body in one try { } catch (RuntimeException e) { }, matching the local convention already used around agents.status in waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED entry naming the exception, in the same diagnostic style as TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED.

Tests: two new tests make send() throw on the /clear call and on the bootstrap-text call respectively (after /clear already succeeded and settled), each asserting status(token) reports FAILED, not IN_PROGRESS. Existing tests already cover the negative control (ROLLED, TURN_NEVER_SETTLED, CLEAR_NEVER_SETTLED) and still pass.

Red/green proof: reverting only the production catch (keeping the tests) makes both new tests fail — the exception now propagates out of confirm() itself, since the test harness runs the continuation synchronously. Restoring the catch turns them green again.

mvn clean install in fleetd/: Tests run: 1866, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

Sweep (report only, not fixed here): LeadRollover has exactly one continuationRunner.accept call site (in confirm()) and exactly one place that writes a non-terminal IN_PROGRESS entry before handing off to it — both already covered by this fix. The two agents.status calls (waitUntilAtTurnBoundary, waitForClearPickupAndSettle) and the one agents.submit call (waitForClearPickupAndSettle) already have their own local try/catch(RuntimeException) around them, predating this change. No other unwrapped agents.* call or non-terminal-state-before-handoff pattern exists in this class.

Fixes fleetd #615. `LeadRollover.runRollover` made two unwrapped `agents.send` calls (`/clear` and the bootstrap text). `HerdrException` is unchecked, and the production `continuationRunner` is a bare virtual thread with no uncaught-exception handler, so a throw from either call killed the continuation silently. `confirm()` had already written `IN_PROGRESS` into `outcomes` before scheduling the continuation, and nothing ever overwrote that entry with a terminal state — so `status(token)` reported `IN_PROGRESS` forever. **Fix:** wrap the whole continuation body in one `try { } catch (RuntimeException e) { }`, matching the local convention already used around `agents.status` in `waitUntilAtTurnBoundary`. On a throw, write a new terminal `RollState.FAILED` entry naming the exception, in the same diagnostic style as `TURN_NEVER_SETTLED` and `CLEAR_NEVER_SETTLED`. **Tests:** two new tests make `send()` throw on the `/clear` call and on the bootstrap-text call respectively (after `/clear` already succeeded and settled), each asserting `status(token)` reports `FAILED`, not `IN_PROGRESS`. Existing tests already cover the negative control (`ROLLED`, `TURN_NEVER_SETTLED`, `CLEAR_NEVER_SETTLED`) and still pass. **Red/green proof:** reverting only the production catch (keeping the tests) makes both new tests fail — the exception now propagates out of `confirm()` itself, since the test harness runs the continuation synchronously. Restoring the catch turns them green again. `mvn clean install` in `fleetd/`: Tests run: 1866, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. **Sweep (report only, not fixed here):** `LeadRollover` has exactly one `continuationRunner.accept` call site (in `confirm()`) and exactly one place that writes a non-terminal `IN_PROGRESS` entry before handing off to it — both already covered by this fix. The two `agents.status` calls (`waitUntilAtTurnBoundary`, `waitForClearPickupAndSettle`) and the one `agents.submit` call (`waitForClearPickupAndSettle`) already have their own local `try/catch(RuntimeException)` around them, predating this change. No other unwrapped `agents.*` call or non-terminal-state-before-handoff pattern exists in this class.
agent added 1 commit 2026-09-22 05:17:31 +02:00
fleetd #615: write FAILED instead of leaving status(token) stuck at IN_PROGRESS
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m25s
388ef5a3c3
LeadRollover.runRollover made two unwrapped agents.send calls. HerdrException
is unchecked, and the production continuationRunner is a bare virtual thread
with no uncaught-exception handler, so a throw from either send call killed
the continuation silently — confirm() had already written IN_PROGRESS into
outcomes before scheduling it, and nothing ever overwrote that entry with a
terminal state.

Wrap the whole continuation body in one try/catch(RuntimeException), matching
the local convention already used around agents.status in
waitUntilAtTurnBoundary. On a throw, write a new terminal RollState.FAILED
entry naming the exception, in the same diagnostic style as
TURN_NEVER_SETTLED and CLEAR_NEVER_SETTLED.

Two new tests make send() throw on the /clear call and on the bootstrap-text
call respectively, each asserting status(token) reports FAILED, not
IN_PROGRESS. Reverting only the production catch (keeping the tests) turns
both red; restoring it turns them green again.
ltms merged commit 203f034528 into main 2026-09-22 05:21:21 +02:00
Sign in to join this conversation.