#280: sweep a GONE/NEVER_READY target's lapsed fleet_ask once, not just on transition #294

Closed
agent wants to merge 0 commits from worker/fix-280-gone-ask-lapse-bca98e-2 into main
Member

#280 — follow-up to #275

Step 1 answer: the gap is real and reachable. I traced it in the code myself:

  • SessionManager.reapIdle (SessionManager.java:840-864) only ever reaps sessions in READY or DONE state (line 844: if (s.state() != MemberSession.State.READY && s.state() != MemberSession.State.DONE) continue;).
  • A session mid-turn is BUSY (onDelivered, SessionManager.java:731-744, sets State.BUSY and nothing clears it until the turn completes). A worker parked in fleet_ask is still mid-turn, so it stays BUSY the whole time it is asking and after the ask lapses.
  • FleetHealthMonitor's failTarget is wired to messages::abandon directly (Fleetd.java:559-561) — it never touches SessionManager state or calls release/onRelease at all.
  • So a GONE member that is never explicitly stopped stays BUSY forever, reapIdle never reaps it (state excluded), and onRelease's sweepAsking=true path (#275's fix) is never reached. Nothing else releases such a session. The gap described in the ticket is real.

Step 2 fix. I weighed the two shapes and chose the narrower one: a single delayed re-check per terminal transition, not a general per-tick retry (CB-580 rejected that shape and I kept it rejected).

  • reportTransition now schedules one bounded, delayed follow-up (ASK_LAPSE_RECHECK_DELAY_SECONDS = 120, chosen to exceed FleetMcp.ASK_DEFAULT_TIMEOUT_MS/FleetApp.MAX_ASK_TIMEOUT_MS's 55-115s window) alongside the existing immediate failTerminalTarget call, only on an actual state change into GONE/NEVER_READY.
  • recheckTerminalTarget fires failTerminalTarget again only if the target is still classified in that same terminal state at the time it runs — guarding against reaching into a member that recovered, or was released and dropped from the roster, in the meantime (which would otherwise risk failing a brand-new unrelated turn).
  • sweepAsking stays false throughout, so the invariant abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer pins is untouched: a ticket whose ask has not yet lapsed is never failed, on the first sweep or the delayed one.

Tests added (FleetHealthMonitorTest):

  • terminalTransitionSchedulesExactlyOneDelayedRecheck — proves the scheduling wiring and that an unchanged tick doesn't queue a second one.
  • recheckIsANoOpOnceTheTargetHasRecovered / recheckIsANoOpForATargetItNeverObserved — prove the terminal-state guard.
  • delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved — the ticket's exact scenario end-to-end through the real MessageService: GONE fires while genuinely ASKING (skipped), the ask lapses on its own, an unchanged tick still doesn't refire, and the delayed recheck finally sweeps the ticket to FAILED.

Mutation proof. I temporarily neutered recheckTerminalTarget's body (no-op) and reran mvn test -Dtest=FleetHealthMonitorTest:

[ERROR] Tests run: 21, Failures: 1, Errors: 0, Skipped: 0
org.opentest4j.AssertionFailedError: expected: <FAILED> but was: <PENDING>
  at FleetHealthMonitorTest.delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved(FleetHealthMonitorTest.java:536)

Only that one test failed; the other 20 in the class stayed green. Restored the fix and reran the full build: mvn clean install → Tests run: 1300, Failures: 0, Errors: 0, Skipped: 0 / BUILD SUCCESS.

Files changed:

  • fleetd/src/main/java/dev/ltms/fleet/health/FleetHealthMonitor.java
  • fleetd/src/test/java/dev/ltms/fleet/health/FleetHealthMonitorTest.java
## #280 — follow-up to #275 **Step 1 answer: the gap is real and reachable.** I traced it in the code myself: - `SessionManager.reapIdle` (SessionManager.java:840-864) only ever reaps sessions in `READY` or `DONE` state (line 844: `if (s.state() != MemberSession.State.READY && s.state() != MemberSession.State.DONE) continue;`). - A session mid-turn is `BUSY` (`onDelivered`, SessionManager.java:731-744, sets `State.BUSY` and nothing clears it until the turn completes). A worker parked in `fleet_ask` is still mid-turn, so it stays `BUSY` the whole time it is asking and after the ask lapses. - `FleetHealthMonitor`'s `failTarget` is wired to `messages::abandon` directly (Fleetd.java:559-561) — it never touches `SessionManager` state or calls `release`/`onRelease` at all. - So a `GONE` member that is never explicitly stopped stays `BUSY` forever, `reapIdle` never reaps it (state excluded), and `onRelease`'s `sweepAsking=true` path (#275's fix) is never reached. Nothing else releases such a session. The gap described in the ticket is real. **Step 2 fix.** I weighed the two shapes and chose the narrower one: **a single delayed re-check per terminal transition**, not a general per-tick retry (CB-580 rejected that shape and I kept it rejected). - `reportTransition` now schedules one bounded, delayed follow-up (`ASK_LAPSE_RECHECK_DELAY_SECONDS = 120`, chosen to exceed `FleetMcp.ASK_DEFAULT_TIMEOUT_MS`/`FleetApp.MAX_ASK_TIMEOUT_MS`'s 55-115s window) alongside the existing immediate `failTerminalTarget` call, only on an actual state change into GONE/NEVER_READY. - `recheckTerminalTarget` fires `failTerminalTarget` again **only if the target is still classified in that same terminal state** at the time it runs — guarding against reaching into a member that recovered, or was released and dropped from the roster, in the meantime (which would otherwise risk failing a brand-new unrelated turn). - `sweepAsking` stays `false` throughout, so the invariant `abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer` pins is untouched: a ticket whose ask has not yet lapsed is never failed, on the first sweep or the delayed one. **Tests added** (`FleetHealthMonitorTest`): - `terminalTransitionSchedulesExactlyOneDelayedRecheck` — proves the scheduling wiring and that an unchanged tick doesn't queue a second one. - `recheckIsANoOpOnceTheTargetHasRecovered` / `recheckIsANoOpForATargetItNeverObserved` — prove the terminal-state guard. - `delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved` — the ticket's exact scenario end-to-end through the real `MessageService`: GONE fires while genuinely ASKING (skipped), the ask lapses on its own, an unchanged tick still doesn't refire, and the delayed recheck finally sweeps the ticket to FAILED. **Mutation proof.** I temporarily neutered `recheckTerminalTarget`'s body (no-op) and reran `mvn test -Dtest=FleetHealthMonitorTest`: ``` [ERROR] Tests run: 21, Failures: 1, Errors: 0, Skipped: 0 org.opentest4j.AssertionFailedError: expected: <FAILED> but was: <PENDING> at FleetHealthMonitorTest.delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved(FleetHealthMonitorTest.java:536) ``` Only that one test failed; the other 20 in the class stayed green. Restored the fix and reran the full build: `mvn clean install` → `Tests run: 1300, Failures: 0, Errors: 0, Skipped: 0` / `BUILD SUCCESS`. **Files changed:** - `fleetd/src/main/java/dev/ltms/fleet/health/FleetHealthMonitor.java` - `fleetd/src/test/java/dev/ltms/fleet/health/FleetHealthMonitorTest.java`
agent added 1 commit 2026-09-04 06:28:41 +02:00
#280: sweep a GONE/NEVER_READY target's lapsed fleet_ask once, not just on transition
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m58s
fa39a5f55e
FleetHealthMonitor's fire-once-per-transition rule (CB-580) means a target
that is genuinely mid-fleet_ask when health first classifies it GONE is
correctly skipped (sweepAsking=false). But nothing re-fires abandon() once
that ask lapses on its own 55-115s later: FleetHealth.decide keeps reporting
GONE every tick, and reportTransition's previous==next guard never lets the
sweep run again. The ticket then sat PENDING forever, the same destination
#275 fixed for an explicit teardown, reached here by a health guess instead.

SessionManager.reapIdle only reaps READY/DONE sessions (SessionManager.java:844),
and a session mid-turn (including mid-ask) stays BUSY the whole time
(onDelivered sets BUSY, nothing clears it until the turn completes) — so
SessionReaper never releases such a session and onRelease's sweepAsking=true
path is never reached.

Fix: schedule one bounded, delayed re-check per terminal transition (not a
per-tick retry — that shape was rejected by CB-580). It fires failTerminalTarget
again after a delay that exceeds the worst-case ask-lapse window, and only if
the target is still classified in the same terminal state at that time, so a
recovered or since-released target is never reached into. sweepAsking stays
false throughout, so a ticket whose ask has not yet lapsed is still never
touched — same invariant abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer pins.

Proven with a mutation: neutering recheckTerminalTarget's body made
delayedRecheckSweepsATicketWhoseAskLapsedAfterGoneWasFirstObserved fail with
"expected: <FAILED> but was: <PENDING>", all 20 other FleetHealthMonitorTest
cases still green; restored and reran clean (1300 tests, 0 failures).
ltms closed this pull request 2026-09-04 06:33:51 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m58s

Pull request closed

Sign in to join this conversation.