#275: abandon() sweeps an ASKING ticket only on a definite teardown #279

Closed
agent wants to merge 0 commits from worker/investigate-275-abandon-asking-fdef52-8 into main
Member

Investigation (fleetd #275)

Claim tested: MessageService.abandon(target, reason) skips a task whose async ticket is
parked in fleet_ask (Phase.ASKING, task.question != null) when the target session is torn
down, and nothing ever retries the sweep, so the ticket sits PENDING/ASKING forever.

Reachability: confirmed, via the public API, no Rendezvous.close reaching-past-the-class.

Sequence (file:line references against MessageService.java before this fix):

  1. fleet_send{wait:false} → sendAsync(target, content) (MessageService.java:922) creates a
    Task and opens the forward rendezvous waiter in send() (:739).
  2. The worker calls fleet_ask → ask(target, question, timeoutMillis) (:797).
    markAsyncQuestion (:1051) sets task.question/task.turnId; resolveQuestion (:806)
    resolves the forward waiter with Kind.QUESTION; send()'s finally (:777-780) then
    closes that waiter. The worker's ask() call is now blocked on the primary's answer
    (default 55s, up to 115s — FleetMcp.ASK_DEFAULT_TIMEOUT_MS / FleetApp.MAX_ASK_TIMEOUT_MS).
  3. The lead calls fleet_stop on this member while it is still asking (a documented, expected
    case — see SessionManager.release's own comment: "a worker that stopped to ask a question or
    refused the turn"). SessionManager.release() unconditionally notifies onRelease
    (Fleetd.java:596, before the fix) → messages.abandon(target, reason).
  4. In abandon() (:582): the forward waiter is already closed (step 2), so the waiter branch
    is a no-op; the matching loop's task.question == null guard (:593, before the fix)
    excludes the still-ASKING task. abandon() returns false — nothing failed.
  5. ~55-115s later, ask()'s own timeout fires, clearAsyncQuestion(turnId, true) (:828) clears
    task.question to null. The task is now question == null, future not done — exactly what
    hasOrphanedDelegation (:364) detects — but the session was already released in step 3, so
    it no longer appears in FleetHealthMonitor's roster
    (FleetHealthMonitor.java:108,
    SessionManager.release removes the registry entry), and DELEGATION_ORPHANED is not in
    terminal() (FleetHealthMonitor.java:194) even when it is computed. Nothing ever calls
    abandon() again for this target.
  6. pruneTerminalTickets (:1037) never removes the task (future.isDone() is false forever),
    and fleet_poll{ticket} reports Phase.PENDING for good.

The same gap is reachable via FleetHealthMonitor.failTerminalTarget (fires once per health-state
transition into GONE/NEVER_READY — FleetHealthMonitor.java:172) landing inside the same ASKING
window, as the ticket's own narrative describes; FleetHealth.decide() checks targetNotFound
(GONE) before orphanedDelegation (FleetHealth.java:16 vs :20), so a target that stays GONE
never even reaches the orphaned check, and no state transition means failTerminalTarget never
refires either way.

A wrinkle found during the fix: blindly dropping question == null in abandon() conflicts
with an existing, deliberately-tested invariant —
MessageServiceTest.abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer — which asserts abandon()
must not fail an ASKING ticket, because the primary may be genuinely mid-answer() for that
same turn, and a FleetHealthMonitor-driven "target not found" reading is a classification from
the live agent list, not a teardown the daemon performed — it could in principle be a stale read
that clears on the next tick. Blanket-widening the guard would make that test fail and would let
an ordinary in-flight fleet_ask be killed by a shaky health signal.

Fix implemented: added abandon(String target, String reason, boolean sweepAsking).

  • sessions.onRelease (an explicit fleet_stop, or the idle reaper) now calls it with
    sweepAsking=true — this caller has certain knowledge the target can never resume (its pane is
    being stopped right now), so it now also fails the ASKING task as WORKER_FAILED, removes the
    asyncTasksByTurn entry, and calls rendezvous.closeAsk(turnId).
  • FleetHealthMonitor's failTerminalTarget keeps calling the existing 2-arg
    abandon(target, reason) (delegates to sweepAsking=false) — unchanged behavior, so
    abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer still passes as-is.
  • Also fixed a related gap while touching this code: the old loop only cleared asyncTasksByTurn
    on a recovered outcome, never on WORKER_FAILED — meaning hasAsyncQuestion(target) could
    keep reporting a dead turn as open. Now cleared unconditionally whenever a matched task's future
    completes and had a turnId.
  • hasOrphanedDelegation is left unchanged (still question == null only) — widening it too
    would flag an ordinary in-flight fleet_ask as orphaned. The sweepAsking fix resolves the real
    gap at the point of teardown, before a task can ever reach the "cleared question, future still
    open" state this method looks for.

Test: abandonWithSweepAskingFailsATornDownTargetsAskingTicket drives the real sequence
(sendAsync → ask → abandon(..., true)) — not a hand-built Task/Rendezvous state — and
asserts the ticket reaches Phase.FAILED, the worker's blocked ask() call itself still lapses to
TIMED_OUT on its own timeout, and a late answer() for the same turnId now correctly reports
STALE_TURN. A second test, abandonWithoutSweepAskingBehavesLikeTheTwoArgOverload, pins the
sweepAsking=false path against the unchanged behavior.

Mutation check performed: reverted the matching filter's widening
((sweepAsking || task.question == null) → task.question == null), reran
abandonWithSweepAskingFailsATornDownTargetsAskingTicket alone — it failed red
(expected: <true> but was: <false>) — then restored the fix and reran the full suite green.

Build

cd fleetd && mvn clean install

BUILD SUCCESS; Tests run: 1276, Failures: 0, Errors: 0, Skipped: 0 (up from 1274 before this
change — the two new tests).

Files changed

  • fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java — abandon overload + javadoc,
    hasOrphanedDelegation javadoc note explaining why it is unchanged.
  • fleetd/src/main/java/dev/ltms/fleet/Fleetd.java — sessions.onRelease now passes
    sweepAsking=true.
  • fleetd/src/test/java/dev/ltms/fleet/msg/MessageServiceTest.java — two new tests.

No caveats beyond the design note above (flagged in case the lead wants the health-monitor path to
someday retry the sweep once an ASKING ticket naturally lapses back to PENDING while its target
stays terminal — that would be a separate, larger change to FleetHealthMonitor, out of this
ticket's scope).

## Investigation (fleetd #275) **Claim tested:** `MessageService.abandon(target, reason)` skips a task whose async ticket is parked in `fleet_ask` (`Phase.ASKING`, `task.question != null`) when the target session is torn down, and nothing ever retries the sweep, so the ticket sits `PENDING`/`ASKING` forever. **Reachability: confirmed, via the public API, no `Rendezvous.close` reaching-past-the-class.** Sequence (file:line references against `MessageService.java` before this fix): 1. `fleet_send{wait:false}` → `sendAsync(target, content)` (`MessageService.java:922`) creates a `Task` and opens the forward rendezvous waiter in `send()` (`:739`). 2. The worker calls `fleet_ask` → `ask(target, question, timeoutMillis)` (`:797`). `markAsyncQuestion` (`:1051`) sets `task.question`/`task.turnId`; `resolveQuestion` (`:806`) resolves the forward waiter with `Kind.QUESTION`; `send()`'s `finally` (`:777-780`) then **closes** that waiter. The worker's `ask()` call is now blocked on the primary's answer (default 55s, up to 115s — `FleetMcp.ASK_DEFAULT_TIMEOUT_MS` / `FleetApp.MAX_ASK_TIMEOUT_MS`). 3. **The lead calls `fleet_stop`** on this member while it is still asking (a documented, expected case — see `SessionManager.release`'s own comment: "a worker that stopped to ask a question or refused the turn"). `SessionManager.release()` unconditionally notifies `onRelease` (`Fleetd.java:596`, before the fix) → `messages.abandon(target, reason)`. 4. In `abandon()` (`:582`): the forward waiter is already closed (step 2), so the `waiter` branch is a no-op; the `matching` loop's `task.question == null` guard (`:593`, before the fix) excludes the still-`ASKING` task. `abandon()` returns `false` — nothing failed. 5. ~55-115s later, `ask()`'s own timeout fires, `clearAsyncQuestion(turnId, true)` (`:828`) clears `task.question` to `null`. The task is now `question == null`, `future` not done — exactly what `hasOrphanedDelegation` (`:364`) detects — **but the session was already released in step 3, so it no longer appears in `FleetHealthMonitor`'s roster** (`FleetHealthMonitor.java:108`, `SessionManager.release` removes the registry entry), and `DELEGATION_ORPHANED` is not in `terminal()` (`FleetHealthMonitor.java:194`) even when it is computed. Nothing ever calls `abandon()` again for this target. 6. `pruneTerminalTickets` (`:1037`) never removes the task (`future.isDone()` is false forever), and `fleet_poll{ticket}` reports `Phase.PENDING` for good. The same gap is reachable via `FleetHealthMonitor.failTerminalTarget` (fires once per health-state *transition* into GONE/NEVER_READY — `FleetHealthMonitor.java:172`) landing inside the same ASKING window, as the ticket's own narrative describes; `FleetHealth.decide()` checks `targetNotFound` (GONE) *before* `orphanedDelegation` (`FleetHealth.java:16` vs `:20`), so a target that stays GONE never even reaches the orphaned check, and no state *transition* means `failTerminalTarget` never refires either way. **A wrinkle found during the fix:** blindly dropping `question == null` in `abandon()` conflicts with an existing, deliberately-tested invariant — `MessageServiceTest.abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer` — which asserts `abandon()` must **not** fail an ASKING ticket, because the primary may be genuinely mid-`answer()` for that same turn, and a `FleetHealthMonitor`-driven "target not found" reading is a classification from the live agent list, not a teardown the daemon performed — it could in principle be a stale read that clears on the next tick. Blanket-widening the guard would make that test fail and would let an ordinary in-flight `fleet_ask` be killed by a shaky health signal. **Fix implemented:** added `abandon(String target, String reason, boolean sweepAsking)`. - `sessions.onRelease` (an explicit `fleet_stop`, or the idle reaper) now calls it with `sweepAsking=true` — this caller has *certain* knowledge the target can never resume (its pane is being stopped right now), so it now also fails the ASKING task as `WORKER_FAILED`, removes the `asyncTasksByTurn` entry, and calls `rendezvous.closeAsk(turnId)`. - `FleetHealthMonitor`'s `failTerminalTarget` keeps calling the existing 2-arg `abandon(target, reason)` (delegates to `sweepAsking=false`) — unchanged behavior, so `abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer` still passes as-is. - Also fixed a related gap while touching this code: the old loop only cleared `asyncTasksByTurn` on a *recovered* outcome, never on `WORKER_FAILED` — meaning `hasAsyncQuestion(target)` could keep reporting a dead turn as open. Now cleared unconditionally whenever a matched task's future completes and had a `turnId`. - `hasOrphanedDelegation` is **left unchanged** (still `question == null` only) — widening it too would flag an ordinary in-flight `fleet_ask` as orphaned. The `sweepAsking` fix resolves the real gap at the point of teardown, before a task can ever reach the "cleared question, future still open" state this method looks for. **Test:** `abandonWithSweepAskingFailsATornDownTargetsAskingTicket` drives the real sequence (`sendAsync` → `ask` → `abandon(..., true)`) — not a hand-built `Task`/`Rendezvous` state — and asserts the ticket reaches `Phase.FAILED`, the worker's blocked `ask()` call itself still lapses to `TIMED_OUT` on its own timeout, and a late `answer()` for the same `turnId` now correctly reports `STALE_TURN`. A second test, `abandonWithoutSweepAskingBehavesLikeTheTwoArgOverload`, pins the `sweepAsking=false` path against the unchanged behavior. Mutation check performed: reverted the `matching` filter's widening (`(sweepAsking || task.question == null)` → `task.question == null`), reran `abandonWithSweepAskingFailsATornDownTargetsAskingTicket` alone — it failed red (`expected: <true> but was: <false>`) — then restored the fix and reran the full suite green. ## Build ``` cd fleetd && mvn clean install ``` `BUILD SUCCESS`; `Tests run: 1276, Failures: 0, Errors: 0, Skipped: 0` (up from 1274 before this change — the two new tests). ## Files changed - `fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java` — `abandon` overload + javadoc, `hasOrphanedDelegation` javadoc note explaining why it is unchanged. - `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java` — `sessions.onRelease` now passes `sweepAsking=true`. - `fleetd/src/test/java/dev/ltms/fleet/msg/MessageServiceTest.java` — two new tests. No caveats beyond the design note above (flagged in case the lead wants the health-monitor path to someday retry the sweep once an ASKING ticket naturally lapses back to PENDING while its target stays terminal — that would be a separate, larger change to `FleetHealthMonitor`, out of this ticket's scope).
agent added 1 commit 2026-09-04 05:14:46 +02:00
#275: abandon() sweeps an ASKING ticket only on a definite teardown
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 2m5s
4887d03d88
Confirmed reachable: a target torn down for good (fleet_stop / the idle
reaper) while its async ticket sits in fleet_ask (Phase.ASKING) got
permanently stuck. resolveQuestion already closes the forward waiter, the
question == null guard excluded the task from abandon()'s sweep, and by the
time the worker's own fleet_ask lapses (~55-115s) the released session no
longer appears in FleetHealthMonitor's roster, so nothing ever calls
abandon() again. fleet_poll{ticket} then reports PENDING forever.

Add abandon(target, reason, sweepAsking) — sessions.onRelease (a definite
teardown: the pane is being stopped right now) passes true and now fails the
ASKING ticket and closes its reverse-rendezvous ask. FleetHealthMonitor's
health-classification call keeps the 2-arg overload (sweepAsking=false):
a GONE/NEVER_READY reading is a guess from the live agent list, not a
teardown it performed, and abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer
already covers why an active ask must survive that guess (the primary may
be mid-answer for the same turn). hasOrphanedDelegation is left unchanged
for the same reason — it must not flag a live, active ask as orphaned.

Proven with a test driving the real public sequence (sendAsync -> ask ->
abandon(..., true)), not a hand-built task map; reverted the widening to
confirm it goes red, then restored it.
ltms closed this pull request 2026-09-04 05:27:42 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 2m5s

Pull request closed

Sign in to join this conversation.