diff --git a/11-Features.md b/11-Features.md index 4cd87a2..2908c26 100644 --- a/11-Features.md +++ b/11-Features.md @@ -3516,3 +3516,44 @@ new `turnId` before the answering thread wakes. The guard is kept as defence in sibling guard whose own comment warns that this ordering is not something to rely on — and the measurement is written into the code next to it, so nobody deletes it as dead code without re-checking the ordering, and nobody trusts it as the only protection either. + +## A ticket is not stranded when a dead member's question lapses + +**What.** A member whose pane dies while it is parked in `fleet_ask` no longer leaves its async +ticket at `PENDING` forever. About two minutes after health first classifies the member as `GONE` or +`NEVER_READY`, one delayed re-check sweeps the ticket to `FAILED`, so `fleet_poll` gives the lead an +answer instead of silence. + +**On.** Always on when fleet health is enabled; no configuration. + +**Why it exists.** An earlier fix closed the *definite teardown* road: an explicit `fleet_stop` or +the idle reaper now fails an ASKING ticket immediately. Health deliberately did not, and that is +right — a `GONE` reading is a guess from the live agent list, not a teardown the daemon performed, +and a guess must never kill a ticket whose worker a live lead could still answer. + +But that left a second road to the same dead end. Health fires its sweep once, on the transition +into `GONE`, and skips the ticket because the member really is still asking. Between 55 and 115 +seconds later the member's own `fleet_ask` lapses and clears its question — and now nothing fires +again. The tick loop returns early when the state has not changed, and the classifier answers `GONE` +before it could ever reach the orphan state. The ticket sat at `PENDING` for good. + +Nothing else rescued it. A member parked in `fleet_ask` is `BUSY`, and the idle reaper only ever +touches `READY` or `DONE` sessions, so a `GONE`-but-never-stopped member is never released and the +teardown fix never runs for it. + +**The gotcha: this is one extra attempt, not a retry loop.** An earlier attempt at this called the +sweep once per tick for as long as a member stayed terminal, and that was rejected. So the +transition schedules exactly **one** delayed follow-up. The delay is 120 seconds, chosen to clear +the 115-second worst case of the ask window; because the ask always starts before the `GONE` reading, +that ordering guarantees the question has lapsed by the time the re-check runs. + +Two independent things keep the re-check from doing harm. It still passes `sweepAsking: false`, so a +member that is genuinely asking again is skipped exactly as on the first attempt. And it fires only +if the member is *still* classified in that same terminal state — a member that recovered, or was +released and dropped from the roster, is left alone rather than having some brand-new unrelated turn +failed underneath it. + +**One thing to know for maintenance.** The 120-second delay is a constant in the code, not a config +knob. If either ask ceiling (`FleetMcp`'s or the REST face's, both 115 seconds today) is ever raised +past it, this delay must be raised with it, or the re-check fires while the question is still open, +finds nothing to sweep, and the single attempt is spent.