Features: #280 — the delayed re-check for a lapsed ask

Dai Ha
2026-09-04 11:33:05 +07:00
parent 443c27d605
commit f7e8434856
+41
@@ -3516,3 +3516,44 @@ new `turnId` before the answering thread wakes. The guard is kept as defence in
sibling guard whose own comment warns that this ordering is not something to rely on — and the
measurement is written into the code next to it, so nobody deletes it as dead code without
re-checking the ordering, and nobody trusts it as the only protection either.
## A ticket is not stranded when a dead member's question lapses
**What.** A member whose pane dies while it is parked in `fleet_ask` no longer leaves its async
ticket at `PENDING` forever. About two minutes after health first classifies the member as `GONE` or
`NEVER_READY`, one delayed re-check sweeps the ticket to `FAILED`, so `fleet_poll` gives the lead an
answer instead of silence.
**On.** Always on when fleet health is enabled; no configuration.
**Why it exists.** An earlier fix closed the *definite teardown* road: an explicit `fleet_stop` or
the idle reaper now fails an ASKING ticket immediately. Health deliberately did not, and that is
right — a `GONE` reading is a guess from the live agent list, not a teardown the daemon performed,
and a guess must never kill a ticket whose worker a live lead could still answer.
But that left a second road to the same dead end. Health fires its sweep once, on the transition
into `GONE`, and skips the ticket because the member really is still asking. Between 55 and 115
seconds later the member's own `fleet_ask` lapses and clears its question — and now nothing fires
again. The tick loop returns early when the state has not changed, and the classifier answers `GONE`
before it could ever reach the orphan state. The ticket sat at `PENDING` for good.
Nothing else rescued it. A member parked in `fleet_ask` is `BUSY`, and the idle reaper only ever
touches `READY` or `DONE` sessions, so a `GONE`-but-never-stopped member is never released and the
teardown fix never runs for it.
**The gotcha: this is one extra attempt, not a retry loop.** An earlier attempt at this called the
sweep once per tick for as long as a member stayed terminal, and that was rejected. So the
transition schedules exactly **one** delayed follow-up. The delay is 120 seconds, chosen to clear
the 115-second worst case of the ask window; because the ask always starts before the `GONE` reading,
that ordering guarantees the question has lapsed by the time the re-check runs.
Two independent things keep the re-check from doing harm. It still passes `sweepAsking: false`, so a
member that is genuinely asking again is skipped exactly as on the first attempt. And it fires only
if the member is *still* classified in that same terminal state — a member that recovered, or was
released and dropped from the roster, is left alone rather than having some brand-new unrelated turn
failed underneath it.
**One thing to know for maintenance.** The 120-second delay is a constant in the code, not a config
knob. If either ask ceiling (`FleetMcp`'s or the REST face's, both 115 seconds today) is ever raised
past it, this delay must be raised with it, or the re-check fires while the question is still open,
finds nothing to sweep, and the single attempt is spent.