Features: #280 — the delayed re-check for a lapsed ask
+41
@@ -3516,3 +3516,44 @@ new `turnId` before the answering thread wakes. The guard is kept as defence in
|
||||
sibling guard whose own comment warns that this ordering is not something to rely on — and the
|
||||
measurement is written into the code next to it, so nobody deletes it as dead code without
|
||||
re-checking the ordering, and nobody trusts it as the only protection either.
|
||||
|
||||
## A ticket is not stranded when a dead member's question lapses
|
||||
|
||||
**What.** A member whose pane dies while it is parked in `fleet_ask` no longer leaves its async
|
||||
ticket at `PENDING` forever. About two minutes after health first classifies the member as `GONE` or
|
||||
`NEVER_READY`, one delayed re-check sweeps the ticket to `FAILED`, so `fleet_poll` gives the lead an
|
||||
answer instead of silence.
|
||||
|
||||
**On.** Always on when fleet health is enabled; no configuration.
|
||||
|
||||
**Why it exists.** An earlier fix closed the *definite teardown* road: an explicit `fleet_stop` or
|
||||
the idle reaper now fails an ASKING ticket immediately. Health deliberately did not, and that is
|
||||
right — a `GONE` reading is a guess from the live agent list, not a teardown the daemon performed,
|
||||
and a guess must never kill a ticket whose worker a live lead could still answer.
|
||||
|
||||
But that left a second road to the same dead end. Health fires its sweep once, on the transition
|
||||
into `GONE`, and skips the ticket because the member really is still asking. Between 55 and 115
|
||||
seconds later the member's own `fleet_ask` lapses and clears its question — and now nothing fires
|
||||
again. The tick loop returns early when the state has not changed, and the classifier answers `GONE`
|
||||
before it could ever reach the orphan state. The ticket sat at `PENDING` for good.
|
||||
|
||||
Nothing else rescued it. A member parked in `fleet_ask` is `BUSY`, and the idle reaper only ever
|
||||
touches `READY` or `DONE` sessions, so a `GONE`-but-never-stopped member is never released and the
|
||||
teardown fix never runs for it.
|
||||
|
||||
**The gotcha: this is one extra attempt, not a retry loop.** An earlier attempt at this called the
|
||||
sweep once per tick for as long as a member stayed terminal, and that was rejected. So the
|
||||
transition schedules exactly **one** delayed follow-up. The delay is 120 seconds, chosen to clear
|
||||
the 115-second worst case of the ask window; because the ask always starts before the `GONE` reading,
|
||||
that ordering guarantees the question has lapsed by the time the re-check runs.
|
||||
|
||||
Two independent things keep the re-check from doing harm. It still passes `sweepAsking: false`, so a
|
||||
member that is genuinely asking again is skipped exactly as on the first attempt. And it fires only
|
||||
if the member is *still* classified in that same terminal state — a member that recovered, or was
|
||||
released and dropped from the roster, is left alone rather than having some brand-new unrelated turn
|
||||
failed underneath it.
|
||||
|
||||
**One thing to know for maintenance.** The 120-second delay is a constant in the code, not a config
|
||||
knob. If either ask ceiling (`FleetMcp`'s or the REST face's, both 115 seconds today) is ever raised
|
||||
past it, this delay must be raised with it, or the re-check fires while the question is still open,
|
||||
finds nothing to sweep, and the single attempt is spent.
|
||||
|
||||
Reference in New Issue
Block a user