From ac507a98a14441ec0b8c5d7d1e81776e2a018e11 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Fri, 4 Sep 2026 17:13:49 +0700 Subject: [PATCH] Features: the last fleet_ask timeout stranding window is closed (fleetd #334) --- 11-Features.md | 34 ++++++++++++++++++++++++++++++++++ 1 file changed, 34 insertions(+) diff --git a/11-Features.md b/11-Features.md index f03e410..3c59d79 100644 --- a/11-Features.md +++ b/11-Features.md @@ -4415,3 +4415,37 @@ The first version copied the backend-error check's leading-chrome loop. Measured left all 1369 tests green, and it must, because the scan only looks for `.`, `!` and `?` and no chrome character is one of those. It is gone. **A step that cannot change the result is worse than no step** — the next reader takes it as evidence that chrome was handled. + +--- + +## The last window where a `fleet_ask` timeout stranded its ticket is closed + +**What it does.** When a worker's `fleet_ask` times out, `ask()` now closes the ask turn +(`rendezvous.closeAsk`) **before** it forgets that task's `turnId` mapping. The order used to be the +other way round, with the close happening later in the shared `finally`. The whole teardown is also +gated on `ticket.fresh()` now, matching the `finally` block and the `NO_WAITER` branch, which were +already gated that way. + +**On.** Always on (fleetd #334, closing what #329 only narrowed). + +**Why it exists.** Between the forget and the close, the ask was still answerable while its task +mapping was already gone. A lead calling `fleet_send{turnId}` in that window got a `REPLIED` answer, +while the ticket stayed `PENDING` with no reply — forever, because nothing revisits it. Measured +with a probe before the fix: `answer=REPLIED phase=PENDING reply=null`. With the new order, an +`answer()` either sees the ask open — and then the mapping is still there — or sees it closed and +returns `STALE_TURN`. There is no state in between, because both steps run on one thread with +nothing yielding. + +The `ticket.fresh()` gate fixes a second door into the same failure. A coalesced duplicate `ask` +passes its own `timeoutMillis`, which says nothing about whether the shared ask is done, so a +duplicate timing out first could lapse an ask the fresh owner was still holding. + +**One thing to know for maintenance.** The two halves are pinned by two different tests, and the +second one only exists because the first did not cover it. Measured on merge: removing the +`ticket.fresh()` gate while keeping the new order left all 1371 tests green. The gate shipped with +the reorder and nothing held it there. `aCoalescedDuplicateAskTimingOutLeavesTheFreshOwnersAskOpen` +now does, and `aLateAnswerDuringAskTimeoutTeardownStillCompletesTheAsyncTicket` covers the ordering. + +`MessageService` now carries six test-only hooks, one per race of this kind. Each exists because its +window is unreachable through the public API — which is also why each bug was invisible — but six is +enough that the next one needs a harder look than "the file already does this".