From dc565af7099c952b6f8fb4a473567f9aefed6e8e Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Mon, 31 Aug 2026 23:30:27 +0700 Subject: [PATCH] Features: one worker reply settles exactly one ticket (CB-137) --- 11-Features.md | 37 +++++++++++++++++++++++++++++++++++++ 1 file changed, 37 insertions(+) diff --git a/11-Features.md b/11-Features.md index 33eab6c..ba6abf1 100644 --- a/11-Features.md +++ b/11-Features.md @@ -72,6 +72,7 @@ six weeks, and the table alone will not carry it. | [Fill in a member id the backend names late](#late-resolved-member-ids) | automatic | CB-209 | `session/SessionManager` | | [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` | | [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` | +| [One reply settles one ticket](#one-worker-reply-settles-exactly-one-ticket) | automatic | CB-137 | `msg/MessageService` | Nearly every knob above lives in one file, on one profile: @@ -2432,6 +2433,42 @@ undeclared external-binary requirement on a headless host. --- +## One worker reply settles exactly one ticket + +**What.** When a worker's `fleet_reply` arrives with no send waiting for it, the bridge holds it and +later hands it to **one** delegation — never to several. If it cannot tell which one a held reply +answers, it queues the reply instead of guessing. Automatic — no knob. + +**On.** Nothing to turn on. + +**Why it exists.** Two separate paths used to assume a target had at most one delegation open, with +no guard for the case where it had two: + +- `fleet_send{turnId}` (the answer to a `fleet_ask`) waits only for its own bounded MCP window — + 25s by default, 120s at most. A resumed turn can easily run longer than that. When the window + expired, the worker's real reply had no waiter left, so it was held instead. The ticket stayed + `PENDING` until `fleet_stop`, which then reported *"the worker session was released before it + replied"* — with a worktree, a branch and a snapshot commit, so it read like lost work. The reply + had in fact arrived. Fixed first: a held reply now completes its own ticket (`388aba7`). +- Teardown then drained that held reply **once** and reused it for every open ticket on the target. + Two open tickets both came back `DONE` with the same text, one of them an answer the worker never + gave for that delegation. Fixed second (`966c58a`): the reply goes to the oldest open ticket and + every other one keeps the ordinary failure path. + +**The gotcha: a confidently wrong status costs more than no status.** Both defects produced a +*plausible* answer, not an error. A lead that trusted the first one would go recover a snapshot of +work that was already merged; a lead that trusted the second would act on a report its worker never +wrote. That is why the ambiguous case queues the reply and logs a WARN naming every candidate ticket +rather than picking one. + +**Known limit, on purpose.** Only the second defect's *broad* case is reachable today. Two open +tickets on one target is reachable and was a real bug. The same ambiguity in the reply path, and the +combination of two open tickets *with* a held reply, are both blocked by existing guards — one send +holds a target's lock and its waiter together, so a held reply and a second open ticket cannot exist +at the same instant. The guards stay as defence in depth, and `MessageService.abandon`'s javadoc +records why no test pins them: a test can force the state only by breaking that lock-and-waiter +invariant from outside the class, which no caller does. + ## Late-resolved member ids **What.** A member whose backend cannot name its own session at spawn time gets its