Features: one worker reply settles exactly one ticket (CB-137)

Dai Ha
2026-08-31 23:30:27 +07:00
parent ff48e2927e
commit dc565af709
+37
@@ -72,6 +72,7 @@ six weeks, and the table alone will not carry it.
| [Fill in a member id the backend names late](#late-resolved-member-ids) | automatic | CB-209 | `session/SessionManager` |
| [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` |
| [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` |
| [One reply settles one ticket](#one-worker-reply-settles-exactly-one-ticket) | automatic | CB-137 | `msg/MessageService` |
Nearly every knob above lives in one file, on one profile:
@@ -2432,6 +2433,42 @@ undeclared external-binary requirement on a headless host.
---
## One worker reply settles exactly one ticket
**What.** When a worker's `fleet_reply` arrives with no send waiting for it, the bridge holds it and
later hands it to **one** delegation — never to several. If it cannot tell which one a held reply
answers, it queues the reply instead of guessing. Automatic — no knob.
**On.** Nothing to turn on.
**Why it exists.** Two separate paths used to assume a target had at most one delegation open, with
no guard for the case where it had two:
- `fleet_send{turnId}` (the answer to a `fleet_ask`) waits only for its own bounded MCP window —
25s by default, 120s at most. A resumed turn can easily run longer than that. When the window
expired, the worker's real reply had no waiter left, so it was held instead. The ticket stayed
`PENDING` until `fleet_stop`, which then reported *"the worker session was released before it
replied"* — with a worktree, a branch and a snapshot commit, so it read like lost work. The reply
had in fact arrived. Fixed first: a held reply now completes its own ticket (`388aba7`).
- Teardown then drained that held reply **once** and reused it for every open ticket on the target.
Two open tickets both came back `DONE` with the same text, one of them an answer the worker never
gave for that delegation. Fixed second (`966c58a`): the reply goes to the oldest open ticket and
every other one keeps the ordinary failure path.
**The gotcha: a confidently wrong status costs more than no status.** Both defects produced a
*plausible* answer, not an error. A lead that trusted the first one would go recover a snapshot of
work that was already merged; a lead that trusted the second would act on a report its worker never
wrote. That is why the ambiguous case queues the reply and logs a WARN naming every candidate ticket
rather than picking one.
**Known limit, on purpose.** Only the second defect's *broad* case is reachable today. Two open
tickets on one target is reachable and was a real bug. The same ambiguity in the reply path, and the
combination of two open tickets *with* a held reply, are both blocked by existing guards — one send
holds a target's lock and its waiter together, so a held reply and a second open ticket cannot exist
at the same instant. The guards stay as defence in depth, and `MessageService.abandon`'s javadoc
records why no test pins them: a test can force the state only by breaking that lock-and-waiter
invariant from outside the class, which no caller does.
## Late-resolved member ids
**What.** A member whose backend cannot name its own session at spawn time gets its