Features: one worker reply settles exactly one ticket (CB-137)
+37
@@ -72,6 +72,7 @@ six weeks, and the table alone will not carry it.
|
||||
| [Fill in a member id the backend names late](#late-resolved-member-ids) | automatic | CB-209 | `session/SessionManager` |
|
||||
| [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` |
|
||||
| [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` |
|
||||
| [One reply settles one ticket](#one-worker-reply-settles-exactly-one-ticket) | automatic | CB-137 | `msg/MessageService` |
|
||||
|
||||
Nearly every knob above lives in one file, on one profile:
|
||||
|
||||
@@ -2432,6 +2433,42 @@ undeclared external-binary requirement on a headless host.
|
||||
---
|
||||
|
||||
|
||||
## One worker reply settles exactly one ticket
|
||||
|
||||
**What.** When a worker's `fleet_reply` arrives with no send waiting for it, the bridge holds it and
|
||||
later hands it to **one** delegation — never to several. If it cannot tell which one a held reply
|
||||
answers, it queues the reply instead of guessing. Automatic — no knob.
|
||||
|
||||
**On.** Nothing to turn on.
|
||||
|
||||
**Why it exists.** Two separate paths used to assume a target had at most one delegation open, with
|
||||
no guard for the case where it had two:
|
||||
|
||||
- `fleet_send{turnId}` (the answer to a `fleet_ask`) waits only for its own bounded MCP window —
|
||||
25s by default, 120s at most. A resumed turn can easily run longer than that. When the window
|
||||
expired, the worker's real reply had no waiter left, so it was held instead. The ticket stayed
|
||||
`PENDING` until `fleet_stop`, which then reported *"the worker session was released before it
|
||||
replied"* — with a worktree, a branch and a snapshot commit, so it read like lost work. The reply
|
||||
had in fact arrived. Fixed first: a held reply now completes its own ticket (`388aba7`).
|
||||
- Teardown then drained that held reply **once** and reused it for every open ticket on the target.
|
||||
Two open tickets both came back `DONE` with the same text, one of them an answer the worker never
|
||||
gave for that delegation. Fixed second (`966c58a`): the reply goes to the oldest open ticket and
|
||||
every other one keeps the ordinary failure path.
|
||||
|
||||
**The gotcha: a confidently wrong status costs more than no status.** Both defects produced a
|
||||
*plausible* answer, not an error. A lead that trusted the first one would go recover a snapshot of
|
||||
work that was already merged; a lead that trusted the second would act on a report its worker never
|
||||
wrote. That is why the ambiguous case queues the reply and logs a WARN naming every candidate ticket
|
||||
rather than picking one.
|
||||
|
||||
**Known limit, on purpose.** Only the second defect's *broad* case is reachable today. Two open
|
||||
tickets on one target is reachable and was a real bug. The same ambiguity in the reply path, and the
|
||||
combination of two open tickets *with* a held reply, are both blocked by existing guards — one send
|
||||
holds a target's lock and its waiter together, so a held reply and a second open ticket cannot exist
|
||||
at the same instant. The guards stay as defence in depth, and `MessageService.abandon`'s javadoc
|
||||
records why no test pins them: a test can force the state only by breaking that lock-and-waiter
|
||||
invariant from outside the class, which no caller does.
|
||||
|
||||
## Late-resolved member ids
|
||||
|
||||
**What.** A member whose backend cannot name its own session at spawn time gets its
|
||||
|
||||
Reference in New Issue
Block a user