CB-629: after a bridge_ask round-trip, the worker's final reply orphans its ticket, which then reports a false failure #137

Closed
opened 2026-08-22 22:01:59 +02:00 by ltms · 2 comments
Owner

Hit live while running CB-622 (#126). No work was lost, but the bridge told the lead something untrue about its own state.

What happened

Three workers were delegated with bridge_send{wait:false}, giving tickets task-1, task-2, task-3. All three did their work, replied, and opened PRs that are now merged.

Only one of them used bridge_ask mid-turn. That is the only one whose ticket did not resolve.

unit ticket used bridge_ask how the reply arrived final ticket state
A task-1 no bridge_poll{ticket} returned the report done
C task-3 no bridge_poll{ticket} returned the report done
B task-2 yes the runtime told the lead to use bridge_poll{target} instead failed

For Unit B the notification said:

Worker term_... returned a reply — run bridge_poll(target=term_...) to collect it

so the reply went to the session inbox rather than resolving the ticket. Later, when the lead tore the member down with bridge_stop, the still-open ticket went terminal and reported:

[failed — the worker session was released before it replied;
 worktree=... branch=worker/cb-622b-717c67-2 snapshot=bdda0461c12a82ea8509a2ea19e406c2c6ca7501]

"before it replied" is false. It replied, in full, and the lead had already collected the reply and merged the PR.

Why it matters

The failure text names a real recovery path — a worktree, a branch, a snapshot commit — so it reads as "your worker died with uncommitted work, go dig it out". A lead that trusts it will either redo finished work or spend time recovering a snapshot of work that is already merged. A lead that has learned to distrust it has lost the value of ticket status entirely.

This is worse than a missing status. A status that is confidently wrong is more expensive than no status.

Likely cause, stated as a hypothesis

The bridge_ask reverse rendezvous appears to detach the turn from its originating ticket. After the lead answers with bridge_send{turnId}, the turn resumes, but the final bridge_reply is routed to the session inbox instead of completing the ticket that started the turn. The ticket then stays open until something else makes it terminal — here, bridge_stop.

I did not read the routing code, so treat the mechanism as unverified. The observation is solid: the only unit with an ask is the only unit whose ticket failed, and the runtime itself redirected the lead to target for exactly that unit.

Scope

  1. Find where a resumed turn's bridge_reply is routed, and make it complete the originating ticket when the turn began as a wait:false delegation.
  2. If keeping the inbox path is deliberate, then the ticket must still resolve as done, not stay open.
  3. Never report "the session was released before it replied" when a reply was delivered. If a reply reached any sink for that turn, the ticket's terminal state is done. Reserve the released-without-replying text for turns with no reply anywhere.
  4. Consider whether the snapshot/worktree recovery hint should be printed at all when a reply was received — it is the part that most strongly implies lost work.

Acceptance criteria

  • A wait:false delegation whose worker calls bridge_ask, gets an answer, and replies, ends with bridge_poll{ticket} returning that reply — not a failure.
  • A bridge_stop on a member that already replied never turns its ticket into a failure.
  • A test drives the full path: delegate async, worker asks, lead answers by turnId, worker replies, then poll the ticket. Note that a test calling the reply sink directly will pass without proving any of this — the whole defect is in which sink the resumed turn reaches, so the test has to start from the real delegation.

Related

  • CB-582 / #61 covers the bridge_ask window; this is a different failure at the end of the same flow.
  • The reply-inbox read is destructive, so a lead that polls by target first and by ticket second sees the ticket failure with no reply left to contradict it. That ordering makes this bug maximally confusing.
Hit live while running CB-622 (#126). No work was lost, but the bridge told the lead something untrue about its own state. ## What happened Three workers were delegated with `bridge_send{wait:false}`, giving tickets `task-1`, `task-2`, `task-3`. All three did their work, replied, and opened PRs that are now merged. Only one of them used `bridge_ask` mid-turn. That is the only one whose ticket did not resolve. | unit | ticket | used `bridge_ask` | how the reply arrived | final ticket state | |---|---|---|---|---| | A | task-1 | no | `bridge_poll{ticket}` returned the report | done | | C | task-3 | no | `bridge_poll{ticket}` returned the report | done | | **B** | **task-2** | **yes** | the runtime told the lead to use `bridge_poll{target}` instead | **failed** | For Unit B the notification said: > Worker `term_...` returned a reply — run `bridge_poll(target=term_...)` to collect it so the reply went to the session inbox rather than resolving the ticket. Later, when the lead tore the member down with `bridge_stop`, the still-open ticket went terminal and reported: ``` [failed — the worker session was released before it replied; worktree=... branch=worker/cb-622b-717c67-2 snapshot=bdda0461c12a82ea8509a2ea19e406c2c6ca7501] ``` **"before it replied" is false.** It replied, in full, and the lead had already collected the reply and merged the PR. ## Why it matters The failure text names a real recovery path — a worktree, a branch, a snapshot commit — so it reads as "your worker died with uncommitted work, go dig it out". A lead that trusts it will either redo finished work or spend time recovering a snapshot of work that is already merged. A lead that has learned to distrust it has lost the value of ticket status entirely. This is worse than a missing status. A status that is confidently wrong is more expensive than no status. ## Likely cause, stated as a hypothesis The `bridge_ask` reverse rendezvous appears to detach the turn from its originating ticket. After the lead answers with `bridge_send{turnId}`, the turn resumes, but the final `bridge_reply` is routed to the session inbox instead of completing the ticket that started the turn. The ticket then stays open until something else makes it terminal — here, `bridge_stop`. I did not read the routing code, so treat the mechanism as unverified. The observation is solid: the only unit with an ask is the only unit whose ticket failed, and the runtime itself redirected the lead to `target` for exactly that unit. ## Scope 1. Find where a resumed turn's `bridge_reply` is routed, and make it complete the originating ticket when the turn began as a `wait:false` delegation. 2. If keeping the inbox path is deliberate, then **the ticket must still resolve as done**, not stay open. 3. **Never report "the session was released before it replied" when a reply was delivered.** If a reply reached any sink for that turn, the ticket's terminal state is done. Reserve the released-without-replying text for turns with no reply anywhere. 4. Consider whether the snapshot/worktree recovery hint should be printed at all when a reply was received — it is the part that most strongly implies lost work. ## Acceptance criteria - A `wait:false` delegation whose worker calls `bridge_ask`, gets an answer, and replies, ends with `bridge_poll{ticket}` returning that reply — not a failure. - A `bridge_stop` on a member that already replied never turns its ticket into a failure. - A test drives the full path: delegate async, worker asks, lead answers by `turnId`, worker replies, then poll the **ticket**. Note that a test calling the reply sink directly will pass without proving any of this — the whole defect is in which sink the resumed turn reaches, so the test has to start from the real delegation. ## Related - CB-582 / #61 covers the `bridge_ask` window; this is a different failure at the end of the same flow. - The reply-inbox read is destructive, so a lead that polls by `target` first and by `ticket` second sees the ticket failure with no reply left to contradict it. That ordering makes this bug maximally confusing.
ltms added this to the 2.0 — one operation centre, many hosts milestone 2026-08-22 22:01:59 +02:00
ltms closed this issue 2026-08-31 09:27:34 +02:00
Author
Owner

Fixed and merged to main as 388aba7. Closing.

The hypothesis in this ticket was close, but the real cause is narrower

This ticket guessed that the fleet_ask reverse rendezvous detaches the turn from its ticket. It does not. The actual cause:

answer() — behind fleet_send{turnId} — waits only for the lead's own bounded MCP call window. A resumed turn doing real work (more edits, a build, a commit, a push, opening a PR) routinely outlives that window. On timeout, answer()'s finally closes the rendezvous waiter. The worker's later fleet_reply then finds no waiter, falls to the session inbox, and the async ticket's future is never completed — so fleet_poll{ticket} stays PENDING until fleet_stop forces it WORKER_FAILED with the false "released before it replied" text.

So the trigger is a timing mismatch, not a routing mistake. That also explains why the observation held so cleanly: the unit that used fleet_ask is the only one whose reply had to survive a second bounded wait.

What shipped, against the ticket's four points

  1. Point 1 — done, the strong way. reply() now finds the async task parked on this exact answered turn and completes it with the real reply. fleet_poll{ticket} returns the reply; the inbox is not involved at all for this case.
  2. Point 2 — not needed, since point 1 was achievable directly.
  3. Point 3 — done as independent defence in depth. abandon() checks for a stranded reply before writing a failure, and reports REPLIED with the real text instead. Kept as a separate check on purpose: if some future path ever strands a reply without completing its ticket, a released session still will not claim the worker never replied.
  4. Point 4 — falls out. Once a task resolves as REPLIED, the worktree/branch/snapshot recovery hint — the part that most strongly implies lost work — is never constructed.

Why completing "the parked task" cannot grab the wrong ticket

Worth recording, because it is the one thing that could have made this fix dangerous. answer() calls clearAsyncQuestion(turnId, forgetTurn = false): the question is cleared but the task keeps its turnId and stays in asyncTasksByTurn. So hasAsyncQuestion(target) still returns true, and send() returns BUSY for any second async send to that target. At most one candidate task can exist per target at a time. I checked this in the code rather than taking it from the report.

Verification

mvn clean install unpiped: 1049 tests, exit 0.

Both new tests drive the full delegation path — async delegate → deliver → ask → answer(turnId, …) with a genuinely expiring timeout → reply → poll the ticket. No test hands a constructed Reply to a sink, which is the trap this ticket explicitly warned about. Both were sabotage-proven (expected: <DONE> but was: <PENDING>).

Still worth doing

Nobody has dogfooded this live — it is verified only through the JUnit fixture's simulated timing, never against a real herdr pane and a real MCP client's call cap. A short live check (an async delegation that asks, gets answered, then takes more than two minutes to reply) would close that gap.

Noted, not investigated

answer()'s handling of a nested second fleet_ask — a worker asking again before replying to the first answer — has pre-existing turnId attribution behaviour that nobody untangled. It predates this change and is untouched by it. Flagging only.

Fixed and merged to `main` as `388aba7`. Closing. ## The hypothesis in this ticket was close, but the real cause is narrower This ticket guessed that the `fleet_ask` reverse rendezvous detaches the turn from its ticket. It does not. The actual cause: `answer()` — behind `fleet_send{turnId}` — waits only for the **lead's own bounded MCP call window**. A resumed turn doing real work (more edits, a build, a commit, a push, opening a PR) routinely outlives that window. On timeout, `answer()`'s `finally` closes the rendezvous waiter. The worker's later `fleet_reply` then finds no waiter, falls to the session inbox, and the async ticket's future is never completed — so `fleet_poll{ticket}` stays PENDING until `fleet_stop` forces it `WORKER_FAILED` with the false "released before it replied" text. So the trigger is a **timing mismatch**, not a routing mistake. That also explains why the observation held so cleanly: the unit that used `fleet_ask` is the only one whose reply had to survive a second bounded wait. ## What shipped, against the ticket's four points 1. **Point 1 — done, the strong way.** `reply()` now finds the async task parked on this exact answered turn and completes it with the real reply. `fleet_poll{ticket}` returns the reply; the inbox is not involved at all for this case. 2. **Point 2 — not needed**, since point 1 was achievable directly. 3. **Point 3 — done as independent defence in depth.** `abandon()` checks for a stranded reply *before* writing a failure, and reports `REPLIED` with the real text instead. Kept as a separate check on purpose: if some future path ever strands a reply without completing its ticket, a released session still will not claim the worker never replied. 4. **Point 4 — falls out.** Once a task resolves as `REPLIED`, the worktree/branch/snapshot recovery hint — the part that most strongly implies lost work — is never constructed. ## Why completing "the parked task" cannot grab the wrong ticket Worth recording, because it is the one thing that could have made this fix dangerous. `answer()` calls `clearAsyncQuestion(turnId, forgetTurn = false)`: the question is cleared but the task keeps its `turnId` and stays in `asyncTasksByTurn`. So `hasAsyncQuestion(target)` still returns true, and `send()` returns `BUSY` for any second async send to that target. At most one candidate task can exist per target at a time. I checked this in the code rather than taking it from the report. ## Verification `mvn clean install` unpiped: **1049 tests**, exit 0. Both new tests drive the full delegation path — async delegate → deliver → `ask` → `answer(turnId, …)` with a genuinely expiring timeout → `reply` → poll the **ticket**. No test hands a constructed `Reply` to a sink, which is the trap this ticket explicitly warned about. Both were sabotage-proven (`expected: <DONE> but was: <PENDING>`). ## Still worth doing Nobody has dogfooded this live — it is verified only through the JUnit fixture's simulated timing, never against a real herdr pane and a real MCP client's call cap. A short live check (an async delegation that asks, gets answered, then takes more than two minutes to reply) would close that gap. ## Noted, not investigated `answer()`'s handling of a **nested** second `fleet_ask` — a worker asking again before replying to the first answer — has pre-existing `turnId` attribution behaviour that nobody untangled. It predates this change and is untouched by it. Flagging only.
Author
Owner

Fixed on main in two parts:

  • 388aba7 — after answer()'s own bounded wait times out, the worker's later fleet_reply completes its own async ticket instead of leaving it PENDING until fleet_stop turns it into a false WORKER_FAILED.
  • 966c58a — the follow-up: one stranded reply settles at most one open ticket. Before this, abandon() drained the reply once and reused it for every open task on the target, so two open tickets both reported DONE with the same text.

Full suite on the merge: Tests run: 1071, Failures: 0, Errors: 0, Skipped: 0.

One latent coupling found while doing this, left alone on purpose. MessageService.pendingAsk(String workerSession) returns the first task whose question is still open for that target. That is correct today only because hasAsyncQuestion guarantees at most one open question per target. If that gate is ever loosened, pendingAsk starts returning an arbitrary match with no warning. It is the same shape as the two defects fixed here — "assume at most one, with no guard" — so it is worth knowing about, but it is not reachable now and adding a guard would be guessing at what the loosened rule should be. Recorded here rather than as its own issue.

Fixed on `main` in two parts: - `388aba7` — after `answer()`'s own bounded wait times out, the worker's later `fleet_reply` completes its own async ticket instead of leaving it `PENDING` until `fleet_stop` turns it into a false `WORKER_FAILED`. - `966c58a` — the follow-up: one stranded reply settles at most one open ticket. Before this, `abandon()` drained the reply once and reused it for every open task on the target, so two open tickets both reported `DONE` with the same text. Full suite on the merge: Tests run: 1071, Failures: 0, Errors: 0, Skipped: 0. **One latent coupling found while doing this, left alone on purpose.** `MessageService.pendingAsk(String workerSession)` returns the *first* task whose question is still open for that target. That is correct today only because `hasAsyncQuestion` guarantees at most one open question per target. If that gate is ever loosened, `pendingAsk` starts returning an arbitrary match with no warning. It is the same shape as the two defects fixed here — "assume at most one, with no guard" — so it is worth knowing about, but it is not reachable now and adding a guard would be guessing at what the loosened rule should be. Recorded here rather than as its own issue.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#137