CB-582: bridge_ask is unusable against an async lead — a 55s window vs a lead that polls every few minutes #61

Closed
opened 2026-08-15 10:35:33 +02:00 by ltms · 1 comment
Owner

What happened

While implementing CB-578 stage B, the worker did exactly what CLAUDE.md tells a member to do. The brief said:

Design that field, but bridge_ask me your proposed name and semantics before you build it, because it is a config surface we will live with.

It called bridge_ask. The ask timed out after ~55s with no answer, and it proceeded on its own judgment, as the timeout instructs. It then flagged this in its bridge_reply.

Nothing was broken on the worker's side. The lead — me — never saw the question. I had delegated with wait:false and was polling bridge_poll{ticket} every few minutes, which is the cadence the charter itself recommends for a long task:

prefer wait:false + bridge_poll for anything non-trivial: a blocking bridge_send is capped by your own MCP client call timeout (~60s), well below the task's real runtime.

The contradiction

The two mechanisms are tuned against each other:

  • bridge_send{wait:false} + poll is the recommended way to run any non-trivial delegation, and it makes the lead's attention sparse — minutes between polls.
  • bridge_ask blocks the worker's turn and expires in ~55s.

So for exactly the delegations the charter says to run asynchronously, bridge_ask can almost never succeed. The window is roughly two orders of magnitude smaller than the lead's poll interval. A member that follows the instruction to ask gets silence and has to guess.

That silence is also invisible to the lead. I polled the ticket repeatedly and only ever saw [pending — worker working]; the question opened and closed between two polls, and I learned it had happened only from the final reply, after the decision was already made and built.

Why it matters

bridge_ask exists for the case where a decision is genuinely the lead's — an ambiguous requirement, two defensible fixes, a config surface the project will live with. Those are precisely the decisions worth waiting for. Today the feature quietly degrades to "the worker decides anyway", which is the outcome asking was supposed to prevent. In this instance the worker chose well, so nothing was lost; that is luck, not design.

Possible directions (not a decision)

  1. Make a pending question visible to bridge_poll{ticket}. Today a poll of a ticket whose worker is mid-ask returns pending — worker working, which is indistinguishable from ordinary progress. A distinct state naming the question would at least let a polling lead answer on its next poll.
  2. Let the ask outlive the turn's blocking window — hold the question, let the worker end its turn, and resume it when the lead answers. This is closer to how wait:false delegation already works.
  3. Widen the ask timeout to something matched to a polling lead. Simplest, and the weakest: it only moves the race rather than removing it.
  4. Push the question at the lead the way CB-307's reply push loop already nudges a lead's own pane when a reply lands. A question is at least as worth a nudge as a reply.

Direction 1 looks like the cheapest real improvement, and it composes with the others.

Acceptance criteria (to be settled when this is picked up)

  1. A lead polling an async ticket can see that its worker is waiting on a question, and can read the question.
  2. A lead can answer it through the existing bridge_send{turnId} path.
  3. A worker whose ask goes unanswered still degrades to today's behaviour — decide and report — rather than blocking a turn indefinitely.
  4. Documented in wiki/11-Features.md, since it changes something an operator observes.

Evidence

  • The worker's own report on PR #60, section "Design note not settled by the brief".
  • CLAUDE.md §Member turn contract point 3, and §Primary step 4, which recommend the two settings that conflict.
## What happened While implementing CB-578 stage B, the worker did exactly what `CLAUDE.md` tells a member to do. The brief said: > Design that field, but **`bridge_ask` me your proposed name and semantics before you build it**, because it is a config surface we will live with. It called `bridge_ask`. The ask **timed out after ~55s with no answer**, and it proceeded on its own judgment, as the timeout instructs. It then flagged this in its `bridge_reply`. Nothing was broken on the worker's side. The lead — me — never saw the question. I had delegated with `wait:false` and was polling `bridge_poll{ticket}` every few minutes, which is the cadence the charter itself recommends for a long task: > prefer `wait:false` + `bridge_poll` for anything non-trivial: a blocking `bridge_send` is capped by *your own* MCP client call timeout (~60s), well below the task's real runtime. ## The contradiction The two mechanisms are tuned against each other: - **`bridge_send{wait:false}` + poll** is the recommended way to run any non-trivial delegation, and it makes the lead's attention *sparse* — minutes between polls. - **`bridge_ask`** blocks the worker's turn and expires in ~55s. So for exactly the delegations the charter says to run asynchronously, `bridge_ask` can almost never succeed. The window is roughly two orders of magnitude smaller than the lead's poll interval. A member that follows the instruction to ask gets silence and has to guess. That silence is also invisible to the lead. I polled the ticket repeatedly and only ever saw `[pending — worker working]`; the question opened and closed between two polls, and I learned it had happened only from the final reply, after the decision was already made and built. ## Why it matters `bridge_ask` exists for the case where a decision is genuinely the lead's — an ambiguous requirement, two defensible fixes, a config surface the project will live with. Those are precisely the decisions worth waiting for. Today the feature quietly degrades to "the worker decides anyway", which is the outcome asking was supposed to prevent. In this instance the worker chose well, so nothing was lost; that is luck, not design. ## Possible directions (not a decision) 1. **Make a pending question visible to `bridge_poll{ticket}`.** Today a poll of a ticket whose worker is mid-ask returns `pending — worker working`, which is indistinguishable from ordinary progress. A distinct state naming the question would at least let a polling lead answer on its next poll. 2. **Let the ask outlive the turn's blocking window** — hold the question, let the worker end its turn, and resume it when the lead answers. This is closer to how `wait:false` delegation already works. 3. **Widen the ask timeout** to something matched to a polling lead. Simplest, and the weakest: it only moves the race rather than removing it. 4. **Push the question at the lead** the way CB-307's reply push loop already nudges a lead's own pane when a reply lands. A question is at least as worth a nudge as a reply. Direction 1 looks like the cheapest real improvement, and it composes with the others. ## Acceptance criteria (to be settled when this is picked up) 1. A lead polling an async ticket can see that its worker is waiting on a question, and can read the question. 2. A lead can answer it through the existing `bridge_send{turnId}` path. 3. A worker whose ask goes unanswered still degrades to today's behaviour — decide and report — rather than blocking a turn indefinitely. 4. Documented in `wiki/11-Features.md`, since it changes something an operator observes. ## Evidence - The worker's own report on PR #60, section "Design note not settled by the brief". - `CLAUDE.md` §Member turn contract point 3, and §Primary step 4, which recommend the two settings that conflict.
ltms added this to the 1.1 — single-host close-out milestone 2026-08-16 16:49:37 +02:00
ltms closed this issue 2026-08-16 18:47:20 +02:00
Author
Owner

Merged into main as 83e2ff0, with a follow-up fix from my review as d56c77b.

Build, unpiped, on the trial merge onto current main: Tests run: 848, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. With my fix: 849.

What shipped

The question is now a third source in the CB-588 per-lead push schedule, not a new push path — so CB-590's one-schedule-per-lead guarantee still holds, and it carries its own per-item nudge count in CB-598's shape. I read that part closely, because it is the code I fixed twice this week: a fresh question stays eligible no matter how depleted an older open question's count is, and stopOrRestart folds the question set into the same raced-in check. Correct.

Two real gaps it also closed on the way, which the ticket did not name:

  • REST GET /tasks/{ticket} dropped turnId on an ASKING phase. A REST caller could see the question and had no way to answer it.
  • bridge_status said nothing about an open question at all.

What it deliberately did not do

It did not widen the ask timeout, and it checked why before deciding: DEFAULT_ASK_TIMEOUT_MS = 55_000 sits just under the worker's own MCP client cap of ~60s, so the daemon can return a clean typed timeout before the client severs the call. Widening the server constant would not buy a longer window — the worker's client kills the call regardless. Right call, and it said so rather than quietly skipping it.

The defect I found before merging

ask() clears its question on three paths — no-waiter, timed out, and answered. It can also leave by throwing: an interrupt while blocked on the answer, or an ExecutionException from the answer future. Those run only the finally block, which tore down the rendezvous turn but not the push loop's copy.

A question that took either path stayed pending for good — named in every nudge until it hit its own cap, then left in pendingQuestions with no remover at all.

Fixed in d56c77b by tearing down at the same point the rendezvous teardown already happens, so the two cannot drift apart again. The new test fails on the pre-fix code with expected: <STOP> but was: <INJECT>; I checked that both ways before committing.

Only the interrupt path is reachable today — nothing currently completes the ask future exceptionally. That is why this is a latent defect rather than a live one.

Standing advice unchanged

This closes the window, it does not remove it. A worker still gets ~55s. Do not brief a worker to "ask me" — decide first, or give it an explicit default. The nudge helps a lead who happens to be injectable; it cannot help one that is mid-turn for a minute.

Closing.

Merged into `main` as `83e2ff0`, with a follow-up fix from my review as `d56c77b`. **Build**, unpiped, on the trial merge onto current `main`: `Tests run: 848, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`. With my fix: **849**. ## What shipped The question is now a **third source in the CB-588 per-lead push schedule**, not a new push path — so CB-590's one-schedule-per-lead guarantee still holds, and it carries its own per-item nudge count in CB-598's shape. I read that part closely, because it is the code I fixed twice this week: a fresh question stays eligible no matter how depleted an older open question's count is, and `stopOrRestart` folds the question set into the same raced-in check. Correct. Two real gaps it also closed on the way, which the ticket did not name: - REST `GET /tasks/{ticket}` dropped `turnId` on an ASKING phase. A REST caller could see the question and had no way to answer it. - `bridge_status` said nothing about an open question at all. ## What it deliberately did not do It did **not** widen the ask timeout, and it checked why before deciding: `DEFAULT_ASK_TIMEOUT_MS = 55_000` sits just under the worker's own MCP client cap of ~60s, so the daemon can return a clean typed timeout before the client severs the call. Widening the server constant would not buy a longer window — the worker's client kills the call regardless. Right call, and it said so rather than quietly skipping it. ## The defect I found before merging `ask()` clears its question on three paths — no-waiter, timed out, and answered. It can also leave by **throwing**: an interrupt while blocked on the answer, or an `ExecutionException` from the answer future. Those run only the `finally` block, which tore down the rendezvous turn but not the push loop's copy. A question that took either path stayed pending for good — named in every nudge until it hit its own cap, then left in `pendingQuestions` with no remover at all. Fixed in `d56c77b` by tearing down at the same point the rendezvous teardown already happens, so the two cannot drift apart again. The new test fails on the pre-fix code with `expected: <STOP> but was: <INJECT>`; I checked that both ways before committing. Only the interrupt path is reachable today — nothing currently completes the ask future exceptionally. That is why this is a latent defect rather than a live one. ## Standing advice unchanged This closes the window, it does not remove it. A worker still gets ~55s. **Do not brief a worker to "ask me"** — decide first, or give it an explicit default. The nudge helps a lead who happens to be injectable; it cannot help one that is mid-turn for a minute. Closing.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#61