CB-582: make a pending bridge_ask question visible on the lead's poll cadence #104

Closed
agent wants to merge 0 commits from worker/cb582-477374-21 into main
Member

Closes gitea#61.

Where the ~55s window comes from

Server-side constants — BridgedApp.java:50 (DEFAULT_ASK_TIMEOUT_MS = 55_000) and BridgeMcp.java:65 (same value). But both are documented in-code (BridgeMcp.java:63-64, BridgedApp.java:49) as deliberately kept just under the worker's own MCP client's ~60s tool-call cap, so the daemon can return a clean typed timeout before the client severs the call. Widening the default further would not buy the lead more time — the worker's own client kills the call first regardless — so I left both DEFAULT_ASK_TIMEOUT_MS and MAX_ASK_TIMEOUT_MS untouched. Item 3 (widen the timeout) is not viable, not just weak.

What shipped

(1) Pending question already mostly visible — bridge_poll(ticket) already showed Phase.ASKING with question+turnId (CB-205). Two real gaps closed here:

  • REST GET /tasks/{ticket} never included turnId on an ASKING phase (only the MCP text did) — fixed in BridgedApp.taskStatus.
  • bridge_status(sessionId) and REST GET /sessions/{id}/status showed nothing about an open question at all — added MessageService.pendingAsk(workerSession) and wired it into both.

(4) Push nudge on question open — MessageService.ask() now calls a new ReplyPushLoop.onQuestionOpened(ticket, target, turnId, question) the moment a question opens for an async (wait:false) ticket, reusing the existing CB-588 per-lead schedule (a third source alongside replies and terminal tickets) rather than a new path. questionClosed(turnId) removes it when answered or lapsed. Bounded by the loop's existing maxReminders cap, same as the other two sources.

An unanswered question still behaves as today: TIMED_OUT, the worker proceeds, bridge_reply says the ask went unanswered — not a hard failure (untouched).

Tests

mvn -f bridged/pom.xml clean install, run unpiped:
Tests run: 845, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS

New coverage: ReplyPushLoopTest (question-nudge decide/inject/cap/coalesce), MessageServiceTest (pendingAsk(), end-to-end nudge-on-ask + nudge-stops-on-answer), BridgeMcpTest + BridgedAppTest (bridge_status / GET .../status showing the open question).

Closes gitea#61. ## Where the ~55s window comes from Server-side constants — `BridgedApp.java:50` (`DEFAULT_ASK_TIMEOUT_MS = 55_000`) and `BridgeMcp.java:65` (same value). But both are documented in-code (BridgeMcp.java:63-64, BridgedApp.java:49) as deliberately kept **just under** the worker's own MCP client's ~60s tool-call cap, so the daemon can return a clean typed timeout before the client severs the call. Widening the default further would not buy the lead more time — the worker's own client kills the call first regardless — so I left both DEFAULT_ASK_TIMEOUT_MS and MAX_ASK_TIMEOUT_MS untouched. Item 3 (widen the timeout) is not viable, not just weak. ## What shipped **(1) Pending question already mostly visible** — bridge_poll(ticket) already showed Phase.ASKING with question+turnId (CB-205). Two real gaps closed here: - REST GET /tasks/{ticket} never included turnId on an ASKING phase (only the MCP text did) — fixed in BridgedApp.taskStatus. - bridge_status(sessionId) and REST GET /sessions/{id}/status showed nothing about an open question at all — added MessageService.pendingAsk(workerSession) and wired it into both. **(4) Push nudge on question open** — MessageService.ask() now calls a new ReplyPushLoop.onQuestionOpened(ticket, target, turnId, question) the moment a question opens for an async (wait:false) ticket, reusing the existing CB-588 per-lead schedule (a third source alongside replies and terminal tickets) rather than a new path. questionClosed(turnId) removes it when answered or lapsed. Bounded by the loop's existing maxReminders cap, same as the other two sources. An unanswered question still behaves as today: TIMED_OUT, the worker proceeds, bridge_reply says the ask went unanswered — not a hard failure (untouched). ## Tests mvn -f bridged/pom.xml clean install, run unpiped: Tests run: 845, Failures: 0, Errors: 0, Skipped: 0 BUILD SUCCESS New coverage: ReplyPushLoopTest (question-nudge decide/inject/cap/coalesce), MessageServiceTest (pendingAsk(), end-to-end nudge-on-ask + nudge-stops-on-answer), BridgeMcpTest + BridgedAppTest (bridge_status / GET .../status showing the open question).
agent added 1 commit 2026-08-16 18:36:23 +02:00
CB-582: make a pending bridge_ask question visible on the lead's poll cadence
CI / build (pull_request) Successful in 1m23s
CI / contract (pull_request) Successful in 1m23s
08968bb1b7
bridge_ask blocks the worker's turn for ~55s by default (BridgedApp.java,
BridgeMcp.java) — a value deliberately kept just under the worker's own MCP
client's ~60s call cap so the daemon can return a clean timeout before the
client severs the call, not a value that can usefully be widened. A lead
following the charter's wait:false + poll cadence is minutes away, so the
window closes long before a poll would ever see the question — and until now
bridge_poll on such a ticket just read as ordinary "pending" progress.

bridge_poll(ticket) already surfaced Phase.ASKING with the question and
turnId (CB-205); this ships the two pieces that were still missing:

- The lead's own pane is now nudged the instant a question opens, reusing
  the CB-588 ReplyPushLoop push mechanism (a third source alongside queued
  replies and terminal tickets) rather than a new path. The nudge is capped
  by the loop's existing maxReminders budget, and stops the moment the
  question is answered or lapses.
- bridge_status(sessionId) and REST GET /sessions/{id}/status now also show
  an open question and how to answer it, via a new
  MessageService.pendingAsk() lookup — covering the case where a lead checks
  status directly rather than the ticket.
- The REST /tasks/{ticket} endpoint was silently missing turnId on an ASKING
  phase (only the MCP layer's formatted text carried it) — fixed as part of
  making the state genuinely visible over both surfaces.

An unanswered question still behaves as today: the worker proceeds and its
reply says the ask went unanswered — not a hard failure.
Owner

Merged locally into main as 83e2ff0 and pushed. Gitea cannot mark a locally merged PR as merged, so I am closing it by hand — merged, not rejected.

Good work, and thank you for checking why the 55s constant is 55s before deciding not to touch it. That reasoning is why I did not have to re-derive it.

One defect I found on review and fixed on top as d56c77b: ask() can leave by throwing (interrupt, or an exceptional answer future), and those paths run only the finally block — which tore down the rendezvous turn but not the push loop's question. It would have stayed pending for good. Detail on #61.

Merged locally into `main` as `83e2ff0` and pushed. Gitea cannot mark a locally merged PR as merged, so I am closing it by hand — **merged, not rejected**. Good work, and thank you for checking *why* the 55s constant is 55s before deciding not to touch it. That reasoning is why I did not have to re-derive it. One defect I found on review and fixed on top as `d56c77b`: `ask()` can leave by throwing (interrupt, or an exceptional answer future), and those paths run only the `finally` block — which tore down the rendezvous turn but not the push loop's question. It would have stayed pending for good. Detail on #61.
ltms closed this pull request 2026-08-16 18:47:45 +02:00
Some checks are pending
CI / build (pull_request) Successful in 1m23s
CI / contract (pull_request) Successful in 1m23s

Pull request closed

Sign in to join this conversation.