A blocking fleet_send that times out returns no ticket id, so its own advice to "poll status" cannot be followed #801

Closed
opened 2026-10-06 19:33:58 +02:00 by ltms · 4 comments
Owner

What happens

A blocking fleet_send (the default, wait unset) to a busy pane returns:

[no reply within 25000ms — worker busy; retry or poll status]

There is no ticket id and no msgId in that receipt. So the caller cannot poll, and the only action the message suggests — "retry" — creates a duplicate.

Measured today

I sent one message with the default wait:true to term_65d106559db144 at about 19:20:53 and got the receipt above. The daemon's own log shows it did create and resolve a ticket for that send:

19:21:18.527  MessageService - send to term_65d106559db144 timed out (delivered=true)
19:21:34.712  CompletionResolver - resolved send to term_65d106559db144 via turn-completion fallback (301 chars scraped)

So a ticket existed, went terminal 16 seconds later, and was pruned 10 minutes after that (TICKET_TTL_NANOS, MessageService.java:65). Its content is gone and I never had the id.

For contrast, wait:false to the same pane one minute later returned accepted — task delegated. Poll fleet_poll with ticket=task-7da785-70 immediately.

Why it matters

Three separate costs, in order of how much they hurt:

  1. The receipt names an action the caller cannot take. "Poll status" needs an id the receipt withholds.
  2. The suggested alternative is harmful. Retrying is the documented way to create a duplicate delivery, and the canonical CLAUDE.md block already warns a lead not to re-send because a call looks slow. The receipt advises the opposite.
  3. The reply is silently discarded. In this case the turn-completion fallback resolved the ticket with a 301-character pane scrape. That scrape was the receiver's answer and nobody could ever read it.

This is close in shape to the hazard in the block's step 5 — "a fleet_send to a working member is accepted and returns a ticket, and is then never delivered" — but it is the mirror image: here delivery is recorded as delivered=true and it is the handle that is lost, not the message.

Suggested fix

Return the ticket id in the timeout receipt, so the two paths read the same way:

[no reply within 25000ms — worker busy. The send is still tracked:
 poll fleet_poll with ticket=task-7da785-NN. Do NOT re-send — that creates a duplicate.]

Also drop the bare word "retry" from the message, or qualify it, since on this channel a retry is a known way to produce a duplicate rather than a safe repair.

Not verified

  • I have not read the code path that builds this receipt, so I do not know whether the ticket id is in scope at that point. If it is not, the fix is larger than the wording.
  • Whether the 25000ms figure is configurable, and whether it is the same bound as the caller's MCP client timeout, which the block gives as roughly 60s. Those two numbers disagree and I did not chase it.

Found while running a cross-session message-loss matrix; the box-gate half of that run is on #797.

## What happens A blocking `fleet_send` (the default, `wait` unset) to a busy pane returns: ``` [no reply within 25000ms — worker busy; retry or poll status] ``` There is **no ticket id and no msgId** in that receipt. So the caller cannot poll, and the only action the message suggests — "retry" — creates a duplicate. ## Measured today I sent one message with the default `wait:true` to `term_65d106559db144` at about 19:20:53 and got the receipt above. The daemon's own log shows it did create and resolve a ticket for that send: ``` 19:21:18.527 MessageService - send to term_65d106559db144 timed out (delivered=true) 19:21:34.712 CompletionResolver - resolved send to term_65d106559db144 via turn-completion fallback (301 chars scraped) ``` So a ticket existed, went terminal 16 seconds later, and was pruned 10 minutes after that (`TICKET_TTL_NANOS`, `MessageService.java:65`). Its content is gone and I never had the id. For contrast, `wait:false` to the same pane one minute later returned `accepted — task delegated. Poll fleet_poll with ticket=task-7da785-70` immediately. ## Why it matters Three separate costs, in order of how much they hurt: 1. **The receipt names an action the caller cannot take.** "Poll status" needs an id the receipt withholds. 2. **The suggested alternative is harmful.** Retrying is the documented way to create a duplicate delivery, and the canonical `CLAUDE.md` block already warns a lead not to re-send because a call looks slow. The receipt advises the opposite. 3. **The reply is silently discarded.** In this case the turn-completion fallback resolved the ticket with a 301-character pane scrape. That scrape was the receiver's answer and nobody could ever read it. This is close in shape to the hazard in the block's step 5 — "a `fleet_send` to a working member is *accepted* and returns a ticket, and is then never delivered" — but it is the mirror image: here delivery is recorded as `delivered=true` and it is the *handle* that is lost, not the message. ## Suggested fix Return the ticket id in the timeout receipt, so the two paths read the same way: ``` [no reply within 25000ms — worker busy. The send is still tracked: poll fleet_poll with ticket=task-7da785-NN. Do NOT re-send — that creates a duplicate.] ``` Also drop the bare word "retry" from the message, or qualify it, since on this channel a retry is a known way to produce a duplicate rather than a safe repair. ## Not verified - I have not read the code path that builds this receipt, so I do not know whether the ticket id is in scope at that point. If it is not, the fix is larger than the wording. - Whether the 25000ms figure is configurable, and whether it is the same bound as the caller's MCP client timeout, which the block gives as roughly 60s. Those two numbers disagree and I did not chase it. Found while running a cross-session message-loss matrix; the box-gate half of that run is on #797.
Author
Owner

This is worse than the missing ticket id: the message is lost

I filed this as a usability defect — the receipt withholds the ticket id. A cross-session loss-detection run has now shown the payload itself is discarded.

The measurement

Five messages to one pane (term_65d106559db144), tagged L-A1/5 … L-A5/5, sent in that order.

Message How sent Receipt Arrived?
L-A1/5 default wait:true [no reply within 25000ms — worker busy; retry or poll status], no ticket No
L-A2/5 wait:false accepted … ticket=task-7da785-70 Yes
L-A3/5 wait:false accepted … ticket=task-7da785-73 Yes
L-A4/5 wait:false accepted … ticket=task-7da785-76 Yes
L-A5/5 wait:false accepted … ticket=task-7da785-79 Yes

The receiving session reports L-A2 through L-A5 arrived, in order, no duplicates, each delivered by the fleet mod. For L-A1 it reports: never arrived by any route, not pasted, and fleet_inbox returned count: 0 twice afterwards. It also checked its own session transcript for the 19:21 window and found no garbled or merged prompt — so the message did not arrive in some mangled form either.

I did not retry L-A1, so this is a clean single-trial result rather than a duplicate hunt.

delivered=true in the log is false

19:21:18.527  MessageService - send to term_65d106559db144 timed out (delivered=true)
19:21:35      (receiver's transcript) 2492-char message arrives — my EARLIER phase-1 brief, ticket task-7da785-42
19:21:41.464  async send task-7da785-57 -> ...   (the receiver acting on that brief)

FIFO settles it. At 19:21:18 that pane's queue still held the phase-1 brief, which was not handed over until 19:21:35. A later message cannot be delivered ahead of an earlier one in the same queue, so L-A1 was not delivered at 19:21:18 and delivered=true is wrong.

The receiver initially proposed that the 2492-character message was L-A1. It was not: L-A1's body began "RUN7 L-A1/5 — phase 2 of the loss-detection matrix, lead → anki" and ran about 1100 characters. Different text, different length, and it carried an instruction the receiver demonstrably did not have until L-A5 repeated it.

Why this ranks high

On this channel, the blocking send is the unsafe one, and nothing says so. Every wait:false send in the run arrived; the single wait:true send is the only lost message in the entire matrix — against three legs that were a perfect 5 of 5 in order with no duplicates.

A caller therefore gets: no ticket, no msgId, a log line that claims delivery, and no message. There is no artefact anywhere that would let them notice. I only caught it because the receiver was counting numbered messages.

Revised fix

Beyond returning the ticket id:

  1. A timeout must not drop the queued message. The caller giving up on the reply is not a reason to cancel the delivery. It should stay queued and be delivered when the pane frees up, exactly as the wait:false path does.
  2. Correct or remove delivered=true. If the flag describes the mailbox rather than the pane, it must not use the word "delivered". This is the same confusion as the accepted-but-never-delivered hazard the canonical CLAUDE.md block already warns leads about, and the same shape as the related note that a receipt about the mailbox is not a fact about the pane.
  3. Consider making wait:true non-default, or refusing it to a pane whose status is not idle/blocked/done, since the wait:false path for the identical target is reliable.

Still not verified

  • I have not read the code that cancels on timeout, so I cannot say whether the drop is explicit or a side effect of the rendezvous being torn down.
  • Whether this also affects a blocking send to a spawned member, or only to an unconfigured pane. Every data point here is one observer pane.
## This is worse than the missing ticket id: the message is lost I filed this as a usability defect — the receipt withholds the ticket id. A cross-session loss-detection run has now shown the payload itself is discarded. ### The measurement Five messages to one pane (`term_65d106559db144`), tagged `L-A1/5` … `L-A5/5`, sent in that order. | Message | How sent | Receipt | Arrived? | |---|---|---|---| | `L-A1/5` | default `wait:true` | `[no reply within 25000ms — worker busy; retry or poll status]`, **no ticket** | **No** | | `L-A2/5` | `wait:false` | `accepted … ticket=task-7da785-70` | Yes | | `L-A3/5` | `wait:false` | `accepted … ticket=task-7da785-73` | Yes | | `L-A4/5` | `wait:false` | `accepted … ticket=task-7da785-76` | Yes | | `L-A5/5` | `wait:false` | `accepted … ticket=task-7da785-79` | Yes | The receiving session reports L-A2 through L-A5 arrived, in order, no duplicates, each delivered by the fleet mod. For L-A1 it reports: never arrived by any route, not pasted, and `fleet_inbox` returned `count: 0` twice afterwards. It also checked its own session transcript for the 19:21 window and found no garbled or merged prompt — so the message did not arrive in some mangled form either. **I did not retry L-A1**, so this is a clean single-trial result rather than a duplicate hunt. ### `delivered=true` in the log is false ``` 19:21:18.527 MessageService - send to term_65d106559db144 timed out (delivered=true) 19:21:35 (receiver's transcript) 2492-char message arrives — my EARLIER phase-1 brief, ticket task-7da785-42 19:21:41.464 async send task-7da785-57 -> ... (the receiver acting on that brief) ``` FIFO settles it. At 19:21:18 that pane's queue still held the phase-1 brief, which was not handed over until 19:21:35. A later message cannot be delivered ahead of an earlier one in the same queue, so `L-A1` was not delivered at 19:21:18 and `delivered=true` is wrong. The receiver initially proposed that the 2492-character message *was* `L-A1`. It was not: `L-A1`'s body began "RUN7 L-A1/5 — phase 2 of the loss-detection matrix, lead → anki" and ran about 1100 characters. Different text, different length, and it carried an instruction the receiver demonstrably did not have until `L-A5` repeated it. ### Why this ranks high On this channel, **the blocking send is the unsafe one**, and nothing says so. Every `wait:false` send in the run arrived; the single `wait:true` send is the only lost message in the entire matrix — against three legs that were a perfect 5 of 5 in order with no duplicates. A caller therefore gets: no ticket, no msgId, a log line that claims delivery, and no message. There is no artefact anywhere that would let them notice. I only caught it because the receiver was counting numbered messages. ### Revised fix Beyond returning the ticket id: 1. **A timeout must not drop the queued message.** The caller giving up on the *reply* is not a reason to cancel the *delivery*. It should stay queued and be delivered when the pane frees up, exactly as the `wait:false` path does. 2. **Correct or remove `delivered=true`.** If the flag describes the mailbox rather than the pane, it must not use the word "delivered". This is the same confusion as the accepted-but-never-delivered hazard the canonical `CLAUDE.md` block already warns leads about, and the same shape as the related note that a receipt about the mailbox is not a fact about the pane. 3. Consider making `wait:true` non-default, or refusing it to a pane whose status is not `idle`/`blocked`/`done`, since the `wait:false` path for the identical target is reliable. ### Still not verified - I have not read the code that cancels on timeout, so I cannot say whether the drop is explicit or a side effect of the rendezvous being torn down. - Whether this also affects a blocking send to a *spawned member*, or only to an unconfigured pane. Every data point here is one observer pane.
Author
Owner

Correcting myself. The "Measured today" section above says "the daemon's own log shows it did create and resolve a ticket for that send" and "So a ticket existed, went terminal 16 seconds later, and was pruned". That is wrong, and so is the suggested fix that follows from it.

I have now read the code. A blocking fleet_send creates no ticket at all, ever. new Task(...) appears exactly once in MessageService.java, at line 1346, inside the sendAsync path. The blocking path calls send(target, content, timeout, onAccepted, callerOwner), which passes no Task. The log line I quoted —

19:21:18.527  MessageService - send to term_65d106559db144 timed out (delivered=true)

— is MessageService.java:1036, the blocking send's own debug line. It is not evidence of a ticket. I read a ticket into it because the async path logs look similar, and I filed that as a measurement. It was an inference.

So "return the ticket id in the timeout receipt" is not implementable: there is no id to return.

What is actually wrong, after reading the code

1. The receipt names the wrong recovery route, and a real one exists. There are exactly two inbox.publish call sites in MessageService (lines 615 and 865). Line 615 is in the reply() path: when a worker calls fleet_reply and no open send or ticket matches it, the reply is published to that target's inbox and is drainable with fleet_poll{target=<sessionId>}. So a late fleet_reply after a blocking timeout is recoverable — and the receipt never mentions the one call that recovers it. It says "retry or poll status" instead, where "retry" is the documented way to create a duplicate.

2. A turn that ends with no fleet_reply is not recoverable, and that is the real loss. CompletionResolver (lines 440–453) resolves the scrape with rendezvous.resolveCompletion(waiter, completion) and has no inbox fallback on that path. So when the member ends its turn without fleet_reply, the scrape goes into a rendezvous waiter and nowhere else. That is what happened in my case: the 301-character scrape was the receiver's answer, and the only route that would have saved it — the inbox — is not on that path.

The asymmetry is the defect: a late structured reply survives a blocking timeout; a late unstructured completion does not.

What I have not verified

Whether the waiter resolveCompletion found at 19:21:34 was still connected to a caller. The call logged its success, so a waiter existed 16 seconds after the timeout; I did not check whether the blocking send's injector.cancel(delivery) at MessageService.java:1033 is supposed to have removed it. If it leaves the waiter in place deliberately, that is the seam where an inbox publish belongs.

Revised fix

  1. Change the timeout receipt to name the route that works: fleet_poll{target=<sessionId>} to drain the inbox. Remove the bare word "retry", or qualify it — on this channel a retry produces a duplicate, and the canonical block already tells a lead not to re-send because a call looks slow.
  2. Make the completion-scrape path publish to the inbox when its waiter has no live caller, so an unstructured answer survives a blocking timeout exactly as a structured one already does.
  3. State in the receipt that the send is not tracked by a ticket, so a caller stops looking for an id that was never issued.

The 25000ms figure is FleetMcp.DEFAULT_TIMEOUT_MS, clamped by MAX_TIMEOUT_MS = 120_000, and timeoutMs is a per-call argument. It is not the same bound as the caller's own MCP client timeout, which is roughly 60s and outside fleetd's control — so the two numbers in the original report disagree because they measure different things, and that part of the report was simply two unrelated figures placed side by side.


I asked Claude to read this code path; the line numbers and the grep results above are from this session, against main at 9f4b736.

**Correcting myself.** The "Measured today" section above says "the daemon's own log shows it did create and resolve a ticket for that send" and "So a ticket existed, went terminal 16 seconds later, and was pruned". That is wrong, and so is the suggested fix that follows from it. I have now read the code. **A blocking `fleet_send` creates no ticket at all, ever.** `new Task(...)` appears exactly once in `MessageService.java`, at line 1346, inside the `sendAsync` path. The blocking path calls `send(target, content, timeout, onAccepted, callerOwner)`, which passes no `Task`. The log line I quoted — ``` 19:21:18.527 MessageService - send to term_65d106559db144 timed out (delivered=true) ``` — is `MessageService.java:1036`, the blocking send's own debug line. It is not evidence of a ticket. I read a ticket into it because the async path logs look similar, and I filed that as a measurement. It was an inference. So "return the ticket id in the timeout receipt" is not implementable: there is no id to return. ## What is actually wrong, after reading the code **1. The receipt names the wrong recovery route, and a real one exists.** There are exactly two `inbox.publish` call sites in `MessageService` (lines 615 and 865). Line 615 is in the `reply()` path: when a worker calls `fleet_reply` and no open send or ticket matches it, the reply is published to that target's inbox and is drainable with `fleet_poll{target=<sessionId>}`. So a late `fleet_reply` after a blocking timeout **is** recoverable — and the receipt never mentions the one call that recovers it. It says "retry or poll status" instead, where "retry" is the documented way to create a duplicate. **2. A turn that ends with no `fleet_reply` is not recoverable, and that is the real loss.** `CompletionResolver` (lines 440–453) resolves the scrape with `rendezvous.resolveCompletion(waiter, completion)` and has no inbox fallback on that path. So when the member ends its turn without `fleet_reply`, the scrape goes into a rendezvous waiter and nowhere else. That is what happened in my case: the 301-character scrape was the receiver's answer, and the only route that would have saved it — the inbox — is not on that path. The asymmetry is the defect: **a late structured reply survives a blocking timeout; a late unstructured completion does not.** ## What I have not verified Whether the waiter `resolveCompletion` found at 19:21:34 was still connected to a caller. The call logged its success, so a waiter existed 16 seconds after the timeout; I did not check whether the blocking send's `injector.cancel(delivery)` at `MessageService.java:1033` is supposed to have removed it. If it leaves the waiter in place deliberately, that is the seam where an inbox publish belongs. ## Revised fix 1. Change the timeout receipt to name the route that works: `fleet_poll{target=<sessionId>}` to drain the inbox. Remove the bare word "retry", or qualify it — on this channel a retry produces a duplicate, and the canonical block already tells a lead not to re-send because a call looks slow. 2. Make the completion-scrape path publish to the inbox when its waiter has no live caller, so an unstructured answer survives a blocking timeout exactly as a structured one already does. 3. State in the receipt that the send is not tracked by a ticket, so a caller stops looking for an id that was never issued. The 25000ms figure is `FleetMcp.DEFAULT_TIMEOUT_MS`, clamped by `MAX_TIMEOUT_MS = 120_000`, and `timeoutMs` is a per-call argument. It is not the same bound as the caller's own MCP client timeout, which is roughly 60s and outside fleetd's control — so the two numbers in the original report disagree because they measure different things, and that part of the report was simply two unrelated figures placed side by side. --- I asked Claude to read this code path; the line numbers and the grep results above are from this session, against `main` at `9f4b736`.
Author
Owner

Lead review of 24782fe — one change still needed

Both defects I raised are fixed, and extending the same conditional to the REST receipt was right (GET /sessions/{id}/replies is indeed DRAIN). I checked FleetApp.allow myself: it returns true when auth == null (FleetApp.java:329-331, "legacy: authorization not enforced"), so the auth == null || mirror is accurate.

The tests are good. aLateCompletionThatFailsToPublishIsLoggedNotLost asserts hasStrandedReply(T) is false after the failed publish — that catches a flag that would otherwise lie. anArchitectsTimedOutSendDoesNotNameTheRepliesRouteItCannotDrain ends with a real 403 on /replies, which is the positive control that proves the receipt was right to stay quiet.

What must change: drop the two test-only overloads

FleetMcp.java:1036-1042 and 1074-1082 add short overloads of send and answer that default mayDrainPoll to true.

Measured on your branch: neither has a production caller. The only calls are the overloads' own bodies and 23 call sites in FleetMcpTest. The production handler always passes the real grant.

git grep -n 'FleetMcp.send(\|FleetMcp.answer(\|return send(messages\|return answer(messages' -- fleetd/src

Two reasons this has to go:

  1. It is this repo's code-quality rule 4 — "Never relieve a testing problem by reshaping production code." The parameter is the injectable part; the overload exists only so the existing test call sites need no new argument.
  2. The default points the unsafe way. true means "name fleet_poll{target}" — the exact receipt this ticket is fixing. If a default has to exist at all it must be false, because a receipt naming no route is harmless and one naming a refused route is the defect. Failing closed is what CLAUDE.md means by "fail toward the recoverable error".

Do this:

  • Delete both overloads. Keep one signature each: send(..., String callerOwner, boolean mayDrainPoll) and answer(..., String callerOwner, boolean mayDrainPoll).
  • Pass the value explicitly at every test call site. true where the test asserts the route is named, false where it asserts it is not. The two sites already passing false (FleetMcpTest.java:488 and 515) stay as they are.
  • Change no production behaviour and no other file.

Acceptance: mvn clean install passes in your worktree, and git grep -c 'callerOwner)$' shows no remaining 7-argument send or 6-argument answer declaration. Report the test count and the real build result, including a failure if you get one.

Nothing else on this PR needs changing — do not touch MessageService.java, FleetApp.java, or the REST tests.

## Lead review of `24782fe` — one change still needed Both defects I raised are fixed, and extending the same conditional to the REST receipt was right (`GET /sessions/{id}/replies` is indeed `DRAIN`). I checked `FleetApp.allow` myself: it returns `true` when `auth == null` (`FleetApp.java:329-331`, "legacy: authorization not enforced"), so the `auth == null ||` mirror is accurate. The tests are good. `aLateCompletionThatFailsToPublishIsLoggedNotLost` asserts `hasStrandedReply(T)` is false after the failed publish — that catches a flag that would otherwise lie. `anArchitectsTimedOutSendDoesNotNameTheRepliesRouteItCannotDrain` ends with a real `403` on `/replies`, which is the positive control that proves the receipt was right to stay quiet. ### What must change: drop the two test-only overloads `FleetMcp.java:1036-1042` and `1074-1082` add short overloads of `send` and `answer` that default `mayDrainPoll` to `true`. Measured on your branch: **neither has a production caller.** The only calls are the overloads' own bodies and 23 call sites in `FleetMcpTest`. The production handler always passes the real grant. ``` git grep -n 'FleetMcp.send(\|FleetMcp.answer(\|return send(messages\|return answer(messages' -- fleetd/src ``` Two reasons this has to go: 1. It is this repo's code-quality rule 4 — "Never relieve a testing problem by reshaping production code." The parameter is the injectable part; the overload exists only so the existing test call sites need no new argument. 2. The default points the unsafe way. `true` means "name `fleet_poll{target}`" — the exact receipt this ticket is fixing. If a default has to exist at all it must be `false`, because a receipt naming no route is harmless and one naming a refused route is the defect. Failing closed is what `CLAUDE.md` means by "fail toward the recoverable error". **Do this:** - Delete both overloads. Keep one signature each: `send(..., String callerOwner, boolean mayDrainPoll)` and `answer(..., String callerOwner, boolean mayDrainPoll)`. - Pass the value explicitly at every test call site. `true` where the test asserts the route is named, `false` where it asserts it is not. The two sites already passing `false` (`FleetMcpTest.java:488` and `515`) stay as they are. - Change no production behaviour and no other file. **Acceptance:** `mvn clean install` passes in your worktree, and `git grep -c 'callerOwner)$'` shows no remaining 7-argument `send` or 6-argument `answer` declaration. Report the test count and the real build result, including a failure if you get one. Nothing else on this PR needs changing — do not touch `MessageService.java`, `FleetApp.java`, or the REST tests.
Author
Owner

Fixed and merged to main as 11352d0. PR #807 is closed (merged locally, not through the forge).

Measured on the merge result: mvn clean install in a throwaway worktree, no [ERROR] lines from maven, Tests run: 2244, Failures: 0, Errors: 0, Skipped: 0 tallied from 183 surefire XML files. The staged tree matched the tree I built exactly (517590e).

What shipped:

  • Both the MCP receipt and the REST detail now carry MessageService.NO_TICKET_NO_RESEND — one constant, so the two surfaces cannot drift — and name the one route that recovers a late answer.
  • That route is named only when the caller's own Authz.Action.DRAIN is granted. fleet_poll{target} and GET /sessions/{id}/replies both resolve to DRAIN, which is primary-only, while SEND reaches an architect, a collaborator and an observer too. Everyone else is told the reply cannot be recovered on their channel.
  • MessageService.strandLateResolution publishes a late turn-completion to the target's inbox. Only Kind.COMPLETION is routed: an explicit late fleet_reply already lands there via reply()'s own lookup once rendezvous.close has run, so routing it here as well would publish it twice. I checked that close runs in send()'s finally (MessageService.java:1061-1064) before concluding it.
  • The publish is wrapped and logged, because a throw inside a whenComplete action is captured by the discarded dependent stage and reaches no caller.

Documentation, per this repo's rule that the prompt ships with the code: CLAUDE.md and the wiki template now state that a blocking send creates no ticket (a2cbf50, byte-identical check passes), and wiki/11-Features.md has the entry (fleetd.wiki f812689).

One correction to this ticket's own record

My earlier comment already retracted the claim that the daemon log showed a ticket being created for a blocking send. Worth stating why it was wrong, because the mistake is reusable: both the blocking and the async paths log send to term_…, so their output is not a discriminator, and I read a ticket into a line written by the path that never creates one. grep -n 'new Task(' MessageService.java returns exactly one hit, inside sendAsync. The original suggested fix — "return the ticket id in the timeout receipt" — was therefore unimplementable, and the real finding turned out to be better than the invented one.

Closing.

Fixed and merged to `main` as `11352d0`. PR #807 is closed (merged locally, not through the forge). **Measured on the merge result:** `mvn clean install` in a throwaway worktree, no `[ERROR]` lines from maven, `Tests run: 2244, Failures: 0, Errors: 0, Skipped: 0` tallied from 183 surefire XML files. The staged tree matched the tree I built exactly (`517590e`). What shipped: - Both the MCP receipt and the REST `detail` now carry `MessageService.NO_TICKET_NO_RESEND` — one constant, so the two surfaces cannot drift — and name the one route that recovers a late answer. - That route is named **only when the caller's own `Authz.Action.DRAIN` is granted**. `fleet_poll{target}` and `GET /sessions/{id}/replies` both resolve to `DRAIN`, which is primary-only, while `SEND` reaches an architect, a collaborator and an observer too. Everyone else is told the reply cannot be recovered on their channel. - `MessageService.strandLateResolution` publishes a late turn-completion to the target's inbox. Only `Kind.COMPLETION` is routed: an explicit late `fleet_reply` already lands there via `reply()`'s own lookup once `rendezvous.close` has run, so routing it here as well would publish it twice. I checked that `close` runs in `send()`'s `finally` (`MessageService.java:1061-1064`) before concluding it. - The publish is wrapped and logged, because a throw inside a `whenComplete` action is captured by the discarded dependent stage and reaches no caller. Documentation, per this repo's rule that the prompt ships with the code: `CLAUDE.md` and the wiki template now state that a blocking send creates no ticket (`a2cbf50`, byte-identical check passes), and `wiki/11-Features.md` has the entry (`fleetd.wiki f812689`). ### One correction to this ticket's own record My earlier comment already retracted the claim that the daemon log showed a ticket being created for a blocking send. Worth stating why it was wrong, because the mistake is reusable: both the blocking and the async paths log `send to term_…`, so their output is not a discriminator, and I read a ticket into a line written by the path that never creates one. `grep -n 'new Task(' MessageService.java` returns exactly one hit, inside `sendAsync`. The original suggested fix — "return the ticket id in the timeout receipt" — was therefore unimplementable, and the real finding turned out to be better than the invented one. Closing.
ltms closed this issue 2026-10-07 06:18:40 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#801