Once the turn-completion scrape resolves a ticket, a later fleet_reply is unroutable — and the error names the queue instead of the cause #781

Open
opened 2026-10-05 19:34:15 +02:00 by ltms · 0 comments
Owner

Measured 2026-10-05. A peer called fleet_reply and got:

reply 754a28ed-ab2f-4a24-a169-e50ef16da61f was returned as unroutable (queue not declared/owned)

Twice, from the same pane (term_65d106559d7e43, an observer on the work config dir), at about 19:12 and 19:23 CEST. A second observer pane (term_65d106559db144, ltms) called the same tool and got delivered — resolved the fleet_send that was waiting for it.

The account is not the cause — the ticket's state is

The obvious reading is "it differs by Claude account", and the reporting peer concluded exactly that: "the difference is between trinotes (work) and me (ltms), 3 tries in total." That is confounded. All three observations share the axis that actually matters, so they are one data point about the account, not three.

My delivery log for the refused pane's ticket:

19:10:28.733  MessageService    - async send task-368d62-6 -> term_65d106559d7e43
19:10:41.093  CompletionResolver - resolved send to term_65d106559d7e43 via
                                   turn-completion fallback (150 chars scraped)

The ticket was resolved by the scrape at 19:10:41, 13 seconds after delivery. The first fleet_reply came about 80 seconds later, the second about 12 minutes later. Both were answering a rendezvous that no longer existed, so its reply queue was gone — hence "queue not declared/owned".

The pane that succeeded replied while its ticket was still live: its message had been held by the prompt-box gate and was injected later, and it replied inside that same turn.

So the rule is: a fleet_reply after the completion fallback has resolved the ticket is unroutable. Authorization is not involved — Authz.java:162 is case REPLY, ASK -> caller.ownsSession(targetSession), and the reporting peer found its own audit.log line 19:13:22 observer:term_65d106559d7e43 REPLY term_65d106559d7e43 allowed. The gate passed; routing failed afterwards.

Why this is worse than it looks

The fallback fires fast — 13 seconds after delivery here. A peer that reads a message, thinks, and composes a considered answer can easily take longer than that. When it does:

  1. its turn has already been resolved with a scrape of its pane, which may be a tool-description line or the tail of the brief rather than an answer;
  2. its real answer is then rejected;
  3. the rejection blames a queue, so the reader looks at the broker instead of at the fallback that closed the rendezvous;
  4. the peer concludes "I cannot reply to a lead", which is false, and writes that wrong rule into its own notes. Both affected peers did exactly this, and one generalised it to "never use fleet_reply", which costs it the channel permanently.

This is the same family as the completion fallback already recorded in #757 and #780: a mechanism that exists to stop a stall also destroys the real answer, and reports success while doing it.

Proposed

  1. Say what actually happened. The refusal must name the cause: this turn was already resolved by the completion fallback, so the reply arrived too late. "queue not declared/owned" is an implementation detail pointing at the wrong subsystem.
  2. Do not discard a late reply. A structured fleet_reply is strictly better evidence than a pane scrape. Prefer it: either keep the rendezvous readable long enough to accept a late reply and supersede the scraped content, or deliver the late reply to the ticket as an amendment the sender can poll.
  3. Reconsider the fallback's trigger delay. Resolving 13 seconds after delivery means any peer that composes a careful answer loses the channel. Whatever the bound is, it should be long enough that a normal reply wins the race, and it should be a named constant.
  4. If a late reply genuinely cannot be accepted, the sender should at least learn that one was attempted and rejected. Today only the replier sees the refusal; the lead sees the scrape and has no idea a real answer existed.

Not measured

I did not reproduce this myself — the two refusals and their reply ids are the peer's report, relayed through a third pane, and the audit.log line is that third pane's read. What I measured is the ticket lifecycle above, from my own fleetd.out, which is what makes the timing argument. I did not check whether the rendezvous is torn down by the resolver itself or expires separately, and I did not test a late reply against a ticket resolved by a real fleet_reply rather than by a scrape.

Measured 2026-10-05. A peer called `fleet_reply` and got: ``` reply 754a28ed-ab2f-4a24-a169-e50ef16da61f was returned as unroutable (queue not declared/owned) ``` Twice, from the same pane (`term_65d106559d7e43`, an observer on the `work` config dir), at about 19:12 and 19:23 CEST. A second observer pane (`term_65d106559db144`, `ltms`) called the same tool and got `delivered — resolved the fleet_send that was waiting for it`. ## The account is not the cause — the ticket's state is The obvious reading is "it differs by Claude account", and the reporting peer concluded exactly that: *"the difference is between trinotes (`work`) and me (`ltms`), 3 tries in total."* That is confounded. All three observations share the axis that actually matters, so they are one data point about the account, not three. My delivery log for the refused pane's ticket: ``` 19:10:28.733 MessageService - async send task-368d62-6 -> term_65d106559d7e43 19:10:41.093 CompletionResolver - resolved send to term_65d106559d7e43 via turn-completion fallback (150 chars scraped) ``` The ticket was **resolved by the scrape at 19:10:41**, 13 seconds after delivery. The first `fleet_reply` came about 80 seconds later, the second about 12 minutes later. Both were answering a rendezvous that no longer existed, so its reply queue was gone — hence "queue not declared/owned". The pane that succeeded replied while its ticket was still live: its message had been held by the prompt-box gate and was injected later, and it replied inside that same turn. So the rule is: **a `fleet_reply` after the completion fallback has resolved the ticket is unroutable.** Authorization is not involved — `Authz.java:162` is `case REPLY, ASK -> caller.ownsSession(targetSession)`, and the reporting peer found its own `audit.log` line `19:13:22 observer:term_65d106559d7e43 REPLY term_65d106559d7e43 allowed`. The gate passed; routing failed afterwards. ## Why this is worse than it looks The fallback fires **fast** — 13 seconds after delivery here. A peer that reads a message, thinks, and composes a considered answer can easily take longer than that. When it does: 1. its turn has already been resolved with a scrape of its pane, which may be a tool-description line or the tail of the brief rather than an answer; 2. its real answer is then rejected; 3. the rejection blames a queue, so the reader looks at the broker instead of at the fallback that closed the rendezvous; 4. the peer concludes "I cannot reply to a lead", which is false, and writes that wrong rule into its own notes. Both affected peers did exactly this, and one generalised it to "never use `fleet_reply`", which costs it the channel permanently. This is the same family as the completion fallback already recorded in #757 and #780: a mechanism that exists to stop a stall also destroys the real answer, and reports success while doing it. ## Proposed 1. **Say what actually happened.** The refusal must name the cause: this turn was already resolved by the completion fallback, so the reply arrived too late. "queue not declared/owned" is an implementation detail pointing at the wrong subsystem. 2. **Do not discard a late reply.** A structured `fleet_reply` is strictly better evidence than a pane scrape. Prefer it: either keep the rendezvous readable long enough to accept a late reply and supersede the scraped content, or deliver the late reply to the ticket as an amendment the sender can poll. 3. **Reconsider the fallback's trigger delay.** Resolving 13 seconds after delivery means any peer that composes a careful answer loses the channel. Whatever the bound is, it should be long enough that a normal reply wins the race, and it should be a named constant. 4. If a late reply genuinely cannot be accepted, the sender should at least learn that one was attempted and rejected. Today only the replier sees the refusal; the lead sees the scrape and has no idea a real answer existed. ## Not measured I did not reproduce this myself — the two refusals and their reply ids are the peer's report, relayed through a third pane, and the `audit.log` line is that third pane's read. What I measured is the ticket lifecycle above, from my own `fleetd.out`, which is what makes the timing argument. I did not check whether the rendezvous is torn down by the resolver itself or expires separately, and I did not test a late reply against a ticket resolved by a real `fleet_reply` rather than by a scrape.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#781