A send to a pane running a subagent is accepted, polls as "pending — worker working", and is never delivered, never timed out, and never reported #780

Open
opened 2026-10-05 19:26:23 +02:00 by ltms · 2 comments
Owner

Measured 2026-10-05 between 19:01 and 19:25 on this host. Reported independently by the anki lead and by the operator, who put it plainly: the fleet channel is the only sanctioned channel, and two sessions ended up talking to me over Claude Code's SendMessage instead because the fleet channel silently failed. That is the bug. "Held, nothing lost" is not a pass.

Every layer reports success or progress. Nothing was delivered.

Layer What it reports The truth
fleet_send accepted — task delegated. Poll … ticket=task-368d62-7 never delivered
fleet_poll{ticket} [pending — worker working] the target never received it; it is busy on unrelated work
fleetd/logs/audit.log SEND allowed authorization only — it records no delivery
fleetd.out no Injector line of any kind for that terminal no injection was ever attempted
the receiver nothing in its conversation or on its pane correct

pending — worker working is the most harmful of these. It reads as "the target is working on your message". The target never saw it.

The cause: a live subagent makes a pane WORKING forever, and the receiver does not know

The anki pane (term_65d106559db144, an observer) started a subagent during a long turn. From herdr's view the pane stays WORKING for as long as that subagent runs. The injector delivers only to idle, blocked or done, so there is no window.

Meanwhile the receiver's own loop believes it is idle and waiting. It told me so, in the same minute that:

fleet_status{sessionId: term_65d106559db144}  ->  working

So the two sides disagree, and neither is lying: the main loop is waiting, the pane is busy. Nothing reconciles them, and nothing tells the receiver that messages are queued for it.

Five sends, one arrival

From the receiver's audit.log (its own read), every one authorized:

18:57:30  observer:trinotes -> anki
19:01:44  observer:trinotes -> anki
19:11:44  leader:opus       -> anki
19:13:30  observer:trinotes -> anki
19:14:12  observer:trinotes -> anki

Exactly one reached the receiver's conversation. My own three (task-368d62-7, -8, -9) have an accept line in fleetd.out and nothing after it.

No timeout, which is what makes it unrecoverable

The member path does eventually give up — a readiness grace that measured 83276ms, 85642ms and 102791ms today, ending in a WARN and a failed ticket (#757). This path has no equivalent. The only signal in 22 minutes was the prompt-box gate's single WARN after 20 consecutive holds:

19:03:18  PromptBox - prompt box of term_...db144 is DRAFT (27 character(s)), holding delivery 1
          ... 39 consecutive holds ...
19:08:35  WARN PromptBox - has held a delivery 20 times in a row — nothing is lost
19:17:06  ReplyPushLoop - lead term_...db144 is WORKING (not injectable), waiting
          ... still going at 19:25:09 ...

Note those are the nudge loop's holds, not my messages' — my messages have no delivery record at all. Two separate things are stuck and only one of them logs.

Also: the prompt-box gate is a second, independent cause

For ten minutes the same pane held every delivery because it had 27 characters of unsubmitted text in its input box. That gate is correct and deliberate — a delivery pastes and submits in one step, so it would submit a half-written line. But combined with the above it means a pane can be unreachable for two unrelated reasons, with one WARN between them.

The asymmetry that made this invisible

Claude Code's SendMessage has no prompt-box gate and no status gate. So a pane can be fully reachable on that channel and completely unreachable on the fleet channel at the same moment, with no error on either side. Both affected sessions drifted onto SendMessage without anyone deciding to, which is how a silent failure becomes a habit.

Proposed, in order of value

  1. Tell the sender the truth. fleet_poll must distinguish never injected from delivered and being worked on. pending — worker working must not be returned for a message that was never put in the pane. Report queued-not-delivered, the target's status, and how long it has been that way.
  2. Refuse or warn at accept time when the target is not injectable, as #757 proposal 1 already argues. The information is in hand: fleet_status answers working synchronously.
  3. Give this path a timeout, like the readiness grace. A message that cannot be delivered must fail loudly rather than wait forever.
  4. Expose queue depth per pane in fleet_list's panes row, so a receiver can see that messages are waiting for it and a sender can see the backlog.
  5. Decide what a subagent means for injectability. A pane busy with a subagent for an hour is not a pane that will be free shortly. Either surface it as a distinct state or deliver anyway when the main loop is at a turn boundary.

Not measured

I did not confirm from inside the receiver that its subagent was still running at 19:25 — the subagent is its own report, and fleet_status answering working is mine. I did not test whether clearing the prompt box alone releases the queue while the pane stays WORKING; the two causes overlapped in this window, and separating them needs a clean run.

Measured 2026-10-05 between 19:01 and 19:25 on this host. Reported independently by the anki lead and by the operator, who put it plainly: the fleet channel is the only sanctioned channel, and two sessions ended up talking to me over Claude Code's `SendMessage` instead **because the fleet channel silently failed.** That is the bug. "Held, nothing lost" is not a pass. ## Every layer reports success or progress. Nothing was delivered. | Layer | What it reports | The truth | |---|---|---| | `fleet_send` | `accepted — task delegated. Poll … ticket=task-368d62-7` | never delivered | | `fleet_poll{ticket}` | `[pending — worker working]` | the target never received it; it is busy on unrelated work | | `fleetd/logs/audit.log` | SEND allowed | authorization only — it records no delivery | | `fleetd.out` | **no Injector line of any kind** for that terminal | no injection was ever attempted | | the receiver | nothing in its conversation or on its pane | correct | `pending — worker working` is the most harmful of these. It reads as "the target is working on your message". The target never saw it. ## The cause: a live subagent makes a pane WORKING forever, and the receiver does not know The anki pane (`term_65d106559db144`, an observer) started a subagent during a long turn. From herdr's view the pane stays `WORKING` for as long as that subagent runs. The injector delivers only to `idle`, `blocked` or `done`, so there is no window. Meanwhile the receiver's own loop believes it is idle and waiting. It told me so, in the same minute that: ``` fleet_status{sessionId: term_65d106559db144} -> working ``` So the two sides disagree, and **neither is lying**: the main loop is waiting, the pane is busy. Nothing reconciles them, and nothing tells the receiver that messages are queued for it. ## Five sends, one arrival From the receiver's `audit.log` (its own read), every one authorized: ``` 18:57:30 observer:trinotes -> anki 19:01:44 observer:trinotes -> anki 19:11:44 leader:opus -> anki 19:13:30 observer:trinotes -> anki 19:14:12 observer:trinotes -> anki ``` Exactly **one** reached the receiver's conversation. My own three (`task-368d62-7`, `-8`, `-9`) have an accept line in `fleetd.out` and nothing after it. ## No timeout, which is what makes it unrecoverable The member path does eventually give up — a readiness grace that measured 83276ms, 85642ms and 102791ms today, ending in a WARN and a failed ticket (#757). **This path has no equivalent.** The only signal in 22 minutes was the prompt-box gate's single WARN after 20 consecutive holds: ``` 19:03:18 PromptBox - prompt box of term_...db144 is DRAFT (27 character(s)), holding delivery 1 ... 39 consecutive holds ... 19:08:35 WARN PromptBox - has held a delivery 20 times in a row — nothing is lost 19:17:06 ReplyPushLoop - lead term_...db144 is WORKING (not injectable), waiting ... still going at 19:25:09 ... ``` Note those are the **nudge** loop's holds, not my messages' — my messages have no delivery record at all. Two separate things are stuck and only one of them logs. ## Also: the prompt-box gate is a second, independent cause For ten minutes the same pane held every delivery because it had **27 characters of unsubmitted text** in its input box. That gate is correct and deliberate — a delivery pastes and submits in one step, so it would submit a half-written line. But combined with the above it means a pane can be unreachable for two unrelated reasons, with one WARN between them. ## The asymmetry that made this invisible Claude Code's `SendMessage` has **no** prompt-box gate and **no** status gate. So a pane can be fully reachable on that channel and completely unreachable on the fleet channel at the same moment, with no error on either side. Both affected sessions drifted onto `SendMessage` without anyone deciding to, which is how a silent failure becomes a habit. ## Proposed, in order of value 1. **Tell the sender the truth.** `fleet_poll` must distinguish *never injected* from *delivered and being worked on*. `pending — worker working` must not be returned for a message that was never put in the pane. Report queued-not-delivered, the target's status, and how long it has been that way. 2. **Refuse or warn at accept time** when the target is not injectable, as #757 proposal 1 already argues. The information is in hand: `fleet_status` answers `working` synchronously. 3. **Give this path a timeout**, like the readiness grace. A message that cannot be delivered must fail loudly rather than wait forever. 4. **Expose queue depth per pane** in `fleet_list`'s `panes` row, so a receiver can see that messages are waiting for it and a sender can see the backlog. 5. **Decide what a subagent means for injectability.** A pane busy with a subagent for an hour is not a pane that will be free shortly. Either surface it as a distinct state or deliver anyway when the main loop is at a turn boundary. ## Not measured I did not confirm from inside the receiver that its subagent was still running at 19:25 — the subagent is its own report, and `fleet_status` answering `working` is mine. I did not test whether clearing the prompt box alone releases the queue while the pane stays `WORKING`; the two causes overlapped in this window, and separating them needs a clean run.
Author
Owner

Correction to the issue body — one inference in it is wrong, and the real situation is worse

The table in the body says:

| fleetd.out | no Injector line of any kind for that terminal | no injection was ever attempted |

The right-hand column is wrong. The absence of an Injector line does not mean the message was never injected.

Proof. task-368d62-7 has exactly one line in fleetd.out for its whole life:

19:11:44.126 DEBUG MessageService - async send task-368d62-7 -> term_65d106559db144

Nothing after it. No injection line, no delivery line, no resolution line. And yet that message was delivered: the receiver read it, called fleet_reply, and I collected the structured reply through fleet_poll{ticket=task-368d62-7}.

So the observable state is:

  • a message that was never delivered → accept line, nothing after
  • a message that WAS delivered and answered → accept line, nothing after

The two are indistinguishable in the log. The pane-delivery path has no observability at all, in either direction. I cannot even state when that message was finally injected, because nothing records it — the receiver guessed "about a minute after the box was cleared" and then retracted that itself for not having measured it, and my log cannot supply the number either.

This makes requirement 1 in the body more important, not less, and it adds a requirement:

6. Log the delivery. When a queued message is actually injected into a pane, say so, with the ticket id and the terminal — at the same level the accept is logged. Right now a successful pane delivery is completely silent, so neither an operator nor a lead can answer "did it arrive?" from the daemon's own records. Compare the member path, which logs READY -> BUSY turn=1 and then a resolution; the pane path logs neither.

Anyone diagnosing this must not use "no Injector line" as evidence of non-delivery. I did, in the body above and in a note to a peer, and it was unsound. The correct evidence that a message did not arrive is the receiver's own conversation, plus fleet_status showing the pane was never injectable in the window — not the log's silence.

The rest of the body stands: the pending — worker working reporting, the missing timeout, and the two independent gates (prompt-box draft, and WORKING for a long subagent turn) were all measured and are unaffected by this correction.

## Correction to the issue body — one inference in it is wrong, and the real situation is worse The table in the body says: | `fleetd.out` | **no Injector line of any kind** for that terminal | no injection was ever attempted | **The right-hand column is wrong.** The absence of an Injector line does not mean the message was never injected. Proof. `task-368d62-7` has exactly one line in `fleetd.out` for its whole life: ``` 19:11:44.126 DEBUG MessageService - async send task-368d62-7 -> term_65d106559db144 ``` Nothing after it. No injection line, no delivery line, no resolution line. And yet that message **was delivered**: the receiver read it, called `fleet_reply`, and I collected the structured reply through `fleet_poll{ticket=task-368d62-7}`. So the observable state is: - a message that was never delivered → accept line, nothing after - a message that WAS delivered and answered → accept line, nothing after **The two are indistinguishable in the log.** The pane-delivery path has no observability at all, in either direction. I cannot even state when that message was finally injected, because nothing records it — the receiver guessed "about a minute after the box was cleared" and then retracted that itself for not having measured it, and my log cannot supply the number either. This makes requirement 1 in the body more important, not less, and it adds a requirement: **6. Log the delivery.** When a queued message is actually injected into a pane, say so, with the ticket id and the terminal — at the same level the accept is logged. Right now a successful pane delivery is completely silent, so neither an operator nor a lead can answer "did it arrive?" from the daemon's own records. Compare the member path, which logs `READY -> BUSY turn=1` and then a resolution; the pane path logs neither. Anyone diagnosing this must not use "no Injector line" as evidence of non-delivery. I did, in the body above and in a note to a peer, and it was unsound. The correct evidence that a message did not arrive is the receiver's own conversation, plus `fleet_status` showing the pane was never injectable in the window — not the log's silence. The rest of the body stands: the `pending — worker working` reporting, the missing timeout, and the two independent gates (prompt-box draft, and `WORKING` for a long subagent turn) were all measured and are unaffected by this correction.
Author
Owner

Requirement 7 — deliverable: true is the third receipt that promises something it does not check

Measured just now. Both peer lead panes report themselves ready while every delivery to them is in fact being held:

{"paneId":"wA:p1","workspaceLabel":"trinotes","label":"lead","status":"idle","role":"observer","deliverable":true}
{"paneId":"wB:p1","workspaceLabel":"anki",    "label":"lead","status":"idle","role":"observer","deliverable":true}

At that moment wA:p1's input box held a line starting go and wB:p1's held cleared the input box, so the box gate was holding both. wB:p1 also showed ← 1 agent, a live subagent, so its status gate was holding it too.

The flag is honest about what it computes — it is just not what the name says. Fleetd.deliverableTo:

return target -> presence.isPresent(target) || leads.get().containsKey(target)
        || collaborators.get().containsKey(target);

That is enrolment and nothing else: has this pane connected the bridge MCP, or is it a known lead or collaborator. Its javadoc is explicit that it exists for one purpose, to hold a delivery out of a spawned member's boot window. The prompt-box gate and the status gate are applied later, at delivery time, and neither feeds this predicate.

So a lead diagnosing a silent channel has three receipts and all three say fine:

Receipt What it says What it checks
fleet_send accepted the mailbox took it
fleet_poll pending — worker working nothing about the pane
fleet_list → panes[].deliverable true enrolment only

This is the same defect as the rest of this ticket, at a third site: a receipt that reads a different source than the behaviour. It is the one that cost the most here, because deliverable is the field whose name makes a reader stop looking.

Add to the scope:

  1. panes[].deliverable must mean "a send would land now", or be renamed to what it does check. Renaming is acceptable and may be better — enrolled would be accurate and would not need to consult two more gates on every listing. What must not survive is a field called deliverable that is true for a pane where delivery is being held. If it keeps the name, it must consult the box and status gates; if it keeps the semantics, it needs the name enrolled and the panes row needs a separate field saying whether a send would land, with the reason when it would not (held: prompt box, held: working).

Requirement 4 already asks for queue depth in that row. The two belong together: depth plus reason is the whole diagnosis, and a receiver that can read "3 queued, held: prompt box" fixes it in seconds.

Note for whoever implements this: the box-gate half of the reason has its own defect in #782 — a no-break space the TUI draws is counted as the operator's typing, so the gate can report held: prompt box for a box that is genuinely empty. Do not treat a DRAFT reading as proof the operator typed something until that lands.

## Requirement 7 — `deliverable: true` is the third receipt that promises something it does not check Measured just now. Both peer lead panes report themselves ready while every delivery to them is in fact being held: ``` {"paneId":"wA:p1","workspaceLabel":"trinotes","label":"lead","status":"idle","role":"observer","deliverable":true} {"paneId":"wB:p1","workspaceLabel":"anki", "label":"lead","status":"idle","role":"observer","deliverable":true} ``` At that moment `wA:p1`'s input box held a line starting `go ` and `wB:p1`'s held `cleared the input box`, so the box gate was holding both. `wB:p1` also showed `← 1 agent`, a live subagent, so its status gate was holding it too. The flag is honest about what it computes — it is just not what the name says. `Fleetd.deliverableTo`: ```java return target -> presence.isPresent(target) || leads.get().containsKey(target) || collaborators.get().containsKey(target); ``` That is enrolment and nothing else: has this pane connected the bridge MCP, or is it a known lead or collaborator. Its javadoc is explicit that it exists for one purpose, to hold a delivery out of a *spawned* member's boot window. The prompt-box gate and the status gate are applied later, at delivery time, and neither feeds this predicate. So a lead diagnosing a silent channel has three receipts and all three say fine: | Receipt | What it says | What it checks | |---|---|---| | `fleet_send` | `accepted` | the mailbox took it | | `fleet_poll` | `pending — worker working` | nothing about the pane | | `fleet_list` → `panes[].deliverable` | `true` | enrolment only | This is the same defect as the rest of this ticket, at a third site: a receipt that reads a different source than the behaviour. It is the one that cost the most here, because `deliverable` is the field whose name makes a reader stop looking. Add to the scope: 7. **`panes[].deliverable` must mean "a send would land now", or be renamed to what it does check.** Renaming is acceptable and may be better — `enrolled` would be accurate and would not need to consult two more gates on every listing. What must not survive is a field called `deliverable` that is `true` for a pane where delivery is being held. If it keeps the name, it must consult the box and status gates; if it keeps the semantics, it needs the name `enrolled` and the `panes` row needs a separate field saying whether a send would land, with the reason when it would not (`held: prompt box`, `held: working`). Requirement 4 already asks for queue depth in that row. The two belong together: depth plus reason is the whole diagnosis, and a receiver that can read "3 queued, held: prompt box" fixes it in seconds. Note for whoever implements this: the box-gate half of the reason has its own defect in #782 — a no-break space the TUI draws is counted as the operator's typing, so the gate can report `held: prompt box` for a box that is genuinely empty. Do not treat a `DRAFT` reading as proof the operator typed something until that lands.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#780