A restart de-enrols every hand-opened pane, and the failed send hands the caller a pane scrape instead of the daemon's own reason #757

Open
opened 2026-10-05 09:55:30 +02:00 by ltms · 2 comments
Owner

Measured on this host immediately after redeploying 02eff4c (daemon pid 9796, started 09:52:45, jar b0b40b6ccd4b). This is the first thing #743's new panes array made visible, and it is the limitation that most affects the goal #743 was opened for.

What I measured

fleet_list right after the restart — six observer panes, every one deliverable: false:

{"sessionId":"term_65d106559ba0a2","label":"vms","status":"idle","role":"observer","deliverable":false}
{"sessionId":"term_65d106559d7e43","label":"trinotes","status":"done","role":"observer","deliverable":false}
{"sessionId":"term_65d106559db144","label":"anki","status":"idle","role":"observer","deliverable":false}
{"sessionId":"term_65d106559e8365","label":"adev","status":"done","role":"observer","deliverable":false}
{"sessionId":"term_65d106559f8cc7","label":"70-review-approvals","status":"working","role":"observer","deliverable":false}
{"sessionId":"term_65d10655a01a98","label":"70-tidy","status":"done","role":"observer","deliverable":false}

The lead's own pane reads deliverable: true, because a lead is matched by the lead map rather than by presence.

Before the restart the same trinotes pane took a fleet_send on the first try (task-bfc545-5 -> term_65d106559d7e43 at 08:30:19, delivered). So the panes did not change; the daemon's memory of them did.

Cause, read from the code

inject/MemberPresence is a plain in-memory set:

private final Set<String> present = ConcurrentHashMap.newKeySet();

Nothing persists it and nothing rebuilds it at boot. Enrolment happens only on an MCP request — FleetMcp's contextExtractor → markTrackedCallerPresent. A spawned member is fine: it boots, connects, and is marked present within seconds. A hand-opened pane that has been idle for hours has no reason to issue another MCP request, so it stays unreachable until a person touches that tab.

So the #743 claim "no daemon restart needed" is true for enrolment and misleading about the lifecycle: a restart de-enrols everyone, and only the panes someone happens to interact with come back.

The second half, which is the actionable one

A send to a non-deliverable pane is accepted, not refused:

fleet_send{sessionId: term_65d106559d7e43, wait: false}  ->  "accepted — task delegated. Poll ... ticket=task-e4c8c2-1"

The daemon logged the accept and then nothing:

09:53:48.595 DEBUG [boundedElastic-1] d.l.fleet.msg.MessageService - async send task-e4c8c2-1 -> term_65d106559d7e43

No injector line, no refusal. fleet_poll reported pending — worker done for the whole window I watched. The ticket will sit on ASYNC_TIMEOUT_MS — 30 minutes (msg/MessageService.java:58), not the ~60s the CLAUDE.md block describes for a spawned member's presence gate.

The information needed to fail fast was already in hand: the caller had just read deliverable: false for that exact terminal from fleet_list, computed by Fleetd.deliverableTo — the same predicate the injector will later consult. The send path does not check it at accept time.

That makes a receipt that is true about the mailbox and useless about the pane, which is the shape already recorded in #640's family: a caller reads "accepted", believes it is on its way, and learns 30 minutes later that it never left.

Proposed, in order of value

  1. Refuse, or at least warn, at accept time when deliverableTo is already false for the target. A refusal naming the reason ("not present on the bridge MCP; its agent has not contacted the daemon since the daemon started") is strictly more useful than a 30-minute silence. This is small and does not need the rest.
  2. Re-derive presence at boot for panes herdr already reports as running an agent. The daemon calls agent.list anyway — LeadTabScanner and the new PaneLocator both do — so the fact that a pane is alive and running a Claude is available without waiting for that pane to speak first. Needs care: presence currently means "its MCP client has connected", which is a stronger claim than "herdr says an agent is there", and the Injector relies on the stronger one to avoid pasting into a booting TUI. So this is a design question, not a patch — possibly a second, weaker state rather than a widening of this one.
  3. If neither lands, say so in the docs. I have put the measurement in wiki/11-Features.md already.

Not measured

I did not wait out the full 30 minutes to watch the ticket turn timed_out_working, and I did not test whether touching one of those tabs re-enrols it — both follow from the code but I have not run them. The probe message I sent to trinotes says it needs no action, so it is harmless if it does eventually land.

Measured on this host immediately after redeploying `02eff4c` (daemon pid 9796, started 09:52:45, jar `b0b40b6ccd4b`). This is the first thing #743's new `panes` array made visible, and it is the limitation that most affects the goal #743 was opened for. ## What I measured `fleet_list` right after the restart — six observer panes, **every one** `deliverable: false`: ``` {"sessionId":"term_65d106559ba0a2","label":"vms","status":"idle","role":"observer","deliverable":false} {"sessionId":"term_65d106559d7e43","label":"trinotes","status":"done","role":"observer","deliverable":false} {"sessionId":"term_65d106559db144","label":"anki","status":"idle","role":"observer","deliverable":false} {"sessionId":"term_65d106559e8365","label":"adev","status":"done","role":"observer","deliverable":false} {"sessionId":"term_65d106559f8cc7","label":"70-review-approvals","status":"working","role":"observer","deliverable":false} {"sessionId":"term_65d10655a01a98","label":"70-tidy","status":"done","role":"observer","deliverable":false} ``` The lead's own pane reads `deliverable: true`, because a lead is matched by the lead map rather than by presence. Before the restart the same `trinotes` pane took a `fleet_send` on the first try (`task-bfc545-5 -> term_65d106559d7e43` at 08:30:19, delivered). So the panes did not change; the daemon's memory of them did. ## Cause, read from the code `inject/MemberPresence` is a plain in-memory set: ```java private final Set<String> present = ConcurrentHashMap.newKeySet(); ``` Nothing persists it and nothing rebuilds it at boot. Enrolment happens only on an MCP request — `FleetMcp`'s `contextExtractor` → `markTrackedCallerPresent`. A spawned member is fine: it boots, connects, and is marked present within seconds. A hand-opened pane that has been idle for hours has no reason to issue another MCP request, so it stays unreachable until a person touches that tab. So the #743 claim "no daemon restart needed" is true for *enrolment* and misleading about the lifecycle: **a restart de-enrols everyone**, and only the panes someone happens to interact with come back. ## The second half, which is the actionable one A send to a non-deliverable pane is **accepted**, not refused: ``` fleet_send{sessionId: term_65d106559d7e43, wait: false} -> "accepted — task delegated. Poll ... ticket=task-e4c8c2-1" ``` The daemon logged the accept and then nothing: ``` 09:53:48.595 DEBUG [boundedElastic-1] d.l.fleet.msg.MessageService - async send task-e4c8c2-1 -> term_65d106559d7e43 ``` No injector line, no refusal. `fleet_poll` reported `pending — worker done` for the whole window I watched. The ticket will sit on `ASYNC_TIMEOUT_MS` — **30 minutes** (`msg/MessageService.java:58`), not the ~60s the `CLAUDE.md` block describes for a spawned member's presence gate. The information needed to fail fast was already in hand: the caller had just read `deliverable: false` for that exact terminal from `fleet_list`, computed by `Fleetd.deliverableTo` — the same predicate the injector will later consult. The send path does not check it at accept time. That makes a receipt that is true about the mailbox and useless about the pane, which is the shape already recorded in #640's family: a caller reads "accepted", believes it is on its way, and learns 30 minutes later that it never left. ## Proposed, in order of value 1. **Refuse, or at least warn, at accept time** when `deliverableTo` is already false for the target. A refusal naming the reason ("not present on the bridge MCP; its agent has not contacted the daemon since the daemon started") is strictly more useful than a 30-minute silence. This is small and does not need the rest. 2. **Re-derive presence at boot for panes herdr already reports as running an agent.** The daemon calls `agent.list` anyway — `LeadTabScanner` and the new `PaneLocator` both do — so the fact that a pane is alive and running a Claude is available without waiting for that pane to speak first. Needs care: presence currently means "its MCP client has connected", which is a stronger claim than "herdr says an agent is there", and the `Injector` relies on the stronger one to avoid pasting into a booting TUI. So this is a design question, not a patch — possibly a second, weaker state rather than a widening of this one. 3. If neither lands, say so in the docs. I have put the measurement in `wiki/11-Features.md` already. ## Not measured I did not wait out the full 30 minutes to watch the ticket turn `timed_out_working`, and I did not test whether touching one of those tabs re-enrols it — both follow from the code but I have not run them. The probe message I sent to `trinotes` says it needs no action, so it is harmless if it does eventually land.
Author
Owner

Correction: two claims in the body above are wrong, and the real defect is narrower

The probe ticket finished while I was writing that body, and it falsified the part I had predicted rather than measured. Correcting it here rather than leaving a wrong number in a ticket other people read.

Wrong: "it will sit on ASYNC_TIMEOUT_MS — 30 minutes"

It failed in 94 seconds. Sent 09:53:48, failed 09:55:22:

09:55:22.961 WARN [status-poller] d.ltms.fleet.inject.Injector - readiness grace for
  term_65d106559d7e43 expired after 240 polls (configured=240 polls/60s elapsed=94020ms):
  target never became deliverable, so failing 1 queued message(s) that never reached its pane

The governing clock is the Injector's readiness grace, not the ticket clock. I reached for ASYNC_TIMEOUT_MS because I had been reading it an hour earlier for #588 — the wrong clock, picked by availability. Worse, the canonical CLAUDE.md block already states this correctly in invariant 4: "a send waits on that gate for ~60s and then fails without ever reaching its pane." I had the right answer in the instruction file I maintain and did not connect it.

Wrong: "No injector line, no refusal"

There is an injector line, and it is an excellent one. I grepped at 09:54, before the 09:55:22 event existed, and reported an absence as a fact. An absence measured before the deadline is not an absence.

What survives, and it is still worth fixing

Two things hold up:

  1. The restart de-enrols every pane. All six observer panes read deliverable: false after the redeploy. MemberPresence is an in-memory set with no boot rebuild. Unchanged.

  2. The caller gets a pane scrape instead of the daemon's accurate reason — this is the real defect, and it is sharper than what I first wrote. The daemon knows exactly what happened and says so in the log: "target never became deliverable, so failing 1 queued message(s) that never reached its pane." The caller never sees that sentence. Instead:

09:55:22.974 WARN [turn-failed-term_65d106559d7e43] d.l.f.i.CompletionResolver -
  failing send to term_65d106559d7e43 via turn-stall fallback: hen writes the same cards into each

fleet_poll{ticket} returned [failed — hen writes the same cards into each …] followed by ~4000 characters of that pane's screen. The scrape is not a reply. It contained the pane's unrelated current work, and — worst of all — the pane's answer to a different probe I sent at 08:30, including a verbatim fleet_whoami result. A reader skimming that output would reasonably conclude the pane had just answered and that delivery worked. It did not; the message never reached the pane at all.

So the fix I proposed as item 1 changes shape. Refusing at accept time would still be good, but the cheaper and more valuable fix is: when a send fails because the target was never deliverable, return the Injector's own reason, not a CompletionResolver scrape. The scrape fallback is right for a member that worked and ended its turn without replying; it is actively harmful for a message that was never delivered, because it manufactures the appearance of an answer.

This is the same shape as the clipped-scrape trap already known here: a scrape is not a reply, and a scrape presented where a reply belongs invites exactly the wrong conclusion. The turn-stall path should be able to say "nothing was delivered" rather than handing back whatever is on screen.

Revised proposal

  1. On a readiness-grace expiry, surface the Injector's message to the ticket. Do not scrape.
  2. Optionally refuse at accept time when deliverableTo is already false — still useful, now secondary.
  3. Boot-time presence re-derivation stays a design question, unchanged from the body above.

I have retitled this ticket, because the old title carried the wrong 30-minute number.

## Correction: two claims in the body above are wrong, and the real defect is narrower The probe ticket finished while I was writing that body, and it falsified the part I had predicted rather than measured. Correcting it here rather than leaving a wrong number in a ticket other people read. ### Wrong: "it will sit on `ASYNC_TIMEOUT_MS` — 30 minutes" It failed in **94 seconds**. Sent 09:53:48, failed 09:55:22: ``` 09:55:22.961 WARN [status-poller] d.ltms.fleet.inject.Injector - readiness grace for term_65d106559d7e43 expired after 240 polls (configured=240 polls/60s elapsed=94020ms): target never became deliverable, so failing 1 queued message(s) that never reached its pane ``` The governing clock is the `Injector`'s readiness grace, not the ticket clock. I reached for `ASYNC_TIMEOUT_MS` because I had been reading it an hour earlier for #588 — the wrong clock, picked by availability. Worse, the canonical `CLAUDE.md` block already states this correctly in invariant 4: *"a send waits on that gate for ~60s and then fails without ever reaching its pane."* I had the right answer in the instruction file I maintain and did not connect it. ### Wrong: "No injector line, no refusal" There is an injector line, and it is an excellent one. I grepped at 09:54, before the 09:55:22 event existed, and reported an absence as a fact. An absence measured before the deadline is not an absence. ### What survives, and it is still worth fixing Two things hold up: 1. **The restart de-enrols every pane.** All six observer panes read `deliverable: false` after the redeploy. `MemberPresence` is an in-memory set with no boot rebuild. Unchanged. 2. **The caller gets a pane scrape instead of the daemon's accurate reason** — this is the real defect, and it is sharper than what I first wrote. The daemon knows exactly what happened and says so in the log: *"target never became deliverable, so failing 1 queued message(s) that never reached its pane."* The caller never sees that sentence. Instead: ``` 09:55:22.974 WARN [turn-failed-term_65d106559d7e43] d.l.f.i.CompletionResolver - failing send to term_65d106559d7e43 via turn-stall fallback: hen writes the same cards into each ``` `fleet_poll{ticket}` returned `[failed — hen writes the same cards into each …]` followed by ~4000 characters of that pane's screen. The scrape is **not** a reply. It contained the pane's unrelated current work, and — worst of all — the pane's answer to a *different* probe I sent at 08:30, including a verbatim `fleet_whoami` result. A reader skimming that output would reasonably conclude the pane had just answered and that delivery worked. It did not; the message never reached the pane at all. So the fix I proposed as item 1 changes shape. Refusing at accept time would still be good, but the cheaper and more valuable fix is: **when a send fails because the target was never deliverable, return the `Injector`'s own reason, not a `CompletionResolver` scrape.** The scrape fallback is right for a member that worked and ended its turn without replying; it is actively harmful for a message that was never delivered, because it manufactures the appearance of an answer. This is the same shape as the clipped-scrape trap already known here: a scrape is not a reply, and a scrape presented where a reply belongs invites exactly the wrong conclusion. The `turn-stall` path should be able to say "nothing was delivered" rather than handing back whatever is on screen. ### Revised proposal 1. On a readiness-grace expiry, surface the `Injector`'s message to the ticket. Do not scrape. 2. Optionally refuse at accept time when `deliverableTo` is already false — still useful, now secondary. 3. Boot-time presence re-derivation stays a design question, unchanged from the body above. I have retitled this ticket, because the old title carried the wrong 30-minute number.
ltms changed title from A daemon restart de-enrols every hand-opened pane, and a send to one is then accepted and queued for 30 minutes instead of refused to A restart de-enrols every hand-opened pane, and the failed send hands the caller a pane scrape instead of the daemon's own reason 2026-10-05 10:01:11 +02:00
Author
Owner

Second reproduction today, this time with the cause known in advance. I restarted the daemon myself at 11:19:29 (the #761 redeploy), so the de-enrolment is not a guess here.

The trinotes observer sent twice to the anki pane (term_65d106559db144) and saw nothing arrive. The log explains both sends, and they fail for two different reasons:

11:19:11.338 DEBUG MessageService - async send task-e4c8c2-8 -> term_65d106559db144
   ... daemon unloaded at ~11:19:19 (drain complete: released=0 abandoned=0)
11:20:10.959 DEBUG MessageService - async send task-2c8e2b-1 -> term_65d106559db144
11:21:36.943 WARN  Injector - readiness grace for term_65d106559db144 expired after 240 polls
              (configured=240 polls/60s elapsed=85642ms): target never became deliverable,
              so failing 1 queued message(s) that never reached its pane
11:21:37.689 WARN  CompletionResolver - failing send to term_65d106559db144 via turn-stall
              fallback: ▗ ▗   ▖ ▖  Claude Code v2.1.289
  • task-e4c8c2-8 was accepted 8 seconds before I unloaded the old daemon. It died in the restart.
  • task-2c8e2b-1 was accepted by the new daemon, 41 seconds after it started, against a pane the restart had just de-enrolled. That is this ticket exactly.

Two numbers worth adding to the record. The grace took 85,642 ms over 240 polls against a configured 60 s — close to the 94,020 ms / 240 polls measured on 2026-10-05 morning, so the overshoot is reproducible and is not a one-off.

And the scrape is worse than "unhelpful". The turn-stall fallback returned ▗ ▗ ▖ ▖ Claude Code v2.1.289 — the Claude Code splash screen. The pane had been /cleared, so there was no earlier reply to mistake for a new one this time. But that is luck: the hazard in this ticket is that the body of a failed ticket is a screen scrape, and on a pane with history it can read as a successful answer.

One more thing this run shows: the sender gets no warning at all. The observer was told accepted — task delegated, and the only record of the failure is in a log it cannot read. It also could not read the ticket (see the separate observer/ticket ticket), so from its side both sends simply vanished. That makes the accept-time refusal this ticket asks for more valuable, not less: there is currently no channel by which the sender learns anything.

**Second reproduction today, this time with the cause known in advance.** I restarted the daemon myself at 11:19:29 (the #761 redeploy), so the de-enrolment is not a guess here. The `trinotes` observer sent twice to the `anki` pane (`term_65d106559db144`) and saw nothing arrive. The log explains both sends, and they fail for two *different* reasons: ``` 11:19:11.338 DEBUG MessageService - async send task-e4c8c2-8 -> term_65d106559db144 ... daemon unloaded at ~11:19:19 (drain complete: released=0 abandoned=0) 11:20:10.959 DEBUG MessageService - async send task-2c8e2b-1 -> term_65d106559db144 11:21:36.943 WARN Injector - readiness grace for term_65d106559db144 expired after 240 polls (configured=240 polls/60s elapsed=85642ms): target never became deliverable, so failing 1 queued message(s) that never reached its pane 11:21:37.689 WARN CompletionResolver - failing send to term_65d106559db144 via turn-stall fallback: ▗ ▗ ▖ ▖ Claude Code v2.1.289 ``` - `task-e4c8c2-8` was accepted **8 seconds before** I unloaded the old daemon. It died in the restart. - `task-2c8e2b-1` was accepted by the **new** daemon, 41 seconds after it started, against a pane the restart had just de-enrolled. That is this ticket exactly. **Two numbers worth adding to the record.** The grace took **85,642 ms over 240 polls** against a configured 60 s — close to the 94,020 ms / 240 polls measured on 2026-10-05 morning, so the overshoot is reproducible and is not a one-off. **And the scrape is worse than "unhelpful".** The turn-stall fallback returned `▗ ▗ ▖ ▖ Claude Code v2.1.289` — the Claude Code **splash screen**. The pane had been `/clear`ed, so there was no earlier reply to mistake for a new one this time. But that is luck: the hazard in this ticket is that the body of a `failed` ticket is a screen scrape, and on a pane with history it can read as a successful answer. **One more thing this run shows: the sender gets no warning at all.** The observer was told `accepted — task delegated`, and the only record of the failure is in a log it cannot read. It also could not read the ticket (see the separate observer/ticket ticket), so from its side both sends simply vanished. That makes the accept-time refusal this ticket asks for more valuable, not less: there is currently no channel by which the sender learns anything.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#757