fleet mod: cross-account delivery through fleetd (fleet_inbox pull path) #788

Open
opened 2026-10-06 05:08:47 +02:00 by ltms · 1 comment
Owner

Goal

Claude leads on different Claude accounts on this host (ltms and work config dirs) must reach each other: lead ↔ lead and lead ↔ a hand-opened peer session. opencode is out of scope (separate ticket).

What I measured on 2026-10-06 (mac, Claude Code 2.1.290)

Channel Crosses accounts? How I checked
ListAgents / SendMessage / $.session.send no ltms registry listed 4 sessions, work listed 7, no overlap; $.session.send is documented as "the same delivery the SendMessage tool makes"
mod $.store no one /fleet-peers run per account; each saw only itself. The store is <CLAUDE_CONFIG_DIR>/plugins/store/fleet_inline-cf835e0ef44b.json: the same file name under ltms and work, but two inodes (320151161, 320151280) with different rows. The reference page says "every session on the machine shares" it; that is true only inside one config dir.
mod $.http.fetch → fleetd yes /fleet-whoami (MCP initialize → tools/call fleet_whoami over $.http.fetch) returned {"role":"primary","leader":"opus","sessionId":"term_65d106559b02e1"} under both CLAUDE_CONFIG_DIR=…/ltms and …/work. Caveat: both claude -p runs were children of the lead pane, so both resolved as that pane. This shows fleetd names a caller by process ancestry, not by account. It does not show a second pane resolving as itself (that is the existing MCP path for every pane).

So fleetd is the account-blind hop. The mod already exists in plugin/hooks/register.js (commands /fleet-peers, /fleet-mail, /fleet-whoami; claude plugin test plugin → 10 pass, 0 fail).

What is missing

fleetd only pushes: Injector → herdr → PTY paste. There is no pull endpoint, so a mod cannot collect its mail. Needed:

  1. fleet_inbox MCP tool. Returns, and removes, the messages queued for the caller's own pane. No target argument (invariant 3). Any role may call it for its own pane only.
  2. Injector route. A pane that called fleet_inbox in the last N seconds is mod-served: the Injector queues to that pane's inbox instead of pasting. $.prompt.submit waits for idle by itself, so the box gate (#782) and status gate do not apply to that route.
  3. Mod side. Poll fleet_inbox every few seconds, deliver each message with $.prompt.submit.

Side effect: every poll is an MCP request, so it re-enrols the pane in MemberPresence. A mod-served pane therefore stays deliverable across a daemon restart (#757), without any other change.

Open questions for the implementer

  • What N is safe for "recently polled", and what happens to a message queued for a pane whose mod stops polling (fall back to PTY paste, or fail loudly)?
  • Ticket completion: a message delivered through the inbox must still let fleet_poll report the same states as a pasted one.

Docs (lead does these)

New tool ⇒ the intent→tool table in CLAUDE.md and the wiki template; a wiki/11-Features.md entry.

## Goal Claude leads on different Claude accounts on this host (`ltms` and `work` config dirs) must reach each other: lead ↔ lead and lead ↔ a hand-opened peer session. opencode is out of scope (separate ticket). ## What I measured on 2026-10-06 (mac, Claude Code 2.1.290) | Channel | Crosses accounts? | How I checked | |---|---|---| | `ListAgents` / `SendMessage` / `$.session.send` | **no** | ltms registry listed 4 sessions, work listed 7, no overlap; `$.session.send` is documented as "the same delivery the SendMessage tool makes" | | mod `$.store` | **no** | one `/fleet-peers` run per account; each saw only itself. The store is `<CLAUDE_CONFIG_DIR>/plugins/store/fleet_inline-cf835e0ef44b.json`: the same file name under `ltms` and `work`, but two inodes (320151161, 320151280) with different rows. The reference page says "every session on the machine shares" it; that is true only inside one config dir. | | mod `$.http.fetch` → fleetd | **yes** | `/fleet-whoami` (MCP `initialize` → `tools/call fleet_whoami` over `$.http.fetch`) returned `{"role":"primary","leader":"opus","sessionId":"term_65d106559b02e1"}` under **both** `CLAUDE_CONFIG_DIR=…/ltms` and `…/work`. Caveat: both `claude -p` runs were children of the lead pane, so both resolved as that pane. This shows fleetd names a caller by process ancestry, not by account. It does not show a second pane resolving as itself (that is the existing MCP path for every pane). | So fleetd is the account-blind hop. The mod already exists in `plugin/hooks/register.js` (commands `/fleet-peers`, `/fleet-mail`, `/fleet-whoami`; `claude plugin test plugin` → 10 pass, 0 fail). ## What is missing fleetd only **pushes**: Injector → herdr → PTY paste. There is no pull endpoint, so a mod cannot collect its mail. Needed: 1. **`fleet_inbox` MCP tool.** Returns, and removes, the messages queued for the **caller's own pane**. No target argument (invariant 3). Any role may call it for its own pane only. 2. **Injector route.** A pane that called `fleet_inbox` in the last N seconds is mod-served: the Injector queues to that pane's inbox instead of pasting. `$.prompt.submit` waits for idle by itself, so the box gate (#782) and status gate do not apply to that route. 3. **Mod side.** Poll `fleet_inbox` every few seconds, deliver each message with `$.prompt.submit`. Side effect: every poll is an MCP request, so it re-enrols the pane in `MemberPresence`. A mod-served pane therefore stays deliverable across a daemon restart (#757), without any other change. ## Open questions for the implementer - What N is safe for "recently polled", and what happens to a message queued for a pane whose mod stops polling (fall back to PTY paste, or fail loudly)? - Ticket completion: a message delivered through the inbox must still let `fleet_poll` report the same states as a pasted one. ## Docs (lead does these) New tool ⇒ the intent→tool table in `CLAUDE.md` and the wiki template; a `wiki/11-Features.md` entry.
Author
Owner

Lead review of PR #789 (ac53550) — fix these on the same branch

I read the whole main-source diff. The design is right: the message stays on the Injector queue until the pane takes it, and stop-polling falls back to the paste route. Four things must change.

1. Injector.cancel — false CANCELLED and double delivery (must fix)

cancel does not look at t.inboxOffer.taken(). A pane's fleet_inbox call runs on the MCP thread. So between the pane's drain and the next Injector poll (up to 250 ms), the head message is taken but still QUEUED.

  • False CANCELLED. The pane drains A, then cancel(A) runs. A is QUEUED and still in t.queue, so cancel removes it and returns CANCELLED. But the pane already has A and will act on it.
  • Double delivery. The pane drains A, then cancel(B) runs for a later message B. Cancel sets t.inboxOffer = null, which throws away the only record that A was taken. A is still the queue head, so the next poll calls paneInbox.offer(target, A) again, and the pane gets A twice.

Fix: in cancel, touch the offer only when p is the offered head (t.queue.peek() == p && t.inboxOffer != null). In that case call withdrawAll first, and then check taken(). If it is taken, the pane has A: do not cancel, and return Cancellation.DELIVERED (the next poll records the delivery). If p is not the head, leave t.inboxOffer alone. Add a test for each scenario, and each needs a positive control.

2. Mod — one leaked MCP session per poll (must fix)

fleetTool sends initialize on every call and never sends DELETE. fleetd's transport is HttpServletStreamableServerTransportProvider from mcp-core 2.0.0. It keeps ConcurrentHashMap<String, McpStreamableServerSession> sessions, and only doDelete removes an entry (I checked with javap -p on ~/.m2/.../mcp-core-2.0.0.jar). I found no other expiry. I have not measured the growth on a live daemon. At one poll every 3 s, that is 20 sessions per minute per pane.

Fix: keep the MCP session id in a module-level variable and reuse it. Re-initialize only when a call fails because the session is unknown (check what status fleetd returns for an unknown Mcp-Session-Id, do not guess). Identity is still safe: fleetd resolves the caller per request from the TCP connection, not from the session. Add a test showing that two polls send one initialize, and a test showing re-initialize after an unknown-session answer.

3. MessageService.collectInbox sits between reply's Javadoc and reply (must fix)

The new method and its Javadoc were inserted after the /** … @return which of the three ways … */ block of reply. So reply lost its Javadoc, and that block is now attached to nothing. Move collectInbox above that block.

4. A wrong comment in the mod (must fix)

The catch comment says "Throwing here would kill the timer". The mods docs say the opposite: "If the callback throws, the error goes to the debug log and the timer runs again at the next interval." Keep the catch, but say what it really prevents (a debug-log error every 3 s on a host with no fleetd).

Accepted as reported

  • N = 15 s, and fall back to the paste route.
  • The status gate still applies to the inbox route. Agreed, because changing it risks a completion for the previous turn.
  • No per-message sender. Agreed. It is a separate change.
  • INBOX in the observer exemption: correct. It is gated by ownsSession, the same as REPLY/ASK.

Re-run mvn clean install, claude plugin validate plugin and claude plugin test plugin, and report the counts the same way as last time.

## Lead review of PR #789 (ac53550) — fix these on the same branch I read the whole main-source diff. The design is right: the message stays on the Injector queue until the pane takes it, and stop-polling falls back to the paste route. Four things must change. ### 1. `Injector.cancel` — false CANCELLED and double delivery (must fix) `cancel` does not look at `t.inboxOffer.taken()`. A pane's `fleet_inbox` call runs on the MCP thread. So between the pane's `drain` and the next Injector poll (up to 250 ms), the head message is **taken but still `QUEUED`**. - **False CANCELLED.** The pane drains A, then `cancel(A)` runs. A is `QUEUED` and still in `t.queue`, so cancel removes it and returns `CANCELLED`. But the pane already has A and will act on it. - **Double delivery.** The pane drains A, then `cancel(B)` runs for a later message B. Cancel sets `t.inboxOffer = null`, which throws away the only record that A was taken. A is still the queue head, so the next poll calls `paneInbox.offer(target, A)` again, and the pane gets A twice. Fix: in `cancel`, touch the offer **only when `p` is the offered head** (`t.queue.peek() == p && t.inboxOffer != null`). In that case call `withdrawAll` first, and then check `taken()`. If it is taken, the pane has A: do not cancel, and return `Cancellation.DELIVERED` (the next poll records the delivery). If `p` is not the head, leave `t.inboxOffer` alone. Add a test for each scenario, and each needs a positive control. ### 2. Mod — one leaked MCP session per poll (must fix) `fleetTool` sends `initialize` on every call and never sends DELETE. fleetd's transport is `HttpServletStreamableServerTransportProvider` from `mcp-core` 2.0.0. It keeps `ConcurrentHashMap<String, McpStreamableServerSession> sessions`, and only `doDelete` removes an entry (I checked with `javap -p` on `~/.m2/.../mcp-core-2.0.0.jar`). I found no other expiry. I have not measured the growth on a live daemon. At one poll every 3 s, that is 20 sessions per minute per pane. Fix: keep the MCP session id in a module-level variable and reuse it. Re-initialize only when a call fails because the session is unknown (check what status fleetd returns for an unknown `Mcp-Session-Id`, do not guess). Identity is still safe: fleetd resolves the caller per request from the TCP connection, not from the session. Add a test showing that two polls send one `initialize`, and a test showing re-initialize after an unknown-session answer. ### 3. `MessageService.collectInbox` sits between `reply`'s Javadoc and `reply` (must fix) The new method and its Javadoc were inserted after the `/** … @return which of the three ways … */` block of `reply`. So `reply` lost its Javadoc, and that block is now attached to nothing. Move `collectInbox` above that block. ### 4. A wrong comment in the mod (must fix) The `catch` comment says "Throwing here would kill the timer". The mods docs say the opposite: "If the callback throws, the error goes to the debug log and the timer runs again at the next interval." Keep the `catch`, but say what it really prevents (a debug-log error every 3 s on a host with no fleetd). ### Accepted as reported - N = 15 s, and fall back to the paste route. - The status gate still applies to the inbox route. Agreed, because changing it risks a completion for the previous turn. - No per-message sender. Agreed. It is a separate change. - `INBOX` in the observer exemption: correct. It is gated by `ownsSession`, the same as REPLY/ASK. Re-run `mvn clean install`, `claude plugin validate plugin` and `claude plugin test plugin`, and report the counts the same way as last time.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#788