Compare commits

..

55 Commits

Author SHA1 Message Date
Dai Ha a144dbe0c2 fleetd #809: a lead rollover waits for the successor's MCP connect before it types bootstrapText
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m10s
CI / build (push) Failing after 2m27s
A roll killed the old lead, launched a fresh pane, typed bootstrapText into it
and logged `rolled`. The successor never received the text: herdr accepted the
keystrokes while that pane's Claude Code was still booting, and it connected the
bridge MCP 583ms after the roll reported success. The successor woke with no
handover file and no reason to suspect anything.

The daemon already holds the right signal. MemberPresence is populated from a
Claude's MCP initialize, and Injector already holds a spawned member's first
delivery until the terminal is present there. LeadRollover was the one delivery
path that did not read it: it gated on herdr's agent_status, which reports idle
through the whole boot window.

waitUntilPaneReady now requires both signals — a real turn boundary AND
mcpPresent — and its RELAUNCH_NEVER_READY detail names which of the two was
missing, so the same loss cannot be read as a timeout of the other.

After the send, waitUntilTurnStarted polls for a live turn (WORKING or BLOCKED),
bounded by BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS = 5 rather than
relaunchReadySeconds: that budget sizes a pane booting a CLI, this one an agent
reacting to text it already holds. A miss reports BOOTSTRAP_NOT_CONFIRMED and
re-sends nothing, because the status is polled and a turn shorter than one poll
interval reads as unconfirmed — a second paste into a session that did receive
the first is worse than one unconfirmed roll.

FakeHerdr gains turnStartsOnPrompt(): after each agent.prompt it reports
"working" for the next two agent.get reads, then falls back to its steady state.
Two reads, because one AgentControl.status() costs two — resolveTarget issues its
own agent.get first. Bounded rather than sticky, so a test may roll more than
once; omitting it models herdr accepting a send that never starts a turn.

Verified: mvn -o clean install, 2247 tests, 0 failures. Both new gates were
mutation-checked — dropping the mcpPresent conjunct and short-circuiting the
confirmation each fail exactly one dedicated test.
2026-10-07 08:39:04 +02:00
Dai Ha 0ff0d6cb02 fleetd #808: a late fleet_reply claims its turn's waiter so the completion scrape cannot double-publish
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 2m19s
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 46s
CI / build (push) Failing after 2m3s
A send that times out leaves its captured Rendezvous waiter open for the CB-106
completion fallback (strandLateResolution). If the member then calls fleet_reply,
reply() found no live waiter and queued the answer to the inbox, but never touched
that captured waiter — so the eventual completion fallback still resolved it and
published a second, redundant scrape of the same turn.

reply() now looks up the exact waiter strandLateResolution is tracking for the
target (lateWaiters, a new per-target index onto a per-turn CompletableFuture) and
claims it with Rendezvous.resolveLateReply before falling back to the inbox. The
CompletableFuture's own single-winner complete() is the discriminator: whichever
side resolves it first decides what the other sees. A later completion scrape
against an already-claimed waiter is a no-op in CompletionResolver.resolve's
existing waiter.isDone() check, so nothing is scraped or published a second time.
A turn that never gets an explicit reply is unaffected — its own waiter is still
open when the completion fallback fires, exactly as #801 already fixed.
2026-10-07 07:02:42 +02:00
Dai Ha a2cbf50961 Bridge block: a blocking send creates no ticket (#801)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 2m19s
The paragraph that already warns a blocking fleet_send is capped by the caller's own
MCP timeout now also says it creates no ticket, so a timeout leaves no id to poll and
a resend can deliver the same brief twice, and that the one recovery route --
draining the target's inbox -- is DRAIN, so only a lead may take it.

Kept byte-identical with the wiki template (fleetd.wiki f812689).
2026-10-07 06:17:58 +02:00
Dai Ha 11352d0d06 fleetd #801: a blocking send's timeout receipt names a route the caller may use
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 2m6s
A blocking fleet_send creates no ticket -- new Task( is constructed only in
sendAsync -- so its TIMED_OUT_WORKING / TIMED_OUT_QUEUED / BUSY receipt had no id
to hand back and said "retry or poll status". A retry on that path can deliver the
brief twice, and the one thing that does recover the answer went unnamed.

The receipt now carries NO_TICKET_NO_RESEND and names the target's inbox, on both
the MCP and REST surfaces. It names it only when the caller's own DRAIN is granted:
fleet_poll{target} and GET /sessions/{id}/replies both resolve to DRAIN, which is
primary-only, while SEND is granted to an architect, a collaborator and an observer
as well. A refused caller is told the reply cannot be recovered on its channel
instead of being pointed at a call it cannot make.

MessageService.strandLateResolution routes a late turn-completion to the target's
inbox. send's TimeoutException branch never completes its waiter, and
Rendezvous.close only deregisters it from session lookups, so the CB-106 fallback
can still resolve that exact future after the caller gave up -- previously into
nothing. Only Kind.COMPLETION is routed: an explicit fleet_reply reaches the inbox
through reply()'s own lookup once close() has run, so routing it here as well would
publish it twice. The publish is wrapped, because a throw inside a whenComplete
action is captured by the discarded dependent stage and reaches no caller.

Verified on the merge result, not the branch: mvn clean install, BUILD SUCCESS,
Tests run: 2244, Failures: 0, Errors: 0, Skipped: 0 -- tallied from 183 surefire
XML files, and the branch adds exactly 8 @Test methods. Reviewed in two rounds:
the first receipt named fleet_poll{target} unconditionally and the late publish was
unguarded; the second added send/answer overloads defaulting that grant to true
with no production caller, which c8e31e1 removed.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 06:16:21 +02:00
Dai Ha c8e31e1195 fleetd #801: drop the two mayDrainPoll-defaulting send/answer overloads
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m57s
They had no production caller — only test call sites — and the default
pointed the unsafe way: true means naming fleet_poll{target}, which is
the exact receipt this ticket is fixing. Keep one signature each and
pass mayDrainPoll explicitly at every test call site instead.
2026-10-07 06:12:57 +02:00
Dai Ha 24782fe329 fleetd #801: gate the timeout receipt's recovery route on the caller's own DRAIN grant
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 2m10s
A TIMED_OUT/BUSY receipt named fleet_poll{target}/GET /sessions/{id}/replies
unconditionally, but that route needs Authz.Action.DRAIN, which is primary-only,
while SEND and ANSWER are also granted to an architect, a collaborator, and an
observer. Those callers were told to run a call they are refused, and the late
reply they were pointed at is published somewhere they can never read.

Thread each caller's own DRAIN grant into formatReply (MCP) and writeReply
(REST); name the route only when the grant holds, otherwise say the reply
cannot be recovered on that channel.

Also wrap strandLateResolution's whenComplete callback in try/catch: an
exception thrown there lands in the discarded dependent stage and is never
rethrown, so an unguarded inbox.publish failure would lose a late reply with
no log line at all.
2026-10-07 06:05:28 +02:00
Dai Ha a49d4dcce9 Bridge block: a collaborator may read its own ticket (#804)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 2m9s
Invariant 3 said a collaborator is still refused every ticket, including its
own, and the Collaborator section said you cannot read a ticket at all. Both are
false after #804: a collaborator may poll a ticket its own fleet_send{wait:false}
created, and no other.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 05:40:04 +02:00
Dai Ha 2aacf07841 fleetd #803/#804: a null ticket is refused, and a collaborator may read its own
#803: MessageService.ownsTicket(String, String) guards a null ticket instead of
handing it to ConcurrentHashMap.get, which rejects a null key. fleet_poll{} with
no ticket resolves to TASK_READ with a null id, so the crash sat on the path a
caller reaches by calling the tool wrong.

#804: TASK_READ now grants a collaborator a ticket its own fleet_send{wait:false}
created, on the same ownership classifier the observer grant uses. A collaborator
could already create a ticket -- FleetMcp.sendAsync records any caller as the
creator with no role filter -- so the reply had an owner that was refused it.

Renames NO_OBSERVER_OWNED_TICKET to NO_OWNED_TICKET and observerOwnsTicket to
ticketOwnedByCaller, since neither is observer-specific any more.

The first version also threaded a mayTaskReadNudge boolean through FleetMcp,
PrimaryRegistry and ReplyPushLoop to suppress a ticket nudge to a role Authz
would refuse. With the collaborator granted, that condition is true for every
role reaching the call site, so 820e6f2 removes it.

Verified on the merge result, not the branch: mvn clean install, BUILD SUCCESS,
Tests run: 2236, Failures: 0, Errors: 0, Skipped: 0 -- tallied again from 183
surefire XML files.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 05:39:00 +02:00
Dai Ha f2e025cb13 fleetd #801: name the inbox route in a blocking-send timeout, and stop dropping a late completion
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m10s
The TIMED_OUT_WORKING/TIMED_OUT_QUEUED/BUSY receipt (MCP and REST) named no
recovery route and invited a bare retry, which can duplicate the delivery. It
now names fleet_poll{target=...} (or GET /sessions/{id}/replies) and states
that a blocking send is never tracked by a ticket.

A blocking send's TimeoutException path leaves its rendezvous waiter in place:
Rendezvous.close only deregisters it from future session lookups, while the
turn-completion fallback resolves it through a captured future reference and
can still succeed after the caller gave up. With nobody left listening, that
late completion is now published to the target's inbox, the same route an
unstructured fleet_reply already uses when it arrives with no caller waiting.
2026-10-07 05:37:55 +02:00
Dai Ha 820e6f2620 fleetd #803/#804: remove dead mayTaskReadNudge plumbing
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 2m6s
The #804 grant gave a collaborator TASK_READ for its own ticket, which
dissolved the only case the mayTaskReadNudge gate existed for. With the
real ownership classifier forced to t -> true at its one call site, the
grant reduced to "every non-anonymous role", so the filter could never
suppress a ticket nudge in production. Revert ReplyPushLoop, PrimaryRegistry
and FleetMcp to their pre-plumbing shape, and drop the two tests that only
proved the dead seam, not that any caller could reach it.
2026-10-07 05:36:13 +02:00
Dai Ha 72cc7560f1 fleetd #803/#804: fix null-ticket NPE in ownsTicket, grant collaborator TASK_READ
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 2m11s
#803: MessageService.ownsTicket(String, String) called tasks.get(null) when
fleet_poll{} supplied no ticket, throwing NPE instead of refusing. Guard the
lookup so a null ticket returns false, same as an unknown one.

#804: a collaborator could create a ticket via fleet_send to another
collaborator but could never read it back, since TASK_READ's grant only
covered primary/worker/architect/observer. Extend the grant to a collaborator
confined by the same ownership classifier observer already uses, rename the
classifier parameter now that it serves both roles, and gate the ticket-nudge
in ReplyPushLoop the same way #778 gated the reply/question nudges so no
caller is nudged toward a fleet_poll its role would refuse.
2026-10-07 05:26:34 +02:00
Dai Ha 50c7674b98 Bridge block: invariant 4 — the prompt-box gate is off by default (#797)
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m59s
Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 05:20:03 +02:00
Dai Ha fcce44386b fleetd #797: make the prompt-box gate configurable, default off
PromptBox.clearToSubmit returns true at once when the gate is disabled,
before any pane read, hold log or streak update, so every call site behaves
as if the box were empty whatever its own polarity. promptBoxGateEnabled is
a new top-level config key; null and false both leave the gate off.

The gate held deliveries on the pane's own autocomplete suggestion, which
detection reads as the operator's unsubmitted typing. Detection is unchanged
and still wrong (#802); turn the gate back on once that is fixed.

Merge of PR #805 (branch worker/797-disable-box-gate-5b5478-12).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 05:18:43 +02:00
Dai Ha f3f067f46f fleetd #797: add a config switch to disable the prompt-box delivery gate
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 2m13s
PromptBox currently misreads a pane's own autocomplete suggestion at the
caret as the operator's unsubmitted typing, holding deliveries that were
never actually at risk. Add promptBoxGateEnabled (default off) so the gate
can be switched off without deleting its classification logic — detection
stays fixable separately (fleetd #802).

When disabled, PromptBox.clearToSubmit returns true for every target
without reading the pane, logging a hold, or counting a streak. The three
call sites (ReplyPushLoop, LeadCoordLoop, LeadHeartbeatLoop) each gained a
boolean constructor parameter threaded straight into their PromptBox,
with the previous constructor demoted to a delegator passing true so
every existing caller keeps the gate on by default.
2026-10-07 05:13:50 +02:00
Dai Ha 9f4b736fb9 Bridge block: an observer may read a ticket it created (#778)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m59s
Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 05:10:58 +02:00
Dai Ha 52cd0470b6 fleetd #778: an observer may read a ticket it created
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m59s
Authz grants TASK_READ to an observer only for a ticket whose creating
terminal is its own, checked by a classifier the poll call sites thread in.
ReplyPushLoop no longer nudges a recipient to run a DRAIN or ANSWER its own
role would be refused.

Merge of PR #795 (branch worker/778-e301b0-1).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 05:05:06 +02:00
Dai Ha a3eeace447 handover skill: BOOTSTRAP_NEVER_SENT, and relaunchReadySeconds bounds three waits
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m55s
Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-06 19:28:21 +02:00
Dai Ha b696c31756 fleetd #796: retry bootstrap delivery to a fresh lead pane
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m5s
CI / build (push) Failing after 2m17s
A roll killed the old pane, launched the new one, then lost bootstrapText to a
herdr agent_not_ready refusal. The successor woke with no handover and no way to
tell a failed roll from a cold start.

sendBootstrapWithRetry retries only that refusal, bounded by
relaunchReadySeconds. A persistent refusal now ends the roll with the new
BOOTSTRAP_NEVER_SENT outcome, so fleet_handover{action:"status"} can report it.

Verified at 027d413 in a throwaway worktree: 2210 tests, 0 failures, 0 errors,
clean install.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-06 19:27:36 +02:00
Dai Ha 96ebae28a6 reviewer skill: send fleet_reply before writing the reasoning out
Five reviewer turns were lost in one session. Each wrote a long, correct
analysis to its own pane and ended the turn with no fleet_reply, so the
bridge scraped the pane and the lead received a clipped fragment.

The briefs are half the cause: each carried a five-item checklist of things
to hunt alongside a ~90-word capped output format, which reads as two
contradictory output contracts. The skill now says the checklist is where to
look, not the shape of the answer, and that the reply goes out as soon as
the answer is known.

Measured: tickets task-7da785-29 and -30 resolved at 18:28 via the
turn-completion fallback, 1233 and 1403 chars scraped.
2026-10-06 19:01:38 +02:00
Dai Ha b41aa663f7 Bridge block: invariant 6 — an operator outranks the confirm rule
Invariant 6 as first written told every receiver to confirm, with no
mention of who may override it. A peer cannot impose a rule on another
operator's session.

Measured: the trinotes pane (observer, work config dir) answered the
communication test and then said its operator's standing rule is not to
answer fleet messages, that a peer cannot change that rule, and that it
would ask its operator before following mine. That is the correct reading
and the rule now says so.

Also records the symmetric error: a sender must not read silence as
agreement or as a dead session.

wiki/7-Use-Cases.md synced byte-identical.
2026-10-06 18:48:12 +02:00
Dai Ha 027d413ce9 fleetd #796: document bootstrap retry readiness bound
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m59s
2026-10-06 18:26:01 +02:00
Dai Ha f544906621 Bridge block: a receiver confirms what it receives
Adds invariant 6 to the canonical block. Nothing in it said a received
message must be answered, so a sender could not tell a handled message
from one that never arrived.

Measured today: the anki pane (observer) sent to vms, which reads
deliverable:false because it has not contacted the daemon since the 08:05
restart. The send was accepted, held 83s, and failed with the body 'nst' --
three characters scraped off the vms screen. The sender read that as an
inconclusive result and asked the lead what happened. fleetd #757 covers
the accept-time refusal; this rule covers the half the protocol owns.

Operator's words: 'when a message sent, at least the receiver should
confirm unless explicit told to not reply'.

wiki/7-Use-Cases.md template synced byte-identical (wiki 3ad606f).
2026-10-06 18:18:10 +02:00
Dai Ha c88f01ecb8 fleetd #796: retry bootstrap delivery after agent readiness refusal
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 2m4s
2026-10-06 18:16:47 +02:00
Dai Ha 6022aef612 fleetd #778: let an observer read its own ticket, never nudge a role to run a call it is refused
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Failing after 2m15s
Part 1: Authz.TASK_READ now grants an observer the ticket it created (a new
observerOwnsTicket classifier), via MessageService.ownsTicket(String, String)
(side-effect-free) and a new Authz.permits overload. fleet_status and
fleet_poll{target} (DRAIN) stay refused for an observer. REST and MCP share
the same gate.

Part 2: ReplyPushLoop no longer nudges a reply/question delegator to run
fleet_poll(target=...) or fleet_send(turnId=...) unless Authz would actually
grant that call to its role. FleetMcp.sendHandler asks Authz.permits(DRAIN)/
permits(ANSWER) once, eagerly, while the real Principal is in scope, and
threads the booleans through PrimaryRegistry.recordDelegation into the new
Delegation fields and nudgeMayDrainFor/nudgeMayAnswerFor accessors --
plain booleans, not a Role reference, so no new or widened edge crosses the
frozen auth<->mcp<->msg package-cycle baseline. ReplyPushLoop filters a
reply/question out of its pending set when the grant is false, falling back
to the durable inbox / the ask window lapsing, same as when no nudge target
is known at all. Log lines that called a non-lead nudge target "lead" now
say "nudge target".
2026-10-06 13:55:26 +02:00
Dai Ha 66ab9cab24 fleetd #791: addendum — plugin 0.3.0 has no mount and is installed in four instances
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m59s
2026-10-06 08:03:50 +02:00
Dai Ha f611488c9b fleetd #791: README — the mod still reads FLEETD_MCP_URL as an optional override
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 56s
CI / build (push) Failing after 2m4s
2026-10-06 08:03:06 +02:00
Dai Ha 44d16e7639 fleetd #791: merge PR #794 — fleet plugin 0.3.0, no MCP mount, mod off for members
The plugin no longer carries .mcp.json; the instance or the project mounts
fleet. The mod asks fleet_whoami once and skips the fleet_inbox poll for a
worker or an architect, so members keep paste delivery. The mod reads
FLEETD_MCP_URL with $.env.get (documented in the mods API), defaulting to
http://127.0.0.1:8765/mcp.

Lead verification on 87f7c03: claude plugin validate plugin passed;
claude plugin test plugin 21 pass, 0 fail. With the role gate removed, the
worker and architect tests fail (19 pass, 2 fail), so they are not vacuous.
2026-10-06 08:03:01 +02:00
Dai Ha 87f7c03535 fleetd #791: fleet plugin 0.3.0 — drop the MCP mount, gate the mod off for members
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 2m7s
- Delete plugin/.mcp.json; mounting fleet is now the instance's or the
  project's job, not the plugin's. The setup skill still writes the
  project .mcp.json entry.
- register.js: read FLEETD_MCP_URL via $.env.get, defaulting to
  http://127.0.0.1:8765/mcp, documented at
  https://code.claude.com/docs/en/plugins/mods/api.md ("$.env — get
  and set environment variables").
- Gate the fleet_inbox poll by role: a worker or architect learns its
  role once from fleet_whoami and then never polls fleet_inbox; every
  other role polls as before. An unreachable daemon retries whoami on
  a later tick rather than deciding.
- Bump 0.2.0 -> 0.3.0 in plugin.json and marketplace.json, and update
  README/SKILL.md to drop every claim that the plugin mounts fleetd.
- Add 6 tests for the role gating; all 21 tests and
  `claude plugin validate` pass.
2026-10-06 06:46:26 +02:00
Dai Ha 562354a58d fleetd #790: document observer -> lead send in the bridge block
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m59s
Invariant 3, the observer paragraph and the unconfigured-pane row now say an
observer may send to a lead. Invariant 4 now says a direct send to a lead pane
is not box-gated yet (#793). Wiki template synced (check printed in sync: True).
2026-10-06 06:25:31 +02:00
Dai Ha eaa1f0d170 fleetd #790: merge PR #792 — an observer pane may fleet_send to a lead
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 56s
CI / build (push) Failing after 2m5s
An observer's SEND classifier (CallerResolver.observerSendTarget, renamed from
sendableObserverTarget) now accepts a configured lead terminal as well as
another observer pane. Collaborators, architect slots and spawned members stay
refused. The [fleet_send from observer term_…] prefix is kept. An observer
learns a lead's sessionId from its filtered fleet_list panes rows (role "lead").

Lead verification: mvn clean install on 72ea7de in a throwaway worktree,
exit 0, surefire totals 2208 run / 0 failures / 0 errors / 0 skipped.

Follow-up found in review: #793 (the Injector has no prompt-box gate for a
direct send to a lead pane).
2026-10-06 06:24:33 +02:00
Dai Ha 72ea7def0a fleetd #790: let an observer pane fleet_send to a lead
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 2m7s
An observer could only reach another observer pane, so a hand-opened tab
could answer a lead with fleet_reply but never start a conversation with
one.

CallerResolver.sendableObserverTarget() is renamed observerSendTarget()
and now accepts a configured lead terminal as well as a pane that falls
to the observer floor. It reads the same maps resolve() reads, in the
same order: a live spawned member is refused first, a lead is then
accepted even when the same pane is also bound to an architect slot,
and a collaborator or an architect-slot pane is refused. A collaborator,
an architect and a spawned member stay unreachable.

Authz keeps its one SEND case; only the classifier it is handed widened,
so the MCP gate (FleetMcp.denyFor) and the REST gate (FleetApp.allow)
give the same answer from the same predicate.

fleet_list shows the lead's address in the observer's filtered `panes`
rows rather than by adding the `leads` array: the panes filter already
reads the same classifier as the SEND gate, so reachability and
visibility cannot drift, and the reduced row carries no lead name,
context window or config dir.

Gates traced end to end for an observer -> lead send: denyFor, the
sendTool argument validation (profileTargetError rejects profile names
only), MessageService (records the caller's owner key, no role gate),
the injector's readiness gate (Fleetd.deliverableTo accepts a lead
through the leads map, never a presence entry) and its status gate.
Unchanged: an observer still holds no TASK_READ, so it cannot poll the
ticket a wait:false send returns.

Tests: observer -> lead allowed and observer -> collaborator / architect
slot / spawned member refused over denyFor, each with a positive control
in the same test; the same matrix over the REST route; a real MCP client
through the real MessageService and Injector to the real herdr call,
proving the [fleet_send from observer term_...] prefix reaches a lead's
pane and that the lead passes the production readiness gate; and the
observer's fleet_list panes rows now carrying the lead.

mvn clean install in fleetd/: Tests run: 2208, Failures: 0, Errors: 0,
Skipped: 0 — BUILD SUCCESS. Reverting the widening in CallerResolver
alone turns 7 of the new assertions red across all five touched test
classes, so none of them passes vacuously.
2026-10-06 06:20:43 +02:00
Dai Ha 98f3cd5aa7 fleetd #788: document fleet_inbox in the bridge block and the addendum
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m57s
Block: observer INBOX right, invariant 3 (inbox is only-as-itself),
invariant 4 (a pane that collects its own mail skips the paste and the
box gate, not the status gate), and an intent-table row. The wiki
template carries the same block; the sync check printed in sync: True.
2026-10-06 05:54:13 +02:00
Dai Ha aec5ff2c2f fleetd #788: merge PR #789 — fleet_inbox, so a pane can collect its own mail
A pane that calls fleet_inbox within 15 s is mod-served: the Injector
offers its next message for collection instead of pasting it, and keeps
it queued until the pane takes it. A pane that stops polling gets its
mail pasted again. The fleet mod polls fleet_inbox every 3 s on one
reused MCP session and hands each message to Claude with
$.prompt.submit.

Lead proof at d9fedfb, in a throwaway worktree:
mvn clean install exit 0; surefire XML 182 files, 2202 tests,
0 failures, 0 errors, 0 skipped. claude plugin validate plugin: passed.
claude plugin test plugin: 15 pass, 0 fail.
2026-10-06 05:53:16 +02:00
Dai Ha d9fedfbc5a fleetd #788: review fixes — cancel vs a collected offer, one MCP session
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m59s
Four items from the lead review of PR #789.

1. Injector.cancel touched the inbox offer unconditionally. A pane takes an
   offer on the MCP thread, so between its drain and the next poll the head
   entry is taken but still QUEUED. Cancelling it returned CANCELLED for text
   the pane already held, and cancelling a LATER entry nulled t.inboxOffer and
   so lost the only record that the head was taken, which made the next poll
   offer it a second time. cancel now touches the offer only when the entry is
   the offered head, withdraws first, and answers DELIVERED when the pane took
   it.

2. The mod sent initialize on every fleetTool call and never a DELETE, so
   fleetd's transport kept one session per call -- 20 a minute per pane at a
   3s poll. The mod now holds one MCP session and reopens it only on the 404
   "Session not found" fleetd answers for an id it no longer holds (measured
   against the live daemon, not assumed).

3. MessageService.collectInbox had been inserted between reply's javadoc and
   reply, leaving that block attached to nothing. Moved above it.

4. The timer's catch comment said a throw would kill the timer. It does not;
   the catch keeps every tick from writing an error to the debug log.

mvn clean install: Tests run: 2202, Failures: 0, Errors: 0, Skipped: 0,
BUILD SUCCESS. Counted again over 182 target/surefire-reports/TEST-*.xml:
2202/0/0/0. claude plugin validate plugin and claude plugin test plugin both
exit 0, 15 pass 0 fail.

Three mutation checks, each reverted:
- the exact pre-review cancel body: the two new cancel tests fail with
  "expected: <DELIVERED> but was: <CANCELLED>" and "expected: <true> but was:
  <false>", the other six pass;
- initialize on every call: the two session tests fail (1 vs 2 initializes);
- no 404 retry: only the retry test fails (2 vs 1 initializes).
2026-10-06 05:48:34 +02:00
Dai Ha ac535503ff fleetd #788: fleet_inbox, so a pane can collect its own mail
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 2m15s
A Claude Code session running the fleet mod now pulls its messages from
fleetd and submits them with $.prompt.submit, instead of having them typed
into its terminal by herdr. fleetd names a caller by its pane, so this is
the hop that works between two Claude accounts on one host.

New MCP tool fleet_inbox. It takes no arguments: the pane is the caller's
connection-resolved terminal, so no caller can read another pane's mail.
Authz gets an INBOX action, grouped with REPLY and ASK as only-as-itself.

A pane becomes mod-served by calling fleet_inbox, for a 15s window that
each call renews. The Injector asks that question at the moment it is
about to deliver, and offers the message for collection rather than typing
it. The message stays at the head of the Injector's queue until the pane
takes it, so:

  - delivery is recorded when the pane really has the text, not when it
    was offered, and fleet_poll reports the same states as for a typed
    message;
  - a pane that stops polling leaves the window and its mail is typed
    instead, with nothing stranded and nothing delivered twice. The offer
    is withdrawn under the same monitor that drains it, so an entry is
    removed exactly once;
  - a collected message gets no Enter nudge. Nothing was typed, and an
    Enter in a lead's pane would submit whatever its operator was writing.

cancel, drop and both grace expiries withdraw the offer too, so a message
the caller was told never arrived can never arrive later.

Mod side: a 3s timer collects through the existing fleetTool helper and
submits each message. A fleetd that is down returns from the timer rather
than throwing out of it, which would stop every later poll.

Build: mvn clean install in fleetd/ — Tests run: 2200, Failures: 0,
Errors: 0, Skipped: 0; BUILD SUCCESS. claude plugin validate plugin and
claude plugin test plugin both pass (13 pass, 0 fail).

Two mutations confirm the new tests bind: forcing isModServed false fails
9 of 15, and removing the stop-polling fallback fails exactly
aPaneThatStopsCollectingHasItsMailTypedInstead.
2026-10-06 05:31:08 +02:00
Dai Ha 654b3b5e14 fleetd #788: fleet mod with presence, mail and a fleetd identity probe
Turns the fleet plugin into a Claude Code mod (hooks/hooks.json +
register.js). /fleet-peers and /fleet-mail carry messages through
$.store; /fleet-whoami calls fleet_whoami on the local fleetd over
$.http.fetch.

Measured 2026-10-06 on Claude Code 2.1.290: $.store lives at
<CLAUDE_CONFIG_DIR>/plugins/store/, so the store mailbox does not cross
the ltms and work accounts (two inodes, different rows). /fleet-whoami
reached fleetd from both config dirs, so fleetd is the cross-account hop.
The fleetd pull path that finishes the adapter is #788.

claude plugin validate plugin: passed. claude plugin test plugin:
10 pass, 0 fail.
2026-10-06 05:09:04 +02:00
Dai Ha 1fae8b81a0 fleetd #782: tell a placeholder hint apart from real typed text in the prompt box
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m9s
The box gate held every delivery into an idle lead pane. Claude Code
draws a placeholder hint in the empty input box - the pane's own last
submitted prompt - and draws it faint. PromptBox read the pane with
source 'detection', which carries no escape codes at all, so the
dimness was gone before classify ran and the hint read as a draft.

- AgentControl gains readWithStyling, which sends strip_ansi: false.
  The two-argument read is untouched for its other callers.
- PROBE_SOURCE moves to 'visible', the source that does carry styling.
- markerLength skips leading SGR codes, because the grey drawn on an
  empty box's caret sits before the marker glyph.
- boxContent decodes the line into (char, faint) pairs and drops every
  character inside a faint span.
- boxContent also uses Character.isSpaceChar, not only isWhitespace.
  This was cosmetic on 'detection', which trimmed a trailing U+00A0;
  on 'visible' an empty box arrives as '<caret>\u00a0' and would
  otherwise read as DRAFT with 1 character.

Verified in a throwaway worktree at 4ef8156: mvn clean install, BUILD
SUCCESS, 179 test classes, 2184 tests, 0 failures, 0 errors, 0 skipped,
counted from fleetd/target/surefire-reports/*.xml. PromptBoxTest is 21
tests, up from 15.

Measurements behind the four fixtures, taken from seven live panes with
'herdr agent read <paneId> --source visible --ansi': the two held panes
carried SGR 2 around their hint text, the five empty boxes carried no
faint span, and the caret grey on one empty box was a 24-bit colour
sitting before the marker.

Two checks on the move to 'visible' that the fix depends on:

- The active-turn marker 'esc to interrupt' still arrives as one
  contiguous string with styling kept, so the UNREADABLE guard for a
  live turn is intact.
- Across all seven panes the only intensity codes emitted are 2 and 0.
  SGR 22 (normal intensity) never appears, and no compound code mixes
  an intensity with a colour, so exact-matching those two codes is
  enough for this build. A build that emitted 22 would leave faint on
  and could read a drafted box as empty; tracked on #782.

PR #784's HOLD_GIVE_UP_STREAK is deliberately not here. I asked for it
and it was wrong: the hold cadence measures 15.07s, so 1200 holds is
about five hours, and giving up submits whatever sits in the box.
2026-10-06 04:28:02 +02:00
Dai Ha 4ef815682b fleetd #782: tell a placeholder hint apart from real typed text in the prompt box
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 2m9s
PromptBox probed the detection region, which carries no ANSI styling, so a dim
placeholder hint (the pane's last submitted prompt) read as the operator's draft
and held the delivery forever.

Move the probe to the visible region, read with strip_ansi:false
(AgentControl.readWithStyling), skip leading SGR escapes before matching the
box marker, use Character.isSpaceChar so a non-breaking space in an empty box
counts as padding, and exclude any character drawn inside a faint (SGR 2) span
from the box content. An all-faint box now classifies as EMPTY instead of
UNREADABLE.
2026-10-06 04:22:59 +02:00
Dai Ha 6596458ce6 fleetd #780: queued sends no longer lie to fleet_poll or fleet_send
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m2s
Three false receipts on the delivery path to a pane are fixed:

- fleet_poll reported "pending — worker working" for a message still
  sitting in the injector queue. MessageService.pendingDetail now reads
  Injector.queuedWaitMillis and says "queued, not yet delivered".
- fleet_send's accept text carried no warning when the target was not
  injectable. FleetMcp.sendAsync now appends one.
- a queued message had no timeout at all. Injector gains
  QUEUE_WAIT_GRACE_POLLS (4800, ~20 min at the 250ms poll), mirroring
  READINESS_GRACE_POLLS, and fails the queue through onTurnFailed.

Verified in a throwaway worktree at 1cc0322: mvn clean install, BUILD
SUCCESS, 179 test classes, 2179 tests, 0 failures, 0 errors, 0 skipped,
counted from fleetd/target/surefire-reports/*.xml. All four new tests
are present in the XML and passed.

Reviewed by two reviewers. Two findings, both confirmed in the code and
both left open rather than fixed here:

- FleetMcp.notInjectableWarning reads only AgentStatus.injectable(), not
  the Fleetd.deliverableTo presence gate the Injector actually uses, so
  an idle pane that has not connected its bridge MCP still gets a clean
  accept receipt. This is the headline false receipt of #780 and is the
  reason #780 stays open.
- the queueStalled path does not call forget.accept. The proposed fix is
  rejected: a target stuck UNKNOWN may be alive but unclassifiable, and
  forgetting its presence would make it permanently undeliverable, since
  presence is only re-established by an MCP request an idle pane has no
  reason to make. Presence is already cleared on session release
  (SessionManager.java:477).
2026-10-06 04:03:43 +02:00
Dai Ha 1cc0322782 fleetd #780: tell the truth about a queued-but-undelivered send
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m10s
A send to a pane that never frees up was accepted, then polled as
"pending — worker working" forever, with no WARN and no timeout. That
text is only ever correct for a message the injector already handed
off; a message still sitting in its queue now reports as queued, with
the target's live status and how long it has been waiting.

Injector.enqueue() now stamps Pending.enqueuedAtMillis so
queuedWaitMillis(target) can tell "never attempted" apart from
"delivered, now being worked on". MessageService.pendingDetail() uses
it to pick the poll wording. FleetMcp.sendAsync() adds a best-effort
warning to the accept text when the target is not injectable at send
time, per the brief's stated preference for a warning over a hard
refusal (a short busy spell is normal).

Injector gets a new bound, QUEUE_WAIT_GRACE_POLLS (4800 polls, ~20min
at the 250ms poll interval), mirroring READINESS_GRACE_POLLS: a
message that sits at the head of the queue with no delivery attempt
for that long fails via onTurnFailed, the same path a readiness
timeout uses, without touching presence — the target is busy, not
gone.

Fixed InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants's
stub clock, which now sees one extra, legitimate nowMillis read from
enqueue()'s new stamp.

Tests: aQueuedButNeverInjectedMessageDoesNotPollAsWorking +
aDeliveredMessageStillPollsAsWorkerWorking (poll wording, with
positive control); failsAQueuedMessageWhoseTargetNeverFreesUp +
aTargetThatFreesUpBeforeTheQueueWaitGraceIsDeliveredNormally (timeout,
with positive control).
2026-10-05 19:54:37 +02:00
Dai Ha c043d149cf fleetd #775: name the trigger flag for what it now reads
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 1m57s
The flag tested leader.tab() and was named for it. It now tests whether
fleet.leaders has any entry at all, so the old name states a condition the
code no longer checks.
2026-10-05 14:21:19 +02:00
Dai Ha fc786d0f67 fleetd #775: say which remedy applies to a lead vs a collaborator
The ticket's correction comment pointed out the refusal message still
implied removing a lead's tab helps, when only placement: tab does.
Restate the message so the lead and collaborator remedies are not
conflated, and keep the variable name the correction specified.
2026-10-05 14:19:17 +02:00
Dai Ha b314cb4d51 fleetd #775: pane-placement guard trigger keys on lead existence, not tab:
The guard's lead half now fires whenever fleet.leaders has any entry,
since a lead's tab is always labelled by the fixed LEAD_TAB_LABEL
constant regardless of its own deprecated tab: field. The refusal
message no longer advises removing a lead's tab:, which cannot
satisfy the guard any more.
2026-10-05 14:19:17 +02:00
Dai Ha a3d296f639 fleetd #770: the lead identity check is the label AND the space
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 2m11s
The redeploy skill's check 4 and the script's closing hint both told the
operator that a lead is found by fleet.leaders.*.tab. Identity is now the
fixed 'lead' label together with the lead's configured workspace, so both
would have sent a reader to a key that no longer decides anything.
2026-10-05 14:01:54 +02:00
Dai Ha 79423787d0 fleetd #770: lead identity keys on the space, not the tab label
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 1m57s
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m52s
The lead tab label becomes a fixed constant (Leader.LEAD_TAB_LABEL = "lead");
fleet.leaders.<name>.tab is now optional legacy, matched case-insensitively
alongside the constant via Leader.acceptedLabels(). The uniqueness boundary
between leads moves from the exact tab text to the workspace: FleetConfig
refuses two leaders that share a workspace, LeadTabScanner indexes lead
labels per space (collaborators stay space-agnostic), and
LeadLauncher.leadNameOf/countLeads require both the accepted label and the
lead's own space to match, so a legacy-labelled tab in the wrong space never
counts and a daemon restart never double-spawns a second lead next to a live
one. Config validation also refuses a fleet.tabLabel template or a
collaborator tab that can render as the fixed lead label.
2026-10-05 13:49:14 +02:00
Dai Ha 364b229db9 fleetd #771 step 1: report each pane's workspace label in fleet_list
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m57s
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m50s
Adds workspaceLabel next to workspaceId on fleet_list's panes row, read
from herdr's workspace.list via a new PaneLocator.workspaceLabelsByWorkspaceId().
An observer's reduced row still excludes it; a herdr failure or an unknown
workspaceId yields workspaceLabel: null without costing the rest of fleet_list.
2026-10-05 11:37:41 +02:00
Dai Ha ac436efefb fleetd #761: invariant 4 — a lead's own pane needs an empty input box
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m58s
2026-10-05 11:17:33 +02:00
Dai Ha 7dad054045 Merge remote-tracking branch 'origin/worker/761-draft-aware-lead-nudge-d05850-4' into vfy/761
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 2m16s
2026-10-05 11:09:40 +02:00
Dai Ha 4e414c475f fleetd #761: read the input box marker the live TUI actually draws
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 2m1s
The gate matched the box line as "│ >", which appears zero times on a current
Claude Code pane. EMPTY was unreachable, so every lead nudge was held forever.

Match the box as a line *starting* with "❯" or with "│ >", keeping the older
bordered layout readable. A marker further along a line is transcript text — a
caret the operator quoted — so it no longer counts, and the last matching line
is still the live box because the detection region carries scrollback above it.

Look for the generating marker only from the box line down, for the same
reason: an earlier turn's "esc to interrupt" survives in that scrollback, and
holding on it would be the same unreachable-EMPTY failure by another route.

Fixtures: add IDLE_PROMPT_CARET / DRAFTED_PROMPT_CARET from a live pane and
point every lead-pane fake at them. The bordered constants stay, now covering
the older client. Against the old marker these fixtures fail 55 tests across
PromptBoxTest and the three loop tests, which is the production bug reproduced.

cd fleetd && mvn clean install -> BUILD SUCCESS, MVN_EXIT=0,
Tests run: 2164, Failures: 0, Errors: 0, Skipped: 0 (179 surefire XML files).
2026-10-05 11:07:40 +02:00
Dai Ha 9084667493 fleetd #761: hold a lead nudge while the operator's prompt box holds a draft
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m13s
herdr's agent.prompt pastes AND submits in one call, so a nudge arriving
while the operator is mid-sentence submitted their unfinished line with the
nudge glued to it. AgentStatus.injectable() cannot see this: it describes
the agent, and an idle agent reports the same status whether its input box
is empty or holds a half-typed line.

New herdr/PromptBox reads the pane's `detection` region — the same region
StatusRefiner uses, and the one the input box is drawn in — and clears a
delivery only when the box is positively empty. A box with characters, a
pane it cannot recognise, and a failed read all hold, because a held nudge
is recoverable and a submitted half-line is not. Whitespace and a cursor
block count as empty. After 20 consecutive holds for one target it logs one
warning, so a box that never clears is visible rather than silent; the
warning repeats only after the box has cleared again.

Wired into the three paths that nudge a lead's own pane:
  - ReplyPushLoop.decide -> WAIT_BUSY (the pending work is re-read next tick)
  - LeadHeartbeatLoop.tick -> a new DRAFT_HELD action that spends neither the
    quiet budget nor the one context notice per HIGH stretch
  - LeadCoordLoop.tick -> the peer message stays held and unacked

Not wired into inject/Injector: no human types into a spawned member's pane,
so it would buy nothing and cost a herdr agent.read per member poll. The
heartbeat reads the pane only for a tick that would otherwise send.

Tests. FakeHerdr gains detectionText(), because one readText cannot be both
a worker's transcript (what the completion scrape reads) and a lead's empty
prompt; it falls back to readText so no existing fixture changes meaning.
Fixtures for the lead-nudge paths now state what their pane shows, since the
behaviour depends on it. 11 new behavioural tests across the three loops plus
InjectorTest, and 13 for the classifier; all 10 loop tests were run against
the unpatched loops first and fail there.

mvn clean install: BUILD SUCCESS, Tests run: 2161, Failures: 0, Errors: 0,
Skipped: 0 (summed from target/surefire-reports).
2026-10-05 10:49:50 +02:00
Dai Ha 2289e94223 fleetd #759: fix the fleet_reply comment and the hand-copied role list
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m59s
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m59s
Finding 4: the comment above fleet_reply's handler claimed the authz check
asks whether the caller is a worker at all. It actually checks terminal
ownership (Authz.java REPLY/ASK -> caller.ownsSession), which is why an
observer can reply on its own pane with no role test involved.

Finding 5: fleet_list's tool description hardcoded 'architect/dev/reviewer',
missing hunter. Added MemberRole.wireNames() (pulled out of parse()'s error
message builder, which now calls it too) and used it in the description so
the list can't drift again.
2026-10-05 10:44:55 +02:00
Dai Ha 92adfcfae5 fleetd #756/#758: the canonical block no longer says panes are unlistable
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 2m13s
The table row for an unconfigured pane told every session "Neither
fleet_list nor ListAgents lists these". PR #762 made that false for
fleet_list, and a stale note of this shape is the worst kind: it tells a
future session it cannot do the thing at the moment doing it is the job.

Two edits:

- The observer definition now says where an observer finds a target id,
  which is the one thing it could not learn before.
- The table row names the panes array, says the row is full for a lead, an
  architect or a collaborator and filtered for an observer, and keeps the
  herdr tab list join as the fallback. ListAgents still lists none of them.

wiki/7-Use-Cases.md regenerated from this block in the wiki submodule at
4872227; the sync check prints in sync: True.
2026-10-05 10:39:32 +02:00
Dai Ha 5f7f388e69 Merge remote-tracking branch 'origin/worker/756-758-observer-pane-discovery-7e6ffd-1' into vfy/762
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m56s
2026-10-05 10:29:49 +02:00
Dai Ha 39accf73e6 fleetd #759: fix the third copy of the role claim, and drop two references that rot
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m58s
ConnectionIdentity's Caller record carried the same wrong rule as the two
places PR #760 fixed: it called the terminal a worker's, and read a null
terminal as the primary. The brief for #760 named only two of the three spots.

MemberPresence pointed at FleetMcp.markTrackedCallerPresent by name inside
{@code}, which the compiler does not check, and inject has no dependency on
mcp so a {@link} would add a cross-package reference. States the principle
instead, which stays true whichever roles qualify. Drops a ticket key.
2026-10-05 10:26:36 +02:00
Dai Ha 5c2f296bc3 fleetd #756/#758: panes array reads the architect-slot gate and widens to observer
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m9s
paneRole now reads CallerResolver#boundToArchitectSlot (made public, no
second definition) so a slot-bound pane with no live member reports
"architect", matching what sendableObserverTarget already allowed as a
SEND target.

panesVisibleTo now admits an observer, since an observer holds SEND to
another observer pane. Its panes rows are filtered to
sendableObserverTarget and reduced to sessionId/label/status/role/
deliverable; every other caller's rows are unchanged.
2026-10-05 10:24:00 +02:00
73 changed files with 5694 additions and 676 deletions
+2 -2
View File
@@ -8,8 +8,8 @@
{
"name": "fleet",
"source": "./plugin",
"description": "Mount the fleetd MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.2.0",
"description": "Apply standard Claude Code settings so a session can orchestrate delegated workers, and run the fleet mod for cross-session messaging. Mounting the fleetd MCP gateway is the instance's or the project's job, not this plugin's. Ships no credentials.",
"version": "0.3.0",
"author": {
"name": "LTMS"
}
+10 -4
View File
@@ -188,10 +188,14 @@ Run the control line too. A broken pattern returns a clean `0` that reads exactl
If the first number has grown past 20, somebody has rolled under the restart path, and this section
should be replaced with what they measured.
**Three separate timeouts bound a roll.** `leadRollover.relaunchReadySeconds` (default 45,
`FleetConfig.java:1483`) bounds **each** of two waits that run after the relaunch, so the worst case
there is about twice that number, not 45 seconds in total. A third bound gives your old pane 10
seconds to die (`LeadRollover.PANE_DEATH_TIMEOUT_SECONDS`).
**Four separate waits bound a roll, under three different budgets.**
`leadRollover.relaunchReadySeconds` (default 45, `FleetConfig.java:1493`) bounds **each** of the
first three waits that run after the relaunch, so the worst case there is about three times that
number, not 45 seconds in total. The third of them retries `bootstrapText` while herdr answers
`agent_not_ready`. The fourth wait is separate and much shorter — 5 seconds
(`LeadRollover.BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS`): it waits for the fresh pane's status to reach
`WORKING` after the send, which confirms the successor actually started a turn on the text. A third
bound gives your old pane 10 seconds to die (`LeadRollover.PANE_DEATH_TIMEOUT_SECONDS`).
`fleet_handover{action: "status", token}` answers with one of these:
@@ -204,6 +208,8 @@ seconds to die (`LeadRollover.PANE_DEATH_TIMEOUT_SECONDS`).
| `RELAUNCH_FAILED` | launching the fresh pane failed |
| `RELAUNCH_NEVER_READY` | the fresh pane never became ready within `relaunchReadySeconds` |
| `RELAUNCH_NOT_RECOGNISED` | the fresh terminal never resolved as a lead |
| `BOOTSTRAP_NEVER_SENT` | the fresh pane was ready, but herdr refused `bootstrapText` with `agent_not_ready` for the whole bound, so the successor never learned where the handover file is |
| `BOOTSTRAP_NOT_CONFIRMED` | herdr accepted `bootstrapText`, but the fresh pane never started a turn on it within `BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS` (5s); the send is unconfirmed and nothing was re-sent. A turn shorter than one 250ms status poll can be missed, so check the successor itself before treating the text as lost |
| `FAILED` | the roll threw; `runRollover`'s catch records this rather than leaving it stuck |
Only `TURN_NEVER_SETTLED` guarantees your context is intact. The other failures can leave you
+6 -3
View File
@@ -59,9 +59,12 @@ if the script is unavailable or a step fails, this is what it was protecting you
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
4. **Re-check identity afterwards.** Call `fleet_whoami` and confirm it still answers `primary`. A
lead is found by two things together: its tab is labelled `lead`, and that tab sits in the space
named by `fleet.leaders.<name>.workspace`. Both must match, so a renamed tab *and* a space whose
label differs from the config each demote the lead to worker, which refuses every orchestration
call. A `tab:` still in config is accepted as a second label for that lead, and the daemon logs
one deprecation warning naming it at startup.
5. **Prove the new jar is the one running.** Confirm a *fresh* `fleetd listening` line at the end of
`fleetd/fleetd.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
+9
View File
@@ -37,6 +37,15 @@ confident-but-wrong finding. Anything you could settle by reading more code is y
## 4. The finding — what goes in `fleet_reply`
**Call `fleet_reply` as soon as you know your answer, before you write the reasoning out.** The
four lines below are the whole deliverable, and they are short on purpose. Analysis you type into
your terminal reaches nobody: when a turn ends with no `fleet_reply`, the bridge scrapes the pane
and the lead receives a clipped fragment instead of a finding. A long, correct analysis and no
`fleet_reply` is a failed turn, and it is the most common way this role fails.
If the delegation also handed you a list of things to check, that list is where to *look*. It is
not the shape of the answer. Work the list, then still send these four lines.
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
+77 -18
View File
@@ -33,8 +33,13 @@ gate uses. A worker also carries its `sessionId`, `profile`, `worktree` and `bra
carries the slot name it was bound to; a collaborator carries its registry name and its own
`sessionId`, and **no `leader` key** — a collaborator is a named peer, not a primary. An **observer**
carries only its own `sessionId`: a pane the daemon could not place as any of the above, authorized
to `READ`/`METRICS`, to `REPLY`/`ASK` on its own pane, and to `SEND` only to a target that resolves
as an observer too — never to a lead, a collaborator, or a spawned member, and never a ticket.
to `READ`/`METRICS`, to `REPLY`/`ASK`/`INBOX` on its own pane, to `SEND` only to a lead or to a target
that resolves as an observer too — never to a collaborator, an architect, or a spawned member — and to
`TASK_READ` only a ticket its own `fleet_send{wait:false}` created. Nothing wider: not another
caller's ticket, and not a session's status, which stays refused. It finds such a target in
`fleet_list`'s `panes` array, which for an observer is filtered to
exactly what it may send to and reduced to `sessionId`, `label`, `status`, `role` and `deliverable`;
a lead's row reads `role: "lead"`.
Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
@@ -65,16 +70,54 @@ and the sender silently receives nothing. Fail toward the recoverable error.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead, architect,
collaborator, or observer** — and a collaborator may send only to a lead or another collaborator,
never to a spawned member's terminal, while an observer may send only to another observer pane;
reply/ask are only-as-itself — any peer may answer for its own pane,
and for no other. A call outside your role is refused, not queued.
never to a spawned member's terminal, while an observer may send only to a lead or another observer
pane;
reply/ask/inbox are only-as-itself — any peer may answer, or collect its mail, for its own
pane, and for no other. **A ticket is owned by the caller that created it, not by a role**, so an
observer or a collaborator may poll a ticket its own `fleet_send{wait:false}` created and no
other — never another caller's, and never a session's status, which stays refused for both. A
call outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
**A lead's own pane has a second gate, and it is off by default.** The multiplexer pastes and
submits in one step, so a delivery that lands while the operator is typing would submit their
half-written line. A heartbeat, a ticket nudge and lead-to-lead mail can therefore wait until the
box is clear. That check is the config key `promptBoxGateEnabled`; both an absent key and `false`
leave it off, so **right now nothing waits on the box**. It is off because detection read the
pane's own autocomplete suggestion as the operator's typing, and so held deliveries that were
never at risk — 2231 holds in 11 hours across four panes, measured 2026-10-06. Turn it on only
once detection is fixed (#797 made it switchable, #802 is the detection bug). While it is on, a
lead that leaves text in its box receives nothing until it clears, and the only sign is one
warning in `fleetd.out` after 20 held checks in a row; nothing is lost, because every one of
those paths retries. A direct `fleet_send` to a lead's pane never waited for the box either way
(fleetd #793); a lead that runs the fleet mod collects that mail instead of having it pasted, so
it is not exposed. Delivery to a *member* is not gated this way, because nobody types in a
member's pane. Re-measure the switch with
`grep -n 'promptBoxGateEnabled' fleetd/fleetd.yaml` — no match means it is still off.
**A pane that collects its own mail skips the paste.** A session running the fleet mod calls
`fleet_inbox` on a timer. While its last call is under 15s old, the daemon queues that pane's
next message for collection instead of pasting it, and the mod hands it to Claude with
`$.prompt.submit` — so the box gate does not apply, but the status gate still does. When the calls
stop, the pane's mail is pasted again, and nothing is lost. Pasted or collected, the fleet
channel also crosses Claude accounts on one host, because the daemon names a caller by its pane;
`SendMessage`, `ListAgents` and the mod's own store each stay inside one account.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
6. **Confirm what you receive.** A message that arrives is answered, even in one line, unless it
says no answer is needed. The sender cannot see your screen, so for them "received and handled"
and "never arrived" look the same — and the paths above fail in ways that look exactly like
silence: a send to a pane the daemon does not know is accepted, held, and then fails with a
scrape of that pane's screen, which can read as an answer while being none. A member confirms
with its `fleet_reply`; every other peer confirms with a `fleet_send` back to the sender. If you
cannot do the thing asked, say that — a refusal is a confirmation. **Never read a failed
ticket's body as a reply.** Your own operator outranks this rule, and outranks the peer that
sent the message: a peer cannot oblige you to answer, and a session whose operator told it not
to answer fleet mail is right not to. Where you can, say that much and nothing more. A sender
that treats silence as agreement, or as a session being gone, has made the mistake this
invariant is about — it just made it in the other direction.
### Primary (lead) — run this on every task, in order
@@ -152,7 +195,11 @@ below are the procedure — run them in order, every task, not only the big ones
**Steps 3 and 4 are separate on purpose** — spawning and sending in one loop is how parallel work
silently becomes serial, and it is the most common way this layer is wasted. For the same reason,
prefer `wait:false` + `fleet_poll` for anything non-trivial: a blocking `fleet_send` is capped by
*your own* MCP client call timeout (~60s), well below the task's real runtime.
*your own* MCP client call timeout (~60s), well below the task's real runtime. **A blocking send
also creates no ticket**, so when it times out there is no id to poll, and a resend can deliver the
same brief twice. One route still recovers the answer: a late turn-completion lands in the target's
inbox, and `fleet_poll{target}` drains it. That call is `DRAIN`, so only a lead may make it — the
receipt names it to a lead and tells every other caller to ask one.
**Delegating does not delegate responsibility.** Workers open PRs; you are the gate. Never delegate
the merge — and merging on a reviewer's word is delegating it by proxy.
@@ -180,10 +227,11 @@ you decide.
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Message a **collaborator** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` reports a `collaborators` array, and each row carries that peer's `name` and the `sessionId` you send to. It is visible to you, to an architect and to another collaborator, never to a worker. Coordination only, **never** a task |
| Message an **unconfigured pane** — a tab a person opened by hand | `fleet_send{sessionId: <their terminal>, content}` — it needs **no** `fleet.collaborators` entry and no restart, because a pane becomes deliverable the moment its agent connects the bridge MCP. **Neither `fleet_list` nor `ListAgents` lists these**, so read the terminal id by joining `herdr tab list` to `GET /agents` on `tab_id`. Such a pane resolves as an `observer`: it can answer you with `fleet_reply`, and it can `fleet_send` to another observer pane, but never to you. Coordination only, **never** a task |
| Message an **unconfigured pane** — a tab a person opened by hand | `fleet_send{sessionId: <their terminal>, content}` — it needs **no** `fleet.collaborators` entry and no restart, because a pane becomes deliverable the moment its agent connects the bridge MCP. `fleet_list`'s `panes` array reports every such pane with its label and the terminal id to send to — the full row for you, an architect or a collaborator; filtered and reduced for an observer. **`ListAgents` still never lists these**, and joining `herdr tab list` to `GET /agents` on `tab_id` stays the read-only fallback if the array is missing. Such a pane resolves as an `observer`: it can answer you with `fleet_reply`, and it can `fleet_send` to you or to another observer pane, but never to a collaborator or a member. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Collect the mail queued for your OWN pane, instead of having it pasted | `fleet_inbox` — no arguments, any role, own pane only. The fleet mod calls it on a timer; you rarely call it by hand |
| Tear down a member | `fleet_stop{paneId}` |
| Replace your OWN lead session when its context is full | `fleet_handover{action:"open", reason?}` → write the handover file it names → `fleet_handover{action:"confirm", token, operatorConfirmed}`. Primary-only. **In that order**: the file must be modified *after* `open`, or `confirm` refuses it as stale. There is no terminal parameter — the pane is always your own, so you can never roll another lead. `{action:"cancel", token}` drops a pending request |
@@ -258,19 +306,21 @@ simply complies has thrown away the reason there are two of you.
`fleet_whoami` answered `collaborator`, so your pane's tab matches a `fleet.collaborators.<name>.tab`
entry. You are **not** a member: nothing delegates to you, you have no brief, no worktree and no
ticket, and **you owe no `fleet_reply`** — the turn contract above is for a session a lead spawned,
and it does not apply to you. Read it only to understand what the members around you are doing.
ticket raised against you, and **you owe no `fleet_reply`** — the turn contract above is for a
session a lead spawned, and it does not apply to you. Read it only to understand what the members around you are doing.
What you may do: observe the fleet (`fleet_list`, `fleet_profiles`, `fleet_whoami`), and send to a
lead or to another collaborator. What you may not: spawn, stop or drain anything, roll a lead's
session, answer a member's `fleet_ask`, poll a ticket, or send to a spawned member's terminal. Each
of those is refused at the gate, not queued.
What you may do: observe the fleet (`fleet_list`, `fleet_profiles`, `fleet_whoami`), send to a lead
or to another collaborator, and poll a ticket your own `fleet_send{wait:false}` created. What you may
not: spawn, stop or drain anything, roll a lead's session, answer a member's `fleet_ask`, read
another caller's ticket or any session's status, or send to a spawned member's terminal. Each of
those is refused at the gate, not queued.
Two limits worth knowing before you hit them. **You cannot reach a worker** — not even to help one —
because a worker belongs to the lead that spawned it, and routing around that would make you a
second orchestrator with no plan. Send to the lead instead. And **you cannot read a ticket**, so you
cannot collect a delegation's reply: `fleet_poll` refuses you at the role gate, and a ticket also
records the terminal that created it, so even a leaked id reads nothing.
second orchestrator with no plan. Send to the lead instead. And **you cannot read another caller's
ticket** — only one your own `fleet_send{wait:false}` created. A ticket records the terminal that
created it, so a leaked id reads nothing, and a lead's delegation stays the lead's: you cannot
collect its reply.
Being named buys you a channel, not authority. Your `fleet_send` to a lead is coordination between
peers: the lead owes you no obedience, and you owe it none.
@@ -336,8 +386,17 @@ must obey belongs in the charter, not here.
`handover` (write the file a fresh lead session inherits when the outgoing one hands off,
fleetd #480).
- **This repo is also a Claude Code marketplace, and ships a plugin.** `.claude-plugin/marketplace.json`
points at `plugin/`, which carries the MCP mount and the `setup` skill
(`/claude-bridge:setup` — make any project bridge-ready). It was added in CB-527 and then went
points at `plugin/`, which carries the `setup` skill (`/fleet:setup` — make any project
bridge-ready) and no MCP mount: the instance or the project mounts `fleet`. `plugin/` is also a Claude Code mod (`plugin/hooks/`):
it polls `fleet_inbox` and hands each message to Claude with `$.prompt.submit`. Measured
2026-10-06: `fleet@fleetd` 0.3.0 is installed at user scope in the `gx10`, `ltms`, `ollama` and
`work` instances, so every session started from them runs the mod, and a spawned member on those
config dirs loads it too — but the mod skips the inbox poll for a `worker` or an `architect`, so a
member still gets its brief pasted. The install made a copy under each instance's
`plugins/cache/fleetd/fleet/0.3.0`, so assume an edit to `plugin/` reaches sessions only after a
version bump and `claude plugin update fleet@fleetd` per instance. Re-measure with
`grep -l '"fleet@fleetd"' ~/.ccs/instances/*/plugins/installed_plugins.json`; delete this sentence
if the plugin is uninstalled. It was added in CB-527 and then went
unmentioned by every instruction file, so it drifted and a later session planned it from scratch
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
+6
View File
@@ -174,6 +174,12 @@ bind:
# idleSleepGuard:
# enabled: false
# Whether PromptBox checks a lead's input box before a delivery pastes into it. Default off —
# omitting this key (or setting it false) leaves every delivery proceeding without reading the
# pane. Detection currently misreads the pane's own autocomplete suggestion at the caret as the
# operator's unsubmitted typing, so the check holds deliveries that were never actually at risk.
# promptBoxGateEnabled: false
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -963,12 +963,15 @@ public final class Fleetd {
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
* the same {@code leads} supplier {@code main} already builds for
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
* @param mcpPresent whether a terminal's Claude has connected the bridge MCP client,
* normally {@code MemberPresence::isPresent} — the signal {@link
* LeadRollover}'s own readiness wait requires alongside a turn boundary
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
* absent from the startup config
*/
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents,
WorkspaceControl leadSpaces, LeadLauncher launcher, ConfigRef config,
Supplier<Map<String, String>> liveLeadTerminals) {
Supplier<Map<String, String>> liveLeadTerminals, Predicate<String> mcpPresent) {
if (cfg.leadRollover() == null) {
return null;
}
@@ -983,7 +986,7 @@ public final class Fleetd {
return leader == null ? null : leader.cwd();
};
return new LeadRollover(leadAgents, leadSpaces, launcher, () -> config.get().leadRollover(),
leadWorkspace, leadNameForTerminal, liveLeadTerminals);
leadWorkspace, leadNameForTerminal, liveLeadTerminals, mcpPresent);
}
/**
@@ -254,18 +254,28 @@ final class FleetdAssembly {
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
// configured lead's own exact `tab:` label. fleetd #669: the same scan also recognises a
// configured collaborator's tab, so one herdr pass answers both.
// Discover leads by their tab labels, one scanner per configured lead's own space. fleetd
// #669: the same scan also recognises a configured collaborator's tab, so one herdr pass
// answers both.
final Supplier<Map<String, String>> leads;
final Supplier<Map<String, String>> collaboratorTerminals;
var leaders = cfg.fleet().leaders();
var collaboratorsConfig = cfg.fleet().collaborators();
if (!leaders.isEmpty() || !collaboratorsConfig.isEmpty()) {
Map<String, String> tabToName = new LinkedHashMap<>();
Map<String, Map<String, String>> leadLabelsBySpace = new LinkedHashMap<>();
Map<String, String> spaceByLeadName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
if (leader == null) {
return;
}
spaceByLeadName.put(name, leader.workspace());
Map<String, String> labelsHere = leadLabelsBySpace
.computeIfAbsent(leader.workspace(), k -> new LinkedHashMap<>());
leader.acceptedLabels().forEach(label -> labelsHere.put(label, name));
if (leader.tab() != null && !leader.tab().isBlank()) {
log.warn("lead '{}' (fleet.leaders.{}) still configures tab: \"{}\" — deprecated, "
+ "the lead tab label is now fixed to '{}'",
name, name, leader.tab(), FleetConfig.Leader.LEAD_TAB_LABEL);
}
});
Map<String, String> collaboratorTabToName = new LinkedHashMap<>();
@@ -283,13 +293,13 @@ final class FleetdAssembly {
? 10
: leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
LeadTabScanner scanner = new LeadTabScanner(herdr, tabToName, collaboratorTabToName, Set.of(),
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
LeadTabScanner scanner = new LeadTabScanner(herdr, leadLabelsBySpace, collaboratorTabToName,
Set.of(), TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
leads = scanner;
collaboratorTerminals = scanner::collaborators;
log.info("lead/collaborator scan: tabs {} host a lead, tabs {} host a collaborator "
+ "(rescan every {}s, shared fleet space)",
tabToName.keySet(), collaboratorTabToName.keySet(), scanIntervalSeconds);
log.info("lead/collaborator scan: space per lead {}, tabs {} host a collaborator "
+ "(rescan every {}s)",
spaceByLeadName, collaboratorTabToName.keySet(), scanIntervalSeconds);
} else {
leads = () -> leadTerminals;
collaboratorTerminals = Map::of;
@@ -411,7 +421,7 @@ final class FleetdAssembly {
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
pushScheduler, maxReminders, backoffMs, metrics, cfg.promptBoxGateEnabled());
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
// above at the real push loop, now that it exists.
pushLoopRef.set(pushLoop);
@@ -438,7 +448,8 @@ final class FleetdAssembly {
Fleetd.leadContextSource(leadContextGauge, router.leadAgents(), leads,
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders),
Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)),
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm,
cfg.promptBoxGateEnabled());
heartbeat.start();
} else {
heartbeat = null;
@@ -446,7 +457,7 @@ final class FleetdAssembly {
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), router.leadSpaces(),
leadLauncher, config, leads);
leadLauncher, config, leads, presence::isPresent);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
@@ -535,7 +546,7 @@ final class FleetdAssembly {
if (leadMailbox != null) {
var leadCoordScheduler = ports.newScheduler("bridge-leadcoord-");
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
LEAD_COORD_INTERVAL_MS);
LEAD_COORD_INTERVAL_MS, cfg.promptBoxGateEnabled());
leadCoordLoop.start(); // FIFTH of the recurring background loops to start (optional).
leadCoordSchedulerRef = leadCoordScheduler;
} else {
@@ -32,6 +32,13 @@ public final class Authz {
REPLY,
/** A worker's mid-turn question to the primary. */
ASK,
/**
* Collect the messages queued for the caller's OWN pane, instead of having them typed into
* its terminal. Grouped with {@link #REPLY} and {@link #ASK} below as an only-as-itself
* action: the pane is always the caller's connection-resolved terminal, never an argument,
* so no caller can collect another pane's mail.
*/
INBOX,
/** Collect held replies from a session's inbox. */
DRAIN,
/** Read-only roster, profile, and identity observation: no ticket, task, or turn state. */
@@ -73,55 +80,90 @@ public final class Authz {
/**
* The fail-closed classifier for an observer's {@code SEND}: answers no for every target, so
* the grant is refused unless a caller supplies a real one. {@code
* CallerResolver#sendableObserverTarget()} is the real one, read from the same maps {@code
* CallerResolver#resolve} consults, so a target that classifier calls known is one {@code
* resolve} would actually resolve as {@link Role#OBSERVER}.
* CallerResolver#observerSendTarget()} is the real one, read from the same maps {@code
* CallerResolver#resolve} consults, so a target that classifier accepts is one {@code resolve}
* would actually resolve as a lead ({@link Role#PRIMARY}) or as {@link Role#OBSERVER}.
*/
public static final Predicate<String> NO_KNOWN_OBSERVER_TARGET = target -> false;
public static final Predicate<String> NO_OBSERVER_SEND_TARGET = target -> false;
/**
* The fail-closed classifier for a ticket-scoped {@code TASK_READ} (an observer's or a
* collaborator's): answers no for every ticket, so the grant is refused unless a caller
* supplies a real one. {@code MessageService#ownsTicket(String, String)} is the real one, so
* a ticket this classifier accepts is one that caller's own {@code fleet_send(wait:false)}
* actually created.
*/
public static final Predicate<String> NO_OWNED_TICKET = ticket -> false;
/**
* Convenience form for a caller with no classifier to supply. Fails closed: a collaborator's
* or an observer's {@code SEND} is refused, as if no terminal were a configured lead,
* collaborator, or observer target — the same decision {@link #NO_KNOWN_LEAD_OR_COLLABORATOR}
* and {@link #NO_KNOWN_OBSERVER_TARGET} give explicitly. Every other action's result is
* identical to the five-argument form's, since none of them consult either classifier.
* or an observer's {@code SEND}, and an observer's or a collaborator's {@code TASK_READ}, are
* refused, as if no terminal were a configured lead, collaborator, or observer-reachable
* target, and no ticket were the caller's own — the same decision
* {@link #NO_KNOWN_LEAD_OR_COLLABORATOR}, {@link #NO_OBSERVER_SEND_TARGET} and
* {@link #NO_OWNED_TICKET} give explicitly. Every other action's result is identical to the
* full form's, since none of them consult any of the three classifiers.
*
* <p>Its default classifiers deny every collaborator and every observer, so a caller
* enforcing authorization must use the five-argument form instead.
* enforcing authorization must use the full form instead.
*/
public static boolean permits(Principal caller, Action action, String targetSession) {
return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR, NO_KNOWN_OBSERVER_TARGET);
return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR,
NO_OBSERVER_SEND_TARGET, NO_OWNED_TICKET);
}
/**
* As {@link #permits(Principal, Action, String)}, with a real classifier for a collaborator's
* {@code SEND}. An observer's {@code SEND} still fails closed ({@link #NO_KNOWN_OBSERVER_TARGET}) —
* a caller enforcing both grants must use the five-argument form.
* {@code SEND}. An observer's {@code SEND}, and an observer's or a collaborator's
* {@code TASK_READ}, still fail closed — a caller enforcing all three grants must use the
* full form.
*/
public static boolean permits(Principal caller, Action action, String targetSession,
Predicate<String> knownLeadOrCollaborator) {
return permits(caller, action, targetSession, knownLeadOrCollaborator, NO_KNOWN_OBSERVER_TARGET);
return permits(caller, action, targetSession, knownLeadOrCollaborator,
NO_OBSERVER_SEND_TARGET, NO_OWNED_TICKET);
}
/**
* As {@link #permits(Principal, Action, String, Predicate)}, with a real classifier for an
* observer's {@code SEND} too. An observer's or a collaborator's {@code TASK_READ} still
* fails closed — a caller enforcing all three grants must use the full form.
*/
public static boolean permits(Principal caller, Action action, String targetSession,
Predicate<String> knownLeadOrCollaborator,
Predicate<String> observerSendTarget) {
return permits(caller, action, targetSession, knownLeadOrCollaborator, observerSendTarget,
NO_OWNED_TICKET);
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
*
* @param targetSession the session id in the request path; only consulted for the
* worker-scoped actions ({@code REPLY}, {@code ASK}), for a
* collaborator's {@code SEND}, and for an observer's
* {@code SEND}, ignored otherwise, may be {@code null}
* worker-scoped actions ({@code REPLY}, {@code ASK},
* {@code INBOX}), for a collaborator's {@code SEND}, for an
* observer's {@code SEND}, and — carrying a ticket id instead
* of a session id — for an observer's or a collaborator's
* {@code TASK_READ}, ignored otherwise, may be {@code null}
* @param knownLeadOrCollaborator whether a terminal is a configured lead or collaborator —
* consulted only for a collaborator's {@code SEND}, to confine
* it to another named peer and never a spawned member's
* terminal
* @param knownObserverTarget whether a terminal is one this daemon would itself resolve as
* {@link Role#OBSERVER} — consulted only for an observer's
* {@code SEND}, to confine it to another observer pane and never
* a lead, a collaborator, or a spawned member
* @param observerSendTarget whether a terminal is one this daemon would itself resolve as
* a lead ({@link Role#PRIMARY}) or as {@link Role#OBSERVER} —
* consulted only for an observer's {@code SEND}, to confine it
* to a lead or another observer pane and never a collaborator,
* an architect, or a spawned member
* @param ticketOwnedByCaller whether {@code targetSession} (here, a ticket id) was created
* by this same caller — consulted only for an observer's or a
* collaborator's {@code TASK_READ}, to confine it to a ticket
* its own {@code fleet_send(wait:false)} created, never another
* caller's
*/
public static boolean permits(Principal caller, Action action, String targetSession,
Predicate<String> knownLeadOrCollaborator,
Predicate<String> knownObserverTarget) {
Predicate<String> observerSendTarget,
Predicate<String> ticketOwnedByCaller) {
if (caller == null || caller.isAnonymous()) {
return false; // authenticated as nothing ⇒ authorized for nothing
}
@@ -135,12 +177,12 @@ public final class Authz {
// Delivering a turn to a local session is open to the primary and the architect
// unconditionally. A collaborator may reach only a target that is itself a configured
// lead or collaborator, never a spawned member's terminal. An observer may reach only
// a target that would itself resolve as an observer, never a lead, a collaborator, or
// a spawned member. A worker is excluded from every case — sending would be it
// escalating into the orchestrator role.
// a target that would itself resolve as a lead or as another observer, never a
// collaborator, an architect, or a spawned member. A worker is excluded from every
// case — sending would be it escalating into the orchestrator role.
case SEND -> caller.isPrimary() || caller.isArchitect()
|| (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession))
|| (caller.isObserver() && knownObserverTarget.test(targetSession));
|| (caller.isObserver() && observerSendTarget.test(targetSession));
// Resolving a worker's blocked question is part of delegating to it, open to the same
// two roles that may stand up that delegation in the first place. Not a collaborator:
@@ -159,7 +201,7 @@ public final class Authz {
// collaborator's own pane passes through the same check, so each can answer a funnel
// that delegated to it. An unnamed primary (token/loopback, no pane) owns nothing and
// is still excluded.
case REPLY, ASK -> caller.ownsSession(targetSession);
case REPLY, ASK, INBOX -> caller.ownsSession(targetSession);
// READ is roster, profile, and identity observation — fleet_list, fleet_profiles, and
// fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no
@@ -170,11 +212,17 @@ public final class Authz {
|| caller.isCollaborator() || caller.isObserver();
// Ticket polling and session status, open to every role READ is open to except a
// collaborator or an observer. MessageService compares a ticket's creator to the
// caller on every read as well, so dropping this gate would not expose another
// session's reply — it would move the refusal later and widen what a caller that
// never orchestrates can probe.
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// collaborator or an observer — with one exception: an observer or a collaborator may
// poll a ticket its own fleet_send(wait:false) created, confined by
// ticketOwnedByCaller. Session status is not ticket-scoped, so that call site
// supplies no real classifier here and an observer's or a collaborator's TASK_READ on
// it stays refused. MessageService compares a ticket's creator to the caller on every
// read as well, so dropping this gate would not expose another session's reply — it
// would move the refusal later and widen what a caller that never orchestrates can
// probe.
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect()
|| ((caller.isObserver() || caller.isCollaborator())
&& ticketOwnedByCaller.test(targetSession));
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
@@ -271,21 +271,31 @@ public final class CallerResolver {
}
/**
* Whether {@code target} names a terminal this resolver would itself resolve as {@link
* Role#OBSERVER} — the classifier an observer's {@code SEND} is checked against, read from the
* same maps and functions {@link #resolve} consults so a target this accepts is exactly one
* {@code resolve} would hand back {@link Role#OBSERVER} for, and the reverse.
* Whether {@code target} names a terminal an observer may {@code SEND} to: one this resolver
* would itself resolve as a lead ({@link Role#PRIMARY}) or as {@link Role#OBSERVER}. Read from
* the same maps and functions {@link #resolve} consults, and in the same order, so a target
* this accepts is exactly one {@code resolve} would hand back one of those two roles for, and
* the reverse.
*
* <p>A live spawned member is refused first, whatever a tab map says about its terminal — the
* order {@link #resolve} itself uses. A pane named as a lead is then accepted even when it is
* also bound to an architect slot, because that is the role {@code resolve} gives it.
*/
public Predicate<String> sendableObserverTarget() {
public Predicate<String> observerSendTarget() {
return target -> target != null
&& spawnedMemberRole.apply(target) == null
&& !leadTerminals.get().containsKey(target)
&& !boundToArchitectSlot(target)
&& !collaboratorTerminals.get().containsKey(target);
&& (leadTerminals.get().containsKey(target)
|| (!boundToArchitectSlot(target)
&& !collaboratorTerminals.get().containsKey(target)));
}
/** Whether {@code terminal} is bound to a configured slot the live roster still confirms as an architect. */
private boolean boundToArchitectSlot(String terminal) {
/**
* Whether {@code terminal} is bound to a configured slot the live roster still confirms as an
* architect — the one classifier {@link #observerSendTarget()} and {@code FleetMcp}'s
* {@code panes} row both read, so a pane's reported role and its {@code SEND} reachability can
* never drift apart.
*/
public boolean boundToArchitectSlot(String terminal) {
String slot = architectTerminals.get().get(terminal);
return slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT;
}
@@ -40,8 +40,9 @@ public enum Role {
* fleet.collaborators.<name>.tab} registry). Never spawned — identity comes from the
* connection, never a request argument, exactly like {@link #WORKER} and {@link #ARCHITECT}.
* May {@code SEND} only to a configured lead or collaborator, {@code REPLY}/{@code ASK} only
* as its own pane, and {@code READ}/{@code METRICS}; may not {@code SPAWN}/{@code STOP}/
* {@code DRAIN}/{@code HANDOVER}, poll a ticket ({@code TASK_READ}), or reach the
* as its own pane, {@code READ}/{@code METRICS}, and {@code TASK_READ} only a ticket its own
* {@code fleet_send(wait:false)} created; may not {@code SPAWN}/{@code STOP}/{@code DRAIN}/
* {@code HANDOVER}, read a session's status or another caller's ticket, or reach the
* coordination broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
COLLABORATOR,
@@ -51,10 +52,11 @@ public enum Role {
* configured lead, not a bound architect slot, not a configured collaborator tab. Unforgeable
* like a worker's — derived from the connection's pane, never from a request argument, and
* honoured regardless of auth mode. May {@code READ} and {@code METRICS}, {@code REPLY}/
* {@code ASK} only as its own pane, and {@code SEND} only to a target that would itself
* resolve as {@code OBSERVER}; may not {@code SPAWN}/{@code STOP}/{@code DRAIN}/
* {@code HANDOVER}, poll a ticket ({@code TASK_READ}), or reach the coordination broker
* ({@code COORD_SEND}/{@code COORD_READ}).
* {@code ASK} only as its own pane, {@code SEND} only to a target that would itself resolve
* as a lead ({@link #PRIMARY}) or as {@code OBSERVER}, and {@code TASK_READ} only a ticket its
* own {@code fleet_send(wait:false)} created; may not {@code SPAWN}/{@code STOP}/
* {@code DRAIN}/{@code HANDOVER}, read a session's status or another caller's ticket, or reach
* the coordination broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
OBSERVER,
@@ -183,8 +183,9 @@ import java.util.function.Supplier;
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
* {@code models:} was added as deferred, again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere, and again after
* {@code leadRollover:} was added (fleetd #480).</strong>
* {@code FleetConfig} has 26 top-level record components: 5 cold, 13 deferred, 3 split, 5
* {@code leadRollover:} was added (fleetd #480), and again after {@code promptBoxGateEnabled:}
* was added as deferred.</strong>
* {@code FleetConfig} has 27 top-level record components: 5 cold, 14 deferred, 3 split, 5
* hot-excluded. Five of them are named nowhere in this file, and the reason is the same for all
* five: {@code placement}, {@code memberCredentials}, {@code memberLoginShell}, {@code models} and
* {@code leadRollover} are <strong>hot</strong> and correctly absent — all five are read live off
@@ -288,7 +289,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
static final Set<String> DEFERRED_KEYS = Set.of(
"guard", "worktreeRoot", "worktreeGroup", "memberSkills", "primary", "configReload",
"leadHeartbeat", "lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs",
"quarantineCooldownSeconds", "profiles", "idleSleepGuard");
"quarantineCooldownSeconds", "profiles", "idleSleepGuard", "promptBoxGateEnabled");
private final Path path;
private final AtomicReference<FleetConfig> current;
@@ -531,6 +532,12 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.idleSleepGuard(), fresh.idleSleepGuard())) {
changed.add("idleSleepGuard");
}
// ReplyPushLoop, LeadCoordLoop and LeadHeartbeatLoop each build their own PromptBox once, at
// startup, off this value — a running loop keeps whichever gate setting it was built with
// regardless of a later edit here.
if (!Objects.equals(old.promptBoxGateEnabled(), fresh.promptBoxGateEnabled())) {
changed.add("promptBoxGateEnabled");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
@@ -28,6 +28,7 @@ import java.util.Comparator;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.regex.Pattern;
@@ -142,6 +143,13 @@ import java.util.regex.PatternSyntaxException;
* @param leadRollover opt-in lead rollover (fleetd #480): {@code null} ⇒ off, and no
* {@code dev.ltms.fleet.lead.LeadRollover} is constructed at all — an upgraded
* daemon never clears a lead's pane on its own initiative. See {@link LeadRollover}.
* @param promptBoxGateEnabled whether {@link dev.ltms.fleet.herdr.PromptBox} checks a lead's input
* box before a delivery pastes into it. {@code null} (the default) and
* {@code false} both leave the check off: a delivery proceeds without reading
* the pane, and the box is never classified. Detection currently misreads the
* pane's own autocomplete suggestion at the caret as the operator's
* unsubmitted typing, so the check holds deliveries that were never actually
* at risk; set this to {@code true} only once that is fixed.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -170,7 +178,23 @@ public record FleetConfig(
String memberSkills,
IdleSleepGuard idleSleepGuard,
Models models,
LeadRollover leadRollover) {
LeadRollover leadRollover,
Boolean promptBoxGateEnabled) {
/** Back-compat form before the {@code promptBoxGateEnabled} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard,
Models models, LeadRollover leadRollover) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, models, leadRollover, null);
}
/** Back-compat form before the {@code leadRollover:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
@@ -330,6 +354,7 @@ public record FleetConfig(
} else {
profiles = Map.of();
}
promptBoxGateEnabled = promptBoxGateEnabled != null && promptBoxGateEnabled;
}
@JsonIgnoreProperties(ignoreUnknown = true)
@@ -1092,8 +1117,8 @@ public record FleetConfig(
}
/**
* One entry of the CB-530 {@code leaders:} registry — a pane that orchestrates rather than one
* that is orchestrated.
* One entry of the {@code leaders:} registry — a pane that orchestrates rather than one that is
* orchestrated.
*
* <p>Why a registry and not a second {@code primary:}: {@code primary.terminal} is singular by
* construction, so a session in any other pane resolves as a worker. That is correct while one
@@ -1103,46 +1128,49 @@ public record FleetConfig(
* <p>{@code kind} and {@code model} are descriptive only: they document what runs in the pane
* and are reported back by {@code fleet_whoami}.
*
* <p><b>A lead is now also creatable (CB-557).</b> Before, nothing spawned one — a lead
* pre-existed, which is why it had to be recognised by configuration rather than created. With
* {@code profile} and {@code instances} the daemon may stand one up when none is live, so the
* pane no longer has to exist before the daemon does. Recognition still comes first: a lead
* already running in its configured {@code tab} is adopted, and only the shortfall is launched.
* <p>A lead with {@code profile} and {@code instances} set may be launched by the daemon when
* none is live; recognition always comes first, so only the shortfall is launched.
*
* <p><b>{@code tab} replaced {@code terminal} (CB-579).</b> A herdr {@code terminal_id} changes
* every time the lead's session restarts, so pinning one cost a config edit and a daemon restart
* per restart. A tab is stable: a human opens it once, it holds exactly one pane, and its label
* survives restarts of the agent inside it — so identity is now the tab label alone.
* <p>Every lead's tab is labelled {@link #LEAD_TAB_LABEL}, a fixed constant — not a per-entry
* config value. {@code workspace} is therefore what tells one lead from another: two leaders
* sharing one space would both resolve to the one tab named {@code lead} there, so only one
* could ever be found. {@code tab} is a deprecated legacy label, still matched within this
* lead's own space alongside the constant.
*
* @param profile the {@code profiles:} entry to launch this lead on when one must
* be created; {@code null} ⇒ recognise-only, never create.
* <p>fleetd #176: also the field {@code Fleetd.leadSeatLookup} reads
* to learn which account this lead's own live session shares — set it
* <p>Also the field {@code Fleetd.leadSeatLookup} reads to learn
* which account this lead's own live session shares — set it
* (safely, even on an already-running recognise-only lead: naming a
* profile here never starts anything beyond {@code instances}) so a
* {@code subscription: true} worker profile sharing its
* {@code effectiveCredentialId()} has this lead's seat subtracted from
* {@code fleet_list}'s {@code free}. {@code null} here also means this
* lead's seat cannot be derived and is not counted.
* @param tab the exact tab label hosting this lead, matched case-insensitively;
* the only field identity depends on. Required — a lead with no
* {@code tab} can never be discovered, launched or not
* @param tab deprecated legacy tab label, matched case-insensitively within
* this lead's own space alongside {@link #LEAD_TAB_LABEL}. Optional —
* {@code null}/blank means only the constant is accepted
* @param instances how many of this lead should be live (default 1). The daemon
* launches only the shortfall, so a restart adopts rather than doubles
* @param tabPrefix lead-tab naming convention checked against member labels. Lead
* identity uses {@code tab}. Default {@code "lead:"}
* identity uses {@link #acceptedLabels()}. Default {@code "lead:"}
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
* worst case before a newly-labelled tab is recognised. Default 10
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
* @param model the model or selector it runs, for operators reading the roster
* @param workspace the space this lead's tab lives in — the uniqueness boundary
* identity now depends on. Default {@link #DEFAULT_WORKSPACE}
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Leader(String profile, String tab, Integer instances, String tabPrefix,
Integer scanIntervalSeconds, String kind, String model,
String workspace, String cwd) {
/** The tab label every lead is found by, and an auto-launched instance is created with. */
public static final String LEAD_TAB_LABEL = "lead";
/**
* Where an auto-launched lead's tab is created (CB-558). It defaults to the SAME shared
* Where an auto-launched lead's tab is created. It defaults to the SAME shared
* {@code "fleet"} space the members use, so the operator sees one "session" with many tabs.
* The scanner no longer excludes member spaces — it tells a lead from a member by the exact
* tab label, so a lead sharing the members' space is still discovered (see LeadLauncher).
@@ -1170,9 +1198,23 @@ public record FleetConfig(
return profile != null && !profile.isBlank() && instances > 0;
}
/** The tab label an auto-launched instance of this lead gets — its configured {@code tab}. */
/** The tab label an auto-launched instance of this lead gets. */
public String tabLabel() {
return tab;
return LEAD_TAB_LABEL;
}
/**
* The normalised labels (stripped, lower-cased) a tab in this lead's own space may carry to
* be recognised as this lead: {@link #LEAD_TAB_LABEL} first, plus the deprecated {@code tab}
* when configured and different. Every matcher in this class's callers reads this method —
* none re-derives the set.
*/
public List<String> acceptedLabels() {
String normalizedTab = (tab == null) ? null : tab.toLowerCase(Locale.ROOT);
if (normalizedTab == null || normalizedTab.equals(LEAD_TAB_LABEL)) {
return List.of(LEAD_TAB_LABEL);
}
return List.of(LEAD_TAB_LABEL, normalizedTab);
}
}
@@ -1448,21 +1490,27 @@ public record FleetConfig(
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
* {@code DONE}) before ending that pane's process at all. See the
* paragraph above.
* @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
* @param relaunchReadySeconds default 45 — bound on EACH of three separate waits that run after
* the old lead's pane has been torn down and a fresh one launched: first,
* for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
* typing into a pane that has not finished booting loses the keystrokes;
* second, for the fresh terminal to show up as a recognised lead, which is
* bookkeeping rather than a safety gate, so a timeout on this second wait
* does not withhold {@code bootstrapText} — it is sent once the pane is
* ready regardless. Recognition comes from the same periodically-refreshed
* scan {@code LeadTabScanner} already keeps ({@code scanIntervalSeconds},
* 10s live), so a budget has to clear more than one scan interval to leave
* any real margin for the CLI's own boot time; 20 was rejected for exactly
* that reason — at a 10s scan interval it only buys two scans. 45 buys
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
* ready) withholds {@code bootstrapText}.
* {@code DONE}, never merely {@code BLOCKED}) AND to show up as present in
* {@code MemberPresence} (its Claude has connected the bridge MCP client) —
* the safety gate, since herdr's own {@code agent_status} reports idle
* through the whole boot window, and typing into that window loses the
* keystrokes; second, for the fresh terminal to show up as a recognised
* lead, which is bookkeeping rather than a safety gate, so a timeout on this
* second wait does not withhold {@code bootstrapText} — it is sent once the
* pane is ready regardless. Recognition comes from the same
* periodically-refreshed scan {@code LeadTabScanner} already keeps ({@code
* scanIntervalSeconds}, 10s live), so a budget has to clear more than one
* scan interval to leave any real margin for the CLI's own boot time; 20 was
* rejected for exactly that reason — at a 10s scan interval it only buys two
* scans. 45 buys roughly four; third, to retry sending {@code bootstrapText}
* while herdr reports {@code agent_not_ready}. A timeout on the first wait
* or the third one withholds {@code bootstrapText}. The post-send wait for
* the successor to act on the text is NOT bounded by this key — it has its
* own, much smaller one in {@code LeadRollover}, because this budget sizes
* a pane booting a CLI rather than an agent reacting to text it holds.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* fresh lead's pane once it reaches a real turn boundary after relaunch,
* telling the fresh session where to read the handover and carry on. Left
@@ -1918,7 +1966,7 @@ public record FleetConfig(
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
"idleSleepGuard", "models", "leadRollover");
"idleSleepGuard", "models", "leadRollover", "promptBoxGateEnabled");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -2715,10 +2763,13 @@ public record FleetConfig(
// own compact constructor defaults the fields of a block that IS present. Defaulting it
// here would construct a LeadRollover object (via Fleetd.java's presence gate) for every
// config that never mentioned it.
// promptBoxGateEnabled is left as-is: the compact constructor already normalized it to a
// real true/false, and false (the gate off) is already the safe default — there is no
// further defaulting to do here.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
idleSleepGuard, models, leadRollover);
idleSleepGuard, models, leadRollover, promptBoxGateEnabled);
}
/**
@@ -2747,23 +2798,42 @@ public record FleetConfig(
/**
* Reject a member tab-label template that could render as a configured lead or collaborator
* tab or match a lead-tab naming convention, and reject two {@code fleet.leaders} or
* {@code fleet.collaborators} entries — across either registry — that share one exact tab.
* tab, as the fixed lead tab label, or that matches a lead-tab naming convention; reject two
* {@code fleet.leaders} entries that share one space; reject two {@code fleet.collaborators}
* entries — or a lead and a collaborator — that share one exact tab; and reject a collaborator
* tab equal to the fixed lead tab label.
*
* <p>{@code fleet.collaborators} has no {@code tabPrefix}: identity is matched on the exact
* {@code tab} alone, so only the exact-render check applies there, not the prefix check.
*
* @throws IllegalStateException when the fleet template or a profile {@code tabLabel} override
* can render as a configured lead or collaborator tab or match a
* lead-tab prefix, or when two entries — of either registry, or
* one of each — carry the same exact {@code tab}
* (case-insensitively)
* can render as a configured lead or collaborator tab, as the
* fixed lead tab label, or match a lead-tab prefix; when two
* leaders share one space; when two collaborators (or a lead and
* a collaborator) carry the same exact {@code tab}
* (case-insensitively); or when a collaborator's {@code tab}
* equals the fixed lead tab label
*/
public void validateLeadTabPrefixes() {
if (fleet == null) {
return;
}
List<String> bad = new ArrayList<>();
if (templateCanRenderAs(fleet.tabLabel(), Leader.LEAD_TAB_LABEL)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as \""
+ Leader.LEAD_TAB_LABEL + "\", the fixed lead tab label");
}
profiles().entrySet().stream()
.map(Map.Entry::getKey)
.sorted()
.forEach(p -> {
String label = profiles().get(p).tabLabel();
if (templateCanRenderAs(label, Leader.LEAD_TAB_LABEL)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which can render as \"" + Leader.LEAD_TAB_LABEL
+ "\", the fixed lead tab label");
}
});
fleet.leaders().forEach((leadName, leader) -> {
if (leader == null) {
return;
@@ -2798,6 +2868,10 @@ public record FleetConfig(
return;
}
String tab = collaborator.tab();
if (tab != null && tab.equalsIgnoreCase(Leader.LEAD_TAB_LABEL)) {
bad.add("fleet.collaborators." + collabName + ".tab=\"" + tab + "\" is the fixed "
+ "lead tab label — a collaborator there would shadow a lead");
}
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
+ "collaborator '" + collabName + "' (\"" + tab + "\")");
@@ -2827,18 +2901,20 @@ public record FleetConfig(
for (int i = 0; i < leadNames.size(); i++) {
String nameA = leadNames.get(i);
Leader a = fleet.leaders().get(nameA);
if (a == null || a.tab() == null || a.tab().isBlank()) {
if (a == null) {
continue;
}
for (int j = i + 1; j < leadNames.size(); j++) {
String nameB = leadNames.get(j);
Leader b = fleet.leaders().get(nameB);
if (b == null || b.tab() == null || b.tab().isBlank()) {
if (b == null) {
continue;
}
if (a.tab().equalsIgnoreCase(b.tab())) {
collisions.add("lead '" + nameA + "' and lead '" + nameB + "' both use tab \""
+ a.tab() + "\"");
if (a.workspace().equalsIgnoreCase(b.workspace())) {
collisions.add("lead '" + nameA + "' and lead '" + nameB + "' share workspace \""
+ a.workspace() + "\" — both would resolve to the tab named \""
+ Leader.LEAD_TAB_LABEL + "\" in that space, so only one could ever be "
+ "found");
}
}
}
@@ -2882,9 +2958,9 @@ public record FleetConfig(
return;
}
throw new IllegalStateException("refusing to start: " + String.join("; ", collisions)
+ ". Tab identity is matched exactly, so only one of two entries sharing a tab can "
+ "ever be found — the other is silently unreachable. Give each lead and "
+ "collaborator its own exact tab.");
+ ". Identity is matched exactly, so only one of two entries sharing a space or a "
+ "tab can ever be found — the other is silently unreachable. Give each lead its "
+ "own space, and each collaborator its own exact tab.");
}
private static boolean templateCanRenderAs(String template, String tab) {
@@ -2913,30 +2989,32 @@ public record FleetConfig(
}
/**
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
* or {@code fleet.collaborators} entry names a {@code tab}. A pane-placed member lands inside
* the focused tab rather than its own, so it can land inside a lead's or collaborator's own
* labelled tab. {@link dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead or collaborator
* purely by that tab's label — it does not exclude the member space — so a member that ends up
* there, while its pane carries no entry in the spawned-member roster, is read back as that
* lead or collaborator and granted that identity's authority.
* Reject a profile that places its members by {@code "pane"} while {@code fleet.leaders} has
* any entry, or any {@code fleet.collaborators} entry names a {@code tab}. A pane-placed
* member lands inside the focused tab rather than its own, so it can land inside a lead's or
* collaborator's own labelled tab. {@link dev.ltms.fleet.herdr.LeadTabScanner} identifies a
* lead or collaborator purely by that tab's label — it does not exclude the member space — so
* a member that ends up there, while its pane carries no entry in the spawned-member roster,
* is read back as that lead or collaborator and granted that identity's authority.
*
* <p>Only an entry with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* <p>Every {@code fleet.leaders} entry is in scope regardless of its own {@code tab} field:
* {@link Leader#acceptedLabels()} always includes {@link Leader#LEAD_TAB_LABEL}. Only a
* collaborator with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
*
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
* {@code fleet.leaders} or {@code fleet.collaborators} entry
* names a non-blank {@code tab}
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while
* {@code fleet.leaders} is non-empty, or any
* {@code fleet.collaborators} entry names a non-blank
* {@code tab}
*/
public void validatePanePlacementAgainstLeadTabs() {
if (fleet == null) {
return;
}
boolean anyLeaderHasTab = fleet.leaders().values().stream()
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
boolean anyLead = !fleet.leaders().isEmpty();
boolean anyCollaboratorHasTab = fleet.collaborators().values().stream()
.anyMatch(c -> c != null && c.tab() != null && !c.tab().isBlank());
if (!anyLeaderHasTab && !anyCollaboratorHasTab) {
if (!anyLead && !anyCollaboratorHasTab) {
return;
}
List<String> bad = new ArrayList<>();
@@ -2953,8 +3031,9 @@ public record FleetConfig(
+ "pane-placed member can land inside that labelled tab, and while its pane "
+ "carries no entry in the spawned-member roster, it is read back as the lead or "
+ "collaborator and granted that identity's authority. Set placement: tab for "
+ "each named profile, or remove the tab from every fleet.leaders and "
+ "fleet.collaborators entry.");
+ "each named profile — the only fix when a lead triggered this, since a lead's "
+ "tab label is fixed regardless of its own tab: field. A collaborator's tab can "
+ "still be removed instead.");
}
/**
@@ -3061,14 +3140,13 @@ public record FleetConfig(
* so duplicates are unrepresentable by construction once loaded — and {@link #load(Path)}
* already rejects a duplicated slot name at parse time, before the map collapses.
*
* <p>Also rejects a {@code fleet.collaborators} entry with no (or a blank) {@code tab}. A
* {@code profile}-less lead is still useful recognise-only — {@code tab} is the only field
* that matters to it either way. A collaborator carries no other field at all, so a blank
* {@code tab} leaves nothing for the entry to mean.
* <p>A lead's {@code profile} is optional — a {@code profile}-less lead is still useful
* recognise-only. Also rejects a {@code fleet.collaborators} entry with no (or a blank)
* {@code tab}: a collaborator carries no other field at all, so a blank {@code tab} leaves
* nothing for the entry to mean.
*
* @throws IllegalStateException when a slot names no profile or an unknown one, when a lead
* can be neither found nor created, or when a collaborator names
* no tab, naming the offending entry
* @throws IllegalStateException when a slot or a lead references an unknown profile, or a
* collaborator names no tab, naming the offending entry
*/
public void validateMembers() {
if (fleet == null) {
@@ -3101,10 +3179,6 @@ public record FleetConfig(
+ "', which is not a configured profiles: entry (have: " + profiles.keySet()
+ ").");
}
if (leader.tab() == null || leader.tab().isBlank()) {
bad.add("fleet.leaders." + name + " has no tab: — a lead is now found (and, if "
+ "auto-launched, labelled) purely by its tab, so every entry must name one.");
}
});
fleet.collaborators().forEach((name, collaborator) -> {
if (collaborator == null) {
@@ -137,6 +137,18 @@ public final class AgentControl {
return result.path("read").path("text").asText("");
}
/**
* Read an agent's terminal with its ANSI styling kept, instead of the stripped text {@link
* #read} returns. Needed when a caller must tell apart text the pane draws dim (a placeholder
* hint) from text drawn plain (the operator's own typing).
*
* @param source one of {@code visible|recent|recent_unwrapped|detection}
*/
public String readWithStyling(String target, String source) {
JsonNode result = agentCall("agent.read", target, Map.of("source", source, "strip_ansi", false));
return result.path("read").path("text").asText("");
}
/** Current agent record (status, session UUID, pane). */
public Agent get(String target) {
return Agent.from(agentCall("agent.get", target, Map.of()).get("agent"));
@@ -25,14 +25,12 @@ import java.util.function.Supplier;
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>CB-579 — matched by name, not prefix.</strong> This used to strip one shared
* {@code tabPrefix} off a label to derive the lead's name, and merged a config-supplied
* {@code terminal_id} pin over every scan result so the pin could never expire. Both are gone: each
* lead now configures its own exact {@code tab} label ({@code fleet.leaders.<name>.tab}), so this
* class is handed a {@code tab → name} map up front and matches labels against it exactly
* (case-insensitively). There is no merge step — a scan result is the whole answer. That is the
* fix for the bug this replaces: a {@code terminal_id} pin surviving in config after the pane it
* named was gone, so the daemon kept treating a dead session as a live lead forever.
* <p><strong>Matched by label within a space, not by a shared prefix.</strong> Each lead's accepted
* labels (the fixed {@code lead} label, plus a deprecated {@code tab} when still configured) are
* matched exactly (case-insensitively) against tabs in that lead's own space only — a tab named
* {@code lead} in one space never resolves to another space's lead. A scan result is the whole
* answer; nothing is merged in from configuration between scans, so a tab that is gone drops out on
* the very next scan instead of lingering forever.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
@@ -103,7 +101,8 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private record Entry(String name, Kind kind) {}
private final HerdrClient herdr;
private final Map<String, Entry> tabToEntry;
private final Map<String, Map<String, String>> leadLabelsBySpace;
private final Map<String, String> collaboratorTabToName;
private final Set<String> excludedWorkspaceLabels;
private final long ttlNanos;
private final LongSupplier clock;
@@ -124,15 +123,17 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabToName every configured lead's exact tab label → its name
* ({@code fleet.leaders.<name>.tab}), matched case-insensitively
* @param leadLabelsBySpace each configured lead's accepted tab labels, keyed by the
* lead's own space label, then by label, to its name — matched
* case-insensitively on both the space and the label. A tab
* matches a lead only within that lead's own space
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
public LeadTabScanner(HerdrClient herdr, Map<String, Map<String, String>> leadLabelsBySpace,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this(herdr, tabToName, Map.of(), excludedWorkspaceLabels, ttlNanos, clock);
this(herdr, leadLabelsBySpace, Map.of(), excludedWorkspaceLabels, ttlNanos, clock);
}
/**
@@ -140,14 +141,15 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
* for configured collaborator tabs in the same pass.
*
* @param collaboratorTabToName every configured collaborator's exact tab label → its name
* ({@code fleet.collaborators.<name>.tab}), matched the same way as
* {@code tabToName}
* ({@code fleet.collaborators.<name>.tab}), matched
* case-insensitively in any space
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
public LeadTabScanner(HerdrClient herdr, Map<String, Map<String, String>> leadLabelsBySpace,
Map<String, String> collaboratorTabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabToEntry = buildTabIndex(tabToName, collaboratorTabToName);
this.leadLabelsBySpace = buildLeadIndex(leadLabelsBySpace);
this.collaboratorTabToName = normalizedLabelMap(collaboratorTabToName);
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.ttlNanos = ttlNanos;
@@ -155,30 +157,46 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
/**
* Keys stripped and lower-cased once, so every lookup is a plain map hit. Leads and
* collaborators merge into a single index, so {@link #scan()} matches both kinds in one pass
* over the tab list; a label naming both a lead and a collaborator takes the lead entry —
* leads are put last, so a colliding key's lead entry is the one that overwrites — since a lead
* can already do everything a collaborator can. Config validation already refuses a lead and a
* collaborator sharing one exact tab, so this ordering is defence in depth, not the control.
* Space and label keys stripped and lower-cased once, so every lookup is a plain map hit. A
* space with no usable labels is simply absent — {@link #leadLabelsFor} then finds nothing for
* it, which is also what a space with a {@code null} label gets.
*/
private static Map<String, Entry> buildTabIndex(Map<String, String> tabToName,
Map<String, String> collaboratorTabToName) {
Map<String, Entry> out = new LinkedHashMap<>();
putNormalized(out, collaboratorTabToName, Kind.COLLABORATOR);
putNormalized(out, tabToName, Kind.LEAD);
private static Map<String, Map<String, String>> buildLeadIndex(
Map<String, Map<String, String>> leadLabelsBySpace) {
Map<String, Map<String, String>> out = new LinkedHashMap<>();
if (leadLabelsBySpace == null) {
return Map.of();
}
leadLabelsBySpace.forEach((space, labelsToName) -> {
if (space == null || space.isBlank()) {
return;
}
Map<String, String> normalized = normalizedLabelMap(labelsToName);
if (!normalized.isEmpty()) {
out.put(space.strip().toLowerCase(Locale.ROOT), normalized);
}
});
return Collections.unmodifiableMap(out);
}
private static void putNormalized(Map<String, Entry> out, Map<String, String> tabToName, Kind kind) {
if (tabToName == null) {
return;
private static Map<String, String> normalizedLabelMap(Map<String, String> labelToName) {
Map<String, String> out = new LinkedHashMap<>();
if (labelToName != null) {
labelToName.forEach((label, name) -> {
if (label != null && !label.isBlank() && name != null && !name.isBlank()) {
out.put(label.strip().toLowerCase(Locale.ROOT), name);
}
});
}
tabToName.forEach((tab, name) -> {
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
out.put(tab.strip().toLowerCase(Locale.ROOT), new Entry(name, kind));
}
});
return Collections.unmodifiableMap(out);
}
/** The accepted lead labels configured for {@code spaceLabel}, or an empty map for no match. */
private Map<String, String> leadLabelsFor(String spaceLabel) {
if (spaceLabel == null) {
return Map.of();
}
return leadLabelsBySpace.getOrDefault(spaceLabel.strip().toLowerCase(Locale.ROOT), Map.of());
}
/**
@@ -241,9 +259,10 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
continue;
}
Map<String, String> leadLabelsHere = leadLabelsFor(ws.label());
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
Entry entry = entryOf(tab.label());
Entry entry = entryOf(tab.label(), leadLabelsHere);
if (entry != null && tab.tabId() != null) {
entryByTab.put(tab.tabId(), entry);
}
@@ -296,19 +315,27 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
/**
* The entry a tab label declares, or {@code null} if it names neither a configured lead nor a
* configured collaborator.
* The entry a tab label declares within one space, or {@code null} if it names neither a lead
* accepted in {@code leadLabelsHere} nor a configured collaborator.
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToEntry} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix. The match strips a trailing
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
* not yet closed keeps resolving normally while that reconcile is pending.
* <p>Exact match (case-insensitive, ends stripped) — no prefix stripping, so an operator's
* {@code "lead: something-else"} tab is never mistaken for a configured lead just because it
* shares a prefix. The match strips a trailing {@link PendingCloseMarker} first, so a tab
* {@code LeadLauncher} has flagged as maybe-dead but not yet closed keeps resolving normally
* while that reconcile is pending. A lead match wins over a collaborator match for the same
* label — a lead can already do everything a collaborator can, and config validation refuses a
* lead and a collaborator sharing one exact tab in the first place.
*/
private Entry entryOf(String label) {
private Entry entryOf(String label, Map<String, String> leadLabelsHere) {
if (label == null) {
return null;
}
return tabToEntry.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
String normalized = PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT);
String leadName = leadLabelsHere.get(normalized);
if (leadName != null) {
return new Entry(leadName, Kind.LEAD);
}
String collaboratorName = collaboratorTabToName.get(normalized);
return collaboratorName == null ? null : new Entry(collaboratorName, Kind.COLLABORATOR);
}
}
@@ -172,6 +172,26 @@ public final class PaneLocator {
return out;
}
/**
* Every workspace ("space") herdr tracks across every searched daemon, keyed by workspace id,
* to its display label — the human-readable name behind {@code fleet_list}'s {@code panes} row,
* next to herdr's own internal {@code workspaceId}. Collapses to one scan in the single-daemon
* deployment, the same as {@link #terminalForPid}. A workspace herdr reports with no label maps
* to a {@code null} value here; a workspace with no {@code workspace_id} is skipped.
*/
public Map<String, String> workspaceLabelsByWorkspaceId() {
Map<String, String> out = new LinkedHashMap<>();
for (HerdrClient herdr : herdrs) {
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace workspace = Workspace.from(w);
if (workspace.workspaceId() != null) {
out.put(workspace.workspaceId(), workspace.label());
}
}
}
return out;
}
/** Whether a pane owns one of the scanned pid's ancestors, or the check of it failed outright. */
private enum Ownership { OWNS, DOES_NOT_OWN, UNKNOWN }
@@ -0,0 +1,250 @@
package dev.ltms.fleet.herdr;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Whether an agent pane's input box is clear for a delivery.
*
* <p>{@link AgentControl#send} pastes its text and submits it in the same call, so a delivery into a
* pane whose input box already holds characters submits those characters too. {@link
* AgentStatus#injectable()} cannot see that: it describes the agent, and an agent waiting at its
* prompt reports the same status whether its box is empty or holds a half-typed line. This reads the
* box itself.
*
* <p>Only a box that is positively empty clears the gate. A box with content, a pane this cannot
* recognise, and a failed read all hold the delivery, because a held delivery is recoverable and a
* submitted half-line is not. Every caller must therefore be a path that retries.
*
* <p>A pane that holds for {@link #HOLD_WARN_STREAK} consecutive checks gets one warning, so a box
* that never clears is visible instead of silent. The warning repeats only after the box has cleared
* again.
*
* <p>The check itself can be switched off (see the two-arg constructor): a disabled gate clears
* every target without reading its pane, logging a hold, or counting a streak.
*/
public final class PromptBox {
private static final Logger log = LoggerFactory.getLogger(PromptBox.class);
/**
* herdr {@code agent.read} source, read with its ANSI styling kept. The input box is always
* drawn here, carrying transcript scrollback above it, which is why only the last box line is
* the live one. Styling must survive the read because the pane draws a placeholder hint — the
* pane's own last submitted prompt — in the same spot as unsubmitted text, dimmed; only the
* escape codes tell the two apart.
*/
static final String PROBE_SOURCE = "visible";
/** Consecutive holds for one target before one warning is logged. */
static final int HOLD_WARN_STREAK = 20;
/**
* Input box markers, each matched only as a line's first characters once any leading ANSI
* escape codes are skipped: the caret the current TUI draws, and the bordered box an older one
* drew. A marker further along a line is transcript text, such as a caret inside something the
* operator quoted.
*/
private static final List<String> BOX_MARKERS = List.of("❯", "│ >");
/** Marker of a turn that is still generating; a box drawn under it is not a settled prompt. */
private static final String ACTIVE_TURN_MARKER = "esc to interrupt";
/** Block glyphs a terminal capture can leave in an otherwise empty box for the cursor cell. */
private static final String CURSOR_GLYPHS = "█▉▊▋▌▍▎▏";
/** An SGR escape sequence, e.g. {@code ESC[2m} (faint) or {@code ESC[0m} (reset). */
private static final Pattern SGR = Pattern.compile("\u001b\\[([0-9;]*)m");
/** The SGR code that dims text — herdr's placeholder hint is drawn inside a span of this. */
private static final String FAINT_CODE = "2";
/** The SGR code (or an empty code list) that clears every attribute, including faint. */
private static final List<String> RESET_CODES = List.of("", "0");
/** What a box holds: nothing, unsubmitted characters, or a pane this cannot read as a box. */
public enum State { EMPTY, DRAFT, UNREADABLE }
/** A box reading: its state, and how many characters it holds ({@code 0} unless {@code DRAFT}). */
public record Reading(State state, int characters) {
}
private final AgentControl agents;
private final boolean gateEnabled;
/** Consecutive holds per target, so a box that never clears can be warned about once. */
private final Map<String, Integer> holdStreaks = new ConcurrentHashMap<>();
/** Equivalent to {@link #PromptBox(AgentControl, boolean)} with the gate on. */
public PromptBox(AgentControl agents) {
this(agents, true);
}
/**
* @param gateEnabled {@code false} makes {@link #clearToSubmit} return {@code true} for every
* target without reading its pane, logging a hold, or counting a streak —
* every delivery proceeds exactly as if the box were always empty.
*/
public PromptBox(AgentControl agents, boolean gateEnabled) {
this.agents = agents;
this.gateEnabled = gateEnabled;
}
/**
* Whether {@code target}'s input box is empty, so a delivery would submit only its own text.
* {@code false} means hold and come back; it never means the delivery failed. Always
* {@code true} when the gate is disabled.
*/
public boolean clearToSubmit(String target) {
if (!gateEnabled) {
return true;
}
Reading reading = inspect(target);
if (reading.state() == State.EMPTY) {
holdStreaks.remove(target);
return true;
}
int streak = holdStreaks.merge(target, 1, Integer::sum);
if (streak == HOLD_WARN_STREAK) {
log.warn("prompt box of {} has held a delivery {} times in a row ({}, {} character(s) in the box)"
+ " — nothing is lost, delivery resumes once the box is empty",
target, streak, reading.state(), reading.characters());
} else {
log.debug("prompt box of {} is {} ({} character(s)), holding delivery {}",
target, reading.state(), reading.characters(), streak);
}
return false;
}
/** Read and classify {@code target}'s pane. A read failure reads as {@link State#UNREADABLE}. */
private Reading inspect(String target) {
String pane;
try {
pane = agents.readWithStyling(target, PROBE_SOURCE);
} catch (RuntimeException e) {
log.debug("prompt box read for {} failed, holding delivery: {}", target, e.getMessage());
return new Reading(State.UNREADABLE, 0);
}
return classify(pane);
}
/**
* Classify a Claude Code TUI pane region. Pure, so it is unit-testable without herdr.
*
* <p>{@link State#EMPTY} needs two positive signals: the pane's last box line holds nothing after
* its marker, and nothing below that line says a turn is still generating. Everything else is
* {@link State#UNREADABLE} — a blank capture, or a region with no box line at all — so a pane this
* does not understand holds the delivery rather than guessing it is safe.
*
* <p>The generating marker is looked for only from the box line down. Above it is scrollback, where
* an earlier turn's marker survives; treating that as a live turn would make {@link State#EMPTY}
* unreachable and hold every delivery forever.
*
* <p>Whitespace, a trailing box border and a cursor block count as nothing. A placeholder hint —
* text the pane draws faint, in the same spot as unsubmitted text — also counts as nothing: only
* a character drawn outside a faint span is the operator's own typing.
*/
static Reading classify(String pane) {
if (pane == null || pane.isBlank()) return new Reading(State.UNREADABLE, 0);
int box = lastBoxLineStart(pane);
if (box < 0) return new Reading(State.UNREADABLE, 0);
String fromBox = pane.substring(box);
if (fromBox.toLowerCase().contains(ACTIVE_TURN_MARKER)) return new Reading(State.UNREADABLE, 0);
String content = boxContent(firstLine(fromBox));
return content.isEmpty() ? new Reading(State.EMPTY, 0) : new Reading(State.DRAFT, content.length());
}
/** Offset of the last line starting with a box marker, or {@code -1} if the region has none. */
private static int lastBoxLineStart(String pane) {
int found = -1;
for (int start = 0; start <= pane.length(); ) {
int end = pane.indexOf('\n', start);
String line = pane.substring(start, end < 0 ? pane.length() : end);
if (markerLength(line) > 0) found = start;
if (end < 0) break;
start = end + 1;
}
return found;
}
/**
* Length of the marker prefix — any leading SGR escape codes, then a box marker — this line
* starts with, or {@code 0} if it starts with neither. The colour drawn on the caret itself
* (e.g. an empty box's grey) sits before the marker glyph, so it must be skipped before the
* marker can match.
*/
private static int markerLength(String line) {
int skip = leadingEscapeLength(line);
for (String marker : BOX_MARKERS) {
if (line.startsWith(marker, skip)) return skip + marker.length();
}
return 0;
}
/** Length of the run of SGR escape codes starting at the beginning of {@code line}. */
private static int leadingEscapeLength(String line) {
Matcher m = SGR.matcher(line);
int pos = 0;
while (m.find(pos) && m.start() == pos) pos = m.end();
return pos;
}
private static String firstLine(String text) {
int newline = text.indexOf('\n');
return newline < 0 ? text : text.substring(0, newline);
}
/** One rendered character of a box line, and whether it was drawn inside a faint (dim) span. */
private record Glyph(char c, boolean faint) {
}
/**
* The text the box holds: its own line after the marker, with border, padding, cursor and any
* faint (placeholder-hint) text left out — only a character drawn outside a faint span is the
* operator's own typing.
*/
private static String boxContent(String boxLine) {
List<Glyph> glyphs = renderedGlyphs(boxLine.substring(markerLength(boxLine)));
int end = glyphs.size();
while (end > 0 && isBoxPadding(glyphs.get(end - 1).c())) end--;
if (end > 0 && glyphs.get(end - 1).c() == '│') end--;
StringBuilder content = new StringBuilder();
for (int i = 0; i < end; i++) {
Glyph glyph = glyphs.get(i);
if (glyph.faint() || isBoxPadding(glyph.c())) continue;
content.append(glyph.c());
}
return content.toString();
}
/** Decode {@code text} into its rendered characters, tracking the faint (SGR 2) span each sits in. */
private static List<Glyph> renderedGlyphs(String text) {
List<Glyph> glyphs = new ArrayList<>();
Matcher m = SGR.matcher(text);
boolean faint = false;
int i = 0;
while (i < text.length()) {
if (m.find(i) && m.start() == i) {
String codes = m.group(1);
if (RESET_CODES.contains(codes)) faint = false;
else if (FAINT_CODE.equals(codes)) faint = true;
i = m.end();
continue;
}
glyphs.add(new Glyph(text.charAt(i), faint));
i++;
}
return glyphs;
}
private static boolean isBoxPadding(char c) {
return Character.isWhitespace(c) || Character.isSpaceChar(c) || CURSOR_GLYPHS.indexOf(c) >= 0;
}
}
@@ -94,6 +94,17 @@ public final class Injector {
*/
private static final int READINESS_GRACE_POLLS = 240;
/**
* How many consecutive polls a message may sit queued with no delivery attempt at all before
* it is failed and the queue cleared — covers every reason the head of the queue is never
* reached, including a target that stays busy ({@code working}) or unclassifiable
* ({@code unknown}) for the whole window. Set well above an ordinary turn so a worker
* genuinely mid-task is never cut off, and below a caller's own overall timeout so a target
* that never frees up fails with this specific reason instead of riding out that longer wait
* silently.
*/
private static final int QUEUE_WAIT_GRACE_POLLS = 4800;
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Fleetd} passes this to every {@link StatusPoller} it constructs,
@@ -121,6 +132,12 @@ public final class Injector {
* already uses for the same purpose.
*/
private final LongSupplier nowMillis;
/**
* Mail offered to panes that collect it themselves. Owned here because this is the single
* writer of delivery state, and the offer must be made and taken back under the same target
* monitor that guards the queue the message is still sitting on.
*/
private final PaneInbox paneInbox;
private final ConcurrentHashMap<String, Target> targets = new ConcurrentHashMap<>();
/** Delivery only; completion signalling is a no-op and every target is treated as available. */
@@ -192,6 +209,7 @@ public final class Injector {
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
this.paneInbox = new PaneInbox(nowMillis);
}
public Injector(HerdrRouter router, TurnListener turnListener, Predicate<String> ready,
@@ -221,6 +239,7 @@ public final class Injector {
this.ready = ready;
this.forget = forget;
this.nowMillis = nowMillis;
this.paneInbox = new PaneInbox(nowMillis);
}
private AgentControl agentsFor(String target) {
@@ -296,13 +315,16 @@ public final class Injector {
final String text;
final TurnToken token;
final CompletableFuture<Void> delivered;
final long enqueuedAtMillis;
volatile State state = State.QUEUED; // written under the owning Target monitor
Pending(String target, String text, TurnToken token, CompletableFuture<Void> delivered) {
Pending(String target, String text, TurnToken token, CompletableFuture<Void> delivered,
long enqueuedAtMillis) {
this.target = target;
this.text = text;
this.token = token;
this.delivered = delivered;
this.enqueuedAtMillis = enqueuedAtMillis;
}
String text() {
@@ -329,10 +351,15 @@ public final class Injector {
int unknownSincePostTurn; // the same, for the post-turn housekeeping phase (fleetd #306)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
long notReadySinceMillis; // wall-clock time of the FIRST non-ready sample in the current notReadySincePoll streak (fleetd #501); reset alongside it
int queueWaitSincePoll; // consecutive polls the queue has held an undelivered message with no attempt made
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
/** The head of {@link #queue} as offered to a mod-served pane, or {@code null}. */
PaneInbox.Entry inboxOffer;
/** Whether the delivery the pickup latch is waiting on was collected rather than typed. */
boolean deliveredViaInbox;
synchronized void add(Pending p) {
queue.add(p);
@@ -349,7 +376,7 @@ public final class Injector {
*/
public Delivery enqueue(String target, String text, TurnToken token) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(target, text, token, delivered);
Pending p = new Pending(target, text, token, delivered, nowMillis.getAsLong());
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
@@ -361,7 +388,8 @@ public final class Injector {
/**
* Cancel this exact queued delivery. The target monitor serializes this operation with
* {@link #onStatus}: if delivery wins that race, this returns {@link Cancellation#DELIVERED}
* rather than claiming the message remained queued.
* rather than claiming the message remained queued. A message a mod-served pane has already
* collected answers the same way, even though no poll has recorded that delivery yet.
*/
public Cancellation cancel(Delivery delivery) {
Pending p = delivery.pending;
@@ -370,7 +398,22 @@ public final class Injector {
return cancellationOf(p);
}
synchronized (t) {
if (p.state != Pending.State.QUEUED || !t.queue.remove(p)) {
if (p.state != Pending.State.QUEUED) {
return cancellationOf(p);
}
if (t.queue.peek() == p && t.inboxOffer != null) {
// This exact entry is the one offered to a mod-served pane. The pane takes an
// offer on its own thread, so withdraw first and then read the outcome: a taken
// offer means the pane already holds this text, and the next poll records that
// delivery. Cancelling it would tell the caller nothing arrived while the pane
// acts on it.
paneInbox.withdrawAll(p.target);
if (t.inboxOffer.taken()) {
return Cancellation.DELIVERED;
}
t.inboxOffer = null;
}
if (!t.queue.remove(p)) {
return cancellationOf(p);
}
p.state = Pending.State.CANCELLED;
@@ -415,6 +458,7 @@ public final class Injector {
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
List<Pending> queueStalled = null; // queued messages failed because the queue never drained
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
@@ -454,10 +498,14 @@ public final class Injector {
t.injectableSincePickup = 0;
t.awaitingCompletion = false;
t.turnObserved = false;
} else {
} else if (!t.deliveredViaInbox) {
// Delivered but still idle → the worker hasn't picked it up; the submit
// keystroke likely raced the paste (esp. right as the TUI became ready).
// Re-nudge Enter (CB-113) until the worker starts (WORKING) or the grace ends.
//
// A pane that collected the message submits it itself, and nothing was
// typed into it. Pressing Enter there would submit whatever its operator
// has in the prompt box instead.
resubmit = true;
}
}
@@ -482,45 +530,77 @@ public final class Injector {
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
t.notReadySinceMillis = 0;
// fleetd #551: poll and record BEFORE the irreversible send, not after.
// The entry comes off the queue and its state is set to ATTEMPTED here,
// unconditionally — so a Throwable escaping the send call below (caught or
// not) can never leave the entry QUEUED at the head of t.queue (the fleetd
// #546 hazard, since peek() alone would let the next onStatus round re-enter
// this block and send the same text again), and no path can write a
// confident DELIVERED or NOT_DELIVERED before we actually know which one
// happened.
t.queue.poll();
p.state = Pending.State.ATTEMPTED;
try {
agentsFor(target).send(target, p.text());
p.state = Pending.State.DELIVERED;
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
sent = p;
} catch (Throwable e) {
// fleetd #551: leave p.state == ATTEMPTED (recorded above, before the
// call) rather than downgrading it to NOT_DELIVERED here — reaching this
// catch does not prove the text never reached the pane. Three of the
// four HerdrException throw sites in HerdrCodec fire only after herdr
// has already replied (so it processed the request), and the fourth (a
// transport IOException) leaves it genuinely unknown whether herdr even
// received the bytes — see #551 comment 16867. The one exception is a
// herdr `*_not_found` error: that family is already read as "definitely
// absent, not merely inconclusive" everywhere else in this codebase
// (StatusPoller, AgentControl's own retry, WorkspaceControl,
// HerdrPeerLauncher, FleetApp, ReplyPushLoop) because it means the
// target pane/agent does not exist at all, so nothing could have been
// pasted anywhere — #551 keeps the new state consistent with that
// existing vocabulary rather than inventing a second one.
if (e instanceof HerdrException he && he.code() != null
&& he.code().endsWith("_not_found")) {
p.state = Pending.State.NOT_DELIVERED;
boolean modServed = paneInbox.isModServed(target);
if (t.inboxOffer != null && !modServed) {
// The pane stopped collecting its mail, so take the offer back and
// fall through to the terminal route below. withdrawAll leaves an
// entry the pane took first alone, so the branch under it still sees
// that as the delivery it is.
paneInbox.withdrawAll(target);
if (!t.inboxOffer.taken()) {
t.inboxOffer = null;
}
}
if (t.inboxOffer == null && modServed) {
t.inboxOffer = paneInbox.offer(target, p.text());
}
if (t.inboxOffer != null) {
// The message stays at the head of the queue until the pane takes
// it: nothing has reached the pane yet, so nothing may be recorded
// as delivered and nothing may be failed.
if (t.inboxOffer.taken()) {
t.inboxOffer = null;
t.queue.poll();
p.state = Pending.State.DELIVERED;
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
t.deliveredViaInbox = true;
sent = p;
}
} else {
// fleetd #551: poll and record BEFORE the irreversible send, not after.
// The entry comes off the queue and its state is set to ATTEMPTED here,
// unconditionally — so a Throwable escaping the send call below (caught or
// not) can never leave the entry QUEUED at the head of t.queue (the fleetd
// #546 hazard, since peek() alone would let the next onStatus round re-enter
// this block and send the same text again), and no path can write a
// confident DELIVERED or NOT_DELIVERED before we actually know which one
// happened.
t.queue.poll();
p.state = Pending.State.ATTEMPTED;
try {
agentsFor(target).send(target, p.text());
p.state = Pending.State.DELIVERED;
t.awaitingPickup = true;
t.awaitingCompletion = true;
t.turnObserved = false;
t.injectableSincePickup = 0;
t.deliveredViaInbox = false;
sent = p;
} catch (Throwable e) {
// fleetd #551: leave p.state == ATTEMPTED (recorded above, before the
// call) rather than downgrading it to NOT_DELIVERED here — reaching this
// catch does not prove the text never reached the pane. Three of the
// four HerdrException throw sites in HerdrCodec fire only after herdr
// has already replied (so it processed the request), and the fourth (a
// transport IOException) leaves it genuinely unknown whether herdr even
// received the bytes — see #551 comment 16867. The one exception is a
// herdr `*_not_found` error: that family is already read as "definitely
// absent, not merely inconclusive" everywhere else in this codebase
// (StatusPoller, AgentControl's own retry, WorkspaceControl,
// HerdrPeerLauncher, FleetApp, ReplyPushLoop) because it means the
// target pane/agent does not exist at all, so nothing could have been
// pasted anywhere — #551 keeps the new state consistent with that
// existing vocabulary rather than inventing a second one.
if (e instanceof HerdrException he && he.code() != null
&& he.code().endsWith("_not_found")) {
p.state = Pending.State.NOT_DELIVERED;
}
sent = p;
sendError = e;
}
sent = p;
sendError = e;
}
} else if (p != null) {
// fleetd #501: stamp the wall-clock time of the FIRST non-ready sample in
@@ -539,6 +619,8 @@ public final class Injector {
for (Pending pending : notReady) {
pending.state = Pending.State.NOT_DELIVERED;
}
paneInbox.withdrawAll(target);
t.inboxOffer = null;
// fleetd #501: t.notReadySincePoll — the loop's own counter, already in
// scope — is printed here instead of the READINESS_GRACE_POLLS constant.
// On this branch the counter has JUST reached the threshold, so the two
@@ -606,6 +688,30 @@ public final class Injector {
}
}
// A message still queued and never attempted this poll is bounded on its own, whatever
// the reason the head of the queue was never reached — a target stuck WORKING or
// UNKNOWN for the whole window hits this even though neither branch above ever looks at
// the queue. Any poll that did attempt the head (`sent != null`, success or failure
// alike) counts as progress and resets the streak, even if messages remain behind it.
if (!t.queue.isEmpty() && sent == null) {
if (++t.queueWaitSincePoll >= QUEUE_WAIT_GRACE_POLLS) {
queueStalled = new ArrayList<>(t.queue);
for (Pending pending : queueStalled) {
pending.state = Pending.State.NOT_DELIVERED;
}
paneInbox.withdrawAll(target);
t.inboxOffer = null;
log.warn("queue for {} never drained after {} polls (limit={} polls/{}s): "
+ "failing {} queued message(s) that were never attempted",
target, t.queueWaitSincePoll, QUEUE_WAIT_GRACE_POLLS,
QUEUE_WAIT_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, queueStalled.size());
t.queue.clear();
t.queueWaitSincePoll = 0;
}
} else {
t.queueWaitSincePoll = 0;
}
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (isQuiescent(t)) {
@@ -676,6 +782,19 @@ public final class Injector {
forget.accept(target);
turnListener.onTurnFailed(target);
}
if (queueStalled != null) {
// The target is not gone — it may still be genuinely busy — so this does not call
// forget.accept: that would clear presence/readiness state for a worker that is
// simply taking a long turn. It still resolves the awaiting send's own waiter via
// onTurnFailed (mirroring notReady above), so a caller learns this specific message
// never reached the pane instead of riding out its own much longer timeout.
RuntimeException cause = new IllegalStateException(
target + " never freed up to receive this message within the queue wait grace");
for (Pending p : queueStalled) {
p.delivered().completeExceptionally(cause);
}
turnListener.onTurnFailed(target, cause.getMessage());
}
if (turnCompleted) {
if (startPostTurn) {
// fleetd #553: the listener call is wrapped so `t.postTurnPending` (set true inside
@@ -795,6 +914,24 @@ public final class Injector {
}
}
/**
* Hand {@code terminal} every message held for it and stamp it as collecting its own mail.
* While that stamp is fresh this injector offers that pane's messages instead of typing them;
* once it goes stale the pane's queued mail takes the terminal route again.
*
* <p>Returns the messages in the order they were queued, and an empty list when there are
* none — an empty collection still counts as collecting, so a pane that polls on a timer stays
* mod-served between messages.
*/
public List<String> collectInbox(String terminal) {
return paneInbox.drain(terminal);
}
/** Whether {@code terminal} has collected its mail recently enough to be offered the next one. */
public boolean isModServed(String terminal) {
return paneInbox.isModServed(terminal);
}
/**
* Targets the poller must keep sampling: those with a queued message, an awaited pickup, or an
* awaited turn completion (so the {@code working → idle} boundary is observed).
@@ -812,6 +949,23 @@ public final class Injector {
.collect(Collectors.toSet());
}
/**
* How long the oldest still-queued, never-attempted message for {@code target} has been
* waiting, or {@code null} when nothing is queued (including when the head has already been
* attempted or delivered). A caller uses this to tell a message that genuinely never reached
* the pane apart from one that was delivered and is now simply being worked on.
*/
public Long queuedWaitMillis(String target) {
Target t = targets.get(target);
if (t == null) {
return null;
}
synchronized (t) {
Pending head = t.queue.peek();
return head != null ? nowMillis.getAsLong() - head.enqueuedAtMillis : null;
}
}
/**
* Forget a target whose worker is gone, failing every still-queued message so awaiting callers
* unblock instead of hanging forever. If a message had already been <em>delivered</em> but its
@@ -831,6 +985,11 @@ public final class Injector {
p.state = Pending.State.NOT_DELIVERED;
}
t.queue.clear();
// The pane is gone, so drop its offered mail and its poll stamp together: a terminal
// id can be reused, and a stale stamp would make the next pane under it look mod-served
// before it has ever collected anything.
paneInbox.forget(target);
t.inboxOffer = null;
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
@@ -5,16 +5,16 @@ import java.util.Set;
/**
* Tracks which peers are <em>available</em> — their Claude has booted and connected its MCP
* client to the bridge (CB-113). For a spawned member this is the reliable readiness signal,
* client to the bridge. For a spawned member this is the reliable readiness signal,
* unlike herdr's {@code agent_status}, which reports {@code idle} while its Claude is still
* booting. Delivering into that boot window pastes into a not-yet-ready TUI (the text is lost)
* and wedges that member's delivery state, so the {@link Injector} holds a spawned member's
* first delivery until it is present here.
*
* <p>Populated from the MCP transport, for the peers whose deliverability depends on a proven
* live MCP contact; {@code FleetMcp.markTrackedCallerPresent} decides which peers those are. A
* peer that never mounts the bridge MCP is never marked present — its sends stay queued until
* they time out, which is correct (it could not have replied anyway).
* <p>Populated from the MCP transport, for the peers whose deliverability rests on proving a live
* MCP contact rather than on a configured registry entry. A peer that never mounts the bridge MCP
* is never marked present — its sends stay queued until they time out, which is correct (it could
* not have replied anyway).
*/
public class MemberPresence {
@@ -0,0 +1,141 @@
package dev.ltms.fleet.inject;
import java.util.ArrayDeque;
import java.util.ArrayList;
import java.util.Deque;
import java.util.List;
import java.util.Objects;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
/**
* Mail held for a pane that collects it itself instead of having it typed into its terminal.
*
* <p>A pane becomes <em>mod-served</em> by calling {@code fleet_inbox}: {@link #drain} stamps the
* pane as polling, and {@link #isModServed} answers {@code true} while that stamp is younger than
* {@link #MOD_SERVED_WINDOW_MILLIS}. Nothing else sets it, so a pane that has never polled is
* never mod-served and its mail takes the terminal route.
*
* <p>An offered entry is removed exactly once, by {@link #drain} or by {@link #withdrawAll}, and
* both run under the owning pane's monitor. So an entry the pane took is never also withdrawn, and
* an entry that was withdrawn can never still be collected — which is what lets the {@link
* Injector} keep one message both offered here and queued for the terminal without risking two
* deliveries of it.
*
* <p>This class holds no queue of its own beyond what is currently offered: the {@link Injector}
* keeps the message on its own queue until the pane takes it, so a pane that stops polling strands
* nothing.
*/
public class PaneInbox {
/**
* How long after a {@link #drain} a pane still counts as mod-served. It must exceed the mod's
* own poll interval by enough that a few missed polls are not read as a pane that stopped,
* while staying short enough that a pane which really stopped falls back to the terminal route
* promptly.
*/
public static final long MOD_SERVED_WINDOW_MILLIS = 15_000;
/** One message held for a pane until that pane collects it. */
public static final class Entry {
private final String text;
private boolean taken; // written under the owning pane's monitor
private Entry(String text) {
this.text = text;
}
/** The message text, as it will be handed to the pane. */
public String text() {
return text;
}
/** Whether the pane has collected this entry. Once {@code true} it never goes back. */
public synchronized boolean taken() {
return taken;
}
private synchronized void markTaken() {
taken = true;
}
}
private final LongSupplier nowMillis;
private final ConcurrentHashMap<String, Long> lastPolledAtMillis = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Deque<Entry>> offered = new ConcurrentHashMap<>();
public PaneInbox() {
this(System::currentTimeMillis);
}
public PaneInbox(LongSupplier nowMillis) {
this.nowMillis = Objects.requireNonNull(nowMillis, "nowMillis");
}
/**
* Whether {@code terminal} collected its mail within {@link #MOD_SERVED_WINDOW_MILLIS}. A
* terminal that has never collected any is never mod-served.
*/
public boolean isModServed(String terminal) {
Long at = terminal == null ? null : lastPolledAtMillis.get(terminal);
return at != null && nowMillis.getAsLong() - at <= MOD_SERVED_WINDOW_MILLIS;
}
/** Hold {@code text} for {@code terminal} to collect, and return the entry holding it. */
public Entry offer(String terminal, String text) {
Entry entry = new Entry(text);
Deque<Entry> queue = offered.computeIfAbsent(terminal, _ -> new ArrayDeque<>());
synchronized (queue) {
queue.add(entry);
}
return entry;
}
/**
* Take back every entry {@code terminal} has not collected. An entry the pane took first stays
* taken — this never un-delivers one.
*/
public void withdrawAll(String terminal) {
Deque<Entry> queue = terminal == null ? null : offered.get(terminal);
if (queue == null) {
return;
}
synchronized (queue) {
queue.removeIf(e -> !e.taken());
}
}
/**
* Collect everything held for {@code terminal}, in the order it was offered, and stamp the pane
* as polling. Each returned entry is marked taken, so the {@link Injector} can tell a message
* the pane really has from one it merely offered.
*/
public List<String> drain(String terminal) {
if (terminal == null || terminal.isBlank()) {
return List.of();
}
lastPolledAtMillis.put(terminal, nowMillis.getAsLong());
Deque<Entry> queue = offered.get(terminal);
if (queue == null) {
return List.of();
}
List<String> collected = new ArrayList<>();
synchronized (queue) {
for (Entry e : queue) {
e.markTaken();
collected.add(e.text());
}
queue.clear();
}
return collected;
}
/** Forget a pane that is gone, so neither its poll stamp nor its offered mail lingers. */
public void forget(String terminal) {
if (terminal == null) {
return;
}
lastPolledAtMillis.remove(terminal);
offered.remove(terminal);
}
}
@@ -18,6 +18,7 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
@@ -297,15 +298,11 @@ public final class LeadLauncher {
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
* tab carries a different label, not because any workspace is excluded from this count.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
* <em>not</em> live: a running agent in a tab, in that lead's own space, carrying one of its
* {@link FleetConfig.Leader#acceptedLabels()}. A member sitting in the same shared workspace is
* not counted as a lead because its tab carries a different label, not because any workspace is
* excluded from this count. A tab matching a lead's label in a <em>different</em> space is not
* counted either — space is the uniqueness boundary between leads.
*
* <p>fleetd #359 review finding 1: a labelled tab with nothing running in it is split into
* {@code toClose} (already flagged pending-close by a previous reconcile, and still dead — two
@@ -319,12 +316,11 @@ public final class LeadLauncher {
}
private Map<String, LeadCount> countLeads(Map<String, FleetConfig.Leader> leaders) {
// A lead and the members share ONE workspace now (the operator asked for a single "session"
// with many tabs), so a workspace can no longer be excluded wholesale — the lead lives in the
// member workspace by design. The sole discriminator is the exact tab label: a lead carries
// its configured `fleet.leaders.<name>.tab` ("lead: opus"), while a member carries its
// profile's `worker: {profile} #{n}` template. These never collide, so an exact-label match
// separates them without needing to know which workspace anyone is in.
// A lead and the members share ONE workspace (the operator asked for a single "session" with
// many tabs), so a workspace can no longer be excluded wholesale — the lead lives in the
// member workspace by design. The discriminator is the tab label together with the space: a
// member's tab never carries one of a lead's accepted labels, and a lead's own label only
// counts within that lead's configured space.
Map<String, String> nameByTab = new LinkedHashMap<>();
Set<String> flaggedTabIds = new LinkedHashSet<>();
for (Workspace ws : spaces.listWorkspaces()) {
@@ -332,7 +328,7 @@ public final class LeadLauncher {
continue;
}
for (Tab tab : spaces.listTabs(ws.workspaceId())) {
String declared = leadNameOf(tab.label(), leaders);
String declared = leadNameOf(tab.label(), ws.label(), leaders);
if (declared != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), declared);
if (PendingCloseMarker.isFlagged(tab.label())) {
@@ -382,21 +378,24 @@ public final class LeadLauncher {
}
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
* The configured lead a tab names, or {@code null} for a label or space that names none.
*
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead. A
* trailing {@link PendingCloseMarker} is stripped first, so a tab this class flagged on a
* previous reconcile is still recognised as the same lead's tab on this one.
* <p>A match requires both: the label (case-insensitively, trailing {@link PendingCloseMarker}
* stripped) must be one of the lead's {@link FleetConfig.Leader#acceptedLabels()}, and {@code
* space} must be that lead's own {@link FleetConfig.Leader#workspace()}. The same label in a
* different space names no lead — space is the uniqueness boundary between leads.
*/
private String leadNameOf(String label, Map<String, FleetConfig.Leader> leaders) {
if (label == null) {
private String leadNameOf(String label, String space, Map<String, FleetConfig.Leader> leaders) {
if (label == null || space == null) {
return null;
}
String l = PendingCloseMarker.strip(label);
String l = PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT);
for (Map.Entry<String, FleetConfig.Leader> e : leaders.entrySet()) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
FleetConfig.Leader lead = e.getValue();
if (lead == null || !lead.workspace().equalsIgnoreCase(space)) {
continue;
}
if (lead.acceptedLabels().contains(l)) {
return e.getKey();
}
}
@@ -22,6 +22,7 @@ import java.util.concurrent.TimeUnit;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Predicate;
import java.util.function.Supplier;
/**
@@ -66,16 +67,24 @@ import java.util.function.Supplier;
* <li>launch a fresh lead with {@code LeadLauncher#relaunch}. <strong>If every attempt fails,
* {@code bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_FAILED}.</li>
* <li>wait for the fresh pane to reach a real turn boundary ({@code IDLE} or {@code DONE},
* never merely {@code BLOCKED}), bounded by {@code relaunchReadySeconds}. This is the
* safety gate: typing into a pane that has not actually finished booting loses the
* keystrokes. <strong>If the pane never becomes ready, {@code bootstrapText} is never
* sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
* never merely {@code BLOCKED}) AND to be present in {@link
* dev.ltms.fleet.inject.MemberPresence} — its Claude has actually connected the bridge MCP
* client — bounded by {@code relaunchReadySeconds}. This is the safety gate: herdr's own
* {@code agent_status} reports {@code idle} through the whole boot window, so the turn
* boundary alone cannot tell a booted pane from one still starting up, and typing into that
* window loses the keystrokes. <strong>If the pane never becomes ready, {@code
* bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
* <li>wait for the fresh terminal to be recognised as a live lead — present in the live-lead
* terminal map — bounded by {@code relaunchReadySeconds}. This is bookkeeping, not a
* safety gate: {@code bootstrapText} is sent either way once the pane is ready, whether or
* not this wait itself times out — see {@link RollState#RELAUNCH_NOT_RECOGNISED}.</li>
* <li>{@code agents.send(newTerminal, cfg.bootstrapTextFor(p.handoverPath()))} — sent to the
* FRESH terminal, never the one that was just torn down.</li>
* <li>wait for the fresh pane to enter a live turn — the successor actually starting a turn on
* the sent text — bounded by {@link #BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS}. A send that herdr
* accepts is not proof of delivery; only the successor starting a turn is.
* <strong>If this never happens, the roll is reported as unconfirmed and nothing is
* re-sent</strong> — see {@link RollState#BOOTSTRAP_NOT_CONFIRMED}.</li>
* </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the lead has already been replaced"</em> — the
@@ -138,6 +147,12 @@ public final class LeadRollover {
* CLI boot is.
*/
static final int PANE_DEATH_TIMEOUT_SECONDS = 10;
/**
* Bound on the post-send wait for the successor to start a turn. Far smaller than {@code
* relaunchReadySeconds}, which sizes a pane booting a CLI; this one sizes an agent that already
* holds the text reacting to it, which takes one status poll.
*/
static final int BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS = 5;
/**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
@@ -232,7 +247,7 @@ public final class LeadRollover {
/**
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
* five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* that has been approved but has not finished yet, and two answers for a token that names no
* active work at all: still pending confirmation, or nothing known about this token at all.
*/
@@ -256,7 +271,8 @@ public final class LeadRollover {
* that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
* #RELAUNCH_NEVER_READY}, {@link #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it
* #RELAUNCH_NEVER_READY}, {@link #RELAUNCH_NOT_RECOGNISED}, {@link #BOOTSTRAP_NEVER_SENT},
* {@link #BOOTSTRAP_NOT_CONFIRMED}, or {@link #FAILED}) once it
* finishes — including by throwing, which {@link #runRollover}'s catch turns into {@link
* #FAILED} instead of leaving this entry stuck forever.
*/
@@ -305,6 +321,20 @@ public final class LeadRollover {
* an operator should check why the tab was not recognised.
*/
RELAUNCH_NOT_RECOGNISED,
/**
* A fresh lead was ready, but herdr kept reporting {@code agent_not_ready} while this class
* retried {@code bootstrapText} for {@code relaunchReadySeconds}. The fresh session did not
* receive its handover instruction.
*/
BOOTSTRAP_NEVER_SENT,
/**
* {@code bootstrapText} was sent and herdr accepted it, but the fresh pane's status never
* entered a live turn within {@code BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS} — the successor never
* started a turn on it. A send herdr accepts is not proof the text landed in the TUI;
* nothing is re-sent, because a second paste into a session that did receive the first
* would be worse than one unconfirmed roll.
*/
BOOTSTRAP_NOT_CONFIRMED,
/**
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
@@ -362,6 +392,13 @@ public final class LeadRollover {
* learn when that terminal has been recognised as a live lead — see this class's javadoc.
*/
private final Supplier<Map<String, String>> liveLeadTerminals;
/**
* Whether a terminal's Claude has connected the bridge MCP client — normally {@code
* MemberPresence::isPresent}. {@link #waitUntilPaneReady} requires this alongside a turn
* boundary: herdr's own {@code agent_status} reports {@code idle} through the whole boot
* window, so the turn boundary alone cannot tell a booted pane from one still starting up.
*/
private final Predicate<String> mcpPresent;
private final LongSupplier nowMillis;
private final Runnable pollSleeper;
/**
@@ -404,9 +441,10 @@ public final class LeadRollover {
Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals) {
Supplier<Map<String, String>> liveLeadTerminals,
Predicate<String> mcpPresent) {
this(agents, spaces, launcher, configSupplier, leadWorkspace, leadNameForTerminal,
liveLeadTerminals, System::currentTimeMillis,
liveLeadTerminals, mcpPresent, System::currentTimeMillis,
() -> sleepUninterruptibly(POLL_INTERVAL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
}
@@ -423,6 +461,7 @@ public final class LeadRollover {
Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals,
Predicate<String> mcpPresent,
LongSupplier nowMillis, Runnable pollSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents;
this.spaces = spaces;
@@ -431,6 +470,7 @@ public final class LeadRollover {
this.leadWorkspace = leadWorkspace;
this.leadNameForTerminal = leadNameForTerminal;
this.liveLeadTerminals = liveLeadTerminals;
this.mcpPresent = mcpPresent;
this.nowMillis = nowMillis;
this.pollSleeper = pollSleeper;
this.continuationRunner = continuationRunner;
@@ -715,23 +755,37 @@ public final class LeadRollover {
ReadinessResult readinessResult = waitUntilPaneReady(newAgent.terminalId(),
cfg.relaunchReadySeconds());
if (!readinessResult.ready()) {
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) never reached a real "
+ "turn boundary — bootstrapText was never sent (token={}, configured={}s "
+ "elapsed={}ms)",
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) was not ready — "
+ "bootstrapText was never sent (token={}, configured={}s elapsed={}ms, "
+ "lastStatus={} mcpConnected={})",
leadName, newAgent.terminalId(), p.token(), cfg.relaunchReadySeconds(),
readinessResult.elapsedMillis());
readinessResult.elapsedMillis(), readinessResult.lastStatus(),
readinessResult.mcpSeen());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NEVER_READY,
"fresh terminal " + newAgent.terminalId() + " never reached a real turn "
+ "boundary (IDLE or DONE) within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ readinessResult.elapsedMillis() + "ms) — bootstrapText was never "
+ "sent"));
"fresh terminal " + newAgent.terminalId() + " was not ready within "
+ "relaunchReadySeconds=" + cfg.relaunchReadySeconds() + "s (measured "
+ "elapsed=" + readinessResult.elapsedMillis() + "ms): last status="
+ readinessResult.lastStatus() + " (a real turn boundary is IDLE or "
+ "DONE), bridge MCP connected=" + readinessResult.mcpSeen()
+ " — bootstrapText was never sent"));
return;
}
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
cfg.relaunchReadySeconds());
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
BootstrapResult bootstrapResult = sendBootstrapWithRetry(newAgent.terminalId(),
cfg.bootstrapTextFor(p.handoverPath()), cfg.relaunchReadySeconds());
if (!bootstrapResult.sent()) {
log.warn("lead-rollover: bootstrapText was never sent to fresh terminal {} for lead '{}' "
+ "after agent_not_ready persisted for {}ms (token={}, configured={}s)",
newAgent.terminalId(), leadName, bootstrapResult.elapsedMillis(), p.token(),
cfg.relaunchReadySeconds());
outcomes.put(p.token(), new RollStatus(RollState.BOOTSTRAP_NEVER_SENT,
"fresh terminal " + newAgent.terminalId() + " kept rejecting bootstrapText with "
+ "agent_not_ready for relaunchReadySeconds=" + cfg.relaunchReadySeconds()
+ "s (measured elapsed=" + bootstrapResult.elapsedMillis() + "ms)"));
return;
}
if (!identityResult.ready()) {
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
+ "was never recognised as a live lead — an operator should check why "
@@ -747,6 +801,25 @@ public final class LeadRollover {
return;
}
ConfirmationResult confirmationResult = waitUntilTurnStarted(newAgent.terminalId(),
BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS);
if (!confirmationResult.started()) {
log.warn("lead-rollover: bootstrapText was sent to fresh terminal {} for lead '{}', but "
+ "it never started a turn on it — the send is unconfirmed; nothing was "
+ "re-sent (token={}, configured={}s elapsed={}ms)",
newAgent.terminalId(), leadName, p.token(), BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS,
confirmationResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.BOOTSTRAP_NOT_CONFIRMED,
"bootstrapText was sent to fresh terminal " + newAgent.terminalId() + ", but "
+ "it was never seen in a live turn within "
+ BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS + "s (measured elapsed="
+ confirmationResult.elapsedMillis() + "ms) — the send is unconfirmed "
+ "and nothing was re-sent. Every other step of this roll succeeded, "
+ "and a turn shorter than one status poll can be missed, so confirm "
+ "against the successor itself before treating the text as lost"));
return;
}
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} oldLead={} newTerminal={} elapsedMs={}",
p.token(), lead, newAgent.terminalId(), rollElapsedMillis);
@@ -755,6 +828,30 @@ public final class LeadRollover {
+ newAgent.terminalId()));
}
/**
* Sends {@code bootstrapText}, retrying only the transient herdr {@code agent_not_ready} refusal
* until {@code readySeconds} elapses. All other failures propagate to {@link #runRollover}.
*/
private BootstrapResult sendBootstrapWithRetry(String terminal, String bootstrapText, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
try {
agents.send(terminal, bootstrapText);
return new BootstrapResult(true, nowMillis.getAsLong() - startMillis);
} catch (HerdrException e) {
if (!"agent_not_ready".equals(e.code())) {
throw e;
}
pollSleeper.run();
}
}
return new BootstrapResult(false, nowMillis.getAsLong() - startMillis);
}
/** The result of {@link #sendBootstrapWithRetry}. */
private record BootstrapResult(boolean sent, long elapsedMillis) {}
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
static final int CAPTURE_RETRIES = 3;
@@ -848,13 +945,19 @@ public final class LeadRollover {
* Poll until {@code newTerminal}'s own pane reaches a real turn boundary ({@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE}, never merely {@link AgentStatus#BLOCKED}) —
* the same exclusion {@link #waitUntilAtTurnBoundary} applies to the calling lead's own turn,
* applied here to the fresh one, so {@code bootstrapText} is never typed into a pane that has
* not actually finished booting — or {@code readySeconds} elapses. A failed status read
* degrades to "not yet ready" and is retried on the next poll.
* applied here to the fresh one — AND {@code mcpPresent} reports that terminal's Claude has
* connected the bridge MCP client, or {@code readySeconds} elapses. Herdr's own {@code
* agent_status} reports {@code idle} through the whole boot window (see {@code
* MemberPresence}'s class javadoc), so the turn-boundary check alone cannot tell a booted pane
* from one still starting up; requiring both signals is what makes {@code bootstrapText} never
* typed into a pane that has not actually finished booting. A failed status read degrades to
* "not yet ready" and is retried on the next poll.
*/
private ReadinessResult waitUntilPaneReady(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
AgentStatus lastStatus = null;
boolean mcpSeen = false;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
@@ -864,16 +967,58 @@ public final class LeadRollover {
newTerminal, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis);
boolean atTurnBoundary = status == AgentStatus.IDLE || status == AgentStatus.DONE;
mcpSeen = mcpPresent.test(newTerminal);
lastStatus = status;
if (atTurnBoundary && mcpSeen) {
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis, status, true);
}
pollSleeper.run();
}
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis);
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis, lastStatus, mcpSeen);
}
/** The measured outcome of {@link #waitUntilPaneReady}. */
private record ReadinessResult(boolean ready, long elapsedMillis) {}
/**
* The measured outcome of {@link #waitUntilPaneReady}. {@code lastStatus} is the last status read
* ({@code null} when every read failed) and {@code mcpSeen} the last MCP presence answer, so a
* caller reporting a timeout can name which of the two signals was missing.
*/
private record ReadinessResult(boolean ready, long elapsedMillis, AgentStatus lastStatus,
boolean mcpSeen) {}
/**
* Poll until {@code newTerminal} is in a live turn ({@link AgentStatus#WORKING} or {@link
* AgentStatus#BLOCKED}, a turn that is paused rather than finished) — the successor acting on
* the {@code bootstrapText} just sent — or {@code readySeconds} elapses. Herdr accepting the
* send is not proof the text landed in the TUI; only the successor acting on it is. A failed
* status read degrades to "not yet confirmed" and is retried on the next poll.
*
* <p>The status is polled, so a turn that starts and ends inside one poll interval is missed
* and reads as unconfirmed. That is why the caller only reports and never re-sends: a false
* alarm costs one check, a second paste into a session that did receive the first does not.</p>
*/
private ConfirmationResult waitUntilTurnStarted(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(newTerminal);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while confirming {} started its turn: {}",
newTerminal, e.toString());
status = null;
}
if (status == AgentStatus.WORKING || status == AgentStatus.BLOCKED) {
return new ConfirmationResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new ConfirmationResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilTurnStarted}. */
private record ConfirmationResult(boolean started, long elapsedMillis) {}
/**
* Poll until {@code newTerminal} is present in {@link #liveLeadTerminals} or {@code
@@ -46,9 +46,10 @@ public final class ConnectionIdentity {
}
/**
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
* the pane scan behind {@code terminal} ran to completion ({@link #scanComplete}).
* The caller resolved from the connection: the {@code terminal} of the pane it connects from
* (or {@code null} when the connection maps to no pane), its {@code pid} (or {@code -1} if not
* resolvable), and whether the pane scan behind {@code terminal} ran to completion
* ({@link #scanComplete}).
*/
public record Caller(String terminal, long pid, boolean scanComplete) {
@@ -8,6 +8,7 @@ import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.auth.Role;
import dev.ltms.fleet.guard.GuardException;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.inject.MemberPresence;
@@ -120,9 +121,9 @@ public final class FleetMcp {
/**
* Kept as a field (rather than only captured by the {@code contextExtractor} closure) so
* {@link #denyFor} can read {@link CallerResolver#knownLeadOrCollaborator()} and {@link
* CallerResolver#sendableObserverTarget()} — the classifiers a collaborator's and an
* observer's {@code SEND} are each checked against, built from the same maps {@link #identity}-
* based resolution reads.
* CallerResolver#observerSendTarget()} — the classifiers a collaborator's and an observer's
* {@code SEND} are each checked against, built from the same maps {@link #identity}-based
* resolution reads.
*/
private final CallerResolver callers;
private final Metrics metrics; // CB-502: null → auth failures not counted
@@ -308,17 +309,30 @@ public final class FleetMcp {
/**
* Pane-discovery facts for {@code fleet_list}'s {@code panes} row — every herdr tab's display
* label, and the one deliverability gate the status-gated injector itself reads.
* label, the one deliverability gate the status-gated injector itself reads, whether a
* terminal is bound to a configured architect slot, and whether an observer caller may
* {@code SEND} to it.
*
* @param tabLabels tab id → its display label, read lazily (only once the row is actually
* assembled) since it costs a herdr {@code workspace.list}/{@code tab.list}
* scan; a tab herdr reports with no label maps to a {@code null} value
* @param workspaceLabels workspace id → its display label (the herdr "space" name), read lazily
* the same way as {@code tabLabels}; a workspace herdr reports with no label,
* or one the lookup cannot find, maps to a {@code null} value
* @param deliverable the same gate {@link dev.ltms.fleet.Fleetd#deliverableTo} builds for the
* injector, keyed by terminal id — never a second, separately-derived check
* @param architectSlot the same classifier {@link CallerResolver#boundToArchitectSlot} resolves
* a caller against — never a second, separately-derived check
* @param observerSendTarget the same predicate {@link CallerResolver#observerSendTarget}
* builds for the {@code SEND} gate — never a second, separately-derived
* check
*/
public record PaneSource(Supplier<Map<String, String>> tabLabels, Predicate<String> deliverable) {
/** Inert source — no labels, and every target reports non-deliverable. */
public static PaneSource none() { return new PaneSource(Map::of, _ -> false); }
public record PaneSource(Supplier<Map<String, String>> tabLabels, Supplier<Map<String, String>> workspaceLabels,
Predicate<String> deliverable, Predicate<String> architectSlot, Predicate<String> observerSendTarget) {
/** Inert source — no labels, no deliverable targets, no architect slots, nothing sendable. */
public static PaneSource none() {
return new PaneSource(Map::of, Map::of, _ -> false, _ -> false, _ -> false);
}
}
/**
@@ -502,7 +516,10 @@ public final class FleetMcp {
// Answering a worker's fleet_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a), callerOwner);
// ANSWER is primary-or-architect, and an architect is refused DRAIN too, so a
// timeout receipt must not name fleet_poll{target} to an architect caller.
boolean mayDrainPoll = Authz.permits(caller, Authz.Action.DRAIN, target);
return answer(messages, turnId, content, timeoutMs(a), callerOwner, mayDrainPoll);
}
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
@@ -514,15 +531,22 @@ public final class FleetMcp {
// name would never match there anyway, but passing null for them keeps the intent
// explicit rather than relying on that lookup to filter it out.
String delegatorName = caller.isPrimary() ? caller.name() : null;
Runnable onAccepted = () ->
primaryRegistry.recordDelegation(target, callerTerminal, delegatorName);
// A nudge about this target must never tell its recipient to run a call Authz
// would refuse it — asked once, here, while the real Principal is still in
// scope, rather than re-derived from a bare terminal string later.
boolean mayDrainNudge = Authz.permits(caller, Authz.Action.DRAIN, target);
boolean mayAnswerNudge = Authz.permits(caller, Authz.Action.ANSWER, target);
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, callerTerminal,
delegatorName, mayDrainNudge, mayAnswerNudge);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, target, content, onAccepted, workers.profiles(), caller)
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles(), callerOwner);
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles(), callerOwner,
mayDrainNudge);
};
// fleet_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
// fleet_reply's identity is the CONNECTION, never an argument. The authz check is
// terminal ownership, not a role test: the caller may reply only for its own pane,
// which is why no role appears in the check at all.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> replyHandler =
(exchange, req) -> {
String self = callerTerminal(exchange);
@@ -538,6 +562,16 @@ public final class FleetMcp {
if (denied != null) return denied;
return ask(messages, self, str(req.arguments(), "question"), timeoutMs(req.arguments()));
};
// fleet_inbox: the caller collects the mail queued for its OWN pane. Identity is the
// CONNECTION, never an argument, and the authz check is terminal ownership — the same
// shape as fleet_reply above.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> inboxHandler =
(exchange, req) -> {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_inbox", req.arguments()), self);
if (denied != null) return denied;
return inbox(messages, self);
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> statusHandler =
(exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_status", req.arguments()), null);
@@ -549,10 +583,19 @@ public final class FleetMcp {
Map<String, Object> a = req.arguments();
String target = str(a, "target");
String coordId = str(a, "coordId");
String ticket = str(a, "ticket");
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
Authz.Action action = toolAction("fleet_poll", a);
// TASK_READ reads a ticket, not a DRAIN target -- the gate must see the
// ticket id, and an observer's or a collaborator's grant is confined to a
// ticket it created.
String authzTarget = action == Authz.Action.TASK_READ ? ticket : target;
Predicate<String> ticketOwnedByCaller = action == Authz.Action.TASK_READ
? t -> messages.ownsTicket(t, principal(exchange).ownerKey())
: Authz.NO_OWNED_TICKET;
McpSchema.CallToolResult denied = deny(exchange, action, authzTarget, ticketOwnedByCaller);
if (denied != null) return denied;
return poll(messages, leadChannel, str(a, "ticket"), target, coordId,
return poll(messages, leadChannel, ticket, target, coordId,
principal(exchange).ownerKey());
};
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
@@ -589,7 +632,9 @@ public final class FleetMcp {
// A Supplier: the label lookup costs a herdr scan, and must stay behind
// panesVisible so it only runs for a caller that receives the row at all.
PaneSource panes = new PaneSource(() -> identity.panes().tabLabelsByTabId(),
Fleetd.deliverableTo(presence, callers::leads, callers::collaborators));
() -> identity.panes().workspaceLabelsByWorkspaceId(),
Fleetd.deliverableTo(presence, callers::leads, callers::collaborators),
callers::boundToArchitectSlot, callers.observerSendTarget());
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
callerTerminal(exchange),
@@ -597,7 +642,7 @@ public final class FleetMcp {
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)),
leadsVisibleTo(principal(exchange)), membersVisibleTo(principal(exchange)),
panes, panesVisibleTo(principal(exchange)));
panes, panesVisibleTo(principal(exchange)), principal(exchange).isObserver());
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
(exchange, req) -> {
@@ -641,6 +686,7 @@ public final class FleetMcp {
McpSchema.Tool fleetProfiles = profilesTool();
McpSchema.Tool fleetWhoami = whoamiTool();
McpSchema.Tool fleetHandover = handoverTool();
McpSchema.Tool fleetInbox = inboxTool();
// fleetd #469: the tool schemas above are already named from FleetTool.wireName(), but
// this is the check that a schema was not accidentally dropped, duplicated, or added
@@ -651,7 +697,7 @@ public final class FleetMcp {
Set<String> registeredToolNames = Set.of(fleetSend.name(), fleetReply.name(), fleetAsk.name(),
fleetStatus.name(), fleetPoll.name(), fleetAck.name(), fleetSpawn.name(),
fleetList.name(), fleetStop.name(), fleetProfiles.name(), fleetWhoami.name(),
fleetHandover.name());
fleetHandover.name(), fleetInbox.name());
if (!registeredToolNames.equals(FleetTool.wireNames())) {
throw new IllegalStateException("fleetd #469: registered MCP tools " + registeredToolNames
+ " do not match the canonical tool set " + FleetTool.wireNames()
@@ -673,6 +719,7 @@ public final class FleetMcp {
.toolCall(fleetProfiles, profilesHandler)
.toolCall(fleetWhoami, whoamiHandler)
.toolCall(fleetHandover, handoverHandler)
.toolCall(fleetInbox, inboxHandler)
.build();
this.metrics = metrics;
}
@@ -716,6 +763,16 @@ public final class FleetMcp {
return denyFor(principal(exchange), action, target);
}
/**
* As {@link #deny(McpSyncServerExchange, Authz.Action, String)}, with a real classifier for
* an observer's or a collaborator's {@code TASK_READ} — the only call site that can supply
* one is a ticket poll, which knows the ticket id and the caller's owner key.
*/
private McpSchema.CallToolResult deny(McpSyncServerExchange exchange, Authz.Action action,
String target, Predicate<String> ticketOwnedByCaller) {
return denyFor(principal(exchange), action, target, ticketOwnedByCaller);
}
/**
* The policy half of {@link #deny}: everything except pulling the caller out of the MCP
* exchange. Kept separate so the authorization decision — the actual control — is unit-testable
@@ -729,13 +786,22 @@ public final class FleetMcp {
* @return {@code null} when the call may proceed, or the error result to return when it may not
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target) {
return denyFor(caller, action, target, Authz.NO_OWNED_TICKET);
}
/**
* As {@link #denyFor(Principal, Authz.Action, String)}, with a real classifier for an
* observer's or a collaborator's {@code TASK_READ}.
*/
McpSchema.CallToolResult denyFor(Principal caller, Authz.Action action, String target,
Predicate<String> ticketOwnedByCaller) {
// The enforcement switch lives HERE rather than in the exchange-facing wrapper: any future
// tool that calls this directly must not be able to skip the gate by accident.
if (!authorizationEnforced) {
return null; // AuthorizationMode.UNENFORCED: authorization not enforced (fleetd #518)
}
if (Authz.permits(caller, action, target, callers.knownLeadOrCollaborator(),
callers.sendableObserverTarget())) {
callers.observerSendTarget(), ticketOwnedByCaller)) {
if (action != Authz.Action.READ && action != Authz.Action.TASK_READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
@@ -797,13 +863,15 @@ public final class FleetMcp {
}
/**
* Who may see {@code fleet_list}'s {@code leads} array — exactly the roles that may
* {@link Authz.Action#SEND} to a lead: the primary, an architect, and a collaborator. A
* collaborator's own {@code fleet_whoami} carries no lead address, and {@code leads} is the
* only place this tool gives one, so a collaborator needs this array to use the send it
* already holds. A worker can never {@code SEND} at all, so it still sees neither this array
* nor {@code members}; a worker's own facts come from {@code fleet_whoami} instead. Split
* out for the same reason as {@link #coordinatorVisibleTo} and
* Who may see {@code fleet_list}'s {@code leads} array — the primary, an architect, and a
* collaborator. A collaborator's own {@code fleet_whoami} carries no lead address, and this is
* the only place this tool gives one, so a collaborator needs this array to use the
* {@link Authz.Action#SEND} it already holds. An observer holds that send to a lead too, but
* learns the address from its filtered {@code panes} rows instead: a {@code leads} row carries
* a lead's name, its configured context window and its config dir, which are the fleet's own
* shape rather than an address. A worker can never {@code SEND} at all, so it still sees
* neither this array nor {@code members}; a worker's own facts come from {@code fleet_whoami}
* instead. Split out for the same reason as {@link #coordinatorVisibleTo} and
* {@link #collaboratorsVisibleTo}: the decision must be unit-testable without fabricating an
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather
* than inlining the check.
@@ -824,13 +892,15 @@ public final class FleetMcp {
}
/**
* Who may see {@code fleet_list}'s {@code panes} array — every herdr-tracked agent pane on the
* host, carrying a tab label and a member's {@code cwd}. Visible to exactly the roles that may
* {@link Authz.Action#SEND} to a named peer; a plain worker or an observer holds {@code READ}
* but never {@code SEND}, so it does not see this array.
* Who may see {@code fleet_list}'s {@code panes} array — every role that may
* {@link Authz.Action#SEND} to some other pane. A plain worker holds {@code READ} but never
* {@code SEND}, so it still does not see this array. An observer does hold {@code SEND}, to a
* lead or another observer pane, so it sees the array too — and it is the only place this tool
* gives it a lead's address. {@code listFleet} filters its rows to
* {@link CallerResolver#observerSendTarget} and reduces each one; see {@code paneRows}.
*/
static boolean panesVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator() || caller.isObserver();
}
/**
@@ -966,12 +1036,17 @@ public final class FleetMcp {
* {@code fleet_send}: delegate {@code content} to a worker session and block for its reply.
* The configured profiles are required so a profile name can never bypass target validation.
*
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
* <p>(CB-548): {@code onAccepted} records delegator ownership the instant the send is
* accepted, so a BUSY interloper never claims a turn it did not win. {@code null} disables
* recording.
*
* @param mayDrainPoll whether this caller's {@code Authz.Action.DRAIN} is granted — a
* TIMED_OUT or BUSY receipt must name {@code fleet_poll} only when this
* is {@code true}, since that route is refused otherwise
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted, Set<String> profiles,
String callerOwner) {
String callerOwner, boolean mayDrainPoll) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
@@ -981,7 +1056,8 @@ public final class FleetMcp {
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout, onAccepted, callerOwner), timeout);
return formatReply(messages.send(sessionId, content, timeout, onAccepted, callerOwner), timeout,
sessionId, mayDrainPoll);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
@@ -992,14 +1068,20 @@ public final class FleetMcp {
* {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
* {@code callerOwner} must match the turn's recorded owner or this is refused.
*
* @param mayDrainPoll whether this caller's {@code Authz.Action.DRAIN} is granted — see
* {@link #send(MessageService, String, String, Long, Runnable, Set, String,
* boolean)}'s same parameter. {@code ANSWER} is primary-or-architect, and
* an architect is refused {@code DRAIN}, so this cannot default to true.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs,
String callerOwner) {
String callerOwner, boolean mayDrainPoll) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout, callerOwner), timeout);
String target = messages.sessionForTurn(turnId);
return formatReply(messages.answer(turnId, content, timeout, callerOwner), timeout, target, mayDrainPoll);
}
/**
@@ -1026,8 +1108,15 @@ public final class FleetMcp {
};
}
/** Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and {@link #answer}. */
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout) {
/**
* Render a {@link MessageService.Reply} as a tool result — shared by {@link #send} and
* {@link #answer}. {@code target} is the worker session the send went to (may be {@code null}
* if it could not be resolved), named in a TIMED_OUT_* receipt so the caller knows which
* inbox to poll — but only when {@code mayDrainPoll} is {@code true}: naming a route the
* caller's own {@code Authz.Action.DRAIN} grant would refuse is worse than naming none.
*/
private static McpSchema.CallToolResult formatReply(MessageService.Reply r, long timeout, String target,
boolean mayDrainPoll) {
return switch (r.outcome()) {
case REPLIED -> text(r.text());
// The worker's turn finished but it never called fleet_reply — hand back the scraped
@@ -1050,8 +1139,16 @@ public final class FleetMcp {
+ "answered (turnId stale)");
case NOT_TURN_OWNER -> error("this turn belongs to a different delegation — only the caller "
+ "that opened it may answer it");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> {
String base = "[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + ". "
+ MessageService.NO_TICKET_NO_RESEND + ". ";
yield text(mayDrainPoll
? base + "Poll fleet_poll{target=\""
+ (target != null ? target : "<the worker session you sent to>")
+ "\"} to drain its inbox for a late reply.]"
: base + "A late reply cannot be recovered on this channel — ask the lead.]");
}
// fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call,
// so the message may already be sitting in the pane. Do not invite a blind retry the way
// the case above does; a resend on this route can double-deliver the same brief.
@@ -1089,7 +1186,29 @@ public final class FleetMcp {
return targetError;
}
String ticket = messages.sendAsync(sessionId, content, onAccepted, creator);
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket);
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket
+ notInjectableWarning(messages, sessionId));
}
/**
* A synchronous status check at accept time: when {@code sessionId} is not currently
* idle/blocked/done, this message is queued rather than reaching the pane right away. Empty
* when the target is injectable or its status could not be read — a best-effort warning, not
* a reason to withhold the accept receipt.
*/
private static String notInjectableWarning(MessageService messages, String sessionId) {
AgentStatus status;
try {
status = messages.status(sessionId);
} catch (RuntimeException e) {
return "";
}
if (status.injectable()) {
return "";
}
return "\n\nWarning: " + sessionId + " is currently " + status.name().toLowerCase()
+ ", not idle/blocked/done — this message is queued, not yet delivered, and will "
+ "wait until the target frees up. Poll fleet_poll to see when it lands.";
}
/** A configured profile is never a send target; other unknown values may be herdr-owned panes. */
@@ -1246,6 +1365,7 @@ public final class FleetMcp {
case SEND -> sendAction(str(arguments, "coordId"), str(arguments, "turnId"));
case REPLY -> Authz.Action.REPLY;
case ASK -> Authz.Action.ASK;
case INBOX -> Authz.Action.INBOX;
case STATUS -> Authz.Action.TASK_READ;
case LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
@@ -1389,6 +1509,33 @@ public final class FleetMcp {
return text(outcome.description());
}
/**
* {@code fleet_inbox}: hand the caller the messages queued for its own pane, so it submits them
* itself instead of having them typed into its terminal. {@code callerTerminal} is resolved
* from the connection; there is deliberately no pane argument, so no caller can collect another
* pane's mail.
*
* <p>Calling this is also what marks the pane as collecting its own mail, for a window the
* injector re-checks before every delivery. A pane that stops calling it stops being served
* this way and its queued mail is typed instead, so an empty answer is a normal result that
* must still be requested on a timer.
*
* <p>Returns a JSON object with {@code count} and {@code messages} (in the order they were
* queued), so a caller can tell "no mail" apart from a failure.
*/
static McpSchema.CallToolResult inbox(MessageService messages, String callerTerminal) {
if (callerTerminal == null) {
return error("fleet_inbox could not identify the calling pane from the connection, so "
+ "there is no inbox to collect");
}
List<String> collected = messages.collectInbox(callerTerminal);
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", callerTerminal);
m.put("count", collected.size());
m.put("messages", collected);
return text(json(m));
}
/** {@code fleet_ack}: acknowledge (remove) a specific reply from the inbox. */
static McpSchema.CallToolResult ack(MessageService messages, String target, String msgId) {
if (isBlank(target) || isBlank(msgId)) {
@@ -2069,6 +2216,32 @@ public final class FleetMcp {
CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible,
PaneSource panes, boolean panesVisible) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
leadSeats, contextGauge, leadConfigDirs, leads, selfTerm, collaborators, collaboratorsVisible,
coordination, callerIsPrimary, leadsVisible, membersVisible, panes, panesVisible, false);
}
/**
* As above, plus: an observer sees the {@code panes} array too, but filtered to
* {@link CallerResolver#observerSendTarget} and each row reduced to the five fields an
* observer may learn — see {@code paneRows}/{@code paneRow}.
*
* @param callerIsObserver whether the {@code fleet_list} caller is an observer; every wrapper
* overload above passes {@code false}, so a test that wants the
* filtered, reduced view must call this overload with an explicit
* {@code true}
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
LoopHealthSource loopHealth,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs,
Map<String, String> leads, String selfTerm,
Map<String, String> collaborators, boolean collaboratorsVisible,
CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible,
PaneSource panes, boolean panesVisible, boolean callerIsObserver) {
try {
// Neither row's assembly (leadView/memberCapacityView probing herdr for live status)
// runs unless at least one of them needs the live-agent lookup backing it.
@@ -2118,7 +2291,7 @@ public final class FleetMcp {
}
// gate BEFORE assembling the row, so the key is absent rather than present-and-empty.
if (panesVisible) {
result.put("panes", paneRows(live, roster, leads, collaborators, panes));
result.put("panes", paneRows(live, roster, leads, collaborators, panes, callerIsObserver));
}
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
@@ -2458,21 +2631,39 @@ public final class FleetMcp {
}
}
/** As {@link #tabLabelsOrEmpty}, for {@code panes.workspaceLabels()}. */
private static Map<String, String> workspaceLabelsOrEmpty(PaneSource panes) {
try {
return panes.workspaceLabels().get();
} catch (HerdrException e) {
return Map.of();
}
}
/**
* One row per herdr-tracked agent pane, sorted by terminal id for a stable order. {@code live}
* is the same terminal-keyed {@link Agent} map {@code leadView}/{@code memberCapacityView}
* already read, so a pane neither configured as a lead nor spawned as a member — a hand-opened
* tab — still gets a row here.
*
* <p>For an observer caller ({@code observerView}), the rows are filtered to
* {@link PaneSource#observerSendTarget} before being built, and each row is reduced — see
* {@code paneRow}. A lead's pane survives that filter, so it is where an observer reads a
* lead's {@code sessionId}.
*/
private static List<Map<String, Object>> paneRows(Map<String, Agent> live, List<MemberSession> roster,
Map<String, String> leads, Map<String, String> collaborators, PaneSource panes) {
Map<String, String> leads, Map<String, String> collaborators, PaneSource panes,
boolean observerView) {
final Map<String, String> tabLabels = tabLabelsOrEmpty(panes);
final Map<String, String> workspaceLabels = workspaceLabelsOrEmpty(panes);
Map<String, MemberSession> byTerminal = roster.stream()
.filter(s -> s.terminalId() != null)
.collect(Collectors.toMap(MemberSession::terminalId, Function.identity(), (_, b) -> b));
return live.values().stream()
.filter(a -> !observerView || panes.observerSendTarget().test(a.terminalId()))
.sorted(Comparator.comparing(Agent::terminalId))
.map(a -> paneRow(a, byTerminal.get(a.terminalId()), leads, collaborators, tabLabels, panes))
.map(a -> paneRow(a, byTerminal.get(a.terminalId()), leads, collaborators, tabLabels,
workspaceLabels, panes, observerView))
.toList();
}
@@ -2481,38 +2672,57 @@ public final class FleetMcp {
* daemon never spawned as a member (a hand-opened tab, or a configured lead)
* @param tabLabels tab id → its herdr display label; a tab absent here, or carrying a
* {@code null} label itself, projects as a {@code null} "label"
* @param workspaceLabels workspace id → its herdr display label (the space name); a workspace
* absent here, or carrying a {@code null} label itself, projects as a
* {@code null} "workspaceLabel"
* @param observerView an observer's row carries only {@code sessionId}, {@code label},
* {@code status}, {@code role}, {@code deliverable} — never {@code paneId}
* (the {@code fleet_stop} handle), {@code workspaceId},
* {@code workspaceLabel}, {@code tabId}, {@code agentType}, or {@code cwd}
* (a member's worktree path is the lead's business)
*/
private static Map<String, Object> paneRow(Agent a, MemberSession session, Map<String, String> leads,
Map<String, String> collaborators, Map<String, String> tabLabels, PaneSource panes) {
Map<String, String> collaborators, Map<String, String> tabLabels,
Map<String, String> workspaceLabels, PaneSource panes, boolean observerView) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", a.terminalId());
m.put("paneId", a.paneId());
m.put("workspaceId", a.workspaceId());
m.put("tabId", a.tabId());
if (!observerView) {
m.put("paneId", a.paneId());
m.put("workspaceId", a.workspaceId());
m.put("workspaceLabel", a.workspaceId() == null ? null : workspaceLabels.get(a.workspaceId()));
m.put("tabId", a.tabId());
}
m.put("label", a.tabId() == null ? null : tabLabels.get(a.tabId()));
m.put("agentType", a.agentType());
if (!observerView) {
m.put("agentType", a.agentType());
}
m.put("status", a.status() == null ? "unknown" : a.status().name().toLowerCase());
m.put("role", paneRole(a.terminalId(), session, leads, collaborators));
m.put("role", paneRole(a.terminalId(), session, leads, collaborators, panes));
m.put("deliverable", panes.deliverable().test(a.terminalId()));
if (session != null && session.cwd() != null) {
if (!observerView && session != null && session.cwd() != null) {
m.put("cwd", session.cwd());
}
return m;
}
/**
* The role this pane resolves as: a spawned member's own {@link MemberRole}, else "lead" or
* "collaborator" for a configured but currently-unoccupied slot, else "observer" for a pane
* this daemon neither spawned nor configured.
* The role this pane resolves as: a spawned member's own {@link MemberRole}, else "lead" for a
* configured but currently-unoccupied lead pane, else "architect" for a pane bound to a
* configured architect slot with no live member session, else "collaborator" for a configured
* but currently-unoccupied collaborator tab, else "observer" for a pane this daemon neither
* spawned nor configured.
*/
private static String paneRole(String terminal, MemberSession session, Map<String, String> leads,
Map<String, String> collaborators) {
Map<String, String> collaborators, PaneSource panes) {
if (session != null) {
return session.role().wireName();
}
if (leads.containsKey(terminal)) {
return "lead";
}
if (panes.architectSlot().test(terminal)) {
return "architect";
}
if (collaborators.containsKey(terminal)) {
return "collaborator";
}
@@ -2715,17 +2925,20 @@ public final class FleetMcp {
return tool(FleetTool.LIST.wireName(),
"List the whole fleet the bridge tracks, in two parts. 'members' is visible to "
+ "the primary and an architect only. 'leads' is visible to those two AND a "
+ "collaborator — exactly the roles that may fleet_send to a lead, so a "
+ "collaborator can learn a lead's sessionId before using the send it already "
+ "holds. A worker holds READ to call this tool at all, but gets neither "
+ "collaborator, so a collaborator can learn a lead's sessionId before using "
+ "the send it already holds. An observer may fleet_send to a lead too, but "
+ "reads that sessionId from its own 'panes' rows instead: a 'panes' row for "
+ "an observer is filtered to the panes it may send to — a lead's pane and "
+ "another observer's — and reduced to sessionId, label, status, role and "
+ "deliverable. A worker holds READ to call this tool at all, but gets neither "
+ "array, never an empty one; a worker reads its own session, "
+ "profile, state, worktree, branch and owner from fleet_whoami instead. "
+ "'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "sessions delegated to — each with sessionId, paneId, role (" + MemberRole.wireNames()
+ "), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner/agentSessionId, and live herdr status. agentSessionId, "
+ "when present, is the id to pass as fleet_spawn's resumeSessionId to relaunch "
+ "onto that same conversation. It is ABSENT — not a guess — for a member fleetd "
@@ -2803,6 +3016,18 @@ public final class FleetMcp {
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool inboxTool() {
return tool(FleetTool.INBOX.wireName(),
"Collect the messages the fleet has queued for YOUR OWN pane, then act on each one. "
+ "Takes no arguments: the pane is resolved from your connection, so you can "
+ "never read another session's mail. Any role may call it for itself. Call it "
+ "on a timer — each call is also what tells the daemon you collect your own "
+ "mail, so it offers your next message here instead of typing it into your "
+ "terminal; stop calling it and your mail is typed instead. An empty "
+ "'messages' array is the normal answer when nothing is waiting.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool handoverTool() {
return tool(FleetTool.HANDOVER.wireName(),
"Replace your OWN lead session once its context is full: write a handover file, "
@@ -44,7 +44,8 @@ public enum FleetTool {
STOP("fleet_stop"),
PROFILES("fleet_profiles"),
WHOAMI("fleet_whoami"),
HANDOVER("fleet_handover");
HANDOVER("fleet_handover"),
INBOX("fleet_inbox");
private final String wireName;
@@ -48,8 +48,13 @@ public final class PrimaryRegistry {
*/
private final ConcurrentHashMap<String, Delegation> leadByTarget = new ConcurrentHashMap<>();
/** A recorded delegator: the terminal learned from call traffic, and its name, if it has one. */
private record Delegation(String terminal, String name) {
/**
* A recorded delegator: the terminal learned from call traffic, its name if it has one, and
* whether a nudge may tell it to {@code fleet_poll(target=…)} or {@code fleet_send(turnId=…)}
* — the real authorization decision for this delegator's role, captured once by the caller at
* delegation time rather than re-derived from a bare terminal string later.
*/
private record Delegation(String terminal, String name, boolean mayDrainNudge, boolean mayAnswerNudge) {
}
/**
@@ -128,14 +133,28 @@ public final class PrimaryRegistry {
/**
* As {@link #recordDelegation(String, String)}, additionally recording the delegating lead's
* name when the caller carries one. See {@link #record(String, String)} for why the name
* matters.
* name when the caller carries one, and granting a nudge about this target everything a
* primary may be told to run. See {@link #record(String, String)} for why the name matters,
* and {@link #recordDelegation(String, String, String, boolean, boolean)} for why every other
* caller must supply its own, real grant instead of this default.
*/
public void recordDelegation(String target, String leadTerminal, String leadName) {
recordDelegation(target, leadTerminal, leadName, true, true);
}
/**
* As {@link #recordDelegation(String, String, String)}, with the delegating caller's own
* {@code DRAIN} and {@code ANSWER} grants — the real {@code Authz.permits} decision for its
* role, asked once by the caller at delegation time, since a terminal string alone cannot be
* resolved back to a role here. A nudge about this target must offer only what
* {@link #nudgeMayDrainFor(String)}/{@link #nudgeMayAnswerFor(String)} report back as true.
*/
public void recordDelegation(String target, String leadTerminal, String leadName,
boolean mayDrainNudge, boolean mayAnswerNudge) {
if (target == null || target.isBlank() || leadTerminal == null || leadTerminal.isBlank()) {
return;
}
leadByTarget.put(target, new Delegation(leadTerminal, blankToNull(leadName)));
leadByTarget.put(target, new Delegation(leadTerminal, blankToNull(leadName), mayDrainNudge, mayAnswerNudge));
}
/** Forget a worker's delegating lead — call on release, so a torn-down session leaks nothing. */
@@ -167,6 +186,32 @@ public final class PrimaryRegistry {
return currentPrimaryTerminal();
}
/**
* Whether a nudge about {@code target} may tell its recipient to {@code fleet_poll(target=…)}
* — present on exactly the same condition as {@link #nudgeTargetFor(String)}, carrying the
* grant recorded for a delegation, or {@code true} for the singleton primary fallback, which
* is always genuinely the primary.
*/
public Optional<Boolean> nudgeMayDrainFor(String target) {
Delegation delegation = target == null ? null : leadByTarget.get(target);
if (delegation != null) {
return Optional.of(delegation.mayDrainNudge());
}
return currentPrimaryTerminal().isPresent() ? Optional.of(true) : Optional.empty();
}
/**
* As {@link #nudgeMayDrainFor(String)}, for whether a nudge may tell its recipient to
* {@code fleet_send(turnId=…)}.
*/
public Optional<Boolean> nudgeMayAnswerFor(String target) {
Delegation delegation = target == null ? null : leadByTarget.get(target);
if (delegation != null) {
return Optional.of(delegation.mayAnswerNudge());
}
return currentPrimaryTerminal().isPresent() ? Optional.of(true) : Optional.empty();
}
/**
* The known primary terminal, or empty if not yet learned (and not pinned) — the raw value as
* it was recorded, with no attempt to resolve a named lead's current pane. Callers that need a
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.PromptBox;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -23,7 +24,8 @@ import java.util.function.Supplier;
* <p><strong>Status-gated, exactly like {@link ReplyPushLoop}.</strong> A pane may only be injected
* into at a turn boundary ({@link AgentStatus#injectable()} — idle, blocked or done); pasting into
* a live turn corrupts it. So a tick that finds the lead busy simply does nothing and comes back
* later.
* later. The same holds for a lead whose prompt box holds unsubmitted text ({@link PromptBox}) —
* delivering there would submit the operator's half-typed line along with the message.
*
* <p><strong>Ack only after delivery.</strong> A message is acked — removed from the broker — only
* once {@link AgentControl#send} has actually put it in the pane. Anything not delivered (no lead
@@ -52,6 +54,7 @@ public final class LeadCoordLoop {
private final LeadChannel channel;
private final AgentControl agents;
private final PromptBox promptBox;
private final Supplier<Map<String, String>> leads;
private final ScheduledExecutorService scheduler;
private final long intervalMs;
@@ -60,6 +63,12 @@ public final class LeadCoordLoop {
private volatile boolean running;
/** Equivalent to {@link #LeadCoordLoop(LeadChannel, AgentControl, Supplier, ScheduledExecutorService, long, boolean)} with the prompt-box gate on. */
public LeadCoordLoop(LeadChannel channel, AgentControl agents, Supplier<Map<String, String>> leads,
ScheduledExecutorService scheduler, long intervalMs) {
this(channel, agents, leads, scheduler, intervalMs, true);
}
/**
* @param channel this daemon's own lead mailbox
* @param agents herdr control, for the status gate and the pane injection
@@ -68,11 +77,13 @@ public final class LeadCoordLoop {
* the tab scan after startup becomes reachable without a restart
* @param scheduler the loop's own scheduler; the caller owns its shutdown
* @param intervalMs how long between ticks
* @param promptBoxGateEnabled passed straight to {@link PromptBox#PromptBox(AgentControl, boolean)}
*/
public LeadCoordLoop(LeadChannel channel, AgentControl agents, Supplier<Map<String, String>> leads,
ScheduledExecutorService scheduler, long intervalMs) {
ScheduledExecutorService scheduler, long intervalMs, boolean promptBoxGateEnabled) {
this.channel = channel;
this.agents = agents;
this.promptBox = new PromptBox(agents, promptBoxGateEnabled);
this.leads = leads;
this.scheduler = scheduler;
this.intervalMs = intervalMs;
@@ -149,6 +160,11 @@ public final class LeadCoordLoop {
lead, status, held.size());
return;
}
if (!promptBox.clearToSubmit(lead)) {
log.debug("lead coordination: lead {} has unsubmitted text in its prompt box, holding {} message(s)",
lead, held.size());
return;
}
try {
agents.send(lead, DELIVERY_FORMAT.formatted(msg.from(), msg.content()));
} catch (RuntimeException e) {
@@ -2,6 +2,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.PromptBox;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
@@ -33,7 +34,7 @@ import java.util.function.Supplier;
* no such block it is never constructed, so upgrading the daemon cannot silently acquire a behaviour
* that spends the operator's model subscription on its own initiative (constraint 1).
*
* <p>Four invariants keep it from becoming a runaway subscription burner:
* <p>Five invariants keep it from becoming a runaway subscription burner:
* <ol>
* <li><b>Status-gated</b> — a {@code WORKING} lead is making progress and is never touched; only an
* injectable (idle/done/blocked) lead is even considered (constraint 2).</li>
@@ -45,6 +46,8 @@ import java.util.function.Supplier;
* <li><b>Never races {@link ReplyPushLoop}</b> — while that loop is actively nudging any target this
* loop stands down, so two competing injections never start two turns in the same pane
* (constraint 6).</li>
* <li><b>Never submits the operator's draft</b> — a nudge is held while the lead's prompt box holds
* unsubmitted text ({@link PromptBox}), because the delivery pastes and submits in one call.</li>
* </ol>
*
* <p><b>fleetd #609 — context-high notice.</b> Optionally ({@code contextHighNudge}, opt-in like the
@@ -65,6 +68,7 @@ public final class LeadHeartbeatLoop {
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final PromptBox promptBox;
private final ReplyInbox inbox;
private final Supplier<List<MemberSession>> roster;
private final ReplyPushLoop pushLoop;
@@ -133,8 +137,21 @@ public final class LeadHeartbeatLoop {
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
LeadContextSource contextSource, boolean contextHighNudge,
boolean requireOperatorConfirm) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, metrics, contextSource, contextHighNudge,
requireOperatorConfirm, true);
}
/** As above, plus the prompt-box gate's on/off switch — see {@link PromptBox#PromptBox(AgentControl, boolean)}. */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics,
LeadContextSource contextSource, boolean contextHighNudge,
boolean requireOperatorConfirm, boolean promptBoxGateEnabled) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.promptBox = new PromptBox(agents, promptBoxGateEnabled);
this.inbox = inbox;
this.roster = roster;
this.pushLoop = pushLoop;
@@ -187,7 +204,9 @@ public final class LeadHeartbeatLoop {
/** Idle past the quiet period with nothing pending and the cap exhausted — stop until new state appears. */
QUIET_DONE,
/** {@link ReplyPushLoop} is actively nudging — stand aside rather than start a competing turn. */
STAND_DOWN
STAND_DOWN,
/** The lead's prompt box holds unsubmitted text — hold the nudge rather than submit that text. */
DRAFT_HELD
}
/**
@@ -330,6 +349,7 @@ public final class LeadHeartbeatLoop {
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
quietCount, status, pushLoop.isActive(), leadKnown, fleet,
reading.state(), contextNotified);
d = holdIfOperatorIsTyping(d);
applyDecision(d);
switch (d.action()) {
case INJECT -> injectNudge(d, fleet, reading);
@@ -337,11 +357,33 @@ public final class LeadHeartbeatLoop {
countNudge("exhausted");
contextNotified = d.contextNotified();
}
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> contextNotified = d.contextNotified();
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN, DRAFT_HELD -> contextNotified = d.contextNotified();
}
scheduleNext();
}
/**
* Turn a decision to inject into {@link Action#DRAFT_HELD} when the lead's prompt box holds text
* the operator has not submitted. The pane read happens only for a decision that would otherwise
* send, so a busy or debouncing lead costs no extra herdr call.
*
* <p>The held decision carries this tick's idle window but the <em>pre-tick</em> quiet count and
* context latch: nothing reached the pane, so neither the quiet budget nor the one context notice
* per stretch may be spent on it.
*/
private Decision holdIfOperatorIsTyping(Decision d) {
if (d.action() != Action.INJECT) {
return d;
}
var lead = primaryRegistry.currentPrimaryTerminal();
if (lead.isEmpty() || promptBox.clearToSubmit(lead.get())) {
return d;
}
log.debug("idle-heartbeat: lead {} has unsubmitted text in its prompt box, holding the nudge",
lead.get());
return new Decision(Action.DRAFT_HELD, d.idleSinceNanos(), quietCount, contextNotified);
}
/**
* Persist the idle/quiet state a decision returned, so the next tick starts from it.
*
@@ -134,6 +134,14 @@ public final class MessageService {
NOT_TURN_OWNER
}
/**
* The fact shared by the MCP and REST receipts for {@link Outcome#TIMED_OUT_WORKING},
* {@link Outcome#TIMED_OUT_QUEUED} and {@link Outcome#BUSY}: a blocking send is never tracked
* by a ticket, so there is no id to poll by, and resending it risks a duplicate delivery.
*/
public static final String NO_TICKET_NO_RESEND =
"this send created no ticket, and a resend can duplicate the delivery";
/**
* @param outcome how the send ended (or paused)
* @param text the worker's answer when {@link #completed()} (a structured {@code fleet_reply}
@@ -353,6 +361,16 @@ public final class MessageService {
* target ({@link #send} opening a fresh waiter) or a teardown ({@link #abandon}).
*/
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
/**
* The captured waiter a timed-out send left open for a late turn-completion scrape
* ({@link #strandLateResolution}), keyed by target so a late {@code fleet_reply} can find and
* claim the exact one its own turn belongs to ({@link #reply}). The waiter's own {@code
* complete()} call is the discriminator: whichever caller resolves it first — this reply, or
* the eventual completion scrape — decides what the other one sees, and only the winner's kind
* ever reaches the inbox. An entry removes itself once its waiter resolves.
*/
private final ConcurrentHashMap<String, CompletableFuture<Rendezvous.Resolution>> lateWaiters =
new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
/**
* Minted once per {@code MessageService} instance and folded into every ticket id (see
@@ -514,6 +532,20 @@ public final class MessageService {
return false;
}
/**
* Hand {@code session} every message queued for it that it has not collected yet, and record
* that it collects its own mail. While that record is fresh, delivery to that session is
* offered for collection instead of typed into its terminal; once it goes stale, the terminal
* route takes over again with nothing lost.
*
* <p>The messages are returned in the order they were queued, and are removed by this call.
* An empty list is an ordinary answer: a session polling on a timer keeps itself collecting
* between messages.
*/
public List<String> collectInbox(String session) {
return injector.collectInbox(session);
}
/**
* Route a worker's explicit {@code fleet_reply}: resolve an open send, complete an async ticket
* still parked waiting on this exact turn's answer, or — only once neither applies — queue it in
@@ -598,6 +630,14 @@ public final class MessageService {
+ "answers, queuing to the inbox instead of guessing", session, candidates.size(),
tickets);
}
// Claim the matching timed-out send's captured waiter, if one is still open, before
// publishing this reply. The waiter is this exact turn's own CompletableFuture, so
// completing it here means a completion-fallback scrape that resolves the same waiter
// afterward finds it already answered and does not publish a second entry for this turn.
CompletableFuture<Rendezvous.Resolution> lateWaiter = lateWaiters.get(session);
if (lateWaiter != null) {
rendezvous.resolveLateReply(lateWaiter, content);
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// CB-640: record the stranding itself (not just the reply text) so fleet health can see a
// worker whose replies keep missing their waiter, not only the queue depth this leaves behind.
@@ -794,6 +834,7 @@ public final class MessageService {
// CB-640: the session is gone — nothing will ever accept or deliver into it now.
strandedReplies.remove(target);
queuedDeliveries.remove(target);
lateWaiters.remove(target);
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
@@ -1036,6 +1077,7 @@ public final class MessageService {
queuedDeliveries.put(target, Boolean.TRUE);
outcome = Outcome.TIMED_OUT_QUEUED;
}
strandLateResolution(target, reply);
return recorded(new Reply(outcome, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
@@ -1053,6 +1095,42 @@ public final class MessageService {
}
}
/**
* Once a blocking send has given up on {@code waiter} and already reported a TIMED_OUT_*
* outcome, the CB-106 completion fallback can still resolve it later — it completes the exact
* captured future ({@link Rendezvous#resolveCompletion}), not a session lookup, so
* {@link Rendezvous#close} does not stop it. With nobody left awaiting {@code waiter}, route
* that resolution to {@code target}'s inbox instead, exactly where a late explicit
* {@code fleet_reply} already lands (see {@link #reply}), so {@code fleet_poll{target}} can
* recover it. A no-op if {@code waiter} never resolves, or already has by the time this runs.
*
* <p>Only {@link Rendezvous.Kind#COMPLETION} is routed here. An explicit {@code fleet_reply}
* ({@link Rendezvous.Kind#REPLY}) still reaches {@code waiter} through {@link #reply}'s own
* session lookup, not through this captured reference, so leaving it out here avoids a second,
* racing publish of the same reply.
*/
private void strandLateResolution(String target, CompletableFuture<Rendezvous.Resolution> waiter) {
lateWaiters.put(target, waiter);
waiter.whenComplete((resolution, error) -> {
lateWaiters.remove(target, waiter);
if (resolution != null && resolution.kind() == Rendezvous.Kind.COMPLETION) {
try {
// inbox.publish reaches a broker and can throw. An exception thrown inside a
// whenComplete callback is swallowed into the discarded dependent stage, not
// rethrown to any caller — log it here, or a broker failure is invisible and the
// late reply is simply lost.
inbox.publish(target, UUID.randomUUID().toString(), resolution.text());
strandedReplies.put(target, Boolean.TRUE);
if (pushLoop != null) {
pushLoop.onReplyQueued(target);
}
} catch (RuntimeException e) {
log.warn("strandLateResolution: publishing a late turn-completion for {} failed", target, e);
}
}
});
}
/**
* A worker's mid-turn question (CB-205 reverse rendezvous): surface {@code question} to the
* primary by resolving its open blocking {@code fleet_send}, then block this (worker) call until
@@ -1153,6 +1231,14 @@ public final class MessageService {
}
}
/**
* The worker session an open {@code fleet_ask} turn belongs to, or {@code null} if
* {@code turnId} is unknown or has already lapsed.
*/
public String sessionForTurn(String turnId) {
return rendezvous.askSession(turnId);
}
/**
* The primary's answer to a worker's {@code fleet_ask} (CB-205): resolve the worker's blocked
* question identified by {@code turnId}, then — like a fresh {@link #send} — block for the worker's
@@ -1426,7 +1512,7 @@ public final class MessageService {
return new TaskView(ticket, Phase.ASKING, question.text(), null,
"worker is waiting for your answer", question.turnId());
}
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
return new TaskView(ticket, Phase.PENDING, null, null, pendingDetail(task.target), null);
}
// CB-588: the ticket is terminal and being handed to the caller right here — tell the push
// loop it is collected so a later tick's nudge never names a ticket the lead already has.
@@ -1478,6 +1564,16 @@ public final class MessageService {
|| Objects.equals(callerOwner, task.creatorOwner);
}
/**
* Whether {@code callerOwner} owns {@code ticket}, with no side effect — unlike
* {@link #poll(String, String)}, this never fires the ticket's collection hook. A ticket this
* daemon has never heard of is owned by nobody.
*/
public boolean ownsTicket(String ticket, String callerOwner) {
Task task = ticket == null ? null : tasks.get(ticket);
return task != null && ownsTicket(task, callerOwner);
}
/**
* Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
@@ -1499,6 +1595,21 @@ public final class MessageService {
return task != null && task.completedNanos != null;
}
/**
* Detail text for a {@link Phase#PENDING} poll of a plain (non-asking) delegation.
* Distinguishes a message still sitting in the injector's queue, never delivered, from one
* that already reached the pane and is simply being worked on — so a caller cannot read
* "worker working" as "received" when it was not.
*/
private String pendingDetail(String target) {
Long queuedMillis = injector.queuedWaitMillis(target);
if (queuedMillis != null) {
return "queued, not yet delivered (target is " + liveStatus(target) + "; queued "
+ (queuedMillis / 1000) + "s)";
}
return "worker " + liveStatus(target);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
@@ -304,6 +304,20 @@ public final class Rendezvous {
return waiter != null && waiter.complete(new Resolution(Kind.BACKEND_EXHAUSTED, reason));
}
/**
* Resolve a specific captured {@code waiter} as a reply — for a {@code fleet_reply} that
* arrives after the send which opened {@code waiter} already gave up on it, so the ordinary
* session-keyed {@link #resolve} finds no live waiter to match. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured waiter
* (CB-116): a no-op if that waiter was already resolved — first resolution wins, and whichever
* one wins is what any later resolution attempt on the same waiter will see.
*
* @return true if this call resolved the waiter, false if it was null or already resolved
*/
public boolean resolveLateReply(CompletableFuture<Resolution> waiter, String content) {
return waiter != null && waiter.complete(new Resolution(Kind.REPLY, content));
}
private boolean complete(String session, Resolution resolution) {
ForwardWaiter waiter = waiters.get(session);
return waiter != null && waiter.future().complete(resolution);
@@ -3,6 +3,7 @@ package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PromptBox;
import dev.ltms.fleet.mcp.PrimaryRegistry;
import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
@@ -50,6 +51,11 @@ import java.util.stream.Collectors;
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
* {@link #decide}) — whichever the durable inbox / pending set doesn't already answer via
* {@code STOP}.
*
* <p>A lead that is injectable is nudged only when its prompt box is also empty
* ({@link PromptBox}): the delivery pastes and submits in one call, so a nudge into a box holding
* the operator's half-typed line would submit that line too. A nudge held for that reason waits for
* the next tick like any other, and the pending work is re-read then.
*/
public final class ReplyPushLoop {
@@ -81,6 +87,7 @@ public final class ReplyPushLoop {
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final PromptBox promptBox;
private final ReplyInbox inbox;
private final ScheduledExecutorService scheduler;
private final int maxReminders;
@@ -123,8 +130,16 @@ public final class ReplyPushLoop {
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics) {
this(primaryRegistry, agents, inbox, scheduler, maxReminders, backoffMs, metrics, true);
}
/** As above, with the prompt-box gate's on/off switch — see {@link PromptBox#PromptBox(AgentControl, boolean)}. */
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
int maxReminders, long backoffMs, Metrics metrics, boolean promptBoxGateEnabled) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.promptBox = new PromptBox(agents, promptBoxGateEnabled);
this.inbox = inbox;
this.scheduler = scheduler;
this.maxReminders = maxReminders;
@@ -170,13 +185,20 @@ public final class ReplyPushLoop {
pendingReplies.remove(target, owning);
continue;
}
// Never nudge a caller to run a DRAIN its own role could never perform -- the durable
// inbox stays the backstop for it, same as when no nudge target is known at all.
if (!owning.mayDrainNudge()) continue;
result.add(target);
}
return result;
}
/** A pending reply target: which lead to nudge, and how many nudges have named it so far. */
private record ReplyEntry(String lead, int nudgeCount) {
/**
* A pending reply target: which nudge target to notify, how many nudges have named it so
* far, and whether that nudge target's own role may actually run the
* {@code fleet_poll(target=…)} a reply nudge would tell it to run.
*/
private record ReplyEntry(String lead, int nudgeCount, boolean mayDrainNudge) {
}
/**
@@ -199,11 +221,12 @@ public final class ReplyPushLoop {
/**
* An open question awaiting the lead's answer: which ticket it belongs to, which worker asked,
* which lead to nudge, the question text, and how many nudges have named it so far (CB-598 —
* tracked per question, not per lead per source).
* which lead to nudge, the question text, how many nudges have named it so far (CB-598 —
* tracked per question, not per lead per source), and whether that nudge target's own role may
* actually run the {@code fleet_send(turnId=…)} a question nudge would tell it to run.
*/
private record PendingQuestion(String turnId, String ticket, String target, String lead,
String question, int nudgeCount) {
String question, int nudgeCount, boolean mayAnswerNudge) {
}
private record IncidentLead(String incidentId, String lead) {
@@ -221,7 +244,11 @@ public final class ReplyPushLoop {
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
private List<PendingQuestion> pendingQuestionsFor(String lead) {
return pendingQuestions.values().stream().filter(q -> lead.equals(q.lead())).toList();
// Never nudge a caller to run an ANSWER its own role could never perform -- the 55-second
// ask window lapses on its own, same as when no nudge target is known at all.
return pendingQuestions.values().stream()
.filter(q -> lead.equals(q.lead()) && q.mayAnswerNudge())
.toList();
}
/** Question turnIds still open for {@code lead} — a plain snapshot for race comparison. */
@@ -377,7 +404,7 @@ public final class ReplyPushLoop {
boolean hasIncidentWork = !pendingIncidentKeysFor(lead).isEmpty();
boolean hasUnmappedTargetWork = !pendingUnmappedTargetKeysFor(lead).isEmpty();
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork && !hasIncidentWork && !hasUnmappedTargetWork) {
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
log.debug("push: nothing pending for nudge target {}, stopping reminder", lead);
return Action.STOP;
}
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
@@ -386,7 +413,7 @@ public final class ReplyPushLoop {
boolean incidentEligible = hasIncidentWork && incidentReminderCount < maxReminders;
boolean unmappedTargetEligible = hasUnmappedTargetWork && unmappedTargetReminderCount < maxReminders;
if (!replyEligible && !ticketEligible && !questionEligible && !incidentEligible && !unmappedTargetEligible) {
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
log.debug("push: reminder cap ({}) reached for nudge target {} on every source with pending work, stopping",
maxReminders, lead);
countNudge("exhausted");
return Action.STOP;
@@ -395,14 +422,18 @@ public final class ReplyPushLoop {
try {
status = agents.status(lead);
} catch (RuntimeException e) {
log.debug("push: status check failed for lead {}, will retry", lead, e);
log.debug("push: status check failed for nudge target {}, will retry", lead, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
if (!status.injectable()) {
log.debug("push: nudge target {} is {} (not injectable), waiting", lead, status);
return Action.WAIT_BUSY;
}
log.debug("push: lead {} is {} (not injectable), waiting", lead, status);
return Action.WAIT_BUSY;
if (!promptBox.clearToSubmit(lead)) {
log.debug("push: nudge target {} has unsubmitted text in its prompt box, waiting", lead);
return Action.WAIT_BUSY;
}
return Action.INJECT;
}
/**
@@ -454,14 +485,14 @@ public final class ReplyPushLoop {
if (lead.isEmpty() || isLive(lead.get())) {
return lead;
}
log.debug("push: lead {} delegated to for {} is no longer live, forgetting the stale binding "
log.debug("push: nudge target {} delegated to for {} is no longer live, forgetting the stale binding "
+ "and falling back", lead.get(), target);
primaryRegistry.forgetDelegation(target);
Optional<String> fallback = primaryRegistry.nudgeTargetFor(target);
if (fallback.isEmpty() || isLive(fallback.get())) {
return fallback;
}
log.debug("push: fallback lead {} for {} is also not live, skipping this tick", fallback.get(), target);
log.debug("push: fallback nudge target {} for {} is also not live, skipping this tick", fallback.get(), target);
return Optional.empty();
}
@@ -484,9 +515,9 @@ public final class ReplyPushLoop {
} catch (RuntimeException e) {
boolean gone = e instanceof HerdrException he && "agent_not_found".equals(he.code());
if (gone) {
log.debug("push: lead {} no longer exists ({})", lead, e.toString());
log.debug("push: nudge target {} no longer exists ({})", lead, e.toString());
} else {
log.debug("push: liveness check for lead {} was inconclusive ({}); treating as live "
log.debug("push: liveness check for nudge target {} was inconclusive ({}); treating as live "
+ "rather than risk destroying a live binding", lead, e.toString());
}
return !gone;
@@ -505,11 +536,12 @@ public final class ReplyPushLoop {
public void onReplyQueued(String target) {
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
log.debug("push: no nudge target is known to be waiting on {}, skipping reminder", target);
return;
}
boolean mayDrainNudge = primaryRegistry.nudgeMayDrainFor(target).orElse(false);
pendingReplies.compute(target, (t, existing) ->
new ReplyEntry(lead.get(), existing == null ? 0 : existing.nudgeCount()));
new ReplyEntry(lead.get(), existing == null ? 0 : existing.nudgeCount(), mayDrainNudge));
startOrCoalesce(lead.get());
}
@@ -533,7 +565,7 @@ public final class ReplyPushLoop {
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
log.debug("push: no nudge target is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
return;
}
@@ -568,11 +600,13 @@ public final class ReplyPushLoop {
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
log.debug("push: no nudge target is known to be waiting on {}'s question (turnId {}), skipping nudge",
target, turnId);
return;
}
pendingQuestions.put(turnId, new PendingQuestion(turnId, ticket, target, lead.get(), question, 0));
boolean mayAnswerNudge = primaryRegistry.nudgeMayAnswerFor(target).orElse(false);
pendingQuestions.put(turnId,
new PendingQuestion(turnId, ticket, target, lead.get(), question, 0, mayAnswerNudge));
startOrCoalesce(lead.get());
}
@@ -596,7 +630,7 @@ public final class ReplyPushLoop {
for (String target : targets) {
var lead = resolveLiveLead(target);
if (lead.isEmpty()) {
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
log.warn("push: backend incident {} has no known nudge target for worker {}", incidentId, target);
continue;
}
targetsByLead.computeIfAbsent(lead.get(), _ -> new ArrayList<>()).add(target);
@@ -632,10 +666,10 @@ public final class ReplyPushLoop {
/** Start a reminder schedule for {@code lead}, or join the one already running. */
private void startOrCoalesce(String lead) {
if (activeLeads.putIfAbsent(lead, Boolean.TRUE) != null) {
log.debug("push: reminder loop already active for lead {}, work coalesced in", lead);
log.debug("push: reminder loop already active for nudge target {}, work coalesced in", lead);
return;
}
log.debug("push: starting reminder loop for lead {}", lead);
log.debug("push: starting reminder loop for nudge target {}", lead);
scheduleNext(lead);
}
@@ -740,11 +774,11 @@ public final class ReplyPushLoop {
|| pendingIncidentKeysFor(lead).stream().anyMatch(i -> !incidentsBefore.contains(i))
|| pendingUnmappedTargetKeysFor(lead).stream().anyMatch(i -> !unmappedTargetsBefore.contains(i));
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
log.debug("push: new work for nudge target {} raced the reminder loop's stop — restarting", lead);
scheduleNext(lead);
return;
}
log.debug("push: reminder loop ended for lead {}", lead);
log.debug("push: reminder loop ended for nudge target {}", lead);
}
/** Send one combined nudge covering everything currently pending for {@code lead}. */
@@ -761,13 +795,13 @@ public final class ReplyPushLoop {
List<PendingUnmappedTarget> unmappedTargets = pendingUnmappedTargetsFor(lead);
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty() && incidents.isEmpty()
&& unmappedTargets.isEmpty()) {
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
log.debug("push: pending work for nudge target {} drained before the nudge could be sent", lead);
return;
}
String nudge = formatNudge(replyTargets, tickets, questions, incidents, unmappedTargets);
try {
agents.send(lead, nudge);
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
log.debug("push: nudge sent to nudge target {} (reply {}/{}, ticket {}/{}, question {}/{}; "
+ "{} reply target(s), {} ticket(s), {} question(s))",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders,
@@ -784,7 +818,7 @@ public final class ReplyPushLoop {
}
}
} catch (RuntimeException e) {
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
log.warn("push: failed to nudge target {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders, e.toString());
}
@@ -802,7 +836,8 @@ public final class ReplyPushLoop {
List<PendingQuestion> questions, List<PendingIncident> incidents,
List<PendingUnmappedTarget> unmappedTargets) {
for (String target : replyTargets) {
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
pendingReplies.computeIfPresent(target,
(t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1, e.mayDrainNudge()));
}
for (PendingTicket ticket : tickets) {
pendingTickets.computeIfPresent(ticket.ticket(),
@@ -811,7 +846,7 @@ public final class ReplyPushLoop {
for (PendingQuestion question : questions) {
pendingQuestions.computeIfPresent(question.turnId(), (id, e) ->
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
e.nudgeCount() + 1));
e.nudgeCount() + 1, e.mayAnswerNudge()));
}
for (PendingIncident incident : incidents) {
pendingIncidents.computeIfPresent(incident.key(), (id, e) -> new PendingIncident(e.key(),
@@ -1,6 +1,8 @@
package dev.ltms.fleet.peer;
import java.util.Locale;
import java.util.stream.Collectors;
import java.util.stream.Stream;
/**
* What a member is <em>for</em> — the contract it runs under.
@@ -100,6 +102,11 @@ public enum MemberRole {
return null;
}
/** The wire name of every role, joined with {@code ", "} in declaration order. */
public static String wireNames() {
return Stream.of(values()).map(MemberRole::wireName).collect(Collectors.joining(", "));
}
/**
* Parse a config/wire spelling, case-insensitively.
*
@@ -118,14 +125,7 @@ public enum MemberRole {
}
}
}
StringBuilder valid = new StringBuilder();
for (MemberRole r : values()) {
if (!valid.isEmpty()) {
valid.append(", ");
}
valid.append(r.wireName());
}
throw new IllegalArgumentException(
"unknown member role '" + s + "'; valid roles are: " + valid);
"unknown member role '" + s + "'; valid roles are: " + wireNames());
}
}
@@ -285,13 +285,26 @@ public final class FleetApp {
/**
* As {@link #permitsFor(Principal, Authz.Action, String, Predicate)}, also threading the
* classifier an observer's {@code SEND} is checked against; pass {@link #auth}'s own
* {@code sendableObserverTarget()} to exercise the real production gate, as {@link #allow}
* does.
* {@code observerSendTarget()} to exercise the real production gate, as {@link #allow} does.
*/
static boolean permitsFor(Principal caller, Authz.Action action, String target,
Predicate<String> knownLeadOrCollaborator,
Predicate<String> knownObserverTarget) {
return Authz.permits(caller, action, target, knownLeadOrCollaborator, knownObserverTarget);
Predicate<String> observerSendTarget) {
return Authz.permits(caller, action, target, knownLeadOrCollaborator, observerSendTarget);
}
/**
* As {@link #permitsFor(Principal, Authz.Action, String, Predicate, Predicate)}, also
* threading the classifier an observer's or a collaborator's {@code TASK_READ} is checked
* against; pass a real {@code MessageService#ownsTicket(String, String)}-backed predicate to
* exercise the real production gate for a ticket poll, as {@link #allow} does.
*/
static boolean permitsFor(Principal caller, Authz.Action action, String target,
Predicate<String> knownLeadOrCollaborator,
Predicate<String> observerSendTarget,
Predicate<String> ticketOwnedByCaller) {
return Authz.permits(caller, action, target, knownLeadOrCollaborator, observerSendTarget,
ticketOwnedByCaller);
}
/**
@@ -303,11 +316,22 @@ public final class FleetApp {
* yours" (a worker reaching for another worker's session, or for orchestration).
*/
private boolean allow(Context ctx, Authz.Action action, String target) {
return allow(ctx, action, target, Authz.NO_OWNED_TICKET);
}
/**
* As {@link #allow(Context, Authz.Action, String)}, with a real classifier for an observer's
* or a collaborator's {@code TASK_READ} — the only call site that can supply one is a ticket
* poll, which knows the ticket id and the caller's owner key.
*/
private boolean allow(Context ctx, Authz.Action action, String target,
Predicate<String> ticketOwnedByCaller) {
if (auth == null) {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (permitsFor(caller, action, target, auth.knownLeadOrCollaborator(), auth.sendableObserverTarget())) {
if (permitsFor(caller, action, target, auth.knownLeadOrCollaborator(), auth.observerSendTarget(),
ticketOwnedByCaller)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS
&& action != Authz.Action.TASK_READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
@@ -718,9 +742,15 @@ public final class FleetApp {
content = FleetMcp.attributeIfObserver(caller, content);
timeout = Math.clamp(timeout, 1, MAX_MESSAGE_TIMEOUT_MS);
// Whether THIS caller's own Authz.Action.DRAIN is granted — a TIMED_OUT_*/BUSY receipt
// must not name GET /sessions/{id}/replies when that same caller would be refused it.
// Mirrors allow()'s own legacy bypass: with no CallerResolver configured, that route is
// not gated at all, so every caller may reach it.
boolean mayDrainPoll = auth == null || Authz.permits(caller, Authz.Action.DRAIN, id);
// Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout, callerOwner), timeout);
writeReply(ctx, id, messages.answer(turnId, content, timeout, callerOwner), timeout, mayDrainPoll);
return;
}
@@ -732,7 +762,7 @@ public final class FleetApp {
}
try {
writeReply(ctx, id, messages.send(id, content, timeout, callerOwner), timeout);
writeReply(ctx, id, messages.send(id, content, timeout, callerOwner), timeout, mayDrainPoll);
} catch (HerdrException e) {
herdrError(ctx, e);
}
@@ -742,8 +772,13 @@ public final class FleetApp {
* Render a {@link MessageService.Reply} onto the response — shared by a normal send and a
* fleet_ask answer. A structured/scraped completion is 200; a worker's mid-turn question a 202
* (with its {@code turnId}); a stale answer a 409; every other non-terminal outcome a typed 202.
*
* @param mayDrainPoll whether this caller's {@code Authz.Action.DRAIN} is granted — a
* TIMED_OUT or BUSY {@code detail} names {@code GET /sessions/{id}/replies}
* only when this is {@code true}, since that route shares the same grant.
*/
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout) {
private void writeReply(Context ctx, String id, MessageService.Reply reply, long timeout,
boolean mayDrainPoll) {
switch (reply.outcome()) {
case QUESTION -> ctx.status(202).json(Map.of(
"sessionId", id, "status", "question",
@@ -792,7 +827,10 @@ public final class FleetApp {
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
&& reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
: "no reply within " + timeout + "ms; " + MessageService.NO_TICKET_NO_RESEND + "; "
+ (mayDrainPoll
? "GET /sessions/" + id + "/replies drains its inbox for a late reply"
: "a late reply cannot be recovered on this channel — ask the lead")));
}
}
@@ -929,11 +967,15 @@ public final class FleetApp {
/** Poll an async (wait:false) delegation by ticket. 404 for an unknown/expired ticket. */
private void taskStatus(Context ctx) {
if (!allow(ctx, routeAction("GET /tasks/{ticket}"), null)) {
String ticket = ctx.pathParam("ticket");
Principal caller = ctx.attribute(CALLER);
String callerOwner = caller == null ? null : caller.ownerKey();
// The gate must see the ticket id, not a literal null -- an observer's or a
// collaborator's grant is confined to a ticket its own caller created.
if (!allow(ctx, routeAction("GET /tasks/{ticket}"), ticket, t -> messages.ownsTicket(t, callerOwner))) {
return;
}
Principal caller = ctx.attribute(CALLER);
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.ownerKey());
MessageService.TaskView v = messages.poll(ticket, callerOwner);
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
@@ -125,6 +125,7 @@ class FleetdAssemblyTurnRegistrarBehaviouralTest {
primary:
tab: "lead: primary"
profile: sonnet
workspace: "ltms"
profiles:
sonnet:
subscription: true
@@ -91,6 +91,11 @@ class FleetdBackendErrorSinkTest {
return MAPPER.createObjectNode().set("agent", MAPPER.createObjectNode()
.put("terminal_id", "term_primary").put("agent_status", "idle"));
}
if ("agent.read".equals(method)) {
// The lead-nudge paths read the input box before pasting into it.
return MAPPER.createObjectNode().set("read",
MAPPER.createObjectNode().put("text", FakeHerdr.IDLE_PROMPT_CARET));
}
if ("agent.prompt".equals(method)) {
prompts.add(params);
sendLatch.countDown();
@@ -171,6 +171,7 @@ class FleetdLeadContextSourceWindowAssemblyTest {
%s:
tab: "%s"
profile: %s
workspace: "ltms"
profiles:
%s:
subscription: true
@@ -172,6 +172,7 @@ class FleetdLeadRolloverAssemblyTest {
opus:
tab: "lead: opus"
cwd: "%s"
workspace: "ltms"
leadRollover:
handoverPath: handover.md
requireOperatorConfirm: false
@@ -203,6 +204,7 @@ class FleetdLeadRolloverAssemblyTest {
tab: "lead: opus"
cwd: "%s"
profile: opus
workspace: "ltms"
profiles:
opus:
subscription: true
@@ -334,6 +336,11 @@ class FleetdLeadRolloverAssemblyTest {
assertNotNull(rollover, "leadRollover: is present in this test's config, so a real "
+ "LeadRollover must have been built");
// The readiness gate needs the fresh pane's Claude to have connected the bridge MCP; no real
// Claude boots here, so stand in for that connect. term_new_1 is the first terminal this
// fake's agent.start hands out.
runtime.sessions().asPresence().markPresent("term_new_1");
LeadRollover.PendingRollover pending = rollover.open("term_a", "bootstrapText relaunch test");
Thread.sleep(50);
Files.writeString(Path.of(pending.handoverPath()), "handover content for bootstrapText test");
@@ -398,7 +405,8 @@ class FleetdLeadRolloverAssemblyTest {
WorkspaceControl spaces = new WorkspaceControl(new FakeHerdr());
LeadLauncher launcher = new LeadLauncher(agents, spaces, config.get());
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of);
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of,
_ -> true);
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
@@ -83,7 +83,7 @@ class FleetdLeadRolloverWorkspaceLookupTest {
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
() -> Map.of("term_opus", "opus"));
() -> Map.of("term_opus", "opus"), _ -> true);
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object");
@@ -114,7 +114,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of);
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of,
_ -> true);
assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
@@ -155,7 +156,7 @@ class FleetdLeadRolloverWorkspaceLookupTest {
Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
() -> liveLeadTerminals);
() -> liveLeadTerminals, _ -> true);
assertNotNull(rollover);
// The lead is "discovered" only now — mutate the SAME backing map the supplier reads from.
@@ -143,6 +143,7 @@ class FleetdLeadSeatAssemblyTest {
opus:
tab: "lead: opus"
profile: sonnet
workspace: "ltms"
profiles:
sonnet:
subscription: true
@@ -47,6 +47,28 @@ class AuthzTest {
+ "relies on this gate refusing it first");
}
/**
* {@code fleet_inbox} collects the mail queued for the caller's own pane, so it is gated on
* terminal ownership and on nothing else: every role may do it for itself, and no role may do
* it for another pane.
*/
@Test
void collectingAnInboxIsOnlyEverForTheCallersOwnPane() {
for (Principal self : new Principal[]{WORKER_A, ARCH_DESIGN, COLLABORATOR, OBSERVER}) {
assertTrue(Authz.permits(self, INBOX, self.terminal()),
self.role() + " must be able to collect the mail for its own pane");
assertFalse(Authz.permits(self, INBOX, "term_someone_else"),
self.role() + " must not be able to collect another pane's mail");
}
// A lead carries a pane too, so it collects its own mail on the same rule.
assertTrue(Authz.permits(Principal.leader("opus", "term_lead", 800), INBOX, "term_lead"));
// The unnamed primary owns no pane, so there is no inbox it could be asking for.
assertFalse(Authz.permits(PRIMARY, INBOX, "term_a"),
"a caller with no pane of its own has no inbox to collect");
assertFalse(Authz.permits(WORKER_A, INBOX, null),
"a missing terminal must never match an owner");
}
@Test
void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
@@ -237,18 +259,39 @@ class AuthzTest {
}
/**
* Every action denied to a collaborator, asserted denied even when the classifier would
* accept any target — proving none of these is actually gated on the classifier at all.
* Every action denied to a collaborator, asserted denied even when the
* {@code knownLeadOrCollaborator} classifier would accept any target — proving none of these
* is actually gated on that classifier at all. {@code TASK_READ} is excluded here and given
* its own matrix below, since — unlike every action in this loop — its grant is conditional
* on the separate ticket-ownership classifier, not on this one.
*/
@Test
void aCollaboratorIsDeniedLifecycleCoordinationAndTicketPolling() {
void aCollaboratorIsDeniedLifecycleAndCoordination() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN, HANDOVER, ANSWER, COORD_SEND,
COORD_READ, TASK_READ}) {
COORD_READ}) {
assertFalse(Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
"a collaborator must not " + a + " even when the classifier accepts every target");
}
}
/**
* {@code TASK_READ} is conditional for a collaborator too (fleetd #804), on the same
* ticket-ownership classifier an observer's grant uses: it may poll a ticket its own
* {@code fleet_send(wait:false)} created, and nothing else.
*/
@Test
void aCollaboratorMayTaskReadOnlyATicketTheClassifierSaysItOwns() {
assertTrue(Authz.permits(COLLABORATOR, TASK_READ, "own-ticket", Authz.NO_KNOWN_LEAD_OR_COLLABORATOR,
Authz.NO_OBSERVER_SEND_TARGET, target -> true),
"the classifier accepting the ticket as the caller's own must grant TASK_READ");
assertFalse(Authz.permits(COLLABORATOR, TASK_READ, "someone_elses_ticket",
Authz.NO_KNOWN_LEAD_OR_COLLABORATOR, Authz.NO_OBSERVER_SEND_TARGET, target -> false),
"the classifier refusing the ticket must deny TASK_READ");
assertFalse(Authz.permits(COLLABORATOR, TASK_READ, "own-ticket"),
"the three-argument convenience form fails closed, so TASK_READ is refused without "
+ "a real classifier");
}
@Test
void aCollaboratorIsNotCountedAsPrimaryWorkerOrArchitect() {
assertFalse(COLLABORATOR.isPrimary());
@@ -288,16 +331,17 @@ class AuthzTest {
}
/**
* Every action beyond READ/METRICS/REPLY/ASK/SEND, asserted denied for an observer —
* including {@code TASK_READ}, which is the entire point of this role: an unconfigured pane
* must not be able to poll a ticket or read another session's status. {@code SEND} is excluded
* here and given its own matrix below, since — unlike every action in this loop — its grant is
* conditional on the target, not fixed.
* Every action beyond READ/METRICS/REPLY/ASK/INBOX/SEND, asserted denied for an observer.
* The exempt set is the three only-as-itself actions plus the two open reads. {@code SEND}
* and {@code TASK_READ} are excluded here and given their own matrices below, since — unlike
* every action in this loop — each one's grant is conditional on more than the caller's role
* alone.
*/
@Test
void anObserverIsDeniedEverythingBeyondReadMetricsReplyAskAndSend() {
void anObserverIsDeniedEverythingBeyondReadMetricsReplyAskInboxSendAndTaskRead() {
for (Authz.Action a : Authz.Action.values()) {
if (a == READ || a == METRICS || a == REPLY || a == ASK || a == SEND) {
if (a == READ || a == METRICS || a == REPLY || a == ASK || a == INBOX || a == SEND
|| a == TASK_READ) {
continue;
}
assertFalse(Authz.permits(OBSERVER, a, "term_observer", target -> true),
@@ -305,6 +349,49 @@ class AuthzTest {
}
}
/**
* {@code TASK_READ} is conditional for an observer, not fixed: it may poll a ticket its own
* {@code fleet_send(wait:false)} created, confined by {@code ticketOwnedByCaller}, and nothing
* else — a session's status is not ticket-scoped, so no call site can ever supply a predicate
* that grants it, and the 3-/4-/5-argument convenience forms stay fail-closed for it too.
*/
@Test
void anObserverMayTaskReadOnlyATicketTheClassifierSaysItOwns() {
assertTrue(Authz.permits(OBSERVER, TASK_READ, "own-ticket", Authz.NO_KNOWN_LEAD_OR_COLLABORATOR,
Authz.NO_OBSERVER_SEND_TARGET, target -> true),
"the classifier accepting the ticket as the caller's own must grant TASK_READ");
assertFalse(Authz.permits(OBSERVER, TASK_READ, "someone_elses_ticket",
Authz.NO_KNOWN_LEAD_OR_COLLABORATOR, Authz.NO_OBSERVER_SEND_TARGET, target -> false),
"the classifier refusing the ticket must deny TASK_READ");
assertFalse(Authz.permits(OBSERVER, TASK_READ, "own-ticket"),
"the three-argument convenience form fails closed, so TASK_READ is refused without "
+ "a real classifier");
assertFalse(Authz.permits(OBSERVER, TASK_READ, "own-ticket", Authz.NO_KNOWN_LEAD_OR_COLLABORATOR),
"the four-argument form still fails closed for TASK_READ");
assertFalse(Authz.permits(OBSERVER, TASK_READ, "own-ticket", Authz.NO_KNOWN_LEAD_OR_COLLABORATOR,
Authz.NO_OBSERVER_SEND_TARGET),
"the five-argument form still fails closed for TASK_READ, since it supplies no "
+ "ticketOwnedByCaller classifier either");
}
/**
* Control for the test above: every other action's result for an observer does not move when
* the ticket-ownership classifier does. Only {@code TASK_READ} is wired to it.
*/
@Test
void theTicketOwnershipClassifierMovesOnlyTaskReadForAnObserver() {
for (Authz.Action a : Authz.Action.values()) {
if (a == TASK_READ) {
continue;
}
assertEquals(
Authz.permits(OBSERVER, a, "term_observer"),
Authz.permits(OBSERVER, a, "term_observer", Authz.NO_KNOWN_LEAD_OR_COLLABORATOR,
Authz.NO_OBSERVER_SEND_TARGET, target -> true),
a + " must not depend on the ticket-ownership classifier at all");
}
}
@Test
void anObserverIsNotCountedAsAnyOtherRole() {
assertFalse(OBSERVER.isPrimary());
@@ -912,42 +912,65 @@ class CallerResolverTest {
assertFalse(r.knownLeadOrCollaborator().test("term_a"));
}
// ── fleetd #743: sendableObserverTarget() reads the same maps and functions resolve() does ────
// ── observerSendTarget() reads the same maps and functions resolve() does, in the same order ──
/**
* A terminal this resolver recognises as none of the privileged roles is exactly the one
* {@code resolve} would itself hand back {@link Role#OBSERVER} for.
*/
@Test
void sendableObserverTargetIsTrueForATerminalKnownAsNoOtherRole() {
void observerSendTargetIsTrueForATerminalKnownAsNoOtherRole() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_lead", "opus-5.0"), new MemberRegistry(null),
t -> "term_worker".equals(t) ? MemberRole.DEV : null,
() -> Map.of("term_collab", "ops"));
assertTrue(r.sendableObserverTarget().test("term_other"));
assertTrue(r.observerSendTarget().test("term_other"));
}
/**
* A configured lead's own terminal is reachable: an observer may open a conversation with a
* lead, and this is the classifier that grant is checked against.
*/
@Test
void sendableObserverTargetIsFalseForALeadTerminal() {
void observerSendTargetIsTrueForALeadTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_lead", "opus-5.0"), new MemberRegistry(null), t -> null, Map::of);
assertFalse(r.sendableObserverTarget().test("term_lead"),
"a lead's own terminal must never be a sendable observer target");
assertTrue(r.observerSendTarget().test("term_lead"),
"a lead's own terminal must be reachable from an observer pane");
}
/**
* A pane named as a lead AND bound to an architect slot resolves as the lead, because
* {@code resolve} reads the lead map first — so the classifier must accept it, or a pane's
* resolved role and its reachability would disagree.
*/
@Test
void observerSendTargetIsTrueForALeadTerminalThatIsAlsoABoundArchitectSlot() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"),
boundMembers("architect:lead-designer", MemberRole.ARCHITECT), t -> null, Map::of);
assertEquals(Role.PRIMARY, r.resolve("127.0.0.1", 42, null).role(),
"premise: the lead map is read before the architect registry");
assertTrue(r.observerSendTarget().test("term_a"));
}
@Test
void sendableObserverTargetIsFalseForACollaboratorTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
void observerSendTargetIsFalseForACollaboratorTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_lead", "opus-5.0"),
new MemberRegistry(null), t -> null, () -> Map.of("term_collab", "ops"));
assertFalse(r.sendableObserverTarget().test("term_collab"),
"a collaborator's own terminal must never be a sendable observer target");
assertFalse(r.observerSendTarget().test("term_collab"),
"a collaborator's own terminal must never be reachable from an observer pane");
// CONTROL: the same wiring, a target recognised as no configured role at all.
assertTrue(r.observerSendTarget().test("term_other"));
}
@Test
void sendableObserverTargetIsFalseForALiveSpawnedMembersTerminal() {
void observerSendTargetIsFalseForALiveSpawnedMembersTerminal() {
// Covers both a worker and an architect: spawnedMemberRole.apply(target) is non-null for
// either, and resolve() never falls through to OBSERVER once it is.
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
@@ -958,26 +981,47 @@ class CallerResolverTest {
default -> null;
}, Map::of);
assertFalse(r.sendableObserverTarget().test("term_worker"));
assertFalse(r.sendableObserverTarget().test("term_architect"));
assertFalse(r.observerSendTarget().test("term_worker"));
assertFalse(r.observerSendTarget().test("term_architect"));
// CONTROL: the same wiring, a target the member lookup above answers null for.
assertTrue(r.observerSendTarget().test("term_other"));
}
/**
* A lead map entry does not rescue a terminal a live spawned member occupies: the member
* lookup runs first, exactly as in {@code resolve}.
*/
@Test
void observerSendTargetIsFalseForASpawnedMemberOnATerminalTheLeadMapAlsoNames() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_worker", "opus-5.0"), new MemberRegistry(null),
t -> "term_worker".equals(t) ? MemberRole.DEV : null, Map::of);
assertFalse(r.observerSendTarget().test("term_worker"));
// CONTROL: the same wiring, the same lead map, a terminal no member occupies.
assertTrue(r.observerSendTarget().test("term_other"));
}
@Test
void sendableObserverTargetIsFalseForABoundArchitectSlotWithNoLiveMember() {
void observerSendTargetIsFalseForABoundArchitectSlotWithNoLiveMember() {
// The edge case resolve() itself carries: a terminal bound to a configured architect slot
// but with no live spawned-member session yet.
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
boundMembers("architect:lead-designer", MemberRole.ARCHITECT), t -> null, Map::of);
assertFalse(r.sendableObserverTarget().test("term_a"));
assertFalse(r.observerSendTarget().test("term_a"));
// CONTROL: the same wiring, a terminal the bind above never touched.
assertTrue(r.observerSendTarget().test("term_other"));
}
@Test
void sendableObserverTargetIsFalseForANullTarget() {
void observerSendTargetIsFalseForANullTarget() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, Map::of);
assertFalse(r.sendableObserverTarget().test(null));
assertFalse(r.observerSendTarget().test(null));
// CONTROL: the same wiring, a non-null target.
assertTrue(r.observerSendTarget().test("term_other"));
}
}
@@ -114,6 +114,7 @@ class ConfigRefTopLevelReportingCoverageTest {
// fleetd #480: leadRollover joins placement/memberCredentials/memberLoginShell as
// hot-excluded — never compared by any changed*Keys method, so it stays null like them.
v.put("leadRollover", null);
v.put("promptBoxGateEnabled", true);
assertNamesMatchComponents(v);
return v;
}
@@ -159,6 +160,7 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-b"))));
v.put("leadRollover", null);
v.put("promptBoxGateEnabled", false);
assertNamesMatchComponents(v);
return v;
}
@@ -537,12 +537,13 @@ class FleetConfigTest {
}
/**
* CB-579: {@code tab} is the only field a lead's identity depends on now, so it is required
* whether the entry is creatable or recognise-only — without it the entry can never be found.
* A lead's tab label is a fixed constant, not a per-entry field, so an entry with no {@code
* tab:} is the normal case — it is still found by that constant label in its own {@code
* workspace}, not refused as useless.
*/
@Test
void aLeadWithNoTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("useless-lead.yaml");
void aLeadWithNoTabIsAcceptedAndFoundByTheFixedLabel(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-tab-lead.yaml");
Files.writeString(f, """
bind:
port: 8080
@@ -553,9 +554,9 @@ class FleetConfigTest {
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateMembers);
assertTrue(e.getMessage().contains("ghost"), "the message must name the useless entry");
assertTrue(e.getMessage().contains("tab:"), "the message must say what is missing");
assertDoesNotThrow(cfg::validateMembers);
assertEquals(List.of(FleetConfig.Leader.LEAD_TAB_LABEL),
cfg.fleet().leaders().get("ghost").acceptedLabels());
}
/**
@@ -775,22 +776,45 @@ class FleetConfigTest {
}
/**
* fleetd #677: identity is matched on a lead's exact {@code tab} alone, so two leads sharing
* one tab means only one of them is ever found — the guard must catch this independently of
* the member-template checks above.
* A lead's tab label is fixed, so two leads sharing one {@code workspace} would both resolve
* to the one tab named {@code lead} there — the guard must catch this independently of the
* member-template checks above.
*/
@Test
void twoLeadsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-tab.yaml");
void twoLeadersSharingTheSameWorkspaceRefuseToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-workspace.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "shared tab"
workspace: "shared"
sonnet:
tab: "shared tab"
workspace: "shared"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
assertTrue(e.getMessage().contains("shared"), "the message must name the shared workspace");
}
/** The workspace collision check is case-insensitive, matching how spaces are looked up. */
@Test
void twoLeadersSharingTheSameWorkspaceInDifferentCaseRefuseToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-workspace-case.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
workspace: "Shared"
sonnet:
workspace: "shared"
""");
FleetConfig cfg = FleetConfig.load(f);
@@ -800,52 +824,77 @@ class FleetConfigTest {
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
}
/**
* fleetd #693: the guard matches tabs case-insensitively, because
* {@code LeadTabScanner} keys its tab map on a lowercased label — two tabs differing only in
* case collide there too, and the guard must catch that independently of the exact-match case
* above.
*/
/** Control for the two tests above: distinct workspaces load cleanly, with no {@code tab:} at all. */
@Test
void twoLeadsSharingTheSameTabInDifferentCaseRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-tab-case.yaml");
void twoLeadersWithDistinctWorkspacesAndNoTabAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("distinct-workspaces.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "Shared Tab"
workspace: "space-opus"
sonnet:
tab: "shared tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
}
/** Control for {@link #twoLeadsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void twoLeadsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("distinct-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "opus tab"
sonnet:
tab: "sonnet tab"
workspace: "space-sonnet"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/**
* A member tab-label template that can render exactly as the fixed lead tab label would let a
* member's own tab be read back as a lead — refused outright, with no lead needing to be
* configured at all.
*/
@Test
void aFleetTabLabelTemplateThatCanRenderAsTheFixedLeadTabLabelRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("template-renders-as-lead.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
tabLabel: "lead"
leaders:
opus:
workspace: "fleet"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("fleet.tabLabel"),
"the message must name the offending template");
assertTrue(e.getMessage().contains(FleetConfig.Leader.LEAD_TAB_LABEL),
"the message must name the fixed lead tab label it collides with");
}
/**
* A collaborator's {@code tab} equal to the fixed lead tab label would shadow a lead sharing
* that space — refused outright.
*/
@Test
void aCollaboratorTabEqualToTheFixedLeadTabLabelRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborator-is-lead.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
impostor:
tab: "lead"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("impostor"), "the message must name the offending collaborator");
assertTrue(e.getMessage().contains(FleetConfig.Leader.LEAD_TAB_LABEL),
"the message must name the fixed lead tab label it collides with");
}
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
/**
@@ -874,8 +923,12 @@ class FleetConfigTest {
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
/**
* A lead is found by its fixed tab label regardless of its own {@code tab} field, so a
* {@code fleet.leaders} entry with no {@code tab} must still arm the guard.
*/
@Test
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
void aPanePlacedProfileWithNoLeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab.yaml");
Files.writeString(f, """
bind:
@@ -888,9 +941,59 @@ class FleetConfigTest {
opus:
profile: gx10
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
assertTrue(e.getMessage().contains("the only fix when a lead triggered this"),
"the message must say placement: tab is the only fix for a lead");
assertFalse(e.getMessage().contains("remove the tab from every fleet.leaders"),
"the message must not send the operator in a circle by advising a tab: removal");
}
/**
* {@code placement:} is optional, and {@link FleetConfig.Profile}'s own compact constructor
* defaults an absent or blank value to {@code "tab"}, so a profile naming no placement at all
* is tab-placed and the guard must not fire for it.
*/
@Test
void aProfileWithNoPlacementKeyDefaultsToTabPlacementAndIsAllowed(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("no-placement-key.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10: {}
fleet:
leaders:
opus:
profile: gx10
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.profiles().get("gx10").tabPlacement(),
"Profile's compact constructor defaults an absent placement to \"tab\"");
assertDoesNotThrow(cfg::validatePanePlacementAgainstLeadTabs,
"a profile with no placement: key is tab-placed, not pane-placed");
}
/** Control: no {@code fleet.leaders} entry and no collaborator tab still starts fine. */
@Test
void aPanePlacedProfileWithAnEmptyFleetBlockIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-empty-fleet.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet: {}
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
"no fleet.leaders entry and no collaborator tab means pane placement is safe");
}
@Test
@@ -104,6 +104,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
// here proves it, rather than leaving it null and proving nothing.
v.put("leadRollover", new FleetConfig.LeadRollover(
"/handover/guard.md", true, 3600, 20, 45, "read the handover file"));
v.put("promptBoxGateEnabled", true);
assertNamesMatchComponents(v);
return v;
}
@@ -24,6 +24,41 @@ public final class FakeHerdr implements HerdrClient {
/** The foreground PID of the one agent pane (term_a) in the canned {@code pane.process_info}. */
public static final long WORKER_PID = 4242;
/**
* A {@code detection} region of the current Claude Code TUI, whose input box is empty: a caret
* line between two rules, above the footer.
*/
public static final String IDLE_PROMPT_CARET = """
──────────────────────── lead: opus ─
❯
─────────────────────────────────────
lead: opus · Opus 5 (1M context) · ~/LTMS/claude-bridge
⏵⏵ auto mode on (shift+tab to cycle) · ← 1 agent""";
/** The same region with the operator's unsubmitted line still at the caret. */
public static final String DRAFTED_PROMPT_CARET = """
──────────────────────── lead: opus ─
❯ yes, send it to lead: opus
─────────────────────────────────────
lead: opus · Opus 5 (1M context) · ~/LTMS/claude-bridge
⏵⏵ auto mode on (shift+tab to cycle) · ← 1 agent""";
/** A {@code detection} region of an older TUI, which drew a bordered box, with that box empty. */
public static final String IDLE_PROMPT_BOX = """
⏺ done
╭────────────────────────╮
│ > │
╰────────────────────────╯
⏵⏵ auto mode on""";
/** The older TUI's bordered box, still holding the operator's unsubmitted line. */
public static final String DRAFTED_PROMPT_BOX = """
⏺ done
╭────────────────────────╮
│ > fix the issue when I │
╰────────────────────────╯
⏵⏵ auto mode on""";
private final ObjectMapper mapper = new ObjectMapper();
/**
* Thread-safe on purpose. Background loops — {@link dev.ltms.fleet.msg.ReplyPushLoop} and the
@@ -52,13 +87,18 @@ public final class FakeHerdr implements HerdrClient {
private String agentSendErrorCode = null;
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentStatusOverride; // see agentStatusForNextReads
private volatile int agentStatusOverrideReads; // how many agent.get reads it still covers
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String agentSessionId = null; // agent_session.value on agent.get; null = omitted
private volatile String readText = "worker transcript tail"; // canned agent.read output
/** Canned {@code detection}-source output, or {@code null} to serve {@link #readText} there too. */
private volatile String detectionText = null;
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
private volatile boolean turnStartsOnPrompt; // see turnStartsOnPrompt()
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
/** pane ids that {@link #paneGoneAfterClose} has opted into reporting gone — see that method. */
@@ -176,6 +216,26 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Report {@code status} from the next {@code reads} {@code agent.get} calls, then fall back to
* the steady state. {@code reads} counts raw {@code agent.get} calls, not
* {@code AgentControl.status()} calls, and one {@code status()} costs two: {@code resolveTarget}
* issues its own {@code agent.get} before the status read.
*/
private void agentStatusForNextReads(int reads, String status) {
this.agentStatusOverride = status;
this.agentStatusOverrideReads = reads;
}
/** The status this {@code agent.get} reports, consuming one read of a pending override. */
private String takeAgentStatus() {
if (agentStatusOverrideReads <= 0) {
return agentStatus;
}
agentStatusOverrideReads--;
return agentStatusOverride;
}
/**
* Set the detected agent kind ({@code "agent"} field) that {@code agent.get} reports —
* {@code null} models herdr not (or no longer) detecting a supported backend in the pane, e.g.
@@ -207,12 +267,27 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
/**
* The text {@code agent.read} returns (the CB-106 completion scrape). It serves the
* {@code detection} source as well unless {@link #detectionText} overrides that one.
*/
public FakeHerdr readText(String text) {
this.readText = text;
return this;
}
/**
* Override the text {@code agent.read} returns for the {@code detection} and {@code visible}
* sources only — the prompt/footer tail herdr uses for status detection and input-box probing, a
* different region from the transcript the other sources carry. Needed by a test whose subject
* reads the input box, since one {@link #readText} cannot be both a worker's transcript and a
* lead's empty prompt.
*/
public FakeHerdr detectionText(String text) {
this.detectionText = text;
return this;
}
/** Make delivery ({@code agent.prompt} / {@code agent.send_keys}) fail with this error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
@@ -257,6 +332,17 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Report {@code "working"} from the {@code agent.get} calls that follow each {@code agent.prompt},
* then fall back to {@link #agentStatus} — models the prompted agent starting a turn on the text
* it was just sent. A confirmation wait that polls the status right after a send sees that turn;
* a later wait for a settled pane is not blocked, and a second prompt arms it again, so a test
* may roll more than once. Leave it off to model herdr accepting a send that never starts a turn.
*/
public FakeHerdr turnStartsOnPrompt() {
this.turnStartsOnPrompt = true;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
@@ -331,6 +417,9 @@ public final class FakeHerdr implements HerdrClient {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.prompt failed",
agentSendErrorCode, null);
}
if (turnStartsOnPrompt) {
agentStatusForNextReads(2, "working");
}
yield mapper.readTree(("""
{"type":"agent_prompted","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
@@ -358,10 +447,15 @@ public final class FakeHerdr implements HerdrClient {
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"%s}}""")
.formatted(agentField, agentStatus, sessionField));
.formatted(agentField, takeAgentStatus(), sessionField));
}
case "agent.read" -> {
Object source = params instanceof Map<?, ?> m ? m.get("source") : null;
boolean probeSource = "detection".equals(source) || "visible".equals(source);
String text = probeSource && detectionText != null ? detectionText : readText;
yield mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", text))));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
if (onAgentStart != null) {
onAgentStart.run();
@@ -158,15 +158,24 @@ class LeadTabScannerTest {
return Map.of("lead: opus-5.0", "opus-5.0", "lead: gpt-sol-5.6", "gpt-sol-5.6");
}
/** {@code twoLeads()}'s own space — every lead-label fixture below lives here unless noted. */
private static final String MAIN_SPACE = "main";
/** Wraps a flat label → name map under one space, the shape {@link LeadTabScanner} now takes. */
private static Map<String, Map<String, String>> inSpace(String space, Map<String, String> labelToName) {
return Map.of(space, labelToName);
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> tabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, tabToName, Set.of("fleetd-workers"), TTL, clock::get);
return new LeadTabScanner(herdr, inSpace(MAIN_SPACE, tabToName), Set.of("fleetd-workers"),
TTL, clock::get);
}
private LeadTabScanner scannerWithCollaborators(TopologyHerdr herdr, Map<String, String> tabToName,
Map<String, String> collaboratorTabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, tabToName, collaboratorTabToName,
return new LeadTabScanner(herdr, inSpace(MAIN_SPACE, tabToName), collaboratorTabToName,
Set.of("fleetd-workers"), TTL, clock::get);
}
@@ -237,14 +246,61 @@ class LeadTabScannerTest {
.tab("w1:t2", "w1", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_worker");
LeadTabScanner s = new LeadTabScanner(herdr, Map.of("lead: opus-5.0", "opus-5.0"),
Set.of(), TTL, new AtomicLong()::get);
LeadTabScanner s = new LeadTabScanner(herdr,
inSpace("fleet", Map.of("lead: opus-5.0", "opus-5.0")), Set.of(), TTL,
new AtomicLong()::get);
assertEquals("opus-5.0", s.get().get("term_opus"),
"a lead sharing the members' workspace is still discovered — the label, not the "
+ "workspace, is what matches it");
}
// ── fleetd #770: space is the uniqueness boundary, not the label alone ──────────────────────
/**
* Two leads can share the exact same label (the fixed {@code lead} tab label) as long as they
* sit in different spaces — each tab resolves to its own space's lead, never the other one's.
*/
@Test
void aTabLabelledLeadResolvesToItsOwnSpacesLeadNotTheOtherSpaces() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("wa", "space-a")
.workspace("wb", "space-b")
.tab("wa:t1", "wa", "lead")
.tab("wb:t1", "wb", "lead")
.pane("wa:p1", "wa:t1", "term_a")
.pane("wb:p1", "wb:t1", "term_b");
Map<String, Map<String, String>> leadLabelsBySpace = Map.of(
"space-a", Map.of("lead", "alpha"),
"space-b", Map.of("lead", "beta"));
LeadTabScanner s = new LeadTabScanner(herdr, leadLabelsBySpace, Set.of(), TTL,
new AtomicLong()::get);
Map<String, String> leads = s.get();
assertEquals("alpha", leads.get("term_a"), "space-a's tab must resolve to space-a's lead");
assertEquals("beta", leads.get("term_b"), "space-b's tab must resolve to space-b's lead");
}
/**
* A lead's deprecated legacy {@code tab:} label is still matched, but only within that lead's
* own configured space — exactly the shape {@code FleetdAssembly} builds via {@code
* Leader.acceptedLabels()}.
*/
@Test
void aLegacyTabLabelStillResolvesWithinItsOwnSpace() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "fleet")
.tab("w1:t1", "w1", "lead: opus")
.pane("w1:p1", "w1:t1", "term_opus");
Map<String, Map<String, String>> leadLabelsBySpace =
Map.of("fleet", Map.of("lead", "opus", "lead: opus", "opus"));
LeadTabScanner s = new LeadTabScanner(herdr, leadLabelsBySpace, Set.of(), TTL,
new AtomicLong()::get);
assertEquals("opus", s.get().get("term_opus"),
"the deprecated tab label must still resolve this lead within its own space");
}
@Test
void aLabelWithNoConfiguredEntryIsIgnored() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
@@ -0,0 +1,200 @@
package dev.ltms.fleet.herdr;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Reading a Claude Code input box, so a paste-and-submit delivery never submits the operator's draft. */
class PromptBoxTest {
private static final String EMPTY = FakeHerdr.IDLE_PROMPT_CARET;
private static final String DRAFTED = FakeHerdr.DRAFTED_PROMPT_CARET;
/** An empty box: the pane draws a placeholder hint (its last submitted prompt) at the caret, faint. */
private static final String HINT_1 = "❯\u00a0\u001b[0m\u001b[2mstart md2trilium work for #2\u001b[0m";
/** Same shape as {@link #HINT_1}, a different placeholder hint. */
private static final String HINT_2 = "❯\u00a0\u001b[0m\u001b[2myes, push them\u001b[0m";
/** An empty box with no hint: the caret itself is drawn grey, with escape codes before the marker. */
private static final String EMPTY_GREY_CARET = "\u001b[0m\u001b[38;2;153;153;153m❯\u00a0\u001b[0m";
/** An empty box with no styling at all. */
private static final String EMPTY_PLAIN = "❯\u00a0";
// --- pure classification -------------------------------------------------
@Test
void anEmptyCaretLineIsAnEmptyBox() {
assertEquals(new PromptBox.Reading(PromptBox.State.EMPTY, 0), PromptBox.classify(EMPTY));
}
@Test
void aCaretLineHoldingTextIsADraftAndCountsItsCharacters() {
PromptBox.Reading reading = PromptBox.classify(DRAFTED);
assertEquals(PromptBox.State.DRAFT, reading.state());
assertEquals("yes,sendittolead:opus".length(), reading.characters(),
"padding does not count — only what the operator typed");
}
@Test
void aBorderedBoxIsReadToo() {
assertEquals(PromptBox.State.EMPTY, PromptBox.classify(FakeHerdr.IDLE_PROMPT_BOX).state(),
"an older TUI draws a bordered box, and its panes must still be readable");
assertEquals(PromptBox.State.DRAFT, PromptBox.classify(FakeHerdr.DRAFTED_PROMPT_BOX).state());
}
@Test
void aSingleTypedCharacterIsADraft() {
assertEquals(PromptBox.State.DRAFT, PromptBox.classify("❯ f").state());
assertEquals(PromptBox.State.DRAFT, PromptBox.classify("│ > f │").state());
}
@Test
void aCursorBlockInAnOtherwiseEmptyBoxIsEmpty() {
assertEquals(PromptBox.State.EMPTY, PromptBox.classify("❯ █").state(),
"a terminal capture may leave the cursor cell in an empty box");
assertEquals(PromptBox.State.EMPTY, PromptBox.classify("│ > █ │").state());
}
@Test
void theBoxLineIsReadToItsEndWhateverFollowsIt() {
assertEquals(PromptBox.State.DRAFT, PromptBox.classify("❯ half a line\n ⏵⏵ auto mode on").state());
assertEquals(PromptBox.State.EMPTY, PromptBox.classify("❯\n ⏵⏵ auto mode on").state());
}
@Test
void theLastBoxLineOnThePaneIsTheLiveOne() {
assertEquals(PromptBox.State.DRAFT,
PromptBox.classify("❯ an earlier prompt\n⏺ its answer\n❯ typing now").state(),
"the probed region carries scrollback, so earlier prompts sit above the live box");
assertEquals(PromptBox.State.EMPTY,
PromptBox.classify("❯ an earlier prompt\n⏺ its answer\n❯").state());
}
@Test
void aMarkerPartWayAlongALineIsNotABox() {
assertEquals(PromptBox.State.UNREADABLE, PromptBox.classify("⏺ type ❯ to get a prompt").state(),
"a caret the operator quoted is transcript text, not an input box");
assertEquals(PromptBox.State.EMPTY, PromptBox.classify("⏺ type ❯ to get a prompt\n❯").state(),
"and it must not shadow the real box further down");
}
@Test
void aPaneWithNoBoxIsUnreadable() {
assertEquals(PromptBox.State.UNREADABLE, PromptBox.classify("garbled ansi noise").state());
assertEquals(PromptBox.State.UNREADABLE, PromptBox.classify("").state());
assertEquals(PromptBox.State.UNREADABLE, PromptBox.classify(null).state());
}
@Test
void aGeneratingTurnIsUnreadableEvenWithAnEmptyBox() {
assertEquals(PromptBox.State.UNREADABLE,
PromptBox.classify(EMPTY + "\n ✳ Thinking… (12s · esc to interrupt)").state());
}
@Test
void aGeneratingMarkerInScrollbackAboveTheBoxDoesNotMakeThePaneUnreadable() {
assertEquals(PromptBox.State.EMPTY,
PromptBox.classify(" ✳ Thinking… (12s · esc to interrupt)\n⏺ done\n" + EMPTY).state(),
"that marker survives in scrollback, and holding on it would hold every delivery forever");
}
@Test
void aPlaceholderHintReadsAsAnEmptyBox() {
assertEquals(new PromptBox.Reading(PromptBox.State.EMPTY, 0), PromptBox.classify(HINT_1),
"the hint is the pane's own last prompt, drawn faint — it is not the operator's typing");
assertEquals(new PromptBox.Reading(PromptBox.State.EMPTY, 0), PromptBox.classify(HINT_2));
}
@Test
void anEmptyBoxWithAGreyCaretReadsAsEmpty() {
assertEquals(new PromptBox.Reading(PromptBox.State.EMPTY, 0), PromptBox.classify(EMPTY_GREY_CARET),
"the caret's own colour sits before the marker and must not stop the marker matching");
}
@Test
void anEmptyBoxWithNoStylingAtAllReadsAsEmpty() {
assertEquals(new PromptBox.Reading(PromptBox.State.EMPTY, 0), PromptBox.classify(EMPTY_PLAIN));
}
@Test
void typedTextWithNoStylingIsADraftAndCountsItsCharacters() {
PromptBox.Reading reading = PromptBox.classify("❯\u00a0deploy the thing");
assertEquals(PromptBox.State.DRAFT, reading.state());
assertEquals("deploythething".length(), reading.characters());
}
@Test
void typedTextAfterAFaintHintCountsOnlyTheTextOutsideTheFaintSpan() {
PromptBox.Reading reading =
PromptBox.classify("❯\u00a0\u001b[0m\u001b[2mhint\u001b[0m and typed");
assertEquals(PromptBox.State.DRAFT, reading.state());
assertEquals("andtyped".length(), reading.characters(),
"the faint hint is excluded; only \"and typed\" was drawn plain");
}
// --- the gate ------------------------------------------------------------
@Test
void anEmptyBoxClearsTheGateAndReadsTheVisibleRegionWithStylingKept() {
FakeHerdr herdr = new FakeHerdr().detectionText(EMPTY);
assertTrue(new PromptBox(new AgentControl(herdr)).clearToSubmit("term_a"));
@SuppressWarnings("unchecked")
var params = (java.util.Map<String, Object>) herdr.lastCall("agent.read").params();
assertEquals("visible", params.get("source"),
"the input box is drawn in the visible region, not in transcript scrollback");
assertEquals(false, params.get("strip_ansi"),
"styling must survive the read, or a faint placeholder hint reads as plain typed text");
}
@Test
void aDraftedBoxHoldsTheGate() {
assertFalse(new PromptBox(new AgentControl(new FakeHerdr().detectionText(DRAFTED))).clearToSubmit("term_a"));
}
@Test
void anUnreadablePaneHoldsTheGate() {
assertFalse(new PromptBox(new AgentControl(new FakeHerdr().detectionText("garbled"))).clearToSubmit("term_a"));
}
@Test
void aFailedReadHoldsTheGate() {
assertFalse(new PromptBox(new AgentControl(new FakeHerdr().healthy(false))).clearToSubmit("term_a"),
"a pane this cannot read must never be pasted into");
}
@Test
void theGateClearsAgainOnceTheBoxEmpties() {
FakeHerdr herdr = new FakeHerdr().detectionText(DRAFTED);
PromptBox box = new PromptBox(new AgentControl(herdr));
assertFalse(box.clearToSubmit("term_a"));
herdr.detectionText(EMPTY);
assertTrue(box.clearToSubmit("term_a"));
}
// --- the gate's on/off switch ---------------------------------------------
@Test
void aDisabledGateClearsADraftedBoxWithoutReadingThePane() {
FakeHerdr herdr = new FakeHerdr().detectionText(DRAFTED);
assertTrue(new PromptBox(new AgentControl(herdr), false).clearToSubmit("term_a"),
"a disabled gate must clear every target, draft or not");
assertFalse(herdr.called("agent.read"), "a disabled gate must never read the pane");
}
@Test
void anEnabledGateBehavesExactlyLikeTheSingleArgConstructor() {
FakeHerdr herdr = new FakeHerdr().detectionText(DRAFTED);
assertFalse(new PromptBox(new AgentControl(herdr), true).clearToSubmit("term_a"),
"gateEnabled=true must hold a draft exactly like the default constructor");
assertTrue(herdr.called("agent.read"), "an enabled gate still reads the pane");
}
}
@@ -89,6 +89,11 @@ class BackendOutageFlowTest {
return MAPPER.createObjectNode().set("agent", MAPPER.createObjectNode()
.put("terminal_id", "term_primary").put("agent_status", "idle"));
}
if ("agent.read".equals(method)) {
// The lead-nudge paths read the input box before pasting into it.
return MAPPER.createObjectNode().set("read",
MAPPER.createObjectNode().put("text", FakeHerdr.IDLE_PROMPT_CARET));
}
if ("agent.prompt".equals(method)) {
prompts.add(params);
sendLatch.countDown();
@@ -0,0 +1,200 @@
package dev.ltms.fleet.inject;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.msg.TestTurnTokens;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Which route a message takes: offered to a pane that collects its own mail, or typed into the
* pane's terminal. Driven by feeding {@code onStatus}, so no real polling is involved.
*/
class InjectorModServedDeliveryTest {
/** A pane that collects its own mail. */
private static final String MOD = "term_mod";
/** A pane that does not, used as the control for every "nothing was typed" assertion. */
private static final String PTY = "term_pty";
private final FakeHerdr herdr = new FakeHerdr();
private final AtomicLong clock = new AtomicLong(1_000_000);
private final Injector injector = new Injector(new AgentControl(herdr), TurnListener.NOOP,
_ -> true, _ -> {
}, clock::get);
/** The messages typed into a pane, in order. A collected message must never appear here. */
@SuppressWarnings("unchecked")
private List<String> typed() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void aPaneThatCollectsItsOwnMailIsNeverTypedInto() {
injector.collectInbox(MOD); // the pane says it collects its own mail
CompletableFuture<Void> delivered =
injector.enqueue(MOD, "do the task", TestTurnTokens.inert(MOD)).completion();
injector.onStatus(MOD, AgentStatus.IDLE); // offers it for collection
assertFalse(delivered.isDone(), "an offered message has not reached the pane yet");
assertEquals(List.of("do the task"), injector.collectInbox(MOD), "the pane collects it");
injector.onStatus(MOD, AgentStatus.IDLE); // the next sample records the delivery
assertTrue(delivered.isDone(), "a collected message is a delivered message");
assertEquals(List.of(), typed(), "nothing was typed into a pane that collects its own mail");
// The control: without it, an injector that typed nothing anywhere would pass the line
// above. Same injector, same herdr, a pane that never collected its mail.
injector.enqueue(PTY, "type this", TestTurnTokens.inert(PTY));
injector.onStatus(PTY, AgentStatus.IDLE);
assertEquals(List.of("type this"), typed(), "control: an ordinary pane is still typed into");
}
@Test
void aPaneThatStopsCollectingHasItsMailTypedInstead() {
injector.collectInbox(MOD);
CompletableFuture<Void> delivered =
injector.enqueue(MOD, "do the task", TestTurnTokens.inert(MOD)).completion();
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(List.of(), typed(), "control: while it is still collecting, nothing is typed");
assertFalse(delivered.isDone(), "control: and nothing is reported delivered either");
// The mod stopped calling fleet_inbox, so the pane leaves the window.
clock.addAndGet(PaneInbox.MOD_SERVED_WINDOW_MILLIS + 1);
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(List.of("do the task"), typed(), "the message falls back to the terminal route");
assertTrue(delivered.isDone(), "and is reported delivered once it is typed");
assertEquals(List.of(), injector.collectInbox(MOD),
"a message that was typed must not also still be collectable");
}
@Test
void aMessageOfferedForCollectionIsStillReportedAsNotYetDelivered() {
injector.collectInbox(MOD);
injector.enqueue(MOD, "do the task", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE);
assertFalse(injector.queuedWaitMillis(MOD) == null,
"an offered-but-uncollected message is still waiting, not delivered");
injector.collectInbox(MOD);
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(null, injector.queuedWaitMillis(MOD),
"once collected it is off the queue, the same as a typed message");
}
@Test
void aCollectedMessageIsNotFollowedByAnEnterNudge() {
injector.collectInbox(MOD);
injector.enqueue(MOD, "do the task", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE);
injector.collectInbox(MOD);
injector.onStatus(MOD, AgentStatus.IDLE); // records the delivery, arms the pickup latch
injector.onStatus(MOD, AgentStatus.IDLE); // still idle: the typed route nudges Enter here
assertFalse(herdr.called("agent.send_keys"),
"a pane that collects its own mail submits it itself; an Enter there would submit "
+ "whatever its operator is typing");
// The control: the nudge really does fire on the typed route, so the absence above is
// this route's behaviour and not a harness that never nudges at all.
injector.enqueue(PTY, "type this", TestTurnTokens.inert(PTY));
injector.onStatus(PTY, AgentStatus.IDLE);
injector.onStatus(PTY, AgentStatus.IDLE);
assertTrue(herdr.called("agent.send_keys"), "control: a typed message is nudged");
}
@Test
void aCancelledMessageStopsBeingCollectable() {
injector.collectInbox(MOD);
Injector.Delivery delivery = injector.enqueue(MOD, "retracted", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE); // offered for collection
assertEquals(Injector.Cancellation.CANCELLED, injector.cancel(delivery),
"an offered message has not reached the pane, so it can still be cancelled");
assertEquals(List.of(), injector.collectInbox(MOD),
"a cancelled message the caller was told never arrived must not arrive later");
// The control: an uncancelled message on the same route really is collectable, so the
// empty list above is the cancel working and not the offer never being made.
injector.enqueue(MOD, "kept", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(List.of("kept"), injector.collectInbox(MOD), "control: an offer is collectable");
}
@Test
void aMessageThePaneAlreadyCollectedCannotBeCancelled() {
injector.collectInbox(MOD);
// The control first: an offer the pane has not taken really is cancellable, so the
// different answer below is the collection and not a cancel that gave up on this route.
Injector.Delivery untaken = injector.enqueue(MOD, "retracted", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(Injector.Cancellation.CANCELLED, injector.cancel(untaken),
"control: an uncollected offer is still cancellable");
Injector.Delivery taken = injector.enqueue(MOD, "do the task", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(List.of("do the task"), injector.collectInbox(MOD), "the pane takes the offer");
// No poll has run since the pane took it, so the entry is still at the head and still
// QUEUED: the state alone cannot tell this case from an uncollected offer.
assertEquals(Injector.Cancellation.DELIVERED, injector.cancel(taken),
"the pane holds this text and will act on it, so nothing can be cancelled");
injector.onStatus(MOD, AgentStatus.IDLE);
assertTrue(taken.completion().isDone(), "the next poll records the delivery");
assertEquals(List.of(), typed(), "and nothing was typed into the pane");
}
@Test
void cancellingALaterMessageLeavesACollectedOneDeliveredOnce() {
injector.collectInbox(MOD);
Injector.Delivery first = injector.enqueue(MOD, "first", TestTurnTokens.inert(MOD));
Injector.Delivery second = injector.enqueue(MOD, "second", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE); // offers the head
assertEquals(List.of("first"), injector.collectInbox(MOD), "the pane takes the head");
assertEquals(Injector.Cancellation.CANCELLED, injector.cancel(second),
"a message behind the collected one was never offered, so it cancels");
injector.onStatus(MOD, AgentStatus.IDLE);
assertTrue(first.completion().isDone(),
"cancelling a later message must not lose the record that the head was taken");
assertEquals(List.of(), injector.collectInbox(MOD),
"and the head must not be offered a second time");
// The control: the same injector still hands a later message over, so the empty
// collection above is this one not being re-offered rather than the route going quiet.
injector.onStatus(MOD, AgentStatus.WORKING); // the pane picks the collected message up
injector.onStatus(MOD, AgentStatus.IDLE); // and that turn ends
injector.enqueue(MOD, "third", TestTurnTokens.inert(MOD));
injector.onStatus(MOD, AgentStatus.IDLE);
assertEquals(List.of("third"), injector.collectInbox(MOD), "control: a later message is offered");
assertEquals(List.of(), typed(), "nothing took the terminal route");
}
@Test
void aPaneThatNeverCollectedIsTypedIntoFromTheStart() {
injector.enqueue(PTY, "do the task", TestTurnTokens.inert(PTY));
injector.onStatus(PTY, AgentStatus.IDLE);
assertEquals(List.of("do the task"), typed(), "no poll, no offer: the terminal route applies");
assertFalse(injector.isModServed(PTY));
}
}
@@ -59,6 +59,17 @@ class InjectorTest {
assertEquals(List.of("hello"), sent());
}
@Test
void deliveringToAMemberReadsNoPane() {
// No human types into a spawned member's pane, so its delivery path must not pay for a
// prompt-box read the way a lead's nudge paths do.
injector.enqueue(T, "task", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE);
assertEquals(List.of("task"), sent());
assertFalse(herdr.called("agent.read"), "a member delivery must not read its pane");
}
@Test
void holdsDeliveryUntilTheWorkerIsAvailable() {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
@@ -548,20 +559,21 @@ class InjectorTest {
void readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants() {
// fleetd #501, defect 2: the old line computed "({}s)" as READINESS_GRACE_POLLS *
// POLL_INTERVAL_MILLIS / 1000 — arithmetic on two constants, never a measurement, and wrong
// in the direction that says everything ran on schedule. This stub clock returns two FIXED
// values (1_000ms at the first non-ready sample, 318_412ms at the poll that trips the grace)
// whose difference — 317_412ms — does NOT equal 240 * POLL_INTERVAL_MILLIS (=60_000ms).
// Asserting on that literal, non-derived number is what makes this test able to fail if the
// production code goes back to printing the constant-arithmetic value instead of the
// injected clock's measurement.
long[] readings = {1_000L, 318_412L};
// in the direction that says everything ran on schedule. This stub clock returns three FIXED
// values (an unused enqueue-time stamp, 1_000ms at the first non-ready sample, 318_412ms at
// the poll that trips the grace) whose last two differ — 317_412ms — which does NOT equal
// 240 * POLL_INTERVAL_MILLIS (=60_000ms). Asserting on that literal, non-derived number is
// what makes this test able to fail if the production code goes back to printing the
// constant-arithmetic value instead of the injected clock's measurement.
long[] readings = {0L, 1_000L, 318_412L};
AtomicInteger call = new AtomicInteger(0);
LongSupplier stubClock = () -> {
int i = call.getAndIncrement();
if (i >= readings.length) {
throw new AssertionError("nowMillis read more times than this fixture expects (" + i
+ "); the readiness-not-ready branch should read the clock exactly twice — "
+ "once to stamp the first non-ready sample, once at grace expiry");
+ "); the readiness-not-ready branch should read the clock exactly three times —"
+ " once to stamp the queued message's own enqueue time, once to stamp the "
+ "first non-ready sample, once at grace expiry");
}
return readings[i];
};
@@ -603,6 +615,46 @@ class InjectorTest {
assertEquals(List.of("task"), sent(), "a worker that connects within the grace is delivered to");
}
// ~20 minutes of a continuously busy target at the 250ms prod poll interval; enough to trip
// the queue-wait grace.
private static final int QUEUE_WAIT_SAMPLES = 4800;
@Test
void failsAQueuedMessageWhoseTargetNeverFreesUp() {
// A target that stays WORKING the whole time never reaches the branch that looks at the
// queue at all, so nothing else bounds this. The message must fail rather than wait
// forever, and the caller's future unblocks through the same turn-failure path a readiness
// timeout uses — but presence must not be touched, since this target is merely busy, not
// gone.
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> true, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task", TestTurnTokens.inert(T)).completion();
for (int i = 0; i < QUEUE_WAIT_SAMPLES; i++) inj.onStatus(T, AgentStatus.WORKING);
assertEquals(List.of(), sent(), "a target that never frees up is never delivered to");
assertTrue(f.isCompletedExceptionally(), "the caller's future fails instead of hanging forever");
assertEquals(List.of(T), cap.failed, "the awaiting send resolves through the turn-failure path");
assertEquals(List.of(), cap.completed, "a never-freed target is a failure, not a completion");
assertEquals(List.of(), forgotten, "the target is busy, not gone — presence must not be cleared");
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void aTargetThatFreesUpBeforeTheQueueWaitGraceIsDeliveredNormally() {
// Positive control: a target that is merely busy for a while, then frees up before the
// grace elapses, is still delivered normally — the long-task case this grace must not break.
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP);
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.WORKING); // busy, well under the grace
assertEquals(List.of(), sent());
inj.onStatus(T, AgentStatus.IDLE); // frees up
assertEquals(List.of("task"), sent(), "a target that frees up within the grace is delivered to");
}
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
@@ -0,0 +1,93 @@
package dev.ltms.fleet.inject;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** What a pane that collects its own mail may and may not see. */
class PaneInboxTest {
private static final String A = "term_a";
private static final String B = "term_b";
private final AtomicLong clock = new AtomicLong(1_000_000);
private final PaneInbox inbox = new PaneInbox(clock::get);
@Test
void aPaneCollectsItsOwnMailAndLeavesTheNextPanesWhereItIs() {
inbox.offer(A, "for a");
inbox.offer(B, "for b");
assertEquals(List.of("for a"), inbox.drain(A), "a pane sees its own message");
// The control for the assertion below: B's message really is there to be missed, so a
// drain that returned everything would have shown it above.
assertEquals(List.of("for b"), inbox.drain(B), "the other pane's message stayed put");
}
@Test
void aCollectedMessageIsNotHandedOverASecondTime() {
inbox.offer(A, "deliver once");
assertEquals(List.of("deliver once"), inbox.drain(A), "control: the first call hands it over");
assertEquals(List.of(), inbox.drain(A), "a collected message is gone from the inbox");
}
@Test
void messagesComeBackInTheOrderTheyWereOffered() {
inbox.offer(A, "first");
inbox.offer(A, "second");
assertEquals(List.of("first", "second"), inbox.drain(A));
}
@Test
void aPaneIsModServedOnlyWhileItKeepsCollecting() {
assertFalse(inbox.isModServed(A), "a pane that has never collected is not mod-served");
inbox.drain(A);
assertTrue(inbox.isModServed(A), "control: collecting is what makes a pane mod-served");
clock.addAndGet(PaneInbox.MOD_SERVED_WINDOW_MILLIS);
assertTrue(inbox.isModServed(A), "control: the window edge still counts as collecting");
clock.addAndGet(1);
assertFalse(inbox.isModServed(A), "a pane that stopped collecting leaves the window");
}
@Test
void withdrawingTakesBackOnlyWhatThePaneHasNotCollected() {
PaneInbox.Entry collected = inbox.offer(A, "already taken");
inbox.drain(A);
PaneInbox.Entry pending = inbox.offer(A, "not taken yet");
inbox.withdrawAll(A);
assertTrue(collected.taken(), "withdrawing must not un-deliver a collected message");
assertFalse(pending.taken(), "control: the uncollected entry was never handed over");
assertEquals(List.of(), inbox.drain(A), "a withdrawn message is no longer collectable");
}
@Test
void forgettingAPaneDropsBothItsMailAndItsPollRecord() {
inbox.drain(A);
inbox.offer(A, "for a");
assertTrue(inbox.isModServed(A), "control: the pane is mod-served and holding mail");
inbox.forget(A);
assertFalse(inbox.isModServed(A), "a gone pane must not look mod-served to the next one");
assertEquals(List.of(), inbox.drain(A), "a gone pane's mail does not outlive it");
}
@Test
void aMissingTerminalCollectsNothingAndIsNeverModServed() {
assertEquals(List.of(), inbox.drain(null));
assertEquals(List.of(), inbox.drain(" "));
assertFalse(inbox.isModServed(null));
}
}
@@ -114,13 +114,13 @@ class LeadLauncherTest {
"name is lead-<name>-<nonce>-<seq>: " + startedName(herdr));
}
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
/** The tab is labelled with the fixed lead tab label so the scanner finds the lead on the next resolve. */
@Test
void labelsTheTabWithTheConfiguredTabValue() {
void labelsTheTabWithTheFixedLeadTabLabel() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
assertEquals("lead: opus",
assertEquals("lead",
((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
}
@@ -148,6 +148,25 @@ class LeadLauncherTest {
assertFalse(herdr.called("tab.close"), "a labelled tab WITH a live agent must never be closed");
}
/**
* The tab a live lead actually sits in still carries its deprecated legacy {@code tab:} label,
* not the fixed {@code lead} tab label a freshly auto-launched instance would get. Counting must
* still recognise it as the live lead via {@link FleetConfig.Leader#acceptedLabels()}, or a
* daemon restart would read it as missing and launch a second orchestrator next to the first.
*/
@Test
void aLiveLeadInALegacyLabelledTabIsCountedSoNothingIsLaunched() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wL", "fleet")
.withTab("wL", "wL:t1", "lead: opus")
.withAgent("lead-opus", "term_lead", "wL:p1", "wL:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the legacy-labelled live lead must be counted — nothing may be launched");
assertFalse(herdr.called("agent.start"),
"a tab label fixed to a constant must not blind the count to a legacy-labelled lead");
}
/**
* The reason liveness is not "does the label exist". A tab left labelled by a session that has
* since died must not block the relaunch, or one crash disables auto-launch permanently.
@@ -263,7 +282,7 @@ class LeadLauncherTest {
assertFalse(herdr.called("tab.close"), "a tab running an agent again must never be closed");
assertFalse(herdr.called("agent.start"), "the lead is live again — nothing to relaunch");
assertEquals("wL:t1", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("tab_id"));
assertEquals("lead: opus", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"),
assertEquals("lead", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"),
"the pending-close flag must be cleared once the tab is confirmed live again");
}
@@ -285,7 +304,7 @@ class LeadLauncherTest {
@Test
void aHandOpenedLeadWithTheConfiguredTabLabelCountsAsLive() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wX", "main")
.withWorkspace("wX", "fleet")
.withTab("wX", "wX:t1", "lead: opus")
.withAgent("hand-opened", "term_hand", "wX:p1", "wX:t1");
@@ -597,7 +616,7 @@ class LeadLauncherTest {
"terminalId() must be herdr's own generated id: " + started.terminalId());
}
/** The new tab is labelled with the lead's configured {@code tab:}, and AFTER the start. */
/** The new tab is labelled with the fixed lead tab label, and AFTER the start. */
@Test
void relaunchLabelsTheNewTabAfterStarting() {
FakeHerdr herdr = new FakeHerdr();
@@ -606,7 +625,7 @@ class LeadLauncherTest {
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started);
assertEquals("lead: opus", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
assertEquals("lead", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
int startIndex = indexOfLastCall(herdr, "agent.start");
int renameIndex = indexOfLastCall(herdr, "tab.rename");
@@ -2,10 +2,12 @@ package dev.ltms.fleet.lead;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.DisplayName;
@@ -20,6 +22,7 @@ import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -56,16 +59,18 @@ import static org.junit.jupiter.api.Assertions.*;
* assertable without any thread coordination.
*
* <p><strong>A fixed, non-advancing clock and a full continuation run do not mix.</strong>
* {@link LeadRollover#waitUntilPaneGone} and {@code waitUntilRecognisedAsLead} are bounded loops
* of the shape {@code while (nowMillis.getAsLong() < deadline) { ... pollSleeper.run(); }}, and the
* {@link LeadRollover#waitUntilPaneGone}, {@code waitUntilPaneReady}, {@code
* waitUntilRecognisedAsLead} and {@code waitUntilTurnStarted} are all bounded loops of the shape
* {@code while (nowMillis.getAsLong() < deadline) { ... pollSleeper.run(); }}, and the
* package-private test constructor's {@code pollSleeper} is a no-op — so a loop that does not exit
* on its very first check, against a clock that never advances, never exits at all. Every test
* below that lets the continuation reach one of those two waits therefore either (a) arranges for
* the wait to resolve on its first check — a pane already marked gone via {@link
* FakeHerdr#paneGoneAfterClose}, or a terminal already present in {@code liveLeadTerminals} at a
* settled status — so a fixed clock never needs to advance, or (b) uses a clock that advances on
* every read ({@code () -> clock.addAndGet(500)}), so a bounded timeout genuinely elapses instead
* of spinning forever.
* below that lets the continuation reach one of those waits therefore either (a) arranges for the
* wait to resolve on its first check — a pane already marked gone via {@link
* FakeHerdr#paneGoneAfterClose}, a terminal already present in {@code liveLeadTerminals} at a
* settled status, or {@link FakeHerdr#turnStartsOnPrompt()} reporting {@code "working"} right after
* the send lands (see {@link #herdrReadyForAFullRoll()}) — so a fixed clock never needs to
* advance, or (b) uses a clock that advances on every read ({@code () -> clock.addAndGet(500)}),
* so a bounded timeout genuinely elapses instead of spinning forever.
*/
class LeadRolloverTest {
@@ -124,11 +129,14 @@ class LeadRolloverTest {
/**
* A {@code herdr} already set up so the deferred continuation can run the WHOLE sequence to
* {@link LeadRollover.RollState#ROLLED} on its first pass through every bounded wait: the old
* pane ({@link #OLD_PANE}) reports gone the instant it is closed, and the next {@code
* agent.start} is pinned to {@link #NEW_TERMINAL}/{@link #NEW_PANE}.
* pane ({@link #OLD_PANE}) reports gone the instant it is closed, the next {@code agent.start}
* is pinned to {@link #NEW_TERMINAL}/{@link #NEW_PANE}, and the fresh pane's status flips to
* {@code "working"} right after {@code bootstrapText} is sent — the fixture's stand-in for the
* successor starting a turn on it, which {@code waitUntilTurnStarted} requires.
*/
private static FakeHerdr herdrReadyForAFullRoll() {
return new FakeHerdr().paneGoneAfterClose(OLD_PANE).pinNextStarts(1, NEW_TERMINAL, NEW_PANE);
return new FakeHerdr().paneGoneAfterClose(OLD_PANE).pinNextStarts(1, NEW_TERMINAL, NEW_PANE)
.turnStartsOnPrompt();
}
private static LeadRollover newRollover(HerdrClient herdr, FleetConfig.LeadRollover config,
@@ -144,11 +152,19 @@ class LeadRolloverTest {
private static LeadRollover newRollover(HerdrClient herdr, FleetConfig.LeadRollover config,
LongSupplier nowMillis, Function<String, String> leadWorkspace,
java.util.function.Supplier<Map<String, String>> liveLeadTerminals) {
return newRollover(herdr, config, nowMillis, leadWorkspace, liveLeadTerminals, _ -> true);
}
/** As the five-arg {@link #newRollover}, but with an explicit {@code mcpPresent}. */
private static LeadRollover newRollover(HerdrClient herdr, FleetConfig.LeadRollover config,
LongSupplier nowMillis, Function<String, String> leadWorkspace,
java.util.function.Supplier<Map<String, String>> liveLeadTerminals,
java.util.function.Predicate<String> mcpPresent) {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = new LeadLauncher(agents, spaces, emptyFleetConfig());
return new LeadRollover(agents, spaces, launcher, () -> config, leadWorkspace,
_ -> LEAD_NAME, liveLeadTerminals, nowMillis, () -> { }, Runnable::run);
_ -> LEAD_NAME, liveLeadTerminals, mcpPresent, nowMillis, () -> { }, Runnable::run);
}
/**
@@ -161,11 +177,19 @@ class LeadRolloverTest {
private static LeadRollover newRolloverForAFullRoll(FakeHerdr herdr, FleetConfig.LeadRollover config,
LongSupplier nowMillis,
Map<String, String> liveLeadTerminals) {
return newRolloverForAFullRoll(herdr, config, nowMillis, liveLeadTerminals, _ -> true);
}
/** As the four-arg {@link #newRolloverForAFullRoll}, but with an explicit {@code mcpPresent}. */
private static LeadRollover newRolloverForAFullRoll(FakeHerdr herdr, FleetConfig.LeadRollover config,
LongSupplier nowMillis,
Map<String, String> liveLeadTerminals,
java.util.function.Predicate<String> mcpPresent) {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = fakeLauncher(herdr, fleetConfigWithRelaunchableLead());
return new LeadRollover(agents, spaces, launcher, () -> config, _ -> null,
_ -> LEAD_NAME, () -> liveLeadTerminals, nowMillis, () -> { }, Runnable::run);
_ -> LEAD_NAME, () -> liveLeadTerminals, mcpPresent, nowMillis, () -> { }, Runnable::run);
}
private Path writeHandover(String content) throws IOException {
@@ -463,7 +487,7 @@ class LeadRolloverTest {
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = fakeLauncher(herdr, emptyFleetConfig());
LeadRollover rollover = new LeadRollover(agents, spaces, launcher, () -> null, _ -> null,
_ -> null, Map::of, () -> 1_000L, () -> { }, Runnable::run);
_ -> null, Map::of, _ -> true, () -> 1_000L, () -> { }, Runnable::run);
assertThrows(IllegalStateException.class, () -> rollover.open(LEAD, "context is full"));
}
@@ -586,6 +610,83 @@ class LeadRolloverTest {
rollover.status(pending.token()).state());
}
@Test
@DisplayName("a fresh terminal at a turn boundary but absent from MemberPresence never becomes ready — bootstrapText is never sent")
void freshTerminalAtTurnBoundaryButNotMcpPresentNeverBecomesReadyNeverSendsBootstrapText() throws IOException {
FakeHerdr herdr = herdrReadyForAFullRoll();
Path handover = writeHandover("handover contents");
FleetConfig.LeadRollover config =
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*relaunchReadySeconds*/, "boot text");
AtomicLong clock = new AtomicLong(1_000);
Map<String, String> liveLeadTerminals = new LinkedHashMap<>();
liveLeadTerminals.put(NEW_TERMINAL, LEAD_NAME); // recognised already, if the wait is ever reached
// herdr's own agent_status reports "idle" (a real turn boundary) the instant the fresh pane
// comes up — mcpPresent is the only signal left that can withhold readiness here.
LeadRollover rollover = newRolloverForAFullRoll(herdr, config, () -> clock.addAndGet(500),
liveLeadTerminals, _ -> false);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
+ "deep inside the deferred continuation");
assertEquals(0, promptCallCount(herdr), "a turn boundary alone is not readiness — "
+ "bootstrapText must never be typed into a pane whose Claude has not connected "
+ "the bridge MCP client, even while herdr itself reports it idle");
assertEquals(LeadRollover.RollState.RELAUNCH_NEVER_READY,
rollover.status(pending.token()).state());
}
@Test
@DisplayName("a fresh terminal at a turn boundary AND present in MemberPresence becomes ready — the roll proceeds to a full ROLLED")
void freshTerminalAtTurnBoundaryAndMcpPresentBecomesReadyAndRolls() throws IOException {
FakeHerdr herdr = herdrReadyForAFullRoll();
Path handover = writeHandover("handover contents");
FleetConfig.LeadRollover config = cfg(handover.toString());
AtomicLong clock = new AtomicLong(1_000);
Map<String, String> liveLeadTerminals = new LinkedHashMap<>();
liveLeadTerminals.put(NEW_TERMINAL, LEAD_NAME);
// Present for the fresh terminal specifically, not a blanket "always true" — proves the
// predicate's return value, not merely its presence as an argument, is what the readiness
// wait acts on.
LeadRollover rollover = newRolloverForAFullRoll(herdr, config, fixedClock(clock),
liveLeadTerminals, NEW_TERMINAL::equals);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
assertEquals(1, promptCallCount(herdr), "a turn boundary together with MemberPresence must "
+ "let bootstrapText through");
assertEquals(LeadRollover.RollState.ROLLED, rollover.status(pending.token()).state());
}
@Test
@DisplayName("herdr accepts bootstrapText but the fresh pane never starts a turn on it — the roll reports BOOTSTRAP_NOT_CONFIRMED and never re-sends")
void bootstrapTextSentButNeverConfirmedEndsInBootstrapNotConfirmedWithNoResend() throws IOException {
// Ready for relaunch and recognition (idle status, already-live terminal) but WITHOUT
// turnStartsOnPrompt() — this fixture's status never reaches "working", modelling herdr
// accepting the send while the successor never actually starts a turn on it.
FakeHerdr herdr = new FakeHerdr().paneGoneAfterClose(OLD_PANE).pinNextStarts(1, NEW_TERMINAL, NEW_PANE);
Path handover = writeHandover("handover contents");
FleetConfig.LeadRollover config =
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*relaunchReadySeconds*/, "boot text");
AtomicLong clock = new AtomicLong(1_000);
Map<String, String> liveLeadTerminals = new LinkedHashMap<>();
liveLeadTerminals.put(NEW_TERMINAL, LEAD_NAME);
LeadRollover rollover = newRolloverForAFullRoll(herdr, config, () -> clock.addAndGet(500),
liveLeadTerminals);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
assertEquals(1, promptCallCount(herdr), "bootstrapText is sent exactly once; a send herdr "
+ "accepts but never confirms must never be retried");
assertEquals(LeadRollover.RollState.BOOTSTRAP_NOT_CONFIRMED,
rollover.status(pending.token()).state());
}
@Test
@DisplayName("a full successful roll ends the old pane, relaunches, and sends bootstrapText ONLY to the fresh terminal, in order, and consumes the token")
void successfulRollEndsTheOldPaneRelaunchesAndSendsBootstrapTextToTheFreshTerminal() throws IOException {
@@ -809,7 +910,7 @@ class LeadRolloverTest {
LeadLauncher launcher = fakeLauncher(herdr, fleetConfigWithRelaunchableLead());
Function<String, String> leadWorkspace = terminal -> LEAD.equals(terminal) ? workspace.toString() : null;
LeadRollover rollover = new LeadRollover(agents, spaces, launcher, () -> config, leadWorkspace,
_ -> LEAD_NAME, () -> liveLeadTerminals, fixedClock(clock), () -> { }, Runnable::run);
_ -> LEAD_NAME, () -> liveLeadTerminals, _ -> true, fixedClock(clock), () -> { }, Runnable::run);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
@@ -1035,7 +1136,10 @@ class LeadRolloverTest {
@Test
@DisplayName("[STATUS 5] the bounded outcome history never grows past its cap")
void statusHistoryDoesNotGrowPastItsCap() throws IOException {
FakeHerdr herdr = new FakeHerdr().paneGoneAfterClose(OLD_PANE);
FakeHerdr herdr = new FakeHerdr().paneGoneAfterClose(OLD_PANE).turnStartsOnPrompt();
// Each roll must see one live turn after its send, so the confirmation wait returns on its
// first poll instead of burning its whole bound 250 times; a one-shot keeps the pane
// settled for the next roll's own readiness and turn-boundary waits.
Path handover = writeHandover("handover contents"); // one file, reused by every roll below —
// checkHandover only compares its mtime against each open()'s OWN requestedAtMillis (the
// fake clock, in the low thousands), and the file's real (wall-clock) mtime is always far
@@ -1048,7 +1152,7 @@ class LeadRolloverTest {
herdr.pinNextStarts(rolls, NEW_TERMINAL, NEW_PANE);
Map<String, String> liveLeadTerminals = Map.of(NEW_TERMINAL, LEAD_NAME);
LeadRollover rollover = newRolloverForAFullRoll(herdr, cfg(handover.toString()),
() -> clock.addAndGet(1), liveLeadTerminals);
() -> clock.addAndGet(LeadRollover.POLL_INTERVAL_MS), liveLeadTerminals);
String[] tokens = new String[rolls];
for (int i = 0; i < rolls; i++) {
@@ -1133,8 +1237,8 @@ class LeadRolloverTest {
// keying its own single-flight claim on the terminal itself, exactly like a terminal
// the live roster does not recognise.
return new LeadRollover(agents, spaces, launcher, () -> config, _ -> null,
t -> LEAD.equals(t) ? LEAD_NAME : null, () -> liveLeadTerminals, nowMillis, () -> { },
runner);
t -> LEAD.equals(t) ? LEAD_NAME : null, () -> liveLeadTerminals, _ -> true, nowMillis,
() -> { }, runner);
}
@Test
@@ -1323,6 +1427,67 @@ class LeadRolloverTest {
+ "so an operator reading status() has something to act on: " + status.detail());
}
@Test
@DisplayName("[BOOTSTRAP 1] one agent_not_ready bootstrap refusal is retried and sends the "
+ "handover instruction exactly once")
void transientAgentNotReadyRetriesBootstrapAndRolls() throws IOException {
FakeHerdr fake = herdrReadyForAFullRoll();
AtomicInteger refusedSends = new AtomicInteger();
HerdrClient transientRefusal = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if ("agent.prompt".equals(method) && refusedSends.getAndIncrement() == 0) {
throw new HerdrException("herdr error [agent_not_ready]: agent.prompt failed",
"agent_not_ready", null);
}
return fake.call(method, params);
}
@Override
public void close() {
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
AgentControl agents = new AgentControl(transientRefusal);
WorkspaceControl spaces = new WorkspaceControl(transientRefusal);
LeadRollover rollover = new LeadRollover(agents, spaces,
new LeadLauncher(agents, spaces, fleetConfigWithRelaunchableLead()),
() -> cfg(handover.toString()), _ -> null, _ -> LEAD_NAME,
() -> Map.of(NEW_TERMINAL, LEAD_NAME), _ -> true, fixedClock(clock), () -> { },
Runnable::run);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
assertTrue(rollover.confirm(LEAD, pending.token(), true).accepted());
assertEquals(2, refusedSends.get(), "one agent_not_ready refusal must be followed by one retry");
assertEquals(1, promptCallCount(fake), "only the successful retry reaches herdr delivery");
assertEquals(LeadRollover.RollState.ROLLED, rollover.status(pending.token()).state());
}
@Test
@DisplayName("[BOOTSTRAP 2] persistent agent_not_ready records BOOTSTRAP_NEVER_SENT and "
+ "releases the single-flight claim")
void persistentAgentNotReadyRecordsBootstrapNeverSentAndReleasesClaim() throws IOException {
FakeHerdr fake = herdrReadyForAFullRoll();
fake.agentSendFailsWith("agent_not_ready");
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRolloverForAFullRoll(fake, cfg(handover.toString()),
() -> clock.addAndGet(1_000), Map.of(NEW_TERMINAL, LEAD_NAME));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
assertTrue(rollover.confirm(LEAD, pending.token(), true).accepted());
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.BOOTSTRAP_NEVER_SENT, status.state());
assertNotEquals(LeadRollover.RollState.ROLLED, status.state());
assertNotEquals(LeadRollover.RollState.FAILED, status.state());
LeadRollover.PendingRollover retry = rollover.open(LEAD, "retry after bootstrap timeout");
assertTrue(rollover.confirm(LEAD, retry.token(), true).accepted(),
"the terminal outcome must release the single-flight claim");
}
// ---- fleetd #726 unit 3: confirm() single-flights one roll at a time per lead ---------------
@Test
@@ -1372,7 +1537,8 @@ class LeadRolloverTest {
@DisplayName("[SINGLE-FLIGHT 2] two DIFFERENT lead terminals confirm without either being "
+ "refused, and the continuation runs once per terminal")
void differentLeadTerminalsConfirmIndependently() throws IOException {
FakeHerdr herdr = new FakeHerdr().paneGoneAfterClose(OLD_PANE).pinNextStarts(2, NEW_TERMINAL, NEW_PANE);
FakeHerdr herdr = new FakeHerdr().paneGoneAfterClose(OLD_PANE)
.pinNextStarts(2, NEW_TERMINAL, NEW_PANE).turnStartsOnPrompt();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
Map<String, String> liveLeadTerminals = Map.of(NEW_TERMINAL, LEAD_NAME);
@@ -1483,7 +1649,7 @@ class LeadRolloverTest {
// executor's RejectedExecutionException — the continuation never runs, so runRollover's
// own finally never gets a chance to release the claim either.
LeadRollover rollover = new LeadRollover(agents, spaces, launcher, () -> cfg(handover.toString()),
_ -> null, _ -> LEAD_NAME, Map::of, fixedClock(clock), () -> { },
_ -> null, _ -> LEAD_NAME, Map::of, _ -> true, fixedClock(clock), () -> { },
r -> { throw new RuntimeException("simulated continuationRunner rejection"); });
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
@@ -1529,8 +1695,8 @@ class LeadRolloverTest {
// the SAME lead, just a different terminal id.
Function<String, String> leadNameForTerminal = _ -> LEAD_NAME;
LeadRollover rollover = new LeadRollover(agents, spaces, launcher, () -> cfg(handover.toString()),
_ -> null, leadNameForTerminal, () -> liveLeadTerminals, fixedClock(clock), () -> { },
runner);
_ -> null, leadNameForTerminal, () -> liveLeadTerminals, _ -> true, fixedClock(clock),
() -> { }, runner);
LeadRollover.PendingRollover first = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision firstDecision = rollover.confirm(LEAD, first.token(), true);
@@ -1587,7 +1753,8 @@ class LeadRolloverTest {
roster.put(LEAD, LEAD_NAME);
Map<String, String> liveLeadTerminals = Map.of(NEW_TERMINAL, LEAD_NAME);
LeadRollover rollover = new LeadRollover(agents, spaces, launcher, () -> cfg(handover.toString()),
_ -> null, roster::get, () -> liveLeadTerminals, fixedClock(clock), () -> { }, Runnable::run);
_ -> null, roster::get, () -> liveLeadTerminals, _ -> true, fixedClock(clock), () -> { },
Runnable::run);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
@@ -1637,7 +1804,7 @@ class LeadRolloverTest {
Map<String, String> roster = new HashMap<>();
roster.put(LEAD, LEAD_NAME);
LeadRollover rollover = new LeadRollover(agents, spaces, launcher, () -> cfg(handover.toString()),
_ -> null, roster::get, Map::of, fixedClock(clock), () -> { }, Runnable::run);
_ -> null, roster::get, Map::of, _ -> true, fixedClock(clock), () -> { }, Runnable::run);
LeadRollover.PendingRollover first = rollover.open(LEAD, "context is full");
@@ -56,6 +56,7 @@ class FleetMcpAuthzTest {
private final AgentControl agents = new AgentControl(herdr);
private Metrics metrics;
private FleetMcp mcp;
private MessageService messages;
@AfterEach
void close() {
@@ -82,7 +83,7 @@ class FleetMcpAuthzTest {
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
messages = new MessageService(agents, new Injector(agents), new Rendezvous(),
new InMemoryReplyInbox());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
metrics = FleetMetrics.create(sessions, new InMemoryReplyInbox());
@@ -270,18 +271,19 @@ class FleetMcpAuthzTest {
"a spawned member's own terminal must stay unreachable, even once the classifier is real");
}
// --- fleetd #743: the observer SEND matrix, over MCP's denyFor -------------------------------
// --- the observer SEND matrix, over MCP's denyFor -------------------------------------------
private static final Principal OBSERVER = Principal.observer("term_observer", 700);
/**
* Wires one real {@link CallerResolver} that recognises a lead, a collaborator, and a live
* spawned worker, leaving "term_other_observer" classified as none of them — so the same
* wiring both denies an observer's {@code SEND} to every privileged role and grants it to
* another unclassified pane, proving the refusals are the rule and not a missing fixture.
* wiring denies an observer's {@code SEND} to a collaborator and to a member while granting
* it to a lead and to another unclassified pane, proving the refusals are the rule and not a
* missing fixture.
*/
@Test
void anObserverMaySendOnlyToAnotherObserverNeverToALeadWorkerOrArchitect() {
void anObserverMaySendToALeadOrAnotherObserverButNeverToACollaboratorOrAMember() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_lead_known", "lead-x"), new MemberRegistry(null),
@@ -289,22 +291,23 @@ class FleetMcpAuthzTest {
() -> Map.of("term_collab_known", "ops2"));
FleetMcp m = mcp(true, callers);
assertNotNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_lead_known"),
"an observer must never reach a lead's terminal");
assertNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_lead_known"),
"an observer must reach a lead's terminal, so a peer session can open a "
+ "conversation with a lead");
assertNotNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_collab_known"),
"an observer must never reach a collaborator's terminal");
assertNotNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_a"),
"an observer must never reach a live spawned member's terminal");
// CONTROL: the same wiring, the same denyFor call, a target recognised as none of the
// three privileged roles above -- this is what proves the three refusals above are the
// rule working, not a classifier that refuses every target regardless of what it is.
// configured roles above -- this is what proves the two refusals above are the rule
// working, not a classifier that refuses every target regardless of what it is.
assertNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_other_observer"),
"an observer must reach another pane that resolves as an observer itself");
}
/**
* As {@link #anObserverMaySendOnlyToAnotherObserverNeverToALeadWorkerOrArchitect}, for a
* As {@link #anObserverMaySendToALeadOrAnotherObserverButNeverToACollaboratorOrAMember}, for a
* terminal bound to a configured architect slot but hosting no live spawned-member session --
* the case {@link CallerResolver#resolve} itself treats separately from a live worker/architect.
*/
@@ -324,6 +327,93 @@ class FleetMcpAuthzTest {
assertNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_other_observer"));
}
/**
* An observer's reach is widened for {@code SEND} alone. It still holds no {@code TASK_READ},
* so it cannot poll a ticket or read a lead's session status, and no {@code SPAWN}/
* {@code STOP}/{@code COORD_SEND}, so it cannot drive the fleet it can now message.
*/
@Test
void anObserverReachingALeadStillHoldsNoTicketReadAndNoLifecycle() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_lead_known", "lead-x"), new MemberRegistry(null), t -> null, Map::of);
FleetMcp m = mcp(true, callers);
assertNull(m.denyFor(OBSERVER, Authz.Action.SEND, "term_lead_known"),
"premise: this wiring grants the observer's send to that lead");
assertNotNull(m.denyFor(OBSERVER, Authz.Action.TASK_READ, "term_lead_known"));
assertNotNull(m.denyFor(OBSERVER, Authz.Action.SPAWN, null));
assertNotNull(m.denyFor(OBSERVER, Authz.Action.STOP, null));
assertNotNull(m.denyFor(OBSERVER, Authz.Action.COORD_SEND, null));
}
/**
* An observer may read a ticket its own {@code fleet_send(wait:false)} created, never one
* another caller created — exercised against the real {@link MessageService#ownsTicket}, not
* a stand-in classifier.
*/
@Test
void anObserverMayTaskReadATicketItCreatedButNotOneAnotherCallerCreated() {
FleetMcp m = mcp(true);
String ownTicket = messages.sendAsync("term_a", "do it", null, OBSERVER);
String othersTicket = messages.sendAsync("term_a", "do it too", null, WORKER_A);
assertNull(m.denyFor(OBSERVER, Authz.Action.TASK_READ, ownTicket,
t -> messages.ownsTicket(t, OBSERVER.ownerKey())),
"an observer must read a ticket its own send created");
assertNotNull(m.denyFor(OBSERVER, Authz.Action.TASK_READ, othersTicket,
t -> messages.ownsTicket(t, OBSERVER.ownerKey())),
"an observer must not read a ticket a different caller created");
}
/**
* As {@link #anObserverMayTaskReadATicketItCreatedButNotOneAnotherCallerCreated} (fleetd
* #804): a collaborator's {@code fleet_send(wait:false)} creates a ticket too, and the only
* caller that can ever read it back is the collaborator that created it.
*/
@Test
void aCollaboratorMayTaskReadATicketItCreatedButNotOneAnotherCallerCreated() {
FleetMcp m = mcp(true);
String ownTicket = messages.sendAsync("term_a", "do it", null, COLLABORATOR);
String othersTicket = messages.sendAsync("term_a", "do it too", null, WORKER_A);
assertNull(m.denyFor(COLLABORATOR, Authz.Action.TASK_READ, ownTicket,
t -> messages.ownsTicket(t, COLLABORATOR.ownerKey())),
"a collaborator must read a ticket its own send created");
assertNotNull(m.denyFor(COLLABORATOR, Authz.Action.TASK_READ, othersTicket,
t -> messages.ownsTicket(t, COLLABORATOR.ownerKey())),
"a collaborator must not read a ticket a different caller created");
}
/**
* fleetd #803: {@code fleet_poll{}} with no {@code ticket} argument resolves to
* {@code TASK_READ} with a {@code null} ticket. Before the fix, the real {@code
* ownsTicket(String, String)} classifier reached {@code tasks.get(null)} on a
* {@code ConcurrentHashMap} and threw a {@code NullPointerException} instead of refusing —
* the gate never got to write its {@code AuditLog} refusal, the caller saw an internal error.
* Asserts the refusal itself, not merely that nothing throws: a test that only catches the
* absence of a throw would also pass if this path quietly started granting the call.
*/
@Test
void anObserverPollingWithNoTicketIsRefusedRatherThanThrowing() {
FleetMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(OBSERVER, Authz.Action.TASK_READ, null,
t -> messages.ownsTicket(t, OBSERVER.ownerKey()));
assertNotNull(denied, "a missing ticket must be refused, not treated as owned");
}
/**
* {@code fleet_status} stays refused for an observer even once {@code TASK_READ} opens: it is
* not ticket-scoped, so no call site supplies a real {@code ticketOwnedByCaller} classifier,
* and the default 3-argument {@code denyFor} fails closed.
*/
@Test
void anObserverMayNotReadSessionStatusEvenAfterTaskReadOpensForTickets() {
FleetMcp m = mcp(true);
assertNotNull(m.denyFor(OBSERVER, Authz.Action.TASK_READ, "term_a"),
"a session id is not a ticket the caller could ever own");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing FleetMcpTest cases rely on no authorization being enforced.
@@ -477,9 +567,10 @@ class FleetMcpAuthzTest {
/**
* {@link FleetMcp#leadsVisibleTo} is the whole policy decision for {@code fleet_list}'s
* {@code leads} array: visible to exactly the roles that may {@code SEND} to a lead -- the
* primary, an architect, and a collaborator -- never a worker, which holds {@code READ} but
* can never {@code SEND} at all, and never an anonymous caller.
* {@code leads} array: visible to the primary, an architect, and a collaborator -- never a
* worker, which holds {@code READ} but can never {@code SEND} at all, never an observer, which
* reads a lead's address from its filtered {@code panes} rows instead, and never an anonymous
* caller.
*/
@Test
void primaryArchitectAndCollaboratorMaySeeTheLeadsArray() {
@@ -489,6 +580,9 @@ class FleetMcpAuthzTest {
"a collaborator may SEND to a lead, so it must see the leads array to learn where");
assertFalse(FleetMcp.leadsVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the leads array");
assertFalse(FleetMcp.leadsVisibleTo(Principal.observer("term_obs", 700)),
"an observer may SEND to a lead but learns the address from its panes rows, which "
+ "carry no lead name, context window or config dir");
assertFalse(FleetMcp.leadsVisibleTo(ANON), "authenticated as nothing must not see it either");
}
@@ -512,19 +606,21 @@ class FleetMcpAuthzTest {
/**
* {@link FleetMcp#panesVisibleTo} is the whole policy decision for {@code fleet_list}'s
* {@code panes} array: visible to exactly the roles that may {@code SEND} to a named peer --
* the primary, an architect, and a collaborator -- never a worker, never an unconfigured
* observer pane, and never an anonymous caller.
* {@code panes} array: visible to every role that may {@code SEND} to some other pane -- the
* primary, an architect, a collaborator, and an observer (to a lead or another observer pane,
* with its rows filtered and reduced -- see {@code listFleet}) -- never a worker, never an
* anonymous caller.
*/
@Test
void onlyPrimaryArchitectAndCollaboratorMaySeeThePanesArray() {
void onlyPrimaryArchitectCollaboratorAndObserverMaySeeThePanesArray() {
assertTrue(FleetMcp.panesVisibleTo(PRIMARY), "the primary must see the panes array");
assertTrue(FleetMcp.panesVisibleTo(ARCH_DESIGN), "an architect must see the panes array");
assertTrue(FleetMcp.panesVisibleTo(COLLABORATOR), "a collaborator must see its own peer roster");
assertFalse(FleetMcp.panesVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the panes array");
assertFalse(FleetMcp.panesVisibleTo(Principal.observer("term_obs", 700)),
"an unconfigured observer pane must not see every other pane's label and cwd");
assertTrue(FleetMcp.panesVisibleTo(Principal.observer("term_obs", 700)),
"an observer holds SEND to a lead and to another observer pane, so it must see the "
+ "(filtered, reduced) panes array");
assertFalse(FleetMcp.panesVisibleTo(ANON), "authenticated as nothing must not see it either");
}
@@ -94,7 +94,7 @@ class FleetMcpHandoverTest {
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = new LeadLauncher(agents, spaces, minimalFleetConfig());
return new LeadRollover(agents, spaces, launcher, () -> cfg(handoverPath),
_ -> null, _ -> null, Map::of);
_ -> null, _ -> null, Map::of, _ -> true);
}
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
@@ -0,0 +1,178 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.Fleetd;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.inject.TurnListener;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.session.FakeWorktrees;
import dev.ltms.fleet.session.SessionManager;
import io.modelcontextprotocol.client.McpClient;
import io.modelcontextprotocol.client.McpSyncClient;
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
import io.modelcontextprotocol.spec.McpClientTransport;
import io.modelcontextprotocol.spec.McpSchema;
import org.eclipse.jetty.server.Server;
import org.eclipse.jetty.server.ServerConnector;
import org.eclipse.jetty.servlet.ServletContextHandler;
import org.eclipse.jetty.servlet.ServletHolder;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Predicate;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* An observer's {@code fleet_send} to a lead, driven end to end: a real MCP client over a real
* HTTP transport, resolved by the real {@link CallerResolver} to
* {@link dev.ltms.fleet.auth.Role#OBSERVER}, through the real {@link MessageService} and the real
* {@link Injector} to the point of its real herdr call.
*
* <p>{@link FleetMcpAuthzTest} proves {@code denyFor} grants this case. This is the path that also
* proves the grant is not dead at the injector's readiness gate: that gate is the production
* {@link Fleetd#deliverableTo} predicate, and the lead's terminal carries no
* {@link MemberPresence} entry, so the delivery can only pass by the lead being a configured lead.
*/
class FleetMcpObserverSendToLeadDeliveryTest {
private static final String LEAD = "term_lead_pane";
private static final String COLLABORATOR = "term_collab_pane";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final MemberPresence presence = new MemberPresence();
private Injector injector;
private MessageService messages;
private FleetMcp mcp;
private Server server;
private String baseUrl;
@BeforeEach
void startServer() throws Exception {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "tok");
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
// The fake's pane list carries a second pane, "term_shell", whose shell pid is 9001 and
// which hosts no agent -- so the real resolver lands the caller on the observer floor.
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 9001L);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of(LEAD, "fleet01-lead"), new MemberRegistry(null),
_ -> null, () -> Map.of(COLLABORATOR, "ops"));
Predicate<String> deliverable =
Fleetd.deliverableTo(presence, callers::leads, callers::collaborators);
injector = new Injector(agents, TurnListener.NOOP, deliverable);
messages = new MessageService(agents, injector, rendezvous, new InMemoryReplyInbox());
mcp = new FleetMcp(messages, workers, sessions, identity, presence,
new PrimaryRegistry(null), callers, FleetMcp.AuthorizationMode.ENFORCED,
null, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), List.of(), null);
ServletContextHandler handler = new ServletContextHandler();
handler.setContextPath("/");
handler.addServlet(new ServletHolder(mcp.servlet()), "/mcp");
server = new Server(0);
server.setHandler(handler);
server.start();
baseUrl = "http://127.0.0.1:"
+ ((ServerConnector) server.getConnectors()[0]).getLocalPort();
}
@AfterEach
void tearDown() throws Exception {
if (server != null) {
server.stop();
}
if (mcp != null) {
mcp.close();
}
}
@Test
void anObserversSendToALeadIsAttributedAndReachesTheRealInjector() throws Exception {
assertFalse(presence.isPresent(LEAD),
"premise: the lead's terminal is deliverable only as a configured lead, never "
+ "through a presence entry");
McpSchema.CallToolResult result = sendFleetSend(LEAD, "can we split the review?");
assertFalse(result.isError(), "an observer sending to a lead must be accepted: "
+ textOf(result));
long waiterDeadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(LEAD) && System.currentTimeMillis() < waiterDeadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(LEAD), "the async send must have opened its rendezvous waiter");
injector.onStatus(LEAD, AgentStatus.IDLE); // drives the real delivery attempt to herdr
long deliveryDeadline = System.currentTimeMillis() + 3000;
while (!herdr.called("agent.prompt") && System.currentTimeMillis() < deliveryDeadline) {
Thread.sleep(5);
}
assertTrue(herdr.called("agent.prompt"), "the delivery attempt must have reached herdr");
@SuppressWarnings("unchecked")
Map<String, Object> params = (Map<String, Object>) herdr.lastCall("agent.prompt").params();
assertEquals("[fleet_send from observer term_shell]\ncan we split the review?",
params.get("text"),
"a lead must see the sender's own daemon-resolved terminal, never a raw echo of "
+ "the content and never a client-supplied name");
}
/**
* Control for the test above, over the same live server and the same wiring: the grant is
* specific to a lead target, so a collaborator's terminal is still refused at the handler and
* nothing is ever queued for it.
*/
@Test
void thatSameObserverIsStillRefusedACollaboratorsTerminal() {
McpSchema.CallToolResult result = sendFleetSend(COLLABORATOR, "can we split the review?");
assertTrue(result.isError(), "an observer must not reach a collaborator's terminal");
assertFalse(rendezvous.isWaiting(COLLABORATOR),
"a refused send must never open a waiter for its target");
}
private McpSchema.CallToolResult sendFleetSend(String target, String content) {
McpClientTransport transport = HttpClientStreamableHttpTransport.builder(baseUrl)
.endpoint("/mcp")
.build();
try (McpSyncClient client = McpClient.sync(transport).build()) {
client.initialize();
return client.callTool(McpSchema.CallToolRequest.builder("fleet_send")
.arguments(Map.of("sessionId", target, "content", content, "wait", false))
.build());
}
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
}
@@ -78,7 +78,7 @@ class FleetMcpTest {
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles, null));
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles, null, true));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -104,7 +104,7 @@ class FleetMcpTest {
void sendThenReplyRoundTrips() throws Exception {
// fleet_send blocks; fleet_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of(), null));
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of(), null, true));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
@@ -182,7 +182,7 @@ class FleetMcpTest {
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null, true));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
@@ -280,7 +280,7 @@ class FleetMcpTest {
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L, null));
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L, null, true));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -298,7 +298,7 @@ class FleetMcpTest {
@Test
void aDifferentCallersMcpAnswerIsRefusedForABlockingSendButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "do X", 5000L, null, Set.of(), "term_owner"));
() -> FleetMcp.send(messages, T, "do X", 5000L, null, Set.of(), "term_owner", true));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -313,12 +313,12 @@ class FleetMcpTest {
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker");
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker", true);
assertTrue(hijacked.isError(), "a caller that did not open this turn must get an error, not an answer");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner"));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner", true));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
@@ -357,14 +357,14 @@ class FleetMcpTest {
String turnId = asking.turnId();
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L,
"worker:term_attacker");
"worker:term_attacker", true);
assertTrue(hijacked.isError(), "a caller that did not create this delegation must get an error");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, owner.ownerKey()));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, owner.ownerKey(), true));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
@@ -383,11 +383,71 @@ class FleetMcpTest {
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null);
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null, true);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
/**
* fleetd #801: a blocking send creates no ticket, ever, so the TIMED_OUT_QUEUED receipt must
* name the one route that recovers a late answer — draining the target's inbox — and must not
* invite a resend, which can duplicate the delivery.
*/
@Test
void sendTimesOutQueuedNamesFleetPollTargetAndNotTheBareWordRetry() {
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null, true);
String text = textOf(res);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(text.contains("fleet_poll{target=\"term_a\"}"), "got: " + text);
assertFalse(text.contains("retry"), "got: " + text);
}
/** As above, for the TIMED_OUT_WORKING arm (delivered, but the worker never replies). */
@Test
void sendTimesOutWorkingNamesFleetPollTargetAndNotTheBareWordRetry() throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null, true));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts, never replies
McpSchema.CallToolResult res = send.get(5, TimeUnit.SECONDS);
String text = textOf(res);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(text.contains("fleet_poll{target=\"" + T + "\"}"), "got: " + text);
assertFalse(text.contains("retry"), "got: " + text);
}
/** As above, for the BUSY arm (another send already held the session for the whole window). */
@Test
void sendTimesOutBusyNamesFleetPollTargetAndNotTheBareWordRetry() throws Exception {
CompletableFuture<McpSchema.CallToolResult> first = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "first", 5000L, null, Set.of(), null, true));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "the first send should hold the session lock");
McpSchema.CallToolResult busy = FleetMcp.send(messages, T, "second", 100L, null, Set.of(), null, true);
String text = textOf(busy);
assertNotEquals(Boolean.TRUE, busy.isError(), "a timeout is informational, not a tool error");
assertTrue(text.contains("fleet_poll{target=\"" + T + "\"}"), "got: " + text);
assertFalse(text.contains("retry"), "got: " + text);
// Release the first send so the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
FleetMcp.reply(messages, T, Role.WORKER, "first done");
first.get(5, TimeUnit.SECONDS);
}
/**
* fleetd #571 (ticket CORRECTION 5): {@code formatReply}'s {@code TIMED_OUT_UNCONFIRMED} arm is
* the one message whose whole job is to stop a caller retrying a delivery that may already have
@@ -399,7 +459,7 @@ class FleetMcpTest {
void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception {
herdr.agentSendFailsWith("send_failed");
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null));
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null, true));
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -417,15 +477,59 @@ class FleetMcpTest {
+ "a resend here can double-deliver the same brief: got " + text);
}
/**
* fleetd #801 defect 1: {@code SEND} is granted to callers {@code DRAIN} refuses (an architect,
* a collaborator, an observer), so a TIMED_OUT/BUSY receipt must not name {@code fleet_poll}
* when the caller's own grant would be refused there — it must say the reply is unrecoverable
* instead.
*/
@Test
void sendTimesOutWithoutDrainGrantNamesNoRouteAndAdmitsItCannotBeRecovered() {
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null, false);
String text = textOf(res);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertFalse(text.contains("fleet_poll{target"), "got: " + text);
assertTrue(text.contains("cannot be recovered"), "got: " + text);
}
/** As above, for {@code answer()} — {@code ANSWER} is primary-or-architect, and an architect is refused DRAIN too. */
@Test
void answerTimesOutWithoutDrainGrantNamesNoRoute() throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null, true));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the send must be waiting for the ask to surface to");
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, "term_a", "which config?", 5000L));
McpSchema.CallToolResult q = send.get(6, TimeUnit.SECONDS);
String qt = textOf(q);
String afterMarker = qt.substring(qt.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterMarker.substring(0, afterMarker.indexOf('"'));
// The resumed worker never replies, so the answer's own wait times out.
McpSchema.CallToolResult answer = FleetMcp.answer(messages, turnId, "config.yaml", 150L, null, false);
String text = textOf(answer);
assertNotEquals(Boolean.TRUE, answer.isError(), "a timeout is informational, not a tool error");
assertFalse(text.contains("fleet_poll{target"), "got: " + text);
assertTrue(text.contains("cannot be recovered"), "got: " + text);
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
}
@Test
void sendRejectsMissingArgs() {
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of(), null).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of(), null).isError());
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of(), null, true).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of(), null, true).isError());
}
@Test
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"), null);
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"), null, true);
McpSchema.CallToolResult async = FleetMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
assertTrue(blocking.isError());
@@ -444,7 +548,7 @@ class FleetMcpTest {
assertSendRoundTrips("term_live_member", profiles);
// A herdr-owned pane outside the bridge roster cannot be classified at accept time.
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles, null);
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles, null, true);
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
}
@@ -536,7 +640,7 @@ class FleetMcpTest {
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null));
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null, true));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -558,7 +662,7 @@ class FleetMcpTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null, true));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
@@ -584,7 +688,7 @@ class FleetMcpTest {
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L, null);
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L, null, true);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@@ -1129,7 +1233,8 @@ class FleetMcpTest {
/**
* A pane whose tab herdr reports with a label gets that label and the exact terminal id
* {@code fleet_send} takes as a target, carried as {@code sessionId}.
* {@code fleet_send} takes as a target, carried as {@code sessionId}. fleetd #771: the pane's
* workspace carries its herdr space name too, next to {@code workspaceId}.
*/
@Test
void listReportsAPaneRowWithItsTabLabelAndSendableSessionId() {
@@ -1138,13 +1243,16 @@ class FleetMcpTest {
presence.markPresent("term_a");
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(
() -> new PaneLocator(h).tabLabelsByTabId(),
Fleetd.deliverableTo(presence, Map::of, Map::of));
() -> new PaneLocator(h).workspaceLabelsByWorkspaceId(),
Fleetd.deliverableTo(presence, Map::of, Map::of), _ -> false, _ -> false);
String out = textOf(listFleetWithPanes(h, panes, true));
assertTrue(out.contains("\"panes\":["), out);
assertTrue(out.contains("\"sessionId\":\"term_a\""), out);
assertTrue(out.contains("\"label\":\"trinotes\""), out);
assertTrue(out.contains("\"workspaceLabel\":\"ltms\""),
"term_a's agent lives on workspace w2, whose herdr label is \"ltms\": " + out);
assertTrue(out.contains("\"deliverable\":true"), out);
}
@@ -1159,7 +1267,8 @@ class FleetMcpTest {
FakeHerdr h = new FakeHerdr();
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(
() -> new PaneLocator(h).tabLabelsByTabId(),
Fleetd.deliverableTo(new MemberPresence(), Map::of, Map::of));
() -> new PaneLocator(h).workspaceLabelsByWorkspaceId(),
Fleetd.deliverableTo(new MemberPresence(), Map::of, Map::of), _ -> false, _ -> false);
String out = textOf(listFleetWithPanes(h, panes, true));
@@ -1169,6 +1278,29 @@ class FleetMcpTest {
assertTrue(out.contains("\"deliverable\":false"), out);
}
/**
* fleetd #771: a pane whose agent lives in a workspace that {@code workspace.list} does not
* report (an unknown {@code workspaceId}) still gets a row -- the lookup miss must never throw,
* and must never drop the pane, only report a {@code null} "workspaceLabel".
*/
@Test
void listReportsANullWorkspaceLabelForAnUnknownWorkspaceId() {
// withAgent seeds its pane under workspace_id "wQ", which the fake's workspace.list never
// reports (only "w1"/"w2") -- modelling a workspace the lookup has no entry for.
FakeHerdr h = new FakeHerdr().withAgent("claude-x", "term_unknown_ws", "wQ:p1", "wQ:t1");
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(
() -> new PaneLocator(h).tabLabelsByTabId(),
() -> new PaneLocator(h).workspaceLabelsByWorkspaceId(),
_ -> false, _ -> false, _ -> false);
String out = textOf(listFleetWithPanes(h, panes, true));
assertTrue(out.contains("\"sessionId\":\"term_unknown_ws\""), out);
assertTrue(out.contains("\"workspaceId\":\"wQ\""), out);
assertTrue(out.contains("\"workspaceLabel\":null"),
"an unknown workspaceId must project a null workspaceLabel, not throw or drop the row: " + out);
}
/** A caller this role may not show the array to gets no {@code panes} key at all. */
@Test
void listOmitsThePanesArrayWhenTheCallerMayNotSeeIt() {
@@ -1183,14 +1315,16 @@ class FleetMcpTest {
* The tab-label scan behind {@code panes} shares no failure path with the rest of
* {@code listFleet} -- a {@code workspace.list}/{@code tab.list} failure costs only the
* labels in the {@code panes} row (each renders {@code null}), never the {@code leads}/
* {@code members} arrays, which never needed that scan at all.
* {@code members} arrays, which never needed that scan at all. fleetd #771: the same
* {@code workspace.list} failure costs {@code workspaceLabel} the same way.
*/
@Test
void listStillReportsEveryOtherArrayWhenTheLabelScanFails() {
FakeHerdr h = new FakeHerdr().workspaceListFailsWith("unavailable");
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(
() -> new PaneLocator(h).tabLabelsByTabId(),
Fleetd.deliverableTo(new MemberPresence(), Map::of, Map::of));
() -> new PaneLocator(h).workspaceLabelsByWorkspaceId(),
Fleetd.deliverableTo(new MemberPresence(), Map::of, Map::of), _ -> false, _ -> false);
McpSchema.CallToolResult res = listFleetWithPanes(h, panes, true);
@@ -1199,10 +1333,137 @@ class FleetMcpTest {
assertTrue(out.contains("\"panes\":["), out);
assertTrue(out.contains("\"sessionId\":\"term_a\""), out);
assertTrue(out.contains("\"label\":null"), out);
assertTrue(out.contains("\"workspaceLabel\":null"),
"a workspace.list failure must not cost the leads/members arrays, only a null "
+ "workspaceLabel: " + out);
assertTrue(out.contains("\"leads\":[]"), "a label-scan failure must not cost the leads array: " + out);
assertTrue(out.contains("\"members\":[]"), "a label-scan failure must not cost the members array: " + out);
}
/**
* fleetd #756: a pane bound to a configured architect slot with no live member session must
* report {@code role: "architect"}, read from {@link CallerResolver#boundToArchitectSlot} —
* the same classifier {@link CallerResolver#observerSendTarget} refuses as a {@code SEND}
* target — rather than falling through to {@code "observer"}.
*/
@Test
void listReportsArchitectForASlotBoundPaneWithNoLiveMember() {
FakeHerdr h = new FakeHerdr()
.withAgent("claude-arch", "term_unoccupied_architect", "w2:pArch", "w2:tArch");
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(
Map::of, Map::of, _ -> false, "term_unoccupied_architect"::equals, _ -> false);
String out = textOf(listFleetWithPanes(h, panes, true));
assertTrue(out.contains("\"sessionId\":\"term_unoccupied_architect\""), out);
assertTrue(out.contains("\"role\":\"architect\""),
"a slot-bound pane with no live session must read \"architect\", not the generic "
+ "\"observer\" fallback: " + out);
}
/**
* Calls the canonical {@code listFleet} overload directly with an explicit {@code leads}/
* {@code collaborators} payload and {@code callerIsObserver}, mirroring exactly what the real
* {@code fleet_list} handler computes for an observer caller: {@code panesVisible} true,
* {@code leadsVisible}/{@code membersVisible}/{@code collaboratorsVisible} false.
*/
private static McpSchema.CallToolResult listFleetAsObserver(FakeHerdr h, Map<String, String> leads,
Map<String, String> collaborators, FleetMcp.PaneSource panes) {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
return FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(),
leads, "", collaborators, false,
FleetMcp.CoordinationSource.none(), false, false, false, panes, true, true);
}
/**
* An observer's {@code fleet_list} carries a {@code panes} key, filtered to
* {@link CallerResolver#observerSendTarget} (so a lead's pane survives, while a spawned
* member's pane, a collaborator's pane and an unoccupied architect-slot pane are all absent)
* and every surviving row reduced to exactly {@code sessionId}, {@code label}, {@code status},
* {@code role}, {@code deliverable} — never {@code paneId}, {@code workspaceId},
* {@code workspaceLabel} (a space name is host shape, a stronger disclosure than a pane id, so
* it stays out of the reduced row too), {@code tabId}, {@code agentType}, or {@code cwd}.
*/
@Test
void listFiltersAndReducesThePanesArrayForAnObserver() {
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null));
assertTrue(members.bind("architect:lead-designer", "term_architect_pane"));
CallerResolver callers = CallerResolver.withLeadsAndMembers(null, false, null,
() -> Map.of("term_lead_pane", "fleet01-lead"), members,
t -> "term_member_pane".equals(t) ? MemberRole.DEV : null,
() -> Map.of("term_collab_pane", "ops"));
FakeHerdr h = new FakeHerdr()
.withAgent("claude-sendable", "term_sendable", "w2:pS", "w2:tS")
.withAgent("claude-lead", "term_lead_pane", "w2:pL", "w2:tL")
.withAgent("claude-member", "term_member_pane", "w2:pM", "w2:tM")
.withAgent("claude-collab", "term_collab_pane", "w2:pC", "w2:tC")
.withAgent("claude-arch", "term_architect_pane", "w2:pA", "w2:tA")
.withTab("w2", "w2:tS", "trinotes");
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(
() -> new PaneLocator(h).tabLabelsByTabId(),
() -> new PaneLocator(h).workspaceLabelsByWorkspaceId(), _ -> true,
callers::boundToArchitectSlot, callers.observerSendTarget());
String out = textOf(listFleetAsObserver(h, callers.leads(),
Map.of("term_collab_pane", "ops"), panes));
assertTrue(out.contains("\"panes\":["), out);
assertTrue(out.contains("\"sessionId\":\"term_sendable\""),
"an ordinary unclassified pane must still be sendable and visible: " + out);
assertTrue(out.contains("\"sessionId\":\"term_lead_pane\""),
"a lead's pane is where an observer reads the sessionId its send needs: " + out);
assertTrue(out.contains("\"role\":\"lead\""),
"the lead's row must name the role, so an observer can tell it from a peer pane: " + out);
assertFalse(out.contains("term_member_pane"), "a spawned member's pane must not be enumerated: " + out);
assertFalse(out.contains("term_collab_pane"), "a collaborator's pane must not be enumerated: " + out);
assertFalse(out.contains("term_architect_pane"),
"an unoccupied architect-slot pane must not be enumerated: " + out);
assertFalse(out.contains("\"paneId\""), "an observer's row must never carry paneId: " + out);
assertFalse(out.contains("\"workspaceId\""), "an observer's row must never carry workspaceId: " + out);
assertFalse(out.contains("\"workspaceLabel\""),
"an observer's row must never carry workspaceLabel: " + out);
assertFalse(out.contains("\"tabId\""), "an observer's row must never carry tabId: " + out);
assertFalse(out.contains("\"agentType\""), "an observer's row must never carry agentType: " + out);
assertFalse(out.contains("\"cwd\""), "an observer's row must never carry cwd: " + out);
}
/**
* Control for the test above: a primary's {@code panes} row is unchanged by fleetd #758 —
* {@code callerIsObserver} false keeps every field, including {@code paneId} and a spawned
* member's {@code cwd}.
*/
@Test
void listKeepsTheFullPaneRowForAPrimaryIncludingCwdAndPaneId() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
MemberSession spawned = sessions.acquire("ltms-local", "/worktree/member-1", null, null);
h.withAgent("claude-member", spawned.terminalId(), "w9:pMember", "w9:tMember");
FleetMcp.PaneSource panes = new FleetMcp.PaneSource(Map::of, Map::of, _ -> true, _ -> false, _ -> false);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(),
Map.of(), "", Map.of(), false,
FleetMcp.CoordinationSource.none(), true, true, true, panes, true, false);
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"" + spawned.terminalId() + "\""), out);
assertTrue(out.contains("\"paneId\":\"w9:pMember\""),
"a primary must still see the fleet_stop handle: " + out);
assertTrue(out.contains("\"cwd\":\"/worktree/member-1\""),
"a primary must still see a spawned member's worktree path: " + out);
assertTrue(out.contains("\"role\":\"dev\""), out);
}
/**
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
@@ -266,4 +266,61 @@ class PrimaryRegistryTest {
assertEquals("term_opus_after_roll", reg.nudgeTargetFor("term_never_seen").orElseThrow());
}
// ── fleetd #778: a nudge must offer only what the delegator's own role may run ──────────────
@Test
void theShortDelegationOverloadsGrantEverythingAPrimaryMayRun() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_lead");
assertEquals(Boolean.TRUE, reg.nudgeMayDrainFor("term_worker").orElseThrow());
assertEquals(Boolean.TRUE, reg.nudgeMayAnswerFor("term_worker").orElseThrow());
}
@Test
void theFullOverloadCarriesTheCallersOwnGrants() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_observer", "observer-name", false, false);
assertEquals(Boolean.FALSE, reg.nudgeMayDrainFor("term_worker").orElseThrow());
assertEquals(Boolean.FALSE, reg.nudgeMayAnswerFor("term_worker").orElseThrow());
}
@Test
void theTwoGrantsAreIndependent() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_architect", "architect-name", false, true);
assertEquals(Boolean.FALSE, reg.nudgeMayDrainFor("term_worker").orElseThrow(),
"an architect may not DRAIN");
assertEquals(Boolean.TRUE, reg.nudgeMayAnswerFor("term_worker").orElseThrow(),
"an architect may ANSWER");
}
@Test
void theSingletonPrimaryFallbackIsAlwaysGrantedEverything() {
var reg = new PrimaryRegistry("term_pinned");
assertEquals(Boolean.TRUE, reg.nudgeMayDrainFor("term_never_seen").orElseThrow());
assertEquals(Boolean.TRUE, reg.nudgeMayAnswerFor("term_never_seen").orElseThrow());
}
@Test
void withNoDelegationAndNoPrimaryTheGrantsAreUnknown() {
var reg = new PrimaryRegistry(null);
assertTrue(reg.nudgeMayDrainFor("term_never_seen").isEmpty());
assertTrue(reg.nudgeMayAnswerFor("term_never_seen").isEmpty());
}
@Test
void forgettingADelegationForgetsItsGrantsToo() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_observer", "observer-name", false, false);
reg.forgetDelegation("term_worker");
assertTrue(reg.nudgeMayDrainFor("term_worker").isEmpty());
assertTrue(reg.nudgeMayAnswerFor("term_worker").isEmpty());
}
}
@@ -40,7 +40,7 @@ class LeadCoordLoopTest {
@Test
void deliversAHeldMessageToTheLeadPaneAndAcksIt() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "the merge is blocked"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
loop(channel, herdr, Map.of(LEAD_TERM, SELF)).tick();
@@ -57,7 +57,7 @@ class LeadCoordLoopTest {
@Test
void redeliveryOfAMessageAlreadyWrittenToThePaneIsAckedWithoutAnotherPaneWrite() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "recover me"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
var loop = loop(channel, herdr, Map.of(LEAD_TERM, SELF));
loop.tick();
@@ -71,7 +71,7 @@ class LeadCoordLoopTest {
@Test
void aRedeliveryIsAckedEvenWhileTheLeadIsMidTurn() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "recover me"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
var loop = loop(channel, herdr, Map.of(LEAD_TERM, SELF));
loop.tick();
@@ -91,7 +91,7 @@ class LeadCoordLoopTest {
@Test
void leavesTheMessageUnackedWhenTheLeadIsMidTurn() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "hello"));
var herdr = new FakeHerdr().agentStatus("working");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("working");
loop(channel, herdr, Map.of(LEAD_TERM, SELF)).tick();
@@ -103,7 +103,7 @@ class LeadCoordLoopTest {
@Test
void leavesTheMessageUnackedWhenNoLeadPaneIsKnown() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "hello"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
loop(channel, herdr, Map.of()).tick();
@@ -115,7 +115,8 @@ class LeadCoordLoopTest {
@Test
void leavesTheMessageUnackedWhenHerdrRefusesTheInjection() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "hello"));
var herdr = new FakeHerdr().agentStatus("idle").agentSendFailsWith("agent_not_found");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET)
.agentStatus("idle").agentSendFailsWith("agent_not_found");
loop(channel, herdr, Map.of(LEAD_TERM, SELF)).tick();
@@ -126,7 +127,7 @@ class LeadCoordLoopTest {
@Test
void resolvesTheLeadByNameWhenSeveralAreKnown() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "hello"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
// Two leads on this daemon; only one carries the coord-id the mailbox is owned as.
var leads = new java.util.LinkedHashMap<String, String>();
leads.put("term_other", "some-other-lead");
@@ -142,7 +143,7 @@ class LeadCoordLoopTest {
@Test
void holdsWhenSeveralLeadsAreKnownAndNoneCarriesTheCoordId() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "hello"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
var leads = new java.util.LinkedHashMap<String, String>();
leads.put("term_one", "lead-one");
leads.put("term_two", "lead-two");
@@ -159,7 +160,7 @@ class LeadCoordLoopTest {
var channel = new FakeLeadChannel(SELF)
.hold(new LeadMessage("m1", PEER, SELF, "first"))
.hold(new LeadMessage("m2", PEER, SELF, "second"));
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
var loop = loop(channel, herdr, Map.of(LEAD_TERM, SELF));
loop.tick();
@@ -175,10 +176,66 @@ class LeadCoordLoopTest {
@Test
void anEmptyMailboxNeverTouchesHerdr() {
var channel = new FakeLeadChannel(SELF);
var herdr = new FakeHerdr().agentStatus("idle");
var herdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET).agentStatus("idle");
loop(channel, herdr, Map.of(LEAD_TERM, SELF)).tick();
assertEquals(0, herdr.calls.size(), "an idle fleet must not poll a pane's status every tick");
}
@Test
void aLeadWithUnsubmittedTextInItsPromptBoxKeepsTheMessageHeldAndUnacked() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "the merge is blocked"));
var herdr = new FakeHerdr().detectionText(FakeHerdr.DRAFTED_PROMPT_CARET).agentStatus("idle");
loop(channel, herdr, Map.of(LEAD_TERM, SELF)).tick();
assertEquals(0, prompts(herdr).size(),
"delivery pastes and submits, so it must not land on a half-typed line");
assertEquals(List.of(), channel.acked(), "an undelivered message stays on the broker");
assertFalse(channel.peek().isEmpty(), "and is still held");
}
@Test
void aMessageHeldForADraftIsDeliveredOnALaterTick() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "the merge is blocked"));
var herdr = new FakeHerdr().detectionText(FakeHerdr.DRAFTED_PROMPT_CARET).agentStatus("idle");
var loop = loop(channel, herdr, Map.of(LEAD_TERM, SELF));
loop.tick();
assertEquals(0, prompts(herdr).size());
herdr.detectionText(FakeHerdr.IDLE_PROMPT_CARET);
loop.tick();
assertEquals(1, prompts(herdr).size(), "the held message lands once the box is empty");
assertEquals(List.of("m1"), channel.acked());
}
@Test
void anUnreadablePaneKeepsTheMessageHeld() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "the merge is blocked"));
var herdr = new FakeHerdr().detectionText("garbled ansi noise with no input box").agentStatus("idle");
loop(channel, herdr, Map.of(LEAD_TERM, SELF)).tick();
assertEquals(0, prompts(herdr).size(), "a pane whose box cannot be found may be holding a draft");
assertEquals(List.of(), channel.acked());
}
@Test
void aDisabledGateDeliversAHeldMessageDespiteUnsubmittedTextInThePromptBox() {
var channel = new FakeLeadChannel(SELF).hold(new LeadMessage("m1", PEER, SELF, "the merge is blocked"));
var herdr = new FakeHerdr().detectionText(FakeHerdr.DRAFTED_PROMPT_CARET).agentStatus("idle");
var loop = new LeadCoordLoop(channel, new AgentControl(herdr), () -> Map.of(LEAD_TERM, SELF),
NO_SCHEDULER, 3_000L, false);
loop.tick();
assertEquals(1, prompts(herdr).size(), "a disabled gate delivers even though the box holds a draft");
assertEquals(List.of("m1"), channel.acked());
assertTrue(herdr.calls.stream().noneMatch(c -> c.method().equals("agent.read")),
"a disabled gate never reads the pane");
}
}
@@ -7,6 +7,7 @@ import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadContextGauge;
@@ -24,6 +25,7 @@ import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
@@ -520,12 +522,20 @@ class LeadHeartbeatLoopTest {
private static LeadHeartbeatLoop tickableLoop(FailableHerdrClient herdr, AtomicLong now,
List<MemberSession>[] rosterBox, InMemoryReplyInbox inbox,
int quietNudgeCap, ScheduledExecutorService scheduler) {
return tickableLoop(herdr, now, rosterBox, inbox, quietNudgeCap, scheduler, true);
}
/** As above, with the prompt-box gate's on/off switch. */
private static LeadHeartbeatLoop tickableLoop(FailableHerdrClient herdr, AtomicLong now,
List<MemberSession>[] rosterBox, InMemoryReplyInbox inbox,
int quietNudgeCap, ScheduledExecutorService scheduler,
boolean promptBoxGateEnabled) {
AgentControl agents = new AgentControl(herdr);
PrimaryRegistry registry = new PrimaryRegistry(LEAD);
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, agents, inbox, scheduler, 5, 100_000);
return new LeadHeartbeatLoop(registry, agents, inbox, () -> rosterBox[0], pushLoop, scheduler,
now::get, IDLE_AFTER_NANOS, 100_000L, quietNudgeCap, null,
new LeadHeartbeatLoop.LeadContextSource(t -> highReading()), true);
new LeadHeartbeatLoop.LeadContextSource(t -> highReading()), true, true, promptBoxGateEnabled);
}
@Test
@@ -694,7 +704,10 @@ class LeadHeartbeatLoopTest {
private final String lead;
private final List<String> sentTexts = new ArrayList<>();
private final List<String> promptTargets = new ArrayList<>();
private final AtomicInteger reads = new AtomicInteger();
private boolean throwOnNextSend = false;
/** What {@code agent.read} reports — the loop reads the lead's input box before it nudges. */
private String paneTail = FakeHerdr.IDLE_PROMPT_CARET;
FailableHerdrClient(String lead) {
this.lead = lead;
@@ -704,6 +717,10 @@ class LeadHeartbeatLoopTest {
throwOnNextSend = true;
}
void paneTail(String tail) {
this.paneTail = tail;
}
List<String> sentTexts() {
return List.copyOf(sentTexts);
}
@@ -713,6 +730,10 @@ class LeadHeartbeatLoopTest {
return List.copyOf(promptTargets);
}
int readCount() {
return reads.get();
}
@Override
@SuppressWarnings("unchecked")
public JsonNode call(String method, Object params) {
@@ -722,6 +743,11 @@ class LeadHeartbeatLoopTest {
.put("terminal_id", lead)
.put("agent_status", "idle"));
}
if ("agent.read".equals(method)) {
reads.incrementAndGet();
return MAPPER.createObjectNode()
.set("read", MAPPER.createObjectNode().put("text", paneTail));
}
if ("agent.prompt".equals(method)) {
Map<String, Object> p = params instanceof Map ? (Map<String, Object>) params : Map.of();
if (throwOnNextSend) {
@@ -738,4 +764,68 @@ class LeadHeartbeatLoopTest {
public void close() {
}
}
// ── the operator's own prompt box ────────────────────────────────────────────────────────────
@Test
void tickHoldsTheNudgeWhileTheLeadsPromptBoxHoldsUnsubmittedText() {
var herdr = new FailableHerdrClient(LEAD);
herdr.paneTail(FakeHerdr.DRAFTED_PROMPT_CARET);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick(); // opens the idle window
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick(); // would INJECT, but the operator is mid-sentence
assertEquals(0, herdr.sentTexts().size(),
"a nudge pastes and submits, so it must not land on a half-typed line");
herdr.paneTail(FakeHerdr.IDLE_PROMPT_CARET);
loop.tick();
assertEquals(1, herdr.sentTexts().size(), "the held nudge lands once the box is empty");
assertTrue(herdr.sentTexts().get(0).contains("Your own context is nearly full"),
"and it still carries the notice the held tick did not spend: " + herdr.sentTexts().get(0));
}
@Test
void anUnreadablePaneHoldsTheHeartbeatNudge() {
var herdr = new FailableHerdrClient(LEAD);
herdr.paneTail("garbled ansi noise with no input box");
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler);
loop.tick();
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick();
assertEquals(0, herdr.sentTexts().size(),
"a pane whose box cannot be found may be holding a draft");
}
@Test
void aDisabledGateSendsTheNudgeDespiteUnsubmittedTextInThePromptBox() {
var herdr = new FailableHerdrClient(LEAD);
herdr.paneTail(FakeHerdr.DRAFTED_PROMPT_CARET);
var now = new AtomicLong(NOW);
@SuppressWarnings("unchecked")
List<MemberSession>[] rosterBox = new List[]{List.of()};
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop loop = tickableLoop(herdr, now, rosterBox, inbox, 0, scheduler, false);
loop.tick(); // opens the idle window
now.addAndGet(TimeUnit.SECONDS.toNanos(400));
loop.tick(); // a disabled gate must not hold on the operator's draft
assertEquals(1, herdr.sentTexts().size(),
"a disabled gate must send the nudge even though the box holds a draft");
assertEquals(0, herdr.readCount(), "a disabled gate must never read the lead's pane");
}
}
@@ -0,0 +1,142 @@
package dev.ltms.fleet.msg;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* A delegation whose message the pane collected itself must reach the same ticket states as one
* typed into the pane.
*/
class MessageServiceInboxDeliveryTest {
/** A pane that collects its own mail. */
private static final String MOD = "term_mod";
/** A pane that does not — the control for every route-specific assertion here. */
private static final String PTY = "term_pty";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final Injector injector = new Injector(agents);
private final InMemoryReplyInbox replyInbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox);
@BeforeEach
void setUp() {
replyInbox.own(MOD);
replyInbox.own(PTY);
}
/** The messages typed into a pane, in order. A collected message must never appear here. */
@SuppressWarnings("unchecked")
private List<String> typed() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
/** {@code sendAsync} queues on another thread, so wait for the delivery to exist. */
private void awaitWaiting(String target) throws InterruptedException {
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(target),
"the delegation to " + target + " should have opened its rendezvous waiter");
}
/** A ticket's terminal state is stamped by the delegating thread, so poll until it settles. */
private MessageService.TaskView awaitSettled(String ticket) throws InterruptedException {
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while ((view == null || view.phase() == MessageService.Phase.PENDING)
&& System.currentTimeMillis() < deadline) {
view = messages.poll(ticket);
//noinspection BusyWait
Thread.sleep(5);
}
assertNotNull(view, ticket + " must still be a known ticket");
return view;
}
@Test
void aTicketCollectedByThePaneReachesTheSameStatesAsATypedOne() throws Exception {
// The collecting pane announces itself before anything is delegated to it; the control
// pane never calls fleet_inbox at all.
assertEquals(List.of(), messages.collectInbox(MOD), "nothing is waiting yet");
String collectedTicket = messages.sendAsync(MOD, "long task");
awaitWaiting(MOD);
String typedTicket = messages.sendAsync(PTY, "long task");
awaitWaiting(PTY);
injector.onStatus(MOD, AgentStatus.IDLE); // offers the task for collection
injector.onStatus(PTY, AgentStatus.IDLE); // types the task into the pane
assertEquals(List.of("long task"), typed(),
"only the control pane was typed into; the collecting pane was not");
assertTrue(messages.poll(collectedTicket).detail().contains("queued, not yet delivered"),
"control: an offered-but-uncollected message has not reached its pane");
assertFalse(messages.poll(typedTicket).detail().contains("queued, not yet delivered"),
"control: a typed message has reached its pane");
assertEquals(List.of("long task"), messages.collectInbox(MOD), "the pane collects the task");
injector.onStatus(MOD, AgentStatus.IDLE); // records the delivery
assertFalse(messages.poll(collectedTicket).detail().contains("queued, not yet delivered"),
"a collected message is reported delivered, the same as a typed one");
injector.onStatus(MOD, AgentStatus.WORKING);
injector.onStatus(PTY, AgentStatus.WORKING);
// The reply still arrives through fleet_reply, and lands the same way on both routes.
MessageService.ReplyOutcome collectedReply = messages.reply(MOD, "task done");
MessageService.ReplyOutcome typedReply = messages.reply(PTY, "task done");
assertEquals(typedReply, collectedReply, "both routes resolve their delegation the same way");
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND, collectedReply);
MessageService.TaskView collectedDone = awaitSettled(collectedTicket);
MessageService.TaskView typedDone = awaitSettled(typedTicket);
assertEquals(MessageService.Phase.DONE, collectedDone.phase(),
"a ticket delivered by collection must not strand at PENDING");
assertEquals(MessageService.Phase.DONE, typedDone.phase(),
"control: the typed route reaches the same terminal state");
assertEquals("task done", collectedDone.reply());
assertEquals(typedDone.replySource(), collectedDone.replySource());
assertEquals(List.of("long task"), typed(),
"the collected delegation ran start to finish without typing into its pane");
}
@Test
void collectingAnInboxReachesOnlyTheCallersOwnMail() throws Exception {
assertEquals(List.of(), messages.collectInbox(MOD));
assertEquals(List.of(), messages.collectInbox(PTY));
messages.sendAsync(MOD, "task for the mod pane");
awaitWaiting(MOD);
messages.sendAsync(PTY, "task for the other pane");
awaitWaiting(PTY);
injector.onStatus(MOD, AgentStatus.IDLE);
injector.onStatus(PTY, AgentStatus.IDLE);
assertEquals(List.of("task for the mod pane"), messages.collectInbox(MOD));
// The control: the other pane's task really was waiting for it, so a collect that
// returned everyone's mail would have shown it on the line above.
assertEquals(List.of("task for the other pane"), messages.collectInbox(PTY));
}
}
@@ -107,6 +107,219 @@ class MessageServiceTest {
"an unreplied but finished turn resolves via the completion fallback");
assertEquals("BUILD GREEN: 391 files", reply.text(), "the scraped transcript tail is returned");
assertTrue(reply.completed(), "a scraped completion still counts as completed");
assertTrue(messages.drainReplies(T).isEmpty(),
"a completion resolved by a live caller must not also land in the inbox");
}
/** Poll {@code messages.drainReplies(target)} until it is non-empty or the deadline passes. */
private java.util.List<ReplyInbox.InboxMessage> awaitDrained(String target) throws InterruptedException {
var drained = messages.drainReplies(target);
long deadline = System.currentTimeMillis() + 2000;
while (drained.isEmpty() && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
drained = messages.drainReplies(target);
}
return drained;
}
/**
* Polls {@code target}'s inbox for {@code millis} and returns {@code true} only if nothing new
* ever arrived. Used to assert an absence against an asynchronous completion race: the window
* has to be long enough for a virtual thread already scheduled to actually run.
*/
private boolean noNewReplyWithin(String target, long millis) throws InterruptedException {
long deadline = System.currentTimeMillis() + millis;
while (System.currentTimeMillis() < deadline) {
if (!messages.drainReplies(target).isEmpty()) {
return false;
}
//noinspection BusyWait
Thread.sleep(5);
}
return true;
}
/**
* fleetd #801: {@code send}'s TimeoutException branch never completes {@code reply} itself, and
* {@link Rendezvous#close} only deregisters it from future session lookups — it does not cancel
* the future. The CB-106 completion fallback resolves the exact captured future (CB-116), so it
* can still succeed after the caller gave up and returned TIMED_OUT_WORKING. With nobody left
* awaiting that future, the late scrape must land in the target's inbox instead of being
* silently discarded.
*/
@Test
void aCompletionThatArrivesAfterTheCallerGaveUpLandsInTheInbox() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", 150);
awaitWaiting();
herdr.readText("$ prompt"); // pre-turn baseline
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts — no reply yet
MessageService.Reply timedOut = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, timedOut.outcome(),
"the caller gives up while the worker is still delivered and working");
// The worker's turn finishes AFTER the caller already gave up and nobody is listening.
herdr.readText("late answer, nobody was listening");
injector.onStatus(T, AgentStatus.IDLE); // working -> idle: the completion fallback fires
var drained = awaitDrained(T);
assertEquals(1, drained.size(), "the late completion must land in the inbox, not vanish");
assertEquals("late answer, nobody was listening", drained.get(0).content());
}
/**
* fleetd #808: once a send has timed out, an explicit {@code fleet_reply} for that same turn
* must land in the inbox exactly once — not twice, with the second copy a completion scrape of
* the same answer the reply already delivered.
*/
@Test
void anExplicitReplyAfterATimeoutLandsExactlyOnceNotTwice() throws Exception {
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", 150);
awaitWaiting();
herdr.readText("$ prompt");
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker starts — no reply yet
MessageService.Reply timedOut = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, timedOut.outcome(),
"the caller gives up while the worker is still delivered and working");
assertEquals(MessageService.ReplyOutcome.QUEUED, messages.reply(T, "structured answer"),
"no live waiter is open by now, so the explicit reply is held for a later drain");
var afterReply = awaitDrained(T);
assertEquals(1, afterReply.size(), "the explicit reply must land in the inbox exactly once");
assertEquals("structured answer", afterReply.get(0).content());
// The worker's turn then finishes for real; the completion fallback races the same waiter
// the reply already claimed, and must not publish a second entry for it.
herdr.readText("late scrape that must not also publish");
injector.onStatus(T, AgentStatus.IDLE);
assertTrue(noNewReplyWithin(T, 500),
"a completion scrape for a turn that already got an explicit fleet_reply must not "
+ "publish a second entry");
}
/**
* fleetd #808: the suppression above is per-turn, not per-target. A second, independent turn on
* the same member that never calls {@code fleet_reply} must still have its completion scrape
* land, even though an earlier turn on that same target already got an explicit reply.
*/
@Test
void theSuppressionIsPerTurnNotPerTarget() throws Exception {
// First turn: times out, then gets an explicit fleet_reply.
CompletableFuture<MessageService.Reply> first = sendAsync("first task", 150);
awaitWaiting();
herdr.readText("$ prompt 1");
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, first.get(5, TimeUnit.SECONDS).outcome());
assertEquals(MessageService.ReplyOutcome.QUEUED, messages.reply(T, "first answer"));
var firstDrain = awaitDrained(T);
assertEquals(1, firstDrain.size());
assertEquals("first answer", firstDrain.get(0).content());
// Finishing turn 1 for real also frees the worker for turn 2, and exercises that turn 1's
// own late scrape stays suppressed right up to that point.
herdr.readText("first turn's late scrape — must stay suppressed");
injector.onStatus(T, AgentStatus.IDLE);
assertTrue(noNewReplyWithin(T, 300), "turn 1's own scrape must still be suppressed");
// Second, independent turn on the same target: times out, and no fleet_reply is ever called.
CompletableFuture<MessageService.Reply> second = sendAsync("second task", 150);
awaitWaiting();
herdr.readText("$ prompt 2");
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, second.get(5, TimeUnit.SECONDS).outcome());
herdr.readText("second turn's late scrape — must still land");
injector.onStatus(T, AgentStatus.IDLE);
var secondDrain = awaitDrained(T);
assertEquals(1, secondDrain.size(),
"a later, independent turn's completion must still land even though an earlier "
+ "turn on the same target already got an explicit reply");
assertEquals("second turn's late scrape — must still land", secondDrain.get(0).content());
}
/**
* fleetd #801 defect 2: the publish inside {@code strandLateResolution}'s {@code whenComplete}
* callback runs with no caller left to surface a throw to — {@code AmqpReplyInbox.publish} can
* throw {@code IllegalStateException}, and an exception thrown inside a {@code whenComplete}
* action is captured into the discarded dependent stage, never rethrown. It must be caught and
* logged there, or a broker failure on this path is invisible and the late reply is just lost.
*/
@Test
void aLateCompletionThatFailsToPublishIsLoggedNotLost() throws Exception {
FakeHerdr localHerdr = new FakeHerdr();
Rendezvous localRendezvous = new Rendezvous();
java.util.concurrent.atomic.AtomicLong localClock = new java.util.concurrent.atomic.AtomicLong();
CompletionResolver localCompletion = new CompletionResolver(new AgentControl(localHerdr), localRendezvous,
ExhaustedPatternLookup.none(), ExhaustionSink.none(),
() -> localClock.addAndGet(CompletionResolver.MIN_TURN_NANOS + 1));
Injector localInjector = new Injector(new AgentControl(localHerdr), localCompletion);
ReplyInbox throwingInbox = new ReplyInbox() {
private final InMemoryReplyInbox delegate = new InMemoryReplyInbox();
@Override public void own(String target) { delegate.own(target); }
@Override public void release(String target) { delegate.release(target); }
@Override public void publish(String target, String msgId, String content) {
throw new IllegalStateException("PROBE-801-strandLateResolution");
}
@Override public java.util.List<InboxMessage> peek(String target) { return delegate.peek(target); }
@Override public boolean ack(String target, String msgId) { return delegate.ack(target, msgId); }
};
MessageService localMessages = new MessageService(new AgentControl(localHerdr), localInjector,
localRendezvous, throwingInbox);
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> localMessages.send(T, "do the task", 150, (String) null));
long deadline = System.currentTimeMillis() + 2000;
while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertTrue(localRendezvous.isWaiting(T), "send should have opened its rendezvous waiter");
localHerdr.readText("$ prompt");
localInjector.onStatus(T, AgentStatus.IDLE); // deliver
localInjector.onStatus(T, AgentStatus.WORKING); // worker starts — no reply yet
MessageService.Reply timedOut = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, timedOut.outcome(),
"the caller gives up while the worker is still delivered and working");
ListAppender<ILoggingEvent> appender = attachMessageServiceLog();
try {
// The worker's turn finishes AFTER the caller gave up — the completion fallback fires,
// and its publish into throwingInbox fails.
localHerdr.readText("late answer, publish fails");
localInjector.onStatus(T, AgentStatus.IDLE);
long logDeadline = System.currentTimeMillis() + 2000;
while (appender.list.isEmpty() && System.currentTimeMillis() < logDeadline) {
//noinspection BusyWait
Thread.sleep(5);
}
assertFalse(appender.list.isEmpty(), "a failed late-completion publish must be logged");
assertEquals(Level.WARN, appender.list.get(0).getLevel());
} finally {
detachMessageServiceLog(appender);
}
assertFalse(localMessages.hasStrandedReply(T),
"strandedReplies must not be set when the publish it depends on failed");
}
/**
@@ -1034,6 +1247,39 @@ class MessageServiceTest {
assertEquals("forbidden: this ticket was created by a different session", refused.detail());
}
/**
* {@link MessageService#ownsTicket(String, String)} is the side-effect-free ownership check
* the authorization gate consults ahead of {@link MessageService#poll(String, String)}, which
* also fires the ticket's collection hook. Read-only: polling the same ticket afterwards still
* sees it, which {@link #theOneArgPollOverloadBypassesOwnershipEntirely} nearby does not
* guarantee for every overload.
*/
@Test
void publicOwnsTicketIsASideEffectFreeOwnershipCheck() {
Principal lead = Principal.leader("opus", "term_lead", 1);
String ticket = messages.sendAsync(T, "long task", null, lead);
assertTrue(messages.ownsTicket(ticket, lead.ownerKey()));
assertFalse(messages.ownsTicket(ticket, Principal.anonymous().ownerKey()));
assertFalse(messages.ownsTicket("no-such-ticket", lead.ownerKey()),
"a ticket this daemon never heard of is owned by nobody");
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket, lead.ownerKey()).phase(),
"the ownership check must not have consumed or altered the ticket");
}
/**
* fleetd #803: the gap between "unknown ticket" (above) and "no ticket at all". A null ticket
* used to reach {@code ConcurrentHashMap.get(null)} and throw a {@code NullPointerException}
* instead of returning false.
*/
@Test
void publicOwnsTicketWithANullTicketReturnsFalseRatherThanThrowing() {
Principal lead = Principal.leader("opus", "term_lead", 1);
assertFalse(messages.ownsTicket(null, lead.ownerKey()),
"no ticket at all is owned by nobody, the same as an unknown one");
}
@Test
void architectOwnershipUsesTerminalRatherThanSlot() {
Principal oldArchitect = Principal.architect("opus", "term_OLD", 1);
@@ -1706,6 +1952,43 @@ class MessageServiceTest {
"the reply completed its own ticket directly and never touched the inbox");
}
/**
* A message still sitting in the injector's queue — never attempted, let alone delivered —
* must not poll as the worker working on it. {@code herdr.agentStatus("working")} supplies the
* target's real, independent live status (busy with something else entirely), while no
* {@code injector.onStatus} call is ever made, so the message is never even attempted.
*/
@Test
void aQueuedButNeverInjectedMessageDoesNotPollAsWorking() throws Exception {
herdr.agentStatus("working");
String ticket = messages.sendAsync(T, "task");
awaitWaiting();
MessageService.TaskView view = messages.poll(ticket);
assertEquals(MessageService.Phase.PENDING, view.phase());
assertFalse(view.detail().contains("worker working"),
"a message never injected must not read as the worker working on it: " + view.detail());
assertTrue(view.detail().contains("queued") && view.detail().contains("not yet delivered"),
"must report the message as queued, not delivered: " + view.detail());
}
/**
* Positive control for the test above: once the message is actually delivered, the poll's
* detail is the plain live-status text again.
*/
@Test
void aDeliveredMessageStillPollsAsWorkerWorking() throws Exception {
herdr.agentStatus("working");
String ticket = messages.sendAsync(T, "task");
awaitWaiting();
injectDelivery();
MessageService.TaskView view = messages.poll(ticket);
assertEquals(MessageService.Phase.PENDING, view.phase());
assertEquals("worker working", view.detail(),
"once actually delivered, the detail reports the worker's live status directly");
}
/**
* fleetd #329 (F1). {@code answer()} completes the async ticket by looking {@code turnId} up in
* {@code asyncTasksByTurn} a SECOND time (the first is at :991, purely to re-register the
@@ -2376,7 +2659,7 @@ class MessageServiceTest {
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs, java.util.function.LongSupplier nowNanos) {
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
FakeHerdr leadHerdr = new FakeHerdr();
FakeHerdr leadHerdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET);
AgentControl leadAgents = new AgentControl(leadHerdr);
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
@@ -2413,7 +2696,7 @@ class MessageServiceTest {
java.util.function.LongSupplier nowNanos) {
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
FakeHerdr leadHerdr = new FakeHerdr();
FakeHerdr leadHerdr = new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET);
AgentControl leadAgents = new AgentControl(leadHerdr);
ManualScheduler scheduler = new ManualScheduler();
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
@@ -2672,7 +2955,8 @@ class MessageServiceTest {
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
// A backoff far longer than the test: the schedule is started but no tick ever fires, so
// decide() is read directly and nothing here depends on timing.
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, new AgentControl(new FakeHerdr()), inbox,
ReplyPushLoop pushLoop = new ReplyPushLoop(registry,
new AgentControl(new FakeHerdr().detectionText(FakeHerdr.IDLE_PROMPT_CARET)), inbox,
scheduler, 5, 60_000);
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop);
try {
@@ -3,6 +3,7 @@ package dev.ltms.fleet.msg;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.mcp.PrimaryRegistry;
@@ -338,6 +339,85 @@ class ReplyPushLoopTest {
assertEquals(0, rec.promptTargets().size(), "nobody live was found, so nothing was ever sent");
}
// --- fleetd #778: never nudge a delegator to run a call its own role is refused -------------
@Test
void onReplyQueuedNeverNudgesADelegatorThatMayNotRunDrain() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
registry.recordDelegation(WORKER, OTHER_PRIMARY, "observer", false, false);
loop(1, 50).onReplyQueued(WORKER);
Thread.sleep(200);
assertEquals(0, rec.sendCount(),
"a delegator whose role cannot run fleet_poll(target=...) must never be nudged to run it");
}
@Test
void onReplyQueuedStillNudgesADelegatorThatMayRunDrain() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
registry.recordDelegation(WORKER, OTHER_PRIMARY, "lead", true, true);
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("fleet_poll(target=" + WORKER + ")"),
"a delegator that may run DRAIN must still be nudged about it: " + nudge);
}
@Test
void onQuestionOpenedNeverNudgesADelegatorThatMayNotRunAnswer() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
registry.recordDelegation(WORKER, OTHER_PRIMARY, "observer", false, false);
loop(1, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
Thread.sleep(200);
assertEquals(0, rec.sendCount(),
"a delegator whose role cannot run fleet_send(turnId=...) must never be nudged to run it");
}
@Test
void onQuestionOpenedStillNudgesADelegatorThatMayRunAnswer() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
registry.recordDelegation(WORKER, OTHER_PRIMARY, "architect", false, true);
loop(1, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("fleet_send(turnId="),
"an architect may run ANSWER, so the nudge must still name it: " + nudge);
}
/**
* A delegator barred from DRAIN can still have a ticket genuinely pending for it — tickets are
* out of this unit's scope (fleetd #778 Part 1 already makes a ticket nudge correct for an
* observer). The reply half must still be suppressed even though the ticket half fires.
*/
@Test
void aForbiddenReplyNudgeIsSuppressedWhileAnEligibleTicketStillFires() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
registry.recordDelegation(WORKER, OTHER_PRIMARY, "observer", false, false);
var loop = loop(1, 300); // wide backoff so both calls land before the first tick fires
loop.onReplyQueued(WORKER);
loop.onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
String nudge = rec.sentParams().getFirst().getValue().toString();
assertFalse(nudge.contains("fleet_poll(target=" + WORKER + ")"),
"the reply half is forbidden for this delegator and must never be named: " + nudge);
}
// --- nudge format --------------------------------------------------------------------------
@Test
@@ -1079,6 +1159,91 @@ class ReplyPushLoopTest {
"hitting the ticket reminder cap must count as exhausted");
}
// --- the operator's own prompt box ----------------------------------------------------------
@Test
void aLeadWithUnsubmittedTextInItsPromptBoxIsNotNudged() {
var herdr = new PaneTextHerdrClient(FakeHerdr.DRAFTED_PROMPT_CARET);
agents = new AgentControl(herdr);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(5, 100_000);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0),
"a nudge pastes and submits, so an idle lead mid-sentence must not be nudged");
assertEquals(0, herdr.promptCount(), "nothing reached the pane");
assertFalse(inbox.peek(WORKER).isEmpty(), "and the reply is still waiting to be collected");
}
@Test
void theSameLeadIsNudgedOnceItsPromptBoxIsEmpty() {
var herdr = new PaneTextHerdrClient(FakeHerdr.DRAFTED_PROMPT_CARET);
agents = new AgentControl(herdr);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(5, 100_000);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
herdr.paneText(FakeHerdr.IDLE_PROMPT_CARET);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
"the box emptied, so the held nudge is due");
}
@Test
void anUnrecognisablePaneHoldsTheNudge() {
var herdr = new PaneTextHerdrClient("garbled ansi noise with no input box");
agents = new AgentControl(herdr);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(5, 100_000);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0),
"a pane whose box cannot be found may be holding a draft");
assertEquals(0, herdr.promptCount());
}
@Test
void aFailedPaneReadHoldsTheNudge() {
var herdr = new PaneTextHerdrClient(FakeHerdr.IDLE_PROMPT_CARET).failReads();
agents = new AgentControl(herdr);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(5, 100_000);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0),
"an unreadable box is treated as a draft, never as an empty one");
assertEquals(0, herdr.promptCount());
}
@Test
void aNudgeHeldForADraftIsSentOnALaterTick() throws Exception {
var herdr = new PaneTextHerdrClient(FakeHerdr.DRAFTED_PROMPT_CARET);
agents = new AgentControl(herdr);
loop(5, 50).onTicketTerminal("task-1", WORKER, false);
Thread.sleep(300);
assertEquals(0, herdr.promptCount(), "every tick holds while the operator is typing");
herdr.paneText(FakeHerdr.IDLE_PROMPT_CARET);
assertTrue(herdr.sendLatch.await(3, TimeUnit.SECONDS),
"the nudge lands on the first tick after the box empties");
}
// --- the gate's on/off switch ----------------------------------------------------------------
@Test
void aDisabledGateNudgesALeadWithUnsubmittedTextInItsPromptBox() {
var herdr = new PaneTextHerdrClient(FakeHerdr.DRAFTED_PROMPT_CARET);
agents = new AgentControl(herdr);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(5, 100_000, false);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
"a disabled gate delivers even though the box holds a draft");
assertEquals(0, herdr.readCount(), "a disabled gate never reads the pane");
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
@@ -1093,10 +1258,24 @@ class ReplyPushLoopTest {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, metrics);
}
private ReplyPushLoop loop(int maxReminders, long backoffMs, boolean promptBoxGateEnabled) {
return new ReplyPushLoop(registry, agents, inbox, scheduler, maxReminders, backoffMs, null,
promptBoxGateEnabled);
}
private static AgentControl agentWithStatus(String status) {
return new AgentControl(new FakeHerdrClient(status));
}
/**
* The {@code agent.read} frame every fake here returns: a lead settled at an empty input box. The
* loop reads the box before it nudges, so a fake that answered nothing would read as a pane it
* cannot classify and hold every nudge.
*/
private static JsonNode emptyPromptBoxRead() {
return MAPPER.createObjectNode().set("read", MAPPER.createObjectNode().put("text", FakeHerdr.IDLE_PROMPT_CARET));
}
/** Non-recording (single-threaded) fake — safe for decide() tests. */
private static final class FakeHerdrClient implements HerdrClient {
private final String agentStatus;
@@ -1113,6 +1292,9 @@ class ReplyPushLoopTest {
.put("terminal_id", PRIMARY)
.put("agent_status", agentStatus));
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1142,6 +1324,9 @@ class ReplyPushLoopTest {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1179,6 +1364,9 @@ class ReplyPushLoopTest {
if ("agent.prompt".equals(method)) {
sendLatch.countDown();
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1216,6 +1404,9 @@ class ReplyPushLoopTest {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1266,6 +1457,9 @@ class ReplyPushLoopTest {
promptTargets.add(String.valueOf(p.get("target")));
sendLatch.countDown();
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1316,6 +1510,9 @@ class ReplyPushLoopTest {
if ("agent.prompt".equals(method)) {
promptTargets.add(String.valueOf(p.get("target")));
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1363,6 +1560,9 @@ class ReplyPushLoopTest {
promptTargets.add(String.valueOf(p.get("target")));
sendLatch.countDown();
}
if ("agent.read".equals(method)) {
return emptyPromptBoxRead();
}
return MAPPER.createObjectNode();
}
@@ -1374,4 +1574,66 @@ class ReplyPushLoopTest {
public void close() {
}
}
/**
* Thread-safe fake that always reports {@code idle} and serves a mutable pane tail, so a test can
* change what the lead's input box holds between ticks. Records every {@code agent.prompt}.
*/
private static final class PaneTextHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> prompts =
Collections.synchronizedList(new ArrayList<>());
private final AtomicInteger reads = new AtomicInteger();
private volatile String paneText;
private volatile boolean failReads = false;
volatile CountDownLatch sendLatch = new CountDownLatch(1);
PaneTextHerdrClient(String paneText) {
this.paneText = paneText;
}
void paneText(String text) {
this.paneText = text;
}
PaneTextHerdrClient failReads() {
this.failReads = true;
return this;
}
int promptCount() {
return prompts.size();
}
int readCount() {
return reads.get();
}
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", "idle"));
}
if ("agent.read".equals(method)) {
reads.incrementAndGet();
if (failReads) {
throw new HerdrException("herdr socket read timed out");
}
return MAPPER.createObjectNode()
.set("read", MAPPER.createObjectNode().put("text", paneText));
}
if ("agent.prompt".equals(method)) {
prompts.add(Map.entry(method, params));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
}
@@ -285,6 +285,67 @@ class FleetAppAuthTest {
}
}
/**
* An observer may read a ticket its own {@code fleet_send(wait:false)} created, over the REST
* route and not only the unit-level classifier — and still not a ticket a different caller
* created, even another observer pane's.
*/
@Test
void restPollLetsAnObserverReadItsOwnTicketButNotAnothersOverRest() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin observerApp = startOnSharedServiceAsObserver(messages, herdr, 9001L); // -> term_shell
Javalin otherObserverApp = startOnSharedServiceAsObserver(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
try {
Principal observer = Principal.observer("term_shell", 9001);
String ticket = messages.sendAsync("term_a", "long task", null, observer);
HttpResponse<String> own = send(observerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the observer that created the ticket must read it: " + own.body());
// Unlike a worker's or the primary's mismatched ticket (refused by MessageService's own
// ownership check, inside a 200 response), a different observer is refused at the
// authorization gate itself, since its grant is conditional on the classifier -- so
// this is a 403, never reaching poll().
HttpResponse<String> other = send(otherObserverApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(403, other.statusCode(), other.body());
assertFalse(other.body().contains("\"reply\""),
"a refusal must never carry reply text: " + other.body());
} finally {
observerApp.stop();
otherObserverApp.stop();
}
}
/**
* As {@link #startOnSharedService(MessageService, FakeHerdr, long)}, but resolves every
* connecting pane to the unconfigured-pane {@link Role#OBSERVER} floor instead of a spawned
* worker — no lead, collaborator, or member map claims it.
*/
private Javalin startOnSharedServiceAsObserver(MessageService messages, FakeHerdr herdr, long pid) {
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null), t -> null, Map::of);
Metrics appMetrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
return new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, appMetrics).build().start("127.0.0.1", 0);
}
/**
* {@code POST /sessions/{id}/message} with {@code wait:false} must record the creating
* caller's own terminal on the ticket it returns, so that caller can still poll its own
@@ -694,6 +755,27 @@ class FleetAppAuthTest {
assertEquals(403, toSpawnedMembersTerminal.statusCode(), toSpawnedMembersTerminal.body());
}
/**
* The REST route must give the same answer as MCP for an observer: pid 9001 resolves to
* "term_shell", a herdr pane recognised as no configured role, so the real
* {@link CallerResolver#observerSendTarget()} classifier reaches the known lead and refuses
* the known collaborator -- over the route, not just the unit-level classifier, so a grant
* covering only MCP cannot leave this one behind.
*/
@Test
void anObserverMaySendToAKnownLeadButNotToAKnownCollaboratorOverRest() throws Exception {
int port = startWithRealClassifier(9001L,
Map.of("term_lead_known", "lead-x"), Map.of("term_collab_known", "ops2"));
HttpResponse<String> toLead = send(port, "POST", "/sessions/term_lead_known/message",
"{\"content\":\"hi\",\"wait\":false}", null);
assertEquals(202, toLead.statusCode(), toLead.body());
HttpResponse<String> toCollaborator = send(port, "POST", "/sessions/term_collab_known/message",
"{\"content\":\"hi\",\"wait\":false}", null);
assertEquals(403, toCollaborator.statusCode(), toCollaborator.body());
}
/**
* As {@link #start}, but with explicit lead/collaborator maps and no spawned-member roster, so
* a test can wire the real {@link CallerResolver#knownLeadOrCollaborator()} classifier instead
@@ -724,6 +806,59 @@ class FleetAppAuthTest {
return app.port();
}
/**
* As {@link #startWithRealClassifier}, but with a {@link MemberRegistry} bound to an architect
* slot at {@code terminal} (the defense-in-depth path, no live roster) — lets a test drive a
* real architect {@link Principal} over HTTP.
*/
private int startWithArchitect(long pid, String terminal) {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(
Map.of(), Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null));
assertTrue(members.bind("architect:lead-designer", terminal));
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null, Map::of, members);
metrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
app = new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, metrics).build().start("127.0.0.1", 0);
return app.port();
}
/**
* fleetd #801 defect 1: {@code SEND} is granted to an architect but {@code DRAIN} is not, so a
* TIMED_OUT/BUSY receipt to an architect must not name {@code GET /sessions/{id}/replies} —
* that route refuses the same caller — and must say the reply is unrecoverable instead.
*/
@Test
void anArchitectsTimedOutSendDoesNotNameTheRepliesRouteItCannotDrain() throws Exception {
int port = startWithArchitect(FakeHerdr.WORKER_PID, "term_a"); // FakeHerdr resolves this pid to term_a
HttpResponse<String> res = send(port, "POST", "/sessions/term_b/message",
"{\"content\":\"hi\",\"timeoutMs\":150}", null);
assertEquals(202, res.statusCode(), "an architect is granted SEND: " + res.body());
JsonNode body = new ObjectMapper().readTree(res.body());
String detail = body.get("detail").asText();
assertFalse(detail.contains("/replies"), "got: " + detail);
assertTrue(detail.contains("cannot be recovered"), "got: " + detail);
assertEquals(403, send(port, "GET", "/sessions/term_b/replies", null, null).statusCode(),
"the architect this receipt was sent to is indeed refused DRAIN");
}
// --- loopback-trust: the caller is the primary -------------------------------------------
@Test
@@ -663,8 +663,10 @@ class FleetAppTest {
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
assertEquals(202, res.statusCode());
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
JsonNode body = mapper.readTree(res.body());
assertEquals("queued", body.get("status").asText());
assertFalse(herdr.called("agent.prompt"), "no injection while the worker is mid-turn");
assertDetailNamesTheInboxRouteNotARetry(body);
}
@Test
@@ -674,8 +676,21 @@ class FleetAppTest {
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
assertEquals(202, res.statusCode());
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
JsonNode body = mapper.readTree(res.body());
assertEquals("working", body.get("status").asText());
assertTrue(herdr.called("agent.prompt"), "message was injected");
assertDetailNamesTheInboxRouteNotARetry(body);
}
/**
* fleetd #801: the TIMED_OUT_WORKING / TIMED_OUT_QUEUED / BUSY receipt is never tracked by a
* ticket, so it must name the one route that recovers a late reply — draining the session's
* inbox — and must not invite a bare resend, which can duplicate the delivery.
*/
private void assertDetailNamesTheInboxRouteNotARetry(JsonNode body) {
String detail = body.get("detail").asText();
assertTrue(detail.contains("/sessions/term_a/replies"), "got: " + detail);
assertFalse(detail.contains("retry"), "got: " + detail);
}
/**
+2 -2
View File
@@ -1,7 +1,7 @@
{
"name": "fleet",
"description": "Make a project fleet-ready: mount the fleetd MCP gateway and set up standard Claude Code settings so this session can orchestrate a fleet of delegated workers. Lead-side only — member skills and agents travel in the worktree. Ships no credentials.",
"version": "0.2.0",
"description": "Set up standard Claude Code settings so this session can orchestrate a fleet of delegated workers, and run the fleet mod for cross-session messaging. Lead-side only — member skills and agents travel in the worktree, and mounting the fleetd MCP gateway is now the instance's or the project's job, not this plugin's. Ships no credentials.",
"version": "0.3.0",
"author": {
"name": "LTMS"
},
-8
View File
@@ -1,8 +0,0 @@
{
"mcpServers": {
"fleet": {
"type": "http",
"url": "${FLEETD_MCP_URL}"
}
}
}
+35 -19
View File
@@ -1,7 +1,7 @@
# fleet (Claude Code plugin)
Makes a project **fleet-ready**: mounts the `fleetd` MCP gateway and applies standard Claude Code
settings, so the session can orchestrate a fleet of delegated workers.
Makes a project **fleet-ready**: applies standard Claude Code settings and runs the fleet mod, so
the session can orchestrate a fleet of delegated workers.
**This plugin ships no credentials.** Every secret is referenced by environment-variable *name*;
the values stay with the user. Nothing the plugin writes is unsafe to commit.
@@ -22,9 +22,10 @@ Member-facing assets travel in the worktree, not in this plugin. See fleetd #362
## What it is not
The plugin is the **client-side setup**, not the bridge. `fleetd` is a separate daemon and `herdr`
is a separate PTY multiplexer, each with its own lifecycle and install. The plugin mounts an
already-running daemon and tells you what is missing when one isn't there — it deliberately does
not try to install system services on your behalf.
is a separate PTY multiplexer, each with its own lifecycle and install, and the plugin does not try
to install either on your behalf. It also does not mount the daemon for you — mounting is the
instance's or the project's own `.mcp.json`, and `/fleet:setup` is the one thing in this plugin that
still helps with that (it writes the project-level entry).
## Install
@@ -33,11 +34,20 @@ not try to install system services on your behalf.
/plugin install fleet@fleetd
```
Export the gateway URL — the plugin mounts `${FLEETD_MCP_URL}`, not a hardcoded address, so one
plugin serves hosts that run the daemon on different ports:
Mount the daemon yourself — this plugin carries no mount of its own. Either add the entry below to
your Claude Code instance's own `.claude.json`, so every project you open there gets it, or run
`/fleet:setup` in the project you want to onboard, which writes the same entry into that project's
`.mcp.json`:
```shell
export FLEETD_MCP_URL=http://127.0.0.1:8765/mcp
```json
{
"mcpServers": {
"fleet": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
```
Then, in the project you want to onboard:
@@ -50,18 +60,24 @@ Then, in the project you want to onboard:
| Component | Effect |
|---|---|
| `.mcp.json` | mounts `fleet` at `${FLEETD_MCP_URL}` for any session with the plugin enabled |
| `skills/setup` | `/fleet:setup` — preflight, project settings, credential guidance, and verification |
| `skills/setup` | `/fleet:setup` — preflight, project settings, credential guidance, and verification. Also the only thing in this plugin that helps mount `fleet`: it writes the project `.mcp.json` entry shown above. |
| `hooks/register.js` (the fleet mod) | Cross-account session messaging while the plugin is enabled: `/fleet-peers`, `/fleet-mail`, `/fleet-whoami`, and a background poll that delivers mail fleetd queued for this pane. A spawned worker or architect skips that poll, because it already gets its brief pasted into its pane. |
The server is named **`fleet`** on purpose: that is `PeerLauncher.MCP_MOUNT_NAME` in the daemon and
the name a spawned member's own mount carries. Version 0.1.0 named it `fleetd`, which produced two
mounts of one daemon for anyone who also had a project-level `.mcp.json`. Upgrading from 0.1.0 is a
**breaking change** — a project that pre-allowed `mcp__fleetd__fleet_whoami` in
`.claude/settings.json` must be updated to `mcp__fleet__*`.
Whichever file mounts the daemon, name the server **`fleet`**. That is `PeerLauncher.MCP_MOUNT_NAME`
in the daemon, the name a spawned member's own mount carries, and the name the `mcp__fleet__*`
role heuristic in `CLAUDE.md` keys on.
Because the plugin carries its own `.mcp.json`, an installed plugin needs no project-level MCP
file at all. The setup skill writes one only when you want the mount to work *without* the plugin —
for teammates who haven't installed it, or for CI.
## Upgrading from 0.2.0 — breaking
The plugin no longer mounts the daemon. It used to carry its own `.mcp.json`, pointed at
`${FLEETD_MCP_URL}`, and that file is gone. The mod still reads `FLEETD_MCP_URL`, but only as an
optional override of the address it calls (default `http://127.0.0.1:8765/mcp`). Mount `fleet`
yourself: add the entry under **Install** above to your instance's `.claude.json` or to the
project's own `.mcp.json`, by hand or with `/fleet:setup`.
Version 0.1.0 named the mounted server `fleetd`, which produced two mounts of one daemon for
anyone who also had a project-level `.mcp.json`. A project that pre-allowed
`mcp__fleetd__fleet_whoami` in `.claude/settings.json` must be updated to `mcp__fleet__*`.
## Verifying a setup
+4
View File
@@ -0,0 +1,4 @@
{
"description": "fleet mod: cross-account session messaging and PTY-free delivery",
"modules": ["./register.js"]
}
+281
View File
@@ -0,0 +1,281 @@
// fleet mod — session messaging for the claude-bridge fleet.
//
// $.store is kept under CLAUDE_CONFIG_DIR, so /fleet-peers and /fleet-mail reach
// only sessions that share this session's config dir. fleetd names a caller by
// its pane, not its account, so /fleet-whoami and anything else sent through
// fleetTool reach the fleet from either account.
//
// Delivery uses $.prompt.submit, so nothing is typed into a pane and no prompt
// box is read.
//
// There are two inboxes. The $.store one carries /fleet-mail between sessions
// that share this config dir. fleet_inbox carries what the fleet queued for
// this pane, from any account, and calling it is also what tells fleetd to
// queue here rather than type into the terminal.
const PRESENCE_PREFIX = 'presence:'
const INBOX_PREFIX = 'inbox:'
const PRESENCE_REFRESH_MS = 15_000
const INBOX_POLL_MS = 3_000
// fleetd stops queueing for a pane that goes quiet, so this must stay well
// under the window the daemon allows between calls.
const FLEETD_INBOX_POLL_MS = 3_000
// A session whose presence row is older than this is treated as gone. It must
// exceed PRESENCE_REFRESH_MS by enough that one missed refresh is not a death.
const PRESENCE_STALE_MS = 60_000
const MAX_INBOX = 50
/** The store key holding one session's queued messages. */
function inboxKey(sessionId) {
return INBOX_PREFIX + sessionId
}
/** The store key holding one session's presence row. */
function presenceKey(sessionId) {
return PRESENCE_PREFIX + sessionId
}
/**
* Append one message to a target's inbox.
*
* $.store has no compare-and-swap, so two senders writing in the same instant
* can lose a message. Callers that need delivery confirmed should read the
* inbox back.
*/
async function deliver($, target, message) {
const key = inboxKey(target)
const queued = (await $.store.get(key)) || []
queued.push(message)
// Keep the newest: an unread inbox must not grow without bound.
const kept = queued.slice(-MAX_INBOX)
await $.store.set(key, kept)
return kept.length
}
/** Every session that refreshed its presence row recently, newest first. */
async function livePeers($, now) {
const keys = await $.store.keys()
const rows = []
for (const key of keys) {
if (!key.startsWith(PRESENCE_PREFIX)) continue
const row = await $.store.get(key)
if (!row || typeof row.at !== 'number') continue
if (now - row.at > PRESENCE_STALE_MS) continue
rows.push(row)
}
rows.sort((a, b) => b.at - a.at)
return rows
}
const FLEETD_MCP_DEFAULT = 'http://127.0.0.1:8765/mcp'
const MCP_HEADERS = { 'Content-Type': 'application/json', Accept: 'application/json, text/event-stream' }
// The status fleetd answers, with "Session not found", for an Mcp-Session-Id it no longer holds.
const MCP_SESSION_GONE = 404
// The MCP session every call below shares, and the URL it was opened against. fleetd keeps a
// server-side session per initialize and drops it only on a DELETE, so one initialize per call
// would leave a session behind every time.
let mcpSessionId = null
let fleetdUrl = FLEETD_MCP_DEFAULT
/**
* Open an MCP session on the local fleetd and hold it for later calls.
*
* Cleared first, so a failure here leaves no dead id behind for the next call to reuse. Reads
* FLEETD_MCP_URL fresh on every open, so a session opened after the daemon moves uses the new
* address.
*/
async function openFleetSession($) {
mcpSessionId = null
fleetdUrl = (await $.env.get('FLEETD_MCP_URL')) || FLEETD_MCP_DEFAULT
const init = await $.http.fetch(fleetdUrl, {
method: 'POST',
headers: MCP_HEADERS,
body: JSON.stringify({
jsonrpc: '2.0', id: 1, method: 'initialize',
params: { protocolVersion: '2025-06-18', capabilities: {}, clientInfo: { name: 'fleet-mod', version: '0' } },
}),
})
if (!init.ok) throw new Error('fleetd initialize failed with status ' + init.status)
const opened = init.headers['mcp-session-id']
await $.http.fetch(fleetdUrl, {
method: 'POST',
headers: { ...MCP_HEADERS, 'Mcp-Session-Id': opened },
body: JSON.stringify({ jsonrpc: '2.0', method: 'notifications/initialized' }),
})
mcpSessionId = opened
}
/** Send one tools/call on the session this mod holds, and return the raw HTTP answer. */
function sendFleetToolCall($, tool, args) {
return $.http.fetch(fleetdUrl, {
method: 'POST',
headers: { ...MCP_HEADERS, 'Mcp-Session-Id': mcpSessionId },
body: JSON.stringify({ jsonrpc: '2.0', id: 2, method: 'tools/call', params: { name: tool, arguments: args } }),
})
}
/**
* Call one fleet_* tool on the local fleetd, and return its text result.
*
* fleetd names the caller from the TCP connection on every request, not from the MCP session, so
* the answer is about this session's own pane whichever Claude account the session runs on, and
* reusing one session never changes whose call it is.
*/
async function fleetTool($, tool, args) {
if (mcpSessionId === null) await openFleetSession($)
let call = await sendFleetToolCall($, tool, args)
if (call.status === MCP_SESSION_GONE) {
// A daemon restart drops every session it held. Open a new one and retry once.
await openFleetSession($)
call = await sendFleetToolCall($, tool, args)
}
if (!call.ok) throw new Error('fleetd ' + tool + ' failed with status ' + call.status)
// A tool call answers as a server-sent event: the JSON is on the data: line.
const line = call.text.split('\n').find((l) => l.startsWith('data:'))
const body = JSON.parse(line ? line.slice(5) : call.text)
if (body.error) throw new Error(body.error.message)
return body.result.content.map((c) => c.text).join('\n')
}
export function register(on) {
on('session.start', async ($, e, next) => {
const self = await $.session.id()
// Announce before the first refresh is due, or a session shorter than one
// refresh interval never appears to its peers at all.
await $.store.set(presenceKey(self), {
sessionId: self,
cwd: await $.session.cwd(),
at: await $.clock.now(),
})
// Keep announcing: a row that stops being refreshed is how another session
// learns this one is gone.
$.clock.every(PRESENCE_REFRESH_MS, async () => {
const now = await $.clock.now()
await $.store.set(presenceKey(self), {
sessionId: self,
cwd: await $.session.cwd(),
at: now,
})
})
// Collect this session's mail and hand it to Claude. $.prompt.submit waits
// for the session to be idle, so this never lands mid-turn.
$.clock.every(INBOX_POLL_MS, async () => {
const key = inboxKey(self)
const queued = (await $.store.get(key)) || []
if (queued.length === 0) return
await $.store.set(key, [])
for (const message of queued) {
await $.prompt.submit({
text: 'Message from fleet session ' + message.from + ':\n\n' + message.text,
})
}
})
// A spawned worker or architect already gets its brief pasted into its pane, so this
// poll would be a second, redundant delivery path for it. Every other role collects
// its own mail through this poll.
let role = null
// Collect what the fleet queued for this pane and hand each message to
// Claude. Every call also renews fleetd's record that this pane collects its
// own mail, so an empty answer still has to be asked for.
$.clock.every(FLEETD_INBOX_POLL_MS, async () => {
if (role === null) {
try {
role = JSON.parse(await fleetTool($, 'fleet_whoami', {})).role
} catch {
// A daemon that is down, or a pane fleetd cannot place, is the ordinary
// case on a host with no fleet running. Retry on the next tick.
return
}
}
if (role === 'worker' || role === 'architect') return
let collected
try {
collected = JSON.parse(await fleetTool($, 'fleet_inbox', {}))
} catch {
// The timer survives a throw, so this only keeps every tick from
// writing an error to the debug log.
return
}
for (const text of collected.messages || []) {
await $.prompt.submit({ text: 'Message from the fleet, via fleetd:\n\n' + text })
}
})
await $.command.register({
name: 'fleet-peers',
description: 'List fleet sessions on this machine, including other accounts',
})
await $.command.register({
name: 'fleet-whoami',
description: 'Show who fleetd says this session is',
})
await $.command.register({
name: 'fleet-mail',
description: 'Send a message to a fleet session on this machine',
argumentHint: '<sessionId> <text>',
// Runs even while Claude is working, so a correction is never queued
// behind the turn it is meant to correct.
immediate: true,
})
return next(e)
})
on('command.run', { command: 'fleet-peers' }, async ($) => {
const now = await $.clock.now()
const self = await $.session.id()
const peers = await livePeers($, now)
if (peers.length === 0) return { text: 'No fleet sessions have announced themselves yet.' }
const lines = peers.map((p) => {
const age = Math.round((now - p.at) / 1000)
const mark = p.sessionId === self ? ' (this session)' : ''
return p.sessionId + ' ' + p.cwd + ' seen ' + age + 's ago' + mark
})
return { text: 'Fleet sessions on this machine:\n' + lines.join('\n') }
})
on('command.run', { command: 'fleet-mail' }, async ($, e) => {
const args = (e.args || '').trim()
const split = args.indexOf(' ')
if (split < 1) return { text: 'Usage: /fleet-mail <sessionId> <text>' }
const target = args.slice(0, split)
const text = args.slice(split + 1).trim()
if (text === '') return { text: 'Usage: /fleet-mail <sessionId> <text>' }
const now = await $.clock.now()
const peers = await livePeers($, now)
if (!peers.some((p) => p.sessionId === target)) {
return { text: 'No live fleet session ' + target + '. Run /fleet-peers.' }
}
const self = await $.session.id()
const depth = await deliver($, target, { from: self, text: text, at: now })
return { text: 'Queued for ' + target + ' (' + depth + ' in its inbox).' }
})
on('command.run', { command: 'fleet-whoami' }, async ($) => {
try {
return { text: 'fleetd says: ' + (await fleetTool($, 'fleet_whoami', {})) }
} catch (err) {
return { text: 'fleetd unreachable: ' + err.message }
}
})
// Record what arrives over Claude Code's own channel, so a message delivered
// by the fleet and one delivered by SendMessage can be told apart.
//
// This hook gates delivery, so it must never decide the message's fate. The
// .catch handler passes the message on when the logging above throws.
on('session.receive', async ($, e, next) => {
$.ui.log('fleet: inbound ' + (e.origin && e.origin.kind) + ', ' + String(e.text).length + ' chars')
return next(e)
}).catch(async ($, e, next) => {
if (next.called) return undefined
return next(e)
})
}
+6 -12
View File
@@ -36,16 +36,11 @@ a time.
command -v herdr && herdr --version 2>&1 | head -1 || echo "MISSING: herdr"
command -v ccs && ccs version 2>&1 | head -1 || echo "MISSING: ccs (needed for worker profiles)"
command -v codex && codex --version 2>&1 | head -1 || echo "absent: codex (optional)"
curl -s -m 5 "${FLEETD_MCP_URL%/mcp}/healthz" 2>/dev/null \
|| curl -s -m 5 http://127.0.0.1:8765/healthz \
|| echo "MISSING: fleetd daemon is not reachable"
[ -n "$FLEETD_MCP_URL" ] && echo "FLEETD_MCP_URL is set" || echo "MISSING: FLEETD_MCP_URL"
curl -s -m 5 http://127.0.0.1:8765/healthz || echo "MISSING: fleetd daemon is not reachable"
```
**`FLEETD_MCP_URL` is required.** The plugin's own `.mcp.json` mounts `${FLEETD_MCP_URL}` rather
than a hardcoded address, so one plugin can serve hosts that run the daemon on different ports. If
it is unset the mount does not resolve. The usual value is `http://127.0.0.1:8765/mcp`; tell the
user to export it, do not write it into a file for them.
The usual address is `http://127.0.0.1:8765`. If this daemon runs elsewhere, use that address
instead wherever this skill writes `http://127.0.0.1:8765/mcp` below.
A healthy daemon answers with its status **and the herdr protocol it negotiated**:
@@ -95,10 +90,9 @@ If `.mcp.json` already exists, add only the `fleet` key and leave every other se
If a `fleet` entry is already there with a different URL, **ask** rather than assuming yours is
right — a non-default port usually means a deliberate second daemon.
> **If this plugin is installed, you can skip this step entirely.** The plugin ships its own
> `.mcp.json`, so `fleetd` is already mounted for any session with the plugin enabled. Write the
> project-level file only when the user wants the mount to work *without* the plugin — for
> teammates who have not installed it, or for CI.
This write is the only way this plugin helps mount `fleet` — the plugin carries no mount of its
own. A session that wants the mount without running this skill can instead add the same entry to
its own Claude Code instance's `.claude.json`.
**Before writing it, settle whether `.mcp.json` is committed here:**
+490
View File
@@ -0,0 +1,490 @@
import { expect, mock, test } from 'claude-code/testing'
// The store-backed tests write the row another session would write, because a
// test cannot start a second session.
const PEER = 'peer-session'
const SELF = 'this-session'
// A fixed clock keeps the staleness arithmetic exact.
const NOW = 1_700_000_000_000
/**
* Answer the mods API calls the harness has no implementation for, and stand in
* for Claude Code's own behaviour beneath the mod's gating hooks.
*
* Returns the map backing $.store. The harness puts no `store` namespace on the
* test's own `$`, so a test seeds and inspects the carrier through this map.
*/
function stubEngine(on: any): Map<string, any> {
const store = new Map<string, any>()
on('store.get', (_$: any, e: any) => ({ value: store.get(e.key) }))
on('store.set', (_$: any, e: any) => {
store.set(e.key, e.value)
return { value: undefined }
})
on('store.delete', (_$: any, e: any) => {
store.delete(e.key)
return { value: undefined }
})
on('store.keys', () => ({ value: [...store.keys()] }))
on('clock.now', () => ({ value: NOW }))
on('session.id', () => ({ value: SELF }))
on('session.cwd', () => ({ value: '/Users/x/claude-bridge' }))
on('ui.log', () => ({ value: undefined }))
on('prompt.submit', () => ({ value: undefined }))
on('env.get', () => ({ value: undefined }))
// Claude Code's own delivery, which the mod's receive hook must reach.
on('session.receive', (_$: any, e: any) => e)
return store
}
/** The presence row a session on the other account would write. */
function announce(store: Map<string, any>, sessionId: string, cwd: string, at: number) {
store.set('presence:' + sessionId, { sessionId, cwd, at })
}
test('/fleet-peers lists a session that announced itself', async ($, on) => {
const store = stubEngine(on)
announce(store, PEER, '/Users/x/work-repo', NOW)
const answer = await $.command.run({ command: 'fleet-peers', args: '' })
expect(answer.text).toContain(PEER)
expect(answer.text).toContain('/Users/x/work-repo')
})
test('/fleet-peers hides a session whose presence row went stale', async ($, on) => {
const store = stubEngine(on)
// One second past the 60s staleness cut.
announce(store, PEER, '/Users/x/work-repo', NOW - 61_000)
const answer = await $.command.run({ command: 'fleet-peers', args: '' })
expect(answer.text).not.toContain(PEER)
})
test('/fleet-peers keeps a row one second inside the staleness cut', async ($, on) => {
// The positive control for the test above: without this, a bug that hid
// every row would still satisfy that assertion.
const store = stubEngine(on)
announce(store, PEER, '/Users/x/work-repo', NOW - 59_000)
const answer = await $.command.run({ command: 'fleet-peers', args: '' })
expect(answer.text).toContain(PEER)
})
test('/fleet-mail queues a message in the target session inbox', async ($, on) => {
const store = stubEngine(on)
announce(store, PEER, '/Users/x/work-repo', NOW)
const answer = await $.command.run({
command: 'fleet-mail',
args: PEER + ' rebasing on main is safe now',
})
expect(answer.text).toContain('Queued for ' + PEER)
const inbox = store.get('inbox:' + PEER)
expect(inbox.length).toBe(1)
expect(inbox[0].text).toBe('rebasing on main is safe now')
expect(inbox[0].from).toBe(SELF)
})
test('a second message appends rather than replacing the first', async ($, on) => {
const store = stubEngine(on)
announce(store, PEER, '/Users/x/work-repo', NOW)
await $.command.run({ command: 'fleet-mail', args: PEER + ' first' })
await $.command.run({ command: 'fleet-mail', args: PEER + ' second' })
const inbox = store.get('inbox:' + PEER)
expect(inbox.length).toBe(2)
expect(inbox[0].text).toBe('first')
expect(inbox[1].text).toBe('second')
})
test('/fleet-mail refuses a target that never announced itself', async ($, on) => {
const store = stubEngine(on)
const answer = await $.command.run({ command: 'fleet-mail', args: 'ghost-session hello' })
expect(answer.text).toContain('No live fleet session ghost-session')
// Nothing may be queued for a session we could not confirm.
expect(store.get('inbox:ghost-session')).toBe(undefined)
})
test('/fleet-mail rejects input with no message text', async ($, on) => {
const store = stubEngine(on)
announce(store, PEER, '/Users/x/work-repo', NOW)
for (const args of ['', PEER, PEER + ' ']) {
const answer = await $.command.run({ command: 'fleet-mail', args })
expect(answer.text).toContain('Usage: /fleet-mail')
}
expect(store.get('inbox:' + PEER)).toBe(undefined)
})
test('an inbound peer message is passed on, not consumed', async ($, on) => {
// A receive hook that withheld a message would silently break Claude Code's
// own channel, so assert this mod stays transparent.
stubEngine(on)
const result = await $.session.receive({
text: 'from the other session',
origin: { kind: 'peer' },
})
expect(result?.consumed).toBe(undefined)
})
/** Answer fleetd's three MCP requests the way the live daemon does. */
function stubFleetd(on: any, toolText: string, seen: any[]) {
on('http.fetch', (_$: any, e: any) => {
const body = JSON.parse(e.init.body)
seen.push({ method: body.method, session: e.init.headers['Mcp-Session-Id'] })
if (body.method === 'initialize') {
return { value: { ok: true, status: 200, headers: { 'mcp-session-id': 'sid-1' }, text: '{}' } }
}
if (body.method === 'tools/call') {
const result = { jsonrpc: '2.0', id: 2, result: { content: [{ type: 'text', text: toolText }] } }
return { value: { ok: true, status: 200, headers: {}, text: 'event: message\ndata: ' + JSON.stringify(result) + '\n' } }
}
return { value: { ok: true, status: 202, headers: {}, text: '' } }
})
}
test('/fleet-whoami reads the tool result out of the event stream', async ($, on) => {
stubEngine(on)
const seen: any[] = []
stubFleetd(on, '{"role":"observer","sessionId":"term_x"}', seen)
const answer = await $.command.run({ command: 'fleet-whoami', args: '' })
expect(answer.text).toBe('fleetd says: {"role":"observer","sessionId":"term_x"}')
// The session id from initialize must ride on every later request.
expect(seen.map((r) => r.method)).toEqual(['initialize', 'notifications/initialized', 'tools/call'])
expect(seen[2].session).toBe('sid-1')
})
test('/fleet-whoami reports a daemon that is down instead of throwing', async ($, on) => {
stubEngine(on)
on('http.fetch', () => ({ value: { ok: false, status: 503, headers: {}, text: '' } }))
const answer = await $.command.run({ command: 'fleet-whoami', args: '' })
expect(answer.text).toBe('fleetd unreachable: fleetd initialize failed with status 503')
})
/**
* The session.start hook's own calls, for a test that fires it. Kept apart from stubEngine
* because a timer test drives $.clock through mock.clock(on) instead of a fixed clock.now.
*/
function stubSessionStart(on: any, submitted: string[]): Map<string, any> {
const store = new Map<string, any>()
on('store.get', (_$: any, e: any) => ({ value: store.get(e.key) }))
on('store.set', (_$: any, e: any) => {
store.set(e.key, e.value)
return { value: undefined }
})
on('store.keys', () => ({ value: [...store.keys()] }))
on('session.id', () => ({ value: SELF }))
on('session.cwd', () => ({ value: '/Users/x/claude-bridge' }))
on('command.register', () => ({ value: undefined }))
on('ui.log', () => ({ value: undefined }))
on('env.get', () => ({ value: undefined }))
// The engine skips a prompt.submit hook that answers anything but { text } or { drop }, and
// the mod's callback then throws, so this must hand the text straight back.
on('prompt.submit', (_$: any, e: any) => {
submitted.push(e.text)
return { text: e.text }
})
on('session.start', (_$: any, e: any) => e)
return store
}
/**
* Answer fleetd's MCP requests, with the tool result read fresh on every call.
*
* `fetches` collects every method sent, with a tools/call entry naming its tool (such as
* `tools/call:fleet_inbox`), and `sessions` the Mcp-Session-Id of each tools/call. Each
* initialize hands out the next id, so a reused session and a reopened one differ.
* `sessionGone` makes a tools/call answer the way fleetd answers for a session id it no
* longer holds. `whoamiText` answers a `fleet_whoami` call apart from `toolText`, which
* answers every other tool.
*/
function stubFleetdDynamic(
on: any,
toolText: () => string,
fetches: string[],
sessions: string[] = [],
sessionGone: () => boolean = () => false,
whoamiText: () => string = () => '{"role":"primary"}',
) {
let opened = 0
on('http.fetch', (_$: any, e: any) => {
const body = JSON.parse(e.init.body)
if (body.method === 'initialize') {
fetches.push(body.method)
opened += 1
return { value: { ok: true, status: 200, headers: { 'mcp-session-id': 'sid-' + opened }, text: '{}' } }
}
if (body.method === 'tools/call') {
fetches.push(body.method + ':' + body.params.name)
sessions.push(e.init.headers['Mcp-Session-Id'])
if (sessionGone()) {
const gone = '{"jsonRpcError":{"code":-32603,"message":"Session not found"}}'
return { value: { ok: false, status: 404, headers: {}, text: gone } }
}
const text = body.params.name === 'fleet_whoami' ? whoamiText() : toolText()
const result = { jsonrpc: '2.0', id: 2, result: { content: [{ type: 'text', text }] } }
return { value: { ok: true, status: 200, headers: {}, text: 'data: ' + JSON.stringify(result) + '\n' } }
}
fetches.push(body.method)
return { value: { ok: true, status: 202, headers: {}, text: '' } }
})
}
/** How many of `fetches` were an initialize. */
function initializes(fetches: string[]): number {
return fetches.filter((m) => m === 'initialize').length
}
test('the fleetd inbox poll submits each collected message and names the sender', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
let inbox = { sessionId: 'term_self', count: 0, messages: [] as string[] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), [])
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
// An empty inbox is the ordinary answer and must submit nothing.
await clock.advance(3_000)
expect(submitted.length).toBe(0)
// The control for the line above: the same timer, the same stubs, with mail waiting.
inbox = { sessionId: 'term_self', count: 2, messages: ['rebase is safe now', 'build is green'] }
await clock.advance(3_000)
expect(submitted.length).toBe(2)
expect(submitted[0]).toContain('rebase is safe now')
expect(submitted[0]).toContain('fleetd')
expect(submitted[1]).toContain('build is green')
})
test('a collected message is not submitted a second time', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
// fleetd removes a message when it hands it over, so the next poll answers empty.
let inbox = { sessionId: 'term_self', count: 1, messages: ['do the task'] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), [])
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
expect(submitted.length).toBe(1) // control: the first poll really did deliver it
inbox = { sessionId: 'term_self', count: 0, messages: [] }
await clock.advance(3_000)
expect(submitted.length).toBe(1)
})
test('a fleetd that is down leaves the poll timer running', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
on('http.fetch', (_$: any, e: any) => {
fetches.push(JSON.parse(e.init.body).method)
return { value: { ok: false, status: 503, headers: {}, text: '' } }
})
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
const afterFirst = fetches.length
expect(afterFirst).toBeGreaterThan(0)
expect(submitted.length).toBe(0)
// A throw out of the timer would stop it, so a second tick that still reaches fleetd is the
// positive control for the assertion above: the timer survived the failure.
await clock.advance(3_000)
expect(fetches.length).toBeGreaterThan(afterFirst)
expect(submitted.length).toBe(0)
})
test('a second inbox poll reuses the first MCP session', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const sessions: string[] = []
const inbox = { sessionId: 'term_self', count: 0, messages: [] as string[] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, sessions)
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
await clock.advance(3_000)
// The first poll also reads the role once, so it makes two tools/call (whoami, then
// inbox); the second poll already knows the role and makes only one (inbox).
expect(sessions.length).toBe(3) // control: all three calls really reached fleetd
expect(initializes(fetches)).toBe(1)
expect(sessions.every((s) => s === sessions[0])).toBe(true)
})
test('a session fleetd no longer holds is opened again and the call retried', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const sessions: string[] = []
let gone = false
const inbox = { sessionId: 'term_self', count: 1, messages: ['the daemon restarted'] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, sessions, () => {
// Only the first call after the flag is set is refused; the retry succeeds.
const refuse = gone
gone = false
return refuse
})
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
// The control: an ordinary poll opens one session and needs no second one.
await clock.advance(3_000)
expect(initializes(fetches)).toBe(1)
expect(submitted.length).toBe(1)
gone = true
await clock.advance(3_000)
expect(initializes(fetches)).toBe(2)
expect(sessions[sessions.length - 1]).toBe('sid-2')
expect(submitted.length).toBe(2)
})
/** How many of `fetches` were a tools/call for the named tool. */
function toolCalls(fetches: string[], tool: string): number {
return fetches.filter((m) => m === 'tools/call:' + tool).length
}
test('a worker role stops the poll from calling fleet_inbox', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const inbox = { sessionId: 'term_self', count: 1, messages: ['do the task'] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, [], () => false, () => '{"role":"worker"}')
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_whoami')).toBe(1)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(0)
expect(submitted.length).toBe(0)
})
test('an architect role stops the poll from calling fleet_inbox', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const inbox = { sessionId: 'term_self', count: 1, messages: ['do the task'] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, [], () => false, () => '{"role":"architect"}')
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_whoami')).toBe(1)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(0)
expect(submitted.length).toBe(0)
})
test('a primary role keeps the poll calling fleet_inbox', async ($, on) => {
// The positive control for the two tests above: the same stubs, a role neither
// gates, so a bug that silenced every role would still pass them.
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const inbox = { sessionId: 'term_self', count: 1, messages: ['do the task'] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, [], () => false, () => '{"role":"primary"}')
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(1)
expect(submitted.length).toBe(1)
})
test('an observer role keeps the poll calling fleet_inbox', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const inbox = { sessionId: 'term_self', count: 1, messages: ['do the task'] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, [], () => false, () => '{"role":"observer"}')
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(1)
expect(submitted.length).toBe(1)
})
test('a whoami that fails on the first tick is retried and the poll resumes', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const inbox = { sessionId: 'term_self', count: 1, messages: ['do the task'] }
let whoamiFails = true
on('http.fetch', (_$: any, e: any) => {
const body = JSON.parse(e.init.body)
if (body.method === 'tools/call' && body.params.name === 'fleet_whoami' && whoamiFails) {
fetches.push(body.method + ':' + body.params.name)
return { value: { ok: false, status: 503, headers: {}, text: '' } }
}
if (body.method === 'initialize') {
fetches.push(body.method)
return { value: { ok: true, status: 200, headers: { 'mcp-session-id': 'sid-1' }, text: '{}' } }
}
if (body.method === 'tools/call') {
fetches.push(body.method + ':' + body.params.name)
const text = body.params.name === 'fleet_whoami' ? '{"role":"observer"}' : JSON.stringify(inbox)
const result = { jsonrpc: '2.0', id: 2, result: { content: [{ type: 'text', text }] } }
return { value: { ok: true, status: 200, headers: {}, text: 'data: ' + JSON.stringify(result) + '\n' } }
}
fetches.push(body.method)
return { value: { ok: true, status: 202, headers: {}, text: '' } }
})
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(0) // control: a failed whoami calls no fleet_inbox
expect(submitted.length).toBe(0)
whoamiFails = false
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(1)
expect(submitted.length).toBe(1)
})
test('whoami is called once, not on every tick, once the role is known', async ($, on) => {
const clock = mock.clock(on)
const submitted: string[] = []
stubSessionStart(on, submitted)
const fetches: string[] = []
const inbox = { sessionId: 'term_self', count: 0, messages: [] as string[] }
stubFleetdDynamic(on, () => JSON.stringify(inbox), fetches, [], () => false, () => '{"role":"primary"}')
await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/Users/x/claude-bridge' })
await clock.advance(3_000)
await clock.advance(3_000)
await clock.advance(3_000)
expect(toolCalls(fetches, 'fleet_whoami')).toBe(1)
expect(toolCalls(fetches, 'fleet_inbox')).toBe(3)
})
+3 -2
View File
@@ -1507,6 +1507,7 @@ else
grep -E ' (ERROR|SEVERE) ' "$FRESH_LOG" | tail -5 | sed 's/^/ /'
fi
echo
echo " Next: call fleet_whoami and confirm it still answers 'primary'. A lead whose tab label"
echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration."
echo " Next: call fleet_whoami and confirm it still answers 'primary'. A lead is found by its"
echo " tab being labelled 'lead' AND sitting in fleet.leaders.<name>.workspace; if either stops"
echo " matching, the lead is demoted to worker and refuses orchestration."
echo