Compare commits

...

57 Commits

Author SHA1 Message Date
Dai Ha 9d653e86df fleetd #726: name RELAUNCH_NEVER_READY in the IN_PROGRESS terminal-state list
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 2m7s
The javadoc on RollState.IN_PROGRESS enumerates the terminal states the entry
can be overwritten with, and omitted RELAUNCH_NEVER_READY. That state is
reachable at LeadRollover.java:710, so the list told a reader a state could not
occur when it can. Found by a reviewer on PR #742, outside its assigned scope.

Comment only; no behaviour change.
2026-10-04 21:44:03 +02:00
Dai Ha f6d1131d7a Merge remote-tracking branch 'origin/worker/726-unit2-75cb13-4' 2026-10-04 21:43:34 +02:00
Dai Ha a0505dc614 fleetd #726 unit 2: scope the member-daemon assertion to the roll itself
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m0s
The assembled daemon's own boot-time orphan-worker reap makes a real call on
the member fake before any roll starts. Clear the member fake's recorded
calls once assembly finishes and before the roll begins, so the assertion
measures calls made since the roll started rather than the whole process's
lifetime, and reword its message to say so.
2026-10-04 21:24:49 +02:00
Dai Ha d2f30f1654 fleetd #726 unit 2: cover the bootstrapText relaunch send with a dedicated test
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m46s
LeadRollover already sends every call through the lead-bound AgentControl and
WorkspaceControl it receives at construction, so no production code needed a
routing fix. Add a regression test that drives a full relaunch to the point
where recognition times out and asserts bootstrapText still lands on the lead
daemon and never on the member daemon, the one path the existing assembly test
never reaches.
2026-10-04 21:17:15 +02:00
Dai Ha cd1f04cbb4 Merge remote-tracking branch 'origin/worker/737-owner-key-ff061f-10'
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m48s
2026-10-04 21:11:20 +02:00
Dai Ha 73137f198f Merge remote-tracking branch 'origin/main' into worker/726-unit2-75cb13-4 2026-10-04 20:55:29 +02:00
Dai Ha 4ffe49f3bb fleetd #726 unit 2: replace /clear-based lead rollover with a real process restart
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Failing after 1m56s
LeadRollover's deferred continuation now ends the old lead's pane, relaunches
a fresh one, and bootstraps it, instead of sending /clear into the same
process. The relaunch step runs two separate bounded waits instead of one
combined check: a readiness wait (the fresh pane reaches a real turn
boundary) is the safety gate and withholds bootstrapText on timeout
(RELAUNCH_NEVER_READY); a recognition wait (the fresh terminal shows up in
the live-lead map) is bookkeeping only, so a timeout there still lets
bootstrapText go out (RELAUNCH_NOT_RECOGNISED). clearSettleSeconds is
retired in favor of relaunchReadySeconds (default 45), which bounds both
waits. Updates FleetConfig/FleetMcp operator-facing text to match.
2026-10-04 20:44:42 +02:00
Dai Ha efd9cdb983 fleetd #737 units 1+2: key tickets and turns on a stable owner, not a terminal
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m4s
Add Principal.ownerKey(): role-prefixed, keyed on name for a named lead and
a collaborator (survives a handover's terminal change), on terminal for a
worker, architect and observer, and on a distinct "anonymous" value for an
unauthenticated caller so that case no longer relies on Authz refusing it
first. The unnamed primary keeps a null key, preserving its primary-wide
ticket rule.

Thread that key through Task.creatorOwner, Rendezvous.Owner, poll,
pendingAsk and answer in place of a raw terminal, in both the MCP and REST
surfaces, so a named lead whose terminal changes can still poll and answer
its own delegations while a different lead is refused both.

Mutation evidence (each one-line change killed a named test, then reverted
to green):
- Principal.ownerKey() PRIMARY case made unconditional (dropped the
  null-name guard) -> ownerKeyCoversEveryRole dies:
  "expected: <null> but was: <leader:null>"
- OBSERVER case changed to use the "worker" prefix -> ownerKeyCoversEveryRole
  dies: "expected: <observer:term_observer> but was: <worker:term_observer>"
- prefixed() changed to drop the role prefix entirely -> both
  ownerKeyCoversEveryRole and rolePrefixesKeepLeadAndArchitectKeysDistinct
  die on a lead/architect key collision: "expected: <leader:opus> but was:
  <opus>"
- PRIMARY case changed to key on terminal instead of name ->
  rolePrefixesKeepLeadAndArchitectKeysDistinct and ownerKeyCoversEveryRole
  die: "expected: <leader:opus> but was: <leader:term_lead>"
- ARCHITECT case changed to key on the slot name instead of terminal ->
  same two tests die: "expected: <architect:opus> but was: <architect:design>"

All five mutations were caught by the existing test suite; no test needed
adding.
2026-10-04 20:42:44 +02:00
Dai Ha aabecce901 Merge remote-tracking branch 'origin/worker/736-presence-forget-f35144-9'
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m58s
2026-10-04 20:19:52 +02:00
Dai Ha 6754b4edbc fleetd #736: release clears the member's presence entry
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 1m45s
SessionManager.releaseRemoved() tore down a member's registry row and pane
but never cleared it from MemberPresence, so a terminal stayed marked
"present" for the daemon's lifetime after release/idle-reap/shutdown drain.
Clear it in the method's unconditional finally block, alongside the other
must-always-run teardown step, so every release path (explicit release,
the idle reaper's releaseIfCurrent, and a shutdown drain) forgets it the
same way, and a throw from the dirty-worktree check does not skip it.

MemberPresence.forget(null) throws NullPointerException (verified empirically:
ConcurrentHashMap.remove(null) NPEs on key.hashCode()), so the new call guards
on a non-null, non-blank terminal id rather than relying on forget to no-op.
2026-10-04 20:15:34 +02:00
Dai Ha 787ae0ed7a fleetd #726: tell the handover skill which rollover behaviour is live
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m59s
Unit 2 replaces the /clear continuation with a real process restart, so every
paragraph in the skill that describes /clear goes false the moment the new jar
is deployed. The code is not merged yet, and a merge is not a deployment, so
rewriting those paragraphs now would hand a lead doing a handover tonight a
document that does not match the daemon it is talking to.

Add a dated note instead. It states that the /clear text stays accurate while
the old jar runs, and gives a test a lead can apply with no shell: the
fleet_handover tool description is served by the running daemon, so if it still
says "clear your pane", the old behaviour is live. It also names the two things
that change, including the one that doubles as a second indicator -- the
"never observed as WORKING after 8 consecutive IDLE/DONE polls" warning cannot
appear once the wait that logs it is deleted. The note names the condition for
deleting itself.

Re-measure the roll evidence while here. The skill recorded four
"lead-rollover: rolled" lines from 2026-09-22; the log now holds 20, against a
control of 86 "lead-rollover:" lines, and "Unknown command" still returns 0.
Add the elapsed spread (median 16507 ms, max 48261 ms, two above 45000 ms) with
the caveat that it times the whole roll and is dominated by the wait for the
calling turn to end, so a slow roll is not a failed one.

Markdown only, no code touched, so no build was run.
2026-10-04 20:09:58 +02:00
Dai Ha e3050efe8b fleetd #705: correct the stale reason on the TASK_READ gate
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m48s
The comment said ticket ids are a sequential counter with no owner check, so a
holder could walk every ticket and read another session's reply. PRs #712 and
#716 added that owner check: MessageService.ownsTicket compares a ticket's
creatorTerminal to the caller on every read.

The rule is still right, so only the reason changes. This matters now because
fleetd #737 is deciding ticket ownership across a lead handover, and a reader
who believed the old text could delete the TASK_READ restriction on the grounds
that its stated reason no longer applies.

Comment-only. mvn -o clean install: Tests run: 2083, Failures: 0, Errors: 0,
BUILD SUCCESS. Flagged by the #705 option-1 worker as out of its scope, which
was the right call.
2026-10-04 19:59:45 +02:00
Dai Ha 11998cd626 Merge remote-tracking branch 'origin/worker/705-observer-14c258-6'
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 56s
CI / build (push) Failing after 1m54s
2026-10-04 19:52:34 +02:00
Dai Ha 8e5394f63f fleetd #705 option 1: narrow the unconfigured-pane floor to OBSERVER
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m57s
Adds Role.OBSERVER as the bottom rung CallerResolver falls to when a
herdr pane matches no live roster entry, lead, architect slot, or
collaborator tab. An observer may only READ/METRICS and REPLY/ASK on
its own pane. Widens the presence gate so an observer's MCP contact
still marks it deliverable, matching what already happens for a
worker or architect, so a pane that outlives a daemon restart is not
left permanently undeliverable.

Ships as defence in depth alongside the already-merged ticket-owner
check (#712/#716), which closed the reachable exploit this ticket
reported.
2026-10-04 19:46:54 +02:00
Dai Ha 428a12af62 fleetd #722: reconcile presence that arrives before a session's registry entry
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m48s
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m46s
A member whose MCP contact lands between launcher.spawn() and registry.put()
had its presence marked, but the SPAWNING -> READY transition that markPresent
triggers found no registry entry yet and silently did nothing. The mark then
persisted while registration left the session in SPAWNING, with nothing to
retry the transition. That left the session undeliverable to reclaim/seat
accounting even though it was present and deliverable.

Add SessionManager.reconcilePresence, called right after registry.put in both
the plain-spawn and worktree-spawn paths, to retry the transition for a
terminal already marked present. One private helper serves both call sites.

Tests cover both orderings (contact-then-register and register-then-contact)
for both spawn paths, plus a terminal never marked present staying in
SPAWNING. The contact-then-register tests use a new PresenceRacingLauncher
test double that marks presence from inside spawn(), before acquire()'s own
registry.put runs.
2026-10-04 19:23:02 +02:00
Dai Ha a332dfdb2c Merge remote-tracking branch 'origin/worker/726-ea34a0-2'
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m53s
2026-10-04 19:09:03 +02:00
Dai Ha d0f4ae057b fleetd #726 unit 3 fix: release the single-flight claim when continuationRunner rejects the hand-off
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 1m51s
confirm() takes the per-lead-terminal claim before handing the roll to
continuationRunner, and the only release path was runRollover's own
finally. If continuationRunner.accept itself throws, runRollover never
starts, so that finally never runs, and nothing else ever writes
rollingByTerminal — the claim is held forever and the terminal can never
be rolled again. This differs from fleetd #615, which covers a throw
INSIDE the continuation (runRollover already catches that and still
releases the claim) — this is a throw from the hand-off itself, which
is not reachable with today's virtual-thread runner but would be with
a bounded executor's RejectedExecutionException.

confirm() now catches that throw, releases the claim, and overwrites
the IN_PROGRESS outcome with a terminal FAILED one, matching how a
throw inside the continuation is already surfaced.
2026-10-04 19:06:57 +02:00
Dai Ha 7b3beaa209 Merge remote-tracking branch 'origin/worker/726-10cbf0-1'
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m53s
2026-10-04 19:06:47 +02:00
Dai Ha b3b2bf3da6 fleetd #726 unit 1 review fixes: correct relaunch's javadoc and dedupe its resolve logic
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 1m49s
relaunch's javadoc said it reads the live config; it actually reads the
FleetConfig snapshot this launcher was constructed with (fleet.leaders
is the frozen half), so say that and note a profile/tab edit needs a
daemon restart.

relaunch's recognise-only refusal reused ensureLeads()'s log wording,
which claims the lead 'is not live' — true in ensureLeads()'s context
(reached only after a short live count), false in relaunch's (which
never counts liveness, by design). Dropped that clause.

Pulled the declared/creatable/profile-configured resolution shared by
ensureLeads() and relaunch() into one private resolveLaunchable(name)
helper (returns a new ResolvedLead(lead, profile) record, or null
having logged), so the three refusals and their wording live in one
place instead of two copies that can drift. Behaviour-preserving:
ensureLeads() keeps its own liveness-count logic around the shared
resolve, and the existing 37 LeadLauncherTest cases are unchanged and
still pass.
2026-10-04 18:59:56 +02:00
Dai Ha 8bb2aa0be4 Merge remote-tracking branch 'origin/worker/729-5961c6-3'
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 2m0s
2026-10-04 18:59:54 +02:00
Dai Ha 4ce3149bfd fleetd #726 unit 3: make LeadRollover.confirm() single-flight per lead terminal
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m48s
Two open() calls for the same lead terminal minted two tokens that both
passed confirm()'s ownership check, so both could reach the deferred
continuation and roll the same lead twice. confirm() now claims a
per-lead-terminal slot (an atomic put-if-absent) once every other gate has
passed, refusing a concurrent confirm with the new ROLL_ALREADY_RUNNING
reason; runRollover releases the claim in a finally, on both the success
and the thrown-exception path.
2026-10-04 18:53:53 +02:00
Dai Ha 1fc9e85bf1 fleetd #729: fold a per-boot nonce into every turnId
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 1m57s
askSeq restarts at 0 on every daemon boot, so a turnId (session#n)
minted by one Rendezvous instance could be minted again by a later
instance and resolve to an unrelated ask. Fold a per-instance nonce
into the mint, the same way #719 fixed MessageService's ticket ids.
2026-10-04 18:53:05 +02:00
Dai Ha 38544d467c fleetd #726 unit 1: give LeadLauncher a public single-lead relaunch seam
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m48s
Adds LeadLauncher.relaunch(name), which starts exactly the named lead
from the live config, outside of ensureLeads()'s instances bookkeeping.
It retries the whole launch attempt (not just the agent_name_taken/
agent_pane_busy cases ResilientAgentLaunch already retries inside one
agents.start call) up to RELAUNCH_ATTEMPTS times.

launch() now returns the started Agent (null on failure) instead of a
boolean, so relaunch() and ensureLeads() share the same primitive.
2026-10-04 18:52:38 +02:00
Dai Ha 809b7d9b20 Merge remote-tracking branch 'origin/worker/727-ee14ed-3'
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m56s
2026-10-04 18:29:28 +02:00
Dai Ha 8cf7215d56 Merge remote-tracking branch 'origin/worker/719-bdd95e-4'
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m58s
2026-10-04 18:19:55 +02:00
Dai Ha 1ef93e57cc fleetd #727: give a lead launch the three protections every member spawn gets
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m59s
Extract checkPaneCommandFits, the agent_pane_busy retry, and the agent_name_taken
retry out of HerdrPeerLauncher into a shared dev.ltms.fleet.herdr.ResilientAgentLaunch,
and route LeadLauncher.launch through the same seam instead of a bare agents.start
call. The lead's agent name now carries a per-process nonce and a per-start sequence
number (like a member's), so a stale agent_name_taken from an earlier crashed session
no longer blocks a legitimate relaunch outright.
2026-10-04 18:14:06 +02:00
Dai Ha cf0c9b9316 fleetd #719: make the foreign-id test reach a colliding sequence number
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m49s
other.sendAsync had never been called, so other's tasks map was empty and
poll(ticket) returned null regardless of whether the nonce existed — the
test passed against an empty map, not against a colliding id. Mint once on
other so it reaches the same sequence number as the first instance, making
the test exercise the actual collision the nonce guards against.
2026-10-04 18:14:01 +02:00
Dai Ha 337dbd491e fleetd #719: fold a per-boot nonce into every ticket id
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 2m0s
ticketSeq restarted at zero on every daemon boot with no persistence, so a
ticket id minted in one boot could be reused by a later boot and resolve to
an unrelated Task instead of failing to resolve at all. Mint each
MessageService instance's own short nonce once and fold it into every ticket
(task-<nonce>-<n>), so an id from one instance can never match another's id
space.

Adds a disjoint-id-space test and a foreign-instance-ticket test (with the
positive control) in MessageServiceTest.
2026-10-04 18:05:06 +02:00
Dai Ha 7cf6075b79 Correct the configDir note: the count was wrong, the shared dir is the defect
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 44s
CI / build (push) Failing after 1m57s
The previous commit compared 4 configDir lines against 8 total profiles, but
four of those are opencode and never read CLAUDE_CONFIG_DIR. Every claude-code
profile does set one, so the original claim was right and this file said
otherwise.

The real defect is narrower: opus and sonnet name the operator's own config dir,
so for those members the store is shared, and ClaudeCodeLauncher's javadoc
already records that fleetd and the operator's session write that same file.
2026-10-04 10:35:33 +02:00
Dai Ha fbdcd709c9 The plugin addendum claimed every profile sets configDir; it does not
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 42s
CI / build (push) Failing after 2m1s
GET /profiles reports 8 live profiles and fleetd.yaml carries 4 configDir
lines, two of them pointing at the operator's own instance dir. The bullet
stated the blanket claim as a structural limit, so it would have been believed.
The conclusion it supported is unchanged: member-facing assets travel in the
worktree.
2026-10-04 10:31:07 +02:00
Dai Ha 8d3f10d291 Releasing is reached directly by its test, and its javadoc names the real case
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 59s
CI / build (push) Failing after 2m0s
The javadoc said a losing releaseIfCurrent CAS is the call that arrives with no
known session. releaseIfCurrent is only called by the reaper, always with a
non-null expected, so it always has one; the null-known call is an overlapping
release that finds the registry entry already gone. The same wrong claim was in
a test's failure message.

Releasing, enter and leave drop private so SessionManagerTest binds them at
compile time. The six reflection helpers are gone, and a rename now breaks the
build instead of a test run.
2026-10-04 10:25:17 +02:00
Dai Ha 38f4fd64ee Merge PR #724: fleetd #702 — a pane mid-teardown keeps its member identity 2026-10-04 10:25:08 +02:00
Dai Ha 3fc39b981d Only the delegation's creator can answer its member's question
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 56s
CI / build (push) Failing after 1m39s
fleetd #715 gates fleet_send{turnId} on the caller that created the delegation,
so the shipped block had to say so: the step-5 note already covered who can see
a pending question, not who can answer it.
2026-10-04 10:15:24 +02:00
Dai Ha 70328ca0f8 Merge PR #725: fleetd #715 — gate fleet_send{turnId} on the caller that created the delegation
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m47s
2026-10-04 10:12:38 +02:00
Dai Ha efab9b8c49 fleetd #702: pin the architect-demotion window and split a thin test
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 2m8s
Add a CallerResolverTest case proving a releasing architect resolves as
WORKER while still inside releaseRemoved's teardown window, with a
control resolve outside the window that must stay ARCHITECT so the
in-window assertion cannot pass against a slot that was never bound.

Split spawnedMemberRoleSurvivesAnOverlappingReleaseThatUnmarksEarly into
two SessionManagerTest cases, each reaching Releasing.enter/leave
directly through reflection so a mutation to one invariant (the depth
count in leave, or enter's prior-terminal preservation) can only fail
its own test.
2026-10-04 10:04:00 +02:00
Dai Ha 2374de28e4 fleetd #715: pin that no production caller uses Rendezvous.open(String)
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Failing after 2m4s
Add a source-scrape test mirroring #718's MessageServicePollUsageTest
shape, for the same residual: a convenience overload that defaults the
turn's owner to null, left in place because deleting it would break 81
test-only call sites across 7 unrelated files.

The scanner is exercised against a file known to hold many real
one-argument rendezvous.open( calls before it is ever pointed at
production, using the identical matching logic for both. A pattern that
cannot find the known calls would also find none in production, and
that is exactly the failure mode a -based git grep regex hit earlier
on this ticket: git grep's -E engine does not treat \b as a word
boundary, so that pattern silently matched nothing anywhere, in clean
code and in the 81 real calls alike.
2026-10-04 09:56:05 +02:00
Dai Ha dd18bd1f38 fleetd #715: gate answer() on the caller that owns the turn
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 2m15s
Record the turn's owner on the forward rendezvous waiter (Rendezvous.Owner,
a three-state record: no record / unnamed primary / named terminal). A
fresh fleet_ask copies that owner onto the ask turn; a coalesced duplicate
ask keeps the first owner. answer() compares the answering caller against
the stored owner before taking the session lock or reopening the resumed
waiter, and a mismatch returns the new NOT_TURN_OWNER outcome instead of
STALE_TURN, with no rendezvous/task/question cleanup.

send() and answer() both drop their no-caller overloads; every call site
in MessageService, FleetMcp and FleetApp now threads an explicit caller
terminal through. AuthzTest pins that the unnamed-primary null allowance
is safe only because ANONYMOUS never reaches ANSWER.

Covers all four combinations of capture input x adapter: blocking-send and
async-send delegations, answered through both FleetMcp and FleetApp, each
with a hijack attempt refused and the real owner's answer succeeding as a
control.
2026-10-04 09:46:20 +02:00
Dai Ha efeffb4ab7 fleetd #702: mark a pane as mid-teardown so a caller still resolves as a member
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m51s
A caller from a pane being torn down used to be briefly absent from the
session registry, so CallerResolver fell through to a lead/architect tab map
for the same terminal and could resolve the wrong role.

SessionManager now keeps a depth-counted "releasing" marker per pane, written
before the registry removal and cleared in a finally once release finishes.
A depth count (not a Set) is needed because two threads can race teardown of
the same pane; a Set-based unmark by the losing thread would reopen the
window while the winning thread is still mid-teardown. Both removal sites
(the unconditional release() and the idle reaper's CAS releaseIfCurrent())
go through one shared helper, so all four entry routes (release, the
context-cap release in completeTurn, reapIdle, drainSnapshot) are covered.

SessionManager.spawnedMemberRole is the one reader: it checks the live
registry first (via the existing no-copy findByTerminal), then the releasing
marker. FleetdAssembly now wires this method reference instead of its own
untested inline lambda, which also drops a roster() list copy + stream from
the per-request hot path. CallerResolver is unchanged — its contract already
fit.
2026-10-04 09:27:09 +02:00
Dai Ha ed4f4b08ad callerTerminal's javadoc names which callers carry a terminal
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m59s
It said "null if the primary", which reads as every primary. Only the unnamed
primary has no terminal; a named lead carries one, so a reader deciding what a
null means was given the wrong rule.
2026-10-04 09:11:08 +02:00
Dai Ha 28a1f3d6f5 fleet_status's pending-ask block is shown only to the delegation's creator
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m54s
A lead polling a member no longer sees a question raised under an architect's
brief, so an absent question is not evidence there is none.
2026-10-04 08:57:14 +02:00
Dai Ha b12d70716b Merge PR #723: fleetd #721 — gate fleet_status's pending-ask block by the delegation's creator
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m41s
2026-10-04 08:53:42 +02:00
Dai Ha 0f5985b419 fleetd #721: gate fleet_status's pending-ask block by the delegation's creator terminal
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m55s
fleet_status (MCP) and GET /sessions/{id}/status (REST) handed any TASK_READ
holder another session's open fleet_ask question, its turnId and its ticket,
with no check that the caller created that delegation. MessageService.pendingAsk
now takes the caller's terminal and reuses the existing ownsTicket comparison;
FleetMcp.status and FleetApp.sessionStatus both thread the resolved caller
terminal through. The base status line and REST's ready field are unaffected.
2026-10-04 08:46:38 +02:00
Dai Ha 6e06058b07 Merge PR #720: fleetd #718 — pin that no production caller uses MessageService.poll(String)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m44s
2026-10-04 08:23:58 +02:00
Dai Ha e5f4fb81ab fleetd #718: pin the naming convention the poll-usage scan depends on
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Failing after 2m22s
MessageServicePollUsageTest's messages.poll( receiver anchor only
covers a MessageService reached through a variable, field, or
parameter named "messages". Adds a second assertion in the same
class that every such declaration under src/main/java uses that
name, with its own file-walk and declaration-count controls, so a
future declaration under a different name turns this check red
instead of leaving the original scan silently blind to it.

Also tightens the Authz.permits(Principal, Action, String) javadoc
sentence to read as a plain contract statement.
2026-10-04 08:18:54 +02:00
Dai Ha d28ab0968b fleetd #718: pin that no production caller uses MessageService.poll(String)
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m47s
Adds a source-scrape test over src/main/java that fails if any caller
reaches the fail-open single-argument poll(String) overload instead of
poll(String, String). The scan anchors on the "messages.poll(" receiver
to avoid matching java.util.Queue.poll(), and balances parentheses to
avoid being fooled by a two-argument call whose first argument contains
nested parens.

Also documents Authz.permits(Principal, Action, String) as a test
convenience whose default classifier denies every collaborator.
2026-10-04 08:09:33 +02:00
Dai Ha 9a64d42599 Merge PR #716: fleetd #705 — gate the REST ticket routes by the creating caller
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m57s
Three parts. GET /tasks/{ticket} passed no caller, so it used the overload that skips the ownership
check and any worker could read any ticket; it now passes the caller resolved from the same CALLER
attribute the authorization gate reads. The wait:false send path recorded no creator terminal, so a
REST-created ticket matched no terminal-bearing caller and its own creator was refused; it now
records one. Both handlers gained a scrape guard with its own control assertion.

Resolved one conflict in FleetMcpAuthzTest by keeping both sides: PR #717 and this branch each
appended tests at the same point. The test count is the check on that resolution — 2014 + 3 + 1 =
2018, so no test was dropped.

Verified in a throwaway worktree off main: Tests run: 2018, Failures: 0, BUILD SUCCESS. Three
mutations, each confirmed live with mvn -o compile before the suite ran. FleetMcp:527 is the one
that survived before this work and now kills theFleetPollHandlerActuallyThreadsCallerTerminalIntoPoll.
FleetApp:698 kills the new creator test. FleetApp:899 kills three, including the behavioural test.
Every file restored byte-identical.
2026-10-04 07:48:18 +02:00
Dai Ha f4e0ca41e6 Merge PR #717: fleetd #710 — gate fleet_list's leads and members arrays by caller role
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m52s
A worker now gets neither array, and the key is absent rather than present-and-empty. leads stays
visible to the primary, an architect and a collaborator; members only to the primary and an
architect. That split follows the rule the collaborators array already states: you may list what
you could address. A collaborator may send to a lead, and leads is the only place the bridge gives
it that address, so hiding it would have left a shipped grant unusable.

Verified in a throwaway worktree off main: Tests run: 2014, Failures: 0, BUILD SUCCESS, against a
main baseline of 2008 that I measured myself. Dropping the collaborator clause from leadsVisibleTo
compiled green and then killed exactly two tests, the truth-table row and the behavioural test.
2026-10-04 07:41:21 +02:00
Dai Ha 29a2f97c25 fleetd: thread the creating caller's terminal into the REST fire-and-poll send path
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m53s
sendMessage's wait:false branch now records the resolved caller's own
terminal as the ticket's creatorTerminal, the same way taskStatus already
resolves its caller, so a REST-created ticket's own creator can still poll
it under the ownership check that now gates GET /tasks/{ticket}.
2026-10-04 07:41:12 +02:00
Dai Ha d2db8c7dc9 fleetd #710 PR #717 review: leadsVisibleTo must also admit a collaborator
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m47s
A collaborator may SEND to a lead, and fleet_list's leads array is the
only place this tool gives it a lead's sessionId -- its own
fleet_whoami carries no lead address. leadsVisibleTo now returns true
for caller.isCollaborator() as well as primary and architect, matching
the existing rule for the sibling collaborators array (every role that
may SEND to a named peer). membersVisibleTo is unchanged: a
collaborator may never SEND to a spawned member.

Updated the truth-table tests for both predicates, added a behavioural
test proving a collaborator's fleet_list output contains leads and not
members, and updated fleet_list's tool description.
2026-10-04 07:37:55 +02:00
Dai Ha 95311c6e8e Correct two stale claims in the canonical bridge block
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m5s
CI / build (push) Failing after 2m4s
The block is the instruction surface this repo ships, so a merged change that makes it false is an
incomplete change. Two had gone stale:

- the lead table said fleet_list does not report collaborators. #709 made it report them, visible
  to the primary, an architect and another collaborator, never a worker.
- the collaborator section justified the ticket refusal by saying ticket ids are a plain counter
  with no owner check. #712 added that owner check, so the stated reason no longer held. The role
  gate is what refuses a collaborator; the recorded creator terminal is the second line.

wiki/7-Use-Cases.md carries the same edit and was pushed to its own remote, verified by ref. The
sync check prints True.
2026-10-04 07:34:23 +02:00
Dai Ha 10ab58e4fc fleetd #710: gate fleet_list's leads and members arrays by caller role
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m50s
Omit the leads and members keys entirely (never an empty array) from
fleet_list's result for a worker, matching the existing
coordinatorVisibleTo/collaboratorsVisibleTo pattern: two new named
predicates (leadsVisibleTo, membersVisibleTo) are consulted before
assembling either array, so a worker holding only READ can no longer
read every session on the daemon through this tool. Primary and
architect callers are unaffected.
2026-10-04 07:25:55 +02:00
Dai Ha 0ba597e394 fleetd #705: close the REST ticket-poll door and pin both handlers' caller-terminal wiring
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m47s
GET /tasks/{ticket} now resolves the caller the same way allow(...) does and
threads that terminal into MessageService.poll(ticket, callerTerminal) instead
of the no-check overload, so a worker can no longer read a ticket a different
session created over REST. Adds a source-scrape guard (with its own control
assertion) for both the fleet_poll MCP handler and this REST route, plus a
behavioural test driving GET /tasks/{ticket} with three differently-resolved
callers against one shared MessageService.
2026-10-04 07:23:29 +02:00
Dai Ha c468953963 Merge PR #714: fleetd #711 — re-key two startup refusals onto the hazard that is still real
CI / shell-tests (push) Failing after 8s
CI / build (push) Failing after 1m53s
CI / contract (push) Successful in 50s
The roster is consulted ahead of every tab map, so a live registered member
is no longer read back as a lead. Both refusals still matter, for the narrower
case where the pane is alive and the roster holds no entry for it.

Behaviour unchanged. Verified in a throwaway worktree, not piped:
Tests run: 2008, FleetConfigTest 170, BUILD SUCCESS.
2026-10-04 07:04:21 +02:00
Dai Ha 25d53e6ef7 Merge PR #713: fleetd #710 — remove the unused CallerResolver.members() accessor
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 2m3s
It returned architectTerminals, so its name contradicted its contents and
collided with fleet_list's members array. No production caller.

Verified in a throwaway worktree, not piped: Tests run: 2008, BUILD SUCCESS.
2026-10-04 07:00:35 +02:00
Dai Ha a9a37af957 Merge PR #712: fleetd #705 — gate fleet_poll's ticket lookup by the creating caller
Closes the MCP door only. The REST door (GET /tasks/{ticket}) still calls the
no-check poll overload, so #705 stays open.

Verified in a throwaway worktree, not piped: Tests run: 2008, BUILD SUCCESS.
2026-10-04 07:00:35 +02:00
Dai Ha b205bcc2aa Remove CallerResolver members accessor
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m47s
2026-10-04 06:53:39 +02:00
Dai Ha 11eccc3a1b fleetd #705: gate fleet_poll's ticket lookup by the creating caller's terminal
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m45s
A ticket id is a plain sequential counter, so any session holding
TASK_READ could walk task-1, task-2, ... and read another session's
delegation reply. sendAsync now records the creating caller's terminal
on the Task, and poll refuses a caller whose terminal differs from it.
A caller with no terminal (the unnamed primary) is still allowed
through regardless, since it never carries a herdr pane to compare.
2026-10-04 06:50:26 +02:00
42 changed files with 5143 additions and 1332 deletions
+35 -7
View File
@@ -159,6 +159,29 @@ fails.
`{action: "cancel", token}` drops a pending request without rolling. `{action: "cancel", token}` drops a pending request without rolling.
**A change is coming: the roll will restart the process instead of sending `/clear` (fleetd #726
unit 2, written 2026-10-04).**
Today a roll types `/clear` into your pane. Your `claude` process keeps running, so a newer CLI on
disk is never loaded. Unit 2 replaces that: the daemon ends the old pane, launches a fresh one,
waits for the new terminal to be recognised as a lead, and only then sends the bootstrap text.
**Everything below about `/clear` is accurate while the old jar is running.** Unit 2 was not merged
when this note was written, and a merge is not a deployment.
**How to tell which one is live: read your own tool list.** If `fleet_handover`'s description says
it will "clear your pane", the daemon is serving the old behaviour. If it names a restart, the new
behaviour is live. The description comes from the running daemon, so it cannot disagree with the
code that is actually loaded.
Two things change for you once it is live. The `status` outcomes are different: three new failures
replace the `/clear` ones. And the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
warning described below can no longer appear, because that wait is deleted — so if you still see
it, the old jar is running. `TURN_NEVER_SETTLED` does not change, and still means nothing was
touched.
**Delete this note and rewrite the `/clear` paragraphs once the new jar is live.**
**Things that will surprise you:** **Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll - **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
@@ -200,13 +223,18 @@ fails.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell. - **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log. Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was - **The bootstrap prompt works end to end. Measured 2026-09-22, re-measured 2026-10-04.** This used
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log to say the fix was unproven (fleetd #489) and told you to expect a failure. That is no longer
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43, true. On 2026-09-22 the daemon log held four `lead-rollover: rolled` lines. On 2026-10-04 it holds
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the **20**, against a control of 86 `lead-rollover:` lines. Each roll cleared the old lead and started
handover file, with the configured `bootstrapText` arriving as its first message. No context was a fresh session against the handover file, with the configured `bootstrapText` arriving as its
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log first message. No context was lost. The old `Unknown command: /clearFresh` failure from 2026-09-12
at all. Re-measure both numbers with: does not appear in the log at all.
19 of the 20 carry an `elapsedMs`: median 16507 ms, maximum 48261 ms, and two above 45000 ms. That
figure times the **whole** roll, and the wait for your own turn to end dominates it, so do not
read it as the cost of the clear. Expect a roll to take tens of seconds, and do not treat a slow
one as a failed one. Re-measure all of these with:
```bash ```bash
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
+44 -23
View File
@@ -27,26 +27,31 @@ through its `fleet_*` tools. No session addresses a peer, a broker, or the netwo
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits **Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional. this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `fleet_whoami`.** It returns `primary`, `worker`, `architect`, or `collaborator`, resolved by **Call `fleet_whoami`.** It returns `primary`, `worker`, `architect`, `collaborator`, or `observer`,
the daemon from your connection — unforgeable, and the same resolution its authorization gate uses. resolved by the daemon from your connection — unforgeable, and the same resolution its authorization
A worker also carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the gate uses. A worker also carries its `sessionId`, `profile`, `worktree` and `branch`; an architect
slot name it was bound to; a collaborator carries its registry name and its own `sessionId`, and carries the slot name it was bound to; a collaborator carries its registry name and its own
**no `leader` key** — a collaborator is a named peer, not a primary. Don't infer what you can ask. `sessionId`, and **no `leader` key** — a collaborator is a named peer, not a primary. An **observer**
carries only its own `sessionId`: a pane the daemon could not place as any of the above, authorized
to `READ`/`METRICS` and to `REPLY`/`ASK` on its own pane and nothing more — never `SEND`, never a
ticket. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are a spawned member in the fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run `.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect — on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect,
only `fleet_whoami` does. **And none of them fires for a collaborator at all**: every signal in the or a worker from an **observer** — an observer is just as unspawned as a collaborator and carries
ladder detects a *spawned* member, while a collaborator is a tab a person opened by hand, so it has none of these signals either, so only `fleet_whoami` tells the two apart. **And none of them fires
no charter, no fixed mount name and a normal environment. A collaborator that cannot call for a collaborator at all**: every signal in the ladder detects a *spawned* member, while a
`fleet_whoami` therefore falls to the line below and acts as a worker. That is the safe direction — collaborator is a tab a person opened by hand, so it has no charter, no fixed mount name and a
it under-privileges, and the refusals are loud — but it means a collaborator has no way to learn normal environment. A collaborator — or an observer — that cannot call `fleet_whoami` therefore falls
what it is except by asking. **Still unsure ⇒ act as a worker**, the most restricted member role. The to the line below and acts as a worker. That is the safe direction — it under-privileges, and the
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate — refusals are loud — but it means a collaborator or an observer has no way to learn what it is except
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`, by asking. **Still unsure ⇒ act as a worker**, the most restricted member role this ladder can name.
The two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate
— loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
and the sender silently receives nothing. Fail toward the recoverable error. and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — every role, no exceptions ### Invariants — every role, no exceptions
@@ -100,7 +105,11 @@ below are the procedure — run them in order, every task, not only the big ones
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask` 5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId` `fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never that answers it — but **only to the caller that created that delegation**, so a question raised
under an architect's brief is invisible to you, and seeing none does not mean there is none.
**Only that same creator can answer it.** A `turnId` you came by any other way is refused, so an
architect's worker waits for that architect and not for you.
**A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default. brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and **A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
returns a ticket, and is then never delivered — measured here three times in one session, and the returns a ticket, and is then never delivered — measured here three times in one session, and the
@@ -165,10 +174,10 @@ you decide.
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` | | See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` | | Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` | | Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` | | Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId`, and only the caller that created that delegation |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task | | Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task | | Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Message a **collaborator** on this host | `fleet_send{sessionId: <their terminal>, content}` — but **`fleet_list` does not report collaborators**, so you cannot discover one: it must tell you its `sessionId`, which its own `fleet_whoami` gives it. Coordination only, **never** a task | | Message a **collaborator** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` reports a `collaborators` array, and each row carries that peer's `name` and the `sessionId` you send to. It is visible to you, to an architect and to another collaborator, never to a worker. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused | | Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body | | Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` | | Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
@@ -257,8 +266,8 @@ of those is refused at the gate, not queued.
Two limits worth knowing before you hit them. **You cannot reach a worker** — not even to help one — Two limits worth knowing before you hit them. **You cannot reach a worker** — not even to help one —
because a worker belongs to the lead that spawned it, and routing around that would make you a because a worker belongs to the lead that spawned it, and routing around that would make you a
second orchestrator with no plan. Send to the lead instead. And **you cannot read a ticket**, so you second orchestrator with no plan. Send to the lead instead. And **you cannot read a ticket**, so you
cannot collect a delegation's reply; ticket ids are a plain counter with no owner check, so holding cannot collect a delegation's reply: `fleet_poll` refuses you at the role gate, and a ticket also
one would let you walk every other session's answers. records the terminal that created it, so even a leaked id reads nothing.
Being named buys you a channel, not authority. Your `fleet_send` to a lead is coordination between Being named buys you a channel, not authority. Your `fleet_send` to a lead is coordination between
peers: the lead owes you no obedience, and you owe it none. peers: the lead owes you no obedience, and you owe it none.
@@ -330,10 +339,22 @@ must obey belongs in the charter, not here.
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are (#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own `ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because worktree; and a plugin reaches a member only through `CLAUDE_CONFIG_DIR`, which
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it, `ClaudeCodeLauncher.java:286` exports with `putIfPresent` — so only for a profile that sets
so a member never reads the operator's plugin store. **The plugin is the lead-side surface; `configDir`. Every `claude-code` profile does set one (the four without are `opencode`, which
member-facing assets travel in the worktree.** never reads that variable). **But measured 2026-10-04: two of them point at
`~/.ccs/instances/ltms`, which is the operator's own `CLAUDE_CONFIG_DIR` on this host.** So for an
`opus` or `sonnet` member, "a member never reads the operator's plugin store" is false — it reads
the same store, because that store is the one its `configDir` names. It stays true for `local` and
`local-direct`, which point at `~/.ccs/instances/gx10`. `ClaudeCodeLauncher`'s own javadoc names
the related hazard: that file is rewritten on every spawn, so for those two profiles fleetd and the
operator's live session write the same `.claude.json`, and its compare-and-swap "narrows the
lost-update window, it does not close it". Re-measure which profiles share the operator's dir with
`awk '/^profiles:/{i=1;next} /^[a-z]/{i=0} i&&/^ [a-z-]+:$/{p=$1} i&&/configDir:/{print p,$2}'
fleetd/fleetd.yaml` against `echo $CLAUDE_CONFIG_DIR`; delete this note once no profile names the
operator's dir. **Treat the plugin as the lead-side surface and put member-facing assets in the
worktree** — that conclusion holds either way, because a worktree asset does not depend on which
config dir a member reads.
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/` - **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote). (a submodule with its own remote).
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's - **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
+22 -17
View File
@@ -99,17 +99,17 @@ bind:
# contextHighNudge: false # contextHighNudge: false
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced, # Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its # without an operator doing it by hand. A lead writes a handover file, then asks fleetd to end its
# own pane and bootstrap a fresh session against that file. # own pane, launch a fresh one, and bootstrap that fresh session against the handover file.
# #
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never # Opt-in on purpose — it tears down the lead's own pane on request, so upgrading the daemon must
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even # never acquire that ability for you. Absent block = feature off, and nothing is constructed at all.
# once present, nothing but an explicit confirm() call — one that passes every check — can ever # Even once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that # tear a pane down: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so # fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so it
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot # cannot act on the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See # continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction). # dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order.
# #
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session # handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no # reads must live. No default (an operator-specific path); a present block with no
@@ -118,25 +118,30 @@ bind:
# working directory when that lead has none configured) — never against whatever # working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is # directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not # used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory (fleetd #480 follow-up). # share a working directory.
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes # requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true # # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this # maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING # turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again) # # lead's own turn to end (its pane to report injectable again)
# # before sending /clear at all. If this elapses, /clear is NEVER # # before tearing the old pane down at all. If this elapses, nothing
# # sent — a lead that never goes idle is still doing real work. # # is torn down — a lead that never goes idle is still doing real
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable # # work.
# # again AFTER /clear before giving up (never sends bootstrapText # relaunchReadySeconds: 45 # default 45 — bounds two later waits, after the old pane is gone
# # if this elapses). A separate, second wait from turnSettleSeconds. # # and a fresh one has been launched: first, for the fresh pane to
# # reach a real turn boundary (never sends bootstrapText if THIS one
# # elapses); second, for the new terminal to be recognised as this
# # lead (bootstrapText is sent either way once the first wait
# # passes). A separate, later pair of waits from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to # bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the lead once its pane settles after /clear # # the freshly relaunched lead's pane once it reaches a real turn
# # boundary
# leadRollover: # leadRollover:
# handoverPath: /path/to/handover.md # handoverPath: /path/to/handover.md
# requireOperatorConfirm: true # requireOperatorConfirm: true
# maxDocAgeSeconds: 3600 # maxDocAgeSeconds: 3600
# turnSettleSeconds: 20 # turnSettleSeconds: 20
# clearSettleSeconds: 20 # relaunchReadySeconds: 45
# bootstrapText: "Fresh lead session: read the handover file and carry on." # bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list # Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
@@ -226,9 +226,8 @@ public final class Fleetd {
* is the same way — a person's own tab, matched to a configured name, never spawned. * is the same way — a person's own tab, matched to a configured name, never spawned.
* *
* <p>Neither a lead nor a collaborator is ever enrolled in {@link MemberPresence} — {@code * <p>Neither a lead nor a collaborator is ever enrolled in {@link MemberPresence} — {@code
* FleetMcp} marks presence for every spawned member (worker and architect), deliberately, since * FleetMcp} marks presence only for a worker, an architect, or the unconfigured-pane floor,
* that map doubles as the member roster's availability signal and a lead or collaborator counted * never for a lead or a collaborator. So without the second and third disjuncts a lead or
* there would show up as an available member. So without the second and third disjuncts a lead or
* collaborator is permanently un-deliverable: every send to one sat on the gate for * collaborator is permanently un-deliverable: every send to one sat on the gate for
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane. * {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
* *
@@ -942,21 +941,27 @@ public final class Fleetd {
* cfg.leadHeartbeat()} * cfg.leadHeartbeat()}
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not * @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
* {@code memberAgents}), normally {@code router.leadAgents()} * {@code memberAgents}), normally {@code router.leadAgents()}
* @param leadSpaces the {@link WorkspaceControl} instance that reaches the LEAD's
* workspace, normally {@code router.leadSpaces()} — used to tear down
* a rolled lead's old pane and confirm it is gone
* @param launcher starts the fresh lead a roll relaunches once the old one is gone
* @param config the live {@link ConfigRef}, captured only inside the returned * @param config the live {@link ConfigRef}, captured only inside the returned
* supplier and the workspace lookup — never dereferenced here * supplier and the two lookups below — never dereferenced here
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally * @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
* the same {@code leads} supplier {@code main} already builds for * the same {@code leads} supplier {@code main} already builds for
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot * {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is * @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
* absent from the startup config * absent from the startup config
*/ */
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config, static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents,
WorkspaceControl leadSpaces, LeadLauncher launcher, ConfigRef config,
Supplier<Map<String, String>> liveLeadTerminals) { Supplier<Map<String, String>> liveLeadTerminals) {
if (cfg.leadRollover() == null) { if (cfg.leadRollover() == null) {
return null; return null;
} }
Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal);
Function<String, String> leadWorkspace = terminal -> { Function<String, String> leadWorkspace = terminal -> {
String leadName = liveLeadTerminals.get().get(terminal); String leadName = leadNameForTerminal.apply(terminal);
if (leadName == null) { if (leadName == null) {
return null; return null;
} }
@@ -964,7 +969,8 @@ public final class Fleetd {
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName); FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
return leader == null ? null : leader.cwd(); return leader == null ? null : leader.cwd();
}; };
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace); return new LeadRollover(leadAgents, leadSpaces, launcher, () -> config.get().leadRollover(),
leadWorkspace, leadNameForTerminal, liveLeadTerminals);
} }
/** /**
@@ -51,7 +51,6 @@ import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
import dev.ltms.fleet.power.IdleSleepGuard; import dev.ltms.fleet.power.IdleSleepGuard;
import dev.ltms.fleet.rest.FleetApp; import dev.ltms.fleet.rest.FleetApp;
import dev.ltms.fleet.session.GitWorktrees; import dev.ltms.fleet.session.GitWorktrees;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager; import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.SessionReaper; import dev.ltms.fleet.session.SessionReaper;
import io.javalin.Javalin; import io.javalin.Javalin;
@@ -298,11 +297,15 @@ final class FleetdAssembly {
leadsRef.set(leads); leadsRef.set(leads);
collaboratorTerminalsRef.set(collaboratorTerminals); collaboratorTerminalsRef.set(collaboratorTerminals);
// Constructed unconditionally — it is cheap and side-effect free — so a LeadRollover built
// below can relaunch a lead even on a boot where herdr was down for the ensureLeads() call.
LeadLauncher leadLauncher = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg);
// CB-558: start any declared lead that is not already running. After the scanner is built, // CB-558: start any declared lead that is not already running. After the scanner is built,
// and only when herdr answered — the launcher's whole safety property is that it can count // and only when herdr answered — the launcher's whole safety property is that it can count
// live leads first, and must never guess and risk a second orchestrator. // live leads first, and must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) { if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads(); int launched = leadLauncher.ensureLeads();
if (launched > 0) { if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched); log.info("lead auto-launch: {} lead(s) started", launched);
} }
@@ -441,7 +444,8 @@ final class FleetdAssembly {
heartbeatScheduler.shutdownNow(); heartbeatScheduler.shutdownNow();
} }
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed. // fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads); LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), router.leadSpaces(),
leadLauncher, config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox, MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics); pushLoop, metrics);
@@ -480,13 +484,9 @@ final class FleetdAssembly {
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup()); new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// fleetd #669 Unit D: a live spawned member resolves as its own role, whatever a tab map // fleetd #669 Unit D: a live spawned member resolves as its own role, whatever a tab map
// says about the same terminal — read from the roster meant for a hot path (SessionManager // says about the same terminal. fleetd #702: SessionManager.spawnedMemberRole also answers
// javadoc), never rosterResolved(), since resolve() runs on every request. // for a pane mid-teardown, not only one still in the registry — see its javadoc.
Function<String, MemberRole> spawnedMemberRole = terminal -> sessions.roster().stream() Function<String, MemberRole> spawnedMemberRole = sessions::spawnedMemberRole;
.filter(s -> terminal.equals(s.terminalId()))
.map(MemberSession::role)
.findFirst()
.orElse(null);
// CB-501: one resolver behind both entry paths. Worker identity still comes from the // CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out. // connection and is never token-gated, so enabling token mode cannot lock the fleet out.
@@ -75,6 +75,9 @@ public final class Authz {
* {@code SEND} is refused, as if no terminal were a configured lead or collaborator — the * {@code SEND} is refused, as if no terminal were a configured lead or collaborator — the
* same decision {@link #NO_KNOWN_LEAD_OR_COLLABORATOR} gives explicitly. Every other action's * same decision {@link #NO_KNOWN_LEAD_OR_COLLABORATOR} gives explicitly. Every other action's
* result is identical to the four-argument form's, since none of them consult the classifier. * result is identical to the four-argument form's, since none of them consult the classifier.
*
* <p>Its default classifier denies every collaborator, so a caller enforcing authorization
* must use the four-argument form instead.
*/ */
public static boolean permits(Principal caller, Action action, String targetSession) { public static boolean permits(Principal caller, Action action, String targetSession) {
return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR); return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR);
@@ -135,13 +138,15 @@ public final class Authz {
// fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no // fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no
// other session's turn state. Those live under TASK_READ. METRICS is the separate // other session's turn state. Those live under TASK_READ. METRICS is the separate
// Prometheus scrape. Both are open to every authenticated role, including a // Prometheus scrape. Both are open to every authenticated role, including a
// collaborator. // collaborator and the unconfigured-pane floor.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect() case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect()
|| caller.isCollaborator(); || caller.isCollaborator() || caller.isObserver();
// Ticket polling and session status, open to every role READ is open to except a // Ticket polling and session status, open to every role READ is open to except a
// collaborator: ticket ids are a sequential counter with no owner check, so a holder // collaborator or an observer. MessageService compares a ticket's creator to the
// could walk every ticket and read another session's delegation reply. // caller on every read as well, so dropping this gate would not expose another
// session's reply — it would move the refusal later and widen what a caller that
// never orchestrates can probe.
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect(); case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect // fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
@@ -40,7 +40,7 @@ import java.util.function.Supplier;
* the case the previous step does not catch: a binding with no live spawned-member session.</li> * the case the previous step does not catch: a binding with no live spawned-member session.</li>
* <li>A loopback peer PID that maps to an operator-labelled collaborator tab ⇒ * <li>A loopback peer PID that maps to an operator-labelled collaborator tab ⇒
* {@link Role#COLLABORATOR}, carrying that collaborator's name.</li> * {@link Role#COLLABORATOR}, carrying that collaborator's name.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is * <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#OBSERVER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured * unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li> * regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li> * <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
@@ -249,17 +249,6 @@ public final class CallerResolver {
return leadTerminals.get(); return leadTerminals.get();
} }
/**
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
*
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
*/
public Map<String, String> members() {
return architectTerminals.get();
}
/** /**
* The currently-recognised collaborator tabs, {@code terminal_id → name}. * The currently-recognised collaborator tabs, {@code terminal_id → name}.
* *
@@ -334,7 +323,7 @@ public final class CallerResolver {
// of the above keeps that stronger role. // of the above keeps that stronger role.
return Principal.collaborator(collaborator, c.terminal(), c.pid()); return Principal.collaborator(collaborator, c.terminal(), c.pid());
} }
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated return Principal.observer(c.terminal(), c.pid()); // unforgeable; never token-gated
} }
if (tokenMode) { if (tokenMode) {
@@ -88,6 +88,13 @@ public record Principal(Role role, String terminal, long pid, String name) {
return new Principal(Role.COLLABORATOR, terminal, pid, name); return new Principal(Role.COLLABORATOR, terminal, pid, name);
} }
/**
* The unconfigured-pane floor: a loopback caller whose pane matched no other role.
*/
public static Principal observer(String terminal, long pid) {
return new Principal(Role.OBSERVER, terminal, pid);
}
public boolean isPrimary() { public boolean isPrimary() {
return role == Role.PRIMARY; return role == Role.PRIMARY;
} }
@@ -104,11 +111,15 @@ public record Principal(Role role, String terminal, long pid, String name) {
return role == Role.WORKER; return role == Role.WORKER;
} }
public boolean isObserver() {
return role == Role.OBSERVER;
}
/** /**
* Whether this caller is a spawned member with its own pane. * Whether this caller is a spawned member with its own pane.
* *
* <p>Both workers and architects are spawned members. A lead is excluded because recording it * <p>Both workers and architects are spawned members. A lead is not: it is a peer the
* as present would count it as an available member in the roster. * operator started and named, never a pane this daemon spawned.
*/ */
public boolean isSpawnedMember() { public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT; return role == Role.WORKER || role == Role.ARCHITECT;
@@ -134,12 +145,32 @@ public record Principal(Role role, String terminal, long pid, String name) {
return terminal != null && terminal.equals(sessionId); return terminal != null && terminal.equals(sessionId);
} }
/**
* Stable identity used to own tickets and open turns. The unnamed primary has no owner key so
* it can use the message layer's primary-wide ticket access rule.
*/
public String ownerKey() {
return switch (role) {
case PRIMARY -> name == null ? null : prefixed("leader", name);
case WORKER -> prefixed("worker", terminal);
case ARCHITECT -> prefixed("architect", terminal);
case COLLABORATOR -> prefixed("collaborator", name);
case OBSERVER -> prefixed("observer", terminal);
case ANONYMOUS -> "anonymous";
};
}
private static String prefixed(String role, String identity) {
return role + ":" + identity;
}
/** Short, non-sensitive description for audit lines and error details. */ /** Short, non-sensitive description for audit lines and error details. */
public String describe() { public String describe() {
return switch (role) { return switch (role) {
case WORKER -> "worker:" + terminal; case WORKER -> "worker:" + terminal;
case ARCHITECT -> "architect:" + name; case ARCHITECT -> "architect:" + name;
case COLLABORATOR -> "collaborator:" + name; case COLLABORATOR -> "collaborator:" + name;
case OBSERVER -> "observer:" + terminal;
case PRIMARY -> name == null ? "primary" : "leader:" + name; case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous"; case ANONYMOUS -> "anonymous";
}; };
@@ -45,6 +45,17 @@ public enum Role {
*/ */
COLLABORATOR, COLLABORATOR,
/**
* A loopback pane that resolved to none of the roles above: not a live spawned member, not a
* configured lead, not a bound architect slot, not a configured collaborator tab. Unforgeable
* like a worker's — derived from the connection's pane, never from a request argument, and
* honoured regardless of auth mode. May {@code READ} and {@code METRICS}, and {@code REPLY}/
* {@code ASK} only as its own pane; may not {@code SPAWN}/{@code STOP}/{@code DRAIN}/
* {@code HANDOVER}, {@code SEND}, poll a ticket ({@code TASK_READ}), or reach the coordination
* broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
OBSERVER,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */ /** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS ANONYMOUS
} }
@@ -67,7 +67,7 @@ import java.util.function.Supplier;
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a * {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape) * {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/ * and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every * {@code turnSettleSeconds}/{@code relaunchReadySeconds}/{@code bootstrapText} fresh on every
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()} * {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather * continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
* than capturing them into fields at construction — unlike its closest * than capturing them into fields at construction — unlike its closest
@@ -1416,17 +1416,16 @@ public record FleetConfig(
* before anything exists to call — the same fact already true of adding a brand-new * before anything exists to call — the same fact already true of adding a brand-new
* {@code profiles:} entry. * {@code profiles:} entry.
* *
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is * <p><strong>{@code turnSettleSeconds}:</strong> {@code confirm()} is called FROM the calling
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant * lead's own turn, so its pane is still {@code WORKING} the instant {@code confirm()} validates
* {@code confirm()} validates every gate and schedules the roll. {@code * every gate and schedules the roll. {@code dev.ltms.fleet.lead.LeadRollover}'s deferred
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for * continuation waits up to this many seconds for that SAME pane to report {@code IDLE} or
* that SAME pane to report {@code IDLE} or {@code DONE} — i.e. for the calling turn to actually * {@code DONE} — i.e. for the calling turn to actually end — before it ends the old pane's
* end — before it sends {@code /clear} at all. {@code BLOCKED} does not count: that is a live * process at all. {@code BLOCKED} does not count: that is a live turn merely paused, not one
* turn merely paused, not one that has finished. If that wait times out, no {@code /clear} is * that has finished. If that wait times out, the old pane is never touched: a lead that never
* ever sent: a lead that never goes idle is still doing real work, and clearing it would * goes idle is still doing real work, and the roll ends that pane's whole process — there is no
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which * way back from this once it runs, so this wait is the only thing standing between "still
* bounds the SECOND wait, for the pane to reach {@code IDLE} or {@code DONE} again AFTER * working" and "gone".
* {@code /clear} has already gone out.
* *
* @param handoverPath required when this block is present — where the handover file a fresh * @param handoverPath required when this block is present — where the handover file a fresh
* lead session reads must live. There is no sane non-null default for an * lead session reads must live. There is no sane non-null default for an
@@ -1447,36 +1446,49 @@ public record FleetConfig(
* attempt can never be mistaken for a fresh one. * attempt can never be mistaken for a fresh one.
* @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the * @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report {@code IDLE} or * CALLING lead's own turn to end (its pane to report {@code IDLE} or
* {@code DONE}) before sending {@code /clear} at all. See the paragraph * {@code DONE}) before ending that pane's process at all. See the
* above. * paragraph above.
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to * @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
* report {@code IDLE} or {@code DONE} again after {@code /clear} before * the old lead's pane has been torn down and a fresh one launched: first,
* giving up. A roll that times out here never sends {@code bootstrapText}. * for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
* typing into a pane that has not finished booting loses the keystrokes;
* second, for the fresh terminal to show up as a recognised lead, which is
* bookkeeping rather than a safety gate, so a timeout on this second wait
* does not withhold {@code bootstrapText} — it is sent once the pane is
* ready regardless. Recognition comes from the same periodically-refreshed
* scan {@code LeadTabScanner} already keeps ({@code scanIntervalSeconds},
* 10s live), so a budget has to clear more than one scan interval to leave
* any real margin for the CLI's own boot time; 20 was rejected for exactly
* that reason — at a 10s scan interval it only buys two scans. 45 buys
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
* ready) withholds {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the * @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* lead's pane once it settles after {@code /clear}, telling the fresh * fresh lead's pane once it reaches a real turn boundary after relaunch,
* session where to read the handover and carry on. Left {@code null} here * telling the fresh session where to read the handover and carry on. Left
* when the operator configures none: the default sentence cannot be built * {@code null} here when the operator configures none: the default sentence
* at construction time because it must name the path AFTER {@code * cannot be built at construction time because it must name the path AFTER
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code * {@code dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative
* handoverPath} against the calling lead's workspace, which this record has * {@code handoverPath} against the calling lead's workspace, which this
* no way to know — see {@link #bootstrapTextFor(String)}. * record has no way to know — see {@link #bootstrapTextFor(String)}.
*/ */
@JsonIgnoreProperties(ignoreUnknown = true) @JsonIgnoreProperties(ignoreUnknown = true)
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm, public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
Integer maxDocAgeSeconds, Integer turnSettleSeconds, Integer maxDocAgeSeconds, Integer turnSettleSeconds,
Integer clearSettleSeconds, String bootstrapText) { Integer relaunchReadySeconds, String bootstrapText) {
public LeadRollover { public LeadRollover {
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm; requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds; maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds; turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds;
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds; relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0)
? 45 : relaunchReadySeconds;
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText; bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
} }
/** /**
* The text actually sent to the lead's pane once it settles after {@code /clear}: the * The text actually sent to the fresh lead's pane once it reaches a real turn boundary
* operator's configured {@link #bootstrapText} when one is set, otherwise the default * after relaunch: the operator's configured {@link #bootstrapText} when one is set,
* sentence built from {@code resolvedHandoverPath}. * otherwise the default sentence built from {@code resolvedHandoverPath}.
* *
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover * @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
* #open} already resolved — never the raw configured {@link * #open} already resolved — never the raw configured {@link
@@ -1919,6 +1931,7 @@ public record FleetConfig(
rejectNegativeMaxLoad(yaml); rejectNegativeMaxLoad(yaml);
rejectAutoCompactWindowOutOfRange(yaml); rejectAutoCompactWindowOutOfRange(yaml);
warnConflictingAutoCompactWindows(yaml); warnConflictingAutoCompactWindows(yaml);
warnRetiredClearSettleSecondsKey(yaml);
rejectMalformedProfilePatterns(yaml); rejectMalformedProfilePatterns(yaml);
rejectUnknownKind(yaml); rejectUnknownKind(yaml);
rejectUnknownAuthMode(yaml); rejectUnknownAuthMode(yaml);
@@ -2342,6 +2355,33 @@ public record FleetConfig(
names, String.join(", ", detail)); names, String.join(", ", detail));
} }
/**
* Warn when a {@code leadRollover:} block still sets the retired {@code clearSettleSeconds}
* key. {@link LeadRollover} carries {@code @JsonIgnoreProperties(ignoreUnknown = true)} and no
* longer declares that component, so Jackson drops it with no signal of its own — this raw-YAML
* check is the only place an operator's now-inert setting is reported at all; by the time a
* {@link LeadRollover} instance exists to run a validator against, the key is already gone.
*
* @param yaml the raw config text
*/
static void warnRetiredClearSettleSecondsKey(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return;
}
if (raw == null || !(raw.get("leadRollover") instanceof Map<?, ?> leadRollover)) {
return;
}
if (leadRollover.containsKey("clearSettleSeconds")) {
log.warn("leadRollover.clearSettleSeconds is retired and no longer read. Set "
+ "leadRollover.relaunchReadySeconds instead: it bounds how long to wait, after "
+ "a lead is relaunched, for its pane to become ready and then for it to be "
+ "recognised as a lead. Remove clearSettleSeconds from fleetd.yaml.");
}
}
/** /**
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern} * Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own * (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
@@ -0,0 +1,146 @@
package dev.ltms.fleet.herdr;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.charset.StandardCharsets;
import java.util.List;
import java.util.function.IntFunction;
/**
* The three protections every {@code agent.start} caller needs against herdr's pane-typed launch
* surface (fleetd #220, #727): a byte-limit check on the assembled command line, a bounded retry
* on {@code agent_pane_busy} (the target pane's shell has not reached its prompt yet), and a
* bounded retry on {@code agent_name_taken} with a fresh name each attempt. One implementation —
* every caller of {@code agent.start}, lead or member, goes through this seam rather than carrying
* its own copy.
*/
public final class ResilientAgentLaunch {
private static final Logger log = LoggerFactory.getLogger(ResilientAgentLaunch.class);
private ResilientAgentLaunch() {
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkFits}.
*/
public static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** herdr rejects a duplicate agent {@code name}; a caller retries a bumped name this many times. */
public static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
public static final int SHELL_READY_RETRIES = 20;
/** Raised by {@link #checkFits} when the assembled command cannot fit the pane line. */
public static final class TooLargeException extends RuntimeException {
public TooLargeException(String message) {
super(message);
}
}
/**
* Verify the assembled launch line fits the pane herdr types it into. herdr does not exec the
* launch command — it TYPES it into the pane as one line, and a pty line buffer holds only
* {@value #PANE_COMMAND_BYTE_LIMIT} bytes. Everything past that byte is dropped with no error
* anywhere: herdr answers "agent started", the backend exits on the mangled argument it was
* handed, the pane closes, and the only symptom is a readiness timeout with no reason. That is
* how fleetd #214 broke every claude-code spawn — one 50-byte flag pushed a 978-byte command to
* 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than start something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @param label names the launch in the refusal message (a profile name)
* @param argv the full argv, including the executable at index 0
* @throws TooLargeException naming the limit, the estimate, and the longest argument
*/
public static void checkFits(String label, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new TooLargeException(
"launch command for " + label + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* Start an agent into {@code paneId}, retrying {@code agent_pane_busy} up to {@code retries}
* times with {@code sleeper} run between attempts.
*
* @throws HerdrException the last {@code agent_pane_busy} failure once {@code retries} is
* spent, or immediately for any other herdr failure
*/
public static Agent startAwaitingShellPrompt(AgentControl agents, String name, String kind,
List<String> args, String paneId,
int retries, Runnable sleeper) {
HerdrException busy = null;
for (int attempt = 0; attempt < retries; attempt++) {
try {
return agents.start(name, kind, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
/**
* Start an agent under a freshly generated name each attempt, retrying {@code agent_name_taken}
* up to {@code nameRetries} times — herdr refuses a duplicate {@code name} outright, so a stale
* registry entry (a crashed session, a name the registry has not yet released) must not block a
* legitimate relaunch. Each attempt also carries its own {@link #startAwaitingShellPrompt} retry.
*
* @param nameForAttempt called once per attempt (0-based) to produce that attempt's name
* @throws HerdrException the last {@code agent_name_taken} failure once {@code nameRetries} is
* spent, or immediately for any other herdr failure
*/
public static Agent startUniquelyNamed(AgentControl agents, String kind, List<String> args,
String paneId, IntFunction<String> nameForAttempt,
int nameRetries, int shellReadyRetries, Runnable sleeper) {
HerdrException last = null;
for (int attempt = 0; attempt < nameRetries; attempt++) {
String name = nameForAttempt.apply(attempt);
try {
return startAwaitingShellPrompt(agents, name, kind, args, paneId,
shellReadyRetries, sleeper);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("agent name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
}
@@ -5,6 +5,7 @@ import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException; import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PendingCloseMarker; import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.Tab; import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace; import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl; import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -12,6 +13,7 @@ import dev.ltms.fleet.launch.ClaudeCodeArguments;
import org.slf4j.Logger; import org.slf4j.Logger;
import org.slf4j.LoggerFactory; import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList; import java.util.ArrayList;
import java.util.LinkedHashMap; import java.util.LinkedHashMap;
import java.util.LinkedHashSet; import java.util.LinkedHashSet;
@@ -19,6 +21,7 @@ import java.util.List;
import java.util.Map; import java.util.Map;
import java.util.Objects; import java.util.Objects;
import java.util.Set; import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import dev.ltms.fleet.peer.PeerLauncher; import dev.ltms.fleet.peer.PeerLauncher;
/** /**
@@ -80,9 +83,19 @@ public final class LeadLauncher {
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class); private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
/** Attempts {@link #relaunch(String)} makes before giving up and returning {@code null}. */
static final int RELAUNCH_ATTEMPTS = 3;
private final AgentControl agents; private final AgentControl agents;
private final WorkspaceControl spaces; private final WorkspaceControl spaces;
private final FleetConfig cfg; private final FleetConfig cfg;
private final Runnable sleeper;
// Per-process token mixed into each lead agent name so a fresh daemon process (seq back at 0)
// cannot collide with a same-name lead that outlived a restart — the same scheme
// HerdrPeerLauncher uses for members (fleetd #727).
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
private final AtomicLong nameSeq = new AtomicLong();
/** /**
* @param agents herdr agent control (start, list) * @param agents herdr agent control (start, list)
@@ -90,9 +103,27 @@ public final class LeadLauncher {
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab * @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/ */
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) { public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) {
this(agents, spaces, cfg, () -> sleepUninterruptibly(300));
}
/**
* Test seam: as above, plus an injectable {@code sleeper} for the {@code agent_pane_busy}
* retry (fleetd #727), so a test can prove the retry budget without a real sleep.
*/
LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg, Runnable sleeper) {
this.agents = agents; this.agents = agents;
this.spaces = spaces; this.spaces = spaces;
this.cfg = cfg; this.cfg = cfg;
this.sleeper = sleeper;
}
/** Uninterruptible sleep — the production {@link #sleeper} between {@code agent_pane_busy} retries. */
private static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
} }
/** /**
@@ -173,24 +204,14 @@ public final class LeadLauncher {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted); log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue; continue;
} }
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have fleetd start it.",
name, name);
continue;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile()); ResolvedLead resolved = resolveLaunchable(name);
if (profile == null) { if (resolved == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
continue; continue;
} }
for (int i = running; i < wanted; i++) { for (int i = running; i < wanted; i++) {
if (launch(name, lead, profile)) { if (launch(name, resolved.lead(), resolved.profile()) != null) {
started++; started++;
} }
} }
@@ -198,6 +219,82 @@ public final class LeadLauncher {
return started; return started;
} }
/** A declared lead paired with the profile it launches on — {@link #resolveLaunchable}'s result. */
private record ResolvedLead(FleetConfig.Leader lead, FleetConfig.Profile profile) {
}
/**
* The declared {@code Leader} and its {@code Profile} for {@code name}, read from the config
* snapshot this launcher was constructed with.
*
* @return the resolved pair, or {@code null} (having logged) if {@code name} is not declared
* under {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
* recognise-only lead), or its {@code profile:} is not configured. Shared by
* {@link #ensureLeads()} and {@link #relaunch(String)} so the three refusals and their
* wording live in one place.
*/
private ResolvedLead resolveLaunchable(String name) {
FleetConfig.Leader lead = cfg.fleet().leaders().get(name);
if (lead == null) {
log.warn("lead '{}' is not declared under fleet.leaders — not launching", name);
return null;
}
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' names no profile — it can be recognised but not launched. Add "
+ "`profile:` under fleet.leaders.{} to have fleetd start it.", name, name);
return null;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
return null;
}
return new ResolvedLead(lead, profile);
}
/**
* Start the named lead from the config snapshot this launcher was constructed with — not a
* live read, so a lead's {@code profile:} or {@code tab:} edited in config needs a daemon
* restart to take effect here — outside of {@link #ensureLeads()}'s {@code instances}
* bookkeeping.
*
* @return the started {@link Agent}, or {@code null} if {@code name} is not declared under
* {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
* recognise-only lead), its {@code profile:} is not configured, or every attempt up to
* {@link #RELAUNCH_ATTEMPTS} failed to start it. Never throws.
*
* <p>Does not count how many instances of this lead are already live. {@link #ensureLeads()}'s
* count exists to avoid starting a second orchestrator; the caller of this method has already
* decided to replace the lead and owns that decision.
*
* <p>Retries the whole launch attempt — not only the {@code agent_name_taken}/
* {@code agent_pane_busy} cases {@link ResilientAgentLaunch} already retries inside one
* {@code agents.start} call — up to {@link #RELAUNCH_ATTEMPTS} times, sleeping via the
* injected sleeper between attempts, and returns the agent from the first attempt that
* succeeds.
*/
public Agent relaunch(String name) {
ResolvedLead resolved = resolveLaunchable(name);
if (resolved == null) {
return null;
}
for (int attempt = 1; attempt <= RELAUNCH_ATTEMPTS; attempt++) {
Agent started = launch(name, resolved.lead(), resolved.profile());
if (started != null) {
return started;
}
if (attempt < RELAUNCH_ATTEMPTS) {
sleeper.run();
}
}
return null;
}
/** /**
* How many live leads exist per configured name, and which of that name's labelled tabs are * How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab} * <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
@@ -306,8 +403,17 @@ public final class LeadLauncher {
return null; return null;
} }
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */ /**
private boolean launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) { * Start one lead. Returns null (having logged) rather than throwing on any failure.
*
* <p>Goes through the same {@link ResilientAgentLaunch} seam every member spawn uses
* (fleetd #727): the assembled argv is refused outright if it cannot fit the pane line herdr
* types it into, a stale {@code agent_name_taken} (a crashed session's name the registry has
* not yet released) is retried under a fresh per-attempt name rather than refusing the whole
* relaunch, and a seed pane whose shell has not reached its prompt yet ({@code
* agent_pane_busy}) is retried rather than failing on the first miss.
*/
private Agent launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
String label = lead.tabLabel(); String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank()) String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd(); ? System.getProperty("user.dir") : lead.cwd();
@@ -323,8 +429,12 @@ public final class LeadLauncher {
// Same shape as the member launchers: herdr resolves the executable from `kind`, so // Same shape as the member launchers: herdr resolves the executable from `kind`, so
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed. // argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
List<String> argv = leadArgv(profile); List<String> argv = leadArgv(profile);
Agent started = agents.start("lead-" + name, herdrKind(profile), ResilientAgentLaunch.checkFits(profile.profile(), argv);
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId()); List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
Agent started = ResilientAgentLaunch.startUniquelyNamed(agents, herdrKind(profile), args,
tab.rootPaneId(),
attempt -> "lead-" + name + "-" + nameNonce + "-" + nameSeq.incrementAndGet(),
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
// Label AFTER the start succeeds. A label written before would survive a failed start // Label AFTER the start succeeds. A label written before would survive a failed start
// and then read back as a live lead on the next boot, which is the exact staleness the // and then read back as a live lead on the next boot, which is the exact staleness the
@@ -334,7 +444,7 @@ public final class LeadLauncher {
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}", log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
name, profile.profile(), tab.tab().tabId(), started.paneId(), name, profile.profile(), tab.tab().tabId(), started.paneId(),
started.terminalId(), label, cwd); started.terminalId(), label, cwd);
return true; return started;
} catch (RuntimeException e) { } catch (RuntimeException e) {
log.warn("lead '{}' failed to launch on profile '{}': {}", log.warn("lead '{}' failed to launch on profile '{}': {}",
name, profile.profile(), e.getMessage()); name, profile.profile(), e.getMessage());
@@ -346,7 +456,7 @@ public final class LeadLauncher {
tab.tab().tabId(), cleanup.getMessage()); tab.tab().tabId(), cleanup.getMessage());
} }
} }
return false; return null;
} }
} }
@@ -1,8 +1,11 @@
package dev.ltms.fleet.lead; package dev.ltms.fleet.lead;
import dev.ltms.fleet.config.FleetConfig; import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus; import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.WorkspaceControl;
import org.slf4j.Logger; import org.slf4j.Logger;
import org.slf4j.LoggerFactory; import org.slf4j.LoggerFactory;
@@ -24,60 +27,69 @@ import java.util.function.Supplier;
/** /**
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an * fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once * operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation * every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation ends
* clears the lead's own pane and bootstraps a fresh session against that file. * the lead's own pane, launches a fresh one, and bootstraps that fresh session against the file.
* *
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code * <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link * dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier * #cancel}, and {@link #status} from a tool call.
* version of this paragraph said nothing called this class at all; that stopped being true once
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
* *
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first * <p><strong>{@code confirm()} cannot roll inline.</strong> {@code confirm()} is called BY the
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code * lead, FROM the lead's own turn: the lead's pane is {@code WORKING} for the whole duration of that
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code * call and cannot possibly report a real turn boundary until {@code confirm()} itself returns. So
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING} * {@link #confirm} validates every gate, then does no I/O against the lead's own pane at all — it
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()} * only records that the request is approved and hands a one-shot continuation to {@code
* itself returns. The poll always timed out — but only after the {@code /clear} had already been * continuationRunner} before returning. That continuation is what actually touches the pane, once
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the * the calling turn has ended, in this order:
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
* ever started, and a refusal return value that claimed nothing had happened.
*
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
* at all — it only records that the request is approved and hands a one-shot continuation to
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
* once the calling turn has ended, in this order:
* <ol> * <ol>
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code * <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that * DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If * approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong> * this never happens, nothing else in this list runs: the old pane is never touched.</strong>
* A lead that never goes idle is a lead still doing real work, and clearing it would throw * A lead that never goes idle is a lead still doing real work, and tearing it down would throw
* away live context — exactly the failure this correction exists to prevent.</li> * away live context.</li>
* <li>{@code agents.send(lead, "/clear")}</li> * <li>capture the old pane id (and, through it, the old tab) from {@link AgentControl#get}, with
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds} * a bounded retry — the terminal-to-pane lookup it goes through can itself report a genuinely
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no * live agent as not found (see {@code AgentControl#agentCall}'s own re-resolve-once
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen, * behaviour), and one false negative here must not abort an otherwise-healthy roll. Neither id
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has; * is ever re-resolved from the terminal again after this — once the pane below is closed there
* see {@link #waitForClearPickupAndSettle})</li> * is nothing left to resolve it from.</li>
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li> * <li>resolve the lead's configured name from its terminal, for the relaunch step below.</li>
* <li>end the old session: close the pane (an already-gone pane counts as success; any other
* failure propagates), then close its tab only when the pane was that tab's sole occupant —
* the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for a member.</li>
* <li>confirm the old pane is actually gone by polling {@link
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} for a {@code null} result — never {@link
* AgentControl#status}, and never the live-lead terminal map, each of which answers a
* different question. <strong>If the old pane is never confirmed gone, no relaunch is
* attempted</strong> — see {@link RollState#OLD_PANE_NEVER_DIED}.</li>
* <li>launch a fresh lead with {@code LeadLauncher#relaunch}. <strong>If every attempt fails,
* {@code bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_FAILED}.</li>
* <li>wait for the fresh pane to reach a real turn boundary ({@code IDLE} or {@code DONE},
* never merely {@code BLOCKED}), bounded by {@code relaunchReadySeconds}. This is the
* safety gate: typing into a pane that has not actually finished booting loses the
* keystrokes. <strong>If the pane never becomes ready, {@code bootstrapText} is never
* sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
* <li>wait for the fresh terminal to be recognised as a live lead — present in the live-lead
* terminal map — bounded by {@code relaunchReadySeconds}. This is bookkeeping, not a
* safety gate: {@code bootstrapText} is sent either way once the pane is ready, whether or
* not this wait itself times out — see {@link RollState#RELAUNCH_NOT_RECOGNISED}.</li>
* <li>{@code agents.send(newTerminal, cfg.bootstrapTextFor(p.handoverPath()))} — sent to the
* FRESH terminal, never the one that was just torn down.</li>
* </ol> * </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate * A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may * passed and the roll is scheduled"</em>, never <em>"the lead has already been replaced"</em> — the
* still be mid-turn, possibly for a long time, when the caller gets that answer back. * old pane may still be mid-turn, possibly for a long time, when the caller gets that answer back.
* *
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket * <p><strong>The safety invariant.</strong> "No timer, no scheduler, no background thread" means
* that first defined this class required "no timer, no scheduler, no background thread" so that * that nothing but an explicit {@link #confirm} call can ever tear a lead's pane down.
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant * {@code continuationRunner} launches a single-shot task that exists only because one specific,
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
* continuationRunner} launches a single-shot task that exists only because one specific,
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at * already-approved {@link #confirm} call created it — it is not recurring, it is not started at
* construction time or on any schedule, and no two invocations of it ever share state. A recurring * construction time or on any schedule, and no two invocations of it ever share state. A recurring
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will * heartbeat or timer that could decide on its own initiative to roll a pane is absent from this
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that * class. <strong>Nothing but an explicit {@link #confirm} call that passes every gate can ever tear
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work * a pane down — that call may simply finish its own work slightly later than the method return, as
* slightly later than the method return, as a continuation of the same approved request, rather * a continuation of the same approved request, rather than entirely inside the method body.</strong>
* than entirely inside the method body.</strong>
* *
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480 * <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
* correction.</strong> The first version resolved the pane to clear via {@code * correction.</strong> The first version resolved the pane to clear via {@code
@@ -115,19 +127,17 @@ public final class LeadRollover {
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class); private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */ /** Poll interval shared by every bounded wait in this class. */
static final long SETTLE_POLL_MS = 250; static final long POLL_INTERVAL_MS = 250;
/** /**
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows * How long {@link #waitUntilPaneGone} polls {@link WorkspaceControl#locatePane} before giving
* before releasing rather than wedging the roll — the same constant and the same * up on ever seeing the old pane disappear. Not configurable: once {@link #endOldSession} has
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own * closed the pane (and, usually, its tab), herdr dropping the pane from its own bookkeeping is
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of * expected to show up within one or two polls, not on an operator-tunable timescale the way a
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1} * CLI boot is.
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
* nudging again — so 8 polls produce 7 nudges, not 8.
*/ */
static final int PICKUP_GRACE_POLLS = 8; static final int PANE_DEATH_TIMEOUT_SECONDS = 10;
/** /**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}). * One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
@@ -164,15 +174,21 @@ public final class LeadRollover {
* The handover file's modified time is not after {@link #open}'s request timestamp, or is * The handover file's modified time is not after {@link #open}'s request timestamp, or is
* older than {@code maxDocAgeSeconds}. * older than {@code maxDocAgeSeconds}.
*/ */
HANDOVER_STALE HANDOVER_STALE,
/**
* This lead terminal already has a roll running: an earlier {@link #confirm} call claimed
* it and that roll's continuation has not released it yet. {@code detail} names the lead
* terminal and the token that holds the claim.
*/
ROLL_ALREADY_RUNNING
} }
/** /**
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the * The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been * roll has been handed to a one-shot continuation — <strong>not</strong> that the lead has
* cleared; the continuation may still be waiting for the calling turn to end when this returns. * already been replaced; the continuation may still be waiting for the calling turn to end when
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going * this returns. Whether the deferred roll itself later goes on to tear the old pane down and
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's * relaunch the lead, or refuses at any of its own steps, is logged only (see this class's
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to. * javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
*/ */
public record RollDecision(boolean accepted, RefusalReason reason, String detail) { public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
@@ -208,7 +224,7 @@ public final class LeadRollover {
/** /**
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes * What is known about one token, right now — the answer {@link #status} gives. Distinguishes
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll * five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* that has been approved but has not finished yet, and two answers for a token that names no * that has been approved but has not finished yet, and two answers for a token that names no
* active work at all: still pending confirmation, or nothing known about this token at all. * active work at all: still pending confirmation, or nothing known about this token at all.
*/ */
@@ -231,41 +247,64 @@ public final class LeadRollover {
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll * #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
* that is, in fact, actively running. This is not sticky: the deferred continuation * that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link * overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes * #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into * #RELAUNCH_NEVER_READY}, {@link #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it
* {@link #FAILED} instead of leaving this entry stuck forever. * finishes — including by throwing, which {@link #runRollover}'s catch turns into {@link
* #FAILED} instead of leaving this entry stuck forever.
*/ */
IN_PROGRESS, IN_PROGRESS,
/** /**
* {@link #confirm} was approved and the deferred continuation completed the entire roll: * {@link #confirm} was approved and the deferred continuation completed the entire roll: the
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code * calling lead's turn settled, the old pane was torn down and confirmed gone, a fresh lead
* bootstrapText} was sent. * was launched and recognised, and {@code bootstrapText} was sent to it.
*/ */
ROLLED, ROLLED,
/** /**
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary * {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at * (IDLE or DONE) within {@code turnSettleSeconds} — the old pane was never touched at all.
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this * This is the state that makes a lead's own stuck turn VISIBLE: without it, a lead that hit
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find * this case would have no way to find out, and would carry on believing it was about to be
* out, and would carry on believing it was about to be replaced. See this class's javadoc. * replaced. See this class's javadoc.
*/ */
TURN_NEVER_SETTLED, TURN_NEVER_SETTLED,
/** /**
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled * {@link #confirm} was approved and the calling lead's turn settled, the old pane was closed
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent. * (and its tab, if it was the sole occupant), but {@link
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} kept reporting it as still present for
* the whole pane-death timeout. No relaunch was ever attempted, and {@code bootstrapText}
* was never sent.
*/ */
CLEAR_NEVER_SETTLED, OLD_PANE_NEVER_DIED,
/** /**
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a * The old pane was confirmed gone, but {@code LeadLauncher#relaunch} returned {@code null}
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code * — every launch attempt failed. {@code bootstrapText} was never sent, and no fresh terminal
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it. * exists for this roll to have recognised.
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS} */
* forever, because the production {@code continuationRunner} is a bare virtual thread with RELAUNCH_FAILED,
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a /**
* terminal outcome. {@code detail} names the exception, so a reader has something to act on * A fresh lead was launched, but its pane never reached a real turn boundary ({@code IDLE}
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link * or {@code DONE}, never merely {@code BLOCKED}) within {@code relaunchReadySeconds} — the
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck * CLI never finished booting, or it stayed paused on a startup prompt. {@code bootstrapText}
* lead must {@link #open} a fresh request. * was never sent: typing into a pane that is not actually ready to accept input loses the
* keystrokes.
*/
RELAUNCH_NEVER_READY,
/**
* A fresh lead was launched and its pane reached a real turn boundary, so {@code
* bootstrapText} WAS sent to it, but the terminal was never recognised as a live lead —
* present in the live-lead terminal map — within {@code relaunchReadySeconds}. The session
* itself is alive and bootstrapped; only the daemon's own bookkeeping has not caught up, and
* an operator should check why the tab was not recognised.
*/
RELAUNCH_NOT_RECOGNISED,
/**
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
* #IN_PROGRESS} forever, because the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler and nothing downstream of the throw ever runs to
* write a terminal outcome. {@code detail} names the exception, so a reader has something to
* act on. The roll is dead at this point and does not retry itself; a stuck lead must
* {@link #open} a fresh request.
*/ */
FAILED, FAILED,
/** /**
@@ -286,6 +325,10 @@ public final class LeadRollover {
public record RollStatus(RollState state, String detail) {} public record RollStatus(RollState state, String detail) {}
private final AgentControl agents; private final AgentControl agents;
/** Workspace/tab/pane control — used to tear down the old pane and confirm it is gone. */
private final WorkspaceControl spaces;
/** Starts the fresh lead that replaces the one this roll tears down. */
private final LeadLauncher launcher;
private final Supplier<FleetConfig.LeadRollover> configSupplier; private final Supplier<FleetConfig.LeadRollover> configSupplier;
/** /**
* Terminal id → that lead's configured workspace directory (their {@code * Terminal id → that lead's configured workspace directory (their {@code
@@ -295,8 +338,21 @@ public final class LeadRollover {
* daemon-cwd bug this parameter exists to fix. * daemon-cwd bug this parameter exists to fix.
*/ */
private final Function<String, String> leadWorkspace; private final Function<String, String> leadWorkspace;
/**
* Terminal id → that lead's configured name under {@code fleet.leaders}, or {@code null} when
* the terminal names no currently-recognised lead. The deferred continuation calls this, on the
* OLD terminal, before tearing it down, so it knows which lead to pass to {@link
* LeadLauncher#relaunch}.
*/
private final Function<String, String> leadNameForTerminal;
/**
* The daemon's current terminal id → lead name map, read fresh on every poll. The deferred
* continuation polls this for the FRESH terminal {@link LeadLauncher#relaunch} returns, to
* learn when that terminal has been recognised as a live lead — see this class's javadoc.
*/
private final Supplier<Map<String, String>> liveLeadTerminals;
private final LongSupplier nowMillis; private final LongSupplier nowMillis;
private final Runnable settleSleeper; private final Runnable pollSleeper;
/** /**
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual * Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
* thread per confirmed request — see this class's javadoc for why that is a single-shot task, * thread per confirmed request — see this class's javadoc for why that is a single-shot task,
@@ -305,6 +361,15 @@ public final class LeadRollover {
*/ */
private final Consumer<Runnable> continuationRunner; private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>(); private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/**
* Lead terminal → the token of the roll currently holding that terminal exclusive, for
* {@link #confirm}'s single-flight claim. {@link #confirm} claims an entry here with an
* atomic put-if-absent once every other gate has passed, refusing with {@link
* RefusalReason#ROLL_ALREADY_RUNNING} when a claim is already held; {@link #runRollover}
* releases it in a {@code finally}, on both the success and the thrown-exception path. A
* terminal absent from this map has no roll currently in flight for it.
*/
private final Map<String, String> rollingByTerminal = new ConcurrentHashMap<>();
/** /**
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link * Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order * #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
@@ -323,29 +388,40 @@ public final class LeadRollover {
} }
}); });
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */ /** Production constructor — wall clock, real sleep between polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier, public LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
Function<String, String> leadWorkspace) { Supplier<FleetConfig.LeadRollover> configSupplier,
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis, Function<String, String> leadWorkspace,
() -> sleepUninterruptibly(SETTLE_POLL_MS), Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals) {
this(agents, spaces, launcher, configSupplier, leadWorkspace, leadNameForTerminal,
liveLeadTerminals, System::currentTimeMillis,
() -> sleepUninterruptibly(POLL_INTERVAL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r)); r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
} }
/** /**
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation * Full constructor — an injectable wall-clock supplier, poll sleeper, and continuation runner,
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code * for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares * System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
* against a file's modified time, which only a wall clock is comparable to, and {@code * against a file's modified time, which only a wall clock is comparable to, and {@code
* nanoTime} freezes while the host sleeps (fleetd #386). * nanoTime} freezes while the host sleeps.
*/ */
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier, LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
Function<String, String> leadWorkspace, LongSupplier nowMillis, Supplier<FleetConfig.LeadRollover> configSupplier,
Runnable settleSleeper, Consumer<Runnable> continuationRunner) { Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals,
LongSupplier nowMillis, Runnable pollSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents; this.agents = agents;
this.spaces = spaces;
this.launcher = launcher;
this.configSupplier = configSupplier; this.configSupplier = configSupplier;
this.leadWorkspace = leadWorkspace; this.leadWorkspace = leadWorkspace;
this.leadNameForTerminal = leadNameForTerminal;
this.liveLeadTerminals = liveLeadTerminals;
this.nowMillis = nowMillis; this.nowMillis = nowMillis;
this.settleSleeper = settleSleeper; this.pollSleeper = pollSleeper;
this.continuationRunner = continuationRunner; this.continuationRunner = continuationRunner;
} }
@@ -487,6 +563,16 @@ public final class LeadRollover {
return docCheck; return docCheck;
} }
// Single-flight claim: atomic put-if-absent, taken only after every other gate has
// passed, so a refused confirm() never takes it. A non-null previous value means a
// different, still-running roll already holds this lead terminal.
String holder = rollingByTerminal.putIfAbsent(p.leadTerminal(), token);
if (holder != null) {
return RollDecision.refused(RefusalReason.ROLL_ALREADY_RUNNING,
"lead terminal " + p.leadTerminal() + " already has a roll running under token "
+ holder);
}
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and // Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes` // OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
// while it is STILL present in `pending`; status() checks `outcomes` first (see that // while it is STILL present in `pending`; status() checks `outcomes` first (see that
@@ -494,34 +580,48 @@ public final class LeadRollover {
// remove-then-put ordering would leave in which the token is in neither map. // remove-then-put ordering would leave in which the token is in neither map.
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS, outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
"confirm() approved this roll and handed it to the deferred continuation; it has " "confirm() approved this roll and handed it to the deferred continuation; it has "
+ "not finished yet — still waiting for the calling turn to settle, for " + "not finished yet — still waiting for the calling turn to settle, for the "
+ "/clear to be sent and settle, or for bootstrapText to be sent")); + "old pane to be torn down and confirmed gone, for the fresh lead to be "
+ "recognised, or for bootstrapText to be sent"));
pending.remove(token); pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends", log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal); token, callerTerminal);
continuationRunner.accept(() -> runRollover(p, cfg)); try {
continuationRunner.accept(() -> runRollover(p, cfg));
} catch (RuntimeException e) {
// continuationRunner can reject the hand-off itself (e.g. a bounded executor's
// RejectedExecutionException) before runRollover ever starts, so runRollover's own
// finally — the only other place that releases rollingByTerminal — never runs either.
// Release the claim here and overwrite the IN_PROGRESS entry with a terminal outcome,
// or this lead terminal could never be rolled again and status() would report
// IN_PROGRESS forever for a roll that in fact never started.
log.warn("lead-rollover: continuationRunner rejected token={} lead={}: {} — the roll "
+ "never started; releasing its claim and reporting it as FAILED",
token, callerTerminal, e.toString(), e);
rollingByTerminal.remove(p.leadTerminal(), token);
outcomes.put(token, new RollStatus(RollState.FAILED,
"continuationRunner rejected this roll before it ever started: " + e.toString()
+ " — the roll never ran; open() a fresh rollover request"));
}
return RollDecision.approved(); return RollDecision.approved();
} }
/** /**
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs * The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the * entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged * full order. There is no result to return to by this point, so every outcome is logged only.
* only.
* *
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code * <p><strong>The whole body is wrapped in one {@code try}.</strong> Several calls below —
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} → * {@code agents.get}, {@code agents.close}, {@code agents.send} — can throw an unchecked {@link
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see * dev.ltms.fleet.herdr.HerdrException} (see {@code AgentControl.java}), and the production
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual * {@code continuationRunner} is a bare virtual thread with no uncaught-exception handler (see
* thread with no uncaught-exception handler (see this class's public constructor). Before this * this class's public constructor). An uncaught throw would kill the continuation thread
* fix, either throw killed the continuation thread silently, leaving the {@link * silently, leaving the {@link RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link * stuck forever — {@link #status} would have no way to tell a dead roll from one still
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch} * genuinely running. The {@code catch} below is scoped to the method body rather than to each
* below is scoped to the method body rather than to each {@code send} call individually, so it * call individually, so it also covers every call in this continuation, not a fixed list of
* also covers anything else added to this continuation later, not just today's two call sites — * call sites — the same reasoning that put the write-a-terminal-outcome step at each of this
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other * method's other exits rather than inside the helpers that detect them.</p>
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
* branches below) rather than inside the helpers that detect them.</p>
* *
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link * <p>Only {@link RuntimeException} is caught, matching the local convention {@link
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the * #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
@@ -539,6 +639,12 @@ public final class LeadRollover {
"the roll's continuation threw " + e.toString() + " — the roll is dead and will " "the roll's continuation threw " + e.toString() + " — the roll is dead and will "
+ "not retry itself; check the daemon log for the stack trace, then open() " + "not retry itself; check the daemon log for the stack trace, then open() "
+ "a fresh rollover request")); + "a fresh rollover request"));
} finally {
// Release the single-flight claim on both the normal return and the thrown-exception
// path above — a release only on success would leave this lead terminal unrollable
// forever after one failure. The conditional two-argument remove only clears the
// entry this roll itself holds, never a different roll's claim on the same terminal.
rollingByTerminal.remove(p.leadTerminal(), p.token());
} }
} }
@@ -546,53 +652,237 @@ public final class LeadRollover {
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) { private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal(); String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong(); long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds()); TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
if (!turnResult.settled()) { if (!turnResult.settled()) {
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
// path already does.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after " log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "confirm() — refusing to send /clear at all; the calling lead's own " + "confirm() — the old pane is never touched; the calling lead's own "
+ "turn is still live and clearing it now would destroy live context " + "turn is still live and tearing it down now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)", + "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis()); lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED, outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
"the calling lead's own turn never reached a boundary (IDLE or DONE) within " "the calling lead's own turn never reached a boundary (IDLE or DONE) within "
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed=" + "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this " + turnResult.elapsedMillis() + "ms) — the old pane was never touched. If "
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml")); + "this keeps happening, raise turnSettleSeconds in fleetd.yaml"));
return; return;
} }
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext: // Captured once, here, and never re-resolved from `lead` again below: once the pane is
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the // closed there is nothing left for a terminal lookup to find.
// pane forever (see this class's javadoc). Agent oldAgent = captureAgentWithRetry(lead);
agents.send(lead, "/clear"); String oldPaneId = oldAgent.paneId();
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds()); String leadName = leadNameForTerminal.apply(lead);
if (!clearResult.settled()) {
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait endOldSession(oldPaneId);
// actually ran — an operator reading only that number wrongly believes it is a measured DeathResult deathResult = waitUntilPaneGone(oldPaneId);
// duration. Print the measured elapsed time and nudge count alongside it, each labelled, if (!deathResult.gone()) {
// so the two can be compared at a glance. log.warn("lead-rollover: old pane {} for lead {} was never confirmed gone after being "
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after " + "closed — not attempting a relaunch (token={}, timeout={}s "
+ "/clear — NOT sending bootstrapText (token={}, configured={}s " + "elapsed={}ms)",
+ "elapsed={}ms nudges={})", oldPaneId, lead, p.token(), PANE_DEATH_TIMEOUT_SECONDS, deathResult.elapsedMillis());
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(), outcomes.put(p.token(), new RollStatus(RollState.OLD_PANE_NEVER_DIED,
clearResult.nudges()); "the old pane was closed, but locatePane kept reporting it as still present "
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED, + "after a pane-death timeout=" + PANE_DEATH_TIMEOUT_SECONDS
"/clear was sent, but the pane never re-settled within clearSettleSeconds=" + "s (measured elapsed=" + deathResult.elapsedMillis() + "ms) — no "
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis() + "relaunch was attempted"));
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
return; return;
} }
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
Agent newAgent = launcher.relaunch(leadName);
if (newAgent == null) {
log.warn("lead-rollover: relaunch of lead '{}' (old terminal {}) failed every attempt "
+ "— bootstrapText was never sent (token={})", leadName, lead, p.token());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_FAILED,
"lead '" + leadName + "' could not be relaunched — every attempt failed; "
+ "bootstrapText was never sent"));
return;
}
ReadinessResult readinessResult = waitUntilPaneReady(newAgent.terminalId(),
cfg.relaunchReadySeconds());
if (!readinessResult.ready()) {
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) never reached a real "
+ "turn boundary — bootstrapText was never sent (token={}, configured={}s "
+ "elapsed={}ms)",
leadName, newAgent.terminalId(), p.token(), cfg.relaunchReadySeconds(),
readinessResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NEVER_READY,
"fresh terminal " + newAgent.terminalId() + " never reached a real turn "
+ "boundary (IDLE or DONE) within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ readinessResult.elapsedMillis() + "ms) — bootstrapText was never "
+ "sent"));
return;
}
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
cfg.relaunchReadySeconds());
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
if (!identityResult.ready()) {
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
+ "was never recognised as a live lead — an operator should check why "
+ "the tab was not recognised (token={}, configured={}s elapsed={}ms)",
newAgent.terminalId(), leadName, p.token(), cfg.relaunchReadySeconds(),
identityResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NOT_RECOGNISED,
"bootstrapText was sent to fresh terminal " + newAgent.terminalId() + ", but "
+ "it was never recognised as a live lead within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ identityResult.elapsedMillis() + "ms) — check why the tab was not "
+ "recognised"));
return;
}
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis; long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis); log.info("lead-rollover: rolled token={} oldLead={} newTerminal={} elapsedMs={}",
p.token(), lead, newAgent.terminalId(), rollElapsedMillis);
outcomes.put(p.token(), new RollStatus(RollState.ROLLED, outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
"rolled successfully in " + rollElapsedMillis + "ms")); "rolled successfully in " + rollElapsedMillis + "ms; new terminal="
+ newAgent.terminalId()));
} }
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
static final int CAPTURE_RETRIES = 3;
/**
* {@link AgentControl#get} for {@code lead}, retried up to {@link #CAPTURE_RETRIES} times. The
* terminal-to-pane lookup it goes through can report a genuinely live agent as not found (see
* {@code AgentControl#agentCall}'s own re-resolve-once behaviour), and one such false negative
* must not abort an otherwise-healthy roll. The result is captured once by the caller and never
* looked up again — see this class's javadoc.
*
* @throws RuntimeException the last failure, if every attempt fails — {@link #runRollover}'s
* catch turns that into {@link RollState#FAILED}
*/
private Agent captureAgentWithRetry(String lead) {
RuntimeException last = null;
for (int attempt = 1; attempt <= CAPTURE_RETRIES; attempt++) {
try {
return agents.get(lead);
} catch (RuntimeException e) {
last = e;
log.debug("lead-rollover: agents.get({}) failed on attempt {}/{}: {}",
lead, attempt, CAPTURE_RETRIES, e.toString());
if (attempt < CAPTURE_RETRIES) {
pollSleeper.run();
}
}
}
throw last;
}
/**
* End the old lead's session: close its pane, then close its tab only when the pane was that
* tab's sole occupant — the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for
* a member. An already-gone pane counts as success; any other {@code agents.close} failure
* propagates, so a genuinely failed teardown is never reported as done. A failing
* {@code spaces.closeTab} never propagates — by the time it runs the pane is already closed, so
* it is cosmetic tidying, not a real teardown failure.
*/
private void endOldSession(String paneId) {
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) {
throw e;
}
log.debug("lead-rollover: pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
try {
spaces.closeTab(loc.tabId());
} catch (RuntimeException e) {
log.warn("lead-rollover: tab.close({}) failed — the pane is already torn down, so "
+ "continuing; the tab may need manual cleanup: {}", loc.tabId(), e.getMessage());
}
} else if (loc != null) {
log.debug("lead-rollover: not closing tab {} — it holds {} panes (not a dedicated lead "
+ "tab)", loc.tabId(), loc.tabPaneCount());
}
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
/**
* Poll {@link WorkspaceControl#locatePane} for {@code paneId} until it reports {@code null}
* (the pane is gone) or {@link #PANE_DEATH_TIMEOUT_SECONDS} elapses. Deliberately never calls
* {@link AgentControl#status} and never reads the live-lead terminal map — both answer a
* different question (whether an AGENT is live, not whether this PANE still exists) and
* {@code locatePane} alone catches a {@link HerdrException} from the underlying {@code
* pane.get} and turns it into {@code null} — see this class's javadoc.
*/
private DeathResult waitUntilPaneGone(String paneId) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(PANE_DEATH_TIMEOUT_SECONDS);
while (nowMillis.getAsLong() < deadline) {
if (spaces.locatePane(paneId) == null) {
return new DeathResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new DeathResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilPaneGone}. */
private record DeathResult(boolean gone, long elapsedMillis) {}
/**
* Poll until {@code newTerminal}'s own pane reaches a real turn boundary ({@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE}, never merely {@link AgentStatus#BLOCKED}) —
* the same exclusion {@link #waitUntilAtTurnBoundary} applies to the calling lead's own turn,
* applied here to the fresh one, so {@code bootstrapText} is never typed into a pane that has
* not actually finished booting — or {@code readySeconds} elapses. A failed status read
* degrades to "not yet ready" and is retried on the next poll.
*/
private ReadinessResult waitUntilPaneReady(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(newTerminal);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to be ready: {}",
newTerminal, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilPaneReady}. */
private record ReadinessResult(boolean ready, long elapsedMillis) {}
/**
* Poll until {@code newTerminal} is present in {@link #liveLeadTerminals} or {@code
* readySeconds} elapses. This is bookkeeping, not a safety gate: the pane's own readiness (see
* {@link #waitUntilPaneReady}) is what decides whether {@code bootstrapText} is safe to send —
* a timeout here only means the daemon's own lead-discovery scan has not caught up yet.
*/
private IdentityResult waitUntilRecognisedAsLead(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
if (liveLeadTerminals.get().containsKey(newTerminal)) {
return new IdentityResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new IdentityResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilRecognisedAsLead}. */
private record IdentityResult(boolean ready, long elapsedMillis) {}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */ /** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
public boolean cancel(String token) { public boolean cancel(String token) {
return pending.remove(token) != null; return pending.remove(token) != null;
@@ -682,15 +972,12 @@ public final class LeadRollover {
/** /**
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link * Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by * AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used by
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear} * {@link #runRollover} to wait for the CALLING turn's own pane to settle before the old pane is
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The * touched at all — the {@code turnSettleSeconds} gate that makes tearing it down safe. A failed
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd * status read degrades to "not yet settled" and is retried on the next poll, the same posture
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of * {@code LeadHeartbeatLoop} and {@code HerdrPeerLauncher}'s readiness gate already take toward
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not * an unreadable status.
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
* *
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()} * <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live * answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
@@ -698,16 +985,16 @@ public final class LeadRollover {
* an approval prompt is safe to queue a message behind. This class asks a stricter question — * an approval prompt is safe to queue a message behind. This class asks a stricter question —
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is * "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this * merely paused, not one that has finished. Reusing {@code injectable()} here would let this
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and * wait tear the old pane down while the lead's own {@code confirm()}-calling turn is still live
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds} * and paused on a prompt — exactly the live-context-destroying failure {@code turnSettleSeconds}
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link * exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same * #waitUntilPaneReady} applies the same exclusion of {@code BLOCKED} to the fresh lead's own
* reason, on the second wait.) * turn.)
* *
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real * @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
* boundary was observed, {@code false} if {@code settleSeconds} elapses first. * boundary was observed, {@code false} if {@code settleSeconds} elapses first.
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis} * {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up). * clock, never the configured {@code settleSeconds} budget.
*/ */
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) { private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong(); long startMillis = nowMillis.getAsLong();
@@ -724,142 +1011,11 @@ public final class LeadRollover {
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) { if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis); return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
} }
settleSleeper.run(); pollSleeper.run();
} }
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis); return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
} }
/** /** The measured outcome of {@link #waitUntilAtTurnBoundary}. */
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
* count.
*/
private record TurnSettleResult(boolean settled, long elapsedMillis) {} private record TurnSettleResult(boolean settled, long elapsedMillis) {}
/**
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
* own javadoc already records that the submit accompanying a delivery "can race the paste —
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
* it, landing as one concatenated line.
*
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
* <ul>
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
* turn;</li>
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
* Claude Code prompt is a no-op, so repeating it is safe;</li>
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
* observed, releases rather than wedges the roll instead of nudging again — the same
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
* path ran and how long it actually took;</li>
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
* </ul>
*
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
* a real boundary is reached or {@code settleSeconds} runs out.
*
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
* nudge must not abort the roll.
*
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
* {@code settleSeconds} budget (fleetd #494).
*/
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
int idlePollsAwaitingPickup = 0;
int nudges = 0;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
+ "/clear: {}", target, e.toString());
status = null;
}
if (status == AgentStatus.WORKING) {
pickedUp = true;
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
if (pickedUp) {
// a real WORKING -> IDLE/DONE completion boundary
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
}
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
long elapsedMillis = nowMillis.getAsLong() - startMillis;
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
// instead of wedging the roll — that trade is deliberate and stays. But it is
// also exactly the case that reported false success in the real incident (the
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
// print the MEASURED elapsed time next to the target pane, not just the count.
//
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
// this loop, on the same branch, so on this branch they cannot differ from
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
// difference on this line, and printing the counters does not change that.
// What it does buy: one source of truth instead of two, so a later change to
// the loop (an early return, a second increment site, a different exit
// condition) cannot leave this message reporting a number the loop no longer
// produces. The place where `nudges` genuinely varies with the run — and is
// covered by a test that can tell it apart from a constant — is the
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
+ "releasing rather than wedging the roll (elapsed={}ms)",
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
return new ClearSettleResult(true, elapsedMillis, nudges);
}
try {
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
} catch (RuntimeException e) {
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
target, e.getMessage());
} finally {
nudges++; // an attempted nudge, whether or not the submit call itself threw
}
}
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
// signal nor a boundary — keep polling without nudging or releasing.
settleSleeper.run();
}
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
}
/**
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
* the settle/timeout decision, so callers can log them instead of the configured budget, which
* is not how long the wait actually ran.
*/
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
} }
@@ -334,7 +334,7 @@ public final class FleetMcp {
* caller explicitly saying so — never by omitting a {@link CallerResolver} the way the old * caller explicitly saying so — never by omitting a {@link CallerResolver} the way the old
* {@code callers == null} idiom allowed. {@code callers} itself is required either way: even * {@code callers == null} idiom allowed. {@code callers} itself is required either way: even
* under {@link #UNENFORCED}, the one real {@link CallerResolver} still resolves every caller's * under {@link #UNENFORCED}, the one real {@link CallerResolver} still resolves every caller's
* {@link Principal} (so {@code markSpawnedMemberPresent}/{@code recordPrimarySingleton} see a * {@link Principal} (so {@code markTrackedCallerPresent}/{@code recordPrimarySingleton} see a
* real identity), and {@link #denyFor} is the only thing that changes. * real identity), and {@link #denyFor} is the only thing that changes.
*/ */
public enum AuthorizationMode { ENFORCED, UNENFORCED } public enum AuthorizationMode { ENFORCED, UNENFORCED }
@@ -442,10 +442,10 @@ public final class FleetMcp {
// fall back to here — AuthorizationMode governs enforcement, not identity. // fall back to here — AuthorizationMode governs enforcement, not identity.
Principal p = callers.resolve(req.getRemoteAddr(), req.getRemotePort(), Principal p = callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization")); req.getHeader("Authorization"));
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a // Guards on the ROLE, not on the terminal being null — this marks presence for
// lead, which carries its pane too, while including every spawned member role. // a worker, an architect, or the unconfigured-pane floor, and excludes a lead
// Enrolling a lead would count it as an available member in the roster. // or a collaborator even though each carries its own pane too.
markSpawnedMemberPresent(p, presence); markTrackedCallerPresent(p, presence);
return McpTransportContext.create(Map.of( return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()), CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()), CALLER_PID, Long.toString(p.pid()),
@@ -459,12 +459,14 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_send", req.arguments()), McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_send", req.arguments()),
str(req.arguments(), "sessionId")); str(req.arguments(), "sessionId"));
if (denied != null) return denied; if (denied != null) return denied;
String caller = callerTerminal(exchange); Principal caller = principal(exchange);
String callerTerminal = caller.terminal();
String callerOwner = caller.ownerKey();
// CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback. // CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback.
// An architect delegates as its own pane but must never become the fallback that // An architect delegates as its own pane but must never become the fallback that
// no-delegation inbox nudges target as if it were the primary (the per-target // no-delegation inbox nudges target as if it were the primary (the per-target
// delegation map does not cure the singleton). // delegation map does not cure the singleton).
recordPrimarySingleton(primaryRegistry, caller, principal(exchange)); recordPrimarySingleton(primaryRegistry, callerTerminal, caller);
Map<String, Object> a = req.arguments(); Map<String, Object> a = req.arguments();
String target = str(a, "sessionId"); String target = str(a, "sessionId");
String content = str(a, "content"); String content = str(a, "content");
@@ -480,18 +482,18 @@ public final class FleetMcp {
// Answering a worker's fleet_ask (CB-205): resolve its blocked question and // Answering a worker's fleet_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn. This is the same // block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded. // delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a)); return answer(messages, turnId, content, timeoutMs(a), callerOwner);
} }
// CB-548: delegator ownership (which lead's reply nudge this worker routes to, // CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the // CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
// session lock and queued delivery — via the accepted-delivery callback, never at // session lock and queued delivery — via the accepted-delivery callback, never at
// request time. A concurrent sender that times out BUSY therefore cannot steal a // request time. A concurrent sender that times out BUSY therefore cannot steal a
// live turn's reply routing without ever owning the turn. // live turn's reply routing without ever owning the turn.
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller); Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, callerTerminal);
// wait defaults to true (block for the reply); wait:false is fire-and-poll. // wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait")) return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, target, content, onAccepted, workers.profiles()) ? sendAsync(messages, target, content, onAccepted, workers.profiles(), caller)
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles()); : send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles(), callerOwner);
}; };
// fleet_reply's identity is the CONNECTION, never an argument — so the authz check // fleet_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself. // is "is this caller a worker at all", and it can only ever reply as itself.
@@ -514,7 +516,7 @@ public final class FleetMcp {
(exchange, req) -> { (exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_status", req.arguments()), null); McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_status", req.arguments()), null);
if (denied != null) return denied; if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId")); return status(messages, str(req.arguments(), "sessionId"), principal(exchange).ownerKey());
}; };
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> pollHandler = BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> pollHandler =
(exchange, req) -> { (exchange, req) -> {
@@ -524,7 +526,8 @@ public final class FleetMcp {
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction. // The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target); McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
if (denied != null) return denied; if (denied != null) return denied;
return poll(messages, leadChannel, str(a, "ticket"), target, coordId); return poll(messages, leadChannel, str(a, "ticket"), target, coordId,
principal(exchange).ownerKey());
}; };
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox). // CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read. // Acking removes a reply from the inbox, so it is a drain, not a read.
@@ -562,7 +565,8 @@ public final class FleetMcp {
callerTerminal(exchange), callerTerminal(exchange),
callers.collaborators(), collaboratorsVisibleTo(principal(exchange)), callers.collaborators(), collaboratorsVisibleTo(principal(exchange)),
new CoordinationSource(leadChannel, peers), new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange))); coordinatorVisibleTo(principal(exchange)),
leadsVisibleTo(principal(exchange)), membersVisibleTo(principal(exchange)));
}; };
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler = BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
(exchange, req) -> { (exchange, req) -> {
@@ -759,7 +763,38 @@ public final class FleetMcp {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator(); return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
} }
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */ /**
* Who may see {@code fleet_list}'s {@code leads} array — exactly the roles that may
* {@link Authz.Action#SEND} to a lead: the primary, an architect, and a collaborator. A
* collaborator's own {@code fleet_whoami} carries no lead address, and {@code leads} is the
* only place this tool gives one, so a collaborator needs this array to use the send it
* already holds. A worker can never {@code SEND} at all, so it still sees neither this array
* nor {@code members}; a worker's own facts come from {@code fleet_whoami} instead. Split
* out for the same reason as {@link #coordinatorVisibleTo} and
* {@link #collaboratorsVisibleTo}: the decision must be unit-testable without fabricating an
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather
* than inlining the check.
*/
static boolean leadsVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
}
/**
* Who may see {@code fleet_list}'s {@code members} array — the primary and an architect,
* which may {@link Authz.Action#SEND} to a member. A worker can never {@code SEND} at all,
* and a collaborator may {@code SEND} only to a lead or another collaborator, never to a
* spawned member, so neither sees this array even though {@link #leadsVisibleTo} grants a
* collaborator the sibling one.
*/
static boolean membersVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect();
}
/**
* The terminal of the caller on this call's connection, or {@code null} when that caller carries
* no terminal, which is only the unnamed primary. A named lead, an architect and a worker each
* carry one.
*/
private static String callerTerminal(McpSyncServerExchange exchange) { private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL); Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString(); String s = v == null ? null : v.toString();
@@ -801,9 +836,14 @@ public final class FleetMcp {
return identity; return identity;
} }
/** Mark a connected spawned member available for the injector readiness gate. */ /**
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) { * Mark a caller present for the injector readiness gate, when its deliverability depends on
if (caller.isSpawnedMember()) { * proving a live MCP contact: a worker, an architect, or the unconfigured-pane floor. A lead
* or a collaborator is excluded — each is already deliverable through its own named-registry
* entry.
*/
static void markTrackedCallerPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember() || caller.isObserver()) {
presence.markPresent(caller.terminal()); presence.markPresent(caller.terminal());
} }
} }
@@ -859,8 +899,8 @@ public final class FleetMcp {
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package * fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with. * reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact * {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText} * instance and waits for the real continuation to end the old pane, relaunch a fresh one, and
* through the real {@code router.leadAgents()}. * send {@code bootstrapText} through the real {@code router.leadAgents()}.
*/ */
public LeadRollover leadRollover() { public LeadRollover leadRollover() {
return leadRollover; return leadRollover;
@@ -876,7 +916,8 @@ public final class FleetMcp {
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording. * a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/ */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted, Set<String> profiles) { Long timeoutMs, Runnable onAccepted, Set<String> profiles,
String callerOwner) {
if (isBlank(sessionId) || isBlank(content)) { if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required"); return error("sessionId and content are required");
} }
@@ -886,7 +927,7 @@ public final class FleetMcp {
} }
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs); long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try { try {
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout); return formatReply(messages.send(sessionId, content, timeout, onAccepted, callerOwner), timeout);
} catch (HerdrException e) { } catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage()); return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
} }
@@ -896,13 +937,15 @@ public final class FleetMcp {
* {@code fleet_send} carrying a {@code turnId}: the primary's answer to a worker's * {@code fleet_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as * {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send. * it resumes the same turn — surfaced to the primary identically to a normal send.
* {@code callerOwner} must match the turn's recorded owner or this is refused.
*/ */
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) { static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs,
String callerOwner) {
if (isBlank(turnId) || isBlank(content)) { if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question"); return error("turnId and content are required to answer a worker's question");
} }
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs); long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout); return formatReply(messages.answer(turnId, content, timeout, callerOwner), timeout);
} }
/** /**
@@ -951,6 +994,8 @@ public final class FleetMcp {
+ "\" and content set to your answer; the worker resumes the same turn."); + "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already " case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)"); + "answered (turnId stale)");
case NOT_TURN_OWNER -> error("this turn belongs to a different delegation — only the caller "
+ "that opened it may answer it");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker " case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]"); + r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
// fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call, // fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call,
@@ -972,6 +1017,16 @@ public final class FleetMcp {
*/ */
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content, static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles) { Runnable onAccepted, Set<String> profiles) {
return sendAsync(messages, sessionId, content, onAccepted, profiles, null);
}
/**
* As above, recording {@code creator}'s owner key so a later {@code fleet_poll{ticket}} only
* hands the result back to the same caller — see {@link MessageService#poll(String, String)}.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles,
Principal creator) {
if (isBlank(sessionId) || isBlank(content)) { if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required"); return error("sessionId and content are required");
} }
@@ -979,7 +1034,7 @@ public final class FleetMcp {
if (targetError != null) { if (targetError != null) {
return targetError; return targetError;
} }
String ticket = messages.sendAsync(sessionId, content, onAccepted); String ticket = messages.sendAsync(sessionId, content, onAccepted, creator);
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket); return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket);
} }
@@ -1158,9 +1213,23 @@ public final class FleetMcp {
* exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this * exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this
* is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently * is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently
* ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches. * ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches.
*
* <p>Does not check who owns {@code ticket} — see the overload that takes {@code
* callerTerminal} for that. Callers that do not resolve a caller terminal (tests, or a surface
* with no connection-based identity) use this one.
*/ */
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket, static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId) { String target, String coordId) {
return poll(messages, leadChannel, ticket, target, coordId, null);
}
/**
* As above, refusing a ticket lookup whose caller owner key differs from the key that created it
* — see {@link MessageService#poll(String, String)}. {@code callerOwner} comes from the calling
* connection's resolved principal, never a client-supplied value.
*/
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId, String callerOwner) {
if (!isBlank(coordId)) { if (!isBlank(coordId)) {
return pollHeldPeerMail(leadChannel, coordId); return pollHeldPeerMail(leadChannel, coordId);
} }
@@ -1174,7 +1243,7 @@ public final class FleetMcp {
if (isBlank(ticket)) { if (isBlank(ticket)) {
return error("ticket (or target) is required"); return error("ticket (or target) is required");
} }
MessageService.TaskView v = messages.poll(ticket); MessageService.TaskView v = messages.poll(ticket, callerOwner);
if (v == null) { if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)"); return error("unknown ticket: " + ticket + " (never issued, or expired)");
} }
@@ -1281,16 +1350,19 @@ public final class FleetMcp {
/** /**
* {@code fleet_status}: the live lifecycle status of a worker session, plus — when the worker * {@code fleet_status}: the live lifecycle status of a worker session, plus — when the worker
* is paused mid-turn in an async {@code fleet_ask} (CB-582) — the open question and how to * is paused mid-turn in an async {@code fleet_ask} — the open question and how to answer it, so
* answer it, so a lead on its normal poll cadence does not need the ticket to notice. * a lead on its normal poll cadence does not need the ticket to notice. The question, its
* {@code turnId} and its ticket id are shown only to the caller whose owner key created that
* delegation, or to the unnamed primary; any other caller still sees the base status.
* {@code callerOwner} comes from the calling connection's resolved principal.
*/ */
static McpSchema.CallToolResult status(MessageService messages, String sessionId) { static McpSchema.CallToolResult status(MessageService messages, String sessionId, String callerOwner) {
if (isBlank(sessionId)) { if (isBlank(sessionId)) {
return error("sessionId is required"); return error("sessionId is required");
} }
try { try {
String base = messages.status(sessionId).name().toLowerCase(); String base = messages.status(sessionId).name().toLowerCase();
MessageService.PendingAsk ask = messages.pendingAsk(sessionId); MessageService.PendingAsk ask = messages.pendingAsk(sessionId, callerOwner);
if (ask == null) { if (ask == null) {
return text(base); return text(base);
} }
@@ -1345,6 +1417,14 @@ public final class FleetMcp {
} }
return text(json(m)); return text(json(m));
} }
if (caller.isObserver()) {
// No architect slot, collaborator name, or lead name to report — only the pane itself,
// so a peer that already knows this terminal can still address it.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) { if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately // CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a // still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
@@ -1425,7 +1505,7 @@ public final class FleetMcp {
} }
if (isBlank(callerTerminal)) { if (isBlank(callerTerminal)) {
// An unnamed primary (token/loopback path, no resolved pane) has nowhere for the // An unnamed primary (token/loopback path, no resolved pane) has nowhere for the
// eventual /clear + bootstrap to land — LeadRollover#open would throw // eventual relaunch + bootstrap to land — LeadRollover#open would throw
// IllegalArgumentException for the same reason; refuse cleanly here instead. // IllegalArgumentException for the same reason; refuse cleanly here instead.
return error("fleet_handover requires a named lead pane (a resolved connection terminal) " return error("fleet_handover requires a named lead pane (a resolved connection terminal) "
+ "to open a rollover request against — an unnamed primary has none"); + "to open a rollover request against — an unnamed primary has none");
@@ -1778,7 +1858,7 @@ public final class FleetMcp {
CoordinationSource coordination) { CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false); Map.of(), false, coordination, false, true, true);
} }
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */ /** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1788,7 +1868,7 @@ public final class FleetMcp {
Map<String, String> leads, String selfTerm) { Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage, return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, CoordinationSource.none(), false); Map.of(), false, CoordinationSource.none(), false, true, true);
} }
/** /**
@@ -1806,7 +1886,7 @@ public final class FleetMcp {
CoordinationSource coordination) { CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(), return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false); Map.of(), false, coordination, false, true, true);
} }
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */ /** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1816,7 +1896,7 @@ public final class FleetMcp {
Map<String, String> leads, String selfTerm, CoordinationSource coordination) { Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage, return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false); Map.of(), false, coordination, false, true, true);
} }
/** /**
@@ -1839,7 +1919,7 @@ public final class FleetMcp {
CoordinationSource coordination) { CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage, return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false); Map.of(), false, coordination, false, true, true);
} }
/** /**
@@ -1856,15 +1936,23 @@ public final class FleetMcp {
* {@code false} (fleetd #463: a forgotten argument fails closed, not * {@code false} (fleetd #463: a forgotten argument fails closed, not
* open), so a test that wants the {@code coordinator} row must pass * open), so a test that wants the {@code coordinator} row must pass
* an explicit {@code true} * an explicit {@code true}
* @param leadsVisible whether this caller may see the {@code leads} array (see
* {@link #leadsVisibleTo}) — unlike {@code callerIsPrimary}, this is
* also {@code true} for an architect or a collaborator, so it cannot
* be derived from {@code callerIsPrimary} alone
* @param membersVisible whether this caller may see the {@code members} array (see
* {@link #membersVisibleTo}); also {@code true} for an architect, but
* unlike {@code leadsVisible}, never for a collaborator
*/ */
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages, static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage, QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm, LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) { CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, callerIsPrimary); Map.of(), false, coordination, callerIsPrimary, leadsVisible, membersVisible);
} }
/** /**
@@ -1881,6 +1969,10 @@ public final class FleetMcp {
* {@link #collaboratorsVisibleTo}); every wrapper overload above * {@link #collaboratorsVisibleTo}); every wrapper overload above
* passes {@code false}, so a test that wants the row must call this * passes {@code false}, so a test that wants the row must call this
* overload with an explicit {@code true} * overload with an explicit {@code true}
* @param leadsVisible whether this caller may see the {@code leads} array (see
* {@link #leadsVisibleTo})
* @param membersVisible whether this caller may see the {@code members} array (see
* {@link #membersVisibleTo})
*/ */
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages, static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage, CapacitySource capacity, HealthCoverageSource healthCoverage,
@@ -1890,28 +1982,41 @@ public final class FleetMcp {
LeadConfigDirSource leadConfigDirs, LeadConfigDirSource leadConfigDirs,
Map<String, String> leads, String selfTerm, Map<String, String> leads, String selfTerm,
Map<String, String> collaborators, boolean collaboratorsVisible, Map<String, String> collaborators, boolean collaboratorsVisible,
CoordinationSource coordination, boolean callerIsPrimary) { CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible) {
try { try {
Map<String, Agent> live = workers.list().stream() // Neither row's assembly (leadView/memberCapacityView probing herdr for live status)
.map(Agent.class::cast) // runs unless at least one of them needs the live-agent lookup backing it.
.filter(a -> a.terminalId() != null) Map<String, Agent> live = (leadsVisible || membersVisible)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b)); ? workers.list().stream()
List<Map<String, Object>> leadRows = leads.entrySet().stream() .map(Agent.class::cast)
.sorted(Map.Entry.comparingByValue()) .filter(a -> a.terminalId() != null)
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge, .collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b))
leadConfigDirs)) : Map.of();
.toList();
// fleetd #209: this is the caller-driven fleet_list read that actually reports // fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the // agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
// resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster(). // resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster().
List<MemberSession> roster = sessions.rosterResolved(); List<MemberSession> roster = sessions.rosterResolved();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get()); Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
roster.stream().map(MemberSession::profile).forEach(profiles::add); roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>(); Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out); // READ is permission to enter this tool, not permission to receive every field it can
// build -- gate BEFORE assembling each row, so the key is absent rather than
// present-and-empty; a caller without either row gets its own facts from fleet_whoami.
if (leadsVisible) {
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
result.put("leads", leadRows);
}
if (membersVisible) {
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
result.put("members", out);
}
result.put("healthCoverage", healthCoverage.value().get()); result.put("healthCoverage", healthCoverage.value().get());
result.put("loopHealth", Map.of( result.put("loopHealth", Map.of(
"statusPoller", loopHealth.statusPoller().get().name(), "statusPoller", loopHealth.statusPoller().get().name(),
@@ -2445,7 +2550,14 @@ public final class FleetMcp {
private static McpSchema.Tool listTool() { private static McpSchema.Tool listTool() {
return tool(FleetTool.LIST.wireName(), return tool(FleetTool.LIST.wireName(),
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other " "List the whole fleet the bridge tracks, in two parts. 'members' is visible to "
+ "the primary and an architect only. 'leads' is visible to those two AND a "
+ "collaborator — exactly the roles that may fleet_send to a lead, so a "
+ "collaborator can learn a lead's sessionId before using the send it already "
+ "holds. A worker holds READ to call this tool at all, but gets neither "
+ "array, never an empty one; a worker reads its own session, "
+ "profile, state, worktree, branch and owner from fleet_whoami instead. "
+ "'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), " + "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "name, live status, and 'self': true on your own row; this is how you " + "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the " + "discover a peer lead without being told its address. 'members' are the "
@@ -2531,17 +2643,20 @@ public final class FleetMcp {
private static McpSchema.Tool handoverTool() { private static McpSchema.Tool handoverTool() {
return tool(FleetTool.HANDOVER.wireName(), return tool(FleetTool.HANDOVER.wireName(),
"Replace your OWN lead session once its context is full: write a handover file, " "Replace your OWN lead session once its context is full: write a handover file, "
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead " + "then use this to have fleetd end your pane's process and relaunch a fresh "
+ "session against it. Four actions: 'open' (requests a token and the " + "lead session bootstrapped against it. Four actions: 'open' (requests a "
+ "handoverPath you must write the handover file to before confirming), " + "token and the handoverPath you must write the handover file to before "
+ "'confirm' (validates every gate and — only if every one passes — schedules " + "confirming), 'confirm' (validates every gate and — only if every one "
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's " + "passes — schedules the roll; it does NOT itself end your pane, the roll "
+ "own turn ends), 'cancel' (drops a pending request without rolling), and " + "runs once this call's own turn ends), 'cancel' (drops a pending request "
+ "'status' (read-only: what happened to a token after 'confirm' — still " + "without rolling), and 'status' (read-only: what happened to a token after "
+ "running (approved but not finished yet), the roll completed, the calling " + "'confirm' — still running (approved but not finished yet), the roll "
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or " + "completed, the calling turn never settled within turnSettleSeconds so "
+ "/clear itself never settled so bootstrapText was never sent; never " + "nothing was touched, the old pane never confirmed dead so no relaunch was "
+ "schedules, cancels or retries anything). Primary-only. " + "attempted, the relaunch itself failed, the fresh pane never became ready "
+ "so bootstrapText was never sent, or the fresh pane became ready and was "
+ "bootstrapped but was never recognised as a live lead; never schedules, "
+ "cancels or retries anything). Primary-only. "
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane " + "There is deliberately no terminal/session/leadTerminal parameter: the pane "
+ "to roll is always resolved from YOUR OWN connection, never a value you " + "to roll is always resolved from YOUR OWN connection, never a value you "
+ "pass, so you can only ever roll yourself — never another lead. Requires " + "pass, so you can only ever roll yourself — never another lead. Requires "
@@ -6,6 +6,7 @@ import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus; import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient; import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException; import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.Tab; import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace; import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl; import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -67,16 +68,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class); private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents; private final AgentControl agents;
private final WorkspaceControl spaces; private final WorkspaceControl spaces;
@@ -766,90 +757,26 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so // Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed. // argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size()); List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
checkPaneCommandFits(cfg, argv); try {
HerdrException last = null; ResilientAgentLaunch.checkFits(cfg.profile(), argv);
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) { } catch (ResilientAgentLaunch.TooLargeException e) {
long seq = nameSeq.incrementAndGet(); throw new PeerUnreachableException(e.getMessage());
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
} }
throw last; long[] lastSeq = {0};
} Agent agent = ResilientAgentLaunch.startUniquelyNamed(agents, namePrefix, args, paneId,
attempt -> {
/** lastSeq[0] = nameSeq.incrementAndGet();
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line, return namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + lastSeq[0];
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code },
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
* started", the backend exits on the mangled argument it was handed, the pane closes, and the return new Started(agent, lastSeq[0]);
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
* would have hit anyway, named at the point it is still
* explainable
*/
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new PeerUnreachableException(
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
} }
/** /**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a * The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}. * fleetd choice and not configurable — see {@link ResilientAgentLaunch#checkFits}.
*/ */
static final int PANE_COMMAND_BYTE_LIMIT = 1024; static final int PANE_COMMAND_BYTE_LIMIT = ResilientAgentLaunch.PANE_COMMAND_BYTE_LIMIT;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ---------------------------------------------------------------------- // --- discovery + reap ----------------------------------------------------------------------
@@ -1,5 +1,6 @@
package dev.ltms.fleet.msg; package dev.ltms.fleet.msg;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus; import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter; import dev.ltms.fleet.herdr.HerdrRouter;
@@ -87,7 +88,7 @@ public final class MessageService {
/** /**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the * The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with * question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes. * {@link #answer(String, String, long, String)} and the turn resumes.
*/ */
QUESTION, QUESTION,
/** Timed out after the message was delivered — the worker is still working. */ /** Timed out after the message was delivered — the worker is still working. */
@@ -120,10 +121,16 @@ public final class MessageService {
/** Another send to this session was in flight for the whole window. */ /** Another send to this session was in flight for the whole window. */
BUSY, BUSY,
/** /**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no * An answer ({@link #answer(String, String, long, String)}) referenced a {@code turnId}
* longer open — the worker's {@code fleet_ask} already timed out or was answered. * that is no longer open — the worker's {@code fleet_ask} already timed out or was answered.
*/ */
STALE_TURN STALE_TURN,
/**
* An answer ({@link #answer(String, String, long, String)}) named a {@code turnId} that is
* still open, but the answering caller is not the caller whose accepted delegation opened
* it. Distinct from {@link #STALE_TURN} so a refusal is never reported as a lapsed turn.
*/
NOT_TURN_OWNER
} }
/** /**
@@ -133,7 +140,7 @@ public final class MessageService {
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION}, * {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null} * else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via * @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null} * {@link #answer(String, String, long, String)}), else {@code null}
*/ */
public record Reply(Outcome outcome, String text, String turnId) { public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */ /** A reply with no correlation id (the common terminal outcomes). */
@@ -275,10 +282,16 @@ public final class MessageService {
* {@link #abandon}) can never match again regardless of this flag's value. * {@link #abandon}) can never match again regardless of this flag's value.
*/ */
private volatile boolean askTimedOut; private volatile boolean askTimedOut;
/**
* The owner key of the caller whose {@code fleet_send{wait:false}} created this ticket, or
* {@code null} for the unnamed primary and overloads that do not record a caller.
*/
private final String creatorOwner;
private Task(String ticket, String target, LongSupplier nowNanos) { private Task(String ticket, String target, LongSupplier nowNanos, String creatorOwner) {
this.ticket = ticket; this.ticket = ticket;
this.target = target; this.target = target;
this.creatorOwner = creatorOwner;
this.createdNanos = nowNanos.getAsLong(); this.createdNanos = nowNanos.getAsLong();
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong()); future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
} }
@@ -340,6 +353,14 @@ public final class MessageService {
*/ */
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>(); private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong(); private final AtomicLong ticketSeq = new AtomicLong();
/**
* Minted once per {@code MessageService} instance and folded into every ticket id (see
* {@link #sendAsync(String, String, Runnable, Principal)}). {@link #ticketSeq} alone restarts at
* zero for every instance, so without this a ticket id can be reused across instances and
* resolve to an unrelated {@link Task} with no error; this nonce makes that impossible, because
* an id minted by one instance can never match the id space of another.
*/
private final String ticketBootNonce = UUID.randomUUID().toString().substring(0, 6);
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor( private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory()); Thread.ofVirtual().name("bridge-async-", 0).factory());
@@ -656,7 +677,7 @@ public final class MessageService {
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout"; case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout";
case WORKER_FAILED -> "failed"; case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted"; case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation case STALE_TURN, QUESTION, NOT_TURN_OWNER -> null; // not a completed delegation
}; };
} }
@@ -915,14 +936,17 @@ public final class MessageService {
/** /**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the * Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses. * worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses. {@code callerOwner}
* identifies the caller making this call and is recorded as the turn's owner. It is the only
* caller {@link #answer(String, String, long, String)} will
* later accept an answer from if the worker pauses mid-turn to ask.
*/ */
public Reply send(String target, String content, long timeoutMillis) { public Reply send(String target, String content, long timeoutMillis, String callerOwner) {
return send(target, content, timeoutMillis, null); return send(target, content, timeoutMillis, null, callerOwner);
} }
/** /**
* As {@link #send(String, String, long)}, but with an accepted-delivery hook. * As {@link #send(String, String, long, String)}, but with an accepted-delivery hook.
* *
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send * <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is * lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
@@ -933,12 +957,13 @@ public final class MessageService {
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it * acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook. * never earned. {@code null} disables the hook.
*/ */
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) { public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, String callerOwner) {
return send(target, content, timeoutMillis, onAccepted, null); return send(target, content, timeoutMillis, onAccepted, null, callerOwner);
} }
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */ /** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) { private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task,
String callerOwner) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L; long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock()); ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -958,7 +983,7 @@ public final class MessageService {
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an // open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the // enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind. // failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target); CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target, Rendezvous.Owner.of(callerOwner));
// CB-640: this send now owns target's delivery, so any earlier stranded-reply or // CB-640: this send now owns target's delivery, so any earlier stranded-reply or
// still-queued fact no longer describes the live state — clear both rather than let // still-queued fact no longer describes the live state — clear both rather than let
// them outlive the send that supersedes them. // them outlive the send that supersedes them.
@@ -1137,19 +1162,30 @@ public final class MessageService {
* mid-turn (already picked up), so the answer flows back through its own open {@code fleet_ask} * mid-turn (already picked up), so the answer flows back through its own open {@code fleet_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker * call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost. * is unblocked so a reply that lands the instant it resumes is not lost.
*
* <p>{@code callerOwner} identifies the caller making this call. It is checked against the
* turn's recorded owner (the caller whose
* accepted delegation opened it, see {@link #send(String, String, long, String)} and
* {@link #sendAsync(String, String, Runnable, Principal)}) before anything else runs: a mismatch,
* including a turn with no owner on record at all, returns {@link Outcome#NOT_TURN_OWNER}
* without touching the rendezvous, the session lock, or any async task bookkeeping.
*/ */
public Reply answer(String turnId, String content, long timeoutMillis) { public Reply answer(String turnId, String content, long timeoutMillis, String callerOwner) {
String workerSession = rendezvous.askSession(turnId); String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) { if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered) return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
} }
Rendezvous.Owner owner = rendezvous.askOwner(turnId);
if (!Rendezvous.Owner.permits(owner, callerOwner)) {
return new Reply(Outcome.NOT_TURN_OWNER, null);
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L; long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock()); ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) { if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null); return new Reply(Outcome.BUSY, null);
} }
try { try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession); CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession, owner);
// fleetd #575: this try used to open below, AFTER the Task lookup/registration and the // fleetd #575: this try used to open below, AFTER the Task lookup/registration and the
// STALE_TURN early return that follows it — so that return was covered only by a // STALE_TURN early return that follows it — so that return was covered only by a
// hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the // hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the
@@ -1279,8 +1315,20 @@ public final class MessageService {
* @return the ticket to poll for the eventual result * @return the ticket to poll for the eventual result
*/ */
public String sendAsync(String target, String content, Runnable onAccepted) { public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet(); return sendAsync(target, content, onAccepted, null);
Task task = new Task(ticket, target, nowNanos); }
/**
* As {@link #sendAsync(String, String, Runnable)}, recording {@code creator}'s owner key as this
* ticket's owner. The key is derived here from the resolved principal so callers cannot pass a
* terminal address where an owner identity is required.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted, Principal creator) {
String ticket = "task-" + ticketBootNonce + "-" + ticketSeq.incrementAndGet();
String creatorOwner = creator == null ? null : creator.ownerKey();
Task task = new Task(ticket, target, nowNanos, creatorOwner);
tasks.put(ticket, task); tasks.put(ticket, task);
if (pushLoop != null) { if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a // CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
@@ -1311,7 +1359,7 @@ public final class MessageService {
} }
asyncExecutor.submit(() -> { asyncExecutor.submit(() -> {
try { try {
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task); Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task, creatorOwner);
if (result.outcome() == Outcome.QUESTION) { if (result.outcome() == Outcome.QUESTION) {
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run // Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread. // just after resolveQuestion wakes this thread.
@@ -1343,15 +1391,31 @@ public final class MessageService {
} }
/** /**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket; * As {@link #poll(String, String)}, with no caller owner key — the unnamed primary's ticket rule
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a * checked, so this overload must only be used where the caller's identity is otherwise
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason. * irrelevant.
*/ */
public TaskView poll(String ticket) { public TaskView poll(String ticket) {
return poll(ticket, null);
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket.
* Refuses a {@code callerOwner} that differs from the owner that created the ticket (see
* {@link #sendAsync(String, String, Runnable, Principal)}) with a {@link Phase#FAILED} view that
* carries no reply text. The unnamed primary has a {@code null} owner key and is never refused.
* Otherwise returns a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket, String callerOwner) {
Task task = tasks.get(ticket); Task task = tasks.get(ticket);
if (task == null) { if (task == null) {
return null; return null;
} }
if (!ownsTicket(task, callerOwner)) {
return new TaskView(ticket, Phase.FAILED, null, null,
"forbidden: this ticket was created by a different session", null);
}
CompletableFuture<Reply> f = task.future; CompletableFuture<Reply> f = task.future;
if (!f.isDone()) { if (!f.isDone()) {
Reply question = task.question; Reply question = task.question;
@@ -1387,6 +1451,16 @@ public final class MessageService {
return new TaskView(ticket, Phase.FAILED, null, null, detail, null); return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
} }
/**
* Whether {@code callerOwner} may read {@code task}'s state. A {@code null} caller key is the
* unnamed primary and may read every ticket. Other callers must match the task's owner key. This
* differs from {@link Rendezvous.Owner#permits}: a missing rendezvous owner is not an authenticated
* unnamed primary, so that gate refuses every caller when no owner was recorded.
*/
private static boolean ownsTicket(Task task, String callerOwner) {
return callerOwner == null || callerOwner.equals(task.creatorOwner);
}
/** /**
* Test seam only — carries no production behaviour, and nothing in this class calls it; * Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly. * {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
@@ -1706,16 +1780,17 @@ public final class MessageService {
} }
/** /**
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any * The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any —
* (CB-582) — {@code fleet_status} uses this to show a pending question without the caller * {@code fleet_status} uses this to show a pending question without the caller needing the
* needing the ticket. {@code null} when the session has no open async question (including a * ticket. {@code null} when the session has no open async question (including a session mid a
* session mid a <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see * <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}). * {@link PendingAsk}), or when {@code callerOwner} does not own the task the question
* belongs to (see {@link #ownsTicket(Task, String)}).
*/ */
public PendingAsk pendingAsk(String workerSession) { public PendingAsk pendingAsk(String workerSession, String callerOwner) {
for (Task task : tasks.values()) { for (Task task : tasks.values()) {
Reply q = task.question; Reply q = task.question;
if (q != null && workerSession.equals(task.target)) { if (q != null && workerSession.equals(task.target) && ownsTicket(task, callerOwner)) {
return new PendingAsk(task.ticket, q.text(), q.turnId()); return new PendingAsk(task.ticket, q.text(), q.turnId());
} }
} }
@@ -1,5 +1,6 @@
package dev.ltms.fleet.msg; package dev.ltms.fleet.msg;
import java.util.UUID;
import java.util.concurrent.CompletableFuture; import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap; import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong; import java.util.concurrent.atomic.AtomicLong;
@@ -60,7 +61,7 @@ public final class Rendezvous {
} }
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */ /** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
private record AskWaiter(String session, CompletableFuture<String> answer) { private record AskWaiter(String session, CompletableFuture<String> answer, Owner owner) {
} }
/** /**
@@ -70,11 +71,44 @@ public final class Rendezvous {
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) { public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
} }
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>(); /**
* The caller whose accepted delegation opened a turn — the only caller allowed to answer it.
* A {@code null} owner key means the unnamed primary.
*/
public record Owner(String ownerKey) {
public static final Owner UNNAMED_PRIMARY = new Owner(null);
public static Owner of(String ownerKey) {
return ownerKey == null ? UNNAMED_PRIMARY : new Owner(ownerKey);
}
/**
* Whether {@code callerOwner} matches {@code owner}. A {@code null} owner means no owner was
* recorded, so it matches no caller. {@link #UNNAMED_PRIMARY} records the unnamed primary
* with an owner object whose key is {@code null}.
*/
public static boolean permits(Owner owner, String callerOwner) {
return owner != null && java.util.Objects.equals(owner.ownerKey(), callerOwner);
}
}
/** A registered forward waiter together with the owner its delegation was opened under. */
private record ForwardWaiter(Owner owner, CompletableFuture<Resolution> future) {
}
private final ConcurrentHashMap<String, ForwardWaiter> waiters = new ConcurrentHashMap<>();
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */ /** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>(); private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong(); private final AtomicLong askSeq = new AtomicLong();
/**
* Minted once per {@code Rendezvous} instance and folded into every {@code turnId} (see
* {@link #openAsk(String)}). {@link #askSeq} alone restarts at zero for every instance, so
* without this a {@code turnId} minted by one instance could be minted again by another and
* resolve to an unrelated ask with no error; this nonce makes that impossible, because an id
* minted by one instance can never match the id space of another.
*/
private final String askBootNonce = UUID.randomUUID().toString().substring(0, 6);
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */ /** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>(); private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
@@ -89,8 +123,17 @@ public final class Rendezvous {
* code a double open is impossible; this is a tripwire for the day that no longer holds. * code a double open is impossible; this is a tripwire for the day that no longer holds.
*/ */
public CompletableFuture<Resolution> open(String session) { public CompletableFuture<Resolution> open(String session) {
return open(session, null);
}
/**
* Same as {@link #open(String)}, additionally recording {@code owner} as the caller whose
* delegation opened this waiter. A {@code null} owner records no owner at all — the
* fail-closed default {@link Owner#permits} refuses to everyone.
*/
public CompletableFuture<Resolution> open(String session, Owner owner) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>(); CompletableFuture<Resolution> waiter = new CompletableFuture<>();
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter); ForwardWaiter existing = waiters.putIfAbsent(session, new ForwardWaiter(owner, waiter));
if (existing != null) { if (existing != null) {
throw new IllegalStateException( throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered"); "rendezvous double-open for session " + session + " — a waiter is already registered");
@@ -105,7 +148,13 @@ public final class Rendezvous {
* successful {@code open} after a finished turn requires this close to have happened first). * successful {@code open} after a finished turn requires this close to have happened first).
*/ */
public void close(String session, CompletableFuture<Resolution> waiter) { public void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter); waiters.computeIfPresent(session, (s, w) -> w.future() == waiter ? null : w);
}
/** The owner recorded for {@code session}'s open waiter, or {@code null} if none is open. */
public Owner ownerOf(String session) {
ForwardWaiter w = waiters.get(session);
return w == null ? null : w.owner();
} }
/** Whether a send is currently awaiting a resolution for {@code session}. */ /** Whether a send is currently awaiting a resolution for {@code session}. */
@@ -119,7 +168,8 @@ public final class Rendezvous {
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire. * send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
*/ */
public CompletableFuture<Resolution> currentWaiter(String session) { public CompletableFuture<Resolution> currentWaiter(String session) {
return waiters.get(session); ForwardWaiter w = waiters.get(session);
return w == null ? null : w.future();
} }
/** /**
@@ -145,9 +195,9 @@ public final class Rendezvous {
while (true) { while (true) {
AskWaiter[] minted = { null }; AskWaiter[] minted = { null };
String turnId = openAsksBySession.computeIfAbsent(session, _ -> { String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
String newTurnId = session + "#" + askSeq.incrementAndGet(); String newTurnId = session + "#" + askBootNonce + "-" + askSeq.incrementAndGet();
CompletableFuture<String> answer = new CompletableFuture<>(); CompletableFuture<String> answer = new CompletableFuture<>();
AskWaiter waiter = new AskWaiter(session, answer); AskWaiter waiter = new AskWaiter(session, answer, ownerOf(session));
asks.put(newTurnId, waiter); asks.put(newTurnId, waiter);
minted[0] = waiter; minted[0] = waiter;
return newTurnId; return newTurnId;
@@ -183,6 +233,16 @@ public final class Rendezvous {
return w == null ? null : w.session(); return w == null ? null : w.session();
} }
/**
* The owner recorded for {@code turnId} when its ask turn was freshly opened — the caller
* whose delegation {@link #answerAsk} must match. {@code null} if {@code turnId} is unknown or
* lapsed, or if the ask opened with no forward waiter owner on record.
*/
public Owner askOwner(String turnId) {
AskWaiter w = asks.get(turnId);
return w == null ? null : w.owner();
}
/** /**
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it * Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
* to resume its turn. * to resume its turn.
@@ -245,7 +305,7 @@ public final class Rendezvous {
} }
private boolean complete(String session, Resolution resolution) { private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session); ForwardWaiter waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution); return waiter != null && waiter.future().complete(resolution);
} }
} }
@@ -663,6 +663,8 @@ public final class FleetApp {
if (!allow(ctx, routeAction("POST /sessions/{id}/message"), id)) { if (!allow(ctx, routeAction("POST /sessions/{id}/message"), id)) {
return; return;
} }
Principal caller = ctx.attribute(CALLER);
String callerOwner = caller == null ? null : caller.ownerKey();
JsonNode body; JsonNode body;
try { try {
body = mapper.readTree(ctx.body()); body = mapper.readTree(ctx.body());
@@ -688,19 +690,19 @@ public final class FleetApp {
// Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId. // Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) { if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout); writeReply(ctx, id, messages.answer(turnId, content, timeout, callerOwner), timeout);
return; return;
} }
if (!wait) { if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}. // Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content); String ticket = messages.sendAsync(id, content, null, caller);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted")); ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return; return;
} }
try { try {
writeReply(ctx, id, messages.send(id, content, timeout), timeout); writeReply(ctx, id, messages.send(id, content, timeout, callerOwner), timeout);
} catch (HerdrException e) { } catch (HerdrException e) {
herdrError(ctx, e); herdrError(ctx, e);
} }
@@ -719,6 +721,10 @@ public final class FleetApp {
case STALE_TURN -> ctx.status(409).json(Map.of( case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn", "sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)")); "detail", "that question is no longer open (timed out or already answered)"));
case NOT_TURN_OWNER -> ctx.status(403).json(Map.of(
"sessionId", id, "error", "not_turn_owner",
"detail", "this turn belongs to a different delegation — only the caller that "
+ "opened it may answer it"));
case REPLIED, COMPLETED_UNREPLIED -> { case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured fleet_reply from the CB-106 completion // replySource distinguishes a structured fleet_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying). // fallback (a scrape of the worker's transcript when it finished without replying).
@@ -733,9 +739,10 @@ public final class FleetApp {
// silent fall-through. That is exactly the bug this ticket exists to fix: // silent fall-through. That is exactly the bug this ticket exists to fix:
// `default -> "done"` used to sit here and would have told a REST caller the // `default -> "done"` used to sit here and would have told a REST caller the
// delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery // delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN can never // is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN and
// actually reach this inner switch — the outer switch above always dispatches // NOT_TURN_OWNER can never actually reach this inner switch — the outer switch
// them first — but they still need an arm to keep this switch exhaustive. // above always dispatches them first — but they still need an arm to keep this
// switch exhaustive.
"status", switch (reply.outcome()) { "status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working"; case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued"; case TIMED_OUT_QUEUED -> "queued";
@@ -745,7 +752,7 @@ public final class FleetApp {
case BUSY -> "busy"; case BUSY -> "busy";
case WORKER_FAILED -> "failed"; case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted"; case BACKEND_EXHAUSTED -> "backend_exhausted";
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN -> "done"; // unreachable case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN, NOT_TURN_OWNER -> "done"; // unreachable
}, },
"detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED "detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED
? "no reply within " + timeout + "ms; delivery is unconfirmed — the " ? "no reply within " + timeout + "ms; delivery is unconfirmed — the "
@@ -874,10 +881,12 @@ public final class FleetApp {
body.put("sessionId", id); body.put("sessionId", id);
body.put("status", messages.status(id).name().toLowerCase()); body.put("status", messages.status(id).name().toLowerCase());
body.put("ready", deliverable.test(id)); body.put("ready", deliverable.test(id));
// CB-582: a worker paused mid-turn in an async fleet_ask is otherwise invisible to a // A worker paused mid-turn in an async fleet_ask is otherwise invisible to a status
// status poll — surface the open question and how to answer it, same as fleet_poll's // poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view. // Phase.ASKING view, but only to the caller whose owner key created that delegation, or
MessageService.PendingAsk ask = messages.pendingAsk(id); // to the unnamed primary.
Principal caller = ctx.attribute(CALLER);
MessageService.PendingAsk ask = messages.pendingAsk(id, caller == null ? null : caller.ownerKey());
if (ask != null) { if (ask != null) {
body.put("question", ask.question()); body.put("question", ask.question());
body.put("turnId", ask.turnId()); body.put("turnId", ask.turnId());
@@ -894,7 +903,8 @@ public final class FleetApp {
if (!allow(ctx, routeAction("GET /tasks/{ticket}"), null)) { if (!allow(ctx, routeAction("GET /tasks/{ticket}"), null)) {
return; return;
} }
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket")); Principal caller = ctx.attribute(CALLER);
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.ownerKey());
if (v == null) { if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)")); ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return; return;
@@ -26,6 +26,7 @@ import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong; import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer; import java.util.function.Consumer;
import java.util.function.LongSupplier; import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** /**
* Authoritative in-daemon registry of the worker sessions this {@code fleetd} process spawned. * Authoritative in-daemon registry of the worker sessions this {@code fleetd} process spawned.
@@ -57,6 +58,18 @@ public final class SessionManager implements TurnListener {
* Populated on every spawn path, removed on {@link #release}. * Populated on every spawn path, removed on {@link #release}.
*/ */
private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>(); private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>();
/**
* fleetd #702: a pane mid-teardown, keyed by paneId, held from just before its registry entry
* is removed until {@link #releaseRemoved} finishes. {@link #spawnedMemberRole} consults this
* alongside the registry, so a caller resolving the pane's terminal during that window still
* sees a live member and never falls through to a tab map.
*
* <p>Depth-counted rather than a plain set: two threads can be tearing down the same pane at
* once (the CAS in {@link #releaseIfCurrent} exists for exactly that race), and with a set the
* loser's {@code finally} would unmark the pane while the winner is still mid-teardown,
* reopening the window this exists to close.
*/
private final ConcurrentHashMap<String /*paneId*/, Releasing> releasing = new ConcurrentHashMap<>();
private final MemberPresence presence; private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom(); private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong(); private final AtomicLong nonceSeq = new AtomicLong();
@@ -242,6 +255,10 @@ public final class SessionManager implements TurnListener {
handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0, handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0,
MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId()); MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId());
registry.put(handle.id(), session); registry.put(handle.id(), session);
// A presence contact that already arrived for this terminal found no registry
// entry to transition and gave up silently. Retry it now that one exists; remove
// this call and such a session stays in SPAWNING even though it is present.
reconcilePresence(handle.terminalId());
handles.put(handle.id(), handle); handles.put(handle.id(), handle);
log.debug("acquired session id={} terminal={} profile={} owner={}", log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal()); handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
@@ -303,9 +320,11 @@ public final class SessionManager implements TurnListener {
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work. * is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/ */
private MemberSession release(String paneId, ReleaseCause cause) { private MemberSession release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId); return releaseWindow(paneId, registry.get(paneId), () -> {
releaseRemoved(paneId, removed, handles.remove(paneId), cause); MemberSession removed = registry.remove(paneId);
return removed; releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
});
} }
/** /**
@@ -314,16 +333,85 @@ public final class SessionManager implements TurnListener {
* DONE record from stopping a worker that delivery has made BUSY. * DONE record from stopping a worker that delivery has made BUSY.
*/ */
private boolean releaseIfCurrent(MemberSession expected, ReleaseCause cause) { private boolean releaseIfCurrent(MemberSession expected, ReleaseCause cause) {
if (!registry.remove(expected.paneId(), expected)) { return releaseWindow(expected.paneId(), expected, () -> {
// A lifecycle transition replaced the record between the caller's check and this remove. if (!registry.remove(expected.paneId(), expected)) {
// Log it: this race is by definition unobservable otherwise, and a reaper that silently // A lifecycle transition replaced the record between the caller's check and this
// declines to reap is the hardest kind of behaviour to diagnose after the fact. // remove. Log it: this race is by definition unobservable otherwise, and a reaper
log.debug("skipping reap of pane={}: its registry record changed after the idle check " // that silently declines to reap is the hardest kind of behaviour to diagnose
+ "(most likely a delivery made it BUSY)", expected.paneId()); // after the fact.
return false; log.debug("skipping reap of pane={}: its registry record changed after the idle "
+ "check (most likely a delivery made it BUSY)", expected.paneId());
return false;
}
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
return true;
});
}
/**
* fleetd #702: mark {@code paneId} as mid-teardown — using {@code known}'s terminal/role when
* it is available — for the whole of {@code teardown}, which removes the registry entry and
* then runs {@link #releaseRemoved}. Shared by both registry-removal sites ({@link #release}'s
* unconditional remove and {@link #releaseIfCurrent}'s CAS remove) so neither can leave the
* other's window unmarked.
*
* <p>The mark is written before {@code teardown} runs — so it covers the removal itself, not
* only what comes after it — and cleared in a {@code finally}, so an unchecked throw out of
* {@code teardown} (including one from {@link PeerLauncher#stop}, which declares nothing) can
* never leave the pane marked for the rest of the daemon's life.
*/
private <T> T releaseWindow(String paneId, MemberSession known, Supplier<T> teardown) {
releasing.compute(paneId, (_, prior) -> Releasing.enter(prior, known));
try {
return teardown.get();
} finally {
releasing.compute(paneId, (_, prior) -> prior == null ? null : prior.leave());
} }
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause); }
return true;
/**
* Depth count plus the terminal/role a mid-teardown pane belongs to, for
* {@link #spawnedMemberRole}. The terminal/role come from whichever call into
* {@link #releaseWindow} first knew them: a call that finds the registry entry already gone
* passes a {@code null} session, and must not blank out what the first call recorded.
*/
record Releasing(int depth, String terminalId, MemberRole role) {
static Releasing enter(Releasing prior, MemberSession known) {
int depth = (prior == null ? 0 : prior.depth()) + 1;
String terminalId = known != null ? known.terminalId() : prior == null ? null : prior.terminalId();
MemberRole role = known != null ? known.role() : prior == null ? null : prior.role();
return new Releasing(depth, terminalId, role);
}
Releasing leave() {
return depth <= 1 ? null : new Releasing(depth - 1, terminalId, role);
}
}
/**
* The role of the live spawned member occupying {@code terminal} — whether it is currently in
* the registry, or mid-teardown between {@link #release} removing its registry entry and
* {@link #releaseRemoved} actually stopping its pane (fleetd #702). {@code null} for a terminal
* that is neither: this method is the one reader a caller resolver consults before any tab
* map, so a live or releasing member's identity never falls back to a tab label.
*
* <p>Checks the registry directly via {@link #findByTerminal} rather than {@link #roster()},
* so this hot-path lookup (consulted on every resolve) never pays for a list copy or a stream.
*/
public MemberRole spawnedMemberRole(String terminal) {
MemberSession session = findByTerminal(terminal);
if (session != null) {
return session.role();
}
if (terminal == null) {
return null;
}
for (Releasing r : releasing.values()) {
if (terminal.equals(r.terminalId())) {
return r.role();
}
}
return null;
} }
private void releaseRemoved(String paneId, MemberSession removed, PeerHandle removedHandle, private void releaseRemoved(String paneId, MemberSession removed, PeerHandle removedHandle,
@@ -382,6 +470,12 @@ public final class SessionManager implements TurnListener {
MemberSession resolved = resolveAgentSessionId(removed, removedHandle); MemberSession resolved = resolveAgentSessionId(removed, removedHandle);
notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(), notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(),
resolved.branch(), snapshotRef, resolved.agentSessionId())); resolved.branch(), snapshotRef, resolved.agentSessionId()));
String terminal = removed.terminalId();
if (terminal != null && !terminal.isBlank()) {
// Without this, a terminal stays marked present after its pane is gone, so a
// later send to the same id would read as deliverable instead of refused.
presence.forget(terminal);
}
} }
} }
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed // CB-581: the pane must always stop, even if the dirty check above threw. A session removed
@@ -725,6 +819,10 @@ public final class SessionManager implements TurnListener {
handle.charterReceipt(), handle.charterReceipt(),
handle.agentSessionId()); handle.agentSessionId());
registry.put(handle.id(), session); registry.put(handle.id(), session);
// A presence contact that already arrived for this terminal found no registry entry to
// transition and gave up silently. Retry it now that one exists; remove this call and
// such a session stays in SPAWNING even though it is present.
reconcilePresence(handle.terminalId());
handles.put(handle.id(), handle); handles.put(handle.id(), handle);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}", log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree()); handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
@@ -899,6 +997,20 @@ public final class SessionManager implements TurnListener {
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY); transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
} }
/**
* Completes a newly registered session's {@code SPAWNING -> READY} transition when {@code
* terminalId} was already marked present before this ran. A terminal never marked present is
* left in {@code SPAWNING}; it reaches {@code READY} normally through {@link #onReady} once
* its own contact arrives. Callers must run this only once the session's registry entry is
* already visible — {@link #onReady}'s transition matches against that entry, and reconciling
* before the entry exists finds nothing to transition.
*/
private void reconcilePresence(String terminalId) {
if (terminalId != null && !terminalId.isBlank() && presence.isPresent(terminalId)) {
onReady(terminalId);
}
}
/** /**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn. * Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session * The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
@@ -6,6 +6,8 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr; import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient; import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover; import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.ReplyInbox; import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin; import io.javalin.Javalin;
@@ -34,7 +36,7 @@ import static org.junit.jupiter.api.Assertions.fail;
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor, * fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
* not an independent claim — needs no replacement of its own), {@code * not an independent claim — needs no replacement of its own), {@code
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link * mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code * #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}), and {@code
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link * factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually * #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is * covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
@@ -50,8 +52,8 @@ import static org.junit.jupiter.api.Assertions.fail;
* invisible to this test, even though the two are genuinely different daemons in production. This * invisible to this test, even though the two are genuinely different daemons in production. This
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same * version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate * pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake * lead from member) and asserts the roll's {@code pane.close} call lands on the LEAD fake and
* and never on the MEMBER one. * never on the MEMBER one.
*/ */
class FleetdLeadRolloverAssemblyTest { class FleetdLeadRolloverAssemblyTest {
@@ -177,11 +179,47 @@ class FleetdLeadRolloverAssemblyTest {
return FleetConfig.load(f); return FleetConfig.load(f);
} }
@SuppressWarnings("unchecked") /**
* Unlike {@link #writeConfig}, this names a {@code profile:} for the lead and declares it
* under {@code profiles:}, so {@code LeadLauncher#relaunch} can actually start a fresh agent
* instead of refusing with "names no profile". {@code relaunchReadySeconds} is cut to 2s so
* the recognition wait (expected to time out — see the test) does not cost real test seconds.
*/
private static FleetConfig writeConfigWithRelaunchableLead(Path dir, Path leadCwd) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
fleet:
leaders:
opus:
tab: "lead: opus"
cwd: "%s"
profile: opus
profiles:
opus:
subscription: true
argv: ["ccs", "opus"]
leadRollover:
handoverPath: handover.md
requireOperatorConfirm: false
relaunchReadySeconds: 2
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
return FleetConfig.load(f);
}
@Test @Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation " @DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the open/confirm/continuation "
+ "sequence — /clear, then bootstrapText — through the real herdr router") + "sequence through the real herdr router — ending the old pane, then giving up once it "
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception { + "never reports gone")
void assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace"); Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd); Files.createDirectories(leadCwd);
FleetConfig cfg = writeConfig(dir, leadCwd); FleetConfig cfg = writeConfig(dir, leadCwd);
@@ -204,6 +242,11 @@ class FleetdLeadRolloverAssemblyTest {
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = " + "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
+ "null;` at that call site can never pass this"); + "null;` at that call site can never pass this");
// assembleAndStart's own boot work (the orphan-worker reap) makes a real call on the
// member daemon before the roll ever starts. Clear it here so the assertion below measures
// only what the roll itself does, not what daemon startup does.
member.calls.clear();
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test"); LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString(); String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
assertEquals(expectedHandoverPath, pending.handoverPath()); assertEquals(expectedHandoverPath, pending.handoverPath());
@@ -223,40 +266,109 @@ class FleetdLeadRolloverAssemblyTest {
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token) // constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
// until the real continuation finishes. // until the real continuation finishes.
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token()); LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
+ "the turn-boundary wait settles immediately and the post-/clear wait "
+ "releases via its pickup-grace path — detail: " + status.detail());
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD // FakeHerdr's pane.get is a fixed canned response that never reports a pane as gone, so the
// pane — this is the one thing a source-text pin on the call site could never show. // real router's death poll runs out its whole budget and the roll stops here — proving the
// real teardown call landed on the real LEAD pane without ever reaching a relaunch or a send.
assertEquals(LeadRollover.RollState.OLD_PANE_NEVER_DIED, status.state(),
"the old pane never reports gone against this fake, so the roll must stop with "
+ "OLD_PANE_NEVER_DIED rather than ever relaunching or sending anything — "
+ "detail: " + status.detail());
// Prove the real herdr router actually closed the real LEAD pane — this is the one thing a
// source-text pin on the call site could never show.
boolean closedOldPane = lead.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& c.params() instanceof Map<?, ?> m && "w2:p7".equals(m.get("pane_id")));
assertTrue(closedOldPane, "endOldSession must close the real old pane (w2:p7) through the "
+ "real LEAD herdr client, got calls: " + lead.calls);
// No agent.prompt is ever sent on this path: the roll stops at the pane-death wait, strictly
// before the relaunch and the final send step.
List<FakeHerdr.Call> prompts = lead.calls.stream() List<FakeHerdr.Call> prompts = lead.calls.stream()
.filter(c -> c.method().equals("agent.prompt")) .filter(c -> c.method().equals("agent.prompt"))
.toList(); .toList();
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send " assertTrue(prompts.isEmpty(), "a roll that stops at OLD_PANE_NEVER_DIED must never reach the "
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts); + "send step, got agent.prompt call(s) on the LEAD daemon: " + prompts);
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
"the first send must be the literal /clear housekeeping command");
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
"the second send must be the default bootstrapText naming the resolved handover "
+ "path, got: " + secondText);
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation // fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
// swapping router.leadAgents() for router.memberAgents() at the real call site would move // swapping router.leadAgents() for router.memberAgents() at the real call site would move
// both sends above onto `member` instead, which this assertion catches — the thing the // the pane.close call above onto `member` instead, which this assertion catches — the thing
// single-fake version of this test could never see, because both wrapped the same client. // the single-fake version of this test could never see, because both wrapped the same client.
assertTrue(member.calls.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
+ member.calls.size() + " call(s) recorded on the MEMBER daemon since the roll began: "
+ member.calls);
}
/**
* Exercises the relaunch site {@link #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}
* never reaches: with the old pane confirmed gone, the roll relaunches a fresh lead, and
* {@code bootstrapText} must reach it even though recognition times out (FakeHerdr's
* {@code tab.list} is a fixed canned response that never reflects the relaunch's own
* {@code tab.rename}, so the fresh terminal is never recognised as a live lead). Same
* dual-socket shape as the sibling test: two distinct {@link FakeHerdr} instances, so a
* {@code bootstrapText} send wired to the wrong daemon is visible.
*/
@Test
@DisplayName("[BEHAVIOURAL] bootstrapText reaches the fresh LEAD terminal even when recognition "
+ "times out, and the MEMBER daemon never sees it")
void bootstrapTextReachesTheFreshLeadTerminalEvenWhenRecognitionTimesOut(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
FleetConfig cfg = writeConfigWithRelaunchableLead(dir, leadCwd);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
FakeHerdr lead = new FakeHerdr();
lead.withTab("w2", "w2:t7", "lead: opus");
// Lets the old pane (w2:p7) report gone once pane.close actually reaches it, so the roll
// proceeds to relaunch instead of stopping at OLD_PANE_NEVER_DIED.
lead.paneGoneAfterClose("w2:p7");
FakeHerdr member = new FakeHerdr();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
LeadRollover rollover = runtime.mcp().leadRollover();
assertNotNull(rollover, "leadRollover: is present in this test's config, so a real "
+ "LeadRollover must have been built");
LeadRollover.PendingRollover pending = rollover.open("term_a", "bootstrapText relaunch test");
Thread.sleep(50);
Files.writeString(Path.of(pending.handoverPath()), "handover content for bootstrapText test");
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
assertTrue(decision.accepted(), "confirm() must approve — got: " + decision);
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
assertEquals(LeadRollover.RollState.RELAUNCH_NOT_RECOGNISED, status.state(),
"the fresh terminal is never recognised against this fake's static tab.list, so the "
+ "roll must reach RELAUNCH_NOT_RECOGNISED — not an earlier failure state and "
+ "not ROLLED — detail: " + status.detail());
@SuppressWarnings("unchecked")
List<FakeHerdr.Call> leadPrompts = lead.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertEquals(1, leadPrompts.size(), "exactly one bootstrapText send is expected, on the LEAD "
+ "daemon, once recognition gives up — got: " + leadPrompts);
Object text = ((Map<String, Object>) leadPrompts.get(0).params()).get("text");
assertTrue(text instanceof String && ((String) text).contains("handover.md"),
"the send must be bootstrapText naming the resolved handover path, got: " + text);
// Scoped to agent.prompt specifically, not every MEMBER call: the orphan-worker reap also
// talks to the MEMBER daemon once, unconditionally, at daemon boot — unrelated to this roll.
List<FakeHerdr.Call> memberPrompts = member.calls.stream() List<FakeHerdr.Call> memberPrompts = member.calls.stream()
.filter(c -> c.method().equals("agent.prompt")) .filter(c -> c.method().equals("agent.prompt"))
.toList(); .toList();
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got " assertTrue(memberPrompts.isEmpty(), "bootstrapText must never be sent to the MEMBER daemon, "
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: " + "got: " + memberPrompts);
+ memberPrompts);
} }
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token) private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
throws InterruptedException { throws InterruptedException {
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10); long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(15);
while (System.nanoTime() < deadline) { while (System.nanoTime() < deadline) {
LeadRollover.RollStatus status = rollover.status(token); LeadRollover.RollStatus status = rollover.status(token);
if (status.state() != LeadRollover.RollState.PENDING if (status.state() != LeadRollover.RollState.PENDING
@@ -283,8 +395,10 @@ class FleetdLeadRolloverAssemblyTest {
"""); """);
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml)); ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
AgentControl agents = new AgentControl(new FakeHerdr()); AgentControl agents = new AgentControl(new FakeHerdr());
WorkspaceControl spaces = new WorkspaceControl(new FakeHerdr());
LeadLauncher launcher = new LeadLauncher(agents, spaces, config.get());
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of); LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of);
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in " assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no " + "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
@@ -4,6 +4,8 @@ import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig; import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr; import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover; import dev.ltms.fleet.lead.LeadRollover;
import org.junit.jupiter.api.DisplayName; import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test; import org.junit.jupiter.api.Test;
@@ -54,6 +56,11 @@ class FleetdLeadRolloverWorkspaceLookupTest {
return new AgentControl(new FakeHerdr()); return new AgentControl(new FakeHerdr());
} }
/** None of this class's tests reach the deferred continuation, so a plain fake is enough. */
private static LeadLauncher fakeLauncher(FleetConfig cfg) {
return new LeadLauncher(fakeAgents(), new WorkspaceControl(new FakeHerdr()), cfg);
}
@Test @Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against " @DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
+ "the CALLING lead's configured cwd, not the daemon's own working directory") + "the CALLING lead's configured cwd, not the daemon's own working directory")
@@ -74,7 +81,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
""".formatted(leadCwd.toString())); """.formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml)); ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
() -> Map.of("term_opus", "opus")); () -> Map.of("term_opus", "opus"));
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory " assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object"); + "must construct an object");
@@ -105,7 +113,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan // No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all. // has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of); LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of);
assertNotNull(rollover); assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test"); LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
@@ -144,7 +153,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// below — exactly the natural mistake to make, since leads are discovered by a live tab // below — exactly the natural mistake to make, since leads are discovered by a live tab
// scan that runs AFTER this factory is constructed at startup. // scan that runs AFTER this factory is constructed at startup.
Map<String, String> liveLeadTerminals = new HashMap<>(); Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
() -> liveLeadTerminals); () -> liveLeadTerminals);
assertNotNull(rollover); assertNotNull(rollover);
@@ -15,6 +15,7 @@ class AuthzTest {
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400); private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500); private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600); private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
private static final Principal OBSERVER = Principal.observer("term_observer", 700);
@Test @Test
void anonymousIsAuthorizedForNothing() { void anonymousIsAuthorizedForNothing() {
@@ -30,6 +31,22 @@ class AuthzTest {
assertTrue(Authz.isUnauthenticated(null)); assertTrue(Authz.isUnauthenticated(null));
} }
/**
* {@code MessageService.answer}'s turn-ownership check treats a caller with no terminal as
* matching a turn recorded for the unnamed primary. That rule only stays safe because an
* unauthenticated caller — whose terminal is also {@code null} — never reaches {@code answer}
* at all: {@link #anonymousIsAuthorizedForNothing} already covers every action including
* {@code ANSWER}, but this test names the exact coupling so a future change to either side
* cannot drift without turning this test red.
*/
@Test
void anAnonymousCallerIsRefusedAnswerSoItCanNeverBeMistakenForTheUnnamedPrimary() {
assertFalse(Authz.permits(ANON, ANSWER, null),
"an anonymous caller, whose terminal is also null, must never reach answer() — the "
+ "turn-ownership check's null-terminal match for the unnamed primary owner "
+ "relies on this gate refusing it first");
}
@Test @Test
void orchestrationBelongsToThePrimaryAlone() { void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) { for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
@@ -250,4 +267,49 @@ class AuthzTest {
assertTrue(WORKER_A.isSpawnedMember()); assertTrue(WORKER_A.isSpawnedMember());
assertTrue(ARCH_DESIGN.isSpawnedMember()); assertTrue(ARCH_DESIGN.isSpawnedMember());
} }
// ── the observer matrix ─────────────────────────────────────────────────────────────────────
@Test
void anObserverMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(OBSERVER, READ, null));
assertTrue(Authz.permits(OBSERVER, METRICS, null));
}
@Test
void anObserverMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(OBSERVER, REPLY, "term_observer"), "its own pane is its own");
assertTrue(Authz.permits(OBSERVER, ASK, "term_observer"));
assertFalse(Authz.permits(OBSERVER, REPLY, "term_design"),
"an observer must not reply on another pane");
assertFalse(Authz.permits(OBSERVER, REPLY, null),
"an absent target must not pass the own-session rule");
}
/**
* Every action beyond READ/METRICS/REPLY/ASK, asserted denied for an observer — including
* {@code TASK_READ}, which is the entire point of this role: an unconfigured pane must not be
* able to poll a ticket or read another session's status.
*/
@Test
void anObserverIsDeniedEverythingBeyondReadMetricsReplyAndAsk() {
for (Authz.Action a : Authz.Action.values()) {
if (a == READ || a == METRICS || a == REPLY || a == ASK) {
continue;
}
assertFalse(Authz.permits(OBSERVER, a, "term_observer", target -> true),
"an observer must not " + a + " even when the classifier accepts every target");
}
}
@Test
void anObserverIsNotCountedAsAnyOtherRole() {
assertFalse(OBSERVER.isPrimary());
assertFalse(OBSERVER.isWorker());
assertFalse(OBSERVER.isArchitect());
assertFalse(OBSERVER.isCollaborator());
assertFalse(OBSERVER.isSpawnedMember());
assertTrue(OBSERVER.isObserver());
}
} }
@@ -1,13 +1,25 @@
package dev.ltms.fleet.auth; package dev.ltms.fleet.auth;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr; import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator; import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.mcp.ConnectionIdentity; import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.config.FleetConfig; import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.peer.MemberRole; import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.WorktreeRequest;
import dev.ltms.fleet.session.Worktrees;
import org.junit.jupiter.api.Test; import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map; import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*; import static org.junit.jupiter.api.Assertions.*;
@@ -45,17 +57,23 @@ class CallerResolverTest {
return members; return members;
} }
/**
* With no roster wired up at all (the simple constructor), a loopback pane that owns a herdr
* pane but is not recognised as a live spawned member lands on the {@link Role#OBSERVER} floor
* — unforgeable and never token-gated, exactly like a worker's own identity, because it comes
* from the same connection-derived pane mapping.
*/
@Test @Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() { void aLoopbackPaneWithNoLiveRosterResolvesToObserverRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null); Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret") Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 42, null); .resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, underTrust.role()); assertEquals(Role.OBSERVER, underTrust.role());
assertEquals("term_a", underTrust.terminal()); assertEquals("term_a", underTrust.terminal());
assertEquals(Role.WORKER, underToken.role(), assertEquals(Role.OBSERVER, underToken.role(),
"worker identity is unforgeable and must never be token-gated — otherwise enabling " "the floor is unforgeable and must never be token-gated — otherwise enabling auth "
+ "auth would lock the whole fleet out of fleet_reply"); + "would lock every unconfigured pane out of even READ");
assertEquals("term_a", underToken.terminal()); assertEquals("term_a", underToken.terminal());
} }
@@ -79,20 +97,20 @@ class CallerResolverTest {
} }
@Test @Test
void otherPanesRemainWorkersWhenAPinIsSet() { void otherPanesRemainAtTheFloorWhenAPinIsSet() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else") Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null); .resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role()); assertEquals(Role.OBSERVER, p.role());
assertEquals("term_a", p.terminal()); assertEquals("term_a", p.terminal());
} }
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */ /** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test @Test
void aBlankPinLeavesWorkerResolutionUntouched() { void aBlankPinLeavesFloorResolutionUntouched() {
assertEquals(Role.WORKER, assertEquals(Role.OBSERVER,
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role()); CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER, assertEquals(Role.OBSERVER,
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role()); CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
} }
@@ -213,13 +231,13 @@ class CallerResolverTest {
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker. * mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
*/ */
@Test @Test
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() { void aHerdrErrorOnANonOwningPaneStillResolvesTheRealPane() {
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found"); FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID); ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null); Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
assertEquals(Role.WORKER, p.role()); assertEquals(Role.OBSERVER, p.role());
assertEquals("term_a", p.terminal()); assertEquals("term_a", p.terminal());
} }
@@ -263,12 +281,12 @@ class CallerResolverTest {
} }
@Test @Test
void aPaneAbsentFromTheRegistryIsStillAWorker() { void aPaneAbsentFromTheRegistryFallsToTheObserverFloor() {
Principal p = new CallerResolver(workerIdentity(), false, null, Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6")) Map.of("term_elsewhere", "gpt-sol-5.6"))
.resolve("127.0.0.1", 42, null); .resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role()); assertEquals(Role.OBSERVER, p.role());
assertEquals("term_a", p.terminal()); assertEquals("term_a", p.terminal());
assertNull(p.name()); assertNull(p.name());
} }
@@ -294,16 +312,16 @@ class CallerResolverTest {
} }
@Test @Test
void anEmptyRegistryLeavesEveryPaneAWorker() { void anEmptyRegistryLeavesEveryPaneAtTheObserverFloor() {
Map<String, String> noLeads = null; Map<String, String> noLeads = null;
assertEquals(Role.WORKER, assertEquals(Role.OBSERVER,
new CallerResolver(workerIdentity(), false, null, Map.of()) new CallerResolver(workerIdentity(), false, null, Map.of())
.resolve("127.0.0.1", 42, null).role()); .resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER, assertEquals(Role.OBSERVER,
new CallerResolver(workerIdentity(), false, null, noLeads) new CallerResolver(workerIdentity(), false, null, noLeads)
.resolve("127.0.0.1", 42, null).role()); .resolve("127.0.0.1", 42, null).role());
// CB-531: and the same for the live-registry form, whose supplier may also be absent. // And the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.WORKER, assertEquals(Role.OBSERVER,
CallerResolver.withLeads(workerIdentity(), false, null, null) CallerResolver.withLeads(workerIdentity(), false, null, null)
.resolve("127.0.0.1", 42, null).role()); .resolve("127.0.0.1", 42, null).role());
} }
@@ -376,7 +394,7 @@ class CallerResolverTest {
Map<String, String> live = new java.util.HashMap<>(); Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live); CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role()); assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
@@ -393,7 +411,7 @@ class CallerResolverTest {
mutable.put("term_a", "sneaky"); mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role()); assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
} }
// ── CB-548: architect slots ───────────────────────────────────────────────────────────────── // ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
@@ -423,13 +441,13 @@ class CallerResolverTest {
} }
@Test @Test
void anUnboundPaneStillResolvesAsAWorker() { void anUnboundPaneResolvesToTheObserverFloor() {
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(), MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null)); Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null));
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members).resolve("127.0.0.1", 42, null); Map::of, members).resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role()); assertEquals(Role.OBSERVER, p.role());
assertNull(p.name()); assertNull(p.name());
} }
@@ -455,12 +473,11 @@ class CallerResolverTest {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members); Map::of, members);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role()); assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role()); assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("architect:lead-designer", r.members().get("term_a"));
} }
@Test @Test
@@ -483,12 +500,13 @@ class CallerResolverTest {
} }
@Test @Test
void aBoundNonArchitectSlotStillResolvesAsAWorker() { void aBoundNonArchitectSlotResolvesToTheObserverFloorNotArchitect() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("dev:builder", MemberRole.DEV)) Map::of, boundMembers("dev:builder", MemberRole.DEV))
.resolve("127.0.0.1", 42, null); .resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights"); assertEquals(Role.OBSERVER, p.role(), "a dev binding must never grant architect rights, "
+ "and this construction path wires no roster to recognise it as the live dev it is");
} }
@Test @Test
@@ -499,17 +517,17 @@ class CallerResolverTest {
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " ")); assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
} }
@Test @Test
void aWorkerOnAnyLoopbackSourceAddressIsStillAWorkerNotThePrimary() { void aPaneOnAnyLoopbackSourceAddressIsStillAtTheFloorNotThePrimary() {
// fleetd #305: the escalation. ConnectionIdentity used to accept only 127.0.0.1, so a // fleetd #305: the escalation this guards against. ConnectionIdentity used to accept only
// worker connecting from 127.0.0.2 resolved to no terminal, and this resolver's own // 127.0.0.1, so a pane connecting from 127.0.0.2 resolved to no terminal, and this
// (wider) loopback check then made it the PRIMARY — granting spawn, stop, send and drain. // resolver's own (wider) loopback check then made it the PRIMARY — granting spawn, stop,
// Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds there, so the // send and drain. Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds
// path is real and not theoretical. // there, so the path is real and not theoretical.
CallerResolver r = new CallerResolver(workerIdentity(), false, null); CallerResolver r = new CallerResolver(workerIdentity(), false, null);
for (String src : new String[]{"127.0.0.1", "127.0.0.2", "127.1.2.3", "::ffff:127.0.0.2"}) { for (String src : new String[]{"127.0.0.1", "127.0.0.2", "127.1.2.3", "::ffff:127.0.0.2"}) {
Principal p = r.resolve(src, 55555, null); Principal p = r.resolve(src, 55555, null);
assertEquals(Role.WORKER, p.role(), "a worker must stay a worker from source " + src); assertEquals(Role.OBSERVER, p.role(), "the pane must stay off PRIMARY from source " + src);
assertEquals("term_a", p.terminal(), "worker terminal from source " + src); assertEquals("term_a", p.terminal(), "pane terminal from source " + src);
} }
} }
@@ -542,6 +560,219 @@ class CallerResolverTest {
assertEquals("term_a", p.terminal()); assertEquals("term_a", p.terminal());
} }
/**
* fleetd #702: between {@link SessionManager#release} removing the registry entry and the
* pane actually stopping, a resolve for that terminal must still see the live member and
* never fall through to a tab map — wired through the real {@link SessionManager}, not a
* hand-rolled stand-in for its {@code spawnedMemberRole}.
*
* <p>Reuses the ticket's own test idea: an injected {@code hasUncommitted} resolves the
* releasing terminal from inside {@code release}'s window — a real call landing inside the
* window, so no sleep and no race.
*
* <p>The property asserted is "no tab map is consulted", never "the same role is returned".
* The mandatory control is the lead tab map: it names this exact terminal, and the same
* resolve taken <em>outside</em> the window (before release runs) must still return the lead
* role — without that control, the in-window assertion would also pass on an empty map and
* prove nothing.
*/
@Test
void aPaneMidTeardownResolvesAsItsOwnRoleConsultingNoTabMap() {
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
AtomicReference<Principal> duringWindow = new AtomicReference<>();
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
Worktrees worktrees = new Worktrees() {
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
// Runs from INSIDE release()'s git-status shell-out: the registry entry is
// already gone, but the pane has not stopped yet.
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
return false;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
return Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
};
SessionManager sessions = new SessionManager(workers, worktrees);
// The lead tab map names "term_a" before anything is ever spawned onto it — the control
// this test needs. Built up front so the SAME resolver answers every resolve() call below.
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_a", "the-lead"), null, sessions::spawnedMemberRole, Map::of);
resolverRef.set(resolver);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, before.role(),
"control: with no live or releasing member on this terminal, the lead tab map must "
+ "win — this is what proves the in-window assertion below is not passing on "
+ "an empty map");
assertEquals("the-lead", before.name());
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702", null));
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
sessions.release(s.paneId());
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
assertEquals(Role.WORKER, duringWindow.get().role(),
"inside the window the pane must resolve as its own live-member role, consulting no "
+ "tab map — a lead tab naming the same terminal must not win");
assertEquals("term_a", duringWindow.get().terminal());
}
/**
* {@link SessionManager#releaseRemoved} unbinds the architect slot before the git-status
* shell-out that opens the teardown window, so a resolve landing inside that window must see
* the slot already unbound and resolve {@link Role#WORKER} — never {@link Role#ARCHITECT},
* and never by checking role equality against the live session, which would hold even if the
* unbind ran too late.
*
* <p>The control is the same resolve taken outside the window, while the slot is still bound,
* which must return {@link Role#ARCHITECT} — without it this test would also pass against a
* slot that was never bound, and prove nothing.
*/
@Test
void aReleasingArchitectIsDemotedToWorkerInsideTheTeardownWindow() {
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
AtomicReference<Principal> duringWindow = new AtomicReference<>();
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
Worktrees worktrees = new Worktrees() {
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
// Runs from INSIDE release()'s git-status shell-out: the architect slot is already
// unbound by this point, but the pane has not stopped yet.
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
return false;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
return Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
};
SessionManager sessions = new SessionManager(workers, worktrees);
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("ltms-local")), Map.of(), Map.of(), null));
sessions.setMemberLifecycle(members);
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, members, sessions::spawnedMemberRole, Map::of);
resolverRef.set(resolver);
MemberSession s = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null, "/caller/proj", null,
new WorktreeRequest("fleetd-702d", null));
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
assertEquals(MemberRole.ARCHITECT, s.role(), "sanity: the slot bind succeeded");
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(),
"control: with the slot still bound, the pane must resolve as an architect — this "
+ "is what proves the in-window assertion below is not passing against a "
+ "slot that was never bound");
assertEquals("lead-designer", before.name());
sessions.release(s.paneId());
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
assertEquals(Role.WORKER, duringWindow.get().role(),
"inside the window the architect slot is already unbound, so the result must be "
+ "WORKER — asserting role-equality with the live session here would tempt "
+ "moving the unbind earlier or later, which would be wrong either way");
assertEquals("term_a", duringWindow.get().terminal());
}
/** A spawned architect in the roster resolves ARCHITECT, carrying its bound slot's name. */ /** A spawned architect in the roster resolves ARCHITECT, carrying its bound slot's name. */
@Test @Test
void aSpawnedArchitectInTheRosterResolvesArchitectWithItsSlotName() { void aSpawnedArchitectInTheRosterResolvesArchitectWithItsSlotName() {
@@ -602,14 +833,14 @@ class CallerResolverTest {
assertEquals("term_a", p.terminal()); assertEquals("term_a", p.terminal());
} }
/** Regression: an empty collaborator registry leaves every pane exactly as before. */ /** Regression: an empty collaborator registry leaves every pane at the unconfigured-pane floor. */
@Test @Test
void anEmptyCollaboratorRegistryLeavesEveryPaneAsBefore() { void anEmptyCollaboratorRegistryLeavesEveryPaneAtTheObserverFloor() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of, Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, Map::of) new MemberRegistry(null), t -> null, Map::of)
.resolve("127.0.0.1", 42, null); .resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role()); assertEquals(Role.OBSERVER, p.role());
assertNull(p.name()); assertNull(p.name());
} }
@@ -638,6 +869,39 @@ class CallerResolverTest {
assertFalse(r.knownLeadOrCollaborator().test("term_other")); assertFalse(r.knownLeadOrCollaborator().test("term_other"));
} }
// ── fleetd #705: narrowing the unconfigured-pane floor to OBSERVER ──────────────────────────
/**
* The case this ticket exists for: a pane the resolver cannot place as a live spawned member,
* a lead, a bound architect slot, or a configured collaborator must land on the narrow
* {@link Role#OBSERVER} floor, never the {@link Role#WORKER} the old fallback granted.
*
* <p>The second assertion is the control the ticket requires: a terminal the roster DOES
* recognise as a live spawned member must still resolve its own role. Without it, this test
* would also pass if the fix accidentally turned every caller into an observer.
*/
@Test
void anUnconfiguredPaneResolvesObserverButARegisteredMemberStillResolvesItsOwnRole() {
Principal unconfigured = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, unconfigured.role(),
"a pane matching none of the configured or live-roster roles must fall to the "
+ "floor, not WORKER");
assertEquals("term_a", unconfigured.terminal());
Principal registered = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, new MemberRegistry(null),
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, registered.role(),
"control: a live spawned member must keep resolving its own role, never the "
+ "unconfigured-pane floor");
}
@Test
void describeNamesTheObserverByItsPane() {
assertEquals("observer:term_a", Principal.observer("term_a", 1).describe());
}
@Test @Test
void knownLeadOrCollaboratorIsFalseForASpawnedMembersTerminal() { void knownLeadOrCollaboratorIsFalseForASpawnedMembersTerminal() {
// The exact scenario a collaborator's SEND must never reach: a live spawned member's own // The exact scenario a collaborator's SEND must never reach: a live spawned member's own
@@ -295,9 +295,10 @@ class MemberRegistryLiveTest {
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary()); assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null); Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, after.role(), assertEquals(Role.OBSERVER, after.role(),
"removing the slot from config must demote the bound session to worker on its " "removing the slot from config must demote the bound session on its NEXT request — "
+ "NEXT request — this is the ticket's whole point"); + "this harness wires no live roster for term_a, so the demotion lands on "
+ "the unconfigured-pane floor");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed"); assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
} }
@@ -0,0 +1,31 @@
package dev.ltms.fleet.auth;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
class PrincipalTest {
@Test
void ownerKeyCoversEveryRole() {
assertEquals("leader:opus", Principal.leader("opus", "term_lead", 1).ownerKey());
assertNull(Principal.primary(2).ownerKey());
assertEquals("worker:term_worker", Principal.worker("term_worker", 3).ownerKey());
assertEquals("architect:term_arch", Principal.architect("opus", "term_arch", 4).ownerKey());
assertEquals("collaborator:ops", Principal.collaborator("ops", "term_collab", 5).ownerKey());
assertEquals("observer:term_observer", Principal.observer("term_observer", 6).ownerKey());
assertEquals("anonymous", Principal.anonymous().ownerKey());
}
@Test
void rolePrefixesKeepLeadAndArchitectKeysDistinct() {
String lead = Principal.leader("opus", "term_lead", 1).ownerKey();
String architect = Principal.architect("design", "opus", 2).ownerKey();
assertEquals("leader:opus", lead);
assertEquals("architect:opus", architect);
assertNotEquals(lead, architect);
}
}
@@ -103,7 +103,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value // comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
// here proves it, rather than leaving it null and proving nothing. // here proves it, rather than leaving it null and proving nothing.
v.put("leadRollover", new FleetConfig.LeadRollover( v.put("leadRollover", new FleetConfig.LeadRollover(
"/handover/guard.md", true, 3600, 20, 20, "read the handover file")); "/handover/guard.md", true, 3600, 20, 45, "read the handover file"));
assertNamesMatchComponents(v); assertNamesMatchComponents(v);
return v; return v;
} }
@@ -7,6 +7,7 @@ import java.util.ArrayList;
import java.util.LinkedHashMap; import java.util.LinkedHashMap;
import java.util.List; import java.util.List;
import java.util.Map; import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap; import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList; import java.util.concurrent.CopyOnWriteArrayList;
@@ -59,6 +60,10 @@ public final class FakeHerdr implements HerdrClient {
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable) private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
/** pane ids that {@link #paneGoneAfterClose} has opted into reporting gone — see that method. */
private final Set<String> paneGoneAfterCloseIds = ConcurrentHashMap.newKeySet();
/** pane ids a {@code pane.close} call has actually reached, for {@link #paneGoneAfterCloseIds}. */
private final Set<String> closedPaneIds = ConcurrentHashMap.newKeySet();
public FakeHerdr healthy(boolean h) { public FakeHerdr healthy(boolean h) {
this.healthy = h; this.healthy = h;
@@ -220,6 +225,19 @@ public final class FakeHerdr implements HerdrClient {
return this; return this;
} }
/**
* Make {@code pane.get(paneId)} report the pane gone (a {@code pane_not_found} {@link
* HerdrException}, exactly as {@link WorkspaceControl#locatePane} expects to see once a pane
* has really disappeared) once a {@code pane.close} call for that same {@code paneId} has
* actually reached this fake. Every other pane, and this pane before its own close, keeps
* reporting the default canned {@code pane.get} response — opt-in, by pane id, so no existing
* test's {@code pane.get} behaviour changes.
*/
public FakeHerdr paneGoneAfterClose(String paneId) {
paneGoneAfterCloseIds.add(paneId);
return this;
}
/** /**
* Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake — * Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake —
* i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this * i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this
@@ -422,9 +440,18 @@ public final class FakeHerdr implements HerdrClient {
} }
yield mapper.readTree("{\"type\":\"ok\"}"); yield mapper.readTree("{\"type\":\"ok\"}");
} }
case "pane.get" -> mapper.readTree(""" case "pane.get" -> {
Object paneIdParam = params instanceof Map<?, ?> m ? m.get("pane_id") : null;
String paneIdKey = paneIdParam == null ? null : String.valueOf(paneIdParam);
if (paneIdKey != null && paneGoneAfterCloseIds.contains(paneIdKey)
&& closedPaneIds.contains(paneIdKey)) {
throw new HerdrException("herdr error [pane_not_found]: pane.get failed",
"pane_not_found", null);
}
yield mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9", {"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}"""); "tab_id":"w9:t2","agent_status":"idle"}}""");
}
case "pane.list" -> noPanes case "pane.list" -> noPanes
? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}") ? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}")
: mapper.readTree(""" : mapper.readTree("""
@@ -458,6 +485,9 @@ public final class FakeHerdr implements HerdrClient {
throw new HerdrException("herdr error [" + code + "]: pane.close failed", throw new HerdrException("herdr error [" + code + "]: pane.close failed",
code, null); code, null);
} }
if (paneIdParam != null) {
closedPaneIds.add(String.valueOf(paneIdParam));
}
yield mapper.readTree("{\"type\":\"ok\"}"); yield mapper.readTree("{\"type\":\"ok\"}");
} }
default -> throw new HerdrException("fake has no canned response for " + method); default -> throw new HerdrException("fake has no canned response for " + method);
@@ -1,10 +1,16 @@
package dev.ltms.fleet.lead; package dev.ltms.fleet.lead;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig; import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr; import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.WorkspaceControl; import dev.ltms.fleet.herdr.WorkspaceControl;
import org.junit.jupiter.api.Test; import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashMap; import java.util.LinkedHashMap;
import java.util.List; import java.util.List;
@@ -64,6 +70,21 @@ class LeadLauncherTest {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg); return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
} }
/** As {@link #launcher}, plus a fast no-op sleeper so a busy-retry test never real-sleeps. */
private static LeadLauncher fastLauncher(FakeHerdr herdr, FleetConfig cfg) {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg, () -> { });
}
/** {@link #opusProfile()} with one argv element long enough to overflow the pane line limit. */
private static FleetConfig.Profile hugeArgvProfile() {
return new FleetConfig.Profile(
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "x".repeat(1500)), "tab", "fleet", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
@SuppressWarnings("unchecked") @SuppressWarnings("unchecked")
private static List<String> startedArgs(FakeHerdr herdr) { private static List<String> startedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args"); return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
@@ -86,7 +107,11 @@ class LeadLauncherTest {
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads()); assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started"); assertTrue(herdr.called("agent.start"), "a lead must actually be started");
assertEquals("lead-opus", startedName(herdr)); // fleetd #727: the name carries a per-process nonce and a per-start sequence number — the
// same unique-naming scheme HerdrPeerLauncher uses for members — rather than the fixed
// "lead-opus" a stale registry entry could block a legitimate relaunch under.
assertTrue(startedName(herdr).matches("lead-opus-[0-9a-f]{6}-\\d+"),
"name is lead-<name>-<nonce>-<seq>: " + startedName(herdr));
} }
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */ /** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
@@ -436,4 +461,247 @@ class LeadLauncherTest {
assertEquals(0, launcher(herdr, cfg).ensureLeads()); assertEquals(0, launcher(herdr, cfg).ensureLeads());
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned"); assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
} }
// ── fleetd #727: the same three launch protections every member spawn gets ───────────────────
/**
* A freshly created pane may not have redrawn its prompt yet, so herdr answers
* {@code agent_pane_busy}. The lead launch must wait it out rather than fail on the first miss
* — exactly the retry {@code HerdrPeerLauncher} already gives every member.
*/
@Test
void aLeadLaunchRetriesWhileTheSeedShellBoots() {
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
assertEquals(1, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the lead must still start once the shell is ready");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"two busy rejections, then the successful start");
}
/**
* The busy retry is bounded, not an infinite poll. If the pane never becomes ready the launch
* must eventually give up and log the failure, not hang the daemon's reconcile loop forever —
* proven here by a budget that would still be busy on attempt
* {@value ResilientAgentLaunch#SHELL_READY_RETRIES} and a call count that stops exactly there.
*/
@Test
void aLeadLaunchGivesUpAfterTheBoundedBusyBudgetRatherThanLoopingForever() {
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(999);
assertEquals(0, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a pane that never becomes ready must not be reported as a started lead");
assertEquals(ResilientAgentLaunch.SHELL_READY_RETRIES,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"the retry budget is bounded: it stops after exactly SHELL_READY_RETRIES attempts");
}
/**
* herdr refuses a duplicate agent {@code name} outright. A stale registry entry — a crashed
* lead session, or a name the registry has not yet released — must not permanently block a
* legitimate relaunch, so each retry attempt carries a fresh per-start name.
*/
@Test
@SuppressWarnings("unchecked")
void aLeadLaunchRetriesUnderAFreshNameWhenTheOldNameIsStillTaken() {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the lead must still start once a free name is found");
List<String> names = herdr.calls.stream()
.filter(c -> c.method().equals("agent.start"))
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
.toList();
assertEquals(3, names.size(), "2 rejected + 1 success");
assertEquals(3, Set.copyOf(names).size(), "each attempt must use a distinct name");
}
/**
* The name-collision retry is bounded too. If the name is taken on every attempt, the launch
* must give up rather than keep minting new names forever — the per-start naming scheme means
* every failed attempt was refused outright by herdr (no process exists under a name herdr
* refused), so a bounded, exhausted retry never leaves a second process running: nothing is
* running at all.
*/
@Test
void aLeadLaunchNeverEndsUpWithASecondProcessWhenTheNameStaysTaken() {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(999);
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a name that is never free must not be reported as a started lead");
assertEquals(ResilientAgentLaunch.NAME_RETRIES,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"the retry budget is bounded: it stops after exactly NAME_RETRIES attempts, "
+ "never racing a duplicate into existence");
}
/**
* fleetd #220/#727: herdr types the launch command into the pane as one line, and a pty line
* buffer holds only 1024 bytes — past that the tail is dropped with no error at all, and the
* backend exits on a mangled argument. The lead launch must refuse an over-long command outright
* rather than let it be typed and silently truncated.
*/
@Test
void anOverlongLeadArgvIsRefusedRatherThanTypedAndTruncated() {
FakeHerdr herdr = new FakeHerdr();
Logger logger = (Logger) LoggerFactory.getLogger(LeadLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
int started;
try {
started = launcher(herdr, configWith(lead("opus", "lead: opus", 1), hugeArgvProfile()))
.ensureLeads();
} finally {
logger.detachAppender(appender);
}
assertEquals(0, started, "an over-long command must never be reported as a started lead");
assertFalse(herdr.called("agent.start"),
"nothing may be started — a truncated command is worse than no spawn");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("failed to launch"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a WARN naming the launch failure: "
+ appender.list));
assertTrue(warn.contains("1024"), "names the limit: " + warn);
assertTrue(warn.contains("opus"), "names the profile: " + warn);
assertTrue(warn.contains("x".repeat(60)), "names the culprit argument: " + warn);
}
// ── fleetd #726 unit 1: the single-lead relaunch seam ─────────────────────────────────────
/**
* The returned agent's {@code terminalId()}/{@code paneId()} are the ones the fake
* {@code AgentControl} actually started — not a coincidental field left over from the caller.
* {@code paneId()} echoes the exact {@code pane_id} the launch's own {@code agent.start} call
* carried (protocol 19: the agent starts into the pane it is asked to), and {@code
* terminalId()} is herdr's own generated id, which the fake always shapes as {@code
* term_new_<n>}.
*/
@Test
void relaunchReturnsTheStartedAgent() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started, "a launchable, configured lead must start");
Object startedPaneIdParam = ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("pane_id");
assertEquals(startedPaneIdParam, started.paneId(),
"paneId() must be the pane the agent.start call actually targeted");
assertTrue(started.terminalId() != null && started.terminalId().startsWith("term_new_"),
"terminalId() must be herdr's own generated id: " + started.terminalId());
}
/** The new tab is labelled with the lead's configured {@code tab:}, and AFTER the start. */
@Test
void relaunchLabelsTheNewTabAfterStarting() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started);
assertEquals("lead: opus", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
int startIndex = indexOfLastCall(herdr, "agent.start");
int renameIndex = indexOfLastCall(herdr, "tab.rename");
assertTrue(renameIndex > startIndex,
"the tab must be renamed AFTER the start succeeds, not before: start=" + startIndex
+ " rename=" + renameIndex);
}
private static int indexOfLastCall(FakeHerdr herdr, String method) {
int idx = -1;
List<FakeHerdr.Call> calls = herdr.calls;
for (int i = 0; i < calls.size(); i++) {
if (calls.get(i).method().equals(method)) {
idx = i;
}
}
return idx;
}
@Test
void relaunchOfAnUnknownLeadNameReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("not-declared");
assertNull(started);
assertFalse(herdr.called("agent.start"));
assertFalse(herdr.called("workspace.create"));
assertFalse(herdr.called("tab.create"));
}
@Test
void relaunchOfARecogniseOnlyLeadReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead(null, "lead: dead", 1))).relaunch("opus");
assertNull(started);
assertFalse(herdr.called("agent.start"));
}
@Test
void relaunchWithAnUnconfiguredProfileReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("nope", "lead: opus", 1))).relaunch("opus");
assertNull(started);
assertFalse(herdr.called("agent.start"));
}
/**
* The outer retry {@link LeadLauncher#relaunch(String)} owns, separate from {@code
* ResilientAgentLaunch}'s internal {@code agent_name_taken} retry: a failed attempt must not
* be the end of the whole relaunch. Each of the first two attempts exhausts {@code
* ResilientAgentLaunch.NAME_RETRIES} name attempts (every one of them rejected), so each
* attempt's own tab is created and then closed; the third attempt's first name is free.
*/
@Test
void relaunchRetriesTheWholeAttemptAndSucceedsOnTheThird() {
FakeHerdr herdr = new FakeHerdr()
.agentNameTakenTimes(2 * ResilientAgentLaunch.NAME_RETRIES);
dev.ltms.fleet.herdr.Agent started =
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started, "the third attempt's first name is free — it must succeed");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count(),
"one tab per attempt: three attempts");
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count(),
"the two failed attempts' tabs must be closed");
}
/**
* Every attempt fails outright (a herdr error {@code ResilientAgentLaunch} does not retry at
* all) — {@link LeadLauncher#relaunch(String)} must give up after exactly {@code
* RELAUNCH_ATTEMPTS} and must not leak any of the tabs it created along the way.
*/
@Test
void relaunchGivesUpAfterExactlyRelaunchAttemptsAndLeaksNoTab() {
FakeHerdr herdr = new FakeHerdr().agentStartFailsWith("some_other_error");
dev.ltms.fleet.herdr.Agent started =
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNull(started, "every attempt failed — relaunch must give up, not hang or guess");
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"exactly RELAUNCH_ATTEMPTS attempts, no more, no fewer");
long tabsCreated = herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count();
long tabsClosed = herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count();
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS, tabsCreated);
assertEquals(tabsCreated, tabsClosed, "every tab this method created must be closed — no leaks");
}
} }
File diff suppressed because it is too large Load Diff
@@ -355,10 +355,10 @@ class FleetMcpAuthzTest {
+ "the anchors have drifted, this test is not testing what it claims to"); + "the anchors have drifted, this test is not testing what it claims to");
Pattern trailingArg = Pattern.compile( Pattern trailingArg = Pattern.compile(
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;", "listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*[,)]",
Pattern.DOTALL); Pattern.DOTALL);
Matcher m = trailingArg.matcher(handlerBlock); Matcher m = trailingArg.matcher(handlerBlock);
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the " assertTrue(m.find(), "could not locate listFleet(...)'s coordinator boolean argument in the "
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock); + "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
String trailing = m.group(1); String trailing = m.group(1);
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing, assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
@@ -418,6 +418,145 @@ class FleetMcpAuthzTest {
+ "calling, not pass a literal boolean -- block: " + handlerBlock); + "calling, not pass a literal boolean -- block: " + handlerBlock);
} }
// --- who may see fleet_list's leads and members arrays ---------------------------------------
/**
* {@link FleetMcp#leadsVisibleTo} is the whole policy decision for {@code fleet_list}'s
* {@code leads} array: visible to exactly the roles that may {@code SEND} to a lead -- the
* primary, an architect, and a collaborator -- never a worker, which holds {@code READ} but
* can never {@code SEND} at all, and never an anonymous caller.
*/
@Test
void primaryArchitectAndCollaboratorMaySeeTheLeadsArray() {
assertTrue(FleetMcp.leadsVisibleTo(PRIMARY), "the primary must see the leads array");
assertTrue(FleetMcp.leadsVisibleTo(ARCH_DESIGN), "an architect must see the leads array");
assertTrue(FleetMcp.leadsVisibleTo(COLLABORATOR),
"a collaborator may SEND to a lead, so it must see the leads array to learn where");
assertFalse(FleetMcp.leadsVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the leads array");
assertFalse(FleetMcp.leadsVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* As {@link #primaryArchitectAndCollaboratorMaySeeTheLeadsArray}, for the {@code members}
* array -- but a collaborator may {@code SEND} only to a lead or another collaborator, never
* to a spawned member, so it must not see this one.
*/
@Test
void onlyPrimaryAndArchitectMaySeeTheMembersArray() {
assertTrue(FleetMcp.membersVisibleTo(PRIMARY), "the primary must see the members array");
assertTrue(FleetMcp.membersVisibleTo(ARCH_DESIGN), "an architect must see the members array");
assertFalse(FleetMcp.membersVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the members array");
assertFalse(FleetMcp.membersVisibleTo(COLLABORATOR),
"a collaborator may SEND to a lead, never to a spawned member, so it must not see the members array");
assertFalse(FleetMcp.membersVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* Same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}: the
* predicate above can be perfectly correct while the one production call site never asks it.
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
* {@code listFleet(...)} call asks {@code leadsVisibleTo(principal(exchange))} for the leads
* visibility flag, rather than a literal boolean.
*/
@Test
void theFleetListHandlerActuallyConsultsLeadsVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("leadsVisibleTo(principal(exchange))"),
"the fleet_list handler must ask leadsVisibleTo(principal(exchange)) who is calling, "
+ "not pass a literal boolean -- block: " + handlerBlock);
}
/** As {@link #theFleetListHandlerActuallyConsultsLeadsVisibleTo}, for {@code membersVisibleTo}. */
@Test
void theFleetListHandlerActuallyConsultsMembersVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("membersVisibleTo(principal(exchange))"),
"the fleet_list handler must ask membersVisibleTo(principal(exchange)) who is calling, "
+ "not pass a literal boolean -- block: " + handlerBlock);
}
/**
* {@code fleet_poll{ticket}} must thread the calling connection's owner key into
* {@link MessageService#poll(String, String)}, so a worker cannot read a ticket a different
* session created.
*/
@Test
void theFleetPollHandlerActuallyThreadsCallerOwnerIntoPoll() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("pollHandler =");
assertTrue(start >= 0, "could not find the fleet_poll handler (pollHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("ackHandler =", start);
assertTrue(end > start, "could not find the handler declared after pollHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call poll(...) -- if this fails, the anchors
// above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("poll(messages,"),
"control failed: the scraped pollHandler block contains no poll(messages, ...) call "
+ "at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("principal(exchange).ownerKey()"),
"the fleet_poll handler must thread principal(exchange).ownerKey() into poll(...), not omit "
+ "it or pass a literal null -- block: " + handlerBlock);
}
/**
* {@code fleet_status} must thread the calling connection's owner key into
* {@link FleetMcp#status(MessageService, String, String)}, so a caller that did not create a
* worker's open delegation cannot read its pending question through the status handler either.
*/
@Test
void theFleetStatusHandlerActuallyThreadsCallerOwnerIntoStatus() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("statusHandler =");
assertTrue(start >= 0, "could not find the fleet_status handler (statusHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("pollHandler =", start);
assertTrue(end > start, "could not find the handler declared after statusHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call status(...) -- if this fails, the anchors
// above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("status(messages,"),
"control failed: the scraped statusHandler block contains no status(messages, ...) "
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("principal(exchange).ownerKey()"),
"the fleet_status handler must thread principal(exchange).ownerKey() into status(...), not "
+ "omit it or pass a literal null -- block: " + handlerBlock);
}
// --- which action each tool hands the gate (fleetd #272) ------------------------------------ // --- which action each tool hands the gate (fleetd #272) ------------------------------------
/** /**
@@ -11,6 +11,7 @@ import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator; import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl; import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector; import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover; import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.member.ClaudeCodeLauncher; import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.InMemoryReplyInbox; import dev.ltms.fleet.msg.InMemoryReplyInbox;
@@ -24,6 +25,8 @@ import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test; import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir; import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files; import java.nio.file.Files;
import java.nio.file.Path; import java.nio.file.Path;
import java.util.List; import java.util.List;
@@ -36,14 +39,14 @@ import static org.junit.jupiter.api.Assertions.*;
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls * fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}. * {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
* *
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a * <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms poll, a real
* real virtual-thread continuation runner) rather than its package-private test constructor — * virtual-thread continuation runner) rather than its package-private test constructor — this
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not * test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not need to
* need to control the post-{@code confirm()} continuation's timing: it only asserts the * control the post-{@code confirm()} continuation's timing: it only asserts the SYNCHRONOUS return
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what * value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what {@code
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code * FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this * relaunchReadySeconds} are kept at 1s so a confirmed request's background continuation (which
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a * this class does not wait on or assert against) gives up quickly rather than polling for 20s on a
* daemon virtual thread. * daemon virtual thread.
*/ */
class FleetMcpHandoverTest { class FleetMcpHandoverTest {
@@ -67,10 +70,26 @@ class FleetMcpHandoverTest {
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file"); return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
} }
private static FleetConfig minimalFleetConfig() {
try {
Path yaml = Files.createTempFile("fleet-mcp-handover-test", ".yaml");
Files.writeString(yaml, "bind:\n port: 8080\n");
return FleetConfig.load(yaml);
} catch (IOException e) {
throw new UncheckedIOException(e);
}
}
private LeadRollover newRollover(String handoverPath) { private LeadRollover newRollover(String handoverPath) {
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already // Every handoverPath this test class uses comes from tmp.resolve(...), which is already
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough. // absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null); // None of this class's tests reach the recognition-wait or the relaunch call, so the
// launcher's own functional correctness is irrelevant here — any constructed instance,
// backed by the same fake herdr, is enough.
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = new LeadLauncher(agents, spaces, minimalFleetConfig());
return new LeadRollover(agents, spaces, launcher, () -> cfg(handoverPath),
_ -> null, _ -> null, Map::of);
} }
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */ /** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
@@ -122,6 +122,6 @@ class FleetMcpLeadContextGaugeWiringTest {
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs, FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, Map.of(), false, Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, Map.of(), false,
FleetMcp.CoordinationSource.none(), false); FleetMcp.CoordinationSource.none(), false, true, true);
} }
} }
@@ -77,7 +77,7 @@ class FleetMcpTest {
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception { private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles)); () -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles, null));
long deadline = System.currentTimeMillis() + 3000; long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) { while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
Thread.sleep(5); Thread.sleep(5);
@@ -103,7 +103,7 @@ class FleetMcpTest {
void sendThenReplyRoundTrips() throws Exception { void sendThenReplyRoundTrips() throws Exception {
// fleet_send blocks; fleet_reply resolves it with the worker's structured answer. // fleet_send blocks; fleet_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of())); () -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of(), null));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now // Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip). // queues in the inbox if no waiter is open, which would break the round-trip).
@@ -181,7 +181,7 @@ class FleetMcpTest {
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"')); String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L)); () -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS))); assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000; deadline = System.currentTimeMillis() + 3000;
@@ -279,7 +279,7 @@ class FleetMcpTest {
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content()); assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L)); () -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L, null));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS))); assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) { while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5); Thread.sleep(5);
@@ -288,6 +288,91 @@ class FleetMcpTest {
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS))); assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
} }
// --- fleetd #715: fleet_send{turnId} is gated on the caller that owns the turn -------------
/**
* A blocking {@code fleet_send} from one caller opens the turn; a {@code fleet_send{turnId}}
* from a different caller is refused as an error, and the real owner's answer still succeeds.
*/
@Test
void aDifferentCallersMcpAnswerIsRefusedForABlockingSendButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "do X", 5000L, null, Set.of(), "term_owner"));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
McpSchema.CallToolResult question = send.get(5, TimeUnit.SECONDS);
assertTrue(textOf(question).contains("[question]"), textOf(question));
String questionText = textOf(question);
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker");
assertTrue(hijacked.isError(), "a caller that did not open this turn must get an error, not an answer");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner"));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, T, Role.WORKER, "done");
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
}
/**
* Same hijack and control through the fire-and-poll ({@code sendAsync}) path: the owner comes
* from the resolved principal, not from a caller argument threaded through a live call.
*/
@Test
void aDifferentCallersMcpAnswerIsRefusedForAnAsyncSendButTheRealOwnerSucceeds() throws Exception {
Principal owner = Principal.worker("term_owner", 1);
McpSchema.CallToolResult accepted =
FleetMcp.sendAsync(messages, T, "do it", null, Set.of(), owner);
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
MessageService.TaskView asking = messages.poll(ticket);
deadline = System.currentTimeMillis() + 3000;
while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
asking = messages.poll(ticket);
}
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L,
"worker:term_attacker");
assertTrue(hijacked.isError(), "a caller that did not create this delegation must get an error");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, owner.ownerKey()));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, T, Role.WORKER, "done");
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
}
@Test @Test
void pollUnknownTicketIsAnError() { void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = FleetMcp.poll(messages, "task-999", null); McpSchema.CallToolResult res = FleetMcp.poll(messages, "task-999", null);
@@ -297,7 +382,7 @@ class FleetMcpTest {
@Test @Test
void sendTimesOutWithAWorkingNote() { void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of()); McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error"); assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res)); assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
} }
@@ -313,7 +398,7 @@ class FleetMcpTest {
void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception { void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception {
herdr.agentSendFailsWith("send_failed"); herdr.agentSendFailsWith("send_failed");
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of())); () -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null));
long deadline = System.currentTimeMillis() + 2000; long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) { while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait //noinspection BusyWait
@@ -333,13 +418,13 @@ class FleetMcpTest {
@Test @Test
void sendRejectsMissingArgs() { void sendRejectsMissingArgs() {
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of()).isError()); assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of(), null).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of()).isError()); assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of(), null).isError());
} }
@Test @Test
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() { void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol")); McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"), null);
McpSchema.CallToolResult async = FleetMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol")); McpSchema.CallToolResult async = FleetMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
assertTrue(blocking.isError()); assertTrue(blocking.isError());
@@ -358,7 +443,7 @@ class FleetMcpTest {
assertSendRoundTrips("term_live_member", profiles); assertSendRoundTrips("term_live_member", profiles);
// A herdr-owned pane outside the bridge roster cannot be classified at accept time. // A herdr-owned pane outside the bridge roster cannot be classified at accept time.
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles); McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles, null);
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time"); assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
} }
@@ -450,7 +535,7 @@ class FleetMcpTest {
void askThenAnswerRoundTrips() throws Exception { void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks. // The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of())); () -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null));
long deadline = System.currentTimeMillis() + 3000; long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) { while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait //noinspection BusyWait
@@ -472,7 +557,7 @@ class FleetMcpTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply. // The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync( CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L)); () -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
// The worker's ask returns the answer — it resumes the same turn. // The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS))); assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
@@ -498,7 +583,7 @@ class FleetMcpTest {
@Test @Test
void answerToAStaleTurnIsAnError() { void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L); McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L, null);
assertTrue(res.isError()); assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res)); assertTrue(textOf(res).contains("no longer open"), textOf(res));
} }
@@ -692,7 +777,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res); String out = textOf(res);
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's // fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
@@ -721,7 +806,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res); String out = textOf(res);
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out); assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
@@ -780,19 +865,82 @@ class FleetMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))); SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus") FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1)); .withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary(); Principal worker = Principal.worker("term_a", 1);
McpSchema.CallToolResult res = FleetMcp.listFleet( McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), worker.isPrimary(),
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
String out = textOf(res); String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out); assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out); assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out); assertFalse(out.contains("\"leads\""), "a worker must never see the leads key at all: " + out);
assertTrue(out.contains("\"members\""), out); assertFalse(out.contains("\"members\""), "a worker must never see the members key at all: " + out);
assertTrue(out.contains("\"healthCoverage\""), "the rest of the result must still be present: " + out);
}
/**
* A worker's result must carry neither the {@code leads} nor the {@code members} key, and no
* fragment of either row leaks even though both are fully populated for this call -- a key
* check alone would pass on an implementation that still built the rows and only renamed or
* nested them.
*/
@Test
void listLeaksNoLeadOrMemberRowFragmentToAWorker() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
new WorktreeRequest("cb-304", null));
Principal worker = Principal.worker("term_a", 1);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), worker.isPrimary(),
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
String out = textOf(res);
assertFalse(out.contains("\"leads\""), out);
assertFalse(out.contains("\"members\""), out);
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a worker: " + out);
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a worker: " + out);
assertFalse(out.contains("mac-opus"), "no lead row fragment may leak to a worker: " + out);
assertTrue(out.contains("\"healthCoverage\""), out);
}
/**
* A collaborator may {@code SEND} to a lead, so it must see the {@code leads} array -- the
* only place {@code fleet_whoami} does not already give it a lead's address. It may never
* {@code SEND} to a spawned member, so the {@code members} key must stay absent for it, with
* no fragment of a populated member row leaking either.
*/
@Test
void listShowsLeadsButNotMembersToACollaborator() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
new WorktreeRequest("cb-304", null));
Principal collaborator = Principal.collaborator("ops", "term_collab", 600);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), collaborator.isPrimary(),
FleetMcp.leadsVisibleTo(collaborator), FleetMcp.membersVisibleTo(collaborator));
String out = textOf(res);
assertTrue(out.contains("\"leads\""), "a collaborator must see the leads array: " + out);
assertTrue(out.contains("mac-opus"), "a collaborator must see the lead's name/address: " + out);
assertFalse(out.contains("\"members\""), "a collaborator must never see the members key: " + out);
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a collaborator: " + out);
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a collaborator: " + out);
assertTrue(out.contains("\"healthCoverage\""), out); assertTrue(out.contains("\"healthCoverage\""), out);
} }
@@ -808,16 +956,19 @@ class FleetMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))); SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus") FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1)); .withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary(); Principal architect = Principal.architect("lead-designer", "term_design", 400);
McpSchema.CallToolResult res = FleetMcp.listFleet( McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), architect.isPrimary(),
FleetMcp.leadsVisibleTo(architect), FleetMcp.membersVisibleTo(architect));
String out = textOf(res); String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out); assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
assertTrue(out.contains("\"leads\""), "an architect must still see the leads array: " + out);
assertTrue(out.contains("\"members\""), "an architect must still see the members array: " + out);
} }
/** /**
@@ -842,7 +993,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true)); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true));
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary); assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary); assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
@@ -851,6 +1002,8 @@ class FleetMcpTest {
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary); assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary); assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary); assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"leads\""), "the primary must still see the leads array: " + gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"members\""), "the primary must still see the members array: " + gatedAsPrimary);
} }
@Test @Test
@@ -869,7 +1022,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res); String out = textOf(res);
assertTrue(out.contains("\"msgId\":\"m1\""), out); assertTrue(out.contains("\"msgId\":\"m1\""), out);
@@ -893,7 +1046,7 @@ class FleetMcpTest {
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(), FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(),
Map.of(), "", collaborators, collaboratorsVisible, Map.of(), "", collaborators, collaboratorsVisible,
FleetMcp.CoordinationSource.none(), false); FleetMcp.CoordinationSource.none(), false, true, true);
} }
/** /**
@@ -978,7 +1131,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res); String out = textOf(res);
assertTrue(out.contains("\"pending\":0"), out); assertTrue(out.contains("\"pending\":0"), out);
@@ -1007,7 +1160,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null, workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true); Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res); String out = textOf(res);
assertTrue(out.contains("\"heldDurable\":false"), assertTrue(out.contains("\"heldDurable\":false"),
@@ -1076,7 +1229,7 @@ class FleetMcpTest {
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true); new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true, true, true);
String out = textOf(res); String out = textOf(res);
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out); assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
@@ -1769,8 +1922,8 @@ class FleetMcpTest {
Principal architect = Principal.architect("lead-designer", "term_design", 400); Principal architect = Principal.architect("lead-designer", "term_design", 400);
MemberPresence presence = new MemberPresence(); MemberPresence presence = new MemberPresence();
FleetMcp.markSpawnedMemberPresent(worker, presence); FleetMcp.markTrackedCallerPresent(worker, presence);
FleetMcp.markSpawnedMemberPresent(architect, presence); FleetMcp.markTrackedCallerPresent(architect, presence);
assertTrue(presence.isPresent("term_worker")); assertTrue(presence.isPresent("term_worker"));
assertTrue(presence.isPresent("term_design")); assertTrue(presence.isPresent("term_design"));
@@ -1781,25 +1934,40 @@ class FleetMcpTest {
Principal lead = Principal.leader("opus", "term_lead", 100); Principal lead = Principal.leader("opus", "term_lead", 100);
MemberPresence presence = new MemberPresence(); MemberPresence presence = new MemberPresence();
FleetMcp.markSpawnedMemberPresent(lead, presence); FleetMcp.markTrackedCallerPresent(lead, presence);
FleetMcp.markSpawnedMemberPresent(Principal.anonymous(), presence); FleetMcp.markTrackedCallerPresent(Principal.anonymous(), presence);
assertFalse(presence.isPresent("term_lead")); assertFalse(presence.isPresent("term_lead"));
} }
/**
* The item whose absence would be silent: an observer's own MCP contact must still mark
* presence, or a pane resolving to the unconfigured-pane floor would sit on the injector
* readiness gate forever once something addresses it.
*/
@Test
void anObserverContactMarksPresence() {
Principal observer = Principal.observer("term_observer", 800);
MemberPresence presence = new MemberPresence();
FleetMcp.markTrackedCallerPresent(observer, presence);
assertTrue(presence.isPresent("term_observer"));
}
@Test @Test
void statusReportsLiveAgentStatus() { void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked"); FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked); AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = FleetMcp.status( McpSchema.CallToolResult res = FleetMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a"); new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a", null);
assertNotEquals(Boolean.TRUE, res.isError()); assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res)); assertEquals("blocked", textOf(res));
} }
/** /**
* CB-582: a lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll} * A lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll} — must
* — must also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous * also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
* window it opened with is far shorter than that cadence. * window it opened with is far shorter than that cadence.
*/ */
@Test @Test
@@ -1822,7 +1990,7 @@ class FleetMcpTest {
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline); } while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase()); assertEquals(MessageService.Phase.ASKING, asking.phase());
McpSchema.CallToolResult res = FleetMcp.status(messages, T); McpSchema.CallToolResult res = FleetMcp.status(messages, T, null);
assertNotEquals(Boolean.TRUE, res.isError()); assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res); String out = textOf(res);
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out); assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
@@ -1833,7 +2001,66 @@ class FleetMcpTest {
// Clean up the still-open ask so the background thread does not linger past the test. // Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId(); String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000)); () -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve(T, "done"));
answer.get(5, TimeUnit.SECONDS);
}
/**
* {@code fleet_status}'s pending-ask block (the question, its {@code turnId} and its ticket)
* is shown only to the caller whose owner key created the delegation, or to the unnamed primary.
* A different caller still sees the base status line, but none of the pending-ask fields.
*/
@Test
void statusGatesThePendingAskFieldsByTheDelegationsCreatorOwner() throws Exception {
Principal creator = Principal.worker("term_creator", 1);
String ticket = messages.sendAsync(T, "task that asks", null, creator);
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String other = textOf(FleetMcp.status(messages, T, "worker:term_other"));
assertTrue(other.startsWith("idle"), "the base status must still be shown: " + other);
assertFalse(other.contains("which config file?"),
"a non-creating caller must not see the question text: " + other);
assertFalse(other.contains(asking.turnId()),
"a non-creating caller must not see the turnId: " + other);
assertFalse(other.contains(ticket),
"a non-creating caller must not see the ticket: " + other);
String creatorStatus = textOf(FleetMcp.status(messages, T, creator.ownerKey()));
assertTrue(creatorStatus.contains("which config file?"),
"the creator must see the question: " + creatorStatus);
assertTrue(creatorStatus.contains(asking.turnId()),
"the creator must see the turnId: " + creatorStatus);
assertTrue(creatorStatus.contains(ticket), "the creator must see the ticket: " + creatorStatus);
String unnamed = textOf(FleetMcp.status(messages, T, null));
assertTrue(unnamed.contains("which config file?"),
"a caller with no terminal (the unnamed primary) must see the question: " + unnamed);
// Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, creator.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000; deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) { while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
@@ -1941,6 +2168,27 @@ class FleetMcpTest {
assertTrue(leadOut.contains("\"leader\":\"opus\""), leadOut); assertTrue(leadOut.contains("\"leader\":\"opus\""), leadOut);
} }
/**
* An observer reports its own role and pane, never a {@code leader} key. Without an explicit
* branch it would reach the lead branch by elimination and look right only because the
* {@code leader} key is guarded on a non-null name — this pins the branch rather than the
* accident.
*/
@Test
void whoamiReportsAnObserverNotALead() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = FleetMcp.whoami(
Principal.observer("term_observer", 900), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"observer\""), out);
assertTrue(out.contains("\"sessionId\":\"term_observer\""), out);
assertFalse(out.contains("leader"), out);
}
/** /**
* CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but * CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but
* must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure * must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure
@@ -0,0 +1,280 @@
package dev.ltms.fleet.msg;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Stream;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Pins that no file under {@code src/main/java} calls the fail-open
* {@link MessageService#poll(String)} overload. That overload skips the ownership check in
* {@code MessageService}'s {@code ownsTicket} entirely, so a caller of it can read any session's
* ticket. Every production caller must go through {@link MessageService#poll(String, String)}
* and pass a {@code callerOwner} explicitly, even when it is {@code null}.
*
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
* the risk is a future one-word edit at a call site, not a missing overload.
*
* <p>The scan below finds a violation by its receiver, {@code messages.poll(}, rather than the
* bare method name, so it does not mistake {@link java.util.Queue#poll()} for a violation. That
* anchor only covers a {@code MessageService} reached through a variable or field named
* {@code messages}, so {@link #everyMessageServiceDeclarationIsNamedMessages} pins the naming
* convention the anchor depends on: a declaration under any other name would be invisible to the
* scan above, and must turn this second check red instead of passing silently.
*/
class MessageServicePollUsageTest {
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
@Test
void noProductionFileCallsTheSingleArgumentPollOverload() throws IOException {
List<String> violations = new ArrayList<>();
List<String> twoArgSites = new ArrayList<>();
int filesScanned = scanForPollCalls(PRODUCTION_SOURCE, violations, twoArgSites);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
// violations found" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the absence of violations "
+ "below proves nothing");
assertTrue(violations.isEmpty(), "found a call to the fail-open MessageService.poll(String) "
+ "overload, which skips the ownership check entirely -- pass a callerOwner "
+ "explicitly (even if null) through poll(String, String) instead: " + violations);
// CONTROL: the arity parser actually finds the two genuine two-argument call sites (the
// MCP handler in FleetMcp and the REST handler in FleetApp). If this drops, the parser
// itself is broken, not the production code -- a broken parser (or a scan root that
// reaches no real source) must fail loudly here rather than pass vacuously above.
assertEquals(2, twoArgSites.size(), "control failed: expected exactly the two known "
+ "two-argument messages.poll(...) call sites, found: " + twoArgSites);
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetMcp.java")),
"control failed: did not find the FleetMcp.java messages.poll(ticket, callerOwner) "
+ "site among: " + twoArgSites);
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetApp.java")),
"control failed: did not find the FleetApp.java messages.poll(...) site among: "
+ twoArgSites);
}
/**
* The {@code messages.poll(} anchor above only sees a {@code MessageService} reached through
* a variable, field, or parameter named {@code messages}. This asserts that every such
* declaration under {@code src/main/java} uses that name, so a differently named declaration
* -- invisible to the scan above -- fails loudly here instead of letting that scan pass on a
* call site it never looked at.
*/
@Test
void everyMessageServiceDeclarationIsNamedMessages() throws IOException {
List<String> names = new ArrayList<>();
int filesScanned = scanForDeclarationNames(PRODUCTION_SOURCE, names);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "every
// declaration is named messages" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the result below proves nothing");
// CONTROL: the declaration pattern actually finds real declarations. Zero means the
// pattern is broken, not that every MessageService variable, field, or parameter vanished.
assertTrue(names.size() > 0, "control failed: found zero MessageService declarations under "
+ PRODUCTION_SOURCE + " -- the declaration pattern is broken, update it before "
+ "trusting the naming check below");
List<String> other = names.stream().filter(n -> !n.equals("messages")).distinct().toList();
assertTrue(other.isEmpty(), "found a MessageService declaration not named \"messages\": "
+ other + " -- the messages.poll( scan above only looks for that name, so a call "
+ "through a differently named variable or field is invisible to it; either rename "
+ "the declaration or widen that scan's anchor to cover it");
}
private static int scanForPollCalls(Path root, List<String> violations, List<String> twoArgSites)
throws IOException {
List<Path> files = javaFiles(root);
for (Path file : files) {
scanFileForPollCalls(file, violations, twoArgSites);
}
return files.size();
}
private static int scanForDeclarationNames(Path root, List<String> names) throws IOException {
List<Path> files = javaFiles(root);
Pattern declaration = Pattern.compile("MessageService\\s+([A-Za-z_][A-Za-z0-9_]*)");
for (Path file : files) {
if (file.getFileName().toString().equals("MessageService.java")) {
continue; // the type's own declaration, not a caller holding a reference to it
}
String stripped = stripComments(Files.readString(file));
Matcher m = declaration.matcher(stripped);
while (m.find()) {
int j = m.end();
while (j < stripped.length() && Character.isWhitespace(stripped.charAt(j))) j++;
if (j < stripped.length() && stripped.charAt(j) == '(') {
continue; // a method named like the convention, e.g. "MessageService messages()"
}
names.add(m.group(1));
}
}
return files.size();
}
private static List<Path> javaFiles(Path root) throws IOException {
try (Stream<Path> paths = Files.walk(root)) {
return paths.filter(p -> p.toString().endsWith(".java")).toList();
}
}
/**
* Replaces {@code //} and {@code /* *}{@code /} comment text with nothing, leaving code,
* string/char literals and line breaks untouched -- so a comment that merely mentions
* {@code MessageService} in prose can never be read as a declaration.
*/
private static String stripComments(String source) {
StringBuilder out = new StringBuilder(source.length());
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
out.append(c);
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
if (c == '"') inString = false;
i++;
continue;
}
if (inChar) {
out.append(c);
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
if (c == '\'') inChar = false;
i++;
continue;
}
if (c == '"') { inString = true; out.append(c); i++; continue; }
if (c == '\'') { inChar = true; out.append(c); i++; continue; }
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '/') {
while (i < source.length() && source.charAt(i) != '\n') i++;
continue; // leaves the newline itself for the next iteration to append
}
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '*') {
i += 2;
while (i < source.length() && !(source.charAt(i) == '*' && i + 1 < source.length()
&& source.charAt(i + 1) == '/')) {
if (source.charAt(i) == '\n') out.append('\n');
i++;
}
i += 2;
continue;
}
out.append(c);
i++;
}
return out.toString();
}
private static void scanFileForPollCalls(Path file, List<String> violations, List<String> twoArgSites)
throws IOException {
String source = Files.readString(file);
String needle = "messages.poll(";
int from = 0;
int idx;
while ((idx = source.indexOf(needle, from)) >= 0) {
int argsStart = idx + needle.length();
String args = extractBalancedArgs(source, argsStart, file, idx);
int closeParenIndex = argsStart + args.length();
from = closeParenIndex + 1;
if (args.isBlank()) {
continue; // MessageService has no zero-argument poll() -- nothing to classify
}
String site = file + ":" + lineOf(source, idx);
if (topLevelCommaCount(args) == 0) {
violations.add(site + " -- messages.poll(" + args.trim() + ")");
} else {
twoArgSites.add(site);
}
}
}
/**
* The text between {@code messages.poll(} and its matching close paren: balanced over nested
* calls, and never split by a paren or comma sitting inside a string or char literal.
*/
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
int depth = 1;
boolean inString = false;
boolean inChar = false;
int i = start;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
if (depth == 0) return source.substring(start, i);
}
i++;
}
throw new IllegalStateException(
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
}
/**
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
* separators a human reader would see, not every comma character in the text.
*/
private static int topLevelCommaCount(String args) {
int depth = 0;
int commas = 0;
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < args.length()) {
char c = args.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(' || c == '[' || c == '{') {
depth++;
} else if (c == ')' || c == ']' || c == '}') {
depth--;
} else if (c == ',' && depth == 0) {
commas++;
}
i++;
}
return commas;
}
private static int lineOf(String source, int index) {
int line = 1;
for (int i = 0; i < index; i++) {
if (source.charAt(i) == '\n') line++;
}
return line;
}
}
@@ -4,6 +4,7 @@ import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger; import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent; import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender; import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus; import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr; import dev.ltms.fleet.herdr.FakeHerdr;
@@ -78,7 +79,7 @@ class MessageServiceTest {
} }
private CompletableFuture<MessageService.Reply> sendAsync(String content, long timeoutMillis) { private CompletableFuture<MessageService.Reply> sendAsync(String content, long timeoutMillis) {
return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis)); return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis, null));
} }
private void awaitWaiting() throws InterruptedException { private void awaitWaiting() throws InterruptedException {
@@ -188,7 +189,7 @@ class MessageServiceTest {
localRendezvous, new InMemoryReplyInbox()); localRendezvous, new InMemoryReplyInbox());
String brief = "Implement the requested change. ".repeat(20); String brief = "Implement the requested change. ".repeat(20);
CompletableFuture<MessageService.Reply> send = CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000)); CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000, (String) null));
long deadline = System.currentTimeMillis() + 2000; long deadline = System.currentTimeMillis() + 2000;
while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) { while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5); Thread.sleep(5);
@@ -341,7 +342,7 @@ class MessageServiceTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply. // The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000)); CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
// The worker's ask returns the answer — it resumes the same turn. // The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS); MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
@@ -377,7 +378,7 @@ class MessageServiceTest {
// The primary answers that one turnId; both asks unblock with the same answer. // The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000)); CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS); MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS); MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
@@ -472,7 +473,7 @@ class MessageServiceTest {
"the duplicate's own timeout elapses first"); "the duplicate's own timeout elapses first");
CompletableFuture<MessageService.Reply> answered = CompletableFuture<MessageService.Reply> answered =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500)); CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500, null));
MessageService.AskResult a = fresh.get(5, TimeUnit.SECONDS); MessageService.AskResult a = fresh.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome(), assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome(),
@@ -507,7 +508,7 @@ class MessageServiceTest {
CompletableFuture<MessageService.Reply> lateAnswer = new CompletableFuture<>(); CompletableFuture<MessageService.Reply> lateAnswer = new CompletableFuture<>();
messages.setAskTimeoutRaceHookForTest(() -> messages.setAskTimeoutRaceHookForTest(() ->
lateAnswer.complete(messages.answer(turnId, "too late", 500))); lateAnswer.complete(messages.answer(turnId, "too late", 500, null)));
try { try {
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS); MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.TIMED_OUT, a.outcome()); assertEquals(MessageService.AskOutcome.TIMED_OUT, a.outcome());
@@ -539,7 +540,7 @@ class MessageServiceTest {
@Test @Test
void answeringAnUnknownTurnIsStale() { void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500); MessageService.Reply r = messages.answer(T + "#999", "too late", 500, null);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(), assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang"); "an answer to a turn that never existed (or already lapsed) is stale, not a hang");
} }
@@ -572,7 +573,7 @@ class MessageServiceTest {
messages.setAnswerAskLapseRaceHookForTest(() -> rendezvous.answerAsk(turnId, "raced in first")); messages.setAnswerAskLapseRaceHookForTest(() -> rendezvous.answerAsk(turnId, "raced in first"));
try { try {
MessageService.Reply r = messages.answer(turnId, "too late", 500); MessageService.Reply r = messages.answer(turnId, "too late", 500, null);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(), assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an ask already answered by the race must be seen as lapsed, not double-delivered"); "an ask already answered by the race must be seen as lapsed, not double-delivered");
} finally { } finally {
@@ -596,7 +597,7 @@ class MessageServiceTest {
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() { void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future // Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker. // times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50); MessageService.Reply r = messages.send(T, "never delivered", 50, null);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(), assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working"); "an undelivered send that times out is still queued, not working");
assertNull(r.text()); assertNull(r.text());
@@ -605,7 +606,7 @@ class MessageServiceTest {
@Test @Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception { void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send = CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300)); CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300, null));
awaitWaiting(); awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
@@ -630,7 +631,7 @@ class MessageServiceTest {
void sendTimeoutUsesCancellationDeliveredWhenPickupWinsTheRace() { void sendTimeoutUsesCancellationDeliveredWhenPickupWinsTheRace() {
messages.setTimeoutCancellationRaceHookForTest(() -> injector.onStatus(T, AgentStatus.IDLE)); messages.setTimeoutCancellationRaceHookForTest(() -> injector.onStatus(T, AgentStatus.IDLE));
try { try {
MessageService.Reply reply = messages.send(T, "race delivery", 50); MessageService.Reply reply = messages.send(T, "race delivery", 50, null);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, reply.outcome(), assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, reply.outcome(),
"cancel reporting DELIVERED means the worker received the timed-out message"); "cancel reporting DELIVERED means the worker received the timed-out message");
@@ -652,7 +653,7 @@ class MessageServiceTest {
void sendTimesOutWithAttemptedDeliveryReportsUnconfirmedNotQueued() throws Exception { void sendTimesOutWithAttemptedDeliveryReportsUnconfirmedNotQueued() throws Exception {
herdr.agentSendFailsWith("send_failed"); herdr.agentSendFailsWith("send_failed");
CompletableFuture<MessageService.Reply> send = CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150)); CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150, null));
awaitWaiting(); awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt → ATTEMPTED injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt → ATTEMPTED
@@ -679,7 +680,7 @@ class MessageServiceTest {
// The primary answers, unblocking the worker; but the worker never sends the follow-up // The primary answers, unblocking the worker; but the worker never sends the follow-up
// fleet_reply, so the answering send rides out its short window as still-working. // fleet_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200); MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200, null);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(), assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working"); "an answered worker that never replies times out as still working");
@@ -715,7 +716,7 @@ class MessageServiceTest {
assertEquals(MessageService.Outcome.QUESTION, q.outcome()); assertEquals(MessageService.Outcome.QUESTION, q.outcome());
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000)); CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
ask.get(5, TimeUnit.SECONDS); // worker resumed with the answer ask.get(5, TimeUnit.SECONDS); // worker resumed with the answer
awaitWaiting(); // the answering call has (re)opened its own forward waiter awaitWaiting(); // the answering call has (re)opened its own forward waiter
@@ -723,7 +724,7 @@ class MessageServiceTest {
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS); MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome()); assertEquals(MessageService.Outcome.REPLIED, done.outcome());
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300); MessageService.Reply probe = messages.send(T, "probe after normal reply", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(), assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released after a normal REPLIED answer(), or this bounded " "the session lock must be released after a normal REPLIED answer(), or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work"); + "follow-up send would come back BUSY instead of timing out on its own work");
@@ -756,13 +757,13 @@ class MessageServiceTest {
// The primary answers, unblocking the worker; the worker never sends its follow-up // The primary answers, unblocking the worker; the worker never sends its follow-up
// fleet_reply, so the answering call rides out its short window as still-working. // fleet_reply, so the answering call rides out its short window as still-working.
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200)); CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200, null));
MessageService.Reply answered = answer.get(5, TimeUnit.SECONDS); MessageService.Reply answered = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answered.outcome(), assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answered.outcome(),
"an answered worker that never replies times out as still working"); "an answered worker that never replies times out as still working");
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300); MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(), assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released after a TIMED_OUT_WORKING answer(), or this bounded " "the session lock must be released after a TIMED_OUT_WORKING answer(), or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work"); + "follow-up send would come back BUSY instead of timing out on its own work");
@@ -791,7 +792,7 @@ class MessageServiceTest {
new java.util.concurrent.atomic.AtomicReference<>(); new java.util.concurrent.atomic.AtomicReference<>();
Thread answerer = new Thread(() -> { Thread answerer = new Thread(() -> {
try { try {
messages.answer(turnId, "config.yaml", 5000); messages.answer(turnId, "config.yaml", 5000, null);
caught.set(new AssertionError("expected answer() to throw")); caught.set(new AssertionError("expected answer() to throw"));
} catch (Throwable t) { } catch (Throwable t) {
caught.set(t); caught.set(t);
@@ -813,7 +814,7 @@ class MessageServiceTest {
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300); MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(), assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released when answer()'s reply future fails exceptionally, " "the session lock must be released when answer()'s reply future fails exceptionally, "
+ "or this bounded follow-up send would come back BUSY instead of timing out on its own work"); + "or this bounded follow-up send would come back BUSY instead of timing out on its own work");
@@ -840,7 +841,7 @@ class MessageServiceTest {
new java.util.concurrent.atomic.AtomicReference<>(); new java.util.concurrent.atomic.AtomicReference<>();
Thread answerer = new Thread(() -> { Thread answerer = new Thread(() -> {
try { try {
messages.answer(turnId, "config.yaml", 5000); messages.answer(turnId, "config.yaml", 5000, null);
caught.set(new AssertionError("expected answer() to throw")); caught.set(new AssertionError("expected answer() to throw"));
} catch (Throwable t) { } catch (Throwable t) {
caught.set(t); caught.set(t);
@@ -859,17 +860,147 @@ class MessageServiceTest {
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after interruption", 300); MessageService.Reply probe = messages.send(T, "probe after interruption", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(), assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released when answer()'s wait is interrupted, or this bounded " "the session lock must be released when answer()'s wait is interrupted, or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work"); + "follow-up send would come back BUSY instead of timing out on its own work");
} }
/**
* Drive an async send on {@code T} to a resolved reply, polling as {@code owner} until the
* ticket reports {@link MessageService.Phase#DONE} (or the 2s deadline runs out). {@code owner}
* must be a terminal this ticket's creator check actually accepts, or this loops to the
* deadline and returns a non-{@code DONE} view.
*/
private MessageService.TaskView driveAsyncTicketToDone(String ticket, String owner) throws InterruptedException {
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket, owner);
//noinspection BusyWait
Thread.sleep(5);
}
return view;
}
@Test @Test
void pollReturnsNullForAnUnknownTicket() { void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown"); assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
} }
// --- fleetd #719: a per-boot nonce keeps one instance's ticket ids out of another's space ---
/** A second, fully independent instance — its own agents/injector/rendezvous/inbox, not shared. */
private MessageService newIndependentInstance() {
FakeHerdr otherHerdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
AgentControl otherAgents = new AgentControl(otherHerdr);
Injector otherInjector = new Injector(otherAgents);
return new MessageService(otherAgents, otherInjector, new Rendezvous(), new InMemoryReplyInbox());
}
@Test
void twoInstancesMintDisjointTicketIds() {
MessageService other = newIndependentInstance();
String ticketFromThis = messages.sendAsync(T, "task on first instance", null, null);
String ticketFromOther = other.sendAsync(T, "task on second instance", null, null);
assertNotEquals(ticketFromThis, ticketFromOther,
"each instance mints its own id space, so even a first ticket from each must differ");
}
@Test
void foreignInstanceTicketDoesNotResolve() {
MessageService other = newIndependentInstance();
String ticket = messages.sendAsync(T, "task on first instance", null, null);
// `other` must reach the same sequence number, or this test passes against an empty map
// instead of against a colliding id.
other.sendAsync(T, "task on second instance", null, null);
// control: the id resolves in the instance that minted it, so a null below cannot be
// explained by broken plumbing — only by the ticket being foreign to `other`.
assertNotNull(messages.poll(ticket), "the minting instance must still resolve its own ticket");
assertNull(other.poll(ticket), "a ticket minted by a different instance must not resolve here");
}
// --- ticket ownership -----------------------------------------------------------------------
@Test
void leadBIsRefusedFromLeadAsTicket() throws Exception {
Principal leadA = Principal.leader("opus", "term_a", 1);
Principal leadB = Principal.leader("sol", "term_b", 2);
String ticket = messages.sendAsync(T, "long task", null, leadA);
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "secret async result"), "a reply resolves the async send");
MessageService.TaskView owner = driveAsyncTicketToDone(ticket, leadA.ownerKey());
assertNotNull(owner, "the creator must still be able to read its own ticket");
assertEquals(MessageService.Phase.DONE, owner.phase());
MessageService.TaskView refused = messages.poll(ticket, leadB.ownerKey());
assertNotNull(refused, "a different lead gets a refusal, not silence");
assertNotEquals(MessageService.Phase.DONE, refused.phase(),
"a different lead must never see the ticket as DONE");
assertNull(refused.reply(), "a refusal must never carry the reply text");
assertFalse(String.valueOf(refused).contains("secret async result"),
"the reply text must not appear anywhere in the refused view");
}
@Test
void unnamedPrimaryStillReadsAnyTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task", null,
Principal.leader("opus", "term_lead", 1));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "primary-visible result"), "a reply resolves the async send");
MessageService.TaskView view = driveAsyncTicketToDone(ticket, null);
assertNotNull(view, "the unnamed primary must be able to read any ticket");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("primary-visible result", view.reply());
}
@Test
void namedLeadCanPollItsTicketAfterItsTerminalChanges() throws Exception {
Principal oldLead = Principal.leader("opus", "term_OLD", 1);
Principal newLead = Principal.leader("opus", "term_NEW", 2);
assertNotEquals(oldLead.terminal(), newLead.terminal(), "the test requires different terminals");
String ticket = messages.sendAsync(T, "long task", null, oldLead);
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "own result"), "a reply resolves the async send");
MessageService.TaskView view = driveAsyncTicketToDone(ticket, newLead.ownerKey());
assertNotNull(view, "the same named lead must read the ticket from its new terminal");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("own result", view.reply());
}
@Test
void anonymousOwnerKeyIsRefusedByTheTicketGateItself() {
Principal lead = Principal.leader("opus", "term_lead", 1);
String ticket = messages.sendAsync(T, "long task", null, lead);
MessageService.TaskView refused = messages.poll(ticket, Principal.anonymous().ownerKey());
assertNotNull(refused);
assertEquals(MessageService.Phase.FAILED, refused.phase());
assertEquals("forbidden: this ticket was created by a different session", refused.detail());
}
@Test
void architectOwnershipUsesTerminalRatherThanSlot() {
Principal oldArchitect = Principal.architect("opus", "term_OLD", 1);
Principal newArchitect = Principal.architect("opus", "term_NEW", 2);
String ticket = messages.sendAsync(T, "long task", null, oldArchitect);
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket, oldArchitect.ownerKey()).phase());
assertEquals(MessageService.Phase.FAILED, messages.poll(ticket, newArchitect.ownerKey()).phase());
}
@Test @Test
void pollReportsACompletedTicket() throws Exception { void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task"); String ticket = messages.sendAsync(T, "long task");
@@ -896,11 +1027,11 @@ class MessageServiceTest {
@Test @Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception { void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first = CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000)); CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000, null));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window. // A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100); MessageService.Reply busy = messages.send(T, "second", 100, null);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(), assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang"); "a second send while another holds the session is busy, not a hang");
assertNull(busy.text()); assertNull(busy.text());
@@ -930,13 +1061,13 @@ class MessageServiceTest {
PrimaryRegistry reg = new PrimaryRegistry(null); PrimaryRegistry reg = new PrimaryRegistry(null);
// L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded. // L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded.
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L))); () -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
awaitWaiting(); awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send owns the delegation"); "an accepted send owns the delegation");
// A attempts W while L holds it → BUSY (lock never taken) → its hook never fires. // A attempts W while L holds it → BUSY (lock never taken) → its hook never fires.
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A)); MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A), null);
assertEquals(MessageService.Outcome.BUSY, busy.outcome()); assertEquals(MessageService.Outcome.BUSY, busy.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"a BUSY send must not steal the delegator ownership it never earned"); "a BUSY send must not steal the delegator ownership it never earned");
@@ -957,7 +1088,7 @@ class MessageServiceTest {
void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception { void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null); PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L))); () -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
awaitWaiting(); awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING); injector.onStatus(T, AgentStatus.WORKING);
@@ -968,7 +1099,7 @@ class MessageServiceTest {
// L finished; A's later accepted send takes the delegation over. // L finished; A's later accepted send takes the delegation over.
CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync(
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A))); () -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A), null));
awaitWaiting(); awaitWaiting();
assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(), assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send after the owner finished becomes the new delegator"); "an accepted send after the owner finished becomes the new delegator");
@@ -989,7 +1120,7 @@ class MessageServiceTest {
void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception { void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null); PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L))); () -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
awaitWaiting(); awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation"); assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation");
@@ -1004,7 +1135,7 @@ class MessageServiceTest {
// L answers the ask on the same turn; the answer path must not touch ownership. // L answers the ask on the same turn; the answer path must not touch ownership.
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000)); CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // the answering send reopened its forward waiter awaitWaiting(); // the answering send reopened its forward waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
@@ -1024,7 +1155,7 @@ class MessageServiceTest {
void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() { void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() {
assertThrows(IllegalStateException.class, assertThrows(IllegalStateException.class,
() -> messages.send(T, "doomed", 500, () -> messages.send(T, "doomed", 500,
() -> { throw new IllegalStateException("ownership hook failed"); }), () -> { throw new IllegalStateException("ownership hook failed"); }, null),
"a throwing ownership hook fails the send loudly"); "a throwing ownership hook fails the send loudly");
assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter"); assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter");
@@ -1152,7 +1283,7 @@ class MessageServiceTest {
@Test @Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception { void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send = CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000)); CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
awaitWaiting(); awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned"); assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
@@ -1172,7 +1303,7 @@ class MessageServiceTest {
@Test @Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception { void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send = CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000)); CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
awaitWaiting(); awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer")); assertTrue(rendezvous.resolve(T, "the real answer"));
@@ -1303,7 +1434,7 @@ class MessageServiceTest {
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase()); assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000)); () -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); awaitWaiting();
assertTrue(rendezvous.resolve(T, "done")); assertTrue(rendezvous.resolve(T, "done"));
@@ -1343,7 +1474,7 @@ class MessageServiceTest {
// same turnId must see it as lapsed rather than resolving a question nobody is waiting on. // same turnId must see it as lapsed rather than resolving a question nobody is waiting on.
assertEquals(MessageService.AskOutcome.TIMED_OUT, ask.get(5, TimeUnit.SECONDS).outcome()); assertEquals(MessageService.AskOutcome.TIMED_OUT, ask.get(5, TimeUnit.SECONDS).outcome());
assertEquals(MessageService.Outcome.STALE_TURN, assertEquals(MessageService.Outcome.STALE_TURN,
messages.answer(asking.turnId(), "config.yaml", 200).outcome()); messages.answer(asking.turnId(), "config.yaml", 200, null).outcome());
} }
@Test @Test
@@ -1380,7 +1511,7 @@ class MessageServiceTest {
// The primary answers, but its own bounded wait for the worker's resumed turn is short and // The primary answers, but its own bounded wait for the worker's resumed turn is short and
// expires before the worker (still genuinely working) gets back to it. // expires before the worker (still genuinely working) gets back to it.
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150); MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(), assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait gives up before the worker finishes resuming"); "the primary's own bounded wait gives up before the worker finishes resuming");
@@ -1409,7 +1540,7 @@ class MessageServiceTest {
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000)); CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING); MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150); MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome()); assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome());
@@ -1463,7 +1594,7 @@ class MessageServiceTest {
messages.setFinishAsyncTaskRaceHookForTest(() -> messages.forgetTurnForTest(turnId)); messages.setFinishAsyncTaskRaceHookForTest(() -> messages.forgetTurnForTest(turnId));
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000)); CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // answer() opened its own forward waiter for the resumed worker turn awaitWaiting(); // answer() opened its own forward waiter for the resumed worker turn
@@ -1569,7 +1700,7 @@ class MessageServiceTest {
String turnId = asking.turnId(); String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000)); CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer(), assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer(),
"the worker's own ask() call must have already unblocked with the primary's answer " "the worker's own ask() call must have already unblocked with the primary's answer "
+ "before we force the race below"); + "before we force the race below");
@@ -1616,7 +1747,7 @@ class MessageServiceTest {
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING); MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
String turnId = asking.turnId(); String turnId = asking.turnId();
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150); MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150, null);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(), assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait must give up first, leaving turnId stamped with no " "the primary's own bounded wait must give up first, leaving turnId stamped with no "
@@ -1682,7 +1813,7 @@ class MessageServiceTest {
MessageService.TaskView asking = messages.poll(first); MessageService.TaskView asking = messages.poll(first);
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000)); () -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); awaitWaiting();
assertTrue(rendezvous.resolve(T, "done")); assertTrue(rendezvous.resolve(T, "done"));
@@ -1716,7 +1847,7 @@ class MessageServiceTest {
// The primary answers it — answer() resumes the turn and blocks for what comes next. // The primary answers it — answer() resumes the turn and blocks for what comes next.
CompletableFuture<MessageService.Reply> answer1 = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer1 = CompletableFuture.supplyAsync(
() -> messages.answer(asking1.turnId(), "a1", 5000)); () -> messages.answer(asking1.turnId(), "a1", 5000, null));
assertEquals("a1", ask1.get(5, TimeUnit.SECONDS).answer()); assertEquals("a1", ask1.get(5, TimeUnit.SECONDS).answer());
// Still in the SAME resumed turn — before replying — the worker asks again. // Still in the SAME resumed turn — before replying — the worker asks again.
@@ -1737,7 +1868,7 @@ class MessageServiceTest {
// The primary answers the second question; the worker finally sends its real fleet_reply. // The primary answers the second question; the worker finally sends its real fleet_reply.
CompletableFuture<MessageService.Reply> answer2 = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer2 = CompletableFuture.supplyAsync(
() -> messages.answer(turnId2, "a2", 5000)); () -> messages.answer(turnId2, "a2", 5000, null));
assertEquals("a2", ask2.get(5, TimeUnit.SECONDS).answer()); assertEquals("a2", ask2.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); awaitWaiting();
assertTrue(rendezvous.resolve(T, "done")); assertTrue(rendezvous.resolve(T, "done"));
@@ -1804,20 +1935,20 @@ class MessageServiceTest {
} }
} }
// --- CB-582: fleet_status pendingAsk() ------------------------------------------------------ // --- fleet_status pendingAsk() -----------------------------------------------------------
@Test @Test
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception { void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask"); assertNull(messages.pendingAsk(T, null), "no async ticket at all -> no pending ask");
String ticket = messages.sendAsync(T, "long task"); String ticket = messages.sendAsync(T, "long task");
awaitWaiting(); awaitWaiting();
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question"); assertNull(messages.pendingAsk(T, null), "a plain pending delegation is not a question");
injectDelivery(); injectDelivery();
assertTrue(rendezvous.resolve(T, "done")); assertTrue(rendezvous.resolve(T, "done"));
awaitTicketPhase(ticket, MessageService.Phase.DONE); awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either"); assertNull(messages.pendingAsk(T, null), "a finished ticket carries no open question either");
} }
@Test @Test
@@ -1830,21 +1961,183 @@ class MessageServiceTest {
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000)); CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING); MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.PendingAsk pending = messages.pendingAsk(T); MessageService.PendingAsk pending = messages.pendingAsk(T, null);
assertNotNull(pending, "fleet_status should see the open question"); assertNotNull(pending, "fleet_status should see the open question");
assertEquals(ticket, pending.ticket()); assertEquals(ticket, pending.ticket());
assertEquals("which config file?", pending.question()); assertEquals("which config file?", pending.question());
assertEquals(asking.turnId(), pending.turnId()); assertEquals(asking.turnId(), pending.turnId());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000)); () -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertNull(messages.pendingAsk(T), "an answered question is no longer pending"); assertNull(messages.pendingAsk(T, null), "an answered question is no longer pending");
awaitWaiting(); awaitWaiting();
assertTrue(rendezvous.resolve(T, "done")); assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome()); assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
} }
/**
* A caller's owner key must match the key that created the delegation to see its pending
* question. The unnamed primary always sees it.
*/
@Test
void pendingAskGatesTheQuestionByTheDelegationsCreatorOwner() throws Exception {
Principal creator = Principal.worker("term_creator", 1);
String ticket = messages.sendAsync(T, "task that asks", null, creator);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNull(messages.pendingAsk(T, "worker:term_other"),
"a caller whose key did not create the delegation must not see the question");
MessageService.PendingAsk own = messages.pendingAsk(T, creator.ownerKey());
assertNotNull(own, "the creating caller must see its own open question");
assertEquals("which config file?", own.question());
MessageService.PendingAsk unnamed = messages.pendingAsk(T, null);
assertNotNull(unnamed, "a caller with no terminal (the unnamed primary) must always see the question");
assertEquals("which config file?", unnamed.question());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, creator.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* A task created with no recorded owner (a short {@code sendAsync} overload) must not hand its
* open question to a caller with an owner key. Only the unnamed primary may still see it.
*/
@Test
void pendingAskDeniesATerminalBearingCallerWhenTheTaskRecordsNoCreator() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNull(messages.pendingAsk(T, "worker:term_someone"),
"a caller with an owner key must not see a question whose task records no creator");
assertNotNull(messages.pendingAsk(T, null),
"the unnamed primary must still see it even with no recorded creator");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- fleetd #715: answer() is gated on the caller that owns the turn -----------------------
/**
* A blocking {@code fleet_send} from one caller opens the turn; a different caller's answer is
* refused with no side effect on the rendezvous or the worker's blocked {@code fleet_ask} —
* only the real owner can answer it.
*/
@Test
void aDifferentCallersAnswerIsRefusedForABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, "term_owner"));
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertFalse(rendezvous.isWaiting(T), "the forward waiter closes once the question surfaces");
// The hijack: a different caller answers the SAME turnId.
MessageService.Reply hijacked = messages.answer(q.turnId(), "evil.yaml", 500, "term_attacker");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
"a caller that did not open this turn must be refused, not served");
assertFalse(rendezvous.isWaiting(T), "a refused answer must not open a forward waiter");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(T, rendezvous.askSession(q.turnId()), "a refused answer must leave the ask turn open");
// The control: the real owner answers the same turnId and the turn resumes normally.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(q.turnId(), "config.yaml", 5000, "term_owner"));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* Same hijack and control as the blocking case, but the delegation is opened through the
* fire-and-poll path ({@code sendAsync}) — the owner comes from the ticket's recorded creator
* terminal, not a caller argument threaded through a live blocking call.
*/
@Test
void aDifferentCallersAnswerIsRefusedForAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
Principal owner = Principal.worker("term_owner", 1);
String ticket = messages.sendAsync(T, "task that asks", null, owner);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply hijacked = messages.answer(asking.turnId(), "evil.yaml", 500,
"worker:term_attacker");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
"a caller that did not create this delegation must be refused, not served");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, owner.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
}
@Test
void namedLeadCanSeeAndAnswerAnAskAfterItsTerminalChangesWhileLeadBIsRefused() throws Exception {
Principal oldLead = Principal.leader("opus", "term_OLD", 1);
Principal newLead = Principal.leader("opus", "term_NEW", 2);
Principal leadB = Principal.leader("sol", "term_SOL", 3);
assertNotEquals(oldLead.terminal(), newLead.terminal(), "the test requires different terminals");
String ticket = messages.sendAsync(T, "task that asks", null, oldLead);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNotNull(messages.pendingAsk(T, newLead.ownerKey()),
"the same lead at its new terminal must see the pending ask");
assertNull(messages.pendingAsk(T, leadB.ownerKey()),
"another lead must not see the pending ask");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER,
messages.answer(asking.turnId(), "evil.yaml", 500, leadB.ownerKey()).outcome());
assertFalse(ask.isDone(), "another lead must not resolve the worker's ask");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, newLead.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
}
// --- CB-588: async ticket terminal nudges --------------------------------------------------- // --- CB-588: async ticket terminal nudges ---------------------------------------------------
// //
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes // MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
@@ -2101,7 +2394,7 @@ class MessageServiceTest {
// Clean up the still-open ask so the background thread does not linger past the test. // Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000)); () -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); awaitWaiting();
assertTrue(rendezvous.resolve(T, "done")); assertTrue(rendezvous.resolve(T, "done"));
@@ -2127,7 +2420,7 @@ class MessageServiceTest {
.filter(c -> c.method().equals("agent.prompt")).count(); .filter(c -> c.method().equals("agent.prompt")).count();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync( CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000)); () -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer()); assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); awaitWaiting();
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1); java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
@@ -2520,7 +2813,7 @@ class MessageServiceTest {
void hasQueuedDeliveryIsTrueAfterAnUndeliveredSendTimesOut() { void hasQueuedDeliveryIsTrueAfterAnUndeliveredSendTimesOut() {
// Nothing ever delivers the message (never goes IDLE/BLOCKED), so the send times out with // Nothing ever delivers the message (never goes IDLE/BLOCKED), so the send times out with
// TIMED_OUT_QUEUED — same setup as sendTimesOutBeforeDeliveryIsQueuedNotWorking above. // TIMED_OUT_QUEUED — same setup as sendTimesOutBeforeDeliveryIsQueuedNotWorking above.
MessageService.Reply r = messages.send(T, "never delivered", 50); MessageService.Reply r = messages.send(T, "never delivered", 50, null);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome()); assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome());
assertTrue(messages.hasQueuedDelivery(T), assertTrue(messages.hasQueuedDelivery(T),
@@ -2532,7 +2825,7 @@ class MessageServiceTest {
@Test @Test
void hasQueuedDeliveryClearsOnceTheTargetsNextDeliveryIsAccepted() throws Exception { void hasQueuedDeliveryClearsOnceTheTargetsNextDeliveryIsAccepted() throws Exception {
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome()); assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
assertTrue(messages.hasQueuedDelivery(T)); assertTrue(messages.hasQueuedDelivery(T));
// A fresh send accepts delivery (opens its own waiter) — the stale queued fact is cleared. // A fresh send accepts delivery (opens its own waiter) — the stale queued fact is cleared.
@@ -2549,7 +2842,7 @@ class MessageServiceTest {
@Test @Test
void hasQueuedDeliveryClearsOnAbandon() { void hasQueuedDeliveryClearsOnAbandon() {
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome()); assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
assertTrue(messages.hasQueuedDelivery(T)); assertTrue(messages.hasQueuedDelivery(T));
messages.abandon(T, "session released"); messages.abandon(T, "session released");
@@ -0,0 +1,187 @@
package dev.ltms.fleet.msg;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.stream.Stream;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Pins that no file under {@code src/main/java} calls the fail-open
* {@link Rendezvous#open(String)} overload. That overload opens a forward waiter with no
* recorded {@link Rendezvous.Owner}, so the turn it opens can never be answered by anyone,
* not even the unnamed primary. Every production caller must go through
* {@link Rendezvous#open(String, Rendezvous.Owner)} and record an explicit owner.
*
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
* the risk is a future one-word edit at a call site, not a missing overload.
*
* <p>The scan below finds a violation by its receiver, {@code rendezvous.open(}, classified by
* argument count. The one test method here points that exact scanner at a file known to hold
* many real one-argument calls before it ever looks at production, so a scanner that stops
* matching fails loudly on the known-positive case instead of leaving a clean production result
* looking like evidence it never produced.
*/
class RendezvousOpenUsageTest {
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
private static final Path KNOWN_TEST_CALLER =
Path.of("src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java");
/**
* A pattern that cannot find a known one-argument {@code rendezvous.open(} call would also
* find none in production -- not because production is clean, but because the pattern does
* not match the text it is supposed to catch. That failure mode is exactly what let a
* {@code \b}-based regex read "no callers" under {@code git grep -E} when 81 real ones
* existed: {@code git grep} does not treat {@code \b} as a word boundary, so the pattern
* silently matched nothing anywhere, clean code and real calls alike. This test runs the
* known-positive check first, with the same scanning method the production check then
* depends on, so that mistake fails loudly here instead of reading as a clean result.
*/
@Test
void noProductionFileCallsTheSingleArgumentOpenOverload() throws IOException {
List<String> knownCalls = new ArrayList<>();
int knownFilesScanned = scanForOneArgOpenCalls(KNOWN_TEST_CALLER, knownCalls);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "found
// nothing" having looked at nothing.
assertTrue(knownFilesScanned > 0, "control failed: the scan under " + KNOWN_TEST_CALLER
+ " visited zero .java files -- the path is wrong, so neither result below proves "
+ "anything");
// CONTROL: the scanner actually finds real one-argument rendezvous.open( calls when
// pointed at a file known to hold many. If this is not satisfied, the matching logic
// itself is broken, and the production result below is the scanner failing silently,
// not production code actually being clean.
assertTrue(knownCalls.size() >= 60, "control failed: the scanner found only "
+ knownCalls.size() + " one-argument rendezvous.open( call(s) in " + KNOWN_TEST_CALLER
+ ", which is known to hold many -- the matching logic itself is broken: " + knownCalls);
List<String> violations = new ArrayList<>();
int filesScanned = scanForOneArgOpenCalls(PRODUCTION_SOURCE, violations);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
// violations found" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the absence of violations "
+ "below proves nothing");
assertTrue(violations.isEmpty(), "found a call to the fail-open Rendezvous.open(String) "
+ "overload, which opens a forward waiter with no recorded owner -- record an "
+ "explicit Rendezvous.Owner through open(String, Owner) instead: " + violations);
}
private static int scanForOneArgOpenCalls(Path root, List<String> sites) throws IOException {
List<Path> files = javaFiles(root);
for (Path file : files) {
scanFileForOpenCalls(file, sites);
}
return files.size();
}
private static List<Path> javaFiles(Path root) throws IOException {
try (Stream<Path> paths = Files.walk(root)) {
return paths.filter(p -> p.toString().endsWith(".java")).toList();
}
}
private static void scanFileForOpenCalls(Path file, List<String> sites) throws IOException {
String source = Files.readString(file);
String needle = "rendezvous.open(";
int from = 0;
int idx;
while ((idx = source.indexOf(needle, from)) >= 0) {
int argsStart = idx + needle.length();
String args = extractBalancedArgs(source, argsStart, file, idx);
int closeParenIndex = argsStart + args.length();
from = closeParenIndex + 1;
if (args.isBlank()) {
continue; // Rendezvous has no zero-argument open() -- a javadoc "rendezvous.open()" mention, not a call
}
if (topLevelCommaCount(args) == 0) {
sites.add(file + ":" + lineOf(source, idx) + " -- rendezvous.open(" + args.trim() + ")");
}
}
}
/**
* The text between {@code rendezvous.open(} and its matching close paren: balanced over
* nested calls, and never split by a paren or comma sitting inside a string or char literal.
*/
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
int depth = 1;
boolean inString = false;
boolean inChar = false;
int i = start;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
if (depth == 0) return source.substring(start, i);
}
i++;
}
throw new IllegalStateException(
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
}
/**
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
* separators a human reader would see, not every comma character in the text.
*/
private static int topLevelCommaCount(String args) {
int depth = 0;
int commas = 0;
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < args.length()) {
char c = args.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(' || c == '[' || c == '{') {
depth++;
} else if (c == ')' || c == ']' || c == '}') {
depth--;
} else if (c == ',' && depth == 0) {
commas++;
}
i++;
}
return commas;
}
private static int lineOf(String source, int index) {
int line = 1;
for (int i = 0; i < index; i++) {
if (source.charAt(i) == '\n') line++;
}
return line;
}
}
@@ -132,4 +132,127 @@ class RendezvousTest {
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten"); assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered"); assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
} }
// ── fleetd #715: turn ownership ────────────────────────────────────────────────────────────
@Test
void openWithNoOwnerRecordsNoOwnerAndOpenWithAnOwnerRecordsIt() {
rendezvous.open(W);
assertNull(rendezvous.ownerOf(W), "the no-owner overload records no owner at all");
rendezvous.close(W, rendezvous.currentWaiter(W));
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.ownerOf(W));
}
@Test
void ownerOfIsNullWhenNoWaiterIsOpen() {
assertNull(rendezvous.ownerOf(W), "no waiter open means no owner to report");
}
@Test
void openAskStampsTheForwardWaitersOwnerOntoTheFreshTurnOnly() {
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
Rendezvous.AskTicket fresh = rendezvous.openAsk(W);
assertTrue(fresh.fresh());
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(fresh.turnId()),
"a freshly-opened ask copies the forward waiter's current owner");
Rendezvous.AskTicket coalesced = rendezvous.openAsk(W);
assertFalse(coalesced.fresh());
assertEquals(fresh.turnId(), coalesced.turnId());
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(coalesced.turnId()),
"a coalesced duplicate ask rides the fresh owner's turn, unchanged");
}
@Test
void aSecondAskAfterTheFirstClosesStampsWhateverOwnerIsOpenAtThatLaterMoment() {
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
Rendezvous.AskTicket first = rendezvous.openAsk(W);
rendezvous.closeAsk(first.turnId());
// The forward waiter is reopened under a different owner before the second ask — mirrors
// answer() reopening with the owner it already checked, which can differ turn to turn.
rendezvous.close(W, rendezvous.currentWaiter(W));
rendezvous.open(W, Rendezvous.Owner.of("term_other"));
Rendezvous.AskTicket second = rendezvous.openAsk(W);
assertTrue(second.fresh());
assertEquals(Rendezvous.Owner.of("term_other"), rendezvous.askOwner(second.turnId()),
"a freshly-opened ask always copies whatever owner is open right now, not a stale one");
}
@Test
void askOwnerIsNullForAnUnknownOrLapsedTurn() {
assertNull(rendezvous.askOwner("no-such#1"));
Rendezvous.AskTicket t = rendezvous.openAsk(W);
rendezvous.closeAsk(t.turnId());
assertNull(rendezvous.askOwner(t.turnId()), "a closed ask no longer reports an owner");
}
// ── fleetd #729: per-boot nonce guards turnId against cross-instance reuse ────────────────
@Test
void twoInstancesMintDisjointTurnIds() {
Rendezvous other = new Rendezvous();
Rendezvous.AskTicket fromThis = rendezvous.openAsk(W);
Rendezvous.AskTicket fromOther = other.openAsk(W);
assertNotEquals(fromThis.turnId(), fromOther.turnId(),
"each instance mints its own id space, so even a first ask from each must differ");
}
@Test
void foreignInstanceTurnIdDoesNotResolve() {
Rendezvous other = new Rendezvous();
// `other` must reach the same sequence number as `rendezvous` (two asks each, the first
// closed so the second mints fresh), or this test passes against an empty map instead of
// against a colliding id.
Rendezvous.AskTicket firstFromThis = rendezvous.openAsk(W);
rendezvous.closeAsk(firstFromThis.turnId());
Rendezvous.AskTicket secondFromThis = rendezvous.openAsk(W);
Rendezvous.AskTicket firstFromOther = other.openAsk(W);
other.closeAsk(firstFromOther.turnId());
other.openAsk(W);
// control: the id resolves in the instance that minted it, so a false below cannot be
// explained by broken plumbing — only by the turnId being foreign to `other`.
assertTrue(rendezvous.answerAsk(secondFromThis.turnId(), "answer from this instance"),
"the minting instance must still resolve its own turnId");
assertFalse(other.answerAsk(secondFromThis.turnId(), "answer from other instance"),
"a turnId minted by a different instance must not resolve here");
}
@Test
void openAskStillCoalescesDuplicatesAndStillMintsDistinctIdsPerAsk() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertEquals(t1.turnId(), t2.turnId(),
"a second openAsk while one is open still coalesces onto the same turn");
assertFalse(t2.fresh(), "the coalesced ask is still reported as not fresh");
rendezvous.closeAsk(t1.turnId());
Rendezvous.AskTicket t3 = rendezvous.openAsk(W);
assertNotEquals(t1.turnId(), t3.turnId(), "two asks from the same session still get different turnIds");
}
@Test
void ownerPermitsIsFailClosedOnARecordAndThreeStatesAreDistinct() {
assertFalse(Rendezvous.Owner.permits(null, null),
"no owner on record refuses even the unnamed primary");
assertFalse(Rendezvous.Owner.permits(null, "worker:term_a"),
"no owner on record refuses a caller with an owner key too");
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, null),
"the recorded unnamed primary matches a caller with a null owner key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, "worker:term_a"),
"the recorded unnamed primary does not match another owner key");
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), "worker:term_a"),
"an owner matches the same key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), "worker:term_b"),
"an owner refuses a different key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), null),
"a named owner refuses the unnamed primary");
}
} }
@@ -2,11 +2,13 @@ package dev.ltms.fleet.rest;
import ch.qos.logback.classic.Level; import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent; import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper; import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.auth.CallerResolver; import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.Authz; import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.MemberRegistry; import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal; import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.auth.Role;
import dev.ltms.fleet.config.FleetConfig; import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard; import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl; import dev.ltms.fleet.herdr.AgentControl;
@@ -19,6 +21,7 @@ import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics; import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.msg.MessageService; import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous; import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.FakeWorktrees; import dev.ltms.fleet.session.FakeWorktrees;
import dev.ltms.fleet.session.SessionManager; import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.member.ClaudeCodeLauncher; import dev.ltms.fleet.member.ClaudeCodeLauncher;
@@ -39,6 +42,8 @@ import java.util.List;
import java.util.Locale; import java.util.Locale;
import java.util.Map; import java.util.Map;
import java.util.Set; import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import java.util.function.Predicate; import java.util.function.Predicate;
import java.util.regex.Matcher; import java.util.regex.Matcher;
import java.util.regex.Pattern; import java.util.regex.Pattern;
@@ -139,6 +144,413 @@ class FleetAppAuthTest {
assertThrows(IllegalArgumentException.class, () -> FleetApp.routeAction("GET /healthz")); assertThrows(IllegalArgumentException.class, () -> FleetApp.routeAction("GET /healthz"));
} }
// --- GET /tasks/{ticket} must not be the no-check overload ----------------------------------
/**
* {@code GET /tasks/{ticket}} must resolve its caller the same way {@code allow(...)} does
* and thread that owner key into {@link MessageService#poll(String, String)}, not the
* no-check overload that ignores who is asking.
*/
@Test
void theTaskStatusRouteActuallyThreadsTheCallersOwnerKeyIntoPoll() throws Exception {
String source = Files.readString(REST_SOURCE);
int start = source.indexOf("private void taskStatus(Context ctx) {");
assertTrue(start >= 0, "could not find taskStatus in " + REST_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("private static void herdrError(Context ctx, HerdrException e) {", start);
assertTrue(end > start, "could not find the method declared after taskStatus to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call messages.poll(...) -- if this fails, the
// anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("messages.poll("),
"control failed: the scraped taskStatus block contains no messages.poll( call at all "
+ "-- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("caller.ownerKey()"),
"the taskStatus route must thread the resolved caller's owner key into messages.poll(...), "
+ "not the no-check overload -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
"the taskStatus route must resolve its caller the same way allow(...) does, not via a "
+ "second, separate resolution path -- block: " + handlerBlock);
}
/**
* {@code GET /sessions/{id}/status} must resolve its caller the same way {@code allow(...)}
* does and thread that owner key into {@link MessageService#pendingAsk(String, String)}, not
* the no-check overload that ignores who is asking.
*/
@Test
void theSessionStatusRouteActuallyThreadsTheCallersOwnerKeyIntoPendingAsk() throws Exception {
String source = Files.readString(REST_SOURCE);
int start = source.indexOf("private void sessionStatus(Context ctx) {");
assertTrue(start >= 0, "could not find sessionStatus in " + REST_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("private void taskStatus(Context ctx) {", start);
assertTrue(end > start, "could not find the method declared after sessionStatus to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call messages.pendingAsk(...) -- if this fails,
// the anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("messages.pendingAsk("),
"control failed: the scraped sessionStatus block contains no messages.pendingAsk( "
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("caller.ownerKey()"),
"the sessionStatus route must thread the resolved caller's owner key into "
+ "messages.pendingAsk(...), not the no-check overload -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
"the sessionStatus route must resolve its caller the same way allow(...) does, not via "
+ "a second, separate resolution path -- block: " + handlerBlock);
}
/**
* {@code GET /tasks/{ticket}} refuses a worker whose terminal did not create the ticket,
* while the creating worker and the unnamed primary both still read it. The ticket is minted
* directly on the shared {@link MessageService}, the same way {@code MessageServiceTest}
* drives {@link MessageService#poll(String, String)}, so this exercises only the REST poll
* route's own handling of the ownership already recorded on the ticket.
*/
@Test
void restPollRefusesADifferentWorkerButAllowsTheCreatorAndTheUnnamedPrimary() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
try {
String ticket = messages.sendAsync("term_a", "long task", null,
Principal.worker("term_a", FakeHerdr.WORKER_PID));
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, refused.statusCode());
assertTrue(refused.body().contains("forbidden"),
"a different worker's terminal must be refused, not shown the ticket: " + refused.body());
assertFalse(refused.body().contains("\"reply\""),
"a refusal must never carry reply text: " + refused.body());
HttpResponse<String> own = send(creatorApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the creating worker must read its own ticket: " + own.body());
HttpResponse<String> primary = send(primaryApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, primary.statusCode());
assertFalse(primary.body().contains("forbidden"),
"the unnamed primary must read any ticket: " + primary.body());
} finally {
creatorApp.stop();
otherWorkerApp.stop();
primaryApp.stop();
}
}
/**
* {@code POST /sessions/{id}/message} with {@code wait:false} must record the creating
* caller's own terminal on the ticket it returns, so that caller can still poll its own
* ticket over REST, while a different terminal is refused.
*/
@Test
void restSendAsyncRecordsTheCreatingCallersTerminalSoItCanStillPollItsOwnTicket() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin leadApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-x"));
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
try {
ObjectMapper mapper = new ObjectMapper();
HttpResponse<String> created = send(leadApp.port(), "POST", "/sessions/term_a/message",
"{\"content\":\"long task\",\"wait\":false}", null);
assertEquals(202, created.statusCode(), created.body());
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
HttpResponse<String> own = send(leadApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the session that created the ticket over REST must be able to poll it: " + own.body());
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, refused.statusCode());
assertTrue(refused.body().contains("forbidden: this ticket was created by a different session"),
"a different terminal must still be refused with the ownership detail, not some "
+ "other rejection: " + refused.body());
} finally {
leadApp.stop();
otherWorkerApp.stop();
}
}
/**
* {@code GET /sessions/{id}/status} shows a worker's pending {@code fleet_ask} question, its
* {@code turnId} and its ticket only to the caller whose owner key created that delegation, or
* to the unnamed primary. A different caller still sees the base status line, but none of the
* pending-ask fields.
*/
@Test
void restStatusGatesThePendingAskFieldsByTheDelegationsCreatorOwner() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
try {
ObjectMapper mapper = new ObjectMapper();
String ticket = messages.sendAsync("term_target", "task that asks", null,
Principal.worker("term_a", FakeHerdr.WORKER_PID));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket, null);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
JsonNode other = mapper.readTree(
send(otherWorkerApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("idle", other.get("status").asText(), "the base status must still be shown");
assertFalse(other.has("question"), "a non-creating caller must not see the question: " + other);
assertFalse(other.has("turnId"), "a non-creating caller must not see the turnId: " + other);
assertFalse(other.has("ticket"), "a non-creating caller must not see the ticket: " + other);
JsonNode own = mapper.readTree(
send(creatorApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("which config file?", own.get("question").asText(), "the creator must see the question");
assertEquals(turnId, own.get("turnId").asText(), "the creator must see the turnId");
assertEquals(ticket, own.get("ticket").asText(), "the creator must see the ticket");
JsonNode primary = mapper.readTree(
send(primaryApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("which config file?", primary.get("question").asText(),
"a caller with no terminal (the unnamed primary) must see the question");
// Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, "worker:term_a"));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
answer.get(5, TimeUnit.SECONDS);
} finally {
creatorApp.stop();
otherWorkerApp.stop();
primaryApp.stop();
}
}
/**
* {@code POST /sessions/{id}/message} with a {@code turnId} refuses a caller whose terminal
* did not open the turn, even when that caller otherwise holds ANSWER rights, and leaves the
* turn open for the real owner to resolve. Covers the REST adapter's ANSWER gate for an
* async-send ({@code wait:false}) delegation.
*/
@Test
void restAnswerIsRefusedForADifferentCallerOnAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
try {
ObjectMapper mapper = new ObjectMapper();
HttpResponse<String> created = send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"long task\",\"wait\":false}", null);
assertEquals(202, created.statusCode(), created.body());
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "the async send should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket, null);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
assertEquals(403, hijacked.statusCode(), hijacked.body());
assertTrue(hijacked.body().contains("not_turn_owner"),
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
assertEquals("term_target", rendezvous.askSession(turnId),
"a refused answer must leave the ask turn open");
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
"the real owner's answer must resolve the worker's turn");
} finally {
ownerApp.stop();
attackerApp.stop();
}
}
/**
* As above, but the delegation is a blocking send ({@code wait:true}) instead of a ticket —
* the owning caller's own HTTP request is the one that surfaces the worker's question and
* later carries the real answer. Covers the REST adapter's ANSWER gate for a blocking-send
* delegation.
*/
@Test
void restAnswerIsRefusedForADifferentCallerOnABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
try {
ObjectMapper mapper = new ObjectMapper();
CompletableFuture<HttpResponse<String>> blocking = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"long task\",\"wait\":true,\"timeoutMs\":5000}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "the blocking send should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
HttpResponse<String> questionResponse = blocking.get(5, TimeUnit.SECONDS);
assertEquals(202, questionResponse.statusCode(), questionResponse.body());
JsonNode question = mapper.readTree(questionResponse.body());
assertEquals("question", question.get("status").asText());
String turnId = question.get("turnId").asText();
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
assertEquals(403, hijacked.statusCode(), hijacked.body());
assertTrue(hijacked.body().contains("not_turn_owner"),
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
assertEquals("term_target", rendezvous.askSession(turnId),
"a refused answer must leave the ask turn open");
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
"the real owner's answer must resolve the worker's turn");
} finally {
ownerApp.stop();
attackerApp.stop();
}
}
/**
* As {@link #start}, but shares {@code messages} and {@code herdr} across several app
* instances bound to different pids, each returned as its own started {@link Javalin} rather
* than through the shared {@code app} field, so several differently-resolved callers can
* poll the same ticket.
*/
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid) {
return startOnSharedService(messages, herdr, pid, Map.of());
}
/**
* As {@link #startOnSharedService(MessageService, FakeHerdr, long)}, but {@code leadTerminals}
* resolves the given pid's terminal to a named lead (a caller with SEND permission) instead of
* a plain worker, for a test that needs a terminal-bearing caller able to create a ticket.
*
* <p>Every connecting pane not already claimed by {@code leadTerminals} is wired into the live
* roster as a spawned worker, so a caller's resolved role matches what its own test expects:
* a {@link Role#WORKER}, never the unconfigured-pane {@link Role#OBSERVER} floor a roster-less
* resolver would otherwise fall to.
*/
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid,
Map<String, String> leadTerminals) {
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> leadTerminals, new MemberRegistry(null),
t -> leadTerminals.containsKey(t) ? null : MemberRole.DEV, Map::of);
Metrics appMetrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
return new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, appMetrics).build().start("127.0.0.1", 0);
}
/** /**
* fleetd #669 Unit A: {@code POST /sessions/{id}/message} is two call shapes behind one route, * fleetd #669 Unit A: {@code POST /sessions/{id}/message} is two call shapes behind one route,
* mirroring {@code fleet_send}'s MCP-side split into {@link Authz.Action#SEND} and {@link * mirroring {@code fleet_send}'s MCP-side split into {@link Authz.Action#SEND} and {@link
@@ -0,0 +1,92 @@
package dev.ltms.fleet.session;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import java.util.List;
import java.util.Set;
/**
* {@link PeerLauncher} decorator that marks presence for a spawned terminal before returning its
* handle to the caller — the contact-then-register ordering fleetd #722 covers, where the
* terminal's MCP contact lands before {@link SessionManager#acquire} runs its own
* {@code registry.put}. The presence view is set after construction, via {@link #presence},
* because it is owned by the {@link SessionManager} this launcher is passed into.
*/
final class PresenceRacingLauncher implements PeerLauncher {
private final PeerLauncher delegate;
volatile MemberPresence presence;
PresenceRacingLauncher(PeerLauncher delegate) {
this.delegate = delegate;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle handle = delegate.spawn(req);
presence.markPresent(handle.terminalId());
return handle;
}
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
PeerHandle handle = delegate.spawn(req, decision);
presence.markPresent(handle.terminalId());
return handle;
}
@Override
public Set<Capability> capabilities() {
return delegate.capabilities();
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return delegate.capabilitiesFor(profileName);
}
@Override
public Set<String> profiles() {
return delegate.profiles();
}
@Override
public String defaultProfile() {
return delegate.defaultProfile();
}
@Override
public String effectiveCwd(SpawnRequest req) {
return delegate.effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return delegate.parityOverlay(profileName);
}
@Override
public List<?> list() {
return delegate.list();
}
@Override
public int reapOrphanWorkers() {
return delegate.reapOrphanWorkers();
}
@Override
public void stop(String id) {
delegate.stop(id);
}
@Override
public boolean clearContext(String id) {
return delegate.clearContext(id);
}
}
@@ -422,6 +422,53 @@ class SessionManagerTest {
"turn completion moves BUSY → DONE"); "turn completion moves BUSY → DONE");
} }
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
@Test
void registerThenContactReachesReadyForPlainSpawn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that arrives after registration reaches READY");
}
@Test
void contactThenRegisterStillReachesReadyForPlainSpawn() {
// The racing launcher marks presence for the spawned terminal from inside spawn() —
// before SessionManager.acquire's own registry.put runs — modeling an MCP contact that
// lands in that window.
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
PresenceRacingLauncher race = new PresenceRacingLauncher(workers);
SessionManager sessions = new SessionManager(race);
race.presence = sessions.asPresence();
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that lands before registry.put must still reach READY");
}
@Test
void aTerminalNeverMarkedPresentStaysSpawningAfterRegistration() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertEquals(MemberSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
"registration alone must not advance a terminal that was never marked present");
}
@Test @Test
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() { void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr(); FakeHerdr herdr = new FakeHerdr();
@@ -1405,6 +1452,105 @@ class SessionManagerTest {
+ "dirty check threw"); + "dirty check threw");
} }
// --- fleetd #736: a release must forget the member's presence entry, not just its registry
// row ---------------------------------------------------------------------------------------
@Test
void releaseByPaneIdForgetsThePresenceEntry() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the release");
sessions.release(session.paneId());
assertFalse(sessions.asPresence().isPresent(terminal),
"release must forget the terminal's presence, not just remove its registry row");
}
@Test
void reapIdleForgetsThePresenceEntryToo() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the reap");
clock[0] = 11;
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
assertFalse(sessions.asPresence().isPresent(terminal),
"the idle-reap release path (releaseIfCurrent) goes through the same teardown "
+ "funnel as an explicit release, so it must forget presence too");
}
@Test
void shutdownDrainAlsoForgetsThePresenceEntry() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the drain");
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertFalse(sessions.asPresence().isPresent(terminal),
"a shutdown drain still ends the member's process, so presence must be cleared "
+ "exactly as it is for any other release cause");
}
@Test
void releaseOfAnUnknownPaneIdDoesNotThrow() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
assertDoesNotThrow(() -> sessions.release("no-such-pane"),
"releasing a pane id that was never registered must be a no-op, not a throw");
}
@Test
void releaseStillForgetsPresenceWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-736", null));
String terminal = s.terminalId();
sessions.asPresence().markPresent(terminal);
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a throwing dirty check must not abort the release");
assertFalse(sessions.asPresence().isPresent(terminal),
"presence must be forgotten even when the dirty check throws, which pins the "
+ "forget call to the finally block that runs no matter what happened above");
}
@Test
void releaseLeavesADifferentStillLiveMembersPresenceUntouched() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession released = sessions.acquire("ltms-local", null, "/caller/a", "ownerA");
MemberSession stillLive = sessions.acquire("ltms-local", null, "/caller/b", "ownerB");
sessions.asPresence().markPresent(released.terminalId());
sessions.asPresence().markPresent(stillLive.terminalId());
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
"present before the release of the other member");
sessions.release(released.paneId());
assertFalse(sessions.asPresence().isPresent(released.terminalId()),
"the released terminal is forgotten");
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
"a still-live member's presence must survive an unrelated release");
}
// --- fleetd #316: the dirty check must be re-taken after the worker is stopped, not trusted // --- fleetd #316: the dirty check must be re-taken after the worker is stopped, not trusted
// stale from before it ------------------------------------------------------------------------ // stale from before it ------------------------------------------------------------------------
@@ -2436,4 +2582,160 @@ class SessionManagerTest {
return "threw:" + e.getClass().getName() + ":" + e.getMessage(); return "threw:" + e.getClass().getName() + ":" + e.getMessage();
} }
} }
// --- fleetd #702: spawnedMemberRole must still answer for a pane mid-teardown ----------------
/**
* A {@link Worktrees} test double whose {@code hasUncommitted} runs an injected hook before
* answering. This is what lets a test resolve a releasing pane's terminal from inside the
* window {@link SessionManager#release} opens between removing the registry entry and
* actually stopping the pane — the hook runs synchronously on the release call's own thread,
* at the exact point {@code release} shells out to {@code git status}, so there is no sleep
* and no race to land in it.
*/
private static final class HookedWorktrees implements Worktrees {
private Runnable hook;
private boolean dirty = false;
HookedWorktrees onHasUncommitted(Runnable hook) {
this.hook = hook;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
if (hook != null) {
hook.run();
}
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
return java.util.Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
}
@Test
void spawnedMemberRoleResolvesTheLiveRegisteredRole() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertEquals(s.role(), sessions.spawnedMemberRole(s.terminalId()),
"a registered session resolves to its own role");
assertNull(sessions.spawnedMemberRole("term_unknown"),
"a terminal with no session at all resolves to null");
}
@Test
void spawnedMemberRoleIsNullOnceReleaseFullyCompletes() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, new HookedWorktrees());
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702a", null));
sessions.release(s.paneId());
assertNull(sessions.spawnedMemberRole(s.terminalId()),
"once release has fully finished, the terminal is neither registered nor releasing");
}
@Test
void spawnedMemberRoleStillAnswersBetweenTheRegistryRemovalAndThePaneStop() {
FakeHerdr herdr = new FakeHerdr();
java.util.concurrent.atomic.AtomicReference<MemberRole> duringWindow = new java.util.concurrent.atomic.AtomicReference<>();
HookedWorktrees worktrees = new HookedWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702b", null));
worktrees.onHasUncommitted(() -> {
// This runs from INSIDE release()'s git-status shell-out: the registry entry is
// already gone, but the pane has not stopped yet — a real call landing inside the
// exact window fleetd #702 reports, so no sleep and no race is needed to reach it.
assertTrue(sessions.get(s.paneId()).isEmpty(),
"sanity: the registry entry is already gone at this point");
duringWindow.set(sessions.spawnedMemberRole(s.terminalId()));
});
sessions.release(s.paneId());
assertEquals(s.role(), duringWindow.get(),
"spawnedMemberRole must still answer the live role while the pane is mid-teardown, "
+ "not only while the session is still in the registry");
assertNull(sessions.spawnedMemberRole(s.terminalId()),
"and once release has fully finished, the window is closed too");
}
/**
* {@code Releasing.enter} and {@code Releasing.leave} are called directly here, never through
* {@link SessionManager#release} or {@link SessionManager#spawnedMemberRole}, so a test naming
* one of them exercises only that one — a regression in the other can never hide behind it.
*/
@Test
void releasingLeaveStepsDownADepthGreaterThanOneInsteadOfRemovingIt() {
SessionManager.Releasing depthTwo = new SessionManager.Releasing(2, "term_a", MemberRole.DEV);
SessionManager.Releasing afterLeave = depthTwo.leave();
assertNotNull(afterLeave,
"depth 2 means another release of the SAME pane is still mid-teardown; leave() must "
+ "step the depth down, never remove the marker outright — removing it here "
+ "is what a plain Set would do, and would reopen the window the still-in-"
+ "flight release is relying on staying closed");
assertEquals(1, afterLeave.depth());
assertEquals("term_a", afterLeave.terminalId());
assertEquals(MemberRole.DEV, afterLeave.role());
}
@Test
void releasingEnterPreservesThePriorTerminalWhenTheOverlappingCallHasNoSessionOfItsOwn() {
SessionManager.Releasing prior = new SessionManager.Releasing(1, "term_a", MemberRole.DEV);
SessionManager.Releasing afterEnter = SessionManager.Releasing.enter(prior, null);
assertEquals(2, afterEnter.depth(), "depth still increments whether or not this enter knows its session");
assertEquals("term_a", afterEnter.terminalId(),
"an overlapping release that finds the registry entry already gone has no session of "
+ "its own to pass as known, and must not blank out the terminal the first "
+ "call already recorded — that terminal is what spawnedMemberRole matches "
+ "against");
assertEquals(MemberRole.DEV, afterEnter.role());
}
} }
@@ -159,6 +159,40 @@ class WorktreeSessionManagerTest {
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path"); assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
} }
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
@Test
void registerThenContactReachesReadyForWorktreeSpawn() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-722", null));
sessions.asPresence().markPresent(session.terminalId());
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that arrives after worktree registration reaches READY");
}
@Test
void contactThenRegisterStillReachesReadyForWorktreeSpawn() {
// The racing launcher marks presence for the spawned terminal from inside spawn() —
// before SessionManager.acquireWithWorktree's own registry.put runs — modeling an MCP
// contact that lands in that window.
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
PresenceRacingLauncher race = new PresenceRacingLauncher(workerService(herdr));
SessionManager sessions = new SessionManager(race, worktrees);
race.presence = sessions.asPresence();
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-722", null));
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that lands before worktree registration must still reach READY");
}
@Test @Test
void worktreeArchitectAcquireAlsoBindsItsSlot() { void worktreeArchitectAcquireAlsoBindsItsSlot() {
FakeHerdr herdr = new FakeHerdr(); FakeHerdr herdr = new FakeHerdr();