Compare commits

...

91 Commits

Author SHA1 Message Date
Dai Ha efd9cdb983 fleetd #737 units 1+2: key tickets and turns on a stable owner, not a terminal
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m4s
Add Principal.ownerKey(): role-prefixed, keyed on name for a named lead and
a collaborator (survives a handover's terminal change), on terminal for a
worker, architect and observer, and on a distinct "anonymous" value for an
unauthenticated caller so that case no longer relies on Authz refusing it
first. The unnamed primary keeps a null key, preserving its primary-wide
ticket rule.

Thread that key through Task.creatorOwner, Rendezvous.Owner, poll,
pendingAsk and answer in place of a raw terminal, in both the MCP and REST
surfaces, so a named lead whose terminal changes can still poll and answer
its own delegations while a different lead is refused both.

Mutation evidence (each one-line change killed a named test, then reverted
to green):
- Principal.ownerKey() PRIMARY case made unconditional (dropped the
  null-name guard) -> ownerKeyCoversEveryRole dies:
  "expected: <null> but was: <leader:null>"
- OBSERVER case changed to use the "worker" prefix -> ownerKeyCoversEveryRole
  dies: "expected: <observer:term_observer> but was: <worker:term_observer>"
- prefixed() changed to drop the role prefix entirely -> both
  ownerKeyCoversEveryRole and rolePrefixesKeepLeadAndArchitectKeysDistinct
  die on a lead/architect key collision: "expected: <leader:opus> but was:
  <opus>"
- PRIMARY case changed to key on terminal instead of name ->
  rolePrefixesKeepLeadAndArchitectKeysDistinct and ownerKeyCoversEveryRole
  die: "expected: <leader:opus> but was: <leader:term_lead>"
- ARCHITECT case changed to key on the slot name instead of terminal ->
  same two tests die: "expected: <architect:opus> but was: <architect:design>"

All five mutations were caught by the existing test suite; no test needed
adding.
2026-10-04 20:42:44 +02:00
Dai Ha aabecce901 Merge remote-tracking branch 'origin/worker/736-presence-forget-f35144-9'
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m58s
2026-10-04 20:19:52 +02:00
Dai Ha 6754b4edbc fleetd #736: release clears the member's presence entry
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 1m45s
SessionManager.releaseRemoved() tore down a member's registry row and pane
but never cleared it from MemberPresence, so a terminal stayed marked
"present" for the daemon's lifetime after release/idle-reap/shutdown drain.
Clear it in the method's unconditional finally block, alongside the other
must-always-run teardown step, so every release path (explicit release,
the idle reaper's releaseIfCurrent, and a shutdown drain) forgets it the
same way, and a throw from the dirty-worktree check does not skip it.

MemberPresence.forget(null) throws NullPointerException (verified empirically:
ConcurrentHashMap.remove(null) NPEs on key.hashCode()), so the new call guards
on a non-null, non-blank terminal id rather than relying on forget to no-op.
2026-10-04 20:15:34 +02:00
Dai Ha 787ae0ed7a fleetd #726: tell the handover skill which rollover behaviour is live
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m59s
Unit 2 replaces the /clear continuation with a real process restart, so every
paragraph in the skill that describes /clear goes false the moment the new jar
is deployed. The code is not merged yet, and a merge is not a deployment, so
rewriting those paragraphs now would hand a lead doing a handover tonight a
document that does not match the daemon it is talking to.

Add a dated note instead. It states that the /clear text stays accurate while
the old jar runs, and gives a test a lead can apply with no shell: the
fleet_handover tool description is served by the running daemon, so if it still
says "clear your pane", the old behaviour is live. It also names the two things
that change, including the one that doubles as a second indicator -- the
"never observed as WORKING after 8 consecutive IDLE/DONE polls" warning cannot
appear once the wait that logs it is deleted. The note names the condition for
deleting itself.

Re-measure the roll evidence while here. The skill recorded four
"lead-rollover: rolled" lines from 2026-09-22; the log now holds 20, against a
control of 86 "lead-rollover:" lines, and "Unknown command" still returns 0.
Add the elapsed spread (median 16507 ms, max 48261 ms, two above 45000 ms) with
the caveat that it times the whole roll and is dominated by the wait for the
calling turn to end, so a slow roll is not a failed one.

Markdown only, no code touched, so no build was run.
2026-10-04 20:09:58 +02:00
Dai Ha e3050efe8b fleetd #705: correct the stale reason on the TASK_READ gate
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m48s
The comment said ticket ids are a sequential counter with no owner check, so a
holder could walk every ticket and read another session's reply. PRs #712 and
#716 added that owner check: MessageService.ownsTicket compares a ticket's
creatorTerminal to the caller on every read.

The rule is still right, so only the reason changes. This matters now because
fleetd #737 is deciding ticket ownership across a lead handover, and a reader
who believed the old text could delete the TASK_READ restriction on the grounds
that its stated reason no longer applies.

Comment-only. mvn -o clean install: Tests run: 2083, Failures: 0, Errors: 0,
BUILD SUCCESS. Flagged by the #705 option-1 worker as out of its scope, which
was the right call.
2026-10-04 19:59:45 +02:00
Dai Ha 11998cd626 Merge remote-tracking branch 'origin/worker/705-observer-14c258-6'
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 56s
CI / build (push) Failing after 1m54s
2026-10-04 19:52:34 +02:00
Dai Ha 8e5394f63f fleetd #705 option 1: narrow the unconfigured-pane floor to OBSERVER
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m57s
Adds Role.OBSERVER as the bottom rung CallerResolver falls to when a
herdr pane matches no live roster entry, lead, architect slot, or
collaborator tab. An observer may only READ/METRICS and REPLY/ASK on
its own pane. Widens the presence gate so an observer's MCP contact
still marks it deliverable, matching what already happens for a
worker or architect, so a pane that outlives a daemon restart is not
left permanently undeliverable.

Ships as defence in depth alongside the already-merged ticket-owner
check (#712/#716), which closed the reachable exploit this ticket
reported.
2026-10-04 19:46:54 +02:00
Dai Ha 428a12af62 fleetd #722: reconcile presence that arrives before a session's registry entry
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m48s
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m46s
A member whose MCP contact lands between launcher.spawn() and registry.put()
had its presence marked, but the SPAWNING -> READY transition that markPresent
triggers found no registry entry yet and silently did nothing. The mark then
persisted while registration left the session in SPAWNING, with nothing to
retry the transition. That left the session undeliverable to reclaim/seat
accounting even though it was present and deliverable.

Add SessionManager.reconcilePresence, called right after registry.put in both
the plain-spawn and worktree-spawn paths, to retry the transition for a
terminal already marked present. One private helper serves both call sites.

Tests cover both orderings (contact-then-register and register-then-contact)
for both spawn paths, plus a terminal never marked present staying in
SPAWNING. The contact-then-register tests use a new PresenceRacingLauncher
test double that marks presence from inside spawn(), before acquire()'s own
registry.put runs.
2026-10-04 19:23:02 +02:00
Dai Ha a332dfdb2c Merge remote-tracking branch 'origin/worker/726-ea34a0-2'
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m53s
2026-10-04 19:09:03 +02:00
Dai Ha d0f4ae057b fleetd #726 unit 3 fix: release the single-flight claim when continuationRunner rejects the hand-off
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 1m51s
confirm() takes the per-lead-terminal claim before handing the roll to
continuationRunner, and the only release path was runRollover's own
finally. If continuationRunner.accept itself throws, runRollover never
starts, so that finally never runs, and nothing else ever writes
rollingByTerminal — the claim is held forever and the terminal can never
be rolled again. This differs from fleetd #615, which covers a throw
INSIDE the continuation (runRollover already catches that and still
releases the claim) — this is a throw from the hand-off itself, which
is not reachable with today's virtual-thread runner but would be with
a bounded executor's RejectedExecutionException.

confirm() now catches that throw, releases the claim, and overwrites
the IN_PROGRESS outcome with a terminal FAILED one, matching how a
throw inside the continuation is already surfaced.
2026-10-04 19:06:57 +02:00
Dai Ha 7b3beaa209 Merge remote-tracking branch 'origin/worker/726-10cbf0-1'
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m53s
2026-10-04 19:06:47 +02:00
Dai Ha b3b2bf3da6 fleetd #726 unit 1 review fixes: correct relaunch's javadoc and dedupe its resolve logic
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 1m49s
relaunch's javadoc said it reads the live config; it actually reads the
FleetConfig snapshot this launcher was constructed with (fleet.leaders
is the frozen half), so say that and note a profile/tab edit needs a
daemon restart.

relaunch's recognise-only refusal reused ensureLeads()'s log wording,
which claims the lead 'is not live' — true in ensureLeads()'s context
(reached only after a short live count), false in relaunch's (which
never counts liveness, by design). Dropped that clause.

Pulled the declared/creatable/profile-configured resolution shared by
ensureLeads() and relaunch() into one private resolveLaunchable(name)
helper (returns a new ResolvedLead(lead, profile) record, or null
having logged), so the three refusals and their wording live in one
place instead of two copies that can drift. Behaviour-preserving:
ensureLeads() keeps its own liveness-count logic around the shared
resolve, and the existing 37 LeadLauncherTest cases are unchanged and
still pass.
2026-10-04 18:59:56 +02:00
Dai Ha 8bb2aa0be4 Merge remote-tracking branch 'origin/worker/729-5961c6-3'
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 2m0s
2026-10-04 18:59:54 +02:00
Dai Ha 4ce3149bfd fleetd #726 unit 3: make LeadRollover.confirm() single-flight per lead terminal
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m48s
Two open() calls for the same lead terminal minted two tokens that both
passed confirm()'s ownership check, so both could reach the deferred
continuation and roll the same lead twice. confirm() now claims a
per-lead-terminal slot (an atomic put-if-absent) once every other gate has
passed, refusing a concurrent confirm with the new ROLL_ALREADY_RUNNING
reason; runRollover releases the claim in a finally, on both the success
and the thrown-exception path.
2026-10-04 18:53:53 +02:00
Dai Ha 1fc9e85bf1 fleetd #729: fold a per-boot nonce into every turnId
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 1m57s
askSeq restarts at 0 on every daemon boot, so a turnId (session#n)
minted by one Rendezvous instance could be minted again by a later
instance and resolve to an unrelated ask. Fold a per-instance nonce
into the mint, the same way #719 fixed MessageService's ticket ids.
2026-10-04 18:53:05 +02:00
Dai Ha 38544d467c fleetd #726 unit 1: give LeadLauncher a public single-lead relaunch seam
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m48s
Adds LeadLauncher.relaunch(name), which starts exactly the named lead
from the live config, outside of ensureLeads()'s instances bookkeeping.
It retries the whole launch attempt (not just the agent_name_taken/
agent_pane_busy cases ResilientAgentLaunch already retries inside one
agents.start call) up to RELAUNCH_ATTEMPTS times.

launch() now returns the started Agent (null on failure) instead of a
boolean, so relaunch() and ensureLeads() share the same primitive.
2026-10-04 18:52:38 +02:00
Dai Ha 809b7d9b20 Merge remote-tracking branch 'origin/worker/727-ee14ed-3'
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 47s
CI / build (push) Failing after 1m56s
2026-10-04 18:29:28 +02:00
Dai Ha 8cf7215d56 Merge remote-tracking branch 'origin/worker/719-bdd95e-4'
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m58s
2026-10-04 18:19:55 +02:00
Dai Ha 1ef93e57cc fleetd #727: give a lead launch the three protections every member spawn gets
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m59s
Extract checkPaneCommandFits, the agent_pane_busy retry, and the agent_name_taken
retry out of HerdrPeerLauncher into a shared dev.ltms.fleet.herdr.ResilientAgentLaunch,
and route LeadLauncher.launch through the same seam instead of a bare agents.start
call. The lead's agent name now carries a per-process nonce and a per-start sequence
number (like a member's), so a stale agent_name_taken from an earlier crashed session
no longer blocks a legitimate relaunch outright.
2026-10-04 18:14:06 +02:00
Dai Ha cf0c9b9316 fleetd #719: make the foreign-id test reach a colliding sequence number
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m49s
other.sendAsync had never been called, so other's tasks map was empty and
poll(ticket) returned null regardless of whether the nonce existed — the
test passed against an empty map, not against a colliding id. Mint once on
other so it reaches the same sequence number as the first instance, making
the test exercise the actual collision the nonce guards against.
2026-10-04 18:14:01 +02:00
Dai Ha 337dbd491e fleetd #719: fold a per-boot nonce into every ticket id
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 2m0s
ticketSeq restarted at zero on every daemon boot with no persistence, so a
ticket id minted in one boot could be reused by a later boot and resolve to
an unrelated Task instead of failing to resolve at all. Mint each
MessageService instance's own short nonce once and fold it into every ticket
(task-<nonce>-<n>), so an id from one instance can never match another's id
space.

Adds a disjoint-id-space test and a foreign-instance-ticket test (with the
positive control) in MessageServiceTest.
2026-10-04 18:05:06 +02:00
Dai Ha 7cf6075b79 Correct the configDir note: the count was wrong, the shared dir is the defect
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 44s
CI / build (push) Failing after 1m57s
The previous commit compared 4 configDir lines against 8 total profiles, but
four of those are opencode and never read CLAUDE_CONFIG_DIR. Every claude-code
profile does set one, so the original claim was right and this file said
otherwise.

The real defect is narrower: opus and sonnet name the operator's own config dir,
so for those members the store is shared, and ClaudeCodeLauncher's javadoc
already records that fleetd and the operator's session write that same file.
2026-10-04 10:35:33 +02:00
Dai Ha fbdcd709c9 The plugin addendum claimed every profile sets configDir; it does not
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 42s
CI / build (push) Failing after 2m1s
GET /profiles reports 8 live profiles and fleetd.yaml carries 4 configDir
lines, two of them pointing at the operator's own instance dir. The bullet
stated the blanket claim as a structural limit, so it would have been believed.
The conclusion it supported is unchanged: member-facing assets travel in the
worktree.
2026-10-04 10:31:07 +02:00
Dai Ha 8d3f10d291 Releasing is reached directly by its test, and its javadoc names the real case
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 59s
CI / build (push) Failing after 2m0s
The javadoc said a losing releaseIfCurrent CAS is the call that arrives with no
known session. releaseIfCurrent is only called by the reaper, always with a
non-null expected, so it always has one; the null-known call is an overlapping
release that finds the registry entry already gone. The same wrong claim was in
a test's failure message.

Releasing, enter and leave drop private so SessionManagerTest binds them at
compile time. The six reflection helpers are gone, and a rename now breaks the
build instead of a test run.
2026-10-04 10:25:17 +02:00
Dai Ha 38f4fd64ee Merge PR #724: fleetd #702 — a pane mid-teardown keeps its member identity 2026-10-04 10:25:08 +02:00
Dai Ha 3fc39b981d Only the delegation's creator can answer its member's question
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 56s
CI / build (push) Failing after 1m39s
fleetd #715 gates fleet_send{turnId} on the caller that created the delegation,
so the shipped block had to say so: the step-5 note already covered who can see
a pending question, not who can answer it.
2026-10-04 10:15:24 +02:00
Dai Ha 70328ca0f8 Merge PR #725: fleetd #715 — gate fleet_send{turnId} on the caller that created the delegation
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m47s
2026-10-04 10:12:38 +02:00
Dai Ha efab9b8c49 fleetd #702: pin the architect-demotion window and split a thin test
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 2m8s
Add a CallerResolverTest case proving a releasing architect resolves as
WORKER while still inside releaseRemoved's teardown window, with a
control resolve outside the window that must stay ARCHITECT so the
in-window assertion cannot pass against a slot that was never bound.

Split spawnedMemberRoleSurvivesAnOverlappingReleaseThatUnmarksEarly into
two SessionManagerTest cases, each reaching Releasing.enter/leave
directly through reflection so a mutation to one invariant (the depth
count in leave, or enter's prior-terminal preservation) can only fail
its own test.
2026-10-04 10:04:00 +02:00
Dai Ha 2374de28e4 fleetd #715: pin that no production caller uses Rendezvous.open(String)
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Failing after 2m4s
Add a source-scrape test mirroring #718's MessageServicePollUsageTest
shape, for the same residual: a convenience overload that defaults the
turn's owner to null, left in place because deleting it would break 81
test-only call sites across 7 unrelated files.

The scanner is exercised against a file known to hold many real
one-argument rendezvous.open( calls before it is ever pointed at
production, using the identical matching logic for both. A pattern that
cannot find the known calls would also find none in production, and
that is exactly the failure mode a -based git grep regex hit earlier
on this ticket: git grep's -E engine does not treat \b as a word
boundary, so that pattern silently matched nothing anywhere, in clean
code and in the 81 real calls alike.
2026-10-04 09:56:05 +02:00
Dai Ha dd18bd1f38 fleetd #715: gate answer() on the caller that owns the turn
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 2m15s
Record the turn's owner on the forward rendezvous waiter (Rendezvous.Owner,
a three-state record: no record / unnamed primary / named terminal). A
fresh fleet_ask copies that owner onto the ask turn; a coalesced duplicate
ask keeps the first owner. answer() compares the answering caller against
the stored owner before taking the session lock or reopening the resumed
waiter, and a mismatch returns the new NOT_TURN_OWNER outcome instead of
STALE_TURN, with no rendezvous/task/question cleanup.

send() and answer() both drop their no-caller overloads; every call site
in MessageService, FleetMcp and FleetApp now threads an explicit caller
terminal through. AuthzTest pins that the unnamed-primary null allowance
is safe only because ANONYMOUS never reaches ANSWER.

Covers all four combinations of capture input x adapter: blocking-send and
async-send delegations, answered through both FleetMcp and FleetApp, each
with a hijack attempt refused and the real owner's answer succeeding as a
control.
2026-10-04 09:46:20 +02:00
Dai Ha efeffb4ab7 fleetd #702: mark a pane as mid-teardown so a caller still resolves as a member
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m51s
A caller from a pane being torn down used to be briefly absent from the
session registry, so CallerResolver fell through to a lead/architect tab map
for the same terminal and could resolve the wrong role.

SessionManager now keeps a depth-counted "releasing" marker per pane, written
before the registry removal and cleared in a finally once release finishes.
A depth count (not a Set) is needed because two threads can race teardown of
the same pane; a Set-based unmark by the losing thread would reopen the
window while the winning thread is still mid-teardown. Both removal sites
(the unconditional release() and the idle reaper's CAS releaseIfCurrent())
go through one shared helper, so all four entry routes (release, the
context-cap release in completeTurn, reapIdle, drainSnapshot) are covered.

SessionManager.spawnedMemberRole is the one reader: it checks the live
registry first (via the existing no-copy findByTerminal), then the releasing
marker. FleetdAssembly now wires this method reference instead of its own
untested inline lambda, which also drops a roster() list copy + stream from
the per-request hot path. CallerResolver is unchanged — its contract already
fit.
2026-10-04 09:27:09 +02:00
Dai Ha ed4f4b08ad callerTerminal's javadoc names which callers carry a terminal
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m59s
It said "null if the primary", which reads as every primary. Only the unnamed
primary has no terminal; a named lead carries one, so a reader deciding what a
null means was given the wrong rule.
2026-10-04 09:11:08 +02:00
Dai Ha 28a1f3d6f5 fleet_status's pending-ask block is shown only to the delegation's creator
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 53s
CI / build (push) Failing after 1m54s
A lead polling a member no longer sees a question raised under an architect's
brief, so an absent question is not evidence there is none.
2026-10-04 08:57:14 +02:00
Dai Ha b12d70716b Merge PR #723: fleetd #721 — gate fleet_status's pending-ask block by the delegation's creator
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m41s
2026-10-04 08:53:42 +02:00
Dai Ha 0f5985b419 fleetd #721: gate fleet_status's pending-ask block by the delegation's creator terminal
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m55s
fleet_status (MCP) and GET /sessions/{id}/status (REST) handed any TASK_READ
holder another session's open fleet_ask question, its turnId and its ticket,
with no check that the caller created that delegation. MessageService.pendingAsk
now takes the caller's terminal and reuses the existing ownsTicket comparison;
FleetMcp.status and FleetApp.sessionStatus both thread the resolved caller
terminal through. The base status line and REST's ready field are unaffected.
2026-10-04 08:46:38 +02:00
Dai Ha 6e06058b07 Merge PR #720: fleetd #718 — pin that no production caller uses MessageService.poll(String)
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 50s
CI / build (push) Failing after 1m44s
2026-10-04 08:23:58 +02:00
Dai Ha e5f4fb81ab fleetd #718: pin the naming convention the poll-usage scan depends on
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Failing after 2m22s
MessageServicePollUsageTest's messages.poll( receiver anchor only
covers a MessageService reached through a variable, field, or
parameter named "messages". Adds a second assertion in the same
class that every such declaration under src/main/java uses that
name, with its own file-walk and declaration-count controls, so a
future declaration under a different name turns this check red
instead of leaving the original scan silently blind to it.

Also tightens the Authz.permits(Principal, Action, String) javadoc
sentence to read as a plain contract statement.
2026-10-04 08:18:54 +02:00
Dai Ha d28ab0968b fleetd #718: pin that no production caller uses MessageService.poll(String)
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 1m47s
Adds a source-scrape test over src/main/java that fails if any caller
reaches the fail-open single-argument poll(String) overload instead of
poll(String, String). The scan anchors on the "messages.poll(" receiver
to avoid matching java.util.Queue.poll(), and balances parentheses to
avoid being fooled by a two-argument call whose first argument contains
nested parens.

Also documents Authz.permits(Principal, Action, String) as a test
convenience whose default classifier denies every collaborator.
2026-10-04 08:09:33 +02:00
Dai Ha 9a64d42599 Merge PR #716: fleetd #705 — gate the REST ticket routes by the creating caller
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m57s
Three parts. GET /tasks/{ticket} passed no caller, so it used the overload that skips the ownership
check and any worker could read any ticket; it now passes the caller resolved from the same CALLER
attribute the authorization gate reads. The wait:false send path recorded no creator terminal, so a
REST-created ticket matched no terminal-bearing caller and its own creator was refused; it now
records one. Both handlers gained a scrape guard with its own control assertion.

Resolved one conflict in FleetMcpAuthzTest by keeping both sides: PR #717 and this branch each
appended tests at the same point. The test count is the check on that resolution — 2014 + 3 + 1 =
2018, so no test was dropped.

Verified in a throwaway worktree off main: Tests run: 2018, Failures: 0, BUILD SUCCESS. Three
mutations, each confirmed live with mvn -o compile before the suite ran. FleetMcp:527 is the one
that survived before this work and now kills theFleetPollHandlerActuallyThreadsCallerTerminalIntoPoll.
FleetApp:698 kills the new creator test. FleetApp:899 kills three, including the behavioural test.
Every file restored byte-identical.
2026-10-04 07:48:18 +02:00
Dai Ha f4e0ca41e6 Merge PR #717: fleetd #710 — gate fleet_list's leads and members arrays by caller role
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m52s
A worker now gets neither array, and the key is absent rather than present-and-empty. leads stays
visible to the primary, an architect and a collaborator; members only to the primary and an
architect. That split follows the rule the collaborators array already states: you may list what
you could address. A collaborator may send to a lead, and leads is the only place the bridge gives
it that address, so hiding it would have left a shipped grant unusable.

Verified in a throwaway worktree off main: Tests run: 2014, Failures: 0, BUILD SUCCESS, against a
main baseline of 2008 that I measured myself. Dropping the collaborator clause from leadsVisibleTo
compiled green and then killed exactly two tests, the truth-table row and the behavioural test.
2026-10-04 07:41:21 +02:00
Dai Ha 29a2f97c25 fleetd: thread the creating caller's terminal into the REST fire-and-poll send path
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m53s
sendMessage's wait:false branch now records the resolved caller's own
terminal as the ticket's creatorTerminal, the same way taskStatus already
resolves its caller, so a REST-created ticket's own creator can still poll
it under the ownership check that now gates GET /tasks/{ticket}.
2026-10-04 07:41:12 +02:00
Dai Ha d2db8c7dc9 fleetd #710 PR #717 review: leadsVisibleTo must also admit a collaborator
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m47s
A collaborator may SEND to a lead, and fleet_list's leads array is the
only place this tool gives it a lead's sessionId -- its own
fleet_whoami carries no lead address. leadsVisibleTo now returns true
for caller.isCollaborator() as well as primary and architect, matching
the existing rule for the sibling collaborators array (every role that
may SEND to a named peer). membersVisibleTo is unchanged: a
collaborator may never SEND to a spawned member.

Updated the truth-table tests for both predicates, added a behavioural
test proving a collaborator's fleet_list output contains leads and not
members, and updated fleet_list's tool description.
2026-10-04 07:37:55 +02:00
Dai Ha 95311c6e8e Correct two stale claims in the canonical bridge block
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m5s
CI / build (push) Failing after 2m4s
The block is the instruction surface this repo ships, so a merged change that makes it false is an
incomplete change. Two had gone stale:

- the lead table said fleet_list does not report collaborators. #709 made it report them, visible
  to the primary, an architect and another collaborator, never a worker.
- the collaborator section justified the ticket refusal by saying ticket ids are a plain counter
  with no owner check. #712 added that owner check, so the stated reason no longer held. The role
  gate is what refuses a collaborator; the recorded creator terminal is the second line.

wiki/7-Use-Cases.md carries the same edit and was pushed to its own remote, verified by ref. The
sync check prints True.
2026-10-04 07:34:23 +02:00
Dai Ha 10ab58e4fc fleetd #710: gate fleet_list's leads and members arrays by caller role
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m50s
Omit the leads and members keys entirely (never an empty array) from
fleet_list's result for a worker, matching the existing
coordinatorVisibleTo/collaboratorsVisibleTo pattern: two new named
predicates (leadsVisibleTo, membersVisibleTo) are consulted before
assembling either array, so a worker holding only READ can no longer
read every session on the daemon through this tool. Primary and
architect callers are unaffected.
2026-10-04 07:25:55 +02:00
Dai Ha 0ba597e394 fleetd #705: close the REST ticket-poll door and pin both handlers' caller-terminal wiring
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m47s
GET /tasks/{ticket} now resolves the caller the same way allow(...) does and
threads that terminal into MessageService.poll(ticket, callerTerminal) instead
of the no-check overload, so a worker can no longer read a ticket a different
session created over REST. Adds a source-scrape guard (with its own control
assertion) for both the fleet_poll MCP handler and this REST route, plus a
behavioural test driving GET /tasks/{ticket} with three differently-resolved
callers against one shared MessageService.
2026-10-04 07:23:29 +02:00
Dai Ha c468953963 Merge PR #714: fleetd #711 — re-key two startup refusals onto the hazard that is still real
CI / shell-tests (push) Failing after 8s
CI / build (push) Failing after 1m53s
CI / contract (push) Successful in 50s
The roster is consulted ahead of every tab map, so a live registered member
is no longer read back as a lead. Both refusals still matter, for the narrower
case where the pane is alive and the roster holds no entry for it.

Behaviour unchanged. Verified in a throwaway worktree, not piped:
Tests run: 2008, FleetConfigTest 170, BUILD SUCCESS.
2026-10-04 07:04:21 +02:00
Dai Ha 25d53e6ef7 Merge PR #713: fleetd #710 — remove the unused CallerResolver.members() accessor
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 2m3s
It returned architectTerminals, so its name contradicted its contents and
collided with fleet_list's members array. No production caller.

Verified in a throwaway worktree, not piped: Tests run: 2008, BUILD SUCCESS.
2026-10-04 07:00:35 +02:00
Dai Ha a9a37af957 Merge PR #712: fleetd #705 — gate fleet_poll's ticket lookup by the creating caller
Closes the MCP door only. The REST door (GET /tasks/{ticket}) still calls the
no-check poll overload, so #705 stays open.

Verified in a throwaway worktree, not piped: Tests run: 2008, BUILD SUCCESS.
2026-10-04 07:00:35 +02:00
Dai Ha 7d497aa423 fleetd #711: re-key two pane-tab hazard reasons onto the unregistered-pane condition
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m39s
Both refusals said a member landing in a lead's or collaborator's
labelled tab would be read back as that identity. CallerResolver
consults the spawned-member roster ahead of every tab map, so a live
registered member is never misread this way. State the real condition
instead: the hazard applies only while the pane is alive and carries
no entry in the spawned-member roster.
2026-10-04 06:54:43 +02:00
Dai Ha b205bcc2aa Remove CallerResolver members accessor
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m47s
2026-10-04 06:53:39 +02:00
Dai Ha 11eccc3a1b fleetd #705: gate fleet_poll's ticket lookup by the creating caller's terminal
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m45s
A ticket id is a plain sequential counter, so any session holding
TASK_READ could walk task-1, task-2, ... and read another session's
delegation reply. sendAsync now records the creating caller's terminal
on the Task, and poll refuses a caller whose terminal differs from it.
A caller with no terminal (the unnamed primary) is still allowed
through regardless, since it never carries a herdr pane to compare.
2026-10-04 06:50:26 +02:00
Dai Ha 4939d40362 Merge PR #709: fleetd #703 — fleet_list reports collaborators, narrowed by role
CI / shell-tests (push) Failing after 12s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 1m53s
A collaborator could find and message a lead, but a lead had nowhere to read a
collaborator's sessionId, so the channel only worked once the collaborator
spoke first. fleet_list now carries a collaborators array, and
CallerResolver.collaborators() has its first caller outside the resolver.

Visible to the primary, an architect and a collaborator; absent for a worker.
The rule is that the roles which may see a collaborator row are the roles that
can act on it, and a worker cannot send to a collaborator at all. Narrowed in
the payload the way coordinatorVisibleTo already does it, because fleet_list is
gated on READ and a worker holds READ.

Each row carries the registry name and the sessionId only. A collaborator is
never spawned and has no profile, so a lead row's context and seat fields have
no meaning for it.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 2005, Failures: 0.

Nothing yet pins the handler to collaboratorsVisibleTo, unlike the coordinator
flag, so a literal passed there would regress silently. Tracked in #710.
2026-10-04 06:42:05 +02:00
Dai Ha e7a7711a4e fleetd #703: fleet_list reports a collaborators array so a lead can discover one
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m54s
Threads CallerResolver#collaborators() into fleet_list's canonical listFleet
overload and adds a collaborators array (name, sessionId), visible only to
the primary, an architect, and a collaborator -- the same roles Authz grants
SEND to a named peer -- never a worker. Gives CallerResolver#collaborators()
its first real caller outside the resolver itself.
2026-10-04 06:37:48 +02:00
Dai Ha 1fd2cfa716 Merge PR #708: fleetd #669 — fleetd.example.yaml promised a defence that is not wired
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m10s
CI / build (push) Failing after 2m15s
The example file said two things stop the tab-name convention from becoming a
way to claim leadership, and named workspace exclusion as the first. Production
builds the scanner with an empty excludedWorkspaceLabels, so that defence does
not exist. It now names the two that do: the startup refusal of a colliding
tabPrefix/tabLabel, and CallerResolver asking the live spawned-member roster
before any tab map.

It also said a lead's workspace default is "leads" and must not be a member
workspace. The default is "fleet", the same space the members use, so that
advice was against the shipped shape.

Adds the test nobody wrote: a lead is still discovered when its workspace is
the member workspace, with an empty exclusion set. Corrects a test javadoc that
called itself the guard that matters while covering a parameter production never
passes.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 2000, Failures: 0.
2026-10-04 06:35:49 +02:00
Dai Ha 3e8e314656 Merge PR #706 and PR #707: fleetd #669 — a collaborator is deliverable, and a collaborator config change is reported
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m0s
CI / build (push) Failing after 1m46s
Two follow-ups to #669, verified together in one throwaway worktree.

PR #707 (f820737): ConfigRef reports that a fleet.collaborators change needs a
restart. The map is read once at startup and a reload does not rebuild it, so
before this an operator saw a clean reload and no live effect.

PR #706 (20fc42b): deliverableTo opens for a collaborator. A collaborator is
never in MemberPresence and never found by the lead scan, so every send to one
was authorized, then held on the injector gate for ~60s and dropped without a
keystroke reaching the pane.

mvn clean install on the merged tree: exit 0, BUILD SUCCESS, Tests run: 1999,
Failures: 0, Errors: 0. ConfigRefTest 32, FleetDeliverabilityTest 9.
2026-10-04 06:19:47 +02:00
Dai Ha 20fc42b572 fleetd #669 follow-up: deliverableTo also opens for a collaborator
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Failing after 1m48s
A collaborator's terminal was authorized at Authz but never deliverable: it
is never enrolled in MemberPresence and never discovered by the lead scan,
so a send to a collaborator sat on the injector's readiness gate for
~60s and failed, never typed into the pane. deliverableTo now takes a
third collaborators supplier, read through on each call like the lead
supplier, and FleetdAssembly wires the existing collaboratorTerminals
supplier into it.
2026-10-04 06:13:03 +02:00
Dai Ha f82073717a fleetd #669: report collaborator reload restart
CI / shell-tests (pull_request) Failing after 11s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 2m1s
2026-10-04 06:12:37 +02:00
Dai Ha 16b52fac6c fleetd.example.yaml: the collaborator block works now (#669)
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m49s
The comment above the collaborators: example said the block 'is parsed and
validated today; nothing yet recognises or addresses the tab it names'. That
was true when Unit B landed the config shape. Units C, D and E have since
merged and deployed, so the tab is recognised, the role is authorized, and
the pane routes to the lead herdr daemon.

This is the worst place for that sentence to go stale: it sits in the file an
operator copies to turn the feature on, and it tells them the block does
nothing. Replaced with what the role may and may not do, taken from the rows
in auth/Authz.java rather than from the wiki, plus the #703 limit that a lead
cannot discover a collaborator.

FleetConfigTest (which loads this file) green at 170 tests.
2026-10-04 02:30:28 +02:00
Dai Ha 13b6ae2628 Merge PR #704: fleetd #669 Unit E — route a collaborator's pane to the lead herdr daemon
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 2m5s
A configured collaborator's terminal now routes to the lead herdr daemon
rather than the member one. FleetdAssembly ORs a second terminal map into
the predicate it hands HerdrRouter, and the predicate's field is renamed
isLead -> routeToLead so its name matches the widened contract.

Verified in the lead's own throwaway worktree: mvn clean install green at
1992 tests, 0 failures (174 surefire report files, fresh worktree). The
fix rests on one test, so I reproduced the mutation myself rather than
taking the worker's word: dropping the collaborator clause from the
predicate turns FleetdAssemblyCollaboratorHerdrRoutingTest red at runtime
with two distinct AgentControl identities, while HerdrRouterTest's 4 tests
stay green. That confirms both the kill and that the two-distinct-fakes
setup is actually discriminating.

Also checked by hand: both terminal maps come from LeadTabScanner's one
byKind helper and are keyed terminal_id -> name, so containsKey(target) is
correct for both; no agentsFor call runs between router construction and
the ref being set, so the publication window is harmless and matches the
one leadsRef already had.

One reviewer fanned out against the diff, briefed from the diff rather
than the implementer's rationale. It reported no issue.
2026-10-04 02:19:44 +02:00
Dai Ha 7e838ba8b9 fleetd #669 Unit E: route a collaborator's pane to the lead herdr daemon
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m57s
HerdrRouter.agentsFor picked the member daemon for any terminal the lead
predicate did not recognize, so a configured collaborator's pane (opened
by a person, exactly like a lead's) was routed to the member herdr
daemon instead of the lead one.

FleetdAssembly now combines the leads map and the collaborator-terminals
map into the predicate it hands HerdrRouter. HerdrRouter's isLead field
and constructor parameter are renamed to routeToLead, with its javadoc
naming the real contract: true for any terminal whose pane lives in the
lead daemon, lead or collaborator.
2026-10-04 02:11:17 +02:00
Dai Ha c3e3554bde CLAUDE.md: the collaborator role, after #669 Unit D
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m50s
The block is the instruction surface this repo ships, and Unit D made
four of its statements false. A session can now resolve as
collaborator, read "Which role am I?", and find no section telling it
what it may do.

Each claim below was read from the merged tree at b5bc5d4.

  - fleet_whoami returns four roles, not three. FleetMcp.whoami returns
    before the lead branch, so a collaborator gets its registry name
    and its own sessionId, and no leader key.
  - The fallback ladder said only that it cannot separate a worker from
    an architect. It also fires for no collaborator at all: every rung
    detects a spawned member, and nothing launched a collaborator. It
    therefore falls to "act as a worker", which is the safe direction
    but leaves it unable to learn what it is without asking.
  - Invariant 3 said send is lead or architect. Authz SEND also allows
    a collaborator when the target passes knownLeadOrCollaborator, so
    it may reach a lead or another collaborator and no spawned member.
  - The invariants heading said "both roles"; there are now four.

New Collaborator section, placed after the member turn contract because
a collaborator must be told that contract is not its own. It owes no
fleet_reply: nothing delegates to it. Its two surprising limits are
written down rather than left to be discovered, and both were measured
in Authz: it cannot reach a worker, and it cannot read a ticket,
because TASK_READ is withheld where READ is not.

The intent table gains a row for messaging a collaborator, and that row
states the gap instead of implying discovery works: fleet_list emits
only "leads" and "members", and CallerResolver.collaborators() has no
caller outside the resolver, so a lead cannot find a collaborator
unless the collaborator tells it its sessionId.

The wiki template is byte-identical again, verified with the project's
own check in the main clone, which printed False before the propagation
and True after.
2026-10-04 01:58:03 +02:00
Dai Ha b5bc5d4ab5 Merge PR #701: fleetd #669 Unit D — the resolver, a live spawned member outranks every tab map
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m55s
2026-10-04 01:43:17 +02:00
Dai Ha d83972ede2 fleetd #669 Unit D correction: confirm live slot role before granting a spawned architect
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Failing after 1m44s
The new spawned-member roster step returned Principal.architect from the roster's
own MemberRole.ARCHITECT alone, never reconfirming against memberSlotRoles. A slot
revoked after the bind kept granting ARCHITECT to the already-bound session,
regressing fleetd #424's "config governs what a bound slot still grants" half. The
roster now only answers that the pane is a live spawned member; a confirmed live
slot role still decides whether that grants ARCHITECT, falling through to WORKER
otherwise — mirroring the existing architect-slot step a few lines below.

Also corrects LeadTabScanner.buildTabIndex's javadoc: a colliding lead/collaborator
key resolves to the lead because the lead entry is put last (overwriting), not
because it is put first.
2026-10-04 01:37:13 +02:00
Dai Ha 02c6909546 fleetd #669 Unit D: a live spawned member outranks every tab map
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Failing after 1m43s
CallerResolver.resolve() now checks the spawned-member roster first, ahead
of every tab map, so a live member's own role wins over a lead or
collaborator tab naming the same terminal (closes #661 at the resolver
level). Adds collaborator resolution (Role.COLLABORATOR) and a real
knownLeadOrCollaborator() classifier, wired into both FleetMcp.denyFor and
FleetApp.allow in place of the inert NO_KNOWN_LEAD_OR_COLLABORATOR stand-in.

LeadTabScanner is generalized to match lead and collaborator tabs in one
pass, keeping every existing liveness/caching/grace-scan property for both
kinds. FleetdAssembly wires the collaborator tab map and a roster-backed
spawned-member lookup; a collaborator-only fleet (no leaders configured)
falls back to a 10s scan interval, matching FleetConfig.Leader's own default.
2026-10-04 01:13:25 +02:00
Dai Ha b92a669ddc Merge PR #699: fleetd #669 Unit C — the COLLABORATOR role and its authorization row
CI / shell-tests (push) Failing after 11s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m49s
2026-10-04 00:34:13 +02:00
Dai Ha bb29b001e4 fleetd #669 Unit C: add the COLLABORATOR role and its authorization row
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m48s
Adds Role.COLLABORATOR, Principal.collaborator(), and the Authz.permits grant:
READ/METRICS open, REPLY/ASK via ownership, SEND limited to a configured lead
or collaborator via a new classifier parameter, everything else denied.
isSpawnedMember() stays WORKER || ARCHITECT. fleet_whoami reports role and
collaborator name with no leader key.

Per a same-day ticket correction, the existing 3-argument Authz.permits is
kept (fail-closed default via the new NO_KNOWN_LEAD_OR_COLLABORATOR
classifier) rather than deleted, so the 47 pre-existing call sites in
AuthzTest/CallerResolverTest are untouched; both production gates
(FleetMcp#denyFor, FleetApp#allow via the new permitsFor seam) call the
4-argument form explicitly with the shared constant.

Mutation-verified: removing the classifier conjunct from the SEND arm kills
exactly 3 tests (AuthzTest, FleetMcpAuthzTest, FleetAppAuthTest), one per
gate, no cascade.

Full mvn clean install at this commit: 1974 tests, 0 failures (baseline at
f0ff252 was 1964; +10 are the new collaborator-matrix tests). Independently
counted from target/surefire-reports/*.xml after rm -rf: 173 report files,
aggregate tests=1974 failures=0 errors=0.
2026-10-04 00:22:27 +02:00
Dai Ha f0ff25221e Merge PR #697: fleetd #669 Unit B — recognise-only fleet.collaborators config block
CI / shell-tests (push) Failing after 14s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 2m2s
2026-10-04 00:02:10 +02:00
Dai Ha 780cb342ad fleetd #669: pin the fleet-wide tabLabel-vs-collaborator-tab check
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Failing after 1m58s
Adds a test driving validateLeadTabPrefixes()'s fleet-wide
fleet.tabLabel branch directly against a collaborator tab, with no
profile override involved. Every existing collaborator test exercised
only the profile-override branch below it, leaving the fleet-wide
branch without a pinning test.
2026-10-03 23:55:47 +02:00
Dai Ha 5c563f02c8 fleetd #669: comments-only fixes — drop history/narration and premature behaviour claims
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m54s
The duplicate-collaborator test's javadoc narrated the before/after and a
timestamp; replaced with the present-tense behaviour it protects, per the
project's code-comment rule.

The Collaborator record javadoc and fleetd.example.yaml both claimed the
daemon "learns to address" a collaborator tab. Nothing does yet — that is
Units C-E. Dropped the claim from the record's contract and the example
now says the block is parsed and validated today, nothing more.

Also drops a confidence marker and a cross-class pointer from the
Collaborator record javadoc. No logic, assertion, or file outside this
PR's existing three changed.
2026-10-03 23:47:16 +02:00
Dai Ha ad593c9bb9 fleetd #669: add collaborators to the duplicate-slot-key guard
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Failing after 2m3s
FLEET_POOL_KEYS was missing "collaborators", so a duplicated
fleet.collaborators.<name> key took the skipValue branch in
rejectDuplicateSlotsInPools and silently collapsed last-wins, unlike
the other five pools. Measured: before this change, a test loading a
config with a duplicated collaborator name observed no exception;
after, it fails exactly like the existing architect-pool case.

Also updates the javadoc's "five pools" count to six.
2026-10-03 23:41:13 +02:00
Dai Ha 736fd9cf4b Merge PR #696: fleetd #692 — bound the unbounded userinfo mask in redact()
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m46s
2026-10-03 23:33:37 +02:00
Dai Ha 05244a82b3 fleetd #669 Unit B: recognise-only fleet.collaborators config block
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m46s
Adds fleet.collaborators.<name>.tab (a Collaborator record with only a tab
field — no profile, no instances, no kind; never auto-launched) and widens
the two startup refusals that stop it being a privilege hole:

- validatePanePlacementAgainstLeadTabs() now fires on either a lead tab or
  a collaborator tab, and no longer early-returns when fleet.leaders is
  empty but fleet.collaborators is not.
- validateLeadTabPrefixes() now also refuses a tabLabel template able to
  render as a collaborator tab, two collaborators sharing one exact tab,
  and a collaborator tab equal to a lead tab (case-insensitively).

validateMembers() refuses a collaborator with no (or blank) tab, next to
the existing leader-must-name-a-tab check.

Does not touch Authz, CallerResolver, LeadTabScanner, MemberRole or
FleetMcp — those are later units per the ticket's plan.
2026-10-03 23:33:28 +02:00
Dai Ha ef4996a01e fleetd #692: bound the unbounded userinfo mask at redact()'s line 273
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 2m9s
redact()'s final fallthrough still used the unbounded [^@]* class #638
removed from the verdict-line masker, so a diff line with a URL that has
no userinfo plus a later @ elsewhere had its text between them silently
deleted. Share one bounded implementation, mask_url_userinfo(), between
redact() and mask_verdict_userinfo() so the bound lives in one place.
2026-10-03 23:27:15 +02:00
Dai Ha edbd8d816a Merge PR #694: fleetd #689 — check SEND before reading the request body in sendMessage
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 44s
CI / build (push) Failing after 1m53s
2026-10-03 22:58:17 +02:00
Dai Ha 9425a9b696 fleetd #689: pin the answerGatePasses call site via the audit trail
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Failing after 2m3s
A unit test on the extracted helper proves the helper, not the call
site in sendMessage. allow() logs an AuditLog.allowed() entry for
every granted non-READ/METRICS/TASK_READ action, so a granted turnId
request must log both SEND and ANSWER, and a granted plain request
must log SEND alone. Verified this goes red when the call site is
deleted from sendMessage, and restores to a clean diff.
2026-10-03 22:56:04 +02:00
Dai Ha d0688c8a60 Merge PR #695: fleetd #693 — pin the lead-tab guard's case-insensitivity; fleetd #676 — drop the last stale validator count
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 55s
CI / build (push) Failing after 1m47s
2026-10-03 22:54:04 +02:00
Dai Ha 2c467c2553 fleetd #693, #676: pin lead-tab case-insensitivity, drop stale validator count
CI / shell-tests (pull_request) Failing after 20s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m40s
#693: add a FleetConfigTest case for two fleet.leaders tabs differing only
in case — mutating equalsIgnoreCase to equals left this uncaught before.

#676: rename theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames
to theSweepRunsEveryValidateMethodOnAnUnrelatedClass, since there are
eight validators now, not six, and the count had drifted into the name
and its class-javadoc {@link}.
2026-10-03 22:51:23 +02:00
Dai Ha c6430d8edd fleetd #689: check SEND before reading the request body in sendMessage
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m56s
Authorize twice: the coarse SEND grant first, with no body read, then
parse the body, then check ANSWER too when turnId is present. Restores
the pre-#687 ordering (no attacker-controlled body parse before the
gate) while keeping the SEND/ANSWER split #687 introduced.
2026-10-03 22:47:25 +02:00
Dai Ha bfee23acc3 Merge PR #691: fleetd #677 — refuse two leads sharing one exact tab (supersedes #686)
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 52s
CI / build (push) Failing after 1m37s
2026-10-03 22:45:23 +02:00
Dai Ha 7dec74f1b4 Merge PR #690: fleetd #638 — stop the verdict userinfo mask from crossing / or whitespace (supersedes #685)
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 54s
CI / build (push) Failing after 1m58s
2026-10-03 22:39:42 +02:00
Dai Ha 482598e2a6 fleetd #677: refuse two leads that share one exact tab
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Failing after 2m17s
validateLeadTabPrefixes() compared each lead's tab against the member
tabLabel template, but never against another lead's tab. Two leads
configured with the same exact tab loaded cleanly, even though exact
tab is the only thing lead identity is matched on, so only one of them
could ever be found.

Add an independent pass, case-insensitive, that refuses when two
fleet.leaders entries share one exact tab, with its own exception so
the message stays accurate for this relation.

Decided not to also refuse a lead's tab starting with a sibling's
tabPrefix: tabPrefix plays no role in identity resolution (only the
exact tab does), and the default tabPrefix is "lead:", the same
string used as the conventional lead tab prefix throughout this
codebase's own fixtures (e.g. "lead: opus" / "lead: sol"). Refusing
that case would reject the standard multi-lead setup with no matching
identity hazard.
2026-10-03 22:36:44 +02:00
Dai Ha 5f5d16fbd4 fleetd #638: stop the verdict userinfo mask from crossing / or whitespace
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m57s
mask_verdict_userinfo's character class [^@]* crossed a '/' or a space, so a
verdict line with a URI that has no userinfo plus a later @ (e.g. an email
address in diagnostic prose) had everything between them destroyed. Restrict
the class to [^@/[:space:]]* so the match stops at the end of the URI.

Adds the uncovered-direction test: a URI with no userinfo plus a later @ in
the same line must pass through byte for byte. Also drops the design-rationale
sentence from the helper's comment (now in the PR description).
2026-10-03 22:34:02 +02:00
Dai Ha 7df7985a16 Merge remote-tracking branch 'refs/remotes/pr/686' into worker/677-fix-lead-collision-f69073-12 2026-10-03 22:29:29 +02:00
Dai Ha 724b35b46e Merge remote-tracking branch 'refs/remotes/pr/685' into worker/638-fix-overmask-dbb1bf-11 2026-10-03 22:29:14 +02:00
Dai Ha 2eb2d6112e Merge PR #688: fleetd #675 — pin three unpinned FleetdAssembly constructor args
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 57s
CI / build (push) Failing after 1m35s
2026-10-03 22:27:32 +02:00
Dai Ha 7c458e8bf2 Merge PR #687: fleetd #669 Unit A — split SEND and READ into their real call shapes (closes #678) 2026-10-03 22:27:32 +02:00
Dai Ha 133f03e428 Merge PR #684: fleetd #683 — decouple the completion-fallback test's two 5s budgets 2026-10-03 22:27:27 +02:00
Dai Ha 804279175d fleetd #675: pin assembly loop timing defaults
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m42s
2026-10-03 22:12:46 +02:00
Dai Ha 9dea289975 fleetd #669 Unit A: split SEND and READ into their real call shapes
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 2m7s
SEND covered three different call shapes under one action (local
sessionId delivery, the coordId cross-host broker route, and the
turnId answer-a-blocked-worker form). READ covered both roster/
profile/identity observation and ticket-polling/session-status.
Split each into its own Authz.Action — SEND/COORD_SEND/ANSWER and
READ/TASK_READ — with every new action granted to exactly who held
the combined action before, on both the MCP and REST entry paths.

Also fixes fleetd #678's Authz.java comment: READ no longer claims
"the roster carries no secrets" for ticket replies and pending
questions, because those now live under TASK_READ.
2026-10-03 22:12:13 +02:00
Dai Ha cb4a6869b9 fleetd #677: guard exact lead tab collisions
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Failing after 1m56s
2026-10-03 22:10:29 +02:00
Dai Ha 28b45d97e5 fleetd #638: mask userinfo in the daemon verdict line before it prints
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Failing after 2m4s
Apply the userinfo-only rewrite (sed -E 's#://[^@]*@#://<redacted>@#g') to every
path that prints config-edit.sh's $VERDICT_LINE: restore_and_confirm's two
prints, report_outcome's shared local copy (covering its clean/needs-restart/
refused branches), and check_mode's "last verdict in log" line via
last_verdict_line. Deliberately not routed through redact() — that function's
key:value masking does not match this line's prose, and the rest of the line
(e.g. the pattern quoted in a parse-failure refusal) is the detail an operator
needs to fix the refusal.

No current refusal message echoes a URI, token or password, so this is a guard
against a future validator doing so, not a fix for an observed leak.
2026-10-03 22:06:54 +02:00
51 changed files with 6885 additions and 736 deletions
+35 -7
View File
@@ -159,6 +159,29 @@ fails.
`{action: "cancel", token}` drops a pending request without rolling.
**A change is coming: the roll will restart the process instead of sending `/clear` (fleetd #726
unit 2, written 2026-10-04).**
Today a roll types `/clear` into your pane. Your `claude` process keeps running, so a newer CLI on
disk is never loaded. Unit 2 replaces that: the daemon ends the old pane, launches a fresh one,
waits for the new terminal to be recognised as a lead, and only then sends the bootstrap text.
**Everything below about `/clear` is accurate while the old jar is running.** Unit 2 was not merged
when this note was written, and a merge is not a deployment.
**How to tell which one is live: read your own tool list.** If `fleet_handover`'s description says
it will "clear your pane", the daemon is serving the old behaviour. If it names a restart, the new
behaviour is live. The description comes from the running daemon, so it cannot disagree with the
code that is actually loaded.
Two things change for you once it is live. The `status` outcomes are different: three new failures
replace the `/clear` ones. And the "never observed as WORKING after 8 consecutive IDLE/DONE polls"
warning described below can no longer appear, because that wait is deleted — so if you still see
it, the old jar is running. `TURN_NEVER_SETTLED` does not change, and still means nothing was
touched.
**Delete this note and rewrite the `/clear` paragraphs once the new jar is live.**
**Things that will surprise you:**
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
@@ -200,13 +223,18 @@ fails.
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
handover file, with the configured `bootstrapText` arriving as its first message. No context was
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
at all. Re-measure both numbers with:
- **The bootstrap prompt works end to end. Measured 2026-09-22, re-measured 2026-10-04.** This used
to say the fix was unproven (fleetd #489) and told you to expect a failure. That is no longer
true. On 2026-09-22 the daemon log held four `lead-rollover: rolled` lines. On 2026-10-04 it holds
**20**, against a control of 86 `lead-rollover:` lines. Each roll cleared the old lead and started
a fresh session against the handover file, with the configured `bootstrapText` arriving as its
first message. No context was lost. The old `Unknown command: /clearFresh` failure from 2026-09-12
does not appear in the log at all.
19 of the 20 carry an `elapsedMs`: median 16507 ms, maximum 48261 ms, and two above 45000 ms. That
figure times the **whole** roll, and the wait for your own turn to end dominates it, so do not
read it as the cost of the clear. Expect a roll to take tens of seconds, and do not treat a slow
one as a failed one. Re-measure all of these with:
```bash
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
+69 -19
View File
@@ -27,23 +27,34 @@ through its `fleet_*` tools. No session addresses a peer, a broker, or the netwo
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `fleet_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
**Call `fleet_whoami`.** It returns `primary`, `worker`, `architect`, `collaborator`, or `observer`,
resolved by the daemon from your connection — unforgeable, and the same resolution its authorization
gate uses. A worker also carries its `sessionId`, `profile`, `worktree` and `branch`; an architect
carries the slot name it was bound to; a collaborator carries its registry name and its own
`sessionId`, and **no `leader` key** — a collaborator is a named peer, not a primary. An **observer**
carries only its own `sessionId`: a pane the daemon could not place as any of the above, authorized
to `READ`/`METRICS` and to `REPLY`/`ASK` on its own pane and nothing more — never `SEND`, never a
ticket. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are a spawned member in the
claude-bridge fleet"*) ⇒ **spawned member**; fleet tools prefixed `mcp__fleet__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies — and a member spawned before CB-632 still says `mcp__bridge__*`); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `fleet_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect,
or a worker from an **observer** — an observer is just as unspawned as a collaborator and carries
none of these signals either, so only `fleet_whoami` tells the two apart. **And none of them fires
for a collaborator at all**: every signal in the ladder detects a *spawned* member, while a
collaborator is a tab a person opened by hand, so it has no charter, no fixed mount name and a
normal environment. A collaborator — or an observer — that cannot call `fleet_whoami` therefore falls
to the line below and acts as a worker. That is the safe direction — it under-privileges, and the
refusals are loud — but it means a collaborator or an observer has no way to learn what it is except
by asking. **Still unsure ⇒ act as a worker**, the most restricted member role this ladder can name.
The two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate
— loud and self-correcting — while a member acting as the primary ends its turn with no `fleet_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
### Invariants — every role, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
@@ -51,9 +62,10 @@ and the sender silently receives nothing. Fail toward the recoverable error.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `fleet_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
outside your role is refused, not queued.
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead, architect, or
collaborator** — and a collaborator may send only to a lead or another collaborator, never to a
spawned member's terminal; reply/ask are only-as-itself — any peer may answer for its own pane,
and for no other. A call outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
@@ -93,7 +105,11 @@ below are the procedure — run them in order, every task, not only the big ones
5. **Collect** — `fleet_poll{ticket}` → `fleet_ack{target, msgId}`. Answer a worker's `fleet_ask`
with `fleet_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`fleet_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
that answers it — but **only to the caller that created that delegation**, so a question raised
under an architect's brief is invisible to you, and seeing none does not mean there is none.
**Only that same creator can answer it.** A `turnId` you came by any other way is refused, so an
architect's worker waits for that architect and not for you.
**A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
**A correction cannot reach a busy member.** A `fleet_send` to a working member is *accepted* and
returns a ticket, and is then never delivered — measured here three times in one session, and the
@@ -158,9 +174,10 @@ you decide.
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId`, and only the caller that created that delegation |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Message a **collaborator** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` reports a `collaborators` array, and each row carries that peer's `name` and the `sessionId` you send to. It is visible to you, to an architect and to another collaborator, never to a worker. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
@@ -234,6 +251,27 @@ simply complies has thrown away the reason there are two of you.
7. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
### Collaborator — a named peer, not a member
`fleet_whoami` answered `collaborator`, so your pane's tab matches a `fleet.collaborators.<name>.tab`
entry. You are **not** a member: nothing delegates to you, you have no brief, no worktree and no
ticket, and **you owe no `fleet_reply`** — the turn contract above is for a session a lead spawned,
and it does not apply to you. Read it only to understand what the members around you are doing.
What you may do: observe the fleet (`fleet_list`, `fleet_profiles`, `fleet_whoami`), and send to a
lead or to another collaborator. What you may not: spawn, stop or drain anything, roll a lead's
session, answer a member's `fleet_ask`, poll a ticket, or send to a spawned member's terminal. Each
of those is refused at the gate, not queued.
Two limits worth knowing before you hit them. **You cannot reach a worker** — not even to help one —
because a worker belongs to the lead that spawned it, and routing around that would make you a
second orchestrator with no plan. Send to the lead instead. And **you cannot read a ticket**, so you
cannot collect a delegation's reply: `fleet_poll` refuses you at the role gate, and a ticket also
records the terminal that created it, so even a leaked id reads nothing.
Being named buys you a channel, not authority. Your `fleet_send` to a lead is coordination between
peers: the lead owes you no obedience, and you owe it none.
### Where each rule lives (don't duplicate — extend the right layer)
| Layer | Scope | Reaches |
@@ -301,10 +339,22 @@ must obey belongs in the charter, not here.
(#362). **Read `plugin/` before designing anything about onboarding a project.** Two limits are
structural, not bugs: a plugin cannot carry the role agent files, because
`ClaudeCodeLauncher.java:371` requires `<cwd>/.claude/agents/<role>.md` in the member's own
worktree; and a plugin cannot deliver anything to members at all, because
`ClaudeCodeLauncher.java:285` exports `CLAUDE_CONFIG_DIR` and every Claude profile here sets it,
so a member never reads the operator's plugin store. **The plugin is the lead-side surface;
member-facing assets travel in the worktree.**
worktree; and a plugin reaches a member only through `CLAUDE_CONFIG_DIR`, which
`ClaudeCodeLauncher.java:286` exports with `putIfPresent` — so only for a profile that sets
`configDir`. Every `claude-code` profile does set one (the four without are `opencode`, which
never reads that variable). **But measured 2026-10-04: two of them point at
`~/.ccs/instances/ltms`, which is the operator's own `CLAUDE_CONFIG_DIR` on this host.** So for an
`opus` or `sonnet` member, "a member never reads the operator's plugin store" is false — it reads
the same store, because that store is the one its `configDir` names. It stays true for `local` and
`local-direct`, which point at `~/.ccs/instances/gx10`. `ClaudeCodeLauncher`'s own javadoc names
the related hazard: that file is rewritten on every spawn, so for those two profiles fleetd and the
operator's live session write the same `.claude.json`, and its compare-and-swap "narrows the
lost-update window, it does not close it". Re-measure which profiles share the operator's dir with
`awk '/^profiles:/{i=1;next} /^[a-z]/{i=0} i&&/^ [a-z-]+:$/{p=$1} i&&/configDir:/{print p,$2}'
fleetd/fleetd.yaml` against `echo $CLAUDE_CONFIG_DIR`; delete this note once no profile names the
operator's dir. **Treat the plugin as the lead-side surface and put member-facing assets in the
worktree** — that conclusion holds either way, because a worktree asset does not depend on which
config dir a member reads.
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **A provisioned worktree neutralizes `.mcp.json`, `opencode.json` and `.autoenv`** — the repo's
@@ -359,7 +409,7 @@ Before you call any work done, check the row that matches what you touched:
| `ConnectionIdentity` / how a caller is resolved | the `fleet_whoami` paragraph and the fallback ladder |
| `REPLY_CHARTER`, or a launcher's mount/flags | the fallback ladder (`mcp__fleet__*`), and the layering table's top row |
| the injector / status gating | invariant 4 |
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| worktree provisioning or the parity overlay | the "every role reads this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `fleetd.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
+26 -8
View File
@@ -62,11 +62,12 @@ bind:
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing fleetd places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# Two things stop the tab-name convention from becoming a way to claim leadership: startup REFUSES
# a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel` override, also
# matches, so the two namespaces cannot overlap by accident; and the CallerResolver asks the live
# spawned-member roster BEFORE any tab map, so a live member is never mistaken for a lead no matter
# what its tab says. The label is a NAME, never a capability: what a pane may do is decided by the
# role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open fleet_send driving it) past the quiet period. The fleet is one lead + architects +
@@ -659,9 +660,10 @@ fleet:
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
# # this convention at startup; plays no part in matching a lead
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# workspace: leads # where a launched lead's tab is created (default "fleet",
# # the same shared space the members use). Sharing that space
# # with members is the normal shipped shape: the scanner tells
# # a lead from a member by the exact tab label, not by workspace.
# cwd: /path/to/repo # the launched lead's working directory (default: fleetd's own)
# kind: claude # descriptive; reported by fleet_whoami
# gpt-sol-5.6:
@@ -669,6 +671,22 @@ fleet:
# kind: opencode
# model: openai/gpt-5.6-terra
# A collaborator tab, keyed by name (fleetd #669). A pane whose tab matches resolves as the
# COLLABORATOR role: `fleet_whoami` answers `collaborator`, and the session may observe the fleet
# (`fleet_list`, `fleet_profiles`, `fleet_whoami`) and `fleet_send` to a lead or to another
# collaborator. It may NOT spawn, stop or drain anything, roll a lead's session, answer a
# member's `fleet_ask`, poll a ticket, send across hosts, or send to a spawned member's terminal.
# Each of those is refused at the gate rather than queued.
# Recognise-only, like a profile-less `leaders:` entry above: there is no `profile:`, no
# `instances:` and no `kind:`. `tab:` is REQUIRED and is the only field identity depends on,
# matched case-insensitively — the same GET-THE-VALUE-RIGHT warning above the `leaders:` block
# applies here too.
# A lead cannot discover a collaborator yet (#703): `fleet_list` has no `collaborators` key, so
# the collaborator must speak first, or pass on the `sessionId` its own `fleet_whoami` reports.
# collaborators:
# reviewer-alex:
# tab: "collab: alex"
# architects:
# architect-1:
# profile: opus # a strong model, on the operator's subscription
+14 -12
View File
@@ -217,25 +217,27 @@ public final class Fleetd {
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
* member whose agent has connected the bridge MCP, a lead, <em>or</em> a collaborator.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
* (or labelled its tab) only once it was up, so there is no boot window to guard. A collaborator
* is the same way — a person's own tab, matched to a configured name, never spawned.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code FleetMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
* <p>Neither a lead nor a collaborator is ever enrolled in {@link MemberPresence} — {@code
* FleetMcp} marks presence only for a worker, an architect, or the unconfigured-pane floor,
* never for a lead or a collaborator. So without the second and third disjuncts a lead or
* collaborator is permanently un-deliverable: every send to one sat on the gate for
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
* <p>Both sets are read through their supplier on each call rather than snapshotted, so a lead or
* collaborator discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads,
Supplier<Map<String, String>> collaborators) {
return target -> presence.isPresent(target) || leads.get().containsKey(target)
|| collaborators.get().containsKey(target);
}
/**
@@ -43,6 +43,7 @@ import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.ReplyPushLoop;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
@@ -141,8 +142,12 @@ final class FleetdAssembly {
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
: herdr;
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
// fleetd #669 Unit E: a collaborator's pane is opened by a person, exactly like a lead's,
// so its terminal must also route to the lead herdr daemon rather than the member one.
AtomicReference<Supplier<Map<String, String>>> collaboratorTerminalsRef = new AtomicReference<>(Map::of);
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
target -> leadsRef.get().get().containsKey(target));
target -> leadsRef.get().get().containsKey(target)
|| collaboratorTerminalsRef.get().get().containsKey(target));
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
@@ -250,26 +255,47 @@ final class FleetdAssembly {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
// configured lead's own exact `tab:` label.
// configured lead's own exact `tab:` label. fleetd #669: the same scan also recognises a
// configured collaborator's tab, so one herdr pass answers both.
final Supplier<Map<String, String>> leads;
final Supplier<Map<String, String>> collaboratorTerminals;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
var collaboratorsConfig = cfg.fleet().collaborators();
if (!leaders.isEmpty() || !collaboratorsConfig.isEmpty()) {
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
Map<String, String> collaboratorTabToName = new LinkedHashMap<>();
collaboratorsConfig.forEach((name, collaborator) -> {
if (collaborator != null && collaborator.tab() != null && !collaborator.tab().isBlank()) {
collaboratorTabToName.put(collaborator.tab(), name);
}
});
// A collaborator-only fleet configures no `leaders:` entry to read a scan interval from
// — FleetConfig.Collaborator carries no scanIntervalSeconds of its own. Falling back to
// FleetConfig.Leader's own compact-constructor default keeps a collaborator-only
// deployment on the same rescan cadence as the default lead cadence, instead of
// inventing a second number for the same kind of scan.
int scanIntervalSeconds = leaders.isEmpty()
? 10
: leaders.values().iterator().next().scanIntervalSeconds();
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
LeadTabScanner scanner = new LeadTabScanner(herdr, tabToName, collaboratorTabToName, Set.of(),
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
tabToName.keySet(), scanIntervalSeconds);
leads = scanner;
collaboratorTerminals = scanner::collaborators;
log.info("lead/collaborator scan: tabs {} host a lead, tabs {} host a collaborator "
+ "(rescan every {}s, shared fleet space)",
tabToName.keySet(), collaboratorTabToName.keySet(), scanIntervalSeconds);
} else {
leads = () -> leadTerminals;
collaboratorTerminals = Map::of;
}
leadsRef.set(leads);
collaboratorTerminalsRef.set(collaboratorTerminals);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// and only when herdr answered — the launcher's whole safety property is that it can count
@@ -341,7 +367,7 @@ final class FleetdAssembly {
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads, collaboratorTerminals);
// fleetd #556: registration is wired directly to `completion`, not folded into the
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
// listener) throwing, regardless of call order.
@@ -452,6 +478,11 @@ final class FleetdAssembly {
ConnectionIdentity identity = new ConnectionIdentity(
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
// fleetd #669 Unit D: a live spawned member resolves as its own role, whatever a tab map
// says about the same terminal. fleetd #702: SessionManager.spawnedMemberRole also answers
// for a pane mid-teardown, not only one still in the registry — see its javadoc.
Function<String, MemberRole> spawnedMemberRole = sessions::spawnedMemberRole;
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
final CallerResolver callers;
@@ -461,11 +492,13 @@ final class FleetdAssembly {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting fleetd");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members,
spawnedMemberRole, collaboratorTerminals);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members,
spawnedMemberRole, collaboratorTerminals);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
@@ -1,5 +1,7 @@
package dev.ltms.fleet.auth;
import java.util.function.Predicate;
/**
* The authorization table (CB-505), stated once and enforced on both entry paths.
*
@@ -20,16 +22,22 @@ public final class Authz {
SPAWN,
/** Tear a worker peer down. */
STOP,
/** Deliver a turn to a session (or answer a worker's question). */
/** Deliver a turn to a local session, addressed by {@code sessionId}. */
SEND,
/** Resolve a worker's blocked question and resume its turn, addressed by {@code turnId}. */
ANSWER,
/** Address a peer lead on another daemon over the coordination broker, by {@code coordId}. */
COORD_SEND,
/** A worker's terminal reply for its own turn. */
REPLY,
/** A worker's mid-turn question to the primary. */
ASK,
/** Collect held replies from a session's inbox. */
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
/** Read-only roster, profile, and identity observation: no ticket, task, or turn state. */
READ,
/** Poll a ticket, or read a session's status. */
TASK_READ,
/**
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
*
@@ -54,43 +62,97 @@ public final class Authz {
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
* The fail-closed classifier: answers no for every target, so a collaborator's {@code SEND}
* is refused unless a caller supplies a real one. {@code CallerResolver#knownLeadOrCollaborator()}
* is the real one, read from the same lead and collaborator maps {@code CallerResolver#resolve}
* consults, so a target that classifier calls known is one {@code resolve} would actually
* resolve as a lead or collaborator.
*/
public static final Predicate<String> NO_KNOWN_LEAD_OR_COLLABORATOR = target -> false;
/**
* Convenience form for a caller with no classifier to supply. Fails closed: a collaborator's
* {@code SEND} is refused, as if no terminal were a configured lead or collaborator — the
* same decision {@link #NO_KNOWN_LEAD_OR_COLLABORATOR} gives explicitly. Every other action's
* result is identical to the four-argument form's, since none of them consult the classifier.
*
* @param targetSession the session id in the request path; only consulted for the worker-scoped
* actions ({@code REPLY}, {@code ASK}), ignored otherwise, may be
* {@code null}
* <p>Its default classifier denies every collaborator, so a caller enforcing authorization
* must use the four-argument form instead.
*/
public static boolean permits(Principal caller, Action action, String targetSession) {
return permits(caller, action, targetSession, NO_KNOWN_LEAD_OR_COLLABORATOR);
}
/**
* Whether {@code caller} may perform {@code action} against {@code targetSession}.
*
* @param targetSession the session id in the request path; only consulted for the
* worker-scoped actions ({@code REPLY}, {@code ASK}) and for a
* collaborator's {@code SEND}, ignored otherwise, may be
* {@code null}
* @param knownLeadOrCollaborator whether a terminal is a configured lead or collaborator —
* consulted only for a collaborator's {@code SEND}, to confine
* it to another named peer and never a spawned member's
* terminal
*/
public static boolean permits(Principal caller, Action action, String targetSession,
Predicate<String> knownLeadOrCollaborator) {
if (caller == null || caller.isAnonymous()) {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect and a
// collaborator deliberately do NOT get these, so neither can tear down or stand up
// workers even though one of them coordinates them; and a worker driving any of these
// would be a worker escalating into the orchestrator role.
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
// excluded — sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect();
// Delivering a turn to a local session is open to the primary, the architect, and a
// collaborator whose target is itself a configured lead or collaborator: the architect
// delegates to workers (that is the role's point); a collaborator may reach only
// another named peer, never a spawned member's terminal. A worker is excluded —
// sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect()
|| (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession));
// Resolving a worker's blocked question is part of delegating to it, open to the same
// two roles that may stand up that delegation in the first place. Not a collaborator:
// resuming another session's turn is lifecycle-adjacent, not peer messaging.
case ANSWER -> caller.isPrimary() || caller.isArchitect();
// Leaves the daemon over the coordination broker rather than addressing a local
// session, open to the same two roles as ANSWER. Not a collaborator: it is a
// local-tab peer with no cross-host route.
case COORD_SEND -> caller.isPrimary() || caller.isArchitect();
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
// that can be — a lead answering another lead is replying for its OWN terminal, which
// this already permits — while the rule itself is unchanged, and is what stops anyone
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
// passes through the same check, so it can answer a funnel that delegated to it. An
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
// forging a reply for a rendezvous someone else is waiting on. An architect's or a
// collaborator's own pane passes through the same check, so each can answer a funnel
// that delegated to it. An unnamed primary (token/loopback, no pane) owns nothing and
// is still excluded.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// READ is roster, profile, and identity observation — fleet_list, fleet_profiles, and
// fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no
// other session's turn state. Those live under TASK_READ. METRICS is the separate
// Prometheus scrape. Both are open to every authenticated role, including a
// collaborator and the unconfigured-pane floor.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect()
|| caller.isCollaborator() || caller.isObserver();
// Ticket polling and session status, open to every role READ is open to except a
// collaborator or an observer. MessageService compares a ticket's creator to the
// caller on every read as well, so dropping this gate would not expose another
// session's reply — it would move the refusal later and widen what a caller that
// never orchestrates can probe.
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
// is coordination between leads, not observation of the roster.
// is coordination between leads, not observation of the roster. The same reasoning
// excludes a collaborator.
case COORD_READ -> caller.isPrimary();
};
}
@@ -7,6 +7,7 @@ import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
/**
@@ -19,6 +20,13 @@ import java.util.function.Supplier;
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a pane this gateway itself spawned ⇒ that member's own
* role: {@link Role#WORKER} for a dev, hunter, or reviewer; {@link Role#ARCHITECT} for an
* architect, but only while the live slot role still confirms it (fleetd #424 — a slot
* revoked from config demotes an already-bound session on its very next request, so the
* roster's own role is never granted on its word alone). No tab map is consulted — a live
* spawned member's identity comes from the registry that spawned it, never from a label a
* pane could also carry.</li>
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
* ⇒ {@link Role#PRIMARY}, carrying that lead's
@@ -28,9 +36,11 @@ import java.util.function.Supplier;
* so two leads can work as peers rather than one being demoted.</li>
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
* the generic worker fallback.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* resolved from the <em>live</em> terminal→slot binding (never a request argument). This is
* the case the previous step does not catch: a binding with no live spawned-member session.</li>
* <li>A loopback peer PID that maps to an operator-labelled collaborator tab ⇒
* {@link Role#COLLABORATOR}, carrying that collaborator's name.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#OBSERVER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
@@ -64,6 +74,20 @@ public final class CallerResolver {
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
private final Function<String, String> memberSlotNames;
/**
* terminal_id → the role of the live spawned member occupying it, or {@code null} for a
* terminal no spawned member occupies. Consulted first, ahead of every tab map: a live
* spawned member's identity is its own, whatever a tab map says about the same terminal.
* A function rather than the roster itself, so a resolve on the hot path never scans a list —
* the lookup strategy is the caller's to choose.
*/
private final Function<String, MemberRole> spawnedMemberRole;
/**
* terminal_id → collaborator name; empty when none are configured. A supplier for the same
* reason as {@link #leadTerminals}: a collaborator tab recognised after construction (the tab
* scan discovering a newly-labelled tab) takes effect without a restart.
*/
private final Supplier<Map<String, String>> collaboratorTerminals;
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
CallerResolver(ConnectionIdentity identity) {
@@ -124,17 +148,42 @@ public final class CallerResolver {
/**
* Live registry form that can confirm a bound slot is an architect slot.
*
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
* <p>It keeps terminal bindings and slot roles in the same {@link MemberRegistry}, so a
* configured architect can resolve as an architect. No spawned-member roster or collaborator
* registry is consulted — equivalent to {@link #withLeadsAndMembers(ConnectionIdentity,
* boolean, String, Supplier, MemberRegistry, Function, Supplier)} with both absent. Kept for
* every caller that has neither to offer, so adding them did not churn every construction site.
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members) {
return withLeadsAndMembers(identity, tokenMode, token, leadTerminals, members, null, null);
}
/**
* Live registry form that also resolves a live spawned member to its own role, and a
* configured collaborator tab to {@link Role#COLLABORATOR}.
*
* <p>This is the only public construction path that exercises the full resolution order.
*
* @param spawnedMemberRole terminal_id → the role of the live spawned member occupying
* it, or {@code null} for a terminal no spawned member occupies.
* {@code null} here means no roster is consulted at all (every
* terminal falls through to the tab maps), not that none matches.
* @param collaboratorTerminals terminal_id → collaborator name, live like {@code leadTerminals}
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members,
Function<String, MemberRole> spawnedMemberRole,
Supplier<Map<String, String>> collaboratorTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals,
members == null ? null : members::snapshot,
members == null ? null : members::roleForSlot,
members == null ? null : members::nameForSlot);
members == null ? null : members::nameForSlot,
spawnedMemberRole, collaboratorTerminals);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
@@ -160,6 +209,17 @@ public final class CallerResolver {
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles,
memberSlotNames, null, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames,
Function<String, MemberRole> spawnedMemberRole,
Supplier<Map<String, String>> collaboratorTerminals) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
@@ -172,6 +232,8 @@ public final class CallerResolver {
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
this.spawnedMemberRole = spawnedMemberRole == null ? _ -> null : spawnedMemberRole;
this.collaboratorTerminals = collaboratorTerminals == null ? Map::of : collaboratorTerminals;
}
/**
@@ -188,14 +250,24 @@ public final class CallerResolver {
}
/**
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
* The currently-recognised collaborator tabs, {@code terminal_id → name}.
*
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
* <p>Read from the same supplier {@link #resolve} consults, for the reason given in
* {@link #leads()}. Live for the same reason as {@link #leads()}.
*/
public Map<String, String> members() {
return architectTerminals.get();
public Map<String, String> collaborators() {
return collaboratorTerminals.get();
}
/**
* Whether {@code target} names a terminal this resolver would resolve as a lead or a
* collaborator — the classifier a collaborator's {@code SEND} is checked against, read from the
* exact maps {@link #resolve} consults so a target that would resolve as a lead or collaborator
* is never the one a collaborator is refused to reach, or the reverse.
*/
public Predicate<String> knownLeadOrCollaborator() {
return target -> leadTerminals.get().containsKey(target)
|| collaboratorTerminals.get().containsKey(target);
}
/**
@@ -208,6 +280,25 @@ public final class CallerResolver {
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
MemberRole spawnedRole = spawnedMemberRole.apply(c.terminal());
if (spawnedRole != null) {
// A live spawned member occupies this pane. Its identity is its own, whatever a tab
// map says about the same terminal — checked before every tab map, consulting none
// of them, so a tab label can never override a roster entry for the same terminal.
if (spawnedRole == MemberRole.ARCHITECT) {
// The roster only answers THAT this pane is a live spawned member; config still
// decides WHAT that member's slot grants (fleetd #424). A slot revoked after the
// bind must still demote this session on its very next request, so the roster's
// own ARCHITECT role is confirmed against the live slot role, exactly as the
// architect-slot step below confirms a binding with no live member session.
String slot = architectTerminals.get().get(c.terminal());
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid());
}
String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
// The config names this pane as a lead's own. The pane mapping is exactly as
@@ -221,11 +312,18 @@ public final class CallerResolver {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev, hunter or reviewer into an architect. Checked before
// the worker fallback.
// escalating a dev, hunter or reviewer into an architect. This is the case the
// spawned-member step above does not catch: a binding with no live member session.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
String collaborator = collaboratorTerminals.get().get(c.terminal());
if (collaborator != null) {
// An operator-labelled collaborator tab, confirmed live by the same scan that
// confirms a lead tab. Checked last among the tab maps so a pane also matching one
// of the above keeps that stronger role.
return Principal.collaborator(collaborator, c.terminal(), c.pid());
}
return Principal.observer(c.terminal(), c.pid()); // unforgeable; never token-gated
}
if (tokenMode) {
@@ -11,7 +11,9 @@ package dev.ltms.fleet.auth;
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
* for an architect resolved from the CB-548 {@code architects:} registry, which
* slot it occupies; {@code null} for every other caller, including an unnamed primary
* slot it occupies; for a collaborator resolved from the {@code collaborators:}
* registry, which collaborator it is; {@code null} for every other caller,
* including an unnamed primary
*/
public record Principal(Role role, String terminal, long pid, String name) {
@@ -73,6 +75,26 @@ public record Principal(Role role, String terminal, long pid, String name) {
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
}
/**
* A collaborator: a human-opened tab recognised by its exact label in the
* {@code collaborators:} registry.
*
* <p>Carries {@link Role#COLLABORATOR}. {@code name} is reporting only — it lets
* {@code fleet_whoami} say which collaborator is asking. Identity is the {@code terminal}:
* like a worker's it comes from the connection, so {@code ownsSession} works exactly as it
* does for a worker — a collaborator acts as its own pane and no other.
*/
public static Principal collaborator(String name, String terminal, long pid) {
return new Principal(Role.COLLABORATOR, terminal, pid, name);
}
/**
* The unconfigured-pane floor: a loopback caller whose pane matched no other role.
*/
public static Principal observer(String terminal, long pid) {
return new Principal(Role.OBSERVER, terminal, pid);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
@@ -81,15 +103,23 @@ public record Principal(Role role, String terminal, long pid, String name) {
return role == Role.ARCHITECT;
}
public boolean isCollaborator() {
return role == Role.COLLABORATOR;
}
public boolean isWorker() {
return role == Role.WORKER;
}
public boolean isObserver() {
return role == Role.OBSERVER;
}
/**
* Whether this caller is a spawned member with its own pane.
*
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
* as present would count it as an available member in the roster.
* <p>Both workers and architects are spawned members. A lead is not: it is a peer the
* operator started and named, never a pane this daemon spawned.
*/
public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT;
@@ -115,11 +145,32 @@ public record Principal(Role role, String terminal, long pid, String name) {
return terminal != null && terminal.equals(sessionId);
}
/**
* Stable identity used to own tickets and open turns. The unnamed primary has no owner key so
* it can use the message layer's primary-wide ticket access rule.
*/
public String ownerKey() {
return switch (role) {
case PRIMARY -> name == null ? null : prefixed("leader", name);
case WORKER -> prefixed("worker", terminal);
case ARCHITECT -> prefixed("architect", terminal);
case COLLABORATOR -> prefixed("collaborator", name);
case OBSERVER -> prefixed("observer", terminal);
case ANONYMOUS -> "anonymous";
};
}
private static String prefixed(String role, String identity) {
return role + ":" + identity;
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case ARCHITECT -> "architect:" + name;
case COLLABORATOR -> "collaborator:" + name;
case OBSERVER -> "observer:" + terminal;
case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous";
};
@@ -34,6 +34,28 @@ public enum Role {
*/
ARCHITECT,
/**
* A config-declared, human-opened tab recognised by its exact label (the {@code
* fleet.collaborators.<name>.tab} registry). Never spawned — identity comes from the
* connection, never a request argument, exactly like {@link #WORKER} and {@link #ARCHITECT}.
* May {@code SEND} only to a configured lead or collaborator, {@code REPLY}/{@code ASK} only
* as its own pane, and {@code READ}/{@code METRICS}; may not {@code SPAWN}/{@code STOP}/
* {@code DRAIN}/{@code HANDOVER}, poll a ticket ({@code TASK_READ}), or reach the
* coordination broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
COLLABORATOR,
/**
* A loopback pane that resolved to none of the roles above: not a live spawned member, not a
* configured lead, not a bound architect slot, not a configured collaborator tab. Unforgeable
* like a worker's — derived from the connection's pane, never from a request argument, and
* honoured regardless of auth mode. May {@code READ} and {@code METRICS}, and {@code REPLY}/
* {@code ASK} only as its own pane; may not {@code SPAWN}/{@code STOP}/{@code DRAIN}/
* {@code HANDOVER}, {@code SEND}, poll a ticket ({@code TASK_READ}), or reach the coordination
* broker ({@code COORD_SEND}/{@code COORD_READ}).
*/
OBSERVER,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
@@ -615,7 +615,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
// may bind to, AND what a slot already bound still grants) through its own instance of that
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds.
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
@@ -635,6 +635,11 @@ public final class ConfigRef implements Supplier<FleetConfig> {
+ "live through that same supplier for placement AND through a separate supplier "
+ "on MemberRegistry for spawn-time identity — both already applied");
}
if (!Objects.equals(collaboratorsOf(old), collaboratorsOf(fresh))) {
changed.add("fleet: fleet.collaborators (each collaborator's tab) is read once at "
+ "startup to build the LeadTabScanner's identity map, which is not rebuilt on "
+ "reload, so a collaborator added, removed, or given a new tab: needs a restart");
}
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
// every message here must be traceable to one of the split keys the class doc documents.
// NOTE what this does NOT prove, per the javadoc above: it does not catch a SPLIT_KEYS
@@ -650,6 +655,11 @@ public final class ConfigRef implements Supplier<FleetConfig> {
return cfg.fleet() == null ? Map.of() : cfg.fleet().leaders();
}
/** {@code cfg.fleet().collaborators()}, defensively, in case a caller hands in a non-defaulted config. */
private static Map<String, FleetConfig.Collaborator> collaboratorsOf(FleetConfig cfg) {
return cfg.fleet() == null ? Map.of() : cfg.fleet().collaborators();
}
/**
* {@link FleetConfig.Profile} record components deliberately left out of
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
@@ -1129,11 +1129,8 @@ public record FleetConfig(
* {@code tab} can never be discovered, launched or not
* @param instances how many of this lead should be live (default 1). The daemon
* launches only the shortfall, so a restart adopts rather than doubles
* @param tabPrefix no longer used to find a lead's tab — {@code tab} is matched
* exactly. Its only remaining job is the startup collision guard
* ({@link #validateLeadTabPrefixes()}), which still uses it to refuse
* a worker {@code tabLabel} template that could be misread as a lead.
* Default {@code "lead:"}
* @param tabPrefix lead-tab naming convention checked against member labels. Lead
* identity uses {@code tab}. Default {@code "lead:"}
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
* worst case before a newly-labelled tab is recognised. Default 10
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
@@ -1179,6 +1176,25 @@ public record FleetConfig(
}
}
/**
* A tab fleetd recognises as a collaborator, keyed by name (fleetd #669).
*
* <p>Recognise-only: there is no {@code profile}, no {@code instances} and no {@code kind}.
* Nothing here ever launches a pane.
*
* <p>{@code tabPrefix} is absent. Identity is matched on the exact {@code tab} alone.
*
* @param tab the exact tab label hosting this collaborator, matched case-insensitively; the
* only field identity depends on. Required — an entry with no {@code tab} can
* never be discovered.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Collaborator(String tab) {
public Collaborator {
tab = (tab == null || tab.isBlank()) ? null : tab.strip();
}
}
/**
* One entry of a {@code fleet:} role pool — a role paired with the backend it runs on.
*
@@ -1216,15 +1232,18 @@ public record FleetConfig(
* is exactly compatible with that. The pool is also what replaced {@code defaultProfile:} — an
* unqualified spawn names a role, and the role's pool supplies the candidates.
*
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
* @param architects profiles the {@code architect} role may run on
* @param developers profiles the {@code dev} role may run on
* @param hunters profiles the {@code hunter} role may run on
* @param reviewers profiles the {@code reviewer} role may run on
* @param charters optional launch-charter text keyed by singular role wire name
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
* {@code {model}} and {@code {n}} (a per role+profile counter) are
* substituted. Default {@link #DEFAULT_TAB_LABEL}
* @param leaders panes that orchestrate rather than are orchestrated, keyed by lead name
* @param architects profiles the {@code architect} role may run on
* @param developers profiles the {@code dev} role may run on
* @param hunters profiles the {@code hunter} role may run on
* @param reviewers profiles the {@code reviewer} role may run on
* @param charters optional launch-charter text keyed by singular role wire name
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
* {@code {model}} and {@code {n}} (a per role+profile counter) are
* substituted. Default {@link #DEFAULT_TAB_LABEL}
* @param collaborators tabs fleetd recognises as collaborators (fleetd #669), keyed by name.
* Recognise-only, exactly like a {@code profile}-less {@link Leader}:
* nothing here is ever auto-launched.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Fleet(Map<String, Leader> leaders,
@@ -1233,13 +1252,11 @@ public record FleetConfig(
Map<String, Slot> hunters,
Map<String, Slot> reviewers,
Map<String, String> charters,
String tabLabel) {
String tabLabel,
Map<String, Collaborator> collaborators) {
/**
* Role first, so the tab bar reads as the fleet and so the label shares a namespace with a
* lead's {@code tabPrefix}. Because {@code {role}} comes from a closed enum, a generated
* member label can never begin with {@code "lead:"} — the clash that
* {@link #validateLeadTabPrefixes()} used to have to check for is unrepresentable here.
* Role first, so the tab bar identifies the member's fleet role.
*/
public static final String DEFAULT_TAB_LABEL = "{role}: {profile} #{n}";
@@ -1251,26 +1268,30 @@ public record FleetConfig(
reviewers = unmodifiableOrEmpty(reviewers);
charters = unmodifiableOrEmpty(charters);
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
collaborators = unmodifiableOrEmpty(collaborators);
}
/**
* A fleet with no configured launch charters — the shape every deployment had before
* CB-566, and what most tests want.
* A fleet with no configured launch charters and no collaborators — the shape every
* deployment had before CB-566, and what most tests want.
*
* <p>Kept deliberately, even though an overload that drops a new field is normally the
* shape to avoid. It is safe here because nothing <em>reads</em> a charter through a
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
* {@code collaborators} is dropped the same way and for the same reason: no caller of
* this overload has ever needed to set it, so it defaults to empty here exactly as the
* canonical constructor would default an absent YAML key.
*/
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers,
Map<String, String> charters, String tabLabel) {
this(leaders, architects, developers, null, reviewers, charters, tabLabel);
this(leaders, architects, developers, null, reviewers, charters, tabLabel, null);
}
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
this(leaders, architects, developers, null, reviewers, null, tabLabel);
this(leaders, architects, developers, null, reviewers, null, tabLabel, null);
}
/**
@@ -1916,7 +1937,7 @@ public record FleetConfig(
/** The {@code fleet:} child blocks whose direct children are slot names. */
private static final Set<String> FLEET_POOL_KEYS =
Set.of("leaders", "architects", "developers", "hunters", "reviewers");
Set.of("leaders", "architects", "developers", "hunters", "reviewers", "collaborators");
/**
* Reject a {@code fleet:} role pool whose slot names repeat (CB-548, re-homed by CB-557).
@@ -1926,7 +1947,7 @@ public record FleetConfig(
* daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by
* default, so duplicates are caught here, at parse time, before the map is built.
*
* <p>Only the five pools <em>directly under the top-level {@code fleet:}</em> are considered,
* <p>Only the six pools <em>directly under the top-level {@code fleet:}</em> are considered,
* and only their direct child keys (the slot names). A nested field elsewhere, even one also
* named {@code developers:}, is ignored, so parsing of the rest of the config is unaffected.
*
@@ -2685,31 +2706,21 @@ public record FleetConfig(
}
/**
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
* Reject a member tab-label template that could render as a configured lead or collaborator
* tab or match a lead-tab naming convention, and reject two {@code fleet.leaders} or
* {@code fleet.collaborators} entries — across either registry — that share one exact tab.
*
* <p>The scan reads a tab label and concludes "a lead lives here". fleetd also <em>writes</em>
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
* a member template matches and the daemon starts labelling its own members as leads, promoting
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
* {@link #validatePanePlacementAgainstLeadTabs()} is the check that stops a pane-placed member
* from landing inside a lead's tab in the first place; this check is a second, independent
* guard that catches the hazard even when every profile places members correctly, by refusing
* a label that a scan would still misread as a lead.
* <p>{@code fleet.collaborators} has no {@code tabPrefix}: identity is matched on the exact
* {@code tab} alone, so only the exact-render check applies there, not the prefix check.
*
* <p>CB-557 shrank this check rather than removing it. The default template is
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
* <em>generated</em> label can no longer collide by construction. What remains checkable is what
* an operator still writes by hand: the {@code fleet.tabLabel} template and any per-profile
* {@code tabLabel} override.
*
* <p>Fatal rather than a warning, unlike {@link #warnUnknownTopLevelKeys}: an unknown key means
* a feature does nothing, while this means a feature does the opposite of what it says.
*
* @throws IllegalStateException when the fleet template or any profile's {@code tabLabel}
* override starts with a configured lead prefix
* @throws IllegalStateException when the fleet template or a profile {@code tabLabel} override
* can render as a configured lead or collaborator tab or match a
* lead-tab prefix, or when two entries — of either registry, or
* one of each — carry the same exact {@code tab}
* (case-insensitively)
*/
public void validateLeadTabPrefixes() {
if (fleet == null || fleet.leaders().isEmpty()) {
if (fleet == null) {
return;
}
List<String> bad = new ArrayList<>();
@@ -2717,51 +2728,175 @@ public record FleetConfig(
if (leader == null) {
return;
}
String tab = leader.tab();
String prefix = leader.tabPrefix();
// The fleet-wide template is checked once per prefix: it labels every member that has no
// override, so one bad template promotes the entire fleet, not one profile.
if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
+ "lead '" + leadName + "' (\"" + tab + "\")");
} else if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" starts with the tabPrefix of "
+ "lead '" + leadName + "' (\"" + prefix + "\")");
}
profiles().entrySet().stream()
.filter(e -> startsWithIgnoreCase(e.getValue().tabLabel(), prefix))
.map(Map.Entry::getKey)
.sorted()
.forEach(p -> bad.add("profile '" + p + "' overrides tabLabel with \""
+ profiles().get(p).tabLabel() + "\", which starts with the tabPrefix of "
+ "lead '" + leadName + "' (\"" + prefix + "\")"));
.forEach(p -> {
String label = profiles().get(p).tabLabel();
if (templateCanRenderAs(label, tab)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which can render as the tab of lead '" + leadName
+ "' (\"" + tab + "\")");
} else if (startsWithIgnoreCase(label, prefix)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which starts with the tabPrefix of lead '" + leadName
+ "' (\"" + prefix + "\")");
}
});
});
if (bad.isEmpty()) {
fleet.collaborators().forEach((collabName, collaborator) -> {
if (collaborator == null) {
return;
}
String tab = collaborator.tab();
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
+ "collaborator '" + collabName + "' (\"" + tab + "\")");
}
profiles().entrySet().stream()
.map(Map.Entry::getKey)
.sorted()
.forEach(p -> {
String label = profiles().get(p).tabLabel();
if (templateCanRenderAs(label, tab)) {
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
+ "\", which can render as the tab of collaborator '"
+ collabName + "' (\"" + tab + "\")");
}
});
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
+ ". A member labelled that way, while its pane carries no entry in the "
+ "spawned-member roster, is read back as a lead or collaborator and granted "
+ "that identity's authority. Change one of the two so member tabs cannot be "
+ "confused with a lead's or collaborator's tab.");
}
List<String> collisions = new ArrayList<>();
List<String> leadNames = fleet.leaders().keySet().stream().sorted().toList();
for (int i = 0; i < leadNames.size(); i++) {
String nameA = leadNames.get(i);
Leader a = fleet.leaders().get(nameA);
if (a == null || a.tab() == null || a.tab().isBlank()) {
continue;
}
for (int j = i + 1; j < leadNames.size(); j++) {
String nameB = leadNames.get(j);
Leader b = fleet.leaders().get(nameB);
if (b == null || b.tab() == null || b.tab().isBlank()) {
continue;
}
if (a.tab().equalsIgnoreCase(b.tab())) {
collisions.add("lead '" + nameA + "' and lead '" + nameB + "' both use tab \""
+ a.tab() + "\"");
}
}
}
List<String> collabNames = fleet.collaborators().keySet().stream().sorted().toList();
for (int i = 0; i < collabNames.size(); i++) {
String nameA = collabNames.get(i);
Collaborator a = fleet.collaborators().get(nameA);
if (a == null || a.tab() == null || a.tab().isBlank()) {
continue;
}
for (int j = i + 1; j < collabNames.size(); j++) {
String nameB = collabNames.get(j);
Collaborator b = fleet.collaborators().get(nameB);
if (b == null || b.tab() == null || b.tab().isBlank()) {
continue;
}
if (a.tab().equalsIgnoreCase(b.tab())) {
collisions.add("collaborator '" + nameA + "' and collaborator '" + nameB
+ "' both use tab \"" + a.tab() + "\"");
}
}
}
for (String leadName : leadNames) {
Leader lead = fleet.leaders().get(leadName);
if (lead == null || lead.tab() == null || lead.tab().isBlank()) {
continue;
}
for (String collabName : collabNames) {
Collaborator collaborator = fleet.collaborators().get(collabName);
if (collaborator == null || collaborator.tab() == null
|| collaborator.tab().isBlank()) {
continue;
}
if (lead.tab().equalsIgnoreCase(collaborator.tab())) {
collisions.add("lead '" + leadName + "' and collaborator '" + collabName
+ "' both use tab \"" + lead.tab() + "\"");
}
}
}
if (collisions.isEmpty()) {
return;
}
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
+ ". Every member labelled that way would be read back as a lead and granted "
+ "spawn/stop/send on the whole fleet. Change one of the two so member tabs and "
+ "lead tabs cannot be confused.");
throw new IllegalStateException("refusing to start: " + String.join("; ", collisions)
+ ". Tab identity is matched exactly, so only one of two entries sharing a tab can "
+ "ever be found — the other is silently unreachable. Give each lead and "
+ "collaborator its own exact tab.");
}
private static boolean templateCanRenderAs(String template, String tab) {
if (template == null || template.isBlank() || tab == null || tab.isBlank()) {
return false;
}
var placeholders = Pattern.compile("\\{(?:role|profile|model|n)}").matcher(template);
StringBuilder expression = new StringBuilder("^");
int literalStart = 0;
while (placeholders.find()) {
expression.append(Pattern.quote(template.substring(literalStart, placeholders.start())));
expression.append(".*");
literalStart = placeholders.end();
}
expression.append(Pattern.quote(template.substring(literalStart))).append("$");
return Pattern.compile(expression.toString(), Pattern.CASE_INSENSITIVE).matcher(tab).matches();
}
/** Case-insensitive prefix test that tolerates a null or blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
return false;
}
String stripped = label.strip();
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
}
/**
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
* entry names a {@code tab}. A pane-placed member lands inside the focused tab rather than its
* own, so it can land inside a lead's own labelled tab. {@link
* dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead purely by that tab's label — it does
* not exclude the member space — so a member that ends up there would be read back as the lead
* and granted spawn/stop/send on the whole fleet.
* or {@code fleet.collaborators} entry names a {@code tab}. A pane-placed member lands inside
* the focused tab rather than its own, so it can land inside a lead's or collaborator's own
* labelled tab. {@link dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead or collaborator
* purely by that tab's label — it does not exclude the member space — so a member that ends up
* there, while its pane carries no entry in the spawned-member roster, is read back as that
* lead or collaborator and granted that identity's authority.
*
* <p>Only a leader with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* <p>Only an entry with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
*
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
* {@code fleet.leaders} entry names a non-blank {@code tab}
* {@code fleet.leaders} or {@code fleet.collaborators} entry
* names a non-blank {@code tab}
*/
public void validatePanePlacementAgainstLeadTabs() {
if (fleet == null || fleet.leaders().isEmpty()) {
if (fleet == null) {
return;
}
boolean anyLeaderHasTab = fleet.leaders().values().stream()
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
if (!anyLeaderHasTab) {
boolean anyCollaboratorHasTab = fleet.collaborators().values().stream()
.anyMatch(c -> c != null && c.tab() != null && !c.tab().isBlank());
if (!anyLeaderHasTab && !anyCollaboratorHasTab) {
return;
}
List<String> bad = new ArrayList<>();
@@ -2774,10 +2909,12 @@ public record FleetConfig(
return;
}
throw new IllegalStateException("refusing to start: profile(s) " + bad
+ " use placement: pane while fleet.leaders names a tab. A pane-placed member can "
+ "land inside a lead's labelled tab and be read back as the lead, granted "
+ "spawn/stop/send on the whole fleet. Set placement: tab for each named profile, "
+ "or remove the tab from every fleet.leaders entry.");
+ " use placement: pane while fleet.leaders or fleet.collaborators names a tab. A "
+ "pane-placed member can land inside that labelled tab, and while its pane "
+ "carries no entry in the spawned-member roster, it is read back as the lead or "
+ "collaborator and granted that identity's authority. Set placement: tab for "
+ "each named profile, or remove the tab from every fleet.leaders and "
+ "fleet.collaborators entry.");
}
/**
@@ -2800,15 +2937,6 @@ public record FleetConfig(
}
}
/** Case-insensitive prefix test that tolerates a null/blank label. */
private static boolean startsWithIgnoreCase(String label, String prefix) {
if (label == null || prefix == null || prefix.isBlank()) {
return false;
}
String stripped = label.strip();
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
}
/**
* Reject a subscription profile whose {@code env:} block tries to reseat the Anthropic binding
* (CB-542).
@@ -2893,8 +3021,14 @@ public record FleetConfig(
* so duplicates are unrepresentable by construction once loaded — and {@link #load(Path)}
* already rejects a duplicated slot name at parse time, before the map collapses.
*
* @throws IllegalStateException when a slot names no profile or an unknown one, or when a lead
* can be neither found nor created, naming the offending entry
* <p>Also rejects a {@code fleet.collaborators} entry with no (or a blank) {@code tab}. A
* {@code profile}-less lead is still useful recognise-only — {@code tab} is the only field
* that matters to it either way. A collaborator carries no other field at all, so a blank
* {@code tab} leaves nothing for the entry to mean.
*
* @throws IllegalStateException when a slot names no profile or an unknown one, when a lead
* can be neither found nor created, or when a collaborator names
* no tab, naming the offending entry
*/
public void validateMembers() {
if (fleet == null) {
@@ -2932,6 +3066,16 @@ public record FleetConfig(
+ "auto-launched, labelled) purely by its tab, so every entry must name one.");
}
});
fleet.collaborators().forEach((name, collaborator) -> {
if (collaborator == null) {
return;
}
if (collaborator.tab() == null || collaborator.tab().isBlank()) {
bad.add("fleet.collaborators." + name + " has no tab: — a collaborator is "
+ "recognised purely by its tab, and carries no other field, so every "
+ "entry must name one.");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
@@ -11,12 +11,18 @@ public final class HerdrRouter implements AutoCloseable {
private final AgentControl memberAgents;
private final WorkspaceControl leadSpaces;
private final WorkspaceControl memberSpaces;
private final Predicate<String> isLead;
private final Predicate<String> routeToLead;
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> isLead) {
/**
* @param routeToLead true for a terminal whose pane lives in the lead herdr daemon — a lead's
* own pane or a configured collaborator's, both opened by a person at a
* terminal rather than spawned, so both are found in the lead daemon rather
* than the member one
*/
public HerdrRouter(HerdrClient lead, HerdrClient member, Predicate<String> routeToLead) {
this.lead = Objects.requireNonNull(lead, "lead");
this.member = member != null ? member : lead;
this.isLead = Objects.requireNonNull(isLead, "isLead");
this.routeToLead = Objects.requireNonNull(routeToLead, "routeToLead");
leadAgents = new AgentControl(this.lead);
memberAgents = this.member == this.lead ? leadAgents : new AgentControl(this.member);
leadSpaces = new WorkspaceControl(this.lead);
@@ -27,7 +33,7 @@ public final class HerdrRouter implements AutoCloseable {
public WorkspaceControl leadSpaces() { return leadSpaces; }
public AgentControl memberAgents() { return memberAgents; }
public WorkspaceControl memberSpaces() { return memberSpaces; }
public AgentControl agentsFor(String targetId) { return isLead.test(targetId) ? leadAgents : memberAgents; }
public AgentControl agentsFor(String targetId) { return routeToLead.test(targetId) ? leadAgents : memberAgents; }
HerdrClient leadClient() { return lead; }
HerdrClient memberClient() { return member; }
@@ -96,13 +96,19 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
/** What a matched tab names: a lead or a collaborator. */
private enum Kind { LEAD, COLLABORATOR }
/** One matched tab's name and what it names. */
private record Entry(String name, Kind kind) {}
private final HerdrClient herdr;
private final Map<String, String> tabToName;
private final Map<String, Entry> tabToEntry;
private final Set<String> excludedWorkspaceLabels;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached = Map.of();
private Map<String, Entry> cached = Map.of();
private long scannedAtNanos;
private boolean everScanned;
@@ -126,26 +132,53 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this(herdr, tabToName, Map.of(), excludedWorkspaceLabels, ttlNanos, clock);
}
/**
* As {@link #LeadTabScanner(HerdrClient, Map, Set, long, LongSupplier)}, additionally scanning
* for configured collaborator tabs in the same pass.
*
* @param collaboratorTabToName every configured collaborator's exact tab label → its name
* ({@code fleet.collaborators.<name>.tab}), matched the same way as
* {@code tabToName}
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Map<String, String> collaboratorTabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabToName = normalize(tabToName);
this.tabToEntry = buildTabIndex(tabToName, collaboratorTabToName);
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.ttlNanos = ttlNanos;
this.clock = clock;
}
/** Keys stripped and lower-cased once, so every lookup is a plain map hit. */
private static Map<String, String> normalize(Map<String, String> tabToName) {
if (tabToName == null || tabToName.isEmpty()) {
return Map.of();
/**
* Keys stripped and lower-cased once, so every lookup is a plain map hit. Leads and
* collaborators merge into a single index, so {@link #scan()} matches both kinds in one pass
* over the tab list; a label naming both a lead and a collaborator takes the lead entry —
* leads are put last, so a colliding key's lead entry is the one that overwrites — since a lead
* can already do everything a collaborator can. Config validation already refuses a lead and a
* collaborator sharing one exact tab, so this ordering is defence in depth, not the control.
*/
private static Map<String, Entry> buildTabIndex(Map<String, String> tabToName,
Map<String, String> collaboratorTabToName) {
Map<String, Entry> out = new LinkedHashMap<>();
putNormalized(out, collaboratorTabToName, Kind.COLLABORATOR);
putNormalized(out, tabToName, Kind.LEAD);
return Collections.unmodifiableMap(out);
}
private static void putNormalized(Map<String, Entry> out, Map<String, String> tabToName, Kind kind) {
if (tabToName == null) {
return;
}
Map<String, String> out = new LinkedHashMap<>();
tabToName.forEach((tab, name) -> {
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
out.put(tab.strip().toLowerCase(Locale.ROOT), name);
out.put(tab.strip().toLowerCase(Locale.ROOT), new Entry(name, kind));
}
});
return Collections.unmodifiableMap(out);
}
/**
@@ -156,6 +189,29 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
*/
@Override
public synchronized Map<String, String> get() {
return byKind(refresh(), Kind.LEAD);
}
/**
* The current {@code terminal_id → collaborator name} map, sharing the same scan and cache as
* {@link #get()} — both kinds are matched in one pass, so this never costs a second herdr call.
*/
public synchronized Map<String, String> collaborators() {
return byKind(refresh(), Kind.COLLABORATOR);
}
private static Map<String, String> byKind(Map<String, Entry> entries, Kind kind) {
Map<String, String> out = new LinkedHashMap<>();
entries.forEach((terminal, entry) -> {
if (entry.kind() == kind) {
out.put(terminal, entry.name());
}
});
return Collections.unmodifiableMap(out);
}
/** Rescans if the cache has expired, otherwise returns the cached answer. */
private Map<String, Entry> refresh() {
long now = clock.getAsLong();
if (everScanned && now - scannedAtNanos < ttlNanos) {
return cached;
@@ -165,21 +221,21 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
scannedAtNanos = now;
everScanned = true;
try {
Map<String, String> fresh = scan();
Map<String, Entry> fresh = scan();
if (!fresh.equals(cached)) {
log.info("lead panes: {}", fresh);
log.info("lead/collaborator panes: {}", fresh);
}
cached = fresh;
} catch (HerdrException e) {
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
log.warn("lead-tab scan failed, keeping the {} entr(y/ies) already known: {}",
cached.size(), e.getMessage());
}
return cached;
}
/** One full pass: labelled tabs → live agents in them → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
private Map<String, Entry> scan() {
Map<String, Entry> entryByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace ws = Workspace.from(w);
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
@@ -187,21 +243,22 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
String name = leadNameOf(tab.label());
if (name != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), name);
Entry entry = entryOf(tab.label());
if (entry != null && tab.tabId() != null) {
entryByTab.put(tab.tabId(), entry);
}
}
}
if (nameByTab.isEmpty()) {
if (entryByTab.isEmpty()) {
gracedTerminals = Set.of();
return Map.of();
}
// fleetd #359: a labelled tab is only a lead when herdr also reports a running agent in
// it — the same liveness signal LeadLauncher.countLeads trusts for the identical purpose.
// Without this, a tab left behind by a session that has since died reads as live forever.
// fleetd #359: a labelled tab is only a lead (or collaborator) when herdr also reports a
// running agent in it — the same liveness signal LeadLauncher.countLeads trusts for the
// identical purpose. Without this, a tab left behind by a session that has since died reads
// as live forever.
Set<String> tabsWithAgent = new HashSet<>();
for (JsonNode a : herdr.call("agent.list").path("agents")) {
String tabId = a.path("tab_id").asText(null);
@@ -210,18 +267,18 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
Map<String, Entry> byTerminal = new LinkedHashMap<>();
Set<String> stillGraced = new HashSet<>();
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String tabId = p.path("tab_id").asText(null);
String name = nameByTab.get(tabId);
Entry entry = entryByTab.get(tabId);
String terminal = p.path("terminal_id").asText(null);
if (name == null || terminal == null || terminal.isBlank()) {
if (entry == null || terminal == null || terminal.isBlank()) {
continue;
}
if (tabsWithAgent.contains(tabId)) {
byTerminal.put(terminal, name);
byTerminal.put(terminal, entry);
continue;
}
// No agent reported for this tab, but its tab/pane are still here — this is the
@@ -230,7 +287,7 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
// reported as live; a terminal we never reported live gets none, so the original #359
// fix (a genuinely dead tab is never reported) is unaffected for the common case.
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
byTerminal.put(terminal, name);
byTerminal.put(terminal, entry);
stillGraced.add(terminal);
}
}
@@ -239,18 +296,19 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
/**
* The lead name a tab label declares, or {@code null} if it names none of the configured leads.
* The entry a tab label declares, or {@code null} if it names neither a configured lead nor a
* configured collaborator.
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToEntry} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix. The match strips a trailing
* {@link PendingCloseMarker} first, so a tab {@code LeadLauncher} has flagged as maybe-dead but
* not yet closed keeps resolving normally while that reconcile is pending.
*/
private String leadNameOf(String label) {
private Entry entryOf(String label) {
if (label == null) {
return null;
}
return tabToName.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
return tabToEntry.get(PendingCloseMarker.strip(label).toLowerCase(Locale.ROOT));
}
}
@@ -0,0 +1,146 @@
package dev.ltms.fleet.herdr;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.charset.StandardCharsets;
import java.util.List;
import java.util.function.IntFunction;
/**
* The three protections every {@code agent.start} caller needs against herdr's pane-typed launch
* surface (fleetd #220, #727): a byte-limit check on the assembled command line, a bounded retry
* on {@code agent_pane_busy} (the target pane's shell has not reached its prompt yet), and a
* bounded retry on {@code agent_name_taken} with a fresh name each attempt. One implementation —
* every caller of {@code agent.start}, lead or member, goes through this seam rather than carrying
* its own copy.
*/
public final class ResilientAgentLaunch {
private static final Logger log = LoggerFactory.getLogger(ResilientAgentLaunch.class);
private ResilientAgentLaunch() {
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkFits}.
*/
public static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** herdr rejects a duplicate agent {@code name}; a caller retries a bumped name this many times. */
public static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
public static final int SHELL_READY_RETRIES = 20;
/** Raised by {@link #checkFits} when the assembled command cannot fit the pane line. */
public static final class TooLargeException extends RuntimeException {
public TooLargeException(String message) {
super(message);
}
}
/**
* Verify the assembled launch line fits the pane herdr types it into. herdr does not exec the
* launch command — it TYPES it into the pane as one line, and a pty line buffer holds only
* {@value #PANE_COMMAND_BYTE_LIMIT} bytes. Everything past that byte is dropped with no error
* anywhere: herdr answers "agent started", the backend exits on the mangled argument it was
* handed, the pane closes, and the only symptom is a readiness timeout with no reason. That is
* how fleetd #214 broke every claude-code spawn — one 50-byte flag pushed a 978-byte command to
* 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than start something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @param label names the launch in the refusal message (a profile name)
* @param argv the full argv, including the executable at index 0
* @throws TooLargeException naming the limit, the estimate, and the longest argument
*/
public static void checkFits(String label, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new TooLargeException(
"launch command for " + label + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* Start an agent into {@code paneId}, retrying {@code agent_pane_busy} up to {@code retries}
* times with {@code sleeper} run between attempts.
*
* @throws HerdrException the last {@code agent_pane_busy} failure once {@code retries} is
* spent, or immediately for any other herdr failure
*/
public static Agent startAwaitingShellPrompt(AgentControl agents, String name, String kind,
List<String> args, String paneId,
int retries, Runnable sleeper) {
HerdrException busy = null;
for (int attempt = 0; attempt < retries; attempt++) {
try {
return agents.start(name, kind, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
/**
* Start an agent under a freshly generated name each attempt, retrying {@code agent_name_taken}
* up to {@code nameRetries} times — herdr refuses a duplicate {@code name} outright, so a stale
* registry entry (a crashed session, a name the registry has not yet released) must not block a
* legitimate relaunch. Each attempt also carries its own {@link #startAwaitingShellPrompt} retry.
*
* @param nameForAttempt called once per attempt (0-based) to produce that attempt's name
* @throws HerdrException the last {@code agent_name_taken} failure once {@code nameRetries} is
* spent, or immediately for any other herdr failure
*/
public static Agent startUniquelyNamed(AgentControl agents, String kind, List<String> args,
String paneId, IntFunction<String> nameForAttempt,
int nameRetries, int shellReadyRetries, Runnable sleeper) {
HerdrException last = null;
for (int attempt = 0; attempt < nameRetries; attempt++) {
String name = nameForAttempt.apply(attempt);
try {
return startAwaitingShellPrompt(agents, name, kind, args, paneId,
shellReadyRetries, sleeper);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("agent name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
}
@@ -5,6 +5,7 @@ import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.PendingCloseMarker;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -12,6 +13,7 @@ import dev.ltms.fleet.launch.ClaudeCodeArguments;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
@@ -19,6 +21,7 @@ import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import dev.ltms.fleet.peer.PeerLauncher;
/**
@@ -80,9 +83,19 @@ public final class LeadLauncher {
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
/** Attempts {@link #relaunch(String)} makes before giving up and returning {@code null}. */
static final int RELAUNCH_ATTEMPTS = 3;
private final AgentControl agents;
private final WorkspaceControl spaces;
private final FleetConfig cfg;
private final Runnable sleeper;
// Per-process token mixed into each lead agent name so a fresh daemon process (seq back at 0)
// cannot collide with a same-name lead that outlived a restart — the same scheme
// HerdrPeerLauncher uses for members (fleetd #727).
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
private final AtomicLong nameSeq = new AtomicLong();
/**
* @param agents herdr agent control (start, list)
@@ -90,9 +103,27 @@ public final class LeadLauncher {
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg) {
this(agents, spaces, cfg, () -> sleepUninterruptibly(300));
}
/**
* Test seam: as above, plus an injectable {@code sleeper} for the {@code agent_pane_busy}
* retry (fleetd #727), so a test can prove the retry budget without a real sleep.
*/
LeadLauncher(AgentControl agents, WorkspaceControl spaces, FleetConfig cfg, Runnable sleeper) {
this.agents = agents;
this.spaces = spaces;
this.cfg = cfg;
this.sleeper = sleeper;
}
/** Uninterruptible sleep — the production {@link #sleeper} between {@code agent_pane_busy} retries. */
private static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
/**
@@ -173,24 +204,14 @@ public final class LeadLauncher {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
}
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have fleetd start it.",
name, name);
continue;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
ResolvedLead resolved = resolveLaunchable(name);
if (resolved == null) {
continue;
}
for (int i = running; i < wanted; i++) {
if (launch(name, lead, profile)) {
if (launch(name, resolved.lead(), resolved.profile()) != null) {
started++;
}
}
@@ -198,6 +219,82 @@ public final class LeadLauncher {
return started;
}
/** A declared lead paired with the profile it launches on — {@link #resolveLaunchable}'s result. */
private record ResolvedLead(FleetConfig.Leader lead, FleetConfig.Profile profile) {
}
/**
* The declared {@code Leader} and its {@code Profile} for {@code name}, read from the config
* snapshot this launcher was constructed with.
*
* @return the resolved pair, or {@code null} (having logged) if {@code name} is not declared
* under {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
* recognise-only lead), or its {@code profile:} is not configured. Shared by
* {@link #ensureLeads()} and {@link #relaunch(String)} so the three refusals and their
* wording live in one place.
*/
private ResolvedLead resolveLaunchable(String name) {
FleetConfig.Leader lead = cfg.fleet().leaders().get(name);
if (lead == null) {
log.warn("lead '{}' is not declared under fleet.leaders — not launching", name);
return null;
}
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' names no profile — it can be recognised but not launched. Add "
+ "`profile:` under fleet.leaders.{} to have fleetd start it.", name, name);
return null;
}
FleetConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
return null;
}
return new ResolvedLead(lead, profile);
}
/**
* Start the named lead from the config snapshot this launcher was constructed with — not a
* live read, so a lead's {@code profile:} or {@code tab:} edited in config needs a daemon
* restart to take effect here — outside of {@link #ensureLeads()}'s {@code instances}
* bookkeeping.
*
* @return the started {@link Agent}, or {@code null} if {@code name} is not declared under
* {@code fleet.leaders}, that lead names no {@code profile:} (a {@code tab:}-only,
* recognise-only lead), its {@code profile:} is not configured, or every attempt up to
* {@link #RELAUNCH_ATTEMPTS} failed to start it. Never throws.
*
* <p>Does not count how many instances of this lead are already live. {@link #ensureLeads()}'s
* count exists to avoid starting a second orchestrator; the caller of this method has already
* decided to replace the lead and owns that decision.
*
* <p>Retries the whole launch attempt — not only the {@code agent_name_taken}/
* {@code agent_pane_busy} cases {@link ResilientAgentLaunch} already retries inside one
* {@code agents.start} call — up to {@link #RELAUNCH_ATTEMPTS} times, sleeping via the
* injected sleeper between attempts, and returns the agent from the first attempt that
* succeeds.
*/
public Agent relaunch(String name) {
ResolvedLead resolved = resolveLaunchable(name);
if (resolved == null) {
return null;
}
for (int attempt = 1; attempt <= RELAUNCH_ATTEMPTS; attempt++) {
Agent started = launch(name, resolved.lead(), resolved.profile());
if (started != null) {
return started;
}
if (attempt < RELAUNCH_ATTEMPTS) {
sleeper.run();
}
}
return null;
}
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
@@ -306,8 +403,17 @@ public final class LeadLauncher {
return null;
}
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
/**
* Start one lead. Returns null (having logged) rather than throwing on any failure.
*
* <p>Goes through the same {@link ResilientAgentLaunch} seam every member spawn uses
* (fleetd #727): the assembled argv is refused outright if it cannot fit the pane line herdr
* types it into, a stale {@code agent_name_taken} (a crashed session's name the registry has
* not yet released) is retried under a fresh per-attempt name rather than refusing the whole
* relaunch, and a seed pane whose shell has not reached its prompt yet ({@code
* agent_pane_busy}) is retried rather than failing on the first miss.
*/
private Agent launch(String name, FleetConfig.Leader lead, FleetConfig.Profile profile) {
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
@@ -323,8 +429,12 @@ public final class LeadLauncher {
// Same shape as the member launchers: herdr resolves the executable from `kind`, so
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
List<String> argv = leadArgv(profile);
Agent started = agents.start("lead-" + name, herdrKind(profile),
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId());
ResilientAgentLaunch.checkFits(profile.profile(), argv);
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
Agent started = ResilientAgentLaunch.startUniquelyNamed(agents, herdrKind(profile), args,
tab.rootPaneId(),
attempt -> "lead-" + name + "-" + nameNonce + "-" + nameSeq.incrementAndGet(),
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
// Label AFTER the start succeeds. A label written before would survive a failed start
// and then read back as a live lead on the next boot, which is the exact staleness the
@@ -334,7 +444,7 @@ public final class LeadLauncher {
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
name, profile.profile(), tab.tab().tabId(), started.paneId(),
started.terminalId(), label, cwd);
return true;
return started;
} catch (RuntimeException e) {
log.warn("lead '{}' failed to launch on profile '{}': {}",
name, profile.profile(), e.getMessage());
@@ -346,7 +456,7 @@ public final class LeadLauncher {
tab.tab().tabId(), cleanup.getMessage());
}
}
return false;
return null;
}
}
@@ -164,7 +164,13 @@ public final class LeadRollover {
* The handover file's modified time is not after {@link #open}'s request timestamp, or is
* older than {@code maxDocAgeSeconds}.
*/
HANDOVER_STALE
HANDOVER_STALE,
/**
* This lead terminal already has a roll running: an earlier {@link #confirm} call claimed
* it and that roll's continuation has not released it yet. {@code detail} names the lead
* terminal and the token that holds the claim.
*/
ROLL_ALREADY_RUNNING
}
/**
@@ -305,6 +311,15 @@ public final class LeadRollover {
*/
private final Consumer<Runnable> continuationRunner;
private final Map<String, PendingRollover> pending = new ConcurrentHashMap<>();
/**
* Lead terminal → the token of the roll currently holding that terminal exclusive, for
* {@link #confirm}'s single-flight claim. {@link #confirm} claims an entry here with an
* atomic put-if-absent once every other gate has passed, refusing with {@link
* RefusalReason#ROLL_ALREADY_RUNNING} when a claim is already held; {@link #runRollover}
* releases it in a {@code finally}, on both the success and the thrown-exception path. A
* terminal absent from this map has no roll currently in flight for it.
*/
private final Map<String, String> rollingByTerminal = new ConcurrentHashMap<>();
/**
* Finished tokens → what actually happened, for {@link #status}. Bounded by {@link
* #OUTCOME_HISTORY_CAP}, oldest evicted first ({@code removeEldestEntry} on an insertion-order
@@ -487,6 +502,16 @@ public final class LeadRollover {
return docCheck;
}
// Single-flight claim: atomic put-if-absent, taken only after every other gate has
// passed, so a refused confirm() never takes it. A non-null previous value means a
// different, still-running roll already holds this lead terminal.
String holder = rollingByTerminal.putIfAbsent(p.leadTerminal(), token);
if (holder != null) {
return RollDecision.refused(RefusalReason.ROLL_ALREADY_RUNNING,
"lead terminal " + p.leadTerminal() + " already has a roll running under token "
+ holder);
}
// Record IN_PROGRESS BEFORE removing from `pending` — see RollState#IN_PROGRESS and
// OUTCOME_HISTORY_CAP's javadoc. This ordering means `token` is written into `outcomes`
// while it is STILL present in `pending`; status() checks `outcomes` first (see that
@@ -499,7 +524,23 @@ public final class LeadRollover {
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
continuationRunner.accept(() -> runRollover(p, cfg));
try {
continuationRunner.accept(() -> runRollover(p, cfg));
} catch (RuntimeException e) {
// continuationRunner can reject the hand-off itself (e.g. a bounded executor's
// RejectedExecutionException) before runRollover ever starts, so runRollover's own
// finally — the only other place that releases rollingByTerminal — never runs either.
// Release the claim here and overwrite the IN_PROGRESS entry with a terminal outcome,
// or this lead terminal could never be rolled again and status() would report
// IN_PROGRESS forever for a roll that in fact never started.
log.warn("lead-rollover: continuationRunner rejected token={} lead={}: {} — the roll "
+ "never started; releasing its claim and reporting it as FAILED",
token, callerTerminal, e.toString(), e);
rollingByTerminal.remove(p.leadTerminal(), token);
outcomes.put(token, new RollStatus(RollState.FAILED,
"continuationRunner rejected this roll before it ever started: " + e.toString()
+ " — the roll never ran; open() a fresh rollover request"));
}
return RollDecision.approved();
}
@@ -539,6 +580,12 @@ public final class LeadRollover {
"the roll's continuation threw " + e.toString() + " — the roll is dead and will "
+ "not retry itself; check the daemon log for the stack trace, then open() "
+ "a fresh rollover request"));
} finally {
// Release the single-flight claim on both the normal return and the thrown-exception
// path above — a release only on success would leave this lead terminal unrollable
// forever after one failure. The conditional two-argument remove only clears the
// entry this roll itself holds, never a different roll's claim on the same terminal.
rollingByTerminal.remove(p.leadTerminal(), p.token());
}
}
@@ -114,6 +114,13 @@ public final class FleetMcp {
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
*/
private final ConnectionIdentity identity;
/**
* Kept as a field (rather than only captured by the {@code contextExtractor} closure) so
* {@link #denyFor} can read {@link CallerResolver#knownLeadOrCollaborator()} — the classifier a
* collaborator's {@code SEND} is checked against, built from the same lead and collaborator
* maps {@link #identity}-based resolution reads.
*/
private final CallerResolver callers;
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
@@ -327,7 +334,7 @@ public final class FleetMcp {
* caller explicitly saying so — never by omitting a {@link CallerResolver} the way the old
* {@code callers == null} idiom allowed. {@code callers} itself is required either way: even
* under {@link #UNENFORCED}, the one real {@link CallerResolver} still resolves every caller's
* {@link Principal} (so {@code markSpawnedMemberPresent}/{@code recordPrimarySingleton} see a
* {@link Principal} (so {@code markTrackedCallerPresent}/{@code recordPrimarySingleton} see a
* real identity), and {@link #denyFor} is the only thing that changes.
*/
public enum AuthorizationMode { ENFORCED, UNENFORCED }
@@ -409,6 +416,7 @@ public final class FleetMcp {
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
== AuthorizationMode.ENFORCED;
this.identity = identity;
this.callers = callers;
this.leadChannel = leadChannel;
this.peers = peers == null ? List.of() : List.copyOf(peers);
this.capacity = capacity;
@@ -434,10 +442,10 @@ public final class FleetMcp {
// fall back to here — AuthorizationMode governs enforcement, not identity.
Principal p = callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"));
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
// lead, which carries its pane too, while including every spawned member role.
// Enrolling a lead would count it as an available member in the roster.
markSpawnedMemberPresent(p, presence);
// Guards on the ROLE, not on the terminal being null — this marks presence for
// a worker, an architect, or the unconfigured-pane floor, and excludes a lead
// or a collaborator even though each carries its own pane too.
markTrackedCallerPresent(p, presence);
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
@@ -451,12 +459,14 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_send", req.arguments()),
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
Principal caller = principal(exchange);
String callerTerminal = caller.terminal();
String callerOwner = caller.ownerKey();
// CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback.
// An architect delegates as its own pane but must never become the fallback that
// no-delegation inbox nudges target as if it were the primary (the per-target
// delegation map does not cure the singleton).
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
recordPrimarySingleton(primaryRegistry, callerTerminal, caller);
Map<String, Object> a = req.arguments();
String target = str(a, "sessionId");
String content = str(a, "content");
@@ -472,18 +482,18 @@ public final class FleetMcp {
// Answering a worker's fleet_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a));
return answer(messages, turnId, content, timeoutMs(a), callerOwner);
}
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
// session lock and queued delivery — via the accepted-delivery callback, never at
// request time. A concurrent sender that times out BUSY therefore cannot steal a
// live turn's reply routing without ever owning the turn.
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, callerTerminal);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, target, content, onAccepted, workers.profiles())
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
? sendAsync(messages, target, content, onAccepted, workers.profiles(), caller)
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles(), callerOwner);
};
// fleet_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
@@ -506,7 +516,7 @@ public final class FleetMcp {
(exchange, req) -> {
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_status", req.arguments()), null);
if (denied != null) return denied;
return status(messages, str(req.arguments(), "sessionId"));
return status(messages, str(req.arguments(), "sessionId"), principal(exchange).ownerKey());
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> pollHandler =
(exchange, req) -> {
@@ -516,7 +526,8 @@ public final class FleetMcp {
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
if (denied != null) return denied;
return poll(messages, leadChannel, str(a, "ticket"), target, coordId);
return poll(messages, leadChannel, str(a, "ticket"), target, coordId,
principal(exchange).ownerKey());
};
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
@@ -552,8 +563,10 @@ public final class FleetMcp {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine, outage,
leadSeats, leadContextGauge, leadConfigDirs, callers.leads(),
callerTerminal(exchange),
callers.collaborators(), collaboratorsVisibleTo(principal(exchange)),
new CoordinationSource(leadChannel, peers),
coordinatorVisibleTo(principal(exchange)));
coordinatorVisibleTo(principal(exchange)),
leadsVisibleTo(principal(exchange)), membersVisibleTo(principal(exchange)));
};
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
(exchange, req) -> {
@@ -689,8 +702,8 @@ public final class FleetMcp {
if (!authorizationEnforced) {
return null; // AuthorizationMode.UNENFORCED: authorization not enforced (fleetd #518)
}
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ) {
if (Authz.permits(caller, action, target, callers.knownLeadOrCollaborator())) {
if (action != Authz.Action.READ && action != Authz.Action.TASK_READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return null;
@@ -735,7 +748,53 @@ public final class FleetMcp {
return caller.isPrimary();
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
/**
* fleetd #703: who may see {@code fleet_list}'s {@code collaborators} array — a roster of the
* human-opened tabs this daemon recognises as named peers. Visible to exactly the roles that
* may {@link Authz.Action#SEND} to a named peer ({@code Authz.java}'s {@code SEND} case):
* the primary, an architect, and a collaborator (reaching another collaborator or a lead). A
* worker holds {@code READ} but can never {@code SEND} to a collaborator, so listing them to a
* worker would expose which human tabs exist on the host with no use to that caller. Split out
* for the same reason as {@link #coordinatorVisibleTo}: the decision must be unit-testable
* without fabricating an SDK {@code McpSyncServerExchange}, and the handler must call this
* named predicate rather than inlining the check.
*/
static boolean collaboratorsVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
}
/**
* Who may see {@code fleet_list}'s {@code leads} array — exactly the roles that may
* {@link Authz.Action#SEND} to a lead: the primary, an architect, and a collaborator. A
* collaborator's own {@code fleet_whoami} carries no lead address, and {@code leads} is the
* only place this tool gives one, so a collaborator needs this array to use the send it
* already holds. A worker can never {@code SEND} at all, so it still sees neither this array
* nor {@code members}; a worker's own facts come from {@code fleet_whoami} instead. Split
* out for the same reason as {@link #coordinatorVisibleTo} and
* {@link #collaboratorsVisibleTo}: the decision must be unit-testable without fabricating an
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather
* than inlining the check.
*/
static boolean leadsVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect() || caller.isCollaborator();
}
/**
* Who may see {@code fleet_list}'s {@code members} array — the primary and an architect,
* which may {@link Authz.Action#SEND} to a member. A worker can never {@code SEND} at all,
* and a collaborator may {@code SEND} only to a lead or another collaborator, never to a
* spawned member, so neither sees this array even though {@link #leadsVisibleTo} grants a
* collaborator the sibling one.
*/
static boolean membersVisibleTo(Principal caller) {
return caller.isPrimary() || caller.isArchitect();
}
/**
* The terminal of the caller on this call's connection, or {@code null} when that caller carries
* no terminal, which is only the unnamed primary. A named lead, an architect and a worker each
* carry one.
*/
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
String s = v == null ? null : v.toString();
@@ -777,9 +836,14 @@ public final class FleetMcp {
return identity;
}
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember()) {
/**
* Mark a caller present for the injector readiness gate, when its deliverability depends on
* proving a live MCP contact: a worker, an architect, or the unconfigured-pane floor. A lead
* or a collaborator is excluded — each is already deliverable through its own named-registry
* entry.
*/
static void markTrackedCallerPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember() || caller.isObserver()) {
presence.markPresent(caller.terminal());
}
}
@@ -852,7 +916,8 @@ public final class FleetMcp {
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted, Set<String> profiles) {
Long timeoutMs, Runnable onAccepted, Set<String> profiles,
String callerOwner) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
@@ -862,7 +927,7 @@ public final class FleetMcp {
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
return formatReply(messages.send(sessionId, content, timeout, onAccepted, callerOwner), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
@@ -872,13 +937,15 @@ public final class FleetMcp {
* {@code fleet_send} carrying a {@code turnId}: the primary's answer to a worker's
* {@code fleet_ask} (CB-205). Resolves the worker's blocked question and blocks for its reply as
* it resumes the same turn — surfaced to the primary identically to a normal send.
* {@code callerOwner} must match the turn's recorded owner or this is refused.
*/
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs) {
static McpSchema.CallToolResult answer(MessageService messages, String turnId, String content, Long timeoutMs,
String callerOwner) {
if (isBlank(turnId) || isBlank(content)) {
return error("turnId and content are required to answer a worker's question");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
return formatReply(messages.answer(turnId, content, timeout), timeout);
return formatReply(messages.answer(turnId, content, timeout, callerOwner), timeout);
}
/**
@@ -927,6 +994,8 @@ public final class FleetMcp {
+ "\" and content set to your answer; the worker resumes the same turn.");
case STALE_TURN -> error("that question is no longer open — it timed out or was already "
+ "answered (turnId stale)");
case NOT_TURN_OWNER -> error("this turn belongs to a different delegation — only the caller "
+ "that opened it may answer it");
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> text("[no reply within " + timeout + "ms — worker "
+ r.outcome().name().toLowerCase().replace("timed_out_", "") + "; retry or poll status]");
// fleetd #571: delivery is unknown here — agent.prompt pastes and submits in one call,
@@ -948,6 +1017,16 @@ public final class FleetMcp {
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles) {
return sendAsync(messages, sessionId, content, onAccepted, profiles, null);
}
/**
* As above, recording {@code creator}'s owner key so a later {@code fleet_poll{ticket}} only
* hands the result back to the same caller — see {@link MessageService#poll(String, String)}.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles,
Principal creator) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
@@ -955,7 +1034,7 @@ public final class FleetMcp {
if (targetError != null) {
return targetError;
}
String ticket = messages.sendAsync(sessionId, content, onAccepted);
String ticket = messages.sendAsync(sessionId, content, onAccepted, creator);
return text("accepted — task delegated. Poll fleet_poll with ticket=" + ticket);
}
@@ -1035,31 +1114,21 @@ public final class FleetMcp {
}
/**
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments
* (fleetd #272, widened by fleetd #421).
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments.
*
* <p>{@code fleet_poll} is now <strong>three operations behind one tool name</strong>. With
* <p>{@code fleet_poll} is <strong>three operations behind one tool name</strong>. With
* {@code ticket} it observes an async delegation and changes nothing, which is a {@link
* Authz.Action#READ}. With {@code target} it calls {@link MessageService#drainReplies} on that
* session -- the replies are removed from the inbox and a second call returns nothing -- so it
* is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for removing a
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}. With
* {@code coordId} it reads (never acks) this daemon's own held lead-to-lead mail, which is a
* {@link Authz.Action#COORD_READ} -- <strong>not</strong> {@code READ}, even though nothing is
* consumed: {@code READ}'s grant is open to every authenticated role on the premise that the
* roster carries no secrets, and a lead-to-lead body is not the roster. Mapping a non-destructive
* peer-mail read to {@code READ} would let any worker read every peer lead's mail in full.
*
* <p>Before this method existed (fleetd #272) the handler passed a constant {@code READ} for
* both of the original branches. {@code READ} is open to every authenticated role, so any
* worker could read a peer's id out of {@code fleet_list} and destroy the replies that peer had
* queued for the primary. The gate failed open, and it did so because the required action is a
* function of the arguments while the handler chose it before looking at them.
* Authz.Action#TASK_READ}. With {@code target} it calls {@link MessageService#drainReplies} on
* that session -- the replies are removed from the inbox and a second call returns nothing --
* so it is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for
* removing a single message, and the same one the REST path uses at
* {@code FleetApp.drainReplies}. With {@code coordId} it reads (never acks) this daemon's own
* held lead-to-lead mail, which is a {@link Authz.Action#COORD_READ} -- <strong>not</strong>
* {@code TASK_READ} or {@code READ}: a lead-to-lead body is a different inbox from either, and
* folding it into either would let any worker or architect read every peer lead's mail in full.
*
* <p>The choice lives in this method, and not inline in the handler, so that a test can assert
* the mapping the handler actually uses. {@code FleetMcpAuthzTest} already checked every
* {@link Authz.Action} against every {@link Role} and passed throughout -- it tested the policy
* table, which was correct, while the defect was in which action the caller handed it.
* the mapping the handler actually uses.
*
* <p>Checked first, and exclusively of {@code target}: a call naming {@code coordId} is reading
* a different inbox entirely (this daemon's own lead channel, never a worker's), so it takes
@@ -1072,7 +1141,30 @@ public final class FleetMcp {
if (!isBlank(coordId)) {
return Authz.Action.COORD_READ;
}
return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
return isBlank(target) ? Authz.Action.TASK_READ : Authz.Action.DRAIN;
}
/**
* Which authorization action a {@code fleet_send} call needs, decided by its arguments.
*
* <p>{@code fleet_send} is three call shapes behind one tool name, mirroring {@link
* #pollAction}. With {@code coordId} it addresses a peer lead on another daemon over the
* coordination broker, which is {@link Authz.Action#COORD_SEND}. With {@code turnId} it
* resolves a worker's blocked {@code fleet_ask} and resumes that turn, which is {@link
* Authz.Action#ANSWER}. Otherwise it delivers to a local session by {@code sessionId}, which is
* the plain {@link Authz.Action#SEND}.
*
* <p>Checked in the same order the handler branches: {@code coordId} first and exclusively of
* {@code turnId}, matching {@link #sendToLead}'s own mutual-exclusion check.
*
* @param coordId the {@code coordId} argument of the call, or {@code null}/blank when absent
* @param turnId the {@code turnId} argument of the call, or {@code null}/blank when absent
*/
static Authz.Action sendAction(String coordId, String turnId) {
if (!isBlank(coordId)) {
return Authz.Action.COORD_SEND;
}
return isBlank(turnId) ? Authz.Action.SEND : Authz.Action.ANSWER;
}
/**
@@ -1097,10 +1189,11 @@ public final class FleetMcp {
*/
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
return switch (tool) {
case SEND -> Authz.Action.SEND;
case SEND -> sendAction(str(arguments, "coordId"), str(arguments, "turnId"));
case REPLY -> Authz.Action.REPLY;
case ASK -> Authz.Action.ASK;
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case STATUS -> Authz.Action.TASK_READ;
case LIST, PROFILES, WHOAMI -> Authz.Action.READ;
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case ACK -> Authz.Action.DRAIN;
case SPAWN -> Authz.Action.SPAWN;
@@ -1120,9 +1213,23 @@ public final class FleetMcp {
* exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this
* is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently
* ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches.
*
* <p>Does not check who owns {@code ticket} — see the overload that takes {@code
* callerTerminal} for that. Callers that do not resolve a caller terminal (tests, or a surface
* with no connection-based identity) use this one.
*/
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId) {
return poll(messages, leadChannel, ticket, target, coordId, null);
}
/**
* As above, refusing a ticket lookup whose caller owner key differs from the key that created it
* — see {@link MessageService#poll(String, String)}. {@code callerOwner} comes from the calling
* connection's resolved principal, never a client-supplied value.
*/
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId, String callerOwner) {
if (!isBlank(coordId)) {
return pollHeldPeerMail(leadChannel, coordId);
}
@@ -1136,7 +1243,7 @@ public final class FleetMcp {
if (isBlank(ticket)) {
return error("ticket (or target) is required");
}
MessageService.TaskView v = messages.poll(ticket);
MessageService.TaskView v = messages.poll(ticket, callerOwner);
if (v == null) {
return error("unknown ticket: " + ticket + " (never issued, or expired)");
}
@@ -1243,16 +1350,19 @@ public final class FleetMcp {
/**
* {@code fleet_status}: the live lifecycle status of a worker session, plus — when the worker
* is paused mid-turn in an async {@code fleet_ask} (CB-582) — the open question and how to
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
* is paused mid-turn in an async {@code fleet_ask} — the open question and how to answer it, so
* a lead on its normal poll cadence does not need the ticket to notice. The question, its
* {@code turnId} and its ticket id are shown only to the caller whose owner key created that
* delegation, or to the unnamed primary; any other caller still sees the base status.
* {@code callerOwner} comes from the calling connection's resolved principal.
*/
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
static McpSchema.CallToolResult status(MessageService messages, String sessionId, String callerOwner) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
String base = messages.status(sessionId).name().toLowerCase();
MessageService.PendingAsk ask = messages.pendingAsk(sessionId);
MessageService.PendingAsk ask = messages.pendingAsk(sessionId, callerOwner);
if (ask == null) {
return text(base);
}
@@ -1295,6 +1405,26 @@ public final class FleetMcp {
}
return text(json(m));
}
if (caller.isCollaborator()) {
// A collaborator's name is its slot in the collaborators: registry; sessionId is its
// pane so a peer knows where to reach it. No leader key: a collaborator is not a
// primary for authorization, unlike a lead.
if (caller.name() != null) {
m.put("collaborator", caller.name());
}
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (caller.isObserver()) {
// No architect slot, collaborator name, or lead name to report — only the pane itself,
// so a peer that already knows this terminal can still address it.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
@@ -1727,7 +1857,8 @@ public final class FleetMcp {
Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, loopHealth, quarantine,
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
OutageSource.none(), LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1736,7 +1867,8 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, CoordinationSource.none(), false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, CoordinationSource.none(), false, true, true);
}
/**
@@ -1753,7 +1885,8 @@ public final class FleetMcp {
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, OutageSource.none(),
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1762,7 +1895,8 @@ public final class FleetMcp {
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm, CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
LeadSeatSource.none(), new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
}
/**
@@ -1784,7 +1918,8 @@ public final class FleetMcp {
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine, outage,
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, false);
leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, false, true, true);
}
/**
@@ -1801,14 +1936,23 @@ public final class FleetMcp {
* {@code false} (fleetd #463: a forgotten argument fails closed, not
* open), so a test that wants the {@code coordinator} row must pass
* an explicit {@code true}
* @param leadsVisible whether this caller may see the {@code leads} array (see
* {@link #leadsVisibleTo}) — unlike {@code callerIsPrimary}, this is
* also {@code true} for an architect or a collaborator, so it cannot
* be derived from {@code callerIsPrimary} alone
* @param membersVisible whether this caller may see the {@code members} array (see
* {@link #membersVisibleTo}); also {@code true} for an architect, but
* unlike {@code leadsVisible}, never for a collaborator
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, LoopHealthSource.none(), quarantine,
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm, coordination, callerIsPrimary);
outage, leadSeats, new LeadContextGauge(), LeadConfigDirSource.none(), leads, selfTerm,
Map.of(), false, coordination, callerIsPrimary, leadsVisible, membersVisible);
}
/**
@@ -1818,6 +1962,17 @@ public final class FleetMcp {
* matter); the one caller that matters for caching, {@code fleet_list}'s MCP handler, passes
* its own single long-lived instance instead (see {@code FleetMcp}'s {@code leadContextGauge}
* field).
*
* @param collaborators terminal_id → collaborator name, live from the resolver
* ({@link dev.ltms.fleet.auth.CallerResolver#collaborators()})
* @param collaboratorsVisible whether this caller may see the {@code collaborators} array (see
* {@link #collaboratorsVisibleTo}); every wrapper overload above
* passes {@code false}, so a test that wants the row must call this
* overload with an explicit {@code true}
* @param leadsVisible whether this caller may see the {@code leads} array (see
* {@link #leadsVisibleTo})
* @param membersVisible whether this caller may see the {@code members} array (see
* {@link #membersVisibleTo})
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
@@ -1826,32 +1981,56 @@ public final class FleetMcp {
LeadSeatSource leadSeats, LeadContextGauge contextGauge,
LeadConfigDirSource leadConfigDirs,
Map<String, String> leads, String selfTerm,
CoordinationSource coordination, boolean callerIsPrimary) {
Map<String, String> collaborators, boolean collaboratorsVisible,
CoordinationSource coordination, boolean callerIsPrimary,
boolean leadsVisible, boolean membersVisible) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
// Neither row's assembly (leadView/memberCapacityView probing herdr for live status)
// runs unless at least one of them needs the live-agent lookup backing it.
Map<String, Agent> live = (leadsVisible || membersVisible)
? workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b))
: Map.of();
// fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
// resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster().
List<MemberSession> roster = sessions.rosterResolved();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
// READ is permission to enter this tool, not permission to receive every field it can
// build -- gate BEFORE assembling each row, so the key is absent rather than
// present-and-empty; a caller without either row gets its own facts from fleet_whoami.
if (leadsVisible) {
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm, contextGauge,
leadConfigDirs))
.toList();
result.put("leads", leadRows);
}
if (membersVisible) {
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
result.put("members", out);
}
result.put("healthCoverage", healthCoverage.value().get());
result.put("loopHealth", Map.of(
"statusPoller", loopHealth.statusPoller().get().name(),
"sessionReaper", loopHealth.sessionReaper().get().name()));
// fleetd #703: a collaborator tab is a person's own tab, so the row is assembled and
// included only for the roles that may SEND to a named peer -- gate BEFORE assembling
// it, same reason as the coordinator row just below: the key must be absent for a
// worker, never present-and-empty.
if (collaboratorsVisible) {
result.put("collaborators", collaborators.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> collaboratorRow(e.getKey(), e.getValue()))
.toList());
}
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
// key is absent rather than present-and-empty.
@@ -2163,6 +2342,20 @@ public final class FleetMcp {
return m;
}
/**
* fleetd #703: one row of the {@code collaborators} array — a named peer's registry name and
* the herdr {@code terminal_id} a {@code fleet_send} must target to reach it. No status, no
* context gauge, no profile: a collaborator is never spawned and carries no profile, so those
* fields have no meaning for it, and the scan behind {@code terminal} only reports a tab it
* actually found in herdr, so a tab nobody has open does not appear here at all.
*/
private static Map<String, Object> collaboratorRow(String terminal, String name) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("name", name);
m.put("sessionId", terminal);
return m;
}
/**
* Reads the lead context gauge for one lead. {@code configDir} is this lead's configured
* {@code CLAUDE_CONFIG_DIR} override (see {@link LeadConfigDirSource}), derived from
@@ -2357,7 +2550,14 @@ public final class FleetMcp {
private static McpSchema.Tool listTool() {
return tool(FleetTool.LIST.wireName(),
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
"List the whole fleet the bridge tracks, in two parts. 'members' is visible to "
+ "the primary and an architect only. 'leads' is visible to those two AND a "
+ "collaborator — exactly the roles that may fleet_send to a lead, so a "
+ "collaborator can learn a lead's sessionId before using the send it already "
+ "holds. A worker holds READ to call this tool at all, but gets neither "
+ "array, never an empty one; a worker reads its own session, "
+ "profile, state, worktree, branch and owner from fleet_whoami instead. "
+ "'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to fleet_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the "
@@ -2388,7 +2588,12 @@ public final class FleetMcp {
+ "msgId/from/preview, never the full body), and one row per coordinator.peers "
+ "coord-id ('peers': coordId/reachable, plus pending/consumers when reachable) — "
+ "this is peer DISCOVERY for cross-host leads, distinct from the local 'leads' "
+ "array above. It is omitted entirely when no coordinator is configured.",
+ "array above. It is omitted entirely when no coordinator is configured. A "
+ "'collaborators' array, visible only to the primary, an architect, and a "
+ "collaborator (never a worker), reports every other named peer tab this daemon "
+ "recognises: each row is 'name' (its registry name) and 'sessionId' (the "
+ "terminal id a fleet_send targets to reach it). Absent entirely for a worker, "
+ "whatever is configured.",
objectSchema(Map.of(), List.of()));
}
@@ -6,6 +6,7 @@ import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
@@ -67,16 +68,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
@@ -766,90 +757,26 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
checkPaneCommandFits(cfg, argv);
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
try {
ResilientAgentLaunch.checkFits(cfg.profile(), argv);
} catch (ResilientAgentLaunch.TooLargeException e) {
throw new PeerUnreachableException(e.getMessage());
}
throw last;
}
/**
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line,
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent
* started", the backend exits on the mangled argument it was handed, the pane closes, and the
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
* would have hit anyway, named at the point it is still
* explainable
*/
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new PeerUnreachableException(
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
long[] lastSeq = {0};
Agent agent = ResilientAgentLaunch.startUniquelyNamed(agents, namePrefix, args, paneId,
attempt -> {
lastSeq[0] = nameSeq.incrementAndGet();
return namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + lastSeq[0];
},
ResilientAgentLaunch.NAME_RETRIES, ResilientAgentLaunch.SHELL_READY_RETRIES, sleeper);
return new Started(agent, lastSeq[0]);
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}.
* fleetd choice and not configurable — see {@link ResilientAgentLaunch#checkFits}.
*/
static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
static final int PANE_COMMAND_BYTE_LIMIT = ResilientAgentLaunch.PANE_COMMAND_BYTE_LIMIT;
// --- discovery + reap ----------------------------------------------------------------------
@@ -1,5 +1,6 @@
package dev.ltms.fleet.msg;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrRouter;
@@ -87,7 +88,7 @@ public final class MessageService {
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
* {@link #answer(String, String, long)} and the turn resumes.
* {@link #answer(String, String, long, String)} and the turn resumes.
*/
QUESTION,
/** Timed out after the message was delivered — the worker is still working. */
@@ -120,10 +121,16 @@ public final class MessageService {
/** Another send to this session was in flight for the whole window. */
BUSY,
/**
* An answer ({@link #answer(String, String, long)}) referenced a {@code turnId} that is no
* longer open — the worker's {@code fleet_ask} already timed out or was answered.
* An answer ({@link #answer(String, String, long, String)}) referenced a {@code turnId}
* that is no longer open — the worker's {@code fleet_ask} already timed out or was answered.
*/
STALE_TURN
STALE_TURN,
/**
* An answer ({@link #answer(String, String, long, String)}) named a {@code turnId} that is
* still open, but the answering caller is not the caller whose accepted delegation opened
* it. Distinct from {@link #STALE_TURN} so a refusal is never reported as a lapsed turn.
*/
NOT_TURN_OWNER
}
/**
@@ -133,7 +140,7 @@ public final class MessageService {
* {@link Outcome#COMPLETED_UNREPLIED}), or the question for {@link Outcome#QUESTION},
* else {@code null}
* @param turnId correlation id for a {@link Outcome#QUESTION} (answered via
* {@link #answer(String, String, long)}), else {@code null}
* {@link #answer(String, String, long, String)}), else {@code null}
*/
public record Reply(Outcome outcome, String text, String turnId) {
/** A reply with no correlation id (the common terminal outcomes). */
@@ -275,10 +282,16 @@ public final class MessageService {
* {@link #abandon}) can never match again regardless of this flag's value.
*/
private volatile boolean askTimedOut;
/**
* The owner key of the caller whose {@code fleet_send{wait:false}} created this ticket, or
* {@code null} for the unnamed primary and overloads that do not record a caller.
*/
private final String creatorOwner;
private Task(String ticket, String target, LongSupplier nowNanos) {
private Task(String ticket, String target, LongSupplier nowNanos, String creatorOwner) {
this.ticket = ticket;
this.target = target;
this.creatorOwner = creatorOwner;
this.createdNanos = nowNanos.getAsLong();
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
}
@@ -340,6 +353,14 @@ public final class MessageService {
*/
private final ConcurrentHashMap<String, Boolean> queuedDeliveries = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
/**
* Minted once per {@code MessageService} instance and folded into every ticket id (see
* {@link #sendAsync(String, String, Runnable, Principal)}). {@link #ticketSeq} alone restarts at
* zero for every instance, so without this a ticket id can be reused across instances and
* resolve to an unrelated {@link Task} with no error; this nonce makes that impossible, because
* an id minted by one instance can never match the id space of another.
*/
private final String ticketBootNonce = UUID.randomUUID().toString().substring(0, 6);
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
@@ -656,7 +677,7 @@ public final class MessageService {
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, TIMED_OUT_UNCONFIRMED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation
case STALE_TURN, QUESTION, NOT_TURN_OWNER -> null; // not a completed delegation
};
}
@@ -915,14 +936,17 @@ public final class MessageService {
/**
* Deliver {@code content} to {@code target} (a herdr {@code terminal_id}) and block until the
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses. {@code callerOwner}
* identifies the caller making this call and is recorded as the turn's owner. It is the only
* caller {@link #answer(String, String, long, String)} will
* later accept an answer from if the worker pauses mid-turn to ask.
*/
public Reply send(String target, String content, long timeoutMillis) {
return send(target, content, timeoutMillis, null);
public Reply send(String target, String content, long timeoutMillis, String callerOwner) {
return send(target, content, timeoutMillis, null, callerOwner);
}
/**
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
* As {@link #send(String, String, long, String)}, but with an accepted-delivery hook.
*
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
@@ -933,12 +957,13 @@ public final class MessageService {
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, String callerOwner) {
return send(target, content, timeoutMillis, onAccepted, null, callerOwner);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task,
String callerOwner) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -958,7 +983,7 @@ public final class MessageService {
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target, Rendezvous.Owner.of(callerOwner));
// CB-640: this send now owns target's delivery, so any earlier stranded-reply or
// still-queued fact no longer describes the live state — clear both rather than let
// them outlive the send that supersedes them.
@@ -1137,19 +1162,30 @@ public final class MessageService {
* mid-turn (already picked up), so the answer flows back through its own open {@code fleet_ask}
* call, not a new status-gated delivery. The forward waiter is opened <em>before</em> the worker
* is unblocked so a reply that lands the instant it resumes is not lost.
*
* <p>{@code callerOwner} identifies the caller making this call. It is checked against the
* turn's recorded owner (the caller whose
* accepted delegation opened it, see {@link #send(String, String, long, String)} and
* {@link #sendAsync(String, String, Runnable, Principal)}) before anything else runs: a mismatch,
* including a turn with no owner on record at all, returns {@link Outcome#NOT_TURN_OWNER}
* without touching the rendezvous, the session lock, or any async task bookkeeping.
*/
public Reply answer(String turnId, String content, long timeoutMillis) {
public Reply answer(String turnId, String content, long timeoutMillis, String callerOwner) {
String workerSession = rendezvous.askSession(turnId);
if (workerSession == null) {
return new Reply(Outcome.STALE_TURN, null); // the ask lapsed (timed out or already answered)
}
Rendezvous.Owner owner = rendezvous.askOwner(turnId);
if (!Rendezvous.Owner.permits(owner, callerOwner)) {
return new Reply(Outcome.NOT_TURN_OWNER, null);
}
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(workerSession, _ -> new ReentrantLock());
if (!tryLock(lock, remainingMillis(deadlineNanos))) {
return new Reply(Outcome.BUSY, null);
}
try {
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession);
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(workerSession, owner);
// fleetd #575: this try used to open below, AFTER the Task lookup/registration and the
// STALE_TURN early return that follows it — so that return was covered only by a
// hand-rolled copy of the finally's own cleanup pair, not the finally itself. Widening the
@@ -1279,8 +1315,20 @@ public final class MessageService {
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(ticket, target, nowNanos);
return sendAsync(target, content, onAccepted, null);
}
/**
* As {@link #sendAsync(String, String, Runnable)}, recording {@code creator}'s owner key as this
* ticket's owner. The key is derived here from the resolved principal so callers cannot pass a
* terminal address where an owner identity is required.
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted, Principal creator) {
String ticket = "task-" + ticketBootNonce + "-" + ticketSeq.incrementAndGet();
String creatorOwner = creator == null ? null : creator.ownerKey();
Task task = new Task(ticket, target, nowNanos, creatorOwner);
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
@@ -1311,7 +1359,7 @@ public final class MessageService {
}
asyncExecutor.submit(() -> {
try {
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task, creatorOwner);
if (result.outcome() == Outcome.QUESTION) {
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
@@ -1343,15 +1391,31 @@ public final class MessageService {
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket;
* otherwise a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
* As {@link #poll(String, String)}, with no caller owner key — the unnamed primary's ticket rule
* checked, so this overload must only be used where the caller's identity is otherwise
* irrelevant.
*/
public TaskView poll(String ticket) {
return poll(ticket, null);
}
/**
* Snapshot the state of an async delegation. Returns {@code null} for an unknown/expired ticket.
* Refuses a {@code callerOwner} that differs from the owner that created the ticket (see
* {@link #sendAsync(String, String, Runnable, Principal)}) with a {@link Phase#FAILED} view that
* carries no reply text. The unnamed primary has a {@code null} owner key and is never refused.
* Otherwise returns a {@link Phase#PENDING} view (with the live worker status as detail), a
* {@link Phase#DONE} view carrying the reply, or a {@link Phase#FAILED} view with the reason.
*/
public TaskView poll(String ticket, String callerOwner) {
Task task = tasks.get(ticket);
if (task == null) {
return null;
}
if (!ownsTicket(task, callerOwner)) {
return new TaskView(ticket, Phase.FAILED, null, null,
"forbidden: this ticket was created by a different session", null);
}
CompletableFuture<Reply> f = task.future;
if (!f.isDone()) {
Reply question = task.question;
@@ -1387,6 +1451,16 @@ public final class MessageService {
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/**
* Whether {@code callerOwner} may read {@code task}'s state. A {@code null} caller key is the
* unnamed primary and may read every ticket. Other callers must match the task's owner key. This
* differs from {@link Rendezvous.Owner#permits}: a missing rendezvous owner is not an authenticated
* unnamed primary, so that gate refuses every caller when no owner was recorded.
*/
private static boolean ownsTicket(Task task, String callerOwner) {
return callerOwner == null || callerOwner.equals(task.creatorOwner);
}
/**
* Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
@@ -1706,16 +1780,17 @@ public final class MessageService {
}
/**
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any
* (CB-582) — {@code fleet_status} uses this to show a pending question without the caller
* needing the ticket. {@code null} when the session has no open async question (including a
* session mid a <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}).
* The question {@code workerSession} is currently paused on via {@code fleet_ask}, if any —
* {@code fleet_status} uses this to show a pending question without the caller needing the
* ticket. {@code null} when the session has no open async question (including a session mid a
* <em>blocking</em> {@code fleet_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}), or when {@code callerOwner} does not own the task the question
* belongs to (see {@link #ownsTicket(Task, String)}).
*/
public PendingAsk pendingAsk(String workerSession) {
public PendingAsk pendingAsk(String workerSession, String callerOwner) {
for (Task task : tasks.values()) {
Reply q = task.question;
if (q != null && workerSession.equals(task.target)) {
if (q != null && workerSession.equals(task.target) && ownsTicket(task, callerOwner)) {
return new PendingAsk(task.ticket, q.text(), q.turnId());
}
}
@@ -1,5 +1,6 @@
package dev.ltms.fleet.msg;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;
@@ -60,7 +61,7 @@ public final class Rendezvous {
}
/** A worker's open mid-turn question: the worker session it belongs to and the answer future. */
private record AskWaiter(String session, CompletableFuture<String> answer) {
private record AskWaiter(String session, CompletableFuture<String> answer, Owner owner) {
}
/**
@@ -70,11 +71,44 @@ public final class Rendezvous {
public record AskTicket(String turnId, CompletableFuture<String> answer, boolean fresh) {
}
private final ConcurrentHashMap<String, CompletableFuture<Resolution>> waiters = new ConcurrentHashMap<>();
/**
* The caller whose accepted delegation opened a turn — the only caller allowed to answer it.
* A {@code null} owner key means the unnamed primary.
*/
public record Owner(String ownerKey) {
public static final Owner UNNAMED_PRIMARY = new Owner(null);
public static Owner of(String ownerKey) {
return ownerKey == null ? UNNAMED_PRIMARY : new Owner(ownerKey);
}
/**
* Whether {@code callerOwner} matches {@code owner}. A {@code null} owner means no owner was
* recorded, so it matches no caller. {@link #UNNAMED_PRIMARY} records the unnamed primary
* with an owner object whose key is {@code null}.
*/
public static boolean permits(Owner owner, String callerOwner) {
return owner != null && java.util.Objects.equals(owner.ownerKey(), callerOwner);
}
}
/** A registered forward waiter together with the owner its delegation was opened under. */
private record ForwardWaiter(Owner owner, CompletableFuture<Resolution> future) {
}
private final ConcurrentHashMap<String, ForwardWaiter> waiters = new ConcurrentHashMap<>();
/** Reverse rendezvous (CB-205): worker questions awaiting the primary's answer, keyed by {@code turnId}. */
private final ConcurrentHashMap<String, AskWaiter> asks = new ConcurrentHashMap<>();
private final AtomicLong askSeq = new AtomicLong();
/**
* Minted once per {@code Rendezvous} instance and folded into every {@code turnId} (see
* {@link #openAsk(String)}). {@link #askSeq} alone restarts at zero for every instance, so
* without this a {@code turnId} minted by one instance could be minted again by another and
* resolve to an unrelated ask with no error; this nonce makes that impossible, because an id
* minted by one instance can never match the id space of another.
*/
private final String askBootNonce = UUID.randomUUID().toString().substring(0, 6);
/** Per-session index of the currently-open ask, so duplicate fleet_ask calls coalesce onto one turn. */
private final ConcurrentHashMap<String, String> openAsksBySession = new ConcurrentHashMap<>();
@@ -89,8 +123,17 @@ public final class Rendezvous {
* code a double open is impossible; this is a tripwire for the day that no longer holds.
*/
public CompletableFuture<Resolution> open(String session) {
return open(session, null);
}
/**
* Same as {@link #open(String)}, additionally recording {@code owner} as the caller whose
* delegation opened this waiter. A {@code null} owner records no owner at all — the
* fail-closed default {@link Owner#permits} refuses to everyone.
*/
public CompletableFuture<Resolution> open(String session, Owner owner) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
ForwardWaiter existing = waiters.putIfAbsent(session, new ForwardWaiter(owner, waiter));
if (existing != null) {
throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered");
@@ -105,7 +148,13 @@ public final class Rendezvous {
* successful {@code open} after a finished turn requires this close to have happened first).
*/
public void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
waiters.computeIfPresent(session, (s, w) -> w.future() == waiter ? null : w);
}
/** The owner recorded for {@code session}'s open waiter, or {@code null} if none is open. */
public Owner ownerOf(String session) {
ForwardWaiter w = waiters.get(session);
return w == null ? null : w.owner();
}
/** Whether a send is currently awaiting a resolution for {@code session}. */
@@ -119,7 +168,8 @@ public final class Rendezvous {
* send (see the CB-116 note above) rather than whichever send happens to be waiting when they fire.
*/
public CompletableFuture<Resolution> currentWaiter(String session) {
return waiters.get(session);
ForwardWaiter w = waiters.get(session);
return w == null ? null : w.future();
}
/**
@@ -145,9 +195,9 @@ public final class Rendezvous {
while (true) {
AskWaiter[] minted = { null };
String turnId = openAsksBySession.computeIfAbsent(session, _ -> {
String newTurnId = session + "#" + askSeq.incrementAndGet();
String newTurnId = session + "#" + askBootNonce + "-" + askSeq.incrementAndGet();
CompletableFuture<String> answer = new CompletableFuture<>();
AskWaiter waiter = new AskWaiter(session, answer);
AskWaiter waiter = new AskWaiter(session, answer, ownerOf(session));
asks.put(newTurnId, waiter);
minted[0] = waiter;
return newTurnId;
@@ -183,6 +233,16 @@ public final class Rendezvous {
return w == null ? null : w.session();
}
/**
* The owner recorded for {@code turnId} when its ask turn was freshly opened — the caller
* whose delegation {@link #answerAsk} must match. {@code null} if {@code turnId} is unknown or
* lapsed, or if the ask opened with no forward waiter owner on record.
*/
public Owner askOwner(String turnId) {
AskWaiter w = asks.get(turnId);
return w == null ? null : w.owner();
}
/**
* Resolve a worker's blocked {@code fleet_ask} with the primary's {@code answer}, unblocking it
* to resume its turn.
@@ -245,7 +305,7 @@ public final class Rendezvous {
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
ForwardWaiter waiter = waiters.get(session);
return waiter != null && waiter.future().complete(resolution);
}
}
@@ -49,22 +49,52 @@ import java.util.stream.Collectors;
*/
public final class FleetApp {
/** The authorization action the matching route handler hands to {@link #allow}. */
/**
* The authorization action the matching route handler hands to {@link #allow}, for a route
* whose action does not depend on the request body.
*/
static Authz.Action routeAction(String route) {
return routeAction(route, null);
}
/**
* As above, plus the one route whose action depends on the body: {@code POST
* /sessions/{id}/message} carries a {@code turnId} (the answer-a-blocked-worker shape) or not
* (a plain delivery), mirroring {@code FleetMcp#sendAction}'s split of the same two call
* shapes over MCP. {@code turnId} is ignored by every other route.
*
* @param turnId the request body's {@code turnId}, or {@code null}/blank when absent or not
* applicable to this route
*/
static Authz.Action routeAction(String route, String turnId) {
return switch (route) {
case "GET /metrics" -> Authz.Action.METRICS;
case "POST /members" -> Authz.Action.SPAWN;
case "DELETE /members/{paneId}" -> Authz.Action.STOP;
case "POST /sessions/{id}/message" -> Authz.Action.SEND;
case "POST /sessions/{id}/message" -> turnId == null || turnId.isBlank()
? Authz.Action.SEND : Authz.Action.ANSWER;
case "POST /sessions/{id}/reply" -> Authz.Action.REPLY;
case "GET /sessions/{id}/replies" -> Authz.Action.DRAIN;
case "POST /sessions/{id}/ask" -> Authz.Action.ASK;
case "GET /sessions", "GET /agents", "GET /members", "GET /profiles",
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.READ;
"GET /member-credentials" -> Authz.Action.READ;
case "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.TASK_READ;
default -> throw new IllegalArgumentException("route has no authorization gate: " + route);
};
}
/**
* The second gate for {@code POST /sessions/{id}/message}: checked only when {@code turnId}
* is present and non-blank, against {@link Authz.Action#ANSWER}. A request with no {@code
* turnId} passes this gate unconditionally, without consulting {@code permit} at all, having
* already cleared the coarse {@link Authz.Action#SEND} grant checked ahead of it.
*
* @param permit reports whether the caller holds the named grant
*/
static boolean answerGatePasses(String turnId, Predicate<Authz.Action> permit) {
return turnId == null || turnId.isBlank() || permit.test(Authz.Action.ANSWER);
}
/** Default blocking window for a message; kept under typical HTTP idle timeouts. */
private static final long DEFAULT_MESSAGE_TIMEOUT_MS = 25_000;
private static final long MAX_MESSAGE_TIMEOUT_MS = 120_000;
@@ -236,6 +266,21 @@ public final class FleetApp {
return app;
}
/**
* The authorization decision behind {@link #allow}, taking the caller directly rather than
* pulling it from a servlet {@link Context} — unit-testable without fabricating a live
* request, the same reason {@code FleetMcp#denyFor} is split from {@code FleetMcp#deny}.
*
* @param knownLeadOrCollaborator the classifier a collaborator's {@code SEND} is checked
* against; pass {@link #auth}'s own {@code
* knownLeadOrCollaborator()} to exercise the real production
* gate, as {@link #allow} does
*/
static boolean permitsFor(Principal caller, Authz.Action action, String target,
Predicate<String> knownLeadOrCollaborator) {
return Authz.permits(caller, action, target, knownLeadOrCollaborator);
}
/**
* Gate a handler on the CB-505 authorization table. Returns {@code true} when the request may
* proceed; otherwise writes the error response and returns {@code false}.
@@ -249,8 +294,9 @@ public final class FleetApp {
return true; // legacy: authorization not enforced
}
Principal caller = ctx.attribute(CALLER);
if (Authz.permits(caller, action, target)) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
if (permitsFor(caller, action, target, auth.knownLeadOrCollaborator())) {
if (action != Authz.Action.READ && action != Authz.Action.METRICS
&& action != Authz.Action.TASK_READ) {
AuditLog.allowed(caller, action, target); // reads would drown the trail
}
return true;
@@ -603,26 +649,39 @@ public final class FleetApp {
* status-gated injector and block until the worker returns a structured {@code fleet_reply}.
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
* still land.
*
* <p>Two call shapes share this route, exactly as {@code fleet_send} does over MCP (see
* {@code FleetMcp#sendAction}): a plain delivery to {@code id}, and -- when the body carries
* {@code turnId} -- resolving a worker's blocked question. The coarse {@link
* Authz.Action#SEND} grant is checked first, before the body is read at all; only once that
* passes is the body parsed, and a present {@code turnId} is then checked again against
* {@link Authz.Action#ANSWER}. A body that fails to parse is rejected with 400 and reaches
* neither {@code messages.answer} nor {@code messages.send}.
*/
private void sendMessage(Context ctx) {
String id = ctx.pathParam("id");
if (!allow(ctx, routeAction("POST /sessions/{id}/message"), id)) {
return;
}
String content;
String turnId;
long timeout;
boolean wait;
Principal caller = ctx.attribute(CALLER);
String callerOwner = caller == null ? null : caller.ownerKey();
JsonNode body;
try {
JsonNode body = mapper.readTree(ctx.body());
content = body.path("content").asText("");
turnId = body.path("turnId").asText(null);
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
body = mapper.readTree(ctx.body());
} catch (Exception e) {
body = null;
}
if (body == null) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
return;
}
String turnId = body.path("turnId").asText(null);
if (!answerGatePasses(turnId, action -> allow(ctx, action, id))) {
return;
}
String content = body.path("content").asText("");
long timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
boolean wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
if (content.isBlank()) {
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
return;
@@ -631,19 +690,19 @@ public final class FleetApp {
// Answering a worker's fleet_ask (CB-205): always blocks, and derives the worker from turnId.
if (turnId != null && !turnId.isBlank()) {
writeReply(ctx, id, messages.answer(turnId, content, timeout), timeout);
writeReply(ctx, id, messages.answer(turnId, content, timeout, callerOwner), timeout);
return;
}
if (!wait) {
// Fire-and-poll (CB-107): return a ticket immediately; the caller polls GET /tasks/{ticket}.
String ticket = messages.sendAsync(id, content);
String ticket = messages.sendAsync(id, content, null, caller);
ctx.status(202).json(Map.of("sessionId", id, "ticket", ticket, "status", "accepted"));
return;
}
try {
writeReply(ctx, id, messages.send(id, content, timeout), timeout);
writeReply(ctx, id, messages.send(id, content, timeout, callerOwner), timeout);
} catch (HerdrException e) {
herdrError(ctx, e);
}
@@ -662,6 +721,10 @@ public final class FleetApp {
case STALE_TURN -> ctx.status(409).json(Map.of(
"sessionId", id, "error", "stale_turn",
"detail", "that question is no longer open (timed out or already answered)"));
case NOT_TURN_OWNER -> ctx.status(403).json(Map.of(
"sessionId", id, "error", "not_turn_owner",
"detail", "this turn belongs to a different delegation — only the caller that "
+ "opened it may answer it"));
case REPLIED, COMPLETED_UNREPLIED -> {
// replySource distinguishes a structured fleet_reply from the CB-106 completion
// fallback (a scrape of the worker's transcript when it finished without replying).
@@ -676,9 +739,10 @@ public final class FleetApp {
// silent fall-through. That is exactly the bug this ticket exists to fix:
// `default -> "done"` used to sit here and would have told a REST caller the
// delegation completed for TIMED_OUT_UNCONFIRMED, the one outcome where delivery
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION and STALE_TURN can never
// actually reach this inner switch — the outer switch above always dispatches
// them first — but they still need an arm to keep this switch exhaustive.
// is unknown. REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN and
// NOT_TURN_OWNER can never actually reach this inner switch — the outer switch
// above always dispatches them first — but they still need an arm to keep this
// switch exhaustive.
"status", switch (reply.outcome()) {
case TIMED_OUT_WORKING -> "working";
case TIMED_OUT_QUEUED -> "queued";
@@ -688,7 +752,7 @@ public final class FleetApp {
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN -> "done"; // unreachable
case REPLIED, COMPLETED_UNREPLIED, QUESTION, STALE_TURN, NOT_TURN_OWNER -> "done"; // unreachable
},
"detail", reply.outcome() == MessageService.Outcome.TIMED_OUT_UNCONFIRMED
? "no reply within " + timeout + "ms; delivery is unconfirmed — the "
@@ -817,10 +881,12 @@ public final class FleetApp {
body.put("sessionId", id);
body.put("status", messages.status(id).name().toLowerCase());
body.put("ready", deliverable.test(id));
// CB-582: a worker paused mid-turn in an async fleet_ask is otherwise invisible to a
// status poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view.
MessageService.PendingAsk ask = messages.pendingAsk(id);
// A worker paused mid-turn in an async fleet_ask is otherwise invisible to a status
// poll — surface the open question and how to answer it, same as fleet_poll's
// Phase.ASKING view, but only to the caller whose owner key created that delegation, or
// to the unnamed primary.
Principal caller = ctx.attribute(CALLER);
MessageService.PendingAsk ask = messages.pendingAsk(id, caller == null ? null : caller.ownerKey());
if (ask != null) {
body.put("question", ask.question());
body.put("turnId", ask.turnId());
@@ -837,7 +903,8 @@ public final class FleetApp {
if (!allow(ctx, routeAction("GET /tasks/{ticket}"), null)) {
return;
}
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"));
Principal caller = ctx.attribute(CALLER);
MessageService.TaskView v = messages.poll(ctx.pathParam("ticket"), caller == null ? null : caller.ownerKey());
if (v == null) {
ctx.status(404).json(Map.of("error", "unknown_ticket", "detail", "no such task (or it has expired)"));
return;
@@ -26,6 +26,7 @@ import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* Authoritative in-daemon registry of the worker sessions this {@code fleetd} process spawned.
@@ -57,6 +58,18 @@ public final class SessionManager implements TurnListener {
* Populated on every spawn path, removed on {@link #release}.
*/
private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>();
/**
* fleetd #702: a pane mid-teardown, keyed by paneId, held from just before its registry entry
* is removed until {@link #releaseRemoved} finishes. {@link #spawnedMemberRole} consults this
* alongside the registry, so a caller resolving the pane's terminal during that window still
* sees a live member and never falls through to a tab map.
*
* <p>Depth-counted rather than a plain set: two threads can be tearing down the same pane at
* once (the CAS in {@link #releaseIfCurrent} exists for exactly that race), and with a set the
* loser's {@code finally} would unmark the pane while the winner is still mid-teardown,
* reopening the window this exists to close.
*/
private final ConcurrentHashMap<String /*paneId*/, Releasing> releasing = new ConcurrentHashMap<>();
private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
@@ -242,6 +255,10 @@ public final class SessionManager implements TurnListener {
handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0,
MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId());
registry.put(handle.id(), session);
// A presence contact that already arrived for this terminal found no registry
// entry to transition and gave up silently. Retry it now that one exists; remove
// this call and such a session stays in SPAWNING even though it is present.
reconcilePresence(handle.terminalId());
handles.put(handle.id(), handle);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
@@ -303,9 +320,11 @@ public final class SessionManager implements TurnListener {
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/
private MemberSession release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
return releaseWindow(paneId, registry.get(paneId), () -> {
MemberSession removed = registry.remove(paneId);
releaseRemoved(paneId, removed, handles.remove(paneId), cause);
return removed;
});
}
/**
@@ -314,16 +333,85 @@ public final class SessionManager implements TurnListener {
* DONE record from stopping a worker that delivery has made BUSY.
*/
private boolean releaseIfCurrent(MemberSession expected, ReleaseCause cause) {
if (!registry.remove(expected.paneId(), expected)) {
// A lifecycle transition replaced the record between the caller's check and this remove.
// Log it: this race is by definition unobservable otherwise, and a reaper that silently
// declines to reap is the hardest kind of behaviour to diagnose after the fact.
log.debug("skipping reap of pane={}: its registry record changed after the idle check "
+ "(most likely a delivery made it BUSY)", expected.paneId());
return false;
return releaseWindow(expected.paneId(), expected, () -> {
if (!registry.remove(expected.paneId(), expected)) {
// A lifecycle transition replaced the record between the caller's check and this
// remove. Log it: this race is by definition unobservable otherwise, and a reaper
// that silently declines to reap is the hardest kind of behaviour to diagnose
// after the fact.
log.debug("skipping reap of pane={}: its registry record changed after the idle "
+ "check (most likely a delivery made it BUSY)", expected.paneId());
return false;
}
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
return true;
});
}
/**
* fleetd #702: mark {@code paneId} as mid-teardown — using {@code known}'s terminal/role when
* it is available — for the whole of {@code teardown}, which removes the registry entry and
* then runs {@link #releaseRemoved}. Shared by both registry-removal sites ({@link #release}'s
* unconditional remove and {@link #releaseIfCurrent}'s CAS remove) so neither can leave the
* other's window unmarked.
*
* <p>The mark is written before {@code teardown} runs — so it covers the removal itself, not
* only what comes after it — and cleared in a {@code finally}, so an unchecked throw out of
* {@code teardown} (including one from {@link PeerLauncher#stop}, which declares nothing) can
* never leave the pane marked for the rest of the daemon's life.
*/
private <T> T releaseWindow(String paneId, MemberSession known, Supplier<T> teardown) {
releasing.compute(paneId, (_, prior) -> Releasing.enter(prior, known));
try {
return teardown.get();
} finally {
releasing.compute(paneId, (_, prior) -> prior == null ? null : prior.leave());
}
releaseRemoved(expected.paneId(), expected, handles.remove(expected.paneId()), cause);
return true;
}
/**
* Depth count plus the terminal/role a mid-teardown pane belongs to, for
* {@link #spawnedMemberRole}. The terminal/role come from whichever call into
* {@link #releaseWindow} first knew them: a call that finds the registry entry already gone
* passes a {@code null} session, and must not blank out what the first call recorded.
*/
record Releasing(int depth, String terminalId, MemberRole role) {
static Releasing enter(Releasing prior, MemberSession known) {
int depth = (prior == null ? 0 : prior.depth()) + 1;
String terminalId = known != null ? known.terminalId() : prior == null ? null : prior.terminalId();
MemberRole role = known != null ? known.role() : prior == null ? null : prior.role();
return new Releasing(depth, terminalId, role);
}
Releasing leave() {
return depth <= 1 ? null : new Releasing(depth - 1, terminalId, role);
}
}
/**
* The role of the live spawned member occupying {@code terminal} — whether it is currently in
* the registry, or mid-teardown between {@link #release} removing its registry entry and
* {@link #releaseRemoved} actually stopping its pane (fleetd #702). {@code null} for a terminal
* that is neither: this method is the one reader a caller resolver consults before any tab
* map, so a live or releasing member's identity never falls back to a tab label.
*
* <p>Checks the registry directly via {@link #findByTerminal} rather than {@link #roster()},
* so this hot-path lookup (consulted on every resolve) never pays for a list copy or a stream.
*/
public MemberRole spawnedMemberRole(String terminal) {
MemberSession session = findByTerminal(terminal);
if (session != null) {
return session.role();
}
if (terminal == null) {
return null;
}
for (Releasing r : releasing.values()) {
if (terminal.equals(r.terminalId())) {
return r.role();
}
}
return null;
}
private void releaseRemoved(String paneId, MemberSession removed, PeerHandle removedHandle,
@@ -382,6 +470,12 @@ public final class SessionManager implements TurnListener {
MemberSession resolved = resolveAgentSessionId(removed, removedHandle);
notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(),
resolved.branch(), snapshotRef, resolved.agentSessionId()));
String terminal = removed.terminalId();
if (terminal != null && !terminal.isBlank()) {
// Without this, a terminal stays marked present after its pane is gone, so a
// later send to the same id would read as deliverable instead of refused.
presence.forget(terminal);
}
}
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
@@ -725,6 +819,10 @@ public final class SessionManager implements TurnListener {
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
// A presence contact that already arrived for this terminal found no registry entry to
// transition and gave up silently. Retry it now that one exists; remove this call and
// such a session stays in SPAWNING even though it is present.
reconcilePresence(handle.terminalId());
handles.put(handle.id(), handle);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
@@ -899,6 +997,20 @@ public final class SessionManager implements TurnListener {
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
}
/**
* Completes a newly registered session's {@code SPAWNING -> READY} transition when {@code
* terminalId} was already marked present before this ran. A terminal never marked present is
* left in {@code SPAWNING}; it reaches {@code READY} normally through {@link #onReady} once
* its own contact arrives. Callers must run this only once the session's registry entry is
* already visible — {@link #onReady}'s transition matches against that entry, and reconciling
* before the entry exists finds nothing to transition.
*/
private void reconcilePresence(String terminalId) {
if (terminalId != null && !terminalId.isBlank() && presence.isPresent(terminalId)) {
onReady(terminalId);
}
}
/**
* Lifecycle hook: a message was delivered into the worker — it is now busy on a turn.
* The turn count is bumped and the activity timestamp is refreshed. A {@code DONE} session
@@ -14,10 +14,12 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-534: the injector's readiness gate must open for a lead as well as for a present worker.
* fleetd #669 follow-up: the same gate must also open for a collaborator, which — like a lead —
* is never enrolled in {@link MemberPresence} and never discovered by the lead scan.
*
* <p>The bug these cover was silent and slow: a lead was never marked present (only workers are), so
* every lead→lead delivery sat on the gate for the full readiness grace and failed ~60s later without
* a keystroke ever reaching the pane.
* <p>The bug these cover was silent and slow: a lead (and later a collaborator) was never marked
* present (only workers are) and never counted as a lead, so every send to one sat on the gate for
* the full readiness grace and failed ~60s later without a keystroke ever reaching the pane.
*/
class FleetDeliverabilityTest {
@@ -25,19 +27,25 @@ class FleetDeliverabilityTest {
return () -> m;
}
private static Supplier<Map<String, String>> collaborators(Map<String, String> m) {
return () -> m;
}
@Test
@DisplayName("a worker that has connected its MCP is deliverable")
void presentWorkerIsDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertTrue(Fleetd.deliverableTo(presence, leads(Map.of())).test("term_worker"));
assertTrue(Fleetd.deliverableTo(presence, leads(Map.of()), collaborators(Map.of()))
.test("term_worker"));
}
@Test
@DisplayName("a worker still in its boot window is held back")
void absentWorkerIsNotDeliverable() {
assertFalse(Fleetd.deliverableTo(new MemberPresence(), leads(Map.of())).test("term_booting"));
assertFalse(Fleetd.deliverableTo(new MemberPresence(), leads(Map.of()), collaborators(Map.of()))
.test("term_booting"));
}
@Test
@@ -45,19 +53,31 @@ class FleetDeliverabilityTest {
void leadIsDeliverableWithoutPresence() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")), collaborators(Map.of()));
assertFalse(presence.isPresent("term_lead"), "a lead is never enrolled in worker presence");
assertTrue(deliverable.test("term_lead"), "…and must be deliverable anyway");
}
@Test
@DisplayName("an unknown terminal is deliverable to neither")
@DisplayName("a collaborator is deliverable without ever being marked present or scanned as a lead")
void collaboratorIsDeliverableWithoutPresenceOrLeadStatus() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads(Map.of()),
collaborators(Map.of("term_collab", "kevin")));
assertFalse(presence.isPresent("term_collab"), "a collaborator is never enrolled in worker presence");
assertTrue(deliverable.test("term_collab"), "…and must be deliverable anyway");
}
@Test
@DisplayName("an unknown terminal is deliverable to none of presence, leads, or collaborators")
void strangerIsNotDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertFalse(Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
assertFalse(Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")),
collaborators(Map.of("term_collab", "kevin")))
.test("term_stranger"));
}
@@ -65,22 +85,47 @@ class FleetDeliverabilityTest {
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
void leadSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable = Fleetd.deliverableTo(new MemberPresence(), leads(discovered));
Predicate<String> deliverable =
Fleetd.deliverableTo(new MemberPresence(), leads(discovered), collaborators(Map.of()));
assertFalse(deliverable.test("term_late"));
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
assertTrue(deliverable.test("term_late"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("a collaborator discovered after startup becomes deliverable with no restart")
void collaboratorSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable =
Fleetd.deliverableTo(new MemberPresence(), leads(Map.of()), collaborators(discovered));
assertFalse(deliverable.test("term_late_collab"));
discovered.put("term_late_collab", "kevin"); // the same tab scan picks up a newly labelled collaborator tab
assertTrue(deliverable.test("term_late_collab"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
void forgetDoesNotDisarmALead() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
Fleetd.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")), collaborators(Map.of()));
presence.forget("term_lead"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_lead"));
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a collaborator of its deliverability")
void forgetDoesNotDisarmACollaborator() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads(Map.of()),
collaborators(Map.of("term_collab", "kevin")));
presence.forget("term_collab"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_collab"));
}
}
@@ -0,0 +1,165 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertSame;
/**
* fleetd #669 Unit E. Reaches the real {@link dev.ltms.fleet.herdr.HerdrRouter} that {@link
* FleetdAssembly#assembleAndStart} builds and wires — not a copy built for this test — and proves
* that a configured collaborator's terminal routes to the LEAD herdr daemon.
*
* <p>Two distinct {@link FakeHerdr} instances are required, the same pattern {@code
* FleetdAssemblyConnectionIdentityTest} and {@code FleetdLeadRolloverAssemblyTest} already use:
* with one client shared between {@code herdrSocket} and {@code memberHerdrSocket},
* {@code HerdrRouter} folds {@code leadAgents} and {@code memberAgents} into the same instance
* (see its constructor), and {@code agentsFor} would return that one object regardless of whether
* the collaborator map was ever consulted — invisible to a mutation of the predicate this ticket
* fixes. This test's two sockets resolve to two different fakes, so the assertion only passes when
* the collaborator's terminal is actually recognised and routed to the lead one.
*/
class FleetdAssemblyCollaboratorHerdrRoutingTest {
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
private static final class RecordingResourcePorts implements ResourcePorts {
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
HerdrClient client = herdrsBySocket.get(socketPath);
if (client == null) {
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
}
return client;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: "%s"
memberHerdrSocket: "%s"
idleSleepGuard:
enabled: false
fleet:
collaborators:
reviewer-alex:
tab: "collab: alex"
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
return FleetConfig.load(file);
}
@Test
void assembledRouterRoutesACollaboratorTerminalToTheLeadDaemon(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
RecordingResourcePorts ports = new RecordingResourcePorts();
// The fixed FakeHerdr fixture already ties terminal "term_a" to a live agent on tab
// "w2:t7" (pane "w2:p7") — seeding only the tab LABEL to match the configured collaborator
// is enough to make LeadTabScanner resolve "term_a" as that collaborator. Seeded on the
// LEAD fake only: a collaborator's pane lives in the lead daemon, exactly like a lead's.
FakeHerdr lead = new FakeHerdr().withTab("w2", "w2:t7", "collab: alex");
FakeHerdr member = new FakeHerdr();
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
assertSame(runtime.router().leadAgents(), runtime.router().agentsFor("term_a"),
"a configured collaborator's terminal must route to the LEAD daemon — "
+ "FleetdAssembly must wire the collaborator map into the router's "
+ "predicate, not just LeadTabScanner.get()");
assertSame(runtime.router().memberAgents(), runtime.router().agentsFor("term_shell"),
"control: a terminal naming neither a lead nor a collaborator (term_shell, on "
+ "the unlabelled tab w2:t8) must still route to the member daemon");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
}
}
}
@@ -31,8 +31,7 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
* production boot path passes to {@link LeadTabScanner} at {@code FleetdAssembly.java:265}
* ({@code Set.of()}).
* production boot path passes to {@link LeadTabScanner} ({@code Set.of()}).
*
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
@@ -191,7 +190,7 @@ class FleetdAssemblyLeadTabScannerExclusionTest {
Set<?> excluded = (Set<?>) excludedField.get(leads);
assertTrue(excluded.isEmpty(),
"FleetdAssembly.java:265 must pass an empty excludedWorkspaceLabels to "
"FleetdAssembly must pass an empty excludedWorkspaceLabels to "
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
@@ -0,0 +1,190 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.inject.StatusPoller;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import dev.ltms.fleet.msg.ReplyPushLoop;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotNull;
/**
* Asserts that the assembled loops use the production reminder, coordination, and delivery timing
* defaults when no {@code primary:} block configures the reply-push values.
*/
class FleetdAssemblyTimingDefaultsTest {
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private TestResourcePorts ports;
@AfterEach
void tearDown() {
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
""");
return FleetConfig.load(file);
}
private FleetdRuntime assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new TestResourcePorts();
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
}
private static long longField(Object target, String name) throws Exception {
Field field = target.getClass().getDeclaredField(name);
field.setAccessible(true);
return field.getLong(target);
}
@Test
void productionBootPathUsesTheExpectedLoopTimingDefaults(@TempDir Path dir) throws Exception {
FleetdRuntime runtime = assemble(dir);
ReplyPushLoop pushLoop = runtime.pushLoop();
assertEquals(5, longField(pushLoop, "maxReminders"),
"without primary:, ReplyPushLoop must stop after five reminder attempts");
assertEquals(15_000L, longField(pushLoop, "backoffMs"),
"without primary:, ReplyPushLoop must wait fifteen seconds before the next reminder");
LeadCoordLoop leadCoordLoop = runtime.leadCoordLoop();
assertNotNull(leadCoordLoop, "control: coordinator: must build LeadCoordLoop");
assertEquals(3_000L, longField(leadCoordLoop, "intervalMs"),
"LeadCoordLoop must poll for peer-lead mail every three seconds");
StatusPoller poller = runtime.poller();
assertEquals(Injector.POLL_INTERVAL_MILLIS, longField(poller, "intervalMillis"),
"StatusPoller must use Injector's delivery poll interval");
}
}
@@ -14,6 +14,8 @@ class AuthzTest {
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
private static final Principal OBSERVER = Principal.observer("term_observer", 700);
@Test
void anonymousIsAuthorizedForNothing() {
@@ -29,6 +31,22 @@ class AuthzTest {
assertTrue(Authz.isUnauthenticated(null));
}
/**
* {@code MessageService.answer}'s turn-ownership check treats a caller with no terminal as
* matching a turn recorded for the unnamed primary. That rule only stays safe because an
* unauthenticated caller — whose terminal is also {@code null} — never reaches {@code answer}
* at all: {@link #anonymousIsAuthorizedForNothing} already covers every action including
* {@code ANSWER}, but this test names the exact coupling so a future change to either side
* cannot drift without turning this test red.
*/
@Test
void anAnonymousCallerIsRefusedAnswerSoItCanNeverBeMistakenForTheUnnamedPrimary() {
assertFalse(Authz.permits(ANON, ANSWER, null),
"an anonymous caller, whose terminal is also null, must never reach answer() — the "
+ "turn-ownership check's null-terminal match for the unnamed primary owner "
+ "relies on this gate refusing it first");
}
@Test
void orchestrationBelongsToThePrimaryAlone() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, SEND, DRAIN}) {
@@ -38,6 +56,37 @@ class AuthzTest {
}
}
/**
* {@code fleet_send} is three call shapes behind one action name until {@code
* FleetMcp#sendAction} picks one: a plain local {@link Authz.Action#SEND}, the {@code coordId}
* route ({@link Authz.Action#COORD_SEND}), and the {@code turnId} answer form ({@link
* Authz.Action#ANSWER}). All three carry the same grant as the undivided action did — a worker
* is excluded from every one, exactly as it was excluded from the one combined action before.
*/
@Test
void theThreeSendShapesCarryTheSameGrantAsTheOldUndividedAction() {
for (Authz.Action a : new Authz.Action[]{SEND, COORD_SEND, ANSWER}) {
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary may " + a);
assertTrue(Authz.permits(ARCH_DESIGN, a, "term_a"), "an architect may " + a);
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
"a worker performing " + a + " would be escalating into the orchestrator role");
assertFalse(Authz.permits(ANON, a, "term_a"));
}
}
/**
* {@code fleet_poll{ticket}} and {@code fleet_status} are {@link Authz.Action#TASK_READ}, split
* out of the roster-only {@link Authz.Action#READ} (fleetd #678). The grant is unchanged from
* what the undivided {@code READ} action gave every one of these callers.
*/
@Test
void taskReadCarriesTheSameGrantReadDidBeforeTheSplit() {
assertTrue(Authz.permits(PRIMARY, TASK_READ, null));
assertTrue(Authz.permits(WORKER_A, TASK_READ, null));
assertTrue(Authz.permits(ARCH_DESIGN, TASK_READ, null));
assertFalse(Authz.permits(ANON, TASK_READ, null));
}
@Test
void aWorkerMayReplyAndAskOnlyAsItself() {
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
@@ -135,4 +184,132 @@ class AuthzTest {
assertFalse(Authz.isUnauthenticated(WORKER_A));
assertFalse(Authz.isUnauthenticated(PRIMARY));
}
// ── the collaborator matrix ─────────────────────────────────────────────────────────────────
/**
* {@code SEND} for a collaborator is the one grant that is conditional rather than fixed:
* flipping only the classifier's answer for the target flips only this outcome.
*/
@Test
void aCollaboratorMaySendOnlyWhenTheClassifierAcceptsTheTarget() {
assertTrue(Authz.permits(COLLABORATOR, SEND, "term_lead", target -> true),
"the classifier accepting the target must grant SEND");
assertFalse(Authz.permits(COLLABORATOR, SEND, "term_lead", target -> false),
"the classifier refusing the target must deny SEND");
assertFalse(Authz.permits(COLLABORATOR, SEND, "term_lead"),
"the real production classifier recognises no terminal yet, so SEND is refused today");
}
/**
* Control for the test above: every other action's result for a collaborator does not move
* when the classifier does. Only {@code SEND} is wired to it.
*/
@Test
void theClassifierMovesOnlySendForACollaborator() {
for (Authz.Action a : Authz.Action.values()) {
if (a == SEND) {
continue;
}
assertEquals(
Authz.permits(COLLABORATOR, a, "term_lead"),
Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
a + " must not depend on the classifier at all");
}
}
@Test
void aCollaboratorMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(COLLABORATOR, READ, null));
assertTrue(Authz.permits(COLLABORATOR, METRICS, null));
}
@Test
void aCollaboratorMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(COLLABORATOR, REPLY, "term_collab"),
"its own pane is its own");
assertTrue(Authz.permits(COLLABORATOR, ASK, "term_collab"));
assertFalse(Authz.permits(COLLABORATOR, REPLY, "term_design"),
"a collaborator must not reply on another pane");
assertFalse(Authz.permits(COLLABORATOR, REPLY, null),
"an absent target must not pass the own-session rule");
}
/**
* Every action denied to a collaborator, asserted denied even when the classifier would
* accept any target — proving none of these is actually gated on the classifier at all.
*/
@Test
void aCollaboratorIsDeniedLifecycleCoordinationAndTicketPolling() {
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN, HANDOVER, ANSWER, COORD_SEND,
COORD_READ, TASK_READ}) {
assertFalse(Authz.permits(COLLABORATOR, a, "term_lead", target -> true),
"a collaborator must not " + a + " even when the classifier accepts every target");
}
}
@Test
void aCollaboratorIsNotCountedAsPrimaryWorkerOrArchitect() {
assertFalse(COLLABORATOR.isPrimary());
assertFalse(COLLABORATOR.isWorker());
assertFalse(COLLABORATOR.isArchitect());
assertTrue(COLLABORATOR.isCollaborator());
}
/**
* A collaborator is never spawned, so it must not be enrolled in the presence map as an
* available member. Control: both a worker and an architect — which ARE spawned — still are.
*/
@Test
void isSpawnedMemberIsFalseForACollaboratorButTrueForAWorkerAndAnArchitect() {
assertFalse(COLLABORATOR.isSpawnedMember());
assertTrue(WORKER_A.isSpawnedMember());
assertTrue(ARCH_DESIGN.isSpawnedMember());
}
// ── the observer matrix ─────────────────────────────────────────────────────────────────────
@Test
void anObserverMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(OBSERVER, READ, null));
assertTrue(Authz.permits(OBSERVER, METRICS, null));
}
@Test
void anObserverMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(OBSERVER, REPLY, "term_observer"), "its own pane is its own");
assertTrue(Authz.permits(OBSERVER, ASK, "term_observer"));
assertFalse(Authz.permits(OBSERVER, REPLY, "term_design"),
"an observer must not reply on another pane");
assertFalse(Authz.permits(OBSERVER, REPLY, null),
"an absent target must not pass the own-session rule");
}
/**
* Every action beyond READ/METRICS/REPLY/ASK, asserted denied for an observer — including
* {@code TASK_READ}, which is the entire point of this role: an unconfigured pane must not be
* able to poll a ticket or read another session's status.
*/
@Test
void anObserverIsDeniedEverythingBeyondReadMetricsReplyAndAsk() {
for (Authz.Action a : Authz.Action.values()) {
if (a == READ || a == METRICS || a == REPLY || a == ASK) {
continue;
}
assertFalse(Authz.permits(OBSERVER, a, "term_observer", target -> true),
"an observer must not " + a + " even when the classifier accepts every target");
}
}
@Test
void anObserverIsNotCountedAsAnyOtherRole() {
assertFalse(OBSERVER.isPrimary());
assertFalse(OBSERVER.isWorker());
assertFalse(OBSERVER.isArchitect());
assertFalse(OBSERVER.isCollaborator());
assertFalse(OBSERVER.isSpawnedMember());
assertTrue(OBSERVER.isObserver());
}
}
@@ -1,13 +1,25 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.session.WorktreeRequest;
import dev.ltms.fleet.session.Worktrees;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
@@ -45,17 +57,23 @@ class CallerResolverTest {
return members;
}
/**
* With no roster wired up at all (the simple constructor), a loopback pane that owns a herdr
* pane but is not recognised as a live spawned member lands on the {@link Role#OBSERVER} floor
* — unforgeable and never token-gated, exactly like a worker's own identity, because it comes
* from the same connection-derived pane mapping.
*/
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
void aLoopbackPaneWithNoLiveRosterResolvesToObserverRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
Principal underToken = new CallerResolver(workerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, underTrust.role());
assertEquals(Role.OBSERVER, underTrust.role());
assertEquals("term_a", underTrust.terminal());
assertEquals(Role.WORKER, underToken.role(),
"worker identity is unforgeable and must never be token-gated — otherwise enabling "
+ "auth would lock the whole fleet out of fleet_reply");
assertEquals(Role.OBSERVER, underToken.role(),
"the floor is unforgeable and must never be token-gated — otherwise enabling auth "
+ "would lock every unconfigured pane out of even READ");
assertEquals("term_a", underToken.terminal());
}
@@ -79,20 +97,20 @@ class CallerResolverTest {
}
@Test
void otherPanesRemainWorkersWhenAPinIsSet() {
void otherPanesRemainAtTheFloorWhenAPinIsSet() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals(Role.OBSERVER, p.role());
assertEquals("term_a", p.terminal());
}
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test
void aBlankPinLeavesWorkerResolutionUntouched() {
assertEquals(Role.WORKER,
void aBlankPinLeavesFloorResolutionUntouched() {
assertEquals(Role.OBSERVER,
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
assertEquals(Role.OBSERVER,
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
}
@@ -213,13 +231,13 @@ class CallerResolverTest {
* mid-scan teardown into a refusal — the real match is still found and resolves as a worker.
*/
@Test
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealWorker() {
void aHerdrErrorOnANonOwningPaneStillResolvesTheRealPane() {
FakeHerdr vanishedElsewhere = new FakeHerdr().processInfoFailsForPane("w2:p9", "pane_not_found");
ConnectionIdentity id = new ConnectionIdentity(new PaneLocator(vanishedElsewhere), _ -> FakeHerdr.WORKER_PID);
Principal p = new CallerResolver(id).resolve("127.0.0.1", 55555, null);
assertEquals(Role.WORKER, p.role());
assertEquals(Role.OBSERVER, p.role());
assertEquals("term_a", p.terminal());
}
@@ -263,12 +281,12 @@ class CallerResolverTest {
}
@Test
void aPaneAbsentFromTheRegistryIsStillAWorker() {
void aPaneAbsentFromTheRegistryFallsToTheObserverFloor() {
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals(Role.OBSERVER, p.role());
assertEquals("term_a", p.terminal());
assertNull(p.name());
}
@@ -294,16 +312,16 @@ class CallerResolverTest {
}
@Test
void anEmptyRegistryLeavesEveryPaneAWorker() {
void anEmptyRegistryLeavesEveryPaneAtTheObserverFloor() {
Map<String, String> noLeads = null;
assertEquals(Role.WORKER,
assertEquals(Role.OBSERVER,
new CallerResolver(workerIdentity(), false, null, Map.of())
.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
assertEquals(Role.OBSERVER,
new CallerResolver(workerIdentity(), false, null, noLeads)
.resolve("127.0.0.1", 42, null).role());
// CB-531: and the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.WORKER,
// And the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.OBSERVER,
CallerResolver.withLeads(workerIdentity(), false, null, null)
.resolve("127.0.0.1", 42, null).role());
}
@@ -376,7 +394,7 @@ class CallerResolverTest {
Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
@@ -393,7 +411,7 @@ class CallerResolverTest {
mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
}
// ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
@@ -423,13 +441,13 @@ class CallerResolverTest {
}
@Test
void anUnboundPaneStillResolvesAsAWorker() {
void anUnboundPaneResolvesToTheObserverFloor() {
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null));
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members).resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals(Role.OBSERVER, p.role());
assertNull(p.name());
}
@@ -455,12 +473,11 @@ class CallerResolverTest {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.OBSERVER, r.resolve("127.0.0.1", 42, null).role());
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("architect:lead-designer", r.members().get("term_a"));
}
@Test
@@ -483,12 +500,13 @@ class CallerResolverTest {
}
@Test
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
void aBoundNonArchitectSlotResolvesToTheObserverFloorNotArchitect() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("dev:builder", MemberRole.DEV))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
assertEquals(Role.OBSERVER, p.role(), "a dev binding must never grant architect rights, "
+ "and this construction path wires no roster to recognise it as the live dev it is");
}
@Test
@@ -499,17 +517,17 @@ class CallerResolverTest {
assertThrows(IllegalArgumentException.class, () -> new CallerResolver(id, true, " "));
}
@Test
void aWorkerOnAnyLoopbackSourceAddressIsStillAWorkerNotThePrimary() {
// fleetd #305: the escalation. ConnectionIdentity used to accept only 127.0.0.1, so a
// worker connecting from 127.0.0.2 resolved to no terminal, and this resolver's own
// (wider) loopback check then made it the PRIMARY — granting spawn, stop, send and drain.
// Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds there, so the
// path is real and not theoretical.
void aPaneOnAnyLoopbackSourceAddressIsStillAtTheFloorNotThePrimary() {
// fleetd #305: the escalation this guards against. ConnectionIdentity used to accept only
// 127.0.0.1, so a pane connecting from 127.0.0.2 resolved to no terminal, and this
// resolver's own (wider) loopback check then made it the PRIMARY — granting spawn, stop,
// send and drain. Measured on the Linux fleet host: binding a source of 127.0.0.2 succeeds
// there, so the path is real and not theoretical.
CallerResolver r = new CallerResolver(workerIdentity(), false, null);
for (String src : new String[]{"127.0.0.1", "127.0.0.2", "127.1.2.3", "::ffff:127.0.0.2"}) {
Principal p = r.resolve(src, 55555, null);
assertEquals(Role.WORKER, p.role(), "a worker must stay a worker from source " + src);
assertEquals("term_a", p.terminal(), "worker terminal from source " + src);
assertEquals(Role.OBSERVER, p.role(), "the pane must stay off PRIMARY from source " + src);
assertEquals("term_a", p.terminal(), "pane terminal from source " + src);
}
}
@@ -523,4 +541,375 @@ class CallerResolverTest {
}
}
// ── fleetd #669 Unit D: a live spawned member outranks every tab map ────────────────────────
/**
* Criterion 1: a terminal present in BOTH the spawned-member roster AND the lead tab map
* resolves as its member role, not as a lead — the roster is checked first, consulting no tab
* map at all when it matches.
*/
@Test
void aSpawnedMemberWinsOverALeadTabForTheSamePane() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"), new MemberRegistry(null),
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(),
"a live spawned member's own identity must win over a tab map naming the same pane a lead");
assertEquals("term_a", p.terminal());
}
/**
* fleetd #702: between {@link SessionManager#release} removing the registry entry and the
* pane actually stopping, a resolve for that terminal must still see the live member and
* never fall through to a tab map — wired through the real {@link SessionManager}, not a
* hand-rolled stand-in for its {@code spawnedMemberRole}.
*
* <p>Reuses the ticket's own test idea: an injected {@code hasUncommitted} resolves the
* releasing terminal from inside {@code release}'s window — a real call landing inside the
* window, so no sleep and no race.
*
* <p>The property asserted is "no tab map is consulted", never "the same role is returned".
* The mandatory control is the lead tab map: it names this exact terminal, and the same
* resolve taken <em>outside</em> the window (before release runs) must still return the lead
* role — without that control, the in-window assertion would also pass on an empty map and
* prove nothing.
*/
@Test
void aPaneMidTeardownResolvesAsItsOwnRoleConsultingNoTabMap() {
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
AtomicReference<Principal> duringWindow = new AtomicReference<>();
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
Worktrees worktrees = new Worktrees() {
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
// Runs from INSIDE release()'s git-status shell-out: the registry entry is
// already gone, but the pane has not stopped yet.
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
return false;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
return Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
};
SessionManager sessions = new SessionManager(workers, worktrees);
// The lead tab map names "term_a" before anything is ever spawned onto it — the control
// this test needs. Built up front so the SAME resolver answers every resolve() call below.
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_a", "the-lead"), null, sessions::spawnedMemberRole, Map::of);
resolverRef.set(resolver);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, before.role(),
"control: with no live or releasing member on this terminal, the lead tab map must "
+ "win — this is what proves the in-window assertion below is not passing on "
+ "an empty map");
assertEquals("the-lead", before.name());
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702", null));
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
sessions.release(s.paneId());
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
assertEquals(Role.WORKER, duringWindow.get().role(),
"inside the window the pane must resolve as its own live-member role, consulting no "
+ "tab map — a lead tab naming the same terminal must not win");
assertEquals("term_a", duringWindow.get().terminal());
}
/**
* {@link SessionManager#releaseRemoved} unbinds the architect slot before the git-status
* shell-out that opens the teardown window, so a resolve landing inside that window must see
* the slot already unbound and resolve {@link Role#WORKER} — never {@link Role#ARCHITECT},
* and never by checking role equality against the live session, which would hold even if the
* unbind ran too late.
*
* <p>The control is the same resolve taken outside the window, while the slot is still bound,
* which must return {@link Role#ARCHITECT} — without it this test would also pass against a
* slot that was never bound, and prove nothing.
*/
@Test
void aReleasingArchitectIsDemotedToWorkerInsideTheTeardownWindow() {
FakeHerdr herdr = new FakeHerdr().pinNextStarts(1, "term_a", "w2:p7");
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
AtomicReference<Principal> duringWindow = new AtomicReference<>();
AtomicReference<CallerResolver> resolverRef = new AtomicReference<>();
Worktrees worktrees = new Worktrees() {
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
// Runs from INSIDE release()'s git-status shell-out: the architect slot is already
// unbound by this point, but the pane has not stopped yet.
duringWindow.set(resolverRef.get().resolve("127.0.0.1", 42, null));
return false;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
return Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
};
SessionManager sessions = new SessionManager(workers, worktrees);
MemberRegistry members = new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("ltms-local")), Map.of(), Map.of(), null));
sessions.setMemberLifecycle(members);
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, members, sessions::spawnedMemberRole, Map::of);
resolverRef.set(resolver);
MemberSession s = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null, "/caller/proj", null,
new WorktreeRequest("fleetd-702d", null));
assertEquals("term_a", s.terminalId(), "sanity: the spawn resolved to the pinned pane");
assertEquals(MemberRole.ARCHITECT, s.role(), "sanity: the slot bind succeeded");
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(),
"control: with the slot still bound, the pane must resolve as an architect — this "
+ "is what proves the in-window assertion below is not passing against a "
+ "slot that was never bound");
assertEquals("lead-designer", before.name());
sessions.release(s.paneId());
assertNotNull(duringWindow.get(), "the dirty check must have run and captured a resolve");
assertEquals(Role.WORKER, duringWindow.get().role(),
"inside the window the architect slot is already unbound, so the result must be "
+ "WORKER — asserting role-equality with the live session here would tempt "
+ "moving the unbind earlier or later, which would be wrong either way");
assertEquals("term_a", duringWindow.get().terminal());
}
/** A spawned architect in the roster resolves ARCHITECT, carrying its bound slot's name. */
@Test
void aSpawnedArchitectInTheRosterResolvesArchitectWithItsSlotName() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
boundMembers("architect:lead-designer", MemberRole.ARCHITECT),
t -> "term_a".equals(t) ? MemberRole.ARCHITECT : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role());
assertEquals("lead-designer", p.name());
assertEquals("term_a", p.terminal());
}
/**
* fleetd #424 regression: the roster only answers THAT a pane is a live spawned member; config
* still decides WHAT that member's slot grants. A slot revoked after the bind must still demote
* the session on its very next request, exactly as it would for a pane with no roster entry at
* all — the roster's own ARCHITECT role must never be granted on its word alone.
*
* <p>{@code bind} refuses an unconfigured slot, so the revoked state can only be reached by
* binding while the slot is configured and then swapping the config out from under it, the way
* a live reload does.
*/
@Test
void aRevokedArchitectSlotDemotesALiveSpawnedArchitectToWorker() {
FleetConfig.Fleet configured = new FleetConfig.Fleet(Map.of(),
Map.of("lead-designer", new FleetConfig.Slot("sonnet")), Map.of(), Map.of(), null);
java.util.concurrent.atomic.AtomicReference<FleetConfig.Fleet> live =
new java.util.concurrent.atomic.AtomicReference<>(configured);
MemberRegistry members = MemberRegistry.live(live::get);
assertTrue(members.bind("architect:lead-designer", "term_a"));
live.set(new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), null)); // slot revoked
// Setup controls: the slot is really gone from config, but the occupancy is still there —
// otherwise this test would pass for the wrong reason.
assertNull(members.roleForSlot("architect:lead-designer"), "setup control: the slot must be gone from config");
assertEquals("architect:lead-designer", members.snapshot().get("term_a"),
"setup control: the binding itself must still be there");
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
members, t -> "term_a".equals(t) ? MemberRole.ARCHITECT : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(),
"a revoked slot must demote a live spawned architect on its very next request");
}
/** Criterion 3: a configured collaborator tab that is not a spawned member resolves COLLABORATOR. */
@Test
void aConfiguredCollaboratorTabResolvesToCollaboratorCarryingItsName() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, () -> Map.of("term_a", "ops"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.COLLABORATOR, p.role());
assertEquals("ops", p.name());
assertEquals("term_a", p.terminal());
}
/** Regression: an empty collaborator registry leaves every pane at the unconfigured-pane floor. */
@Test
void anEmptyCollaboratorRegistryLeavesEveryPaneAtTheObserverFloor() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, p.role());
assertNull(p.name());
}
@Test
void describeNamesTheCollaborator() {
assertEquals("collaborator:ops", Principal.collaborator("ops", "term_a", 1).describe());
}
// ── fleetd #669 Unit D: knownLeadOrCollaborator() reads the same maps resolve() does ───────────
@Test
void knownLeadOrCollaboratorIsTrueForALeadTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_lead", "opus-5.0"), new MemberRegistry(null), t -> null, Map::of);
assertTrue(r.knownLeadOrCollaborator().test("term_lead"));
assertFalse(r.knownLeadOrCollaborator().test("term_other"));
}
@Test
void knownLeadOrCollaboratorIsTrueForACollaboratorTerminal() {
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> null, () -> Map.of("term_collab", "ops"));
assertTrue(r.knownLeadOrCollaborator().test("term_collab"));
assertFalse(r.knownLeadOrCollaborator().test("term_other"));
}
// ── fleetd #705: narrowing the unconfigured-pane floor to OBSERVER ──────────────────────────
/**
* The case this ticket exists for: a pane the resolver cannot place as a live spawned member,
* a lead, a bound architect slot, or a configured collaborator must land on the narrow
* {@link Role#OBSERVER} floor, never the {@link Role#WORKER} the old fallback granted.
*
* <p>The second assertion is the control the ticket requires: a terminal the roster DOES
* recognise as a live spawned member must still resolve its own role. Without it, this test
* would also pass if the fix accidentally turned every caller into an observer.
*/
@Test
void anUnconfiguredPaneResolvesObserverButARegisteredMemberStillResolvesItsOwnRole() {
Principal unconfigured = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
assertEquals(Role.OBSERVER, unconfigured.role(),
"a pane matching none of the configured or live-roster roles must fall to the "
+ "floor, not WORKER");
assertEquals("term_a", unconfigured.terminal());
Principal registered = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, new MemberRegistry(null),
t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of)
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, registered.role(),
"control: a live spawned member must keep resolving its own role, never the "
+ "unconfigured-pane floor");
}
@Test
void describeNamesTheObserverByItsPane() {
assertEquals("observer:term_a", Principal.observer("term_a", 1).describe());
}
@Test
void knownLeadOrCollaboratorIsFalseForASpawnedMembersTerminal() {
// The exact scenario a collaborator's SEND must never reach: a live spawned member's own
// terminal, which is neither a configured lead nor a configured collaborator.
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null, Map::of,
new MemberRegistry(null), t -> "term_a".equals(t) ? MemberRole.DEV : null, Map::of);
assertFalse(r.knownLeadOrCollaborator().test("term_a"));
}
}
@@ -295,9 +295,10 @@ class MemberRegistryLiveTest {
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, after.role(),
"removing the slot from config must demote the bound session to worker on its "
+ "NEXT request — this is the ticket's whole point");
assertEquals(Role.OBSERVER, after.role(),
"removing the slot from config must demote the bound session on its NEXT request — "
+ "this harness wires no live roster for term_a, so the demotion lands on "
+ "the unconfigured-pane floor");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
}
@@ -0,0 +1,31 @@
package dev.ltms.fleet.auth;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
class PrincipalTest {
@Test
void ownerKeyCoversEveryRole() {
assertEquals("leader:opus", Principal.leader("opus", "term_lead", 1).ownerKey());
assertNull(Principal.primary(2).ownerKey());
assertEquals("worker:term_worker", Principal.worker("term_worker", 3).ownerKey());
assertEquals("architect:term_arch", Principal.architect("opus", "term_arch", 4).ownerKey());
assertEquals("collaborator:ops", Principal.collaborator("ops", "term_collab", 5).ownerKey());
assertEquals("observer:term_observer", Principal.observer("term_observer", 6).ownerKey());
assertEquals("anonymous", Principal.anonymous().ownerKey());
}
@Test
void rolePrefixesKeepLeadAndArchitectKeysDistinct() {
String lead = Principal.leader("opus", "term_lead", 1).ownerKey();
String architect = Principal.architect("design", "opus", 2).ownerKey();
assertEquals("leader:opus", lead);
assertEquals("architect:opus", architect);
assertNotEquals(lead, architect);
}
}
@@ -107,6 +107,8 @@ class ConfigRefTest {
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty());
assertTrue(out.split().stream().noneMatch(s -> s.contains("fleet.collaborators")),
out.split().toString());
assertEquals("new charter", ref.get().fleet().charterFor(
dev.ltms.fleet.peer.MemberRole.ARCHITECT));
}
@@ -786,14 +788,100 @@ class ConfigRefTest {
assertTrue(out.split().getFirst().startsWith("fleet:"), out.split().toString());
assertTrue(out.split().getFirst().contains("restart"), out.split().toString());
assertTrue(out.split().getFirst().contains("live"), out.split().toString());
assertTrue(out.split().stream().noneMatch(s -> s.contains("fleet.collaborators")),
out.split().toString());
assertTrue(out.summary().contains("partially live"), out.summary());
// The snapshot still carries the new value — the LeadTabScanner's identity map and
// LeadLauncher's auto-launch are what wait for a restart; a reload rebuilds neither.
assertEquals("lead: opus-b", ref.get().fleet().leaders().get("opus").tab());
}
@Test
void addingAFleetCollaboratorIsReportedAsSplit(@TempDir Path dir) throws Exception {
assertCollaboratorChangeIsReported(dir, """
fleet:
collaborators:
alex:
tab: "collaborator: alex"
""", "add a collaborator");
}
@Test
void removingAFleetCollaboratorIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex"
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("fleet: {}\n"));
assertCollaboratorSplit(ref.reload(), "remove a collaborator");
}
@Test
void changingAFleetCollaboratorTabIsReportedAsSplit(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex-a"
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex-b"
"""));
assertCollaboratorSplit(ref.reload(), "change a collaborator tab");
}
@Test
void unchangedFleetCollaboratorsProduceNoSplitReport(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
String config = yaml("""
fleet:
collaborators:
alex:
tab: "collaborator: alex"
""");
Files.writeString(f, config);
ConfigRef ref = refFor(f);
Files.writeString(f, config);
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.split().isEmpty(), out.split().toString());
}
private static void assertCollaboratorChangeIsReported(Path dir, String changed, String action)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml("fleet: {}\n"));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml(changed));
assertCollaboratorSplit(ref.reload(), action);
}
private static void assertCollaboratorSplit(ConfigRef.Outcome out, String action) {
assertTrue(out.applied(), action);
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(1, out.split().size(), out.split().toString());
String report = out.split().getFirst();
assertTrue(report.startsWith("fleet: fleet.collaborators"), report);
assertTrue(report.contains("read once at startup"), report);
assertTrue(report.contains("restart"), report);
}
/**
* fleetd #333: {@code fleet.leaders} is the ONLY frozen part of {@code fleet:}. A reload that
* fleetd #333: {@code fleet.leaders} is a frozen part of {@code fleet:}. A reload that
* changes {@code tabLabel} (or charters, or a role pool) without touching {@code fleet.leaders}
* must stay fully hot with nothing reported — proving {@link ConfigRef#changedSplitKeys}
* compares {@code fleet.leaders} specifically rather than the whole {@code Fleet} record, which
@@ -676,12 +676,8 @@ class FleetConfigTest {
assertEquals(5, hb.quietNudgeCap());
}
/**
* The hazard the guard exists for: fleetd writes worker tab labels and reads lead tab labels.
* Overlap the two and every worker it spawns is read back as a lead.
*/
@Test
void aLeadPrefixThatAProfileTabLabelOverrideAlsoMatchesRefusesToStart(@TempDir Path dir)
void aProfileTabLabelOverrideMatchingALeadTabRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("collide.yaml");
Files.writeString(f, """
@@ -689,32 +685,33 @@ class FleetConfigTest {
port: 8080
profiles:
gx10:
tabLabel: "lead: {profile} #{n}"
tabLabel: "alpha"
fleet:
leaders:
opus:
tab: "lead: opus"
tabPrefix: "lead:"
tab: "alpha"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
assertTrue(e.getMessage().contains("alpha"), "the message must name the offending label");
}
/** A bad fleet-wide template promotes every member, not one profile — so it is checked too. */
@Test
void aFleetTabLabelThatMatchesALeadPrefixRefusesToStart(@TempDir Path dir) throws Exception {
void aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collide-template.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
pha: {}
fleet:
tabLabel: "lead: {role} {profile}"
tabLabel: "al{profile}"
leaders:
opus:
tab: "lead: opus"
tab: "alpha"
""");
FleetConfig cfg = FleetConfig.load(f);
@@ -723,12 +720,28 @@ class FleetConfigTest {
assertTrue(e.getMessage().contains("fleet.tabLabel"));
}
/**
* The point of making role the label's first field: {@code {role}} comes from a closed enum, so
* a generated label cannot begin with {@code "lead:"} however the fleet is configured.
*/
@Test
void theDefaultTabLabelCannotCollideWithTheDefaultLeadPrefix(@TempDir Path dir) throws Exception {
void anExactFleetTabLabelCollisionRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("exact-tab-collision.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
tabLabel: "alpha"
leaders:
alpha:
tab: "alpha"
""");
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> FleetConfig.load(f).validateAll());
assertTrue(e.getMessage().contains("fleet.tabLabel"),
"the message must name the offending label");
assertTrue(e.getMessage().contains("alpha"), "the message must name the colliding lead tab");
}
@Test
void aFleetTabLabelTemplateThatCannotRenderAsALeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("ok.yaml");
Files.writeString(f, """
bind:
@@ -737,17 +750,13 @@ class FleetConfigTest {
gx10:
baseUrl: http://gx00.gw:8000
fleet:
tabLabel: "worker-{profile}"
leaders:
opus:
tab: "lead: opus"
tab: "alpha"
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validateLeadTabPrefixes());
for (MemberRole role : MemberRole.values()) {
assertFalse(FleetConfig.Fleet.DEFAULT_TAB_LABEL
.replace("{role}", role.wireName()).startsWith("lead:"),
"no role renders a label that reads as a lead");
}
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
}
@Test
@@ -765,11 +774,84 @@ class FleetConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
/**
* fleetd #677: identity is matched on a lead's exact {@code tab} alone, so two leads sharing
* one tab means only one of them is ever found — the guard must catch this independently of
* the member-template checks above.
*/
@Test
void twoLeadsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "shared tab"
sonnet:
tab: "shared tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
}
/**
* fleetd #693: the guard matches tabs case-insensitively, because
* {@code LeadTabScanner} keys its tab map on a lowercased label — two tabs differing only in
* case collide there too, and the guard must catch that independently of the exact-match case
* above.
*/
@Test
void twoLeadsSharingTheSameTabInDifferentCaseRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-tab-case.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "Shared Tab"
sonnet:
tab: "shared tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
}
/** Control for {@link #twoLeadsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void twoLeadsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("distinct-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "opus tab"
sonnet:
tab: "sonnet tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
/**
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
* its own, so it can land inside a lead's labelled tab and be read back as that lead.
* its own, so it can land inside a lead's labelled tab and, while its pane carries no entry in
* the spawned-member roster, be read back as that lead.
*/
@Test
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
@@ -919,6 +1001,274 @@ class FleetConfigTest {
"no primary.terminal pin ⇒ nothing registered, even with fleet.leaders configured");
}
// ── fleetd #669: the collaborators registry ────────────────────────────────────────────────
/**
* {@code Fleet} is {@code @JsonIgnoreProperties(ignoreUnknown = true)}, so a config naming
* {@code fleet.collaborators.<name>.tab} loads with no exception whether or not the key is
* ever read into the object model. Asserting only "no exception" would pass both before and
* after the real fix, so this asserts the parsed value is actually reachable from the loaded
* {@code FleetConfig} — the one thing a vacuous "no exception" test cannot tell apart.
*/
@Test
void collaboratorsBlockIsActuallyParsedNotSilentlyDropped(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborators.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
reviewer-alex:
tab: "collab: alex"
""");
FleetConfig cfg = FleetConfig.load(f);
assertEquals("collab: alex", cfg.fleet().collaborators().get("reviewer-alex").tab());
}
/** A collaborator carries no field other than {@code tab}, so a blank one is meaningless. */
@Test
void aCollaboratorWithNoTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("useless-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
ghost: {}
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateMembers);
assertTrue(e.getMessage().contains("ghost"), "the message must name the useless entry");
assertTrue(e.getMessage().contains("tab:"), "the message must say what is missing");
}
/** Control for {@link #aCollaboratorWithNoTabRefusesToStart}: a named tab loads cleanly. */
@Test
void aCollaboratorWithATabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("named-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
reviewer-alex:
tab: "collab: alex"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateMembers);
}
/**
* fleetd #669: the fleet-wide {@code tabLabel} template can render as a collaborator tab, the
* same hazard {@link #aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart} covers on
* the lead side. Drives the fleet-wide branch directly, with no profile override involved.
*/
@Test
void aFleetTabLabelTemplateThatCanRenderAsACollaboratorTabRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("collide-template-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
tabLabel: "al{profile}"
collaborators:
alex:
tab: "alpha"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("fleet.tabLabel"));
assertTrue(e.getMessage().contains("alex"), "the message must name the offending collaborator");
}
/**
* A member tabLabel that can render as a configured collaborator tab is the same hazard as the
* lead case above — while its pane carries no entry in the spawned-member roster, a member
* labelled that way is read back as the collaborator.
*/
@Test
void aProfileTabLabelOverrideMatchingACollaboratorTabRefusesToStart(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("collide-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
tabLabel: "collab-tab"
fleet:
collaborators:
alex:
tab: "collab-tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
assertTrue(e.getMessage().contains("collab-tab"), "the message must name the offending label");
}
/** Control: a profile tabLabel that cannot render as the collaborator tab is allowed. */
@Test
void aProfileTabLabelThatCannotRenderAsACollaboratorTabIsAllowed(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("ok-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
tabLabel: "worker-{profile}"
fleet:
collaborators:
alex:
tab: "collab-tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/** fleetd #669: identity is matched on a collaborator's exact tab, so two sharing one are unreachable. */
@Test
void twoCollaboratorsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("shared-collaborator-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
alex:
tab: "shared tab"
sam:
tab: "Shared Tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("alex"), "the message must name one offending collaborator");
assertTrue(e.getMessage().contains("sam"), "the message must name the other offending collaborator");
}
/** Control for {@link #twoCollaboratorsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void twoCollaboratorsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("distinct-collaborator-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
alex:
tab: "alex tab"
sam:
tab: "sam tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/**
* fleetd #669: a collaborator tab equal to a lead tab crosses a privilege boundary — the worst
* of the three new collisions, since only one of the two identities is ever found.
*/
@Test
void aCollaboratorTabEqualToALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("lead-collaborator-collision.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "shared tab"
collaborators:
alex:
tab: "Shared Tab"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e =
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
assertTrue(e.getMessage().contains("opus"), "the message must name the offending lead");
assertTrue(e.getMessage().contains("alex"), "the message must name the offending collaborator");
}
/** Control for {@link #aCollaboratorTabEqualToALeadTabRefusesToStart}: distinct tabs load cleanly. */
@Test
void aLeadAndACollaboratorWithDistinctTabsAreAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("lead-collaborator-ok.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead tab"
collaborators:
alex:
tab: "collab tab"
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
}
/**
* fleetd #669: a pane-placed member can land in a collaborator's labelled tab exactly as it
* can land in a lead's — {@code validatePanePlacementAgainstLeadTabs()} must fire even when
* {@code fleet.leaders} is empty, which is the early-return the brief flagged as the bug.
*/
@Test
void aPanePlacedProfileWithACollaboratorTabRefusesToStartEvenWithNoLeaders(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("pane-hazard-collaborator.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
collaborators:
alex:
tab: "collab: alex"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
/** Control: a pane-placed profile with no lead or collaborator tab configured is allowed. */
@Test
void aPanePlacedProfileWithNoLeaderOrCollaboratorTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab-at-all.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
collaborators:
alex: {}
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a collaborator with no tab feeds nothing into the scanner, so pane placement is safe");
}
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
@Test
@@ -1191,6 +1541,60 @@ class FleetConfigTest {
assertEquals(Set.of("sonnet"), cfg.fleet().pool(MemberRole.REVIEWER).keySet());
}
/**
* A duplicated name in {@code fleet.collaborators} is refused at parse time, like any other
* {@code fleet:} pool. See {@link #duplicateSlotNamesInOnePoolAreRejectedAtParseTime}.
*/
@Test
void duplicateCollaboratorNamesAreRejectedAtParseTime(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborator-dup.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
collaborators:
alex:
tab: "collab: alex"
alex:
tab: "collab: alex, second"
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> FleetConfig.load(f));
assertTrue(e.getMessage().contains("alex"),
"the refusal names the duplicated entry, was: " + e.getMessage());
assertTrue(e.getMessage().contains("fleet.collaborators"),
"the refusal names the pool the duplicate is in, was: " + e.getMessage());
}
/**
* Control for {@link #duplicateCollaboratorNamesAreRejectedAtParseTime}: the same name reused
* across the collaborators registry and a member role pool is the role × profile matrix doing
* its job in the other pool, not a mistake — only a repeat within one pool loses an entry.
*/
@Test
void theSameNameInCollaboratorsAndAnotherPoolIsNotADuplicate(@TempDir Path dir) throws Exception {
Path f = dir.resolve("collaborator-cross-pool.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
fleet:
architects:
alex:
profile: sonnet
collaborators:
alex:
tab: "collab: alex"
""");
FleetConfig cfg = assertDoesNotThrow(() -> FleetConfig.load(f));
assertEquals(Set.of("alex"), cfg.fleet().pool(MemberRole.ARCHITECT).keySet());
assertEquals(Set.of("alex"), cfg.fleet().collaborators().keySet());
}
@Test
void duplicateKeysOutsideTheFleetPoolsAreUnaffected(@TempDir Path dir) throws Exception {
// The duplicate check is scoped to the fleet pools — a duplicate elsewhere is not this
@@ -17,63 +17,21 @@ import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
* measurement (same technique — remove one call site, run the suite, not read the code) found the
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
* of those thirteen tests called the validator itself directly, never the real caller that was
* supposed to.
* Tests the reflective validator sweep and {@link FleetConfig#validateAll()} reachability.
*
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
* keeps them separate on purpose:
*
* <ol>
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
* happens to declare today, including a class with more of them than {@link FleetConfig}
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysRealValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches every one of today's real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, plus a
* dedicated fixture for {@link FleetConfig#validateLeadRollover()}, which no other test
* drives through {@code validateAll()} — so a single call to {@code validateAll()} is shown
* to reproduce every one of those failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
* a test in this module.
*
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
* FleetConfig#validateAll()} to a hardcoded list of today's method calls leaves the whole
* suite green. Nothing ties {@code validateAll()} to
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's validators, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseValidatorsToday()}: it fails the moment any validator is added
* or removed, which forces whoever changes the set to look at this file.
* <p>{@link #theSweepRunsEveryValidateMethodOnAnUnrelatedClass()} and its neighbours
* prove that {@link FleetConfig#invokeAllValidators} runs each public, no-arg, void
* {@code validateXxx()} method on its target. {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}
* is the canary for the validator set. {@link #validateAllReachesEveryOneOfTodaysRealValidators()}
* is the reachability check for that set.
*/
class FleetConfigValidateAllTest {
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
// ── Claim 1: the reflective sweep is a general mechanism ─────────────────────────────────────
/**
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
* with this shape, not on something special-cased to {@link FleetConfig}.
* Fixture with public, no-arg, void methods named {@code validateXxx}. It proves the sweep uses
* the target's method shape rather than special handling for {@link FleetConfig}.
*/
static class ThreeValidators {
final List<String> ran = new ArrayList<>();
@@ -92,7 +50,7 @@ class FleetConfigValidateAllTest {
}
@Test
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
void theSweepRunsEveryValidateMethodOnAnUnrelatedClass() {
ThreeValidators target = new ThreeValidators();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
@@ -102,12 +60,8 @@ class FleetConfigValidateAllTest {
}
/**
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
* because it exists and matches the shape. This is what makes adding a seventh real validator
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
* call site — there is no "wire it in" step left to forget.
* Fixture with an added valid method. It proves the sweep reaches a method because it matches
* the validator shape.
*/
static class FourValidators {
final List<String> ran = new ArrayList<>();
@@ -135,7 +89,7 @@ class FleetConfigValidateAllTest {
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
sorted(target.ran),
"the fourth method must be reached automatically — proving a class can grow the "
"the added method must be reached automatically — proving a class can grow the "
+ "set of things it validates with no change to the sweep itself");
}
@@ -297,16 +251,16 @@ class FleetConfigValidateAllTest {
port: 8765
""", "auth.mode: token");
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
// validateLeadTabPrefixes: a fleet-wide tabLabel that equals a lead tab.
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "lead: {role} {profile}"
tabLabel: "alpha"
leaders:
opus:
tab: "lead: opus"
tab: "alpha"
""", "fleet.tabLabel");
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
@@ -2,6 +2,8 @@ package dev.ltms.fleet.herdr;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertNotSame;
@@ -28,4 +30,44 @@ class HerdrRouterTest {
assertSame(router.leadAgents(), router.agentsFor("lead"));
assertSame(router.memberAgents(), router.agentsFor("member"));
}
/**
* fleetd #669 Unit E. The predicate shape here is exactly what {@code FleetdAssembly} builds:
* true when the terminal is a known lead OR a known collaborator. Two distinct clients are
* required — with one shared client {@code agentsFor} would return the same object regardless
* of the predicate's answer, and this assertion would pass whether or not the collaborator map
* was ever consulted.
*/
@Test
void collaboratorTerminalRoutesToTheLeadDaemon() {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
Map<String, String> leads = Map.of("term_lead", "primary");
Map<String, String> collaborators = Map.of("term_collab", "reviewer-alex");
HerdrRouter router = new HerdrRouter(lead, member,
id -> leads.containsKey(id) || collaborators.containsKey(id));
assertSame(router.leadAgents(), router.agentsFor("term_collab"),
"a configured collaborator's terminal must route to the LEAD daemon, not the member "
+ "one — its pane is opened by a person, exactly like a lead's");
}
/**
* Companion to {@link #collaboratorTerminalRoutesToTheLeadDaemon}: a terminal that is neither a
* known lead nor a known collaborator must still route to the member daemon. Without this, a
* predicate of {@code _ -> true} would also pass the test above.
*/
@Test
void terminalInNeitherMapStillRoutesToTheMemberDaemon() {
FakeHerdr lead = new FakeHerdr();
FakeHerdr member = new FakeHerdr();
Map<String, String> leads = Map.of("term_lead", "primary");
Map<String, String> collaborators = Map.of("term_collab", "reviewer-alex");
HerdrRouter router = new HerdrRouter(lead, member,
id -> leads.containsKey(id) || collaborators.containsKey(id));
assertSame(router.memberAgents(), router.agentsFor("term_worker"),
"a terminal absent from both maps must stay on the member daemon — the fix widens "
+ "the predicate, it does not make it unconditionally true");
}
}
@@ -163,6 +163,13 @@ class LeadTabScannerTest {
return new LeadTabScanner(herdr, tabToName, Set.of("fleetd-workers"), TTL, clock::get);
}
private LeadTabScanner scannerWithCollaborators(TopologyHerdr herdr, Map<String, String> tabToName,
Map<String, String> collaboratorTabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, tabToName, collaboratorTabToName,
Set.of("fleetd-workers"), TTL, clock::get);
}
@Test
void everyConfiguredTabBecomesALeadNamedByItsEntry() {
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
@@ -198,9 +205,12 @@ class LeadTabScannerTest {
}
/**
* The guard that matters: fleetd labels its own worker tabs, so if a worker space were scanned
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
* worker template never collides.
* Covers the {@code excludedWorkspaceLabels} parameter: a tab in an excluded workspace is never
* matched, whatever its label. Production always constructs this class with an empty set (CB-558,
* {@code FleetdAssembly}), so this parameter plays no part in the live guard against a worker
* tab being mistaken for a lead — that guard is {@code CallerResolver} asking the live
* spawned-member roster before any tab map. This test exists because the parameter still exists
* and is worth covering on its own terms.
*/
@Test
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
@@ -212,6 +222,29 @@ class LeadTabScannerTest {
assertFalse(scanner(herdr, tabToName, new AtomicLong()).get().containsKey("term_impostor"));
}
/**
* The shipped shape (CB-558, {@code FleetdAssembly}): production always constructs this class
* with an empty {@code excludedWorkspaceLabels}, and a lead's {@code workspace:} default is the
* same shared {@code "fleet"} space the members use. A lead tab is still found when it sits in
* the exact same workspace as a member-labelled tab — the scanner tells them apart by the exact
* tab label, not by which workspace either one is in.
*/
@Test
void aLeadIsDiscoveredWhenItsWorkspaceIsTheSameAsTheMemberWorkspace() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "fleet")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_worker");
LeadTabScanner s = new LeadTabScanner(herdr, Map.of("lead: opus-5.0", "opus-5.0"),
Set.of(), TTL, new AtomicLong()::get);
assertEquals("opus-5.0", s.get().get("term_opus"),
"a lead sharing the members' workspace is still discovered — the label, not the "
+ "workspace, is what matches it");
}
@Test
void aLabelWithNoConfiguredEntryIsIgnored() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
@@ -464,4 +497,75 @@ class LeadTabScannerTest {
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
}
// ── fleetd #669 Unit D: a collaborator tab is matched the same way as a lead tab, one pass ─────
@Test
void aConfiguredCollaboratorTabIsReportedByCollaboratorsNotByGet() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "collab: ops")
.pane("w1:p1", "w1:t1", "term_ops");
LeadTabScanner s = scannerWithCollaborators(herdr, Map.of(), Map.of("collab: ops", "ops"),
new AtomicLong());
assertEquals(Map.of("term_ops", "ops"), s.collaborators(),
"a collaborator tab is matched exactly like a lead tab");
assertEquals(Map.of(), s.get(), "a collaborator tab must never also appear as a lead");
}
/**
* Criterion 4, scanner level: a collaborator tab that is labelled but runs no agent is not
* reported — the same #359 liveness cross-check a lead tab gets.
*/
@Test
void aDeadCollaboratorTabIsNotReported() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "collab: ops")
.pane("w1:p1", "w1:t1", "term_ops")
.deadAgent("w1:t1");
LeadTabScanner s = scannerWithCollaborators(herdr, Map.of(), Map.of("collab: ops", "ops"),
new AtomicLong());
assertFalse(s.collaborators().containsKey("term_ops"),
"a dead collaborator tab must never resolve as a live collaborator");
}
/**
* Both kinds are matched in a single pass over the same tab list — not a second scanner, not a
* second scan. Proven by herdr call count: scanning one lead tab and one collaborator tab in the
* same instance costs exactly as many calls as scanning two lead tabs in {@link #twoLeads()}.
*/
@Test
void leadsAndCollaboratorsAreMatchedInOnePassOverTheSameScan() {
TopologyHerdr oneOfEach = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "collab: ops")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_ops");
LeadTabScanner s = scannerWithCollaborators(oneOfEach, Map.of("lead: opus-5.0", "opus-5.0"),
Map.of("collab: ops", "ops"), new AtomicLong());
assertEquals(Map.of("term_opus", "opus-5.0"), s.get());
assertEquals(Map.of("term_ops", "ops"), s.collaborators());
TopologyHerdr twoLeadsBaseline = twoLeads();
scanner(twoLeadsBaseline, twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(twoLeadsBaseline.calls, oneOfEach.calls,
"one lead tab + one collaborator tab must cost exactly as many herdr calls as two "
+ "lead tabs — proof this is one pass, not a second scan");
}
/** Regression: with no collaborators configured, every existing lead-only behaviour is unchanged. */
@Test
void anEmptyCollaboratorMapLeavesCollaboratorsEmptyAndGetUnaffected() {
LeadTabScanner s = scannerWithCollaborators(twoLeads(), twoLeadsConfigured(), Map.of(),
new AtomicLong());
assertEquals(Map.of(), s.collaborators());
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), s.get());
}
}
@@ -1,10 +1,16 @@
package dev.ltms.fleet.lead;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.ResilientAgentLaunch;
import dev.ltms.fleet.herdr.WorkspaceControl;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.LinkedHashMap;
import java.util.List;
@@ -64,6 +70,21 @@ class LeadLauncherTest {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
}
/** As {@link #launcher}, plus a fast no-op sleeper so a busy-retry test never real-sleeps. */
private static LeadLauncher fastLauncher(FakeHerdr herdr, FleetConfig cfg) {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg, () -> { });
}
/** {@link #opusProfile()} with one argv element long enough to overflow the pane line limit. */
private static FleetConfig.Profile hugeArgvProfile() {
return new FleetConfig.Profile(
"opus", null, "claude-opus-5", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "x".repeat(1500)), "tab", "fleet", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
@SuppressWarnings("unchecked")
private static List<String> startedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
@@ -86,7 +107,11 @@ class LeadLauncherTest {
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
assertEquals("lead-opus", startedName(herdr));
// fleetd #727: the name carries a per-process nonce and a per-start sequence number — the
// same unique-naming scheme HerdrPeerLauncher uses for members — rather than the fixed
// "lead-opus" a stale registry entry could block a legitimate relaunch under.
assertTrue(startedName(herdr).matches("lead-opus-[0-9a-f]{6}-\\d+"),
"name is lead-<name>-<nonce>-<seq>: " + startedName(herdr));
}
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
@@ -436,4 +461,247 @@ class LeadLauncherTest {
assertEquals(0, launcher(herdr, cfg).ensureLeads());
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
}
// ── fleetd #727: the same three launch protections every member spawn gets ───────────────────
/**
* A freshly created pane may not have redrawn its prompt yet, so herdr answers
* {@code agent_pane_busy}. The lead launch must wait it out rather than fail on the first miss
* — exactly the retry {@code HerdrPeerLauncher} already gives every member.
*/
@Test
void aLeadLaunchRetriesWhileTheSeedShellBoots() {
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
assertEquals(1, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the lead must still start once the shell is ready");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"two busy rejections, then the successful start");
}
/**
* The busy retry is bounded, not an infinite poll. If the pane never becomes ready the launch
* must eventually give up and log the failure, not hang the daemon's reconcile loop forever —
* proven here by a budget that would still be busy on attempt
* {@value ResilientAgentLaunch#SHELL_READY_RETRIES} and a call count that stops exactly there.
*/
@Test
void aLeadLaunchGivesUpAfterTheBoundedBusyBudgetRatherThanLoopingForever() {
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(999);
assertEquals(0, fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a pane that never becomes ready must not be reported as a started lead");
assertEquals(ResilientAgentLaunch.SHELL_READY_RETRIES,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"the retry budget is bounded: it stops after exactly SHELL_READY_RETRIES attempts");
}
/**
* herdr refuses a duplicate agent {@code name} outright. A stale registry entry — a crashed
* lead session, or a name the registry has not yet released — must not permanently block a
* legitimate relaunch, so each retry attempt carries a fresh per-start name.
*/
@Test
@SuppressWarnings("unchecked")
void aLeadLaunchRetriesUnderAFreshNameWhenTheOldNameIsStillTaken() {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"the lead must still start once a free name is found");
List<String> names = herdr.calls.stream()
.filter(c -> c.method().equals("agent.start"))
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
.toList();
assertEquals(3, names.size(), "2 rejected + 1 success");
assertEquals(3, Set.copyOf(names).size(), "each attempt must use a distinct name");
}
/**
* The name-collision retry is bounded too. If the name is taken on every attempt, the launch
* must give up rather than keep minting new names forever — the per-start naming scheme means
* every failed attempt was refused outright by herdr (no process exists under a name herdr
* refused), so a bounded, exhausted retry never leaves a second process running: nothing is
* running at all.
*/
@Test
void aLeadLaunchNeverEndsUpWithASecondProcessWhenTheNameStaysTaken() {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(999);
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a name that is never free must not be reported as a started lead");
assertEquals(ResilientAgentLaunch.NAME_RETRIES,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"the retry budget is bounded: it stops after exactly NAME_RETRIES attempts, "
+ "never racing a duplicate into existence");
}
/**
* fleetd #220/#727: herdr types the launch command into the pane as one line, and a pty line
* buffer holds only 1024 bytes — past that the tail is dropped with no error at all, and the
* backend exits on a mangled argument. The lead launch must refuse an over-long command outright
* rather than let it be typed and silently truncated.
*/
@Test
void anOverlongLeadArgvIsRefusedRatherThanTypedAndTruncated() {
FakeHerdr herdr = new FakeHerdr();
Logger logger = (Logger) LoggerFactory.getLogger(LeadLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
int started;
try {
started = launcher(herdr, configWith(lead("opus", "lead: opus", 1), hugeArgvProfile()))
.ensureLeads();
} finally {
logger.detachAppender(appender);
}
assertEquals(0, started, "an over-long command must never be reported as a started lead");
assertFalse(herdr.called("agent.start"),
"nothing may be started — a truncated command is worse than no spawn");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("failed to launch"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a WARN naming the launch failure: "
+ appender.list));
assertTrue(warn.contains("1024"), "names the limit: " + warn);
assertTrue(warn.contains("opus"), "names the profile: " + warn);
assertTrue(warn.contains("x".repeat(60)), "names the culprit argument: " + warn);
}
// ── fleetd #726 unit 1: the single-lead relaunch seam ─────────────────────────────────────
/**
* The returned agent's {@code terminalId()}/{@code paneId()} are the ones the fake
* {@code AgentControl} actually started — not a coincidental field left over from the caller.
* {@code paneId()} echoes the exact {@code pane_id} the launch's own {@code agent.start} call
* carried (protocol 19: the agent starts into the pane it is asked to), and {@code
* terminalId()} is herdr's own generated id, which the fake always shapes as {@code
* term_new_<n>}.
*/
@Test
void relaunchReturnsTheStartedAgent() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started, "a launchable, configured lead must start");
Object startedPaneIdParam = ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("pane_id");
assertEquals(startedPaneIdParam, started.paneId(),
"paneId() must be the pane the agent.start call actually targeted");
assertTrue(started.terminalId() != null && started.terminalId().startsWith("term_new_"),
"terminalId() must be herdr's own generated id: " + started.terminalId());
}
/** The new tab is labelled with the lead's configured {@code tab:}, and AFTER the start. */
@Test
void relaunchLabelsTheNewTabAfterStarting() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started);
assertEquals("lead: opus", ((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
int startIndex = indexOfLastCall(herdr, "agent.start");
int renameIndex = indexOfLastCall(herdr, "tab.rename");
assertTrue(renameIndex > startIndex,
"the tab must be renamed AFTER the start succeeds, not before: start=" + startIndex
+ " rename=" + renameIndex);
}
private static int indexOfLastCall(FakeHerdr herdr, String method) {
int idx = -1;
List<FakeHerdr.Call> calls = herdr.calls;
for (int i = 0; i < calls.size(); i++) {
if (calls.get(i).method().equals(method)) {
idx = i;
}
}
return idx;
}
@Test
void relaunchOfAnUnknownLeadNameReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("not-declared");
assertNull(started);
assertFalse(herdr.called("agent.start"));
assertFalse(herdr.called("workspace.create"));
assertFalse(herdr.called("tab.create"));
}
@Test
void relaunchOfARecogniseOnlyLeadReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead(null, "lead: dead", 1))).relaunch("opus");
assertNull(started);
assertFalse(herdr.called("agent.start"));
}
@Test
void relaunchWithAnUnconfiguredProfileReturnsNullAndStartsNothing() {
FakeHerdr herdr = new FakeHerdr();
dev.ltms.fleet.herdr.Agent started =
launcher(herdr, configWith(lead("nope", "lead: opus", 1))).relaunch("opus");
assertNull(started);
assertFalse(herdr.called("agent.start"));
}
/**
* The outer retry {@link LeadLauncher#relaunch(String)} owns, separate from {@code
* ResilientAgentLaunch}'s internal {@code agent_name_taken} retry: a failed attempt must not
* be the end of the whole relaunch. Each of the first two attempts exhausts {@code
* ResilientAgentLaunch.NAME_RETRIES} name attempts (every one of them rejected), so each
* attempt's own tab is created and then closed; the third attempt's first name is free.
*/
@Test
void relaunchRetriesTheWholeAttemptAndSucceedsOnTheThird() {
FakeHerdr herdr = new FakeHerdr()
.agentNameTakenTimes(2 * ResilientAgentLaunch.NAME_RETRIES);
dev.ltms.fleet.herdr.Agent started =
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNotNull(started, "the third attempt's first name is free — it must succeed");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count(),
"one tab per attempt: three attempts");
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count(),
"the two failed attempts' tabs must be closed");
}
/**
* Every attempt fails outright (a herdr error {@code ResilientAgentLaunch} does not retry at
* all) — {@link LeadLauncher#relaunch(String)} must give up after exactly {@code
* RELAUNCH_ATTEMPTS} and must not leak any of the tabs it created along the way.
*/
@Test
void relaunchGivesUpAfterExactlyRelaunchAttemptsAndLeaksNoTab() {
FakeHerdr herdr = new FakeHerdr().agentStartFailsWith("some_other_error");
dev.ltms.fleet.herdr.Agent started =
fastLauncher(herdr, configWith(lead("opus", "lead: opus", 1))).relaunch("opus");
assertNull(started, "every attempt failed — relaunch must give up, not hang or guess");
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS,
herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"exactly RELAUNCH_ATTEMPTS attempts, no more, no fewer");
long tabsCreated = herdr.calls.stream().filter(c -> c.method().equals("tab.create")).count();
long tabsClosed = herdr.calls.stream().filter(c -> c.method().equals("tab.close")).count();
assertEquals(LeadLauncher.RELAUNCH_ATTEMPTS, tabsCreated);
assertEquals(tabsCreated, tabsClosed, "every tab this method created must be closed — no leaks");
}
}
@@ -1305,8 +1305,12 @@ class LeadRolloverTest {
int rolls = LeadRollover.OUTCOME_HISTORY_CAP + 50;
String[] tokens = new String[rolls];
for (int i = 0; i < rolls; i++) {
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full #" + i);
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
// a distinct terminal per iteration: confirm() single-flights per lead terminal, so
// reusing LEAD here would refuse every roll after the first instead of building up
// OUTCOME_HISTORY_CAP+50 simultaneously in-flight entries.
String terminal = "term_cap_" + i;
LeadRollover.PendingRollover pending = rollover.open(terminal, "context is full #" + i);
LeadRollover.RollDecision decision = rollover.confirm(terminal, pending.token(), true);
assertTrue(decision.accepted(), "roll #" + i + " should have been approved: " + decision.detail());
tokens[i] = pending.token();
}
@@ -1447,4 +1451,195 @@ class LeadRolloverTest {
assertTrue(status.detail().contains("HerdrException"), "the detail must name the exception "
+ "so an operator reading status() has something to act on: " + status.detail());
}
// ---- fleetd #726 unit 3: confirm() single-flights per lead terminal -----------------------
@Test
@DisplayName("[SINGLE-FLIGHT 1] a second confirm() for the SAME lead terminal is refused with "
+ "ROLL_ALREADY_RUNNING while the first roll is still in flight, continuationRunner ran "
+ "exactly once, and the claim is released once that first roll finishes")
void secondConfirmForSameTerminalIsRefusedWhileFirstRollIsStillInFlight() throws IOException {
FakeHerdr herdr = new FakeHerdr(); // default idle — the held roll WOULD complete once run
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
HoldingRunner runner = new HoldingRunner();
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
fixedClock(clock), runner);
LeadRollover.PendingRollover first = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision firstDecision = rollover.confirm(LEAD, first.token(), true);
assertTrue(firstDecision.accepted(), "expected approval; got: " + firstDecision.reason()
+ " / " + firstDecision.detail());
assertEquals(1, runner.heldCount(), "sanity: the first roll is held, not run yet");
LeadRollover.PendingRollover second = rollover.open(LEAD, "a second request for the same terminal");
LeadRollover.RollDecision secondDecision = rollover.confirm(LEAD, second.token(), true);
assertFalse(secondDecision.accepted());
assertEquals(LeadRollover.RefusalReason.ROLL_ALREADY_RUNNING, secondDecision.reason());
assertTrue(secondDecision.detail().contains(LEAD), "the refusal detail must name the lead "
+ "terminal: " + secondDecision.detail());
assertTrue(secondDecision.detail().contains(first.token()), "the refusal detail must name "
+ "the token holding the claim: " + secondDecision.detail());
assertEquals(1, runner.heldCount(), "the refused confirm() must never have reached "
+ "continuationRunner — it ran exactly once, for the first roll only");
runner.runNext(); // let the first (and only) held roll finish
assertEquals(2, promptCallCount(herdr), "continuationRunner ran exactly once: /clear then "
+ "bootstrapText, for the first roll only");
// The claim must have been released once the first roll's continuation finished — a
// fresh request for the SAME terminal is now approved.
LeadRollover.PendingRollover third = rollover.open(LEAD, "retry after the first roll finished");
LeadRollover.RollDecision thirdDecision = rollover.confirm(LEAD, third.token(), true);
assertTrue(thirdDecision.accepted(), "the claim must have been released once the first "
+ "roll finished: " + thirdDecision.reason() + " / " + thirdDecision.detail());
}
@Test
@DisplayName("[SINGLE-FLIGHT 2] two DIFFERENT lead terminals confirm without either being "
+ "refused, and the continuation runs once per terminal")
void differentLeadTerminalsConfirmIndependently() throws IOException {
FakeHerdr herdr = new FakeHerdr(); // default idle — both rolls complete
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover a = rollover.open(LEAD, "context is full");
LeadRollover.PendingRollover b = rollover.open(OTHER_LEAD, "context is full too");
LeadRollover.RollDecision aDecision = rollover.confirm(LEAD, a.token(), true);
LeadRollover.RollDecision bDecision = rollover.confirm(OTHER_LEAD, b.token(), true);
assertTrue(aDecision.accepted(), "expected approval; got: " + aDecision.reason() + " / " + aDecision.detail());
assertTrue(bDecision.accepted(), "expected approval; got: " + bDecision.reason() + " / " + bDecision.detail());
assertEquals(4, promptCallCount(herdr), "both rolls ran to completion — /clear and "
+ "bootstrapText, once per terminal");
}
@Test
@DisplayName("[SINGLE-FLIGHT 3] a confirm() refused for OPERATOR_NOT_CONFIRMED does not take "
+ "the single-flight claim — a later valid confirm on the same terminal is still approved")
void refusedConfirmDoesNotTakeTheClaim() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(herdr, cfg(handover.toString(), true), fixedClock(clock));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision refused = rollover.confirm(LEAD, pending.token(), false);
assertFalse(refused.accepted());
assertEquals(LeadRollover.RefusalReason.OPERATOR_NOT_CONFIRMED, refused.reason());
// the same token stays pending after a gate refusal — retrying it with operator
// confirmation must succeed, proving the earlier refusal never took the claim.
LeadRollover.RollDecision approved = rollover.confirm(LEAD, pending.token(), true);
assertTrue(approved.accepted(), "a confirm() refused for an existing gate must never have "
+ "taken the single-flight claim: " + approved.reason() + " / " + approved.detail());
}
@Test
@DisplayName("[SINGLE-FLIGHT 4] a roll whose continuation throws a RuntimeException still "
+ "releases the claim — a fresh confirm() on that terminal is approved afterwards")
void throwingContinuationStillReleasesTheClaim() throws IOException {
FakeHerdr fake = new FakeHerdr(); // default idle — the turn-settle wait passes immediately
HerdrClient throwsOnClear = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) throws HerdrException {
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
throw new HerdrException("simulated herdr transport failure sending /clear");
}
return fake.call(method, params);
}
@Override
public void close() {
fake.close();
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRollover(throwsOnClear, cfg(handover.toString()), fixedClock(clock));
LeadRollover.PendingRollover first = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, first.token(), true);
assertTrue(decision.accepted(), "every synchronous gate passes; the throw happens only "
+ "inside the deferred continuation, which this test's synchronous runner has "
+ "already run to completion by the time confirm() returns");
LeadRollover.RollStatus status = rollover.status(first.token());
assertEquals(LeadRollover.RollState.FAILED, status.state(), "sanity: the continuation threw "
+ "and left a terminal FAILED outcome: " + status.detail());
LeadRollover.PendingRollover second = rollover.open(LEAD, "retry after the failure");
LeadRollover.RollDecision retryDecision = rollover.confirm(LEAD, second.token(), true);
assertTrue(retryDecision.accepted(), "the claim must have been released even though the "
+ "continuation threw — a release only on the success path would leave this lead "
+ "terminal unrollable forever after one failure: " + retryDecision.reason() + " / "
+ retryDecision.detail());
}
@Test
@DisplayName("[SINGLE-FLIGHT 5] cancel(token) after a successful confirm(token) returns false "
+ "and does not release the claim held by that roll")
void cancelAfterConfirmDoesNotReleaseTheClaim() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
HoldingRunner runner = new HoldingRunner();
LeadRollover rollover = newRolloverWithHoldingRunner(herdr, cfg(handover.toString()),
fixedClock(clock), runner);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
assertFalse(rollover.cancel(pending.token()), "confirm() already removed the token from "
+ "pending, so cancel() must find nothing for it");
LeadRollover.PendingRollover another = rollover.open(LEAD,
"a second request while the first roll is still held");
LeadRollover.RollDecision anotherDecision = rollover.confirm(LEAD, another.token(), true);
assertFalse(anotherDecision.accepted(), "cancel() on the already-confirmed token must not "
+ "have released the claim held by that roll");
assertEquals(LeadRollover.RefusalReason.ROLL_ALREADY_RUNNING, anotherDecision.reason());
}
@Test
@DisplayName("[SINGLE-FLIGHT 6] a continuationRunner that rejects the hand-off still releases "
+ "the claim and leaves a terminal FAILED outcome, not a stuck IN_PROGRESS")
void continuationRunnerThatRejectsTheHandOffStillReleasesTheClaim() throws IOException {
FakeHerdr herdr = new FakeHerdr();
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
AgentControl agents = new AgentControl(herdr);
// A continuationRunner that rejects every hand-off, standing in for e.g. a bounded
// executor's RejectedExecutionException — the continuation never runs, so runRollover's
// own finally never gets a chance to release the claim either.
LeadRollover rollover = new LeadRollover(agents, () -> cfg(handover.toString()), _ -> null,
fixedClock(clock), () -> { },
r -> { throw new RuntimeException("simulated continuationRunner rejection"); });
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
assertTrue(decision.accepted(), "every gate passed before the hand-off itself rejected — "
+ "confirm() treats that rejection the same way it treats a throw INSIDE the "
+ "continuation: logged and surfaced only through status(), never as a refusal here: "
+ decision.reason() + " / " + decision.detail());
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.FAILED, status.state(), "a rejected hand-off must leave "
+ "a TERMINAL outcome, never a stuck IN_PROGRESS — status() would otherwise have no "
+ "way to tell a roll that never started from one still genuinely running: got "
+ status.state() + " / " + status.detail());
assertNotEquals(LeadRollover.RollState.IN_PROGRESS, status.state());
LeadRollover.PendingRollover retry = rollover.open(LEAD, "retry after the rejected hand-off");
LeadRollover.RollDecision retryDecision = rollover.confirm(LEAD, retry.token(), true);
assertTrue(retryDecision.accepted(), "the claim must have been released even though "
+ "continuationRunner itself threw before the continuation ever ran — otherwise "
+ "this lead terminal could never be rolled again: " + retryDecision.reason() + " / "
+ retryDecision.detail());
}
}
@@ -63,6 +63,17 @@ class FleetMcpAuthzTest {
/** A fully wired FleetMcp on fakes — constructing it is itself part of what is under test. */
private FleetMcp mcp(boolean enforce) {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
return mcp(enforce, CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)));
}
/**
* As {@link #mcp(boolean)}, with an explicit {@link CallerResolver} — so a test can wire known
* leads/collaborators and drive {@code denyFor}'s real {@code knownLeadOrCollaborator()}
* classifier instead of the default empty one.
*/
private FleetMcp mcp(boolean enforce, CallerResolver callers) {
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
@@ -83,9 +94,7 @@ class FleetMcpAuthzTest {
// to be omitted to reach "legacy" is now always real, and AuthorizationMode is the
// separate, explicit choice that governs enforcement.
mcp = new FleetMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)),
new PrimaryRegistry(null), callers,
enforce ? FleetMcp.AuthorizationMode.ENFORCED : FleetMcp.AuthorizationMode.UNENFORCED,
metrics, FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), null, FleetMcp.OutageSource.none(),
@@ -97,6 +106,7 @@ class FleetMcpAuthzTest {
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal COLLABORATOR = Principal.collaborator("ops", "term_collab", 600);
// --- the table, enforced on THIS path too ---------------------------------------------------
@@ -156,6 +166,45 @@ class FleetMcpAuthzTest {
}
}
/**
* fleetd #669 Unit A: {@code SEND} is split into three actions ({@link Authz.Action#SEND},
* {@link Authz.Action#COORD_SEND}, {@link Authz.Action#ANSWER}), each carrying the same grant
* the one undivided action gave. An architect holds all three, exactly as it held the one.
*/
@Test
void anArchitectMayUseAllThreeSendShapesOverMcp() {
FleetMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
Authz.Action.ANSWER}) {
assertNull(m.denyFor(ARCH_DESIGN, a, "term_a"),
a + " carries the same grant the undivided SEND action gave an architect");
}
}
/** The other half of the same split: a worker is excluded from all three, as it was from one. */
@Test
void aWorkerMayNotUseAnySendShapeOverMcp() {
FleetMcp m = mcp(true);
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
Authz.Action.ANSWER}) {
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
assertNotNull(denied, a + " must stay refused to a worker");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
/**
* fleetd #669 Unit A / #678: {@code TASK_READ} (ticket polling, session status) is split out of
* the roster-only {@code READ}, carrying forward the grant the undivided action gave. A worker
* still has both — it never gained or lost anything by the split.
*/
@Test
void aWorkerKeepsBothReadActionsAfterTheSplit() {
FleetMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, Authz.Action.READ, null));
assertNull(m.denyFor(WORKER_A, Authz.Action.TASK_READ, null));
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
FleetMcp m = mcp(true);
@@ -187,6 +236,39 @@ class FleetMcpAuthzTest {
"the caller IS authenticated — it is just not the right role");
}
/**
* {@code denyFor} passes the real production classifier, not a test-supplied one — no
* terminal is recognised as a configured lead or collaborator, so a collaborator's SEND is
* refused over MCP.
*/
@Test
void aCollaboratorMayNotSendOverMcpWithTheRealProductionClassifier() {
FleetMcp m = mcp(true);
McpSchema.CallToolResult denied = m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_lead");
assertNotNull(denied, "no terminal is recognised as a lead or collaborator yet");
assertTrue(denied.isError());
}
/**
* fleetd #669 Unit D: wires a real {@link CallerResolver} with a known lead and a known
* collaborator tab, and leaves a spawned member's own terminal recognised by neither map — so a
* collaborator's SEND reaches both named peers and is refused for the spawned member's terminal,
* over MCP's {@code denyFor}.
*/
@Test
void aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminal() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> 999_999);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> Map.of("term_lead_known", "lead-x"), new MemberRegistry(null),
t -> null, () -> Map.of("term_collab_known", "ops2"));
FleetMcp m = mcp(true, callers);
assertNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_lead_known"));
assertNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_collab_known"));
assertNotNull(m.denyFor(COLLABORATOR, Authz.Action.SEND, "term_a"),
"a spawned member's own terminal must stay unreachable, even once the classifier is real");
}
@Test
void theLegacyConstructorLeavesTheGateOpen() {
// The 22 pre-existing FleetMcpTest cases rely on no authorization being enforced.
@@ -273,10 +355,10 @@ class FleetMcpAuthzTest {
+ "the anchors have drifted, this test is not testing what it claims to");
Pattern trailingArg = Pattern.compile(
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;",
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*[,)]",
Pattern.DOTALL);
Matcher m = trailingArg.matcher(handlerBlock);
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the "
assertTrue(m.find(), "could not locate listFleet(...)'s coordinator boolean argument in the "
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
String trailing = m.group(1);
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
@@ -284,6 +366,197 @@ class FleetMcpAuthzTest {
+ "calling, not pass a literal boolean -- found: " + trailing);
}
// --- fleetd #703: who may see fleet_list's collaborators array -------------------------------
/**
* fleetd #703: {@link FleetMcp#collaboratorsVisibleTo} is the whole policy decision for
* {@code fleet_list}'s {@code collaborators} array. Visible to exactly the roles that may
* {@code SEND} to a named peer -- the primary, an architect, and a collaborator itself -- never
* a worker, which holds {@code READ} but can never {@code SEND} to a collaborator, and never an
* anonymous caller.
*/
@Test
void onlyPrimaryArchitectAndCollaboratorMaySeeTheCollaboratorsArray() {
assertTrue(FleetMcp.collaboratorsVisibleTo(PRIMARY), "the primary must see the collaborators array");
assertTrue(FleetMcp.collaboratorsVisibleTo(ARCH_DESIGN), "an architect must see the collaborators array");
assertTrue(FleetMcp.collaboratorsVisibleTo(COLLABORATOR), "a collaborator must see its own peer roster");
assertFalse(FleetMcp.collaboratorsVisibleTo(WORKER_A),
"a worker holds READ but can never SEND to a collaborator, so it must not see the array");
assertFalse(FleetMcp.collaboratorsVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* fleetd #703, same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}:
* the predicate above can be perfectly correct while the one production call site never asks it.
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
* {@code listFleet(...)} call both threads {@code callers.collaborators()} into the payload and
* asks {@code collaboratorsVisibleTo(principal(exchange))} for the visibility flag, rather than a
* literal boolean or an empty map.
*/
@Test
void theFleetListHandlerActuallyConsultsCollaboratorsVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
// fails, the anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("callers.collaborators()"),
"the fleet_list handler must thread callers.collaborators() into listFleet(...), not an "
+ "empty or literal map -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("collaboratorsVisibleTo(principal(exchange))"),
"the fleet_list handler must ask collaboratorsVisibleTo(principal(exchange)) who is "
+ "calling, not pass a literal boolean -- block: " + handlerBlock);
}
// --- who may see fleet_list's leads and members arrays ---------------------------------------
/**
* {@link FleetMcp#leadsVisibleTo} is the whole policy decision for {@code fleet_list}'s
* {@code leads} array: visible to exactly the roles that may {@code SEND} to a lead -- the
* primary, an architect, and a collaborator -- never a worker, which holds {@code READ} but
* can never {@code SEND} at all, and never an anonymous caller.
*/
@Test
void primaryArchitectAndCollaboratorMaySeeTheLeadsArray() {
assertTrue(FleetMcp.leadsVisibleTo(PRIMARY), "the primary must see the leads array");
assertTrue(FleetMcp.leadsVisibleTo(ARCH_DESIGN), "an architect must see the leads array");
assertTrue(FleetMcp.leadsVisibleTo(COLLABORATOR),
"a collaborator may SEND to a lead, so it must see the leads array to learn where");
assertFalse(FleetMcp.leadsVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the leads array");
assertFalse(FleetMcp.leadsVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* As {@link #primaryArchitectAndCollaboratorMaySeeTheLeadsArray}, for the {@code members}
* array -- but a collaborator may {@code SEND} only to a lead or another collaborator, never
* to a spawned member, so it must not see this one.
*/
@Test
void onlyPrimaryAndArchitectMaySeeTheMembersArray() {
assertTrue(FleetMcp.membersVisibleTo(PRIMARY), "the primary must see the members array");
assertTrue(FleetMcp.membersVisibleTo(ARCH_DESIGN), "an architect must see the members array");
assertFalse(FleetMcp.membersVisibleTo(WORKER_A),
"a worker holds READ but can never SEND, so it must not see the members array");
assertFalse(FleetMcp.membersVisibleTo(COLLABORATOR),
"a collaborator may SEND to a lead, never to a spawned member, so it must not see the members array");
assertFalse(FleetMcp.membersVisibleTo(ANON), "authenticated as nothing must not see it either");
}
/**
* Same reasoning as {@link #theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}: the
* predicate above can be perfectly correct while the one production call site never asks it.
* This reads {@code FleetMcp.java}'s own source and asserts the {@code fleet_list} handler's
* {@code listFleet(...)} call asks {@code leadsVisibleTo(principal(exchange))} for the leads
* visibility flag, rather than a literal boolean.
*/
@Test
void theFleetListHandlerActuallyConsultsLeadsVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("leadsVisibleTo(principal(exchange))"),
"the fleet_list handler must ask leadsVisibleTo(principal(exchange)) who is calling, "
+ "not pass a literal boolean -- block: " + handlerBlock);
}
/** As {@link #theFleetListHandlerActuallyConsultsLeadsVisibleTo}, for {@code membersVisibleTo}. */
@Test
void theFleetListHandlerActuallyConsultsMembersVisibleTo() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("listHandler =");
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("stopHandler =", start);
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
assertTrue(handlerBlock.contains("listFleet("),
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
+ "the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("membersVisibleTo(principal(exchange))"),
"the fleet_list handler must ask membersVisibleTo(principal(exchange)) who is calling, "
+ "not pass a literal boolean -- block: " + handlerBlock);
}
/**
* {@code fleet_poll{ticket}} must thread the calling connection's owner key into
* {@link MessageService#poll(String, String)}, so a worker cannot read a ticket a different
* session created.
*/
@Test
void theFleetPollHandlerActuallyThreadsCallerOwnerIntoPoll() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("pollHandler =");
assertTrue(start >= 0, "could not find the fleet_poll handler (pollHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("ackHandler =", start);
assertTrue(end > start, "could not find the handler declared after pollHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call poll(...) -- if this fails, the anchors
// above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("poll(messages,"),
"control failed: the scraped pollHandler block contains no poll(messages, ...) call "
+ "at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("principal(exchange).ownerKey()"),
"the fleet_poll handler must thread principal(exchange).ownerKey() into poll(...), not omit "
+ "it or pass a literal null -- block: " + handlerBlock);
}
/**
* {@code fleet_status} must thread the calling connection's owner key into
* {@link FleetMcp#status(MessageService, String, String)}, so a caller that did not create a
* worker's open delegation cannot read its pending question through the status handler either.
*/
@Test
void theFleetStatusHandlerActuallyThreadsCallerOwnerIntoStatus() throws Exception {
String source = Files.readString(MCP_SOURCE);
int start = source.indexOf("statusHandler =");
assertTrue(start >= 0, "could not find the fleet_status handler (statusHandler) in " + MCP_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("pollHandler =", start);
assertTrue(end > start, "could not find the handler declared after statusHandler to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call status(...) -- if this fails, the anchors
// above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("status(messages,"),
"control failed: the scraped statusHandler block contains no status(messages, ...) "
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("principal(exchange).ownerKey()"),
"the fleet_status handler must thread principal(exchange).ownerKey() into status(...), not "
+ "omit it or pass a literal null -- block: " + handlerBlock);
}
// --- which action each tool hands the gate (fleetd #272) ------------------------------------
/**
@@ -297,35 +570,53 @@ class FleetMcpAuthzTest {
* the whole time the defect was live -- the table was right, the action fed to it was wrong.
*/
@Test
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
void pollingByTargetIsADrainAndPollingByTicketIsATaskRead() {
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
"poll by target removes the replies — that is a drain, not an observation");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
"poll by ticket changes nothing");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, null),
"poll by ticket changes nothing, but is not the roster-only READ action");
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(" ", null),
"a blank target is an absent target");
}
/**
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
* priority over target when both happen to be present — it addresses a different inbox entirely.
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ} or {@link
* Authz.Action#TASK_READ}, even though this branch also consumes nothing. A lead-to-lead body
* is not the roster and not a ticket/status read, so folding this branch into either would let
* any worker or architect read every peer lead's mail in full. coordId also takes priority over
* target when both happen to be present — it addresses a different inbox entirely.
*/
@Test
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
void pollingByCoordIdIsACoordReadNeverAPlainOrTaskRead() {
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
"reading held peer mail must not be mapped to the everyone-readable READ action");
"reading held peer mail must not be mapped to a widely-readable action");
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
"a blank target must not fall through to READ/DRAIN when coordId is present");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, " "),
"a blank coordId is an absent coordId, same as target/ticket");
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
"coordId takes priority over target — this is a different inbox, not a drain");
}
/**
* {@code fleet_send} is three call shapes behind one tool name, exactly as {@code fleet_poll}
* is (fleetd #669 Unit A). {@link FleetMcp#sendAction} picks the action from the arguments, not
* the handler, for the same reason {@link FleetMcp#pollAction} does: a test can assert the
* mapping the handler actually uses.
*/
@Test
void sendMapsToThreeDifferentActionsByItsArguments() {
assertEquals(Authz.Action.SEND, FleetMcp.sendAction(null, null),
"a plain delivery, with neither coordId nor turnId, is a local SEND");
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", null),
"coordId addresses a peer lead over the coordination broker");
assertEquals(Authz.Action.ANSWER, FleetMcp.sendAction(null, "turn-1"),
"turnId resolves a worker's blocked question");
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", "turn-1"),
"coordId takes priority over turnId, mirroring sendToLead's own mutual-exclusion check");
}
@Test
void everyRegisteredToolHasItsHandlerActionPinned() {
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
@@ -344,16 +635,22 @@ class FleetMcpAuthzTest {
() -> tool + " is registered but has no pinned authorization action"));
assertEquals(Authz.Action.SEND, FleetMcp.toolAction("fleet_send", Map.of()));
assertEquals(Authz.Action.SEND,
FleetMcp.toolAction("fleet_send", Map.of("sessionId", "term_a", "content", "hi")));
assertEquals(Authz.Action.COORD_SEND,
FleetMcp.toolAction("fleet_send", Map.of("coordId", "mac-opus", "content", "hi")));
assertEquals(Authz.Action.ANSWER,
FleetMcp.toolAction("fleet_send", Map.of("turnId", "turn-1", "content", "hi")));
assertEquals(Authz.Action.REPLY, FleetMcp.toolAction("fleet_reply", Map.of()));
assertEquals(Authz.Action.ASK, FleetMcp.toolAction("fleet_ask", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_status", Map.of()));
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_status", Map.of()));
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_ack", Map.of()));
assertEquals(Authz.Action.SPAWN, FleetMcp.toolAction("fleet_spawn", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_list", Map.of()));
assertEquals(Authz.Action.STOP, FleetMcp.toolAction("fleet_stop", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_profiles", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
assertEquals(Authz.Action.COORD_READ,
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
@@ -121,6 +121,7 @@ class FleetMcpLeadContextGaugeWiringTest {
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), contextGauge, leadConfigDirs,
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, FleetMcp.CoordinationSource.none(), false);
Map.of(LEAD_TERMINAL, LEAD_NAME), LEAD_TERMINAL, Map.of(), false,
FleetMcp.CoordinationSource.none(), false, true, true);
}
}
@@ -10,6 +10,7 @@ import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.inject.LoopWatchdog;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
@@ -76,7 +77,7 @@ class FleetMcpTest {
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles));
() -> FleetMcp.send(messages, target, "hi", 4000L, null, profiles, null));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -102,7 +103,7 @@ class FleetMcpTest {
void sendThenReplyRoundTrips() throws Exception {
// fleet_send blocks; fleet_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of()));
() -> FleetMcp.send(messages, "term_a", "review this", 4000L, null, Set.of(), null));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
@@ -180,7 +181,7 @@ class FleetMcpTest {
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
@@ -278,7 +279,7 @@ class FleetMcpTest {
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L));
() -> FleetMcp.answer(messages, firstTurnId, "config.yaml", 5000L, null));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -287,6 +288,91 @@ class FleetMcpTest {
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
// --- fleetd #715: fleet_send{turnId} is gated on the caller that owns the turn -------------
/**
* A blocking {@code fleet_send} from one caller opens the turn; a {@code fleet_send{turnId}}
* from a different caller is refused as an error, and the real owner's answer still succeeds.
*/
@Test
void aDifferentCallersMcpAnswerIsRefusedForABlockingSendButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "do X", 5000L, null, Set.of(), "term_owner"));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
McpSchema.CallToolResult question = send.get(5, TimeUnit.SECONDS);
assertTrue(textOf(question).contains("[question]"), textOf(question));
String questionText = textOf(question);
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L, "term_attacker");
assertTrue(hijacked.isError(), "a caller that did not open this turn must get an error, not an answer");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, "term_owner"));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, T, Role.WORKER, "done");
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
}
/**
* Same hijack and control through the fire-and-poll ({@code sendAsync}) path: the owner comes
* from the resolved principal, not from a caller argument threaded through a live call.
*/
@Test
void aDifferentCallersMcpAnswerIsRefusedForAnAsyncSendButTheRealOwnerSucceeds() throws Exception {
Principal owner = Principal.worker("term_owner", 1);
McpSchema.CallToolResult accepted =
FleetMcp.sendAsync(messages, T, "do it", null, Set.of(), owner);
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> FleetMcp.ask(messages, T, "which config?", 5000L));
MessageService.TaskView asking = messages.poll(ticket);
deadline = System.currentTimeMillis() + 3000;
while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
asking = messages.poll(ticket);
}
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
McpSchema.CallToolResult hijacked = FleetMcp.answer(messages, turnId, "evil.yaml", 500L,
"worker:term_attacker");
assertTrue(hijacked.isError(), "a caller that did not create this delegation must get an error");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, owner.ownerKey()));
assertEquals("config.yaml", textOf(ask.get(5, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, T, Role.WORKER, "done");
assertEquals("done", textOf(answer.get(5, TimeUnit.SECONDS)));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = FleetMcp.poll(messages, "task-999", null);
@@ -296,7 +382,7 @@ class FleetMcpTest {
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of());
McpSchema.CallToolResult res = FleetMcp.send(messages, "term_a", "hi", 120L, null, Set.of(), null);
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@@ -312,7 +398,7 @@ class FleetMcpTest {
void sendTimesOutWithAnUnconfirmedNoteNotARetryInvitation() throws Exception {
herdr.agentSendFailsWith("send_failed");
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of()));
() -> FleetMcp.send(messages, T, "hi", 150L, null, Set.of(), null));
long deadline = System.currentTimeMillis() + 2000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -332,13 +418,13 @@ class FleetMcpTest {
@Test
void sendRejectsMissingArgs() {
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of()).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of()).isError());
assertTrue(FleetMcp.send(messages, null, "hi", null, null, Set.of(), null).isError());
assertTrue(FleetMcp.send(messages, "term_a", " ", null, null, Set.of(), null).isError());
}
@Test
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"));
McpSchema.CallToolResult blocking = FleetMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"), null);
McpSchema.CallToolResult async = FleetMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
assertTrue(blocking.isError());
@@ -357,7 +443,7 @@ class FleetMcpTest {
assertSendRoundTrips("term_live_member", profiles);
// A herdr-owned pane outside the bridge roster cannot be classified at accept time.
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles);
McpSchema.CallToolResult result = FleetMcp.send(messages, "external-pane", "hi", 10L, null, profiles, null);
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
}
@@ -449,7 +535,7 @@ class FleetMcpTest {
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of()));
() -> FleetMcp.send(messages, "term_a", "do X", 5000L, null, Set.of(), null));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -471,7 +557,7 @@ class FleetMcpTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L));
() -> FleetMcp.answer(messages, turnId, "config.yaml", 5000L, null));
// The worker's ask returns the answer — it resumes the same turn.
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
@@ -497,7 +583,7 @@ class FleetMcpTest {
@Test
void answerToAStaleTurnIsAnError() {
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L);
McpSchema.CallToolResult res = FleetMcp.answer(messages, "term_a#999", "too late", 500L, null);
assertTrue(res.isError());
assertTrue(textOf(res).contains("no longer open"), textOf(res));
}
@@ -691,7 +777,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res);
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
@@ -720,7 +806,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res);
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
@@ -779,19 +865,82 @@ class FleetMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary();
Principal worker = Principal.worker("term_a", 1);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), worker.isPrimary(),
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out);
assertTrue(out.contains("\"members\""), out);
assertFalse(out.contains("\"leads\""), "a worker must never see the leads key at all: " + out);
assertFalse(out.contains("\"members\""), "a worker must never see the members key at all: " + out);
assertTrue(out.contains("\"healthCoverage\""), "the rest of the result must still be present: " + out);
}
/**
* A worker's result must carry neither the {@code leads} nor the {@code members} key, and no
* fragment of either row leaks even though both are fully populated for this call -- a key
* check alone would pass on an implementation that still built the rows and only renamed or
* nested them.
*/
@Test
void listLeaksNoLeadOrMemberRowFragmentToAWorker() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
new WorktreeRequest("cb-304", null));
Principal worker = Principal.worker("term_a", 1);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), worker.isPrimary(),
FleetMcp.leadsVisibleTo(worker), FleetMcp.membersVisibleTo(worker));
String out = textOf(res);
assertFalse(out.contains("\"leads\""), out);
assertFalse(out.contains("\"members\""), out);
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a worker: " + out);
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a worker: " + out);
assertFalse(out.contains("mac-opus"), "no lead row fragment may leak to a worker: " + out);
assertTrue(out.contains("\"healthCoverage\""), out);
}
/**
* A collaborator may {@code SEND} to a lead, so it must see the {@code leads} array -- the
* only place {@code fleet_whoami} does not already give it a lead's address. It may never
* {@code SEND} to a spawned member, so the {@code members} key must stay absent for it, with
* no fragment of a populated member row leaking either.
*/
@Test
void listShowsLeadsButNotMembersToACollaborator() {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_owner",
new WorktreeRequest("cb-304", null));
Principal collaborator = Principal.collaborator("ops", "term_collab", 600);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of("term_lead_x", "mac-opus"), "", FleetMcp.CoordinationSource.none(), collaborator.isPrimary(),
FleetMcp.leadsVisibleTo(collaborator), FleetMcp.membersVisibleTo(collaborator));
String out = textOf(res);
assertTrue(out.contains("\"leads\""), "a collaborator must see the leads array: " + out);
assertTrue(out.contains("mac-opus"), "a collaborator must see the lead's name/address: " + out);
assertFalse(out.contains("\"members\""), "a collaborator must never see the members key: " + out);
assertFalse(out.contains(s.terminalId()), "no member row fragment may leak to a collaborator: " + out);
assertFalse(out.contains("term_owner"), "no member owner fragment may leak to a collaborator: " + out);
assertTrue(out.contains("\"healthCoverage\""), out);
}
@@ -807,16 +956,19 @@ class FleetMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary();
Principal architect = Principal.architect("lead-designer", "term_design", 400);
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), architect.isPrimary(),
FleetMcp.leadsVisibleTo(architect), FleetMcp.membersVisibleTo(architect));
String out = textOf(res);
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
assertTrue(out.contains("\"leads\""), "an architect must still see the leads array: " + out);
assertTrue(out.contains("\"members\""), "an architect must still see the members array: " + out);
}
/**
@@ -841,7 +993,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true));
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true));
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
@@ -850,6 +1002,8 @@ class FleetMcpTest {
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"leads\""), "the primary must still see the leads array: " + gatedAsPrimary);
assertTrue(gatedAsPrimary.contains("\"members\""), "the primary must still see the members array: " + gatedAsPrimary);
}
@Test
@@ -868,7 +1022,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res);
assertTrue(out.contains("\"msgId\":\"m1\""), out);
@@ -879,6 +1033,83 @@ class FleetMcpTest {
assertTrue(out.contains("x".repeat(80) + "…"), "expected an 80-char preview with an ellipsis: " + out);
}
// --- fleetd #703: fleet_list's collaborators array -------------------------------------------
/** Calls the canonical {@code listFleet} overload directly, so a test can set the collaborators
* payload and its visibility independently of a real {@code Principal} / MCP exchange. */
private static McpSchema.CallToolResult listFleetWithCollaborators(FakeHerdr h,
Map<String, String> collaborators, boolean collaboratorsVisible) {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
return FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.LoopHealthSource.none(), FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(),
FleetMcp.LeadSeatSource.none(), new LeadContextGauge(), FleetMcp.LeadConfigDirSource.none(),
Map.of(), "", collaborators, collaboratorsVisible,
FleetMcp.CoordinationSource.none(), false, true, true);
}
/**
* fleetd #703 acceptance A: a visible caller with one configured collaborator gets a
* {@code collaborators} row whose key ({@code CallerResolver.collaborators()}'s
* {@code terminal_id -> name} entry) lands as that row's {@code sessionId}, and {@code leads}/
* {@code members} are unaffected by the new key.
*/
@Test
void listReportsACollaboratorsRowKeyedByTheCollaboratorsTerminalId() {
FakeHerdr h = new FakeHerdr();
String out = textOf(listFleetWithCollaborators(h, Map.of("term_collab", "ops"), true));
assertTrue(out.contains("\"collaborators\":["), out);
assertTrue(out.contains("\"name\":\"ops\""), out);
assertTrue(out.contains("\"sessionId\":\"term_collab\""), out);
assertTrue(out.contains("\"leads\":[]"), out);
assertTrue(out.contains("\"members\":[]"), out);
}
/**
* fleetd #703 acceptance A, control half: with no collaborator configured, the key is absent
* (caller cannot see it) or an empty array (caller can), and {@code leads}/{@code members} are
* unchanged either way.
*/
@Test
void listOmitsOrEmptiesCollaboratorsWhenNoneAreConfigured() {
FakeHerdr h = new FakeHerdr();
String visibleButEmpty = textOf(listFleetWithCollaborators(h, Map.of(), true));
assertTrue(visibleButEmpty.contains("\"collaborators\":[]"), visibleButEmpty);
assertTrue(visibleButEmpty.contains("\"leads\":[]"), visibleButEmpty);
assertTrue(visibleButEmpty.contains("\"members\":[]"), visibleButEmpty);
String notVisible = textOf(listFleetWithCollaborators(h, Map.of(), false));
assertFalse(notVisible.contains("\"collaborators\""), notVisible);
assertTrue(notVisible.contains("\"leads\":[]"), notVisible);
assertTrue(notVisible.contains("\"members\":[]"), notVisible);
}
/**
* fleetd #703 acceptance B: a worker must not see the {@code collaborators} array at all, while
* an architect -- a role that also holds READ, same as a worker -- does see it. Both halves are
* asserted: a test that only checked the worker-hidden half would pass even if the feature were
* never wired up for anyone.
*/
@Test
void listOmitsCollaboratorsForAWorkerAndIncludesThemForAnArchitect() {
FakeHerdr h = new FakeHerdr();
Map<String, String> collaborators = Map.of("term_collab", "ops");
String asWorker = textOf(listFleetWithCollaborators(h, collaborators,
FleetMcp.collaboratorsVisibleTo(Principal.worker("term_w", 1))));
assertFalse(asWorker.contains("\"collaborators\""),
"a worker must not see the collaborators array: " + asWorker);
String asArchitect = textOf(listFleetWithCollaborators(h, collaborators,
FleetMcp.collaboratorsVisibleTo(Principal.architect("design", "term_arch", 2))));
assertTrue(asArchitect.contains("\"collaborators\":["),
"an architect must see the collaborators array: " + asArchitect);
assertTrue(asArchitect.contains("\"sessionId\":\"term_collab\""), asArchitect);
}
/**
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
@@ -900,7 +1131,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res);
assertTrue(out.contains("\"pending\":0"), out);
@@ -929,7 +1160,7 @@ class FleetMcpTest {
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true, true, true);
String out = textOf(res);
assertTrue(out.contains("\"heldDurable\":false"),
@@ -998,7 +1229,7 @@ class FleetMcpTest {
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true);
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true, true, true);
String out = textOf(res);
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
@@ -1691,8 +1922,8 @@ class FleetMcpTest {
Principal architect = Principal.architect("lead-designer", "term_design", 400);
MemberPresence presence = new MemberPresence();
FleetMcp.markSpawnedMemberPresent(worker, presence);
FleetMcp.markSpawnedMemberPresent(architect, presence);
FleetMcp.markTrackedCallerPresent(worker, presence);
FleetMcp.markTrackedCallerPresent(architect, presence);
assertTrue(presence.isPresent("term_worker"));
assertTrue(presence.isPresent("term_design"));
@@ -1703,25 +1934,40 @@ class FleetMcpTest {
Principal lead = Principal.leader("opus", "term_lead", 100);
MemberPresence presence = new MemberPresence();
FleetMcp.markSpawnedMemberPresent(lead, presence);
FleetMcp.markSpawnedMemberPresent(Principal.anonymous(), presence);
FleetMcp.markTrackedCallerPresent(lead, presence);
FleetMcp.markTrackedCallerPresent(Principal.anonymous(), presence);
assertFalse(presence.isPresent("term_lead"));
}
/**
* The item whose absence would be silent: an observer's own MCP contact must still mark
* presence, or a pane resolving to the unconfigured-pane floor would sit on the injector
* readiness gate forever once something addresses it.
*/
@Test
void anObserverContactMarksPresence() {
Principal observer = Principal.observer("term_observer", 800);
MemberPresence presence = new MemberPresence();
FleetMcp.markTrackedCallerPresent(observer, presence);
assertTrue(presence.isPresent("term_observer"));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
AgentControl blockedAgents = new AgentControl(blocked);
McpSchema.CallToolResult res = FleetMcp.status(
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a");
new MessageService(blockedAgents, new Injector(blockedAgents), rendezvous), "term_a", null);
assertNotEquals(Boolean.TRUE, res.isError());
assertEquals("blocked", textOf(res));
}
/**
* CB-582: a lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll}
* — must also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
* A lead polling {@code fleet_status} on its normal cadence — not {@code fleet_poll} — must
* also see a worker's open async {@code fleet_ask} question, since the reverse-rendezvous
* window it opened with is far shorter than that cadence.
*/
@Test
@@ -1744,7 +1990,7 @@ class FleetMcpTest {
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
McpSchema.CallToolResult res = FleetMcp.status(messages, T);
McpSchema.CallToolResult res = FleetMcp.status(messages, T, null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
@@ -1755,7 +2001,66 @@ class FleetMcpTest {
// Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000));
() -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve(T, "done"));
answer.get(5, TimeUnit.SECONDS);
}
/**
* {@code fleet_status}'s pending-ask block (the question, its {@code turnId} and its ticket)
* is shown only to the caller whose owner key created the delegation, or to the unnamed primary.
* A different caller still sees the base status line, but none of the pending-ask fields.
*/
@Test
void statusGatesThePendingAskFieldsByTheDelegationsCreatorOwner() throws Exception {
Principal creator = Principal.worker("term_creator", 1);
String ticket = messages.sendAsync(T, "task that asks", null, creator);
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String other = textOf(FleetMcp.status(messages, T, "worker:term_other"));
assertTrue(other.startsWith("idle"), "the base status must still be shown: " + other);
assertFalse(other.contains("which config file?"),
"a non-creating caller must not see the question text: " + other);
assertFalse(other.contains(asking.turnId()),
"a non-creating caller must not see the turnId: " + other);
assertFalse(other.contains(ticket),
"a non-creating caller must not see the ticket: " + other);
String creatorStatus = textOf(FleetMcp.status(messages, T, creator.ownerKey()));
assertTrue(creatorStatus.contains("which config file?"),
"the creator must see the question: " + creatorStatus);
assertTrue(creatorStatus.contains(asking.turnId()),
"the creator must see the turnId: " + creatorStatus);
assertTrue(creatorStatus.contains(ticket), "the creator must see the ticket: " + creatorStatus);
String unnamed = textOf(FleetMcp.status(messages, T, null));
assertTrue(unnamed.contains("which config file?"),
"a caller with no terminal (the unnamed primary) must see the question: " + unnamed);
// Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, creator.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
@@ -1837,6 +2142,53 @@ class FleetMcpTest {
assertTrue(out.contains("\"sessionId\":\"term_design\""), out);
}
/**
* A collaborator reports its own role and name, never the {@code leader} key a lead gets —
* {@code role} already reads {@code "collaborator"}, so a {@code leader} key alongside it
* would be self-contradicting. Control: the same call shape fed a named lead must still carry
* {@code leader}, so this is not passing because the key stopped being emitted for everyone.
*/
@Test
void whoamiReportsACollaboratorWithNoLeaderKeyButALeadStillGetsOne() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult collabRes = FleetMcp.whoami(
Principal.collaborator("ops", "term_collab", 700), sessions);
assertNotEquals(Boolean.TRUE, collabRes.isError());
String collabOut = textOf(collabRes);
assertTrue(collabOut.contains("\"role\":\"collaborator\""), collabOut);
assertTrue(collabOut.contains("\"collaborator\":\"ops\""), collabOut);
assertTrue(collabOut.contains("\"sessionId\":\"term_collab\""), collabOut);
assertFalse(collabOut.contains("leader"), collabOut);
McpSchema.CallToolResult leadRes = FleetMcp.whoami(
Principal.leader("opus", "term_lead", 100), sessions);
String leadOut = textOf(leadRes);
assertTrue(leadOut.contains("\"leader\":\"opus\""), leadOut);
}
/**
* An observer reports its own role and pane, never a {@code leader} key. Without an explicit
* branch it would reach the lead branch by elimination and look right only because the
* {@code leader} key is guarded on a non-null name — this pins the branch rather than the
* accident.
*/
@Test
void whoamiReportsAnObserverNotALead() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = FleetMcp.whoami(
Principal.observer("term_observer", 900), sessions);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"observer\""), out);
assertTrue(out.contains("\"sessionId\":\"term_observer\""), out);
assertFalse(out.contains("leader"), out);
}
/**
* CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but
* must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure
@@ -0,0 +1,280 @@
package dev.ltms.fleet.msg;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Stream;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Pins that no file under {@code src/main/java} calls the fail-open
* {@link MessageService#poll(String)} overload. That overload skips the ownership check in
* {@code MessageService}'s {@code ownsTicket} entirely, so a caller of it can read any session's
* ticket. Every production caller must go through {@link MessageService#poll(String, String)}
* and pass a {@code callerOwner} explicitly, even when it is {@code null}.
*
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
* the risk is a future one-word edit at a call site, not a missing overload.
*
* <p>The scan below finds a violation by its receiver, {@code messages.poll(}, rather than the
* bare method name, so it does not mistake {@link java.util.Queue#poll()} for a violation. That
* anchor only covers a {@code MessageService} reached through a variable or field named
* {@code messages}, so {@link #everyMessageServiceDeclarationIsNamedMessages} pins the naming
* convention the anchor depends on: a declaration under any other name would be invisible to the
* scan above, and must turn this second check red instead of passing silently.
*/
class MessageServicePollUsageTest {
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
@Test
void noProductionFileCallsTheSingleArgumentPollOverload() throws IOException {
List<String> violations = new ArrayList<>();
List<String> twoArgSites = new ArrayList<>();
int filesScanned = scanForPollCalls(PRODUCTION_SOURCE, violations, twoArgSites);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
// violations found" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the absence of violations "
+ "below proves nothing");
assertTrue(violations.isEmpty(), "found a call to the fail-open MessageService.poll(String) "
+ "overload, which skips the ownership check entirely -- pass a callerOwner "
+ "explicitly (even if null) through poll(String, String) instead: " + violations);
// CONTROL: the arity parser actually finds the two genuine two-argument call sites (the
// MCP handler in FleetMcp and the REST handler in FleetApp). If this drops, the parser
// itself is broken, not the production code -- a broken parser (or a scan root that
// reaches no real source) must fail loudly here rather than pass vacuously above.
assertEquals(2, twoArgSites.size(), "control failed: expected exactly the two known "
+ "two-argument messages.poll(...) call sites, found: " + twoArgSites);
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetMcp.java")),
"control failed: did not find the FleetMcp.java messages.poll(ticket, callerOwner) "
+ "site among: " + twoArgSites);
assertTrue(twoArgSites.stream().anyMatch(s -> s.contains("FleetApp.java")),
"control failed: did not find the FleetApp.java messages.poll(...) site among: "
+ twoArgSites);
}
/**
* The {@code messages.poll(} anchor above only sees a {@code MessageService} reached through
* a variable, field, or parameter named {@code messages}. This asserts that every such
* declaration under {@code src/main/java} uses that name, so a differently named declaration
* -- invisible to the scan above -- fails loudly here instead of letting that scan pass on a
* call site it never looked at.
*/
@Test
void everyMessageServiceDeclarationIsNamedMessages() throws IOException {
List<String> names = new ArrayList<>();
int filesScanned = scanForDeclarationNames(PRODUCTION_SOURCE, names);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "every
// declaration is named messages" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the result below proves nothing");
// CONTROL: the declaration pattern actually finds real declarations. Zero means the
// pattern is broken, not that every MessageService variable, field, or parameter vanished.
assertTrue(names.size() > 0, "control failed: found zero MessageService declarations under "
+ PRODUCTION_SOURCE + " -- the declaration pattern is broken, update it before "
+ "trusting the naming check below");
List<String> other = names.stream().filter(n -> !n.equals("messages")).distinct().toList();
assertTrue(other.isEmpty(), "found a MessageService declaration not named \"messages\": "
+ other + " -- the messages.poll( scan above only looks for that name, so a call "
+ "through a differently named variable or field is invisible to it; either rename "
+ "the declaration or widen that scan's anchor to cover it");
}
private static int scanForPollCalls(Path root, List<String> violations, List<String> twoArgSites)
throws IOException {
List<Path> files = javaFiles(root);
for (Path file : files) {
scanFileForPollCalls(file, violations, twoArgSites);
}
return files.size();
}
private static int scanForDeclarationNames(Path root, List<String> names) throws IOException {
List<Path> files = javaFiles(root);
Pattern declaration = Pattern.compile("MessageService\\s+([A-Za-z_][A-Za-z0-9_]*)");
for (Path file : files) {
if (file.getFileName().toString().equals("MessageService.java")) {
continue; // the type's own declaration, not a caller holding a reference to it
}
String stripped = stripComments(Files.readString(file));
Matcher m = declaration.matcher(stripped);
while (m.find()) {
int j = m.end();
while (j < stripped.length() && Character.isWhitespace(stripped.charAt(j))) j++;
if (j < stripped.length() && stripped.charAt(j) == '(') {
continue; // a method named like the convention, e.g. "MessageService messages()"
}
names.add(m.group(1));
}
}
return files.size();
}
private static List<Path> javaFiles(Path root) throws IOException {
try (Stream<Path> paths = Files.walk(root)) {
return paths.filter(p -> p.toString().endsWith(".java")).toList();
}
}
/**
* Replaces {@code //} and {@code /* *}{@code /} comment text with nothing, leaving code,
* string/char literals and line breaks untouched -- so a comment that merely mentions
* {@code MessageService} in prose can never be read as a declaration.
*/
private static String stripComments(String source) {
StringBuilder out = new StringBuilder(source.length());
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
out.append(c);
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
if (c == '"') inString = false;
i++;
continue;
}
if (inChar) {
out.append(c);
if (c == '\\' && i + 1 < source.length()) { out.append(source.charAt(i + 1)); i += 2; continue; }
if (c == '\'') inChar = false;
i++;
continue;
}
if (c == '"') { inString = true; out.append(c); i++; continue; }
if (c == '\'') { inChar = true; out.append(c); i++; continue; }
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '/') {
while (i < source.length() && source.charAt(i) != '\n') i++;
continue; // leaves the newline itself for the next iteration to append
}
if (c == '/' && i + 1 < source.length() && source.charAt(i + 1) == '*') {
i += 2;
while (i < source.length() && !(source.charAt(i) == '*' && i + 1 < source.length()
&& source.charAt(i + 1) == '/')) {
if (source.charAt(i) == '\n') out.append('\n');
i++;
}
i += 2;
continue;
}
out.append(c);
i++;
}
return out.toString();
}
private static void scanFileForPollCalls(Path file, List<String> violations, List<String> twoArgSites)
throws IOException {
String source = Files.readString(file);
String needle = "messages.poll(";
int from = 0;
int idx;
while ((idx = source.indexOf(needle, from)) >= 0) {
int argsStart = idx + needle.length();
String args = extractBalancedArgs(source, argsStart, file, idx);
int closeParenIndex = argsStart + args.length();
from = closeParenIndex + 1;
if (args.isBlank()) {
continue; // MessageService has no zero-argument poll() -- nothing to classify
}
String site = file + ":" + lineOf(source, idx);
if (topLevelCommaCount(args) == 0) {
violations.add(site + " -- messages.poll(" + args.trim() + ")");
} else {
twoArgSites.add(site);
}
}
}
/**
* The text between {@code messages.poll(} and its matching close paren: balanced over nested
* calls, and never split by a paren or comma sitting inside a string or char literal.
*/
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
int depth = 1;
boolean inString = false;
boolean inChar = false;
int i = start;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
if (depth == 0) return source.substring(start, i);
}
i++;
}
throw new IllegalStateException(
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
}
/**
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
* separators a human reader would see, not every comma character in the text.
*/
private static int topLevelCommaCount(String args) {
int depth = 0;
int commas = 0;
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < args.length()) {
char c = args.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(' || c == '[' || c == '{') {
depth++;
} else if (c == ')' || c == ']' || c == '}') {
depth--;
} else if (c == ',' && depth == 0) {
commas++;
}
i++;
}
return commas;
}
private static int lineOf(String source, int index) {
int line = 1;
for (int i = 0; i < index; i++) {
if (source.charAt(i) == '\n') line++;
}
return line;
}
}
@@ -4,6 +4,7 @@ import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.FakeHerdr;
@@ -78,7 +79,7 @@ class MessageServiceTest {
}
private CompletableFuture<MessageService.Reply> sendAsync(String content, long timeoutMillis) {
return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis));
return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis, null));
}
private void awaitWaiting() throws InterruptedException {
@@ -188,7 +189,7 @@ class MessageServiceTest {
localRendezvous, new InMemoryReplyInbox());
String brief = "Implement the requested change. ".repeat(20);
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000));
CompletableFuture.supplyAsync(() -> localMessages.send(T, brief, 5000, (String) null));
long deadline = System.currentTimeMillis() + 2000;
while (!localRendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
@@ -341,7 +342,7 @@ class MessageServiceTest {
// The primary answers via fleet_send(turnId); this blocks again for the worker's reply.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
// The worker's ask returns the answer — it resumes the same turn.
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
@@ -377,7 +378,7 @@ class MessageServiceTest {
// The primary answers that one turnId; both asks unblock with the same answer.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
MessageService.AskResult a1 = ask1.get(5, TimeUnit.SECONDS);
MessageService.AskResult a2 = ask2.get(5, TimeUnit.SECONDS);
@@ -472,7 +473,7 @@ class MessageServiceTest {
"the duplicate's own timeout elapses first");
CompletableFuture<MessageService.Reply> answered =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500));
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "fleetd.yaml", 500, null));
MessageService.AskResult a = fresh.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.ANSWERED, a.outcome(),
@@ -507,7 +508,7 @@ class MessageServiceTest {
CompletableFuture<MessageService.Reply> lateAnswer = new CompletableFuture<>();
messages.setAskTimeoutRaceHookForTest(() ->
lateAnswer.complete(messages.answer(turnId, "too late", 500)));
lateAnswer.complete(messages.answer(turnId, "too late", 500, null)));
try {
MessageService.AskResult a = ask.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.AskOutcome.TIMED_OUT, a.outcome());
@@ -539,7 +540,7 @@ class MessageServiceTest {
@Test
void answeringAnUnknownTurnIsStale() {
MessageService.Reply r = messages.answer(T + "#999", "too late", 500);
MessageService.Reply r = messages.answer(T + "#999", "too late", 500, null);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an answer to a turn that never existed (or already lapsed) is stale, not a hang");
}
@@ -572,7 +573,7 @@ class MessageServiceTest {
messages.setAnswerAskLapseRaceHookForTest(() -> rendezvous.answerAsk(turnId, "raced in first"));
try {
MessageService.Reply r = messages.answer(turnId, "too late", 500);
MessageService.Reply r = messages.answer(turnId, "too late", 500, null);
assertEquals(MessageService.Outcome.STALE_TURN, r.outcome(),
"an ask already answered by the race must be seen as lapsed, not double-delivered");
} finally {
@@ -596,7 +597,7 @@ class MessageServiceTest {
void sendTimesOutBeforeDeliveryIsQueuedNotWorking() {
// Nothing ever delivers the message and nothing resolves the send, so the reply future
// times out with delivery still incomplete — the message is still queued for the worker.
MessageService.Reply r = messages.send(T, "never delivered", 50);
MessageService.Reply r = messages.send(T, "never delivered", 50, null);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome(),
"an undelivered send that times out is still queued, not working");
assertNull(r.text());
@@ -605,7 +606,7 @@ class MessageServiceTest {
@Test
void sendTimesOutAfterDeliveryIsStillWorking() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300));
CompletableFuture.supplyAsync(() -> messages.send(T, "do the task", 300, null));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver — the delivered future now completes
injector.onStatus(T, AgentStatus.WORKING); // worker starts but never replies
@@ -630,7 +631,7 @@ class MessageServiceTest {
void sendTimeoutUsesCancellationDeliveredWhenPickupWinsTheRace() {
messages.setTimeoutCancellationRaceHookForTest(() -> injector.onStatus(T, AgentStatus.IDLE));
try {
MessageService.Reply reply = messages.send(T, "race delivery", 50);
MessageService.Reply reply = messages.send(T, "race delivery", 50, null);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, reply.outcome(),
"cancel reporting DELIVERED means the worker received the timed-out message");
@@ -652,7 +653,7 @@ class MessageServiceTest {
void sendTimesOutWithAttemptedDeliveryReportsUnconfirmedNotQueued() throws Exception {
herdr.agentSendFailsWith("send_failed");
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150));
CompletableFuture.supplyAsync(() -> messages.send(T, "brief", 150, null));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // triggers the failing delivery attempt → ATTEMPTED
@@ -679,7 +680,7 @@ class MessageServiceTest {
// The primary answers, unblocking the worker; but the worker never sends the follow-up
// fleet_reply, so the answering send rides out its short window as still-working.
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200);
MessageService.Reply answer = messages.answer(q.turnId(), "config.yaml", 200, null);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answer.outcome(),
"an answered worker that never replies times out as still working");
@@ -715,7 +716,7 @@ class MessageServiceTest {
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
ask.get(5, TimeUnit.SECONDS); // worker resumed with the answer
awaitWaiting(); // the answering call has (re)opened its own forward waiter
@@ -723,7 +724,7 @@ class MessageServiceTest {
MessageService.Reply done = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, done.outcome());
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300);
MessageService.Reply probe = messages.send(T, "probe after normal reply", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released after a normal REPLIED answer(), or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work");
@@ -756,13 +757,13 @@ class MessageServiceTest {
// The primary answers, unblocking the worker; the worker never sends its follow-up
// fleet_reply, so the answering call rides out its short window as still-working.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 200, null));
MessageService.Reply answered = answer.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answered.outcome(),
"an answered worker that never replies times out as still working");
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300);
MessageService.Reply probe = messages.send(T, "probe after timed-out-working", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released after a TIMED_OUT_WORKING answer(), or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work");
@@ -791,7 +792,7 @@ class MessageServiceTest {
new java.util.concurrent.atomic.AtomicReference<>();
Thread answerer = new Thread(() -> {
try {
messages.answer(turnId, "config.yaml", 5000);
messages.answer(turnId, "config.yaml", 5000, null);
caught.set(new AssertionError("expected answer() to throw"));
} catch (Throwable t) {
caught.set(t);
@@ -813,7 +814,7 @@ class MessageServiceTest {
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300);
MessageService.Reply probe = messages.send(T, "probe after exceptional failure", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released when answer()'s reply future fails exceptionally, "
+ "or this bounded follow-up send would come back BUSY instead of timing out on its own work");
@@ -840,7 +841,7 @@ class MessageServiceTest {
new java.util.concurrent.atomic.AtomicReference<>();
Thread answerer = new Thread(() -> {
try {
messages.answer(turnId, "config.yaml", 5000);
messages.answer(turnId, "config.yaml", 5000, null);
caught.set(new AssertionError("expected answer() to throw"));
} catch (Throwable t) {
caught.set(t);
@@ -859,17 +860,147 @@ class MessageServiceTest {
ask.get(5, TimeUnit.SECONDS); // drain: the worker resumed with the answer
MessageService.Reply probe = messages.send(T, "probe after interruption", 300);
MessageService.Reply probe = messages.send(T, "probe after interruption", 300, null);
assertNotEquals(MessageService.Outcome.BUSY, probe.outcome(),
"the session lock must be released when answer()'s wait is interrupted, or this bounded "
+ "follow-up send would come back BUSY instead of timing out on its own work");
}
/**
* Drive an async send on {@code T} to a resolved reply, polling as {@code owner} until the
* ticket reports {@link MessageService.Phase#DONE} (or the 2s deadline runs out). {@code owner}
* must be a terminal this ticket's creator check actually accepts, or this loops to the
* deadline and returns a non-{@code DONE} view.
*/
private MessageService.TaskView driveAsyncTicketToDone(String ticket, String owner) throws InterruptedException {
MessageService.TaskView view = null;
long deadline = System.currentTimeMillis() + 2000;
while (view == null || view.phase() != MessageService.Phase.DONE) {
if (System.currentTimeMillis() >= deadline) break;
view = messages.poll(ticket, owner);
//noinspection BusyWait
Thread.sleep(5);
}
return view;
}
@Test
void pollReturnsNullForAnUnknownTicket() {
assertNull(messages.poll("task-999999"), "a ticket that was never minted is unknown");
}
// --- fleetd #719: a per-boot nonce keeps one instance's ticket ids out of another's space ---
/** A second, fully independent instance — its own agents/injector/rendezvous/inbox, not shared. */
private MessageService newIndependentInstance() {
FakeHerdr otherHerdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
AgentControl otherAgents = new AgentControl(otherHerdr);
Injector otherInjector = new Injector(otherAgents);
return new MessageService(otherAgents, otherInjector, new Rendezvous(), new InMemoryReplyInbox());
}
@Test
void twoInstancesMintDisjointTicketIds() {
MessageService other = newIndependentInstance();
String ticketFromThis = messages.sendAsync(T, "task on first instance", null, null);
String ticketFromOther = other.sendAsync(T, "task on second instance", null, null);
assertNotEquals(ticketFromThis, ticketFromOther,
"each instance mints its own id space, so even a first ticket from each must differ");
}
@Test
void foreignInstanceTicketDoesNotResolve() {
MessageService other = newIndependentInstance();
String ticket = messages.sendAsync(T, "task on first instance", null, null);
// `other` must reach the same sequence number, or this test passes against an empty map
// instead of against a colliding id.
other.sendAsync(T, "task on second instance", null, null);
// control: the id resolves in the instance that minted it, so a null below cannot be
// explained by broken plumbing — only by the ticket being foreign to `other`.
assertNotNull(messages.poll(ticket), "the minting instance must still resolve its own ticket");
assertNull(other.poll(ticket), "a ticket minted by a different instance must not resolve here");
}
// --- ticket ownership -----------------------------------------------------------------------
@Test
void leadBIsRefusedFromLeadAsTicket() throws Exception {
Principal leadA = Principal.leader("opus", "term_a", 1);
Principal leadB = Principal.leader("sol", "term_b", 2);
String ticket = messages.sendAsync(T, "long task", null, leadA);
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "secret async result"), "a reply resolves the async send");
MessageService.TaskView owner = driveAsyncTicketToDone(ticket, leadA.ownerKey());
assertNotNull(owner, "the creator must still be able to read its own ticket");
assertEquals(MessageService.Phase.DONE, owner.phase());
MessageService.TaskView refused = messages.poll(ticket, leadB.ownerKey());
assertNotNull(refused, "a different lead gets a refusal, not silence");
assertNotEquals(MessageService.Phase.DONE, refused.phase(),
"a different lead must never see the ticket as DONE");
assertNull(refused.reply(), "a refusal must never carry the reply text");
assertFalse(String.valueOf(refused).contains("secret async result"),
"the reply text must not appear anywhere in the refused view");
}
@Test
void unnamedPrimaryStillReadsAnyTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task", null,
Principal.leader("opus", "term_lead", 1));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "primary-visible result"), "a reply resolves the async send");
MessageService.TaskView view = driveAsyncTicketToDone(ticket, null);
assertNotNull(view, "the unnamed primary must be able to read any ticket");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("primary-visible result", view.reply());
}
@Test
void namedLeadCanPollItsTicketAfterItsTerminalChanges() throws Exception {
Principal oldLead = Principal.leader("opus", "term_OLD", 1);
Principal newLead = Principal.leader("opus", "term_NEW", 2);
assertNotEquals(oldLead.terminal(), newLead.terminal(), "the test requires different terminals");
String ticket = messages.sendAsync(T, "long task", null, oldLead);
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker works
assertTrue(rendezvous.resolve(T, "own result"), "a reply resolves the async send");
MessageService.TaskView view = driveAsyncTicketToDone(ticket, newLead.ownerKey());
assertNotNull(view, "the same named lead must read the ticket from its new terminal");
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("own result", view.reply());
}
@Test
void anonymousOwnerKeyIsRefusedByTheTicketGateItself() {
Principal lead = Principal.leader("opus", "term_lead", 1);
String ticket = messages.sendAsync(T, "long task", null, lead);
MessageService.TaskView refused = messages.poll(ticket, Principal.anonymous().ownerKey());
assertNotNull(refused);
assertEquals(MessageService.Phase.FAILED, refused.phase());
assertEquals("forbidden: this ticket was created by a different session", refused.detail());
}
@Test
void architectOwnershipUsesTerminalRatherThanSlot() {
Principal oldArchitect = Principal.architect("opus", "term_OLD", 1);
Principal newArchitect = Principal.architect("opus", "term_NEW", 2);
String ticket = messages.sendAsync(T, "long task", null, oldArchitect);
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket, oldArchitect.ownerKey()).phase());
assertEquals(MessageService.Phase.FAILED, messages.poll(ticket, newArchitect.ownerKey()).phase());
}
@Test
void pollReportsACompletedTicket() throws Exception {
String ticket = messages.sendAsync(T, "long task");
@@ -896,11 +1027,11 @@ class MessageServiceTest {
@Test
void concurrentSendToSameSessionWhileFirstHoldsItIsBusy() throws Exception {
CompletableFuture<MessageService.Reply> first =
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000));
CompletableFuture.supplyAsync(() -> messages.send(T, "first", 5000, null));
awaitWaiting(); // the first send now holds the session lock, blocked on its reply
// A second send to the SAME session cannot take the lock within its short window.
MessageService.Reply busy = messages.send(T, "second", 100);
MessageService.Reply busy = messages.send(T, "second", 100, null);
assertEquals(MessageService.Outcome.BUSY, busy.outcome(),
"a second send while another holds the session is busy, not a hang");
assertNull(busy.text());
@@ -930,13 +1061,13 @@ class MessageServiceTest {
PrimaryRegistry reg = new PrimaryRegistry(null);
// L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded.
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send owns the delegation");
// A attempts W while L holds it → BUSY (lock never taken) → its hook never fires.
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A));
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A), null);
assertEquals(MessageService.Outcome.BUSY, busy.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"a BUSY send must not steal the delegator ownership it never earned");
@@ -957,7 +1088,7 @@ class MessageServiceTest {
void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
@@ -968,7 +1099,7 @@ class MessageServiceTest {
// L finished; A's later accepted send takes the delegation over.
CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync(
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A)));
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A), null));
awaitWaiting();
assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send after the owner finished becomes the new delegator");
@@ -989,7 +1120,7 @@ class MessageServiceTest {
void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L)));
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L), null));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation");
@@ -1004,7 +1135,7 @@ class MessageServiceTest {
// L answers the ask on the same turn; the answer path must not touch ownership.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // the answering send reopened its forward waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
@@ -1024,7 +1155,7 @@ class MessageServiceTest {
void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() {
assertThrows(IllegalStateException.class,
() -> messages.send(T, "doomed", 500,
() -> { throw new IllegalStateException("ownership hook failed"); }),
() -> { throw new IllegalStateException("ownership hook failed"); }, null),
"a throwing ownership hook fails the send loudly");
assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter");
@@ -1152,7 +1283,7 @@ class MessageServiceTest {
@Test
void abandonFailsASendThatIsStillWaitingOnAReleasedSession() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
awaitWaiting();
assertTrue(messages.abandon(T, "session released"), "a live waiter is abandoned");
@@ -1172,7 +1303,7 @@ class MessageServiceTest {
@Test
void abandonDoesNotOverwriteAnAlreadyResolvedSend() throws Exception {
CompletableFuture<MessageService.Reply> send =
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000));
CompletableFuture.supplyAsync(() -> messages.send(T, "work", 30_000, null));
awaitWaiting();
assertTrue(rendezvous.resolve(T, "the real answer"));
@@ -1303,7 +1434,7 @@ class MessageServiceTest {
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -1343,7 +1474,7 @@ class MessageServiceTest {
// same turnId must see it as lapsed rather than resolving a question nobody is waiting on.
assertEquals(MessageService.AskOutcome.TIMED_OUT, ask.get(5, TimeUnit.SECONDS).outcome());
assertEquals(MessageService.Outcome.STALE_TURN,
messages.answer(asking.turnId(), "config.yaml", 200).outcome());
messages.answer(asking.turnId(), "config.yaml", 200, null).outcome());
}
@Test
@@ -1380,7 +1511,7 @@ class MessageServiceTest {
// The primary answers, but its own bounded wait for the worker's resumed turn is short and
// expires before the worker (still genuinely working) gets back to it.
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait gives up before the worker finishes resuming");
@@ -1409,7 +1540,7 @@ class MessageServiceTest {
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150);
MessageService.Reply answerReply = messages.answer(asking.turnId(), "config.yaml", 150, null);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome());
@@ -1463,7 +1594,7 @@ class MessageServiceTest {
messages.setFinishAsyncTaskRaceHookForTest(() -> messages.forgetTurnForTest(turnId));
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000));
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // answer() opened its own forward waiter for the resumed worker turn
@@ -1569,7 +1700,7 @@ class MessageServiceTest {
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000));
CompletableFuture.supplyAsync(() -> messages.answer(turnId, "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer(),
"the worker's own ask() call must have already unblocked with the primary's answer "
+ "before we force the race below");
@@ -1616,7 +1747,7 @@ class MessageServiceTest {
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
String turnId = asking.turnId();
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150);
MessageService.Reply answerReply = messages.answer(turnId, "config.yaml", 150, null);
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertEquals(MessageService.Outcome.TIMED_OUT_WORKING, answerReply.outcome(),
"the primary's own bounded wait must give up first, leaving turnId stamped with no "
@@ -1682,7 +1813,7 @@ class MessageServiceTest {
MessageService.TaskView asking = messages.poll(first);
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -1716,7 +1847,7 @@ class MessageServiceTest {
// The primary answers it — answer() resumes the turn and blocks for what comes next.
CompletableFuture<MessageService.Reply> answer1 = CompletableFuture.supplyAsync(
() -> messages.answer(asking1.turnId(), "a1", 5000));
() -> messages.answer(asking1.turnId(), "a1", 5000, null));
assertEquals("a1", ask1.get(5, TimeUnit.SECONDS).answer());
// Still in the SAME resumed turn — before replying — the worker asks again.
@@ -1737,7 +1868,7 @@ class MessageServiceTest {
// The primary answers the second question; the worker finally sends its real fleet_reply.
CompletableFuture<MessageService.Reply> answer2 = CompletableFuture.supplyAsync(
() -> messages.answer(turnId2, "a2", 5000));
() -> messages.answer(turnId2, "a2", 5000, null));
assertEquals("a2", ask2.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -1804,20 +1935,20 @@ class MessageServiceTest {
}
}
// --- CB-582: fleet_status pendingAsk() ------------------------------------------------------
// --- fleet_status pendingAsk() -----------------------------------------------------------
@Test
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask");
assertNull(messages.pendingAsk(T, null), "no async ticket at all -> no pending ask");
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question");
assertNull(messages.pendingAsk(T, null), "a plain pending delegation is not a question");
injectDelivery();
assertTrue(rendezvous.resolve(T, "done"));
awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either");
assertNull(messages.pendingAsk(T, null), "a finished ticket carries no open question either");
}
@Test
@@ -1830,21 +1961,183 @@ class MessageServiceTest {
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.PendingAsk pending = messages.pendingAsk(T);
MessageService.PendingAsk pending = messages.pendingAsk(T, null);
assertNotNull(pending, "fleet_status should see the open question");
assertEquals(ticket, pending.ticket());
assertEquals("which config file?", pending.question());
assertEquals(asking.turnId(), pending.turnId());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertNull(messages.pendingAsk(T), "an answered question is no longer pending");
assertNull(messages.pendingAsk(T, null), "an answered question is no longer pending");
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* A caller's owner key must match the key that created the delegation to see its pending
* question. The unnamed primary always sees it.
*/
@Test
void pendingAskGatesTheQuestionByTheDelegationsCreatorOwner() throws Exception {
Principal creator = Principal.worker("term_creator", 1);
String ticket = messages.sendAsync(T, "task that asks", null, creator);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNull(messages.pendingAsk(T, "worker:term_other"),
"a caller whose key did not create the delegation must not see the question");
MessageService.PendingAsk own = messages.pendingAsk(T, creator.ownerKey());
assertNotNull(own, "the creating caller must see its own open question");
assertEquals("which config file?", own.question());
MessageService.PendingAsk unnamed = messages.pendingAsk(T, null);
assertNotNull(unnamed, "a caller with no terminal (the unnamed primary) must always see the question");
assertEquals("which config file?", unnamed.question());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, creator.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* A task created with no recorded owner (a short {@code sendAsync} overload) must not hand its
* open question to a caller with an owner key. Only the unnamed primary may still see it.
*/
@Test
void pendingAskDeniesATerminalBearingCallerWhenTheTaskRecordsNoCreator() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNull(messages.pendingAsk(T, "worker:term_someone"),
"a caller with an owner key must not see a question whose task records no creator");
assertNotNull(messages.pendingAsk(T, null),
"the unnamed primary must still see it even with no recorded creator");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- fleetd #715: answer() is gated on the caller that owns the turn -----------------------
/**
* A blocking {@code fleet_send} from one caller opens the turn; a different caller's answer is
* refused with no side effect on the rendezvous or the worker's blocked {@code fleet_ask} —
* only the real owner can answer it.
*/
@Test
void aDifferentCallersAnswerIsRefusedForABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, "term_owner"));
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertFalse(rendezvous.isWaiting(T), "the forward waiter closes once the question surfaces");
// The hijack: a different caller answers the SAME turnId.
MessageService.Reply hijacked = messages.answer(q.turnId(), "evil.yaml", 500, "term_attacker");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
"a caller that did not open this turn must be refused, not served");
assertFalse(rendezvous.isWaiting(T), "a refused answer must not open a forward waiter");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(T, rendezvous.askSession(q.turnId()), "a refused answer must leave the ask turn open");
// The control: the real owner answers the same turnId and the turn resumes normally.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(q.turnId(), "config.yaml", 5000, "term_owner"));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* Same hijack and control as the blocking case, but the delegation is opened through the
* fire-and-poll path ({@code sendAsync}) — the owner comes from the ticket's recorded creator
* terminal, not a caller argument threaded through a live blocking call.
*/
@Test
void aDifferentCallersAnswerIsRefusedForAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
Principal owner = Principal.worker("term_owner", 1);
String ticket = messages.sendAsync(T, "task that asks", null, owner);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.Reply hijacked = messages.answer(asking.turnId(), "evil.yaml", 500,
"worker:term_attacker");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER, hijacked.outcome(),
"a caller that did not create this delegation must be refused, not served");
assertFalse(ask.isDone(), "a refused answer must not resolve the worker's blocked fleet_ask");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase(),
"a refused answer must not advance the async ticket's phase");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, owner.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
}
@Test
void namedLeadCanSeeAndAnswerAnAskAfterItsTerminalChangesWhileLeadBIsRefused() throws Exception {
Principal oldLead = Principal.leader("opus", "term_OLD", 1);
Principal newLead = Principal.leader("opus", "term_NEW", 2);
Principal leadB = Principal.leader("sol", "term_SOL", 3);
assertNotEquals(oldLead.terminal(), newLead.terminal(), "the test requires different terminals");
String ticket = messages.sendAsync(T, "task that asks", null, oldLead);
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertNotNull(messages.pendingAsk(T, newLead.ownerKey()),
"the same lead at its new terminal must see the pending ask");
assertNull(messages.pendingAsk(T, leadB.ownerKey()),
"another lead must not see the pending ask");
assertEquals(MessageService.Outcome.NOT_TURN_OWNER,
messages.answer(asking.turnId(), "evil.yaml", 500, leadB.ownerKey()).outcome());
assertFalse(ask.isDone(), "another lead must not resolve the worker's ask");
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000, newLead.ownerKey()));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
assertEquals("done", awaitTicketPhase(ticket, MessageService.Phase.DONE).reply());
}
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
//
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
@@ -2101,7 +2394,7 @@ class MessageServiceTest {
// Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
@@ -2127,7 +2420,7 @@ class MessageServiceTest {
.filter(c -> c.method().equals("agent.prompt")).count();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000, null));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
@@ -2520,7 +2813,7 @@ class MessageServiceTest {
void hasQueuedDeliveryIsTrueAfterAnUndeliveredSendTimesOut() {
// Nothing ever delivers the message (never goes IDLE/BLOCKED), so the send times out with
// TIMED_OUT_QUEUED — same setup as sendTimesOutBeforeDeliveryIsQueuedNotWorking above.
MessageService.Reply r = messages.send(T, "never delivered", 50);
MessageService.Reply r = messages.send(T, "never delivered", 50, null);
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, r.outcome());
assertTrue(messages.hasQueuedDelivery(T),
@@ -2532,7 +2825,7 @@ class MessageServiceTest {
@Test
void hasQueuedDeliveryClearsOnceTheTargetsNextDeliveryIsAccepted() throws Exception {
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome());
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
assertTrue(messages.hasQueuedDelivery(T));
// A fresh send accepts delivery (opens its own waiter) — the stale queued fact is cleared.
@@ -2549,7 +2842,7 @@ class MessageServiceTest {
@Test
void hasQueuedDeliveryClearsOnAbandon() {
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50).outcome());
assertEquals(MessageService.Outcome.TIMED_OUT_QUEUED, messages.send(T, "first", 50, null).outcome());
assertTrue(messages.hasQueuedDelivery(T));
messages.abandon(T, "session released");
@@ -0,0 +1,187 @@
package dev.ltms.fleet.msg;
import org.junit.jupiter.api.Test;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.stream.Stream;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Pins that no file under {@code src/main/java} calls the fail-open
* {@link Rendezvous#open(String)} overload. That overload opens a forward waiter with no
* recorded {@link Rendezvous.Owner}, so the turn it opens can never be answered by anyone,
* not even the unnamed primary. Every production caller must go through
* {@link Rendezvous#open(String, Rendezvous.Owner)} and record an explicit owner.
*
* <p>This reads each file's own source text rather than reflecting on compiled bytecode, because
* the risk is a future one-word edit at a call site, not a missing overload.
*
* <p>The scan below finds a violation by its receiver, {@code rendezvous.open(}, classified by
* argument count. The one test method here points that exact scanner at a file known to hold
* many real one-argument calls before it ever looks at production, so a scanner that stops
* matching fails loudly on the known-positive case instead of leaving a clean production result
* looking like evidence it never produced.
*/
class RendezvousOpenUsageTest {
private static final Path PRODUCTION_SOURCE = Path.of("src/main/java");
private static final Path KNOWN_TEST_CALLER =
Path.of("src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java");
/**
* A pattern that cannot find a known one-argument {@code rendezvous.open(} call would also
* find none in production -- not because production is clean, but because the pattern does
* not match the text it is supposed to catch. That failure mode is exactly what let a
* {@code \b}-based regex read "no callers" under {@code git grep -E} when 81 real ones
* existed: {@code git grep} does not treat {@code \b} as a word boundary, so the pattern
* silently matched nothing anywhere, clean code and real calls alike. This test runs the
* known-positive check first, with the same scanning method the production check then
* depends on, so that mistake fails loudly here instead of reading as a clean result.
*/
@Test
void noProductionFileCallsTheSingleArgumentOpenOverload() throws IOException {
List<String> knownCalls = new ArrayList<>();
int knownFilesScanned = scanForOneArgOpenCalls(KNOWN_TEST_CALLER, knownCalls);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "found
// nothing" having looked at nothing.
assertTrue(knownFilesScanned > 0, "control failed: the scan under " + KNOWN_TEST_CALLER
+ " visited zero .java files -- the path is wrong, so neither result below proves "
+ "anything");
// CONTROL: the scanner actually finds real one-argument rendezvous.open( calls when
// pointed at a file known to hold many. If this is not satisfied, the matching logic
// itself is broken, and the production result below is the scanner failing silently,
// not production code actually being clean.
assertTrue(knownCalls.size() >= 60, "control failed: the scanner found only "
+ knownCalls.size() + " one-argument rendezvous.open( call(s) in " + KNOWN_TEST_CALLER
+ ", which is known to hold many -- the matching logic itself is broken: " + knownCalls);
List<String> violations = new ArrayList<>();
int filesScanned = scanForOneArgOpenCalls(PRODUCTION_SOURCE, violations);
// CONTROL: the scan actually walked files -- a wrong root would otherwise report "no
// violations found" having looked at nothing.
assertTrue(filesScanned > 0, "control failed: the scan under " + PRODUCTION_SOURCE
+ " visited zero .java files -- the path is wrong, so the absence of violations "
+ "below proves nothing");
assertTrue(violations.isEmpty(), "found a call to the fail-open Rendezvous.open(String) "
+ "overload, which opens a forward waiter with no recorded owner -- record an "
+ "explicit Rendezvous.Owner through open(String, Owner) instead: " + violations);
}
private static int scanForOneArgOpenCalls(Path root, List<String> sites) throws IOException {
List<Path> files = javaFiles(root);
for (Path file : files) {
scanFileForOpenCalls(file, sites);
}
return files.size();
}
private static List<Path> javaFiles(Path root) throws IOException {
try (Stream<Path> paths = Files.walk(root)) {
return paths.filter(p -> p.toString().endsWith(".java")).toList();
}
}
private static void scanFileForOpenCalls(Path file, List<String> sites) throws IOException {
String source = Files.readString(file);
String needle = "rendezvous.open(";
int from = 0;
int idx;
while ((idx = source.indexOf(needle, from)) >= 0) {
int argsStart = idx + needle.length();
String args = extractBalancedArgs(source, argsStart, file, idx);
int closeParenIndex = argsStart + args.length();
from = closeParenIndex + 1;
if (args.isBlank()) {
continue; // Rendezvous has no zero-argument open() -- a javadoc "rendezvous.open()" mention, not a call
}
if (topLevelCommaCount(args) == 0) {
sites.add(file + ":" + lineOf(source, idx) + " -- rendezvous.open(" + args.trim() + ")");
}
}
}
/**
* The text between {@code rendezvous.open(} and its matching close paren: balanced over
* nested calls, and never split by a paren or comma sitting inside a string or char literal.
*/
private static String extractBalancedArgs(String source, int start, Path file, int callIndex) {
int depth = 1;
boolean inString = false;
boolean inChar = false;
int i = start;
while (i < source.length()) {
char c = source.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
if (depth == 0) return source.substring(start, i);
}
i++;
}
throw new IllegalStateException(
"unbalanced parentheses scanning " + file + ":" + lineOf(source, callIndex));
}
/**
* Commas at paren/bracket/brace depth zero, skipping string and char literals -- the argument
* separators a human reader would see, not every comma character in the text.
*/
private static int topLevelCommaCount(String args) {
int depth = 0;
int commas = 0;
boolean inString = false;
boolean inChar = false;
int i = 0;
while (i < args.length()) {
char c = args.charAt(i);
if (inString) {
if (c == '\\') { i += 2; continue; }
if (c == '"') inString = false;
} else if (inChar) {
if (c == '\\') { i += 2; continue; }
if (c == '\'') inChar = false;
} else if (c == '"') {
inString = true;
} else if (c == '\'') {
inChar = true;
} else if (c == '(' || c == '[' || c == '{') {
depth++;
} else if (c == ')' || c == ']' || c == '}') {
depth--;
} else if (c == ',' && depth == 0) {
commas++;
}
i++;
}
return commas;
}
private static int lineOf(String source, int index) {
int line = 1;
for (int i = 0; i < index; i++) {
if (source.charAt(i) == '\n') line++;
}
return line;
}
}
@@ -132,4 +132,127 @@ class RendezvousTest {
assertNull(rendezvous.askSession(t.turnId()), "a closed ask is forgotten");
assertFalse(rendezvous.answerAsk(t.turnId(), "late"), "a closed ask can no longer be answered");
}
// ── fleetd #715: turn ownership ────────────────────────────────────────────────────────────
@Test
void openWithNoOwnerRecordsNoOwnerAndOpenWithAnOwnerRecordsIt() {
rendezvous.open(W);
assertNull(rendezvous.ownerOf(W), "the no-owner overload records no owner at all");
rendezvous.close(W, rendezvous.currentWaiter(W));
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.ownerOf(W));
}
@Test
void ownerOfIsNullWhenNoWaiterIsOpen() {
assertNull(rendezvous.ownerOf(W), "no waiter open means no owner to report");
}
@Test
void openAskStampsTheForwardWaitersOwnerOntoTheFreshTurnOnly() {
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
Rendezvous.AskTicket fresh = rendezvous.openAsk(W);
assertTrue(fresh.fresh());
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(fresh.turnId()),
"a freshly-opened ask copies the forward waiter's current owner");
Rendezvous.AskTicket coalesced = rendezvous.openAsk(W);
assertFalse(coalesced.fresh());
assertEquals(fresh.turnId(), coalesced.turnId());
assertEquals(Rendezvous.Owner.of("term_lead"), rendezvous.askOwner(coalesced.turnId()),
"a coalesced duplicate ask rides the fresh owner's turn, unchanged");
}
@Test
void aSecondAskAfterTheFirstClosesStampsWhateverOwnerIsOpenAtThatLaterMoment() {
rendezvous.open(W, Rendezvous.Owner.of("term_lead"));
Rendezvous.AskTicket first = rendezvous.openAsk(W);
rendezvous.closeAsk(first.turnId());
// The forward waiter is reopened under a different owner before the second ask — mirrors
// answer() reopening with the owner it already checked, which can differ turn to turn.
rendezvous.close(W, rendezvous.currentWaiter(W));
rendezvous.open(W, Rendezvous.Owner.of("term_other"));
Rendezvous.AskTicket second = rendezvous.openAsk(W);
assertTrue(second.fresh());
assertEquals(Rendezvous.Owner.of("term_other"), rendezvous.askOwner(second.turnId()),
"a freshly-opened ask always copies whatever owner is open right now, not a stale one");
}
@Test
void askOwnerIsNullForAnUnknownOrLapsedTurn() {
assertNull(rendezvous.askOwner("no-such#1"));
Rendezvous.AskTicket t = rendezvous.openAsk(W);
rendezvous.closeAsk(t.turnId());
assertNull(rendezvous.askOwner(t.turnId()), "a closed ask no longer reports an owner");
}
// ── fleetd #729: per-boot nonce guards turnId against cross-instance reuse ────────────────
@Test
void twoInstancesMintDisjointTurnIds() {
Rendezvous other = new Rendezvous();
Rendezvous.AskTicket fromThis = rendezvous.openAsk(W);
Rendezvous.AskTicket fromOther = other.openAsk(W);
assertNotEquals(fromThis.turnId(), fromOther.turnId(),
"each instance mints its own id space, so even a first ask from each must differ");
}
@Test
void foreignInstanceTurnIdDoesNotResolve() {
Rendezvous other = new Rendezvous();
// `other` must reach the same sequence number as `rendezvous` (two asks each, the first
// closed so the second mints fresh), or this test passes against an empty map instead of
// against a colliding id.
Rendezvous.AskTicket firstFromThis = rendezvous.openAsk(W);
rendezvous.closeAsk(firstFromThis.turnId());
Rendezvous.AskTicket secondFromThis = rendezvous.openAsk(W);
Rendezvous.AskTicket firstFromOther = other.openAsk(W);
other.closeAsk(firstFromOther.turnId());
other.openAsk(W);
// control: the id resolves in the instance that minted it, so a false below cannot be
// explained by broken plumbing — only by the turnId being foreign to `other`.
assertTrue(rendezvous.answerAsk(secondFromThis.turnId(), "answer from this instance"),
"the minting instance must still resolve its own turnId");
assertFalse(other.answerAsk(secondFromThis.turnId(), "answer from other instance"),
"a turnId minted by a different instance must not resolve here");
}
@Test
void openAskStillCoalescesDuplicatesAndStillMintsDistinctIdsPerAsk() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
Rendezvous.AskTicket t2 = rendezvous.openAsk(W);
assertEquals(t1.turnId(), t2.turnId(),
"a second openAsk while one is open still coalesces onto the same turn");
assertFalse(t2.fresh(), "the coalesced ask is still reported as not fresh");
rendezvous.closeAsk(t1.turnId());
Rendezvous.AskTicket t3 = rendezvous.openAsk(W);
assertNotEquals(t1.turnId(), t3.turnId(), "two asks from the same session still get different turnIds");
}
@Test
void ownerPermitsIsFailClosedOnARecordAndThreeStatesAreDistinct() {
assertFalse(Rendezvous.Owner.permits(null, null),
"no owner on record refuses even the unnamed primary");
assertFalse(Rendezvous.Owner.permits(null, "worker:term_a"),
"no owner on record refuses a caller with an owner key too");
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, null),
"the recorded unnamed primary matches a caller with a null owner key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.UNNAMED_PRIMARY, "worker:term_a"),
"the recorded unnamed primary does not match another owner key");
assertTrue(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), "worker:term_a"),
"an owner matches the same key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), "worker:term_b"),
"an owner refuses a different key");
assertFalse(Rendezvous.Owner.permits(Rendezvous.Owner.of("worker:term_a"), null),
"a named owner refuses the unnamed primary");
}
}
@@ -1,8 +1,14 @@
package dev.ltms.fleet.rest;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.auth.Role;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -15,9 +21,11 @@ import dev.ltms.fleet.metrics.FleetMetrics;
import dev.ltms.fleet.metrics.Metrics;
import dev.ltms.fleet.msg.MessageService;
import dev.ltms.fleet.msg.Rendezvous;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.session.FakeWorktrees;
import dev.ltms.fleet.session.SessionManager;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.testing.CapturedLog;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -28,10 +36,15 @@ import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import java.util.function.Predicate;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
@@ -122,12 +135,467 @@ class FleetAppAuthTest {
assertEquals(Authz.Action.DRAIN, FleetApp.routeAction("GET /sessions/{id}/replies"));
assertEquals(Authz.Action.ASK, FleetApp.routeAction("POST /sessions/{id}/ask"));
for (String route : Set.of("GET /sessions", "GET /agents", "GET /members", "GET /profiles",
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
"GET /member-credentials")) {
assertEquals(Authz.Action.READ, FleetApp.routeAction(route), route);
}
for (String route : Set.of("GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
assertEquals(Authz.Action.TASK_READ, FleetApp.routeAction(route), route);
}
assertThrows(IllegalArgumentException.class, () -> FleetApp.routeAction("GET /healthz"));
}
// --- GET /tasks/{ticket} must not be the no-check overload ----------------------------------
/**
* {@code GET /tasks/{ticket}} must resolve its caller the same way {@code allow(...)} does
* and thread that owner key into {@link MessageService#poll(String, String)}, not the
* no-check overload that ignores who is asking.
*/
@Test
void theTaskStatusRouteActuallyThreadsTheCallersOwnerKeyIntoPoll() throws Exception {
String source = Files.readString(REST_SOURCE);
int start = source.indexOf("private void taskStatus(Context ctx) {");
assertTrue(start >= 0, "could not find taskStatus in " + REST_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("private static void herdrError(Context ctx, HerdrException e) {", start);
assertTrue(end > start, "could not find the method declared after taskStatus to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call messages.poll(...) -- if this fails, the
// anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("messages.poll("),
"control failed: the scraped taskStatus block contains no messages.poll( call at all "
+ "-- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("caller.ownerKey()"),
"the taskStatus route must thread the resolved caller's owner key into messages.poll(...), "
+ "not the no-check overload -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
"the taskStatus route must resolve its caller the same way allow(...) does, not via a "
+ "second, separate resolution path -- block: " + handlerBlock);
}
/**
* {@code GET /sessions/{id}/status} must resolve its caller the same way {@code allow(...)}
* does and thread that owner key into {@link MessageService#pendingAsk(String, String)}, not
* the no-check overload that ignores who is asking.
*/
@Test
void theSessionStatusRouteActuallyThreadsTheCallersOwnerKeyIntoPendingAsk() throws Exception {
String source = Files.readString(REST_SOURCE);
int start = source.indexOf("private void sessionStatus(Context ctx) {");
assertTrue(start >= 0, "could not find sessionStatus in " + REST_SOURCE
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
int end = source.indexOf("private void taskStatus(Context ctx) {", start);
assertTrue(end > start, "could not find the method declared after sessionStatus to bound the scrape");
String handlerBlock = source.substring(start, end);
// CONTROL: the block we scraped really does call messages.pendingAsk(...) -- if this fails,
// the anchors above moved and the assertions below would otherwise pass on nothing.
assertTrue(handlerBlock.contains("messages.pendingAsk("),
"control failed: the scraped sessionStatus block contains no messages.pendingAsk( "
+ "call at all -- the anchors have drifted, this test is not testing what it claims to");
assertTrue(handlerBlock.contains("caller.ownerKey()"),
"the sessionStatus route must thread the resolved caller's owner key into "
+ "messages.pendingAsk(...), not the no-check overload -- block: " + handlerBlock);
assertTrue(handlerBlock.contains("ctx.attribute(CALLER)"),
"the sessionStatus route must resolve its caller the same way allow(...) does, not via "
+ "a second, separate resolution path -- block: " + handlerBlock);
}
/**
* {@code GET /tasks/{ticket}} refuses a worker whose terminal did not create the ticket,
* while the creating worker and the unnamed primary both still read it. The ticket is minted
* directly on the shared {@link MessageService}, the same way {@code MessageServiceTest}
* drives {@link MessageService#poll(String, String)}, so this exercises only the REST poll
* route's own handling of the ownership already recorded on the ticket.
*/
@Test
void restPollRefusesADifferentWorkerButAllowsTheCreatorAndTheUnnamedPrimary() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
try {
String ticket = messages.sendAsync("term_a", "long task", null,
Principal.worker("term_a", FakeHerdr.WORKER_PID));
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, refused.statusCode());
assertTrue(refused.body().contains("forbidden"),
"a different worker's terminal must be refused, not shown the ticket: " + refused.body());
assertFalse(refused.body().contains("\"reply\""),
"a refusal must never carry reply text: " + refused.body());
HttpResponse<String> own = send(creatorApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the creating worker must read its own ticket: " + own.body());
HttpResponse<String> primary = send(primaryApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, primary.statusCode());
assertFalse(primary.body().contains("forbidden"),
"the unnamed primary must read any ticket: " + primary.body());
} finally {
creatorApp.stop();
otherWorkerApp.stop();
primaryApp.stop();
}
}
/**
* {@code POST /sessions/{id}/message} with {@code wait:false} must record the creating
* caller's own terminal on the ticket it returns, so that caller can still poll its own
* ticket over REST, while a different terminal is refused.
*/
@Test
void restSendAsyncRecordsTheCreatingCallersTerminalSoItCanStillPollItsOwnTicket() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
Javalin leadApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-x"));
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
try {
ObjectMapper mapper = new ObjectMapper();
HttpResponse<String> created = send(leadApp.port(), "POST", "/sessions/term_a/message",
"{\"content\":\"long task\",\"wait\":false}", null);
assertEquals(202, created.statusCode(), created.body());
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
HttpResponse<String> own = send(leadApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, own.statusCode());
assertFalse(own.body().contains("forbidden"),
"the session that created the ticket over REST must be able to poll it: " + own.body());
HttpResponse<String> refused = send(otherWorkerApp.port(), "GET", "/tasks/" + ticket, null, null);
assertEquals(200, refused.statusCode());
assertTrue(refused.body().contains("forbidden: this ticket was created by a different session"),
"a different terminal must still be refused with the ownership detail, not some "
+ "other rejection: " + refused.body());
} finally {
leadApp.stop();
otherWorkerApp.stop();
}
}
/**
* {@code GET /sessions/{id}/status} shows a worker's pending {@code fleet_ask} question, its
* {@code turnId} and its ticket only to the caller whose owner key created that delegation, or
* to the unnamed primary. A different caller still sees the base status line, but none of the
* pending-ask fields.
*/
@Test
void restStatusGatesThePendingAskFieldsByTheDelegationsCreatorOwner() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin creatorApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID); // -> term_a
Javalin otherWorkerApp = startOnSharedService(messages, herdr, 9001L); // -> term_shell
Javalin primaryApp = startOnSharedService(messages, herdr, 999_999L); // no pane -> primary
try {
ObjectMapper mapper = new ObjectMapper();
String ticket = messages.sendAsync("term_target", "task that asks", null,
Principal.worker("term_a", FakeHerdr.WORKER_PID));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket, null);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
JsonNode other = mapper.readTree(
send(otherWorkerApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("idle", other.get("status").asText(), "the base status must still be shown");
assertFalse(other.has("question"), "a non-creating caller must not see the question: " + other);
assertFalse(other.has("turnId"), "a non-creating caller must not see the turnId: " + other);
assertFalse(other.has("ticket"), "a non-creating caller must not see the ticket: " + other);
JsonNode own = mapper.readTree(
send(creatorApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("which config file?", own.get("question").asText(), "the creator must see the question");
assertEquals(turnId, own.get("turnId").asText(), "the creator must see the turnId");
assertEquals(ticket, own.get("ticket").asText(), "the creator must see the ticket");
JsonNode primary = mapper.readTree(
send(primaryApp.port(), "GET", "/sessions/term_target/status", null, null).body());
assertEquals("which config file?", primary.get("question").asText(),
"a caller with no terminal (the unnamed primary) must see the question");
// Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000, "worker:term_a"));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
answer.get(5, TimeUnit.SECONDS);
} finally {
creatorApp.stop();
otherWorkerApp.stop();
primaryApp.stop();
}
}
/**
* {@code POST /sessions/{id}/message} with a {@code turnId} refuses a caller whose terminal
* did not open the turn, even when that caller otherwise holds ANSWER rights, and leaves the
* turn open for the real owner to resolve. Covers the REST adapter's ANSWER gate for an
* async-send ({@code wait:false}) delegation.
*/
@Test
void restAnswerIsRefusedForADifferentCallerOnAnAsyncSendDelegationButTheRealOwnerSucceeds() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
try {
ObjectMapper mapper = new ObjectMapper();
HttpResponse<String> created = send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"long task\",\"wait\":false}", null);
assertEquals(202, created.statusCode(), created.body());
String ticket = mapper.readTree(created.body()).path("ticket").asText(null);
assertNotNull(ticket, "the accepted response carried no ticket: " + created.body());
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "the async send should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket, null);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
String turnId = asking.turnId();
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
assertEquals(403, hijacked.statusCode(), hijacked.body());
assertTrue(hijacked.body().contains("not_turn_owner"),
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
assertEquals("term_target", rendezvous.askSession(turnId),
"a refused answer must leave the ask turn open");
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
"the real owner's answer must resolve the worker's turn");
} finally {
ownerApp.stop();
attackerApp.stop();
}
}
/**
* As above, but the delegation is a blocking send ({@code wait:true}) instead of a ticket —
* the owning caller's own HTTP request is the one that surfaces the worker's question and
* later carries the real answer. Covers the REST adapter's ANSWER gate for a blocking-send
* delegation.
*/
@Test
void restAnswerIsRefusedForADifferentCallerOnABlockingSendDelegationButTheRealOwnerSucceeds() throws Exception {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
Injector injector = new Injector(agents);
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
Javalin ownerApp = startOnSharedService(messages, herdr, FakeHerdr.WORKER_PID, Map.of("term_a", "lead-owner"));
Javalin attackerApp = startOnSharedService(messages, herdr, 9001L, Map.of("term_shell", "lead-attacker"));
try {
ObjectMapper mapper = new ObjectMapper();
CompletableFuture<HttpResponse<String>> blocking = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"long task\",\"wait\":true,\"timeoutMs\":5000}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_target"), "the blocking send should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> messages.ask("term_target", "which config file?", 5000));
HttpResponse<String> questionResponse = blocking.get(5, TimeUnit.SECONDS);
assertEquals(202, questionResponse.statusCode(), questionResponse.body());
JsonNode question = mapper.readTree(questionResponse.body());
assertEquals("question", question.get("status").asText());
String turnId = question.get("turnId").asText();
HttpResponse<String> hijacked = send(attackerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"hijack\",\"turnId\":\"" + turnId + "\"}", null);
assertEquals(403, hijacked.statusCode(), hijacked.body());
assertTrue(hijacked.body().contains("not_turn_owner"),
"a different caller's answer must be refused as not_turn_owner: " + hijacked.body());
assertEquals("term_target", rendezvous.askSession(turnId),
"a refused answer must leave the ask turn open");
CompletableFuture<HttpResponse<String>> owned = CompletableFuture.supplyAsync(() -> {
try {
return send(ownerApp.port(), "POST", "/sessions/term_target/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\"}", null);
} catch (Exception e) {
throw new RuntimeException(e);
}
});
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_target") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve("term_target", "done"));
HttpResponse<String> ownedResponse = owned.get(5, TimeUnit.SECONDS);
assertEquals(200, ownedResponse.statusCode(), ownedResponse.body());
assertEquals("done", mapper.readTree(ownedResponse.body()).path("reply").asText(null),
"the real owner's answer must resolve the worker's turn");
} finally {
ownerApp.stop();
attackerApp.stop();
}
}
/**
* As {@link #start}, but shares {@code messages} and {@code herdr} across several app
* instances bound to different pids, each returned as its own started {@link Javalin} rather
* than through the shared {@code app} field, so several differently-resolved callers can
* poll the same ticket.
*/
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid) {
return startOnSharedService(messages, herdr, pid, Map.of());
}
/**
* As {@link #startOnSharedService(MessageService, FakeHerdr, long)}, but {@code leadTerminals}
* resolves the given pid's terminal to a named lead (a caller with SEND permission) instead of
* a plain worker, for a test that needs a terminal-bearing caller able to create a ticket.
*
* <p>Every connecting pane not already claimed by {@code leadTerminals} is wired into the live
* roster as a spawned worker, so a caller's resolved role matches what its own test expects:
* a {@link Role#WORKER}, never the unconfigured-pane {@link Role#OBSERVER} floor a roster-less
* resolver would otherwise fall to.
*/
private Javalin startOnSharedService(MessageService messages, FakeHerdr herdr, long pid,
Map<String, String> leadTerminals) {
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> leadTerminals, new MemberRegistry(null),
t -> leadTerminals.containsKey(t) ? null : MemberRole.DEV, Map::of);
Metrics appMetrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
return new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, appMetrics).build().start("127.0.0.1", 0);
}
/**
* fleetd #669 Unit A: {@code POST /sessions/{id}/message} is two call shapes behind one route,
* mirroring {@code fleet_send}'s MCP-side split into {@link Authz.Action#SEND} and {@link
* Authz.Action#ANSWER} ({@code FleetMcp#sendAction}). The route never carries a {@code coordId}
* shape — that peer-lead route is MCP-only — so only these two apply here.
*/
@Test
void theMessageRouteIsASendWithNoTurnIdAndAnAnswerWithOne() {
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", null));
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", " "));
assertEquals(Authz.Action.ANSWER, FleetApp.routeAction("POST /sessions/{id}/message", "turn-1"));
}
/**
* fleetd #689: {@code answerGatePasses} is the second, conditional gate behind {@code
* sendMessage}'s coarse {@link Authz.Action#SEND} check. With the {@code ANSWER} grant denied,
* a {@code turnId}-bearing request is refused while a plain one still passes — and the denied
* permit is queried only for the {@code turnId} case, never for the plain one, which is what
* proves this is a genuinely separate, conditional check rather than the {@code SEND} check
* renamed or an unconditional call whose result is ignored. Flipping only the {@code ANSWER}
* grant to allowed then flips only the {@code turnId} shape's outcome.
*/
@Test
void answerGatePassesOnlyWhenTurnIdAbsentOrAnswerGranted() {
List<Authz.Action> queried = new ArrayList<>();
Predicate<Authz.Action> denyAnswer = action -> {
queried.add(action);
return false;
};
assertFalse(FleetApp.answerGatePasses("turn-1", denyAnswer),
"ANSWER denied ⇒ the turnId shape is refused");
assertEquals(List.of(Authz.Action.ANSWER), queried,
"the ANSWER grant, specifically, must be the one consulted");
queried.clear();
assertTrue(FleetApp.answerGatePasses(null, denyAnswer),
"no turnId ⇒ the plain shape passes even though ANSWER is denied");
assertTrue(FleetApp.answerGatePasses(" ", denyAnswer), "a blank turnId is treated as absent");
assertEquals(List.of(), queried, "the plain shape must never consult the permit at all");
assertTrue(FleetApp.answerGatePasses("turn-1", action -> true),
"flipping only the ANSWER grant to allowed flips only the turnId shape's outcome");
}
private static Set<String> routesTheServerRegisters() {
try {
String source = Files.readString(REST_SOURCE).lines()
@@ -147,6 +615,74 @@ class FleetAppAuthTest {
}
}
/**
* {@code permitsFor} is the exact decision {@link FleetApp#allow} makes, passing the real
* production classifier rather than a test-supplied one — built from an empty {@link
* CallerResolver}, so no terminal is recognised as a configured lead or collaborator and a
* collaborator's SEND is refused through the REST gate.
*/
@Test
void aCollaboratorMayNotSendOverRestWithTheRealProductionClassifier() {
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> 700L);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null));
Principal collaborator = Principal.collaborator("ops", "term_collab", 700);
assertFalse(FleetApp.permitsFor(collaborator, Authz.Action.SEND, "term_lead",
callers.knownLeadOrCollaborator()),
"no terminal is recognised as a lead or collaborator yet");
}
/**
* fleetd #669 Unit D: wires a real {@link CallerResolver} with a known lead and a known
* collaborator tab, and a spawned member's own terminal recognised by neither map. A
* collaborator's SEND reaches the known lead and the known collaborator, and is refused for the
* spawned member's terminal — over the REST route, not just the unit-level classifier, so a
* test covering only MCP cannot leave this route open.
*/
@Test
void aCollaboratorMaySendToAKnownLeadOrCollaboratorButNotToASpawnedMembersTerminalOverRest() throws Exception {
int port = startWithRealClassifier(FakeHerdr.WORKER_PID,
Map.of("term_lead_known", "lead-x"), Map.of("term_a", "ops2"));
HttpResponse<String> toLead = send(port, "POST", "/sessions/term_lead_known/message",
"{\"content\":\"hi\",\"wait\":false}", null);
assertEquals(202, toLead.statusCode(), toLead.body());
HttpResponse<String> toSpawnedMembersTerminal = send(port, "POST", "/sessions/term_worker/message",
"{\"content\":\"hi\",\"wait\":false}", null);
assertEquals(403, toSpawnedMembersTerminal.statusCode(), toSpawnedMembersTerminal.body());
}
/**
* As {@link #start}, but with explicit lead/collaborator maps and no spawned-member roster, so
* a test can wire the real {@link CallerResolver#knownLeadOrCollaborator()} classifier instead
* of the default empty one.
*/
private int startWithRealClassifier(long pid, Map<String, String> leadTerminals,
Map<String, String> collaboratorTerminals) {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
agents, new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
Map.of(wcfg.profile(), wcfg), wcfg.profile(),
k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
SessionManager sessions = new SessionManager(workers, new FakeWorktrees());
Injector injector = new Injector(agents);
MessageService messages = new MessageService(agents, injector, new Rendezvous());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, false, null,
() -> leadTerminals, new MemberRegistry(null), t -> null, () -> collaboratorTerminals);
metrics = FleetMetrics.create(sessions, new dev.ltms.fleet.msg.InMemoryReplyInbox());
app = new FleetApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
callers, metrics).build().start("127.0.0.1", 0);
return app.port();
}
// --- loopback-trust: the caller is the primary -------------------------------------------
@Test
@@ -192,6 +728,107 @@ class FleetAppAuthTest {
"draining an inbox is the primary's collection step");
}
/**
* fleetd #669 Unit A: the {@code turnId} shape of {@code POST /sessions/{id}/message} maps to
* {@link Authz.Action#ANSWER}, not the plain {@link Authz.Action#SEND} the test above drives —
* a worker must stay refused on this shape too, exactly as it was refused on the one undivided
* action before the split.
*/
@Test
void aWorkerMayNotAnswerAnotherSessionsBlockedQuestionOverRest() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null);
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null).statusCode(),
"resolving another session's blocked question would be a worker escalating too");
}
/**
* fleetd #689: a caller refused the coarse {@link Authz.Action#SEND} grant is refused on
* {@code SEND} specifically, even on the {@code turnId}-bearing shape that otherwise raises
* the check to {@link Authz.Action#ANSWER} — proving {@code turnId} was never read from the
* body before the refusal (reading it would have changed which action is named in the 403).
* The same caller refused with no body at all gets the identical detail, which could not hold
* if the decision depended on anything read from the body. Control: a caller who IS granted
* reaches past the gate and the body is used normally.
*/
@Test
void aDeniedCallerIsRefusedOnSendEvenWithATurnIdBodyAndNeverReadsTheBody() throws Exception {
int workerPort = start(FakeHerdr.WORKER_PID, false, null); // denied: not primary/architect
HttpResponse<String> withTurnId = send(workerPort, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
assertEquals(403, withTurnId.statusCode());
assertTrue(withTurnId.body().contains("may not SEND"),
"the SEND check must be the one that fired, not ANSWER — ANSWER would only be "
+ "reachable by having already read turnId out of the body");
HttpResponse<String> noBody = send(workerPort, "POST", "/sessions/term_b/message", null, null);
assertEquals(403, noBody.statusCode());
assertTrue(noBody.body().contains("may not SEND"),
"refused identically with no body at all — the refusal cannot depend on body content");
// Control: a primary IS granted SEND, so the same turnId body is read and acted on —
// reaching messages.answer, which reports this unknown turnId as a stale one.
int primaryPort = start(999_999, false, null);
HttpResponse<String> granted = send(primaryPort, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
assertEquals(409, granted.statusCode());
assertTrue(granted.body().contains("stale_turn"), "a granted caller's body IS read and acted on");
}
/**
* fleetd #689 (ticket comment 18353): the only place {@code sendMessage}'s call to {@code
* answerGatePasses} is observable is the audit trail — {@code allow()} logs an {@code
* "allowed"} entry for every granted action except {@code READ}/{@code METRICS}/{@code
* TASK_READ}, and {@code ANSWER} is none of those. A granted {@code turnId} request must
* therefore log both a {@code SEND} and an {@code ANSWER} entry; a granted plain request must
* log {@code SEND} alone. A unit test of the extracted helper pins the helper; this pins the
* call site — deleting the {@code answerGatePasses} call from {@code sendMessage} leaves the
* helper's own test green but turns this one red.
*/
@Test
void aGrantedTurnIdRequestAuditsBothSendAndAnswerButAPlainRequestAuditsSendAlone() throws Exception {
int port = start(999_999, false, null); // primary: granted both SEND and ANSWER
ObjectMapper mapper = new ObjectMapper();
try (CapturedLog audit = CapturedLog.at("audit", Level.INFO)) {
send(port, "POST", "/sessions/term_b/message",
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null);
List<String> allowed = allowedActions(audit, mapper);
assertTrue(allowed.contains("SEND"),
"a turnId request must still clear the coarse SEND grant first");
assertTrue(allowed.contains("ANSWER"),
"a turnId request must ALSO clear the ANSWER grant — this is the call site itself");
}
try (CapturedLog audit = CapturedLog.at("audit", Level.INFO)) {
send(port, "POST", "/sessions/term_b/message",
"{\"content\":\"hi\",\"timeoutMs\":50}", null);
List<String> allowed = allowedActions(audit, mapper);
assertEquals(List.of("SEND"), allowed,
"a plain request must log SEND and nothing else — ANSWER is conditional on "
+ "turnId, not something every request happens to log");
}
}
private static List<String> allowedActions(CapturedLog audit, ObjectMapper mapper) {
return audit.events().stream()
.map(ILoggingEvent::getFormattedMessage)
.map(line -> {
try {
return mapper.readTree(line);
} catch (Exception e) {
throw new AssertionError("audit line is not valid JSON: " + line, e);
}
})
.filter(n -> "allowed".equals(n.path("outcome").asText()))
.map(n -> n.path("action").asText())
.toList();
}
// --- token mode ---------------------------------------------------------------------------
@Test
@@ -675,6 +675,26 @@ class FleetAppTest {
assertEquals(400, postMessage(port, "{}").statusCode());
}
/**
* fleetd #689: a body that fails to parse is rejected with 400 before {@code turnId} is ever
* read from it, so it reaches neither {@code messages.answer} (which needs a {@code turnId})
* nor {@code messages.send} — confirmed here for {@code send} by the fake agent's idle status,
* which would otherwise make an immediate {@code agent.prompt} delivery observable.
*/
@Test
void malformedBodyReturns400AndNeverReachesSendOrAnswer() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // idle ⇒ send would deliver right away if reached
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = postMessage(port, "not json at all");
assertEquals(400, res.statusCode());
JsonNode err = mapper.readTree(res.body());
assertEquals("bad_request", err.get("error").asText());
assertEquals("body must be JSON", err.get("detail").asText());
assertFalse(herdr.called("agent.prompt"),
"a malformed body must never reach messages.send's delivery");
}
@Test
void sessionStatusReportsLiveAgentStatus() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("blocked");
@@ -0,0 +1,92 @@
package dev.ltms.fleet.session;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import java.util.List;
import java.util.Set;
/**
* {@link PeerLauncher} decorator that marks presence for a spawned terminal before returning its
* handle to the caller — the contact-then-register ordering fleetd #722 covers, where the
* terminal's MCP contact lands before {@link SessionManager#acquire} runs its own
* {@code registry.put}. The presence view is set after construction, via {@link #presence},
* because it is owned by the {@link SessionManager} this launcher is passed into.
*/
final class PresenceRacingLauncher implements PeerLauncher {
private final PeerLauncher delegate;
volatile MemberPresence presence;
PresenceRacingLauncher(PeerLauncher delegate) {
this.delegate = delegate;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle handle = delegate.spawn(req);
presence.markPresent(handle.terminalId());
return handle;
}
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
PeerHandle handle = delegate.spawn(req, decision);
presence.markPresent(handle.terminalId());
return handle;
}
@Override
public Set<Capability> capabilities() {
return delegate.capabilities();
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return delegate.capabilitiesFor(profileName);
}
@Override
public Set<String> profiles() {
return delegate.profiles();
}
@Override
public String defaultProfile() {
return delegate.defaultProfile();
}
@Override
public String effectiveCwd(SpawnRequest req) {
return delegate.effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return delegate.parityOverlay(profileName);
}
@Override
public List<?> list() {
return delegate.list();
}
@Override
public int reapOrphanWorkers() {
return delegate.reapOrphanWorkers();
}
@Override
public void stop(String id) {
delegate.stop(id);
}
@Override
public boolean clearContext(String id) {
return delegate.clearContext(id);
}
}
@@ -422,6 +422,53 @@ class SessionManagerTest {
"turn completion moves BUSY → DONE");
}
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
@Test
void registerThenContactReachesReadyForPlainSpawn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that arrives after registration reaches READY");
}
@Test
void contactThenRegisterStillReachesReadyForPlainSpawn() {
// The racing launcher marks presence for the spawned terminal from inside spawn() —
// before SessionManager.acquire's own registry.put runs — modeling an MCP contact that
// lands in that window.
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
PresenceRacingLauncher race = new PresenceRacingLauncher(workers);
SessionManager sessions = new SessionManager(race);
race.presence = sessions.asPresence();
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that lands before registry.put must still reach READY");
}
@Test
void aTerminalNeverMarkedPresentStaysSpawningAfterRegistration() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertEquals(MemberSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
"registration alone must not advance a terminal that was never marked present");
}
@Test
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr();
@@ -1405,6 +1452,105 @@ class SessionManagerTest {
+ "dirty check threw");
}
// --- fleetd #736: a release must forget the member's presence entry, not just its registry
// row ---------------------------------------------------------------------------------------
@Test
void releaseByPaneIdForgetsThePresenceEntry() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the release");
sessions.release(session.paneId());
assertFalse(sessions.asPresence().isPresent(terminal),
"release must forget the terminal's presence, not just remove its registry row");
}
@Test
void reapIdleForgetsThePresenceEntryToo() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the reap");
clock[0] = 11;
assertEquals(1, sessions.reapIdle(10), "READY session past TTL is reaped");
assertFalse(sessions.asPresence().isPresent(terminal),
"the idle-reap release path (releaseIfCurrent) goes through the same teardown "
+ "funnel as an explicit release, so it must forget presence too");
}
@Test
void shutdownDrainAlsoForgetsThePresenceEntry() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertTrue(sessions.asPresence().isPresent(terminal), "present before the drain");
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertFalse(sessions.asPresence().isPresent(terminal),
"a shutdown drain still ends the member's process, so presence must be cleared "
+ "exactly as it is for any other release cause");
}
@Test
void releaseOfAnUnknownPaneIdDoesNotThrow() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
assertDoesNotThrow(() -> sessions.release("no-such-pane"),
"releasing a pane id that was never registered must be a no-op, not a throw");
}
@Test
void releaseStillForgetsPresenceWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-736", null));
String terminal = s.terminalId();
sessions.asPresence().markPresent(terminal);
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a throwing dirty check must not abort the release");
assertFalse(sessions.asPresence().isPresent(terminal),
"presence must be forgotten even when the dirty check throws, which pins the "
+ "forget call to the finally block that runs no matter what happened above");
}
@Test
void releaseLeavesADifferentStillLiveMembersPresenceUntouched() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession released = sessions.acquire("ltms-local", null, "/caller/a", "ownerA");
MemberSession stillLive = sessions.acquire("ltms-local", null, "/caller/b", "ownerB");
sessions.asPresence().markPresent(released.terminalId());
sessions.asPresence().markPresent(stillLive.terminalId());
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
"present before the release of the other member");
sessions.release(released.paneId());
assertFalse(sessions.asPresence().isPresent(released.terminalId()),
"the released terminal is forgotten");
assertTrue(sessions.asPresence().isPresent(stillLive.terminalId()),
"a still-live member's presence must survive an unrelated release");
}
// --- fleetd #316: the dirty check must be re-taken after the worker is stopped, not trusted
// stale from before it ------------------------------------------------------------------------
@@ -2436,4 +2582,160 @@ class SessionManagerTest {
return "threw:" + e.getClass().getName() + ":" + e.getMessage();
}
}
// --- fleetd #702: spawnedMemberRole must still answer for a pane mid-teardown ----------------
/**
* A {@link Worktrees} test double whose {@code hasUncommitted} runs an injected hook before
* answering. This is what lets a test resolve a releasing pane's terminal from inside the
* window {@link SessionManager#release} opens between removing the registry entry and
* actually stopping the pane — the hook runs synchronously on the release call's own thread,
* at the exact point {@code release} shells out to {@code git status}, so there is no sleep
* and no race to land in it.
*/
private static final class HookedWorktrees implements Worktrees {
private Runnable hook;
private boolean dirty = false;
HookedWorktrees onHasUncommitted(Runnable hook) {
this.hook = hook;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
}
@Override
public void deleteBranch(String repoRoot, String branch) {
}
@Override
public boolean hasUncommitted(String worktreePath) {
if (hook != null) {
hook.run();
}
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
return java.util.Optional.empty();
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
}
@Test
void spawnedMemberRoleResolvesTheLiveRegisteredRole() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertEquals(s.role(), sessions.spawnedMemberRole(s.terminalId()),
"a registered session resolves to its own role");
assertNull(sessions.spawnedMemberRole("term_unknown"),
"a terminal with no session at all resolves to null");
}
@Test
void spawnedMemberRoleIsNullOnceReleaseFullyCompletes() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, new HookedWorktrees());
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702a", null));
sessions.release(s.paneId());
assertNull(sessions.spawnedMemberRole(s.terminalId()),
"once release has fully finished, the terminal is neither registered nor releasing");
}
@Test
void spawnedMemberRoleStillAnswersBetweenTheRegistryRemovalAndThePaneStop() {
FakeHerdr herdr = new FakeHerdr();
java.util.concurrent.atomic.AtomicReference<MemberRole> duringWindow = new java.util.concurrent.atomic.AtomicReference<>();
HookedWorktrees worktrees = new HookedWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("fleetd-702b", null));
worktrees.onHasUncommitted(() -> {
// This runs from INSIDE release()'s git-status shell-out: the registry entry is
// already gone, but the pane has not stopped yet — a real call landing inside the
// exact window fleetd #702 reports, so no sleep and no race is needed to reach it.
assertTrue(sessions.get(s.paneId()).isEmpty(),
"sanity: the registry entry is already gone at this point");
duringWindow.set(sessions.spawnedMemberRole(s.terminalId()));
});
sessions.release(s.paneId());
assertEquals(s.role(), duringWindow.get(),
"spawnedMemberRole must still answer the live role while the pane is mid-teardown, "
+ "not only while the session is still in the registry");
assertNull(sessions.spawnedMemberRole(s.terminalId()),
"and once release has fully finished, the window is closed too");
}
/**
* {@code Releasing.enter} and {@code Releasing.leave} are called directly here, never through
* {@link SessionManager#release} or {@link SessionManager#spawnedMemberRole}, so a test naming
* one of them exercises only that one — a regression in the other can never hide behind it.
*/
@Test
void releasingLeaveStepsDownADepthGreaterThanOneInsteadOfRemovingIt() {
SessionManager.Releasing depthTwo = new SessionManager.Releasing(2, "term_a", MemberRole.DEV);
SessionManager.Releasing afterLeave = depthTwo.leave();
assertNotNull(afterLeave,
"depth 2 means another release of the SAME pane is still mid-teardown; leave() must "
+ "step the depth down, never remove the marker outright — removing it here "
+ "is what a plain Set would do, and would reopen the window the still-in-"
+ "flight release is relying on staying closed");
assertEquals(1, afterLeave.depth());
assertEquals("term_a", afterLeave.terminalId());
assertEquals(MemberRole.DEV, afterLeave.role());
}
@Test
void releasingEnterPreservesThePriorTerminalWhenTheOverlappingCallHasNoSessionOfItsOwn() {
SessionManager.Releasing prior = new SessionManager.Releasing(1, "term_a", MemberRole.DEV);
SessionManager.Releasing afterEnter = SessionManager.Releasing.enter(prior, null);
assertEquals(2, afterEnter.depth(), "depth still increments whether or not this enter knows its session");
assertEquals("term_a", afterEnter.terminalId(),
"an overlapping release that finds the registry entry already gone has no session of "
+ "its own to pass as known, and must not blank out the terminal the first "
+ "call already recorded — that terminal is what spawnedMemberRole matches "
+ "against");
assertEquals(MemberRole.DEV, afterEnter.role());
}
}
@@ -159,6 +159,40 @@ class WorktreeSessionManagerTest {
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
// --- fleetd #722: registration and presence must reach READY whichever lands first --------
@Test
void registerThenContactReachesReadyForWorktreeSpawn() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-722", null));
sessions.asPresence().markPresent(session.terminalId());
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that arrives after worktree registration reaches READY");
}
@Test
void contactThenRegisterStillReachesReadyForWorktreeSpawn() {
// The racing launcher marks presence for the spawned terminal from inside spawn() —
// before SessionManager.acquireWithWorktree's own registry.put runs — modeling an MCP
// contact that lands in that window.
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
PresenceRacingLauncher race = new PresenceRacingLauncher(workerService(herdr));
SessionManager sessions = new SessionManager(race, worktrees);
race.presence = sessions.asPresence();
MemberSession session = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-722", null));
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"a presence contact that lands before worktree registration must still reach READY");
}
@Test
void worktreeArchitectAcquireAlsoBindsItsSlot() {
FakeHerdr herdr = new FakeHerdr();
+22 -5
View File
@@ -213,6 +213,17 @@ map_masked_lines() {
done < "$file"
}
# Masks every `scheme://user:pass@host` userinfo on one line of text, replacing just that
# userinfo with `<redacted>` and leaving the rest of the line untouched, byte for byte. The
# pattern stops at the first `/`, whitespace, or `@` reached after `://` — a URI's userinfo
# component cannot contain any of those three characters — so a URI with no userinfo, followed
# later on the same line by an unrelated `@`, never matches. The `g` flag matters: a line can
# carry more than one URI. Shared by every caller that prints a line which may hold a
# credentialed URI, so the bound lives in exactly one place.
mask_url_userinfo() {
printf '%s\n' "$1" | sed -E 's#://[^@/[:space:]]*@#://<redacted>@#g'
}
redact() {
local old_file="$1" new_file="$2"
local line prefix content indent lead key old_line=0 new_line=0 in_hunk=0
@@ -270,7 +281,7 @@ redact() {
continue
fi
fi
printf '%s\n' "$line" | sed -E 's#://[^@]*@#://<redacted>@#g'
mask_url_userinfo "$line"
done
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
}
@@ -580,6 +591,12 @@ install_candidate() {
# -------------------------------------------------------------------------------- the report path
#
# Masks basic-auth userinfo (scheme://user:pass@host) in a daemon verdict line before it reaches
# the terminal.
mask_verdict_userinfo() {
mask_url_userinfo "$1"
}
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
restore_command_line() {
@@ -599,8 +616,8 @@ restore_and_confirm() {
ok "restored from $backup"
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
case "$VERDICT_KIND" in
refused) warn "the RESTORE was also refused by the daemon: $VERDICT_LINE" ;;
*) ok "restore confirmed: $VERDICT_LINE" ;;
refused) warn "the RESTORE was also refused by the daemon: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
*) ok "restore confirmed: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
esac
else
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
@@ -616,7 +633,7 @@ report_outcome() {
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
kind="$VERDICT_KIND"; line="$VERDICT_LINE"
kind="$VERDICT_KIND"; line="$(mask_verdict_userinfo "$VERDICT_LINE")"
else
kind="none"
fi
@@ -673,7 +690,7 @@ check_mode() {
local verdict
verdict="$(last_verdict_line "$LOG")"
if [ -n "$verdict" ]; then
ok "last verdict in log: $verdict"
ok "last verdict in log: $(mask_verdict_userinfo "$verdict")"
else
warn "no reload verdict line found in $LOG"
fi
+131
View File
@@ -233,6 +233,66 @@ test_redaction_holds() {
assert_contains "weight" "$RUN_OUTPUT" "a diff must have been demonstrably printed at all"
}
# redact()'s key-name filter only inspects the KEY, so a diff line whose key does not match
# TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY still reaches the final userinfo
# sed even when its VALUE holds a credentialed URI. "note" is not a sensitive key name, so this
# line must fall all the way through to that sed, not the earlier whole-value branch. The
# trailing prose on both sides of the userinfo is a positive control: it proves the line reached
# the userinfo sed (which touches only the userinfo) rather than the earlier branch (which would
# have replaced the whole value with a bare "<redacted>" and dropped the prose).
test_diff_line_userinfo_is_masked_with_positive_control() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.note=see amqp://alice:wonderland@rabbit.local:5672/vhost for details'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "diff-userinfo-case reload exit code"
assert_not_contains "alice:wonderland" "$RUN_OUTPUT" "the userinfo must never reach the output"
assert_contains "amqp://<redacted>@rabbit.local:5672/vhost" "$RUN_OUTPUT" \
"the userinfo must be MASKED, not deleted — the rest of the value must survive"
assert_contains "note:" "$RUN_OUTPUT" "the key name must still reach the output"
assert_contains "see " "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
assert_contains "for details" "$RUN_OUTPUT" "prose AFTER the userinfo must still reach the output"
}
# A diff line can hold a URL with no userinfo, followed later on the same line by an unrelated @
# (free text in a string value, for example an email address). The line must pass through the
# userinfo sed byte for byte: the match must stop at the end of the URL and must not treat the
# later @ as a second userinfo delimiter.
test_diff_line_uri_without_userinfo_survives_a_later_at_sign() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.note2=see https://docs.local/guide and mail ops@example.com'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "diff-no-userinfo-with-later-at-sign reload exit code"
assert_contains "note2: see https://docs.local/guide and mail ops@example.com" "$RUN_OUTPUT" \
"a URL with no userinfo plus a later @ on the same line must pass through byte for byte"
}
# Two credentialed URIs on one diff line must both be masked — the g flag matters.
test_diff_line_masks_multiple_userinfo_with_g_flag() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.note3=amqp://u1:p1@host1/vhost1 and amqp://u2:p2@host2/vhost2'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "diff-two-userinfo-on-one-line reload exit code"
assert_not_contains "u1:p1" "$RUN_OUTPUT" "the first userinfo must never reach the output"
assert_not_contains "u2:p2" "$RUN_OUTPUT" "the second userinfo must never reach the output"
assert_contains "amqp://<redacted>@host1/vhost1" "$RUN_OUTPUT" "the first URI must be masked"
assert_contains "amqp://<redacted>@host2/vhost2" "$RUN_OUTPUT" "the second URI must be masked"
}
# ------------------------------------------------------- acceptance criterion 9: forgotten value
# `--set .a.b=` is a plausible typo (the value simply forgotten), and it must be refused outright
# rather than silently nulling the field — a null numeric field falls back to its default, which
@@ -629,6 +689,65 @@ test_refusal_shape_from_parse_failure_wording_is_recognised() {
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
}
# A verdict line carrying a credentialed URI has its userinfo masked, with a positive control
# proving the rest of the line still reaches the output unchanged.
test_verdict_userinfo_is_masked_with_positive_control() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reload from %s refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern ("amqp://user:hunter2@host/vhost"): Unclosed character class near index 8\n' \
"$dir/fleetd.yaml" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "refusal-with-userinfo exit code"
assert_not_contains "user:hunter2" "$RUN_OUTPUT" "the userinfo must never reach the output"
assert_contains "amqp://<redacted>@host/vhost" "$RUN_OUTPUT" \
"the userinfo must be MASKED, not deleted — the rest of the quoted value must survive"
# Positive control: the diagnostic prose on both sides of the userinfo must still reach the
# output. Without this, a mutant that drops the whole verdict line would pass identically.
assert_contains "malformed pattern" "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
assert_contains "Unclosed character class near index 8" "$RUN_OUTPUT" \
"prose AFTER the userinfo must still reach the output"
}
# An ordinary refusal line quotes the offending pattern, not a credential, and must survive byte
# for byte: the rewrite is scoped to userinfo only, and the quoted pattern is the detail an
# operator needs to fix the refusal.
test_ordinary_refusal_line_passes_through_unchanged() {
local dir real_line
dir="$(new_fixture)"
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern (\"[unclosed\"): Unclosed character class near index 8"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "ordinary refusal exit code"
assert_contains "$real_line" "$RUN_OUTPUT" \
"an ordinary refusal with no userinfo must pass through byte for byte, unchanged"
}
# A verdict line can hold a URI with NO userinfo and a later, unrelated @ further on in the same
# line (an email address in diagnostic prose, for example). The rewrite must stop at the end of
# the URI and must not treat the later @ as a second userinfo delimiter.
test_uri_without_userinfo_survives_a_later_at_sign() {
local dir real_line
dir="$(new_fixture)"
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: broker.uri amqp://broker.local/vhost unreachable, contact ops@example.com"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 4 "$RUN_RC" "no-userinfo-with-later-at-sign exit code"
assert_contains "$real_line" "$RUN_OUTPUT" \
"a URI with no userinfo plus a later @ in the same line must pass through byte for byte"
}
# --set runs yq over the whole candidate. It warns when that changes more lines than the requested
# pairs, but a simple file with only the intended changed line must stay quiet.
new_fixture_reformat_sensitive() {
@@ -694,6 +813,12 @@ echo "== acceptance criterion 6: the marker works =="
test_marker_skips_lines_before_it
echo "== acceptance criterion 7 (+13: redaction is proven to have run) =="
test_redaction_holds
echo "== fleetd #692: a diff line's userinfo is masked, rest of the value survives =="
test_diff_line_userinfo_is_masked_with_positive_control
echo "== fleetd #692: a diff line's URI with no userinfo survives a later @ in the line =="
test_diff_line_uri_without_userinfo_survives_a_later_at_sign
echo "== fleetd #692: two userinfo URIs on one diff line are both masked =="
test_diff_line_masks_multiple_userinfo_with_g_flag
echo "== acceptance criterion 9: a forgotten value refuses and installs nothing =="
test_forgotten_value_refuses_and_installs_nothing
echo "== acceptance criterion 10: an explicit clear writes a bare null =="
@@ -720,6 +845,12 @@ echo "== extra: --check is read-only and always exits 0 =="
test_check_is_read_only_and_exits_zero
echo "== extra: the parse-failure refusal shape is also recognised =="
test_refusal_shape_from_parse_failure_wording_is_recognised
echo "== verdict-redaction criteria 2+3: verdict userinfo is masked, rest of line survives =="
test_verdict_userinfo_is_masked_with_positive_control
echo "== verdict-redaction criterion 4: an ordinary refusal passes through unchanged =="
test_ordinary_refusal_line_passes_through_unchanged
echo "== fleetd #638: a URI with no userinfo survives a later @ in the same line =="
test_uri_without_userinfo_survives_a_later_at_sign
echo "== acceptance criterion 17: --set warns about yq formatting churn =="
test_set_warns_when_yq_reformats_extra_lines
echo "== acceptance criterion 18: --set stays quiet without formatting churn =="