Compare commits

...

26 Commits

Author SHA1 Message Date
Dai Ha d223a93039 fleetd #184: correct sshAuthSock guidance
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Failing after 1m37s
2026-09-03 16:38:56 +07:00
Dai Ha 96d8191149 Merge #263: free reports what the spawn gate grants (#257)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m13s
2026-09-03 16:33:20 +07:00
Dai Ha f0e7ac73d6 #103: give the operator the string that is actually in the log
The merged fix told the operator to look in journalctl for
"Fleetd.reportRequiredSecrets". That string never appears there — it is a
method name. The logger is d.l.f.Fleetd and the lines read
"startup secret NAME: set|MISSING", so the advice sent the operator looking
for text that does not exist.

Give the grep instead, and say what the report does not cover: it lists only
names a configured profile references, so a secret nothing references is never
reported.
2026-09-03 16:32:23 +07:00
Dai Ha e2fe861b4d Merge #262: the systemd unit names all three secrets it needs (#103) 2026-09-03 16:32:07 +07:00
Dai Ha 279d6f5fbd fleetd #257: free must stop subtracting leadSeatCount
CI / build (pull_request) Successful in 1m43s
CI / contract (pull_request) Successful in 1m51s
fleet_list's free row subtracted leadSeatCount, but the real spawn gate
(CompositePeerLauncher#enforceMaxLoad) only ever compares live against
maxLoad and never reads leadSeatCount. So free could report 0 while a
fleet_spawn on that exact profile still succeeded, and a lead trusting
free:0 gave up on capacity the gate would still grant.

free now always equals max(0, maxLoad - live); leadSeats stays in the
row as an informational fact, never subtracted. Documented in the
fleet_list tool description and fleetd.example.yaml.
2026-09-03 16:28:32 +07:00
Dai Ha 34480cebef Merge #261: compare-and-swap the workspace-trust seed against an external writer (#247)
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m40s
2026-09-03 16:26:37 +07:00
Dai Ha 9d0bf14c46 fleetd #103: document systemd worker secrets
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m31s
2026-09-03 16:23:29 +07:00
Dai Ha 5a3ab5764c fleetd #247: CAS seedTrustDialog against the operator's own live Claude Code
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m49s
TRUST_JSON_LOCK only serialises seedTrustDialog calls this launcher makes
inside its own JVM. It cannot reach the one writer that actually shares the
target file on a real host: the operator's own live Claude Code, whose
CLAUDE_CONFIG_DIR is routinely the very configDir a profile is given, so the
file fleetd writes on every claude-code spawn is that session's own config.
A plain read-modify-write there is a routine lost update: fleetd reads v1,
the operator's session writes v2, fleetd's ATOMIC_MOVE lands v3 built from
v1 and v2 is gone, atomically.

Add a bounded compare-and-swap: before the move, re-read the target's exact
bytes and compare with what the update was built from; on a mismatch,
rebuild from the fresh bytes and retry (up to 5 attempts). Exhausting the
retries writes nothing and logs a WARN naming the file — the member shows
the trust dialog and fails to reach an injectable state instead, which is
visible and recoverable, unlike silently overwriting the operator's live
config. Also warn every time the seed is about to target the default
~/.claude.json (configDir unset), since that is the unsafe default.

This narrows the lost-update window, it does not close it: a write landing
between the final re-read and the ATOMIC_MOVE itself is still lost.
2026-09-03 16:20:20 +07:00
Dai Ha 3bad9f5785 Merge #260: restore the already-gone-worktree teardown test (#116)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 2m9s
2026-09-03 16:13:15 +07:00
Dai Ha 8308c0b68f fleetd #116: recover the already-gone-worktree teardown regression test
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m23s
Ports the intent of the lost CB-576 commit c393600 (worker/cb576-01a04b-17,
never merged, package dev.ltms.bridged.*) onto main's dev.ltms.fleet.*
tree. Adds FakeWorktrees.markGone (tracks which add()'d worktree paths
still "exist", mirroring GitWorktrees.hasUncommitted's Files.exists guard
for the already-gone case) and a regression test,
releaseStillStopsPaneAndNotifiesWhenWorktreeIsGone, asserting that
SessionManager.release still fires notifyReleased, stops the pane, and
falls through to remove() when the worktree is already gone.
2026-09-03 16:11:27 +07:00
Dai Ha 3b3063eb2b #252: the route scrape counted commented-out registrations as live
CI / build (push) Successful in 1m44s
CI / contract (push) Successful in 2m21s
Follow-up to #259, found by verifying the guard rather than trusting it.

The scrape read FleetApp's raw source, so a registration disabled with `//`
still matched. Measured: commenting out `app.get("/tasks/{ticket}", ...)` left
the test GREEN, while deleting the same line was caught. Only the commented-out
shape was blind, and it is the silent direction — the inventory would keep
claiming a route the server no longer serves.

Drop whole-line comments before scraping. Only lines starting with //, * or /*
are dropped, deliberately not every // on a line: that would also cut a string
literal containing // (a URL) and could silently delete a real registration
sharing the line. The remaining gap is a trailing comment beside real code; no
registration here has that shape, and the vacuity test catches a scrape that
loses registrations wholesale.

Proof: with the fix, the comment-out mutation fails naming
"Removed ... [GET /tasks/{ticket}]"; reverted, FleetApp.java confirmed clean.
mvn clean install: 0 compile errors, 1257 tests, BUILD SUCCESS.
2026-09-03 15:59:35 +07:00
Dai Ha df9086263d Merge #259: guard the REST route inventory against drift (#252) 2026-09-03 15:56:15 +07:00
Dai Ha eb568ff451 fleetd #252: guard test for the REST route inventory
CI / contract (pull_request) Successful in 1m16s
CI / build (pull_request) Successful in 1m27s
FleetApp's route list has never been checked against anything and has
already drifted once (GET /member-credentials shipped hours before the
#252 ticket and was missing from its list). Add
RestRouteInventoryTest, modelled on McpContractDocTest, which scrapes
FleetApp.java's app.<verb>("path") calls with a regex and compares
them against an explicit expected inventory, failing loudly with the
added/removed routes when they diverge.
2026-09-03 15:54:45 +07:00
Dai Ha ac790e4cce #258: stop two test fixtures writing the operator's real ~/.claude.json
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m29s
seedTrustDialog targets ~/.claude.json when a profile sets no configDir.
Two IDE-overlay fixtures built a worktree-shaped @TempDir, which opens the
#149 isProvisionedWorktree gate, and left configDir null — so every test run
added two project entries to the operator's real file. 116 had accumulated,
none of them still existing on disk, and 32 of those came from current code.

The #149 gate only closed the opposite case: a fixture with cwd unset falling
back to user.dir. A fixture that builds a worktree on purpose walks straight
through it.

- ideProfile/ideProfileModule now take configDir first and mandatory, so each
  fixture states where the trust seed goes.
- noFixtureSeededTheDefaultClaudeJson snapshots the temp-dir project keys in
  @BeforeAll and fails in @AfterAll on any key this class added. Differential,
  not absolute: an absolute check would fail on every host still carrying the
  historical entries, and such a check gets deleted rather than fixed.

Not done: a blanket -Duser.home redirect in surefire. EnvAllowListScrubTest
tests the credential scrub against the operator's real login chain and guards
with assumeTrue($HOME/.zshrc exists), so the redirect would silently skip two
security tests.

Proof: with the bug put back on one fixture the guard fails and names the path;
reverted and confirmed identical with diff -q. A full suite run under a fake
home now creates no .claude.json at all. mvn clean install: 0 compile errors,
1255 tests, BUILD SUCCESS.
2026-09-03 14:01:45 +07:00
Dai Ha 6b5f3f472f Merge #111: the credential probe reads the policy instead of copying it
CI / contract (push) Successful in 1m19s
CI / build (push) Successful in 1m30s
scripts/probe-member-credentials.sh carried its own NAMES array of 31
names. The live policy has 34. The probe reported 26 blocked against a
policy that blocks 29, exited 0, and printed a table that looked
complete. A verification tool that under-reports is worse than none,
because its clean output stops anyone looking.

Same defect as #114, fixed the same way: DELETE the second copy rather
than correct it. The NAMES array is gone, not updated.

The daemon now serves GET /member-credentials — names and counts, never
a value; MemberCredentialPolicyView reads no environment at all, so
there is nothing to redact by construction. The probe fetches it and
refuses with a non-zero exit when the daemon is unreachable, the policy
is absent or empty, or knownCount disagrees with the length of known[].
No local fallback: a verification tool must not quietly degrade into a
weaker check.

The startup log line and the endpoint now share that one class, so the
counting exists once. That also protects a subtlety I measured before
briefing this: blocked is NOT known - allowed. Live, known=34 and
allow=7, but only 5 of those 7 appear in known, so blocked=29 and the
naive subtraction gives 27. The view reuses creds.blockedSet(), the
existing derivation, so it keeps 29.

Worker's mutation: MemberCredentialPolicyView.of(...) forced to return
ABSENT turned 4 tests red with 0 compile errors — including
MemberCredentialsGapReportTest, which proves the startup log really
does run through this path. Reverted and confirmed with diff -q.

It also caught a bug in its own first draft: jq's // operator treats
false and 0 as missing, so `.present // empty` turned a genuine
"present": false into "unknown". Fixed by reading the fields directly.

NOT yet verified: acceptance criterion 5, the live 34/29/5 run. The
route does not exist until the daemon is redeployed onto this jar, so
that check comes next and is mine, not the worker's.

Merged clean, then built on the merged tree: 1255 tests, 0 failures,
0 compile errors.
2026-09-03 13:32:16 +07:00
Dai Ha 38dec72152 Merge #155: refuse the spawn when allow-list policy cannot be enforced
CI / contract (push) Successful in 46s
CI / build (push) Successful in 2m15s
policy=allow-list is enforced by a ZDOTDIR scrub, and a non-zsh login
shell ignores ZDOTDIR entirely, so no scrub runs. The launcher already
DETECTED this and logged a WARN — then degraded to the weaker overlay
and spawned anyway. The operator asked for the blocking control and
silently got the weaker one, which is the defect the ticket is about.

Detection existed; refusal did not. Under policy=allow-list a non-zsh
shell now throws IllegalArgumentException before any ZDOTDIR or env
work, naming the actual shell and giving three ways out. Under
policy=deny-by-default nothing changes: that overlay is applied to the
pane before any shell runs, so it does not depend on the shell.

Checked against the live config myself, because this refuses spawns and
no worker can see fleetd.yaml:

  memberHerdrSocket : NOT set  -> the shell comes from fleetd's own
                                  $SHELL, not the unset memberLoginShell
  policy            : allow-list
  fleetd's $SHELL   : zsh, proven by behaviour rather than by reading
                      the process env — the daemon log shows the ZDOTDIR
                      scrub generating a directory 147 times, most
                      recently minutes ago, and that only happens when
                      isZshShell() returned true

So the new refusal cannot fire on this host. Had memberHerdrSocket been
set, the unset memberLoginShell would have read as "<unset>", non-zsh,
and refused every spawn — worth knowing before anyone sets that key.

Worker's mutation evidence, re-stated: `if (!zsh)` -> `if (false)` turned
the refusal test RED with 0 compile errors, then reverted clean.

NOT verified: a live non-zsh member spawn. Forcing it means changing the
daemon's own environment, and the value of the test does not justify
that. The unit tests drive the real launcher.spawn entry point.

Merged clean, then built on the merged tree: 1250 tests, 0 failures,
0 compile errors.
2026-09-03 13:29:53 +07:00
Dai Ha 0e8bfb74fc Merge #176 stage 2: group subscription profiles by account, not by name
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m34s
Stage 1 shipped INERT on this host and every test was green. The matcher
compared effectiveCredentialId(), which fell back to the profile's own
NAME when credentialId was unset. This host runs the lead on `opus` and
members on `sonnet`; both are subscription:true with no credentialId, so
it compared "opus" against "sonnet", never matched, and charged 0 seats.

Every stage-1 test put the lead on the SAME profile name as the target,
so the fixture encoded the one shape the live config does not have.

Stage 2 returns a "<subscription>" sentinel when credentialId is unset
and subscription is true. An explicit credentialId still wins, so an
operator with two genuinely separate Claude logins can keep them apart.

Verified by me on the live config shape, not by reasoning:

  opus.effectiveCredentialId()   = <subscription>
  sonnet.effectiveCredentialId() = <subscription>
  seats charged to sonnet = 1     (was 0 before this change)

Only opus and sonnet join the sentinel group on this host; local,
local-direct, gx, xf, sol and terra are unaffected. free is clamped
with Math.max(0, ...), so the subtraction cannot report a negative.

Second, wider consequence, flagged by the worker and confirmed here:
CompositePeerLauncher.credentialIdFor feeds enforceNotQuarantined and
enforceNotCoolingOff, so quarantining one subscription profile now also
refuses spawns on the other. That is correct — one Claude subscription
hitting a usage limit really does take out every profile on it — but it
is a behavioural change beyond fleet_list's numbers.

Checked all 5 logical callers of effectiveCredentialId(); every one
wants "this account", none wants "this exact profile".

Merged clean, then built: 1250 tests, 0 failures, 0 compile errors.
An auto-merge with no conflicts is not a compiling merge, so the build
was run on the merged tree before this landed.
2026-09-03 13:24:19 +07:00
Dai Ha 51f7b0a3ca fleetd #111: probe reads the live memberCredentials policy, no hardcoded name list
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m22s
scripts/probe-member-credentials.sh carried its own hand-maintained NAMES array
(31 names, recorded 2026-08-16), so a name added later to fleetd.yaml's
memberCredentials.known was never checked and the probe still exited 0 with a
clean-looking table. Same drift shape as #114's tool catalogue.

- New dev.ltms.fleet.member.MemberCredentialPolicyView: the single place that
  turns a MemberCredentials policy into names + counts (never a value). Reused
  by Fleetd.reportMemberCredentialsGap (startup log line) and by the new
  GET /member-credentials REST endpoint (FleetApp), so the two can no longer
  drift apart the way the probe and the policy did.
- FleetApp gains one route + handler + a Supplier<MemberCredentialPolicyView>
  constructor param (legacy constructors default to ::absent, so existing call
  sites are unaffected).
- probe-member-credentials.sh now fetches its name list from
  GET /member-credentials instead of carrying one. No local fallback: an
  unreachable daemon, an empty/absent policy, or a knownCount/known[] length
  mismatch all refuse with a non-zero exit rather than silently checking zero
  names. Prints "policy contains N; this run checked N" so the two numbers are
  visibly equal.
2026-09-03 13:21:35 +07:00
Dai Ha 21c539f22e #113: derive the config guard from the record tree, both directions
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m48s
The existing guard walks KNOWN_TOP_LEVEL_KEYS and anchors its regex at
column 0, so it sees only top-level keys. Every nested key was outside
its scope and nothing said so, which is the shape #113 collects: a
checker narrower than it looks, whose green run stops anyone looking.

Two derived guards replace the assumption:

  everyNestedConfigKeyIsDocumentedInTheExample
      walks FleetConfig's record components (17 records, 83 distinct
      key names) and requires each to be documented in the example.

  everyLiveKeyInTheExampleBindsToARecordComponent
      resolves every live key path in the example against the record
      tree, so a documented key that binds to nothing fails here
      instead of being silently ignored in production.

Neither carries a list, so a key added to any nested record is covered
the moment it compiles (criterion 2).

Both mutations run through the real caller, not the helper (criterion 1,
which asks for exactly that):

  removed every mention of paneProbeIntervalSeconds from the example
      -> FAILS, naming health.paneProbeIntervalSeconds
  added a live bind.totallyMadeUpKnob to the example
      -> FAILS, naming bind.totallyMadeUpKnob

0 compile errors in both; both reverted and confirmed with diff -q.
The first attempt at mutation 1 removed only the `key:` line and the
run stayed green — correctly, because the key was still documented in
prose. An incomplete mutation proves nothing, so it was redone.

Denominators (criterion 3): both guards print how many keys they
checked, and the floor for "did the walk descend?" is derived from
KNOWN_TOP_LEVEL_KEYS.size() rather than being a literal.

Scope is stated in the javadoc rather than implied: the guards do not
check a key sits at the right path, do not parse commented prose for
the reverse direction, and do not prove a parsed key is read by
anything. paneProbeIntervalSeconds is parsed and read by nothing, and
these guards pass it -- the example already says so in its own text.

everyOptionalKnobDocumentedInTheExampleBinds keeps its hand-written
list but is re-documented as a value-binding spot check, explicitly
not a coverage guard; coverage now comes from the two derived tests.

broker.uri is documented only in the example's prose convention
(`#  uri  -> ...`), never as a copy-pasteable `uri:` key, because
writing it out invites pasting a password into a file -- the thing
uriEnv exists to avoid. The matcher accepts that convention rather
than pushing the file toward doing it.

Full build: 1236 tests, 0 failures, 0 compile errors.
2026-09-03 13:18:50 +07:00
Dai Ha bd2774b5f1 fleetd #155: refuse a member spawn under memberCredentials.policy=allow-list on a non-zsh shell
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m56s
The ZDOTDIR scrub that enforces policy=allow-list only runs on zsh. The daemon
already detected a non-zsh login shell (isZshShell/warnNonZsh, from #213), but
degraded to the weaker CB-596 overlay and spawned anyway — the exact "control
silently does nothing" defect this ticket is about. Now a non-zsh shell under
policy=allow-list refuses the spawn (IllegalArgumentException, naming the
shell), surfaced by FleetMcp.spawn's existing catch(IllegalArgumentException).
policy=deny-by-default is unaffected in substance (its overlay never depended
on the shell) but now also logs a one-time WARN naming the shell, since the
stronger allow-list control is unavailable there.

The worktreeRoot/worktreeGroup-missing degrade path under memberHerdrSocket
is untouched — that gap is fleetd #213's scope, not this one.
2026-09-03 13:13:23 +07:00
Dai Ha c50f5b2d61 fleetd #176 stage 2: make effectiveCredentialId() subscription-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m44s
Stage 1's lead-seat matcher (leadSeatLookup) was correct but inert on
the live host: the lead runs on profile 'opus', members on 'sonnet',
both subscription:true with no explicit credentialId. Because
effectiveCredentialId() fell back to the profile's own name, opus and
sonnet never matched even though they share one Claude login, so the
matcher charged zero seats.

FleetConfig.Profile.effectiveCredentialId() now falls back to a shared
sentinel (SUBSCRIPTION_CREDENTIAL_ID = "<subscription>") instead of the
profile name when subscription:true and credentialId is unset. An
explicit credentialId still wins, so two separate Claude logins on one
host can still be kept apart.

This is also BackendQuarantine's and BackendOutagePolicy's grouping
key and CompositePeerLauncher's spawn-time enforcement key, so the fix
also links quarantine/cool-off across subscription profiles sharing an
account -- intentional: one usage limit really does take out every
profile on that login, mirroring credentialId: openai-shared already
doing this for off-subscription profiles. Every caller was reviewed;
none wants "this exact profile" over "this account".

Tests added:
- FleetdLeadSeatLookupTest: the live shape itself (lead on a
  DIFFERENT subscription profile than the target, same account,
  neither sets credentialId) -- the case stage 1's suite never covered
- FleetMcpTest: quarantining one subscription profile's shared
  account zeroes free on another sharing it, via the same
  effectiveCredentialId()-driven wiring Fleetd.main uses

Mutation-tested: reverting the subscription branch to the old
fall-back-to-profile-name behavior sends both new tests RED with 0
compile errors; reverting the mutation restores byte-identical
(diff -q) source and green tests.

fleetd.example.yaml's fleetd #176 notes are rewritten for the sentinel
semantics and when to override it with an explicit credentialId.
2026-09-03 13:10:10 +07:00
Dai Ha 01a840cc14 #248 follow-up: drive the real backendErrorSink, not a copy of it
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m29s
BackendOutageFlowTest held a ~30-line hand-copy of the lambda in
Fleetd.main, under a comment promising it mirrored production "EXACTLY".
That promise was the defect. The test proved the copy, so any change to
the real sink left the flow test green.

#248 made Fleetd.backendErrorSink(...) public for exactly this reason.
The test now calls it.

Measured, same mutation in the real sink (an early return after
sessions.onBackendError, dropping the cool-off and the lead nudge):

  old test (hand-copy):  Tests run: 5, Failures: 0  -- blind
  new test (real sink):  Tests run: 5, Failures: 4  -- catches it

0 compile errors in both runs, so both are real results. Production
reverted and confirmed with diff -q.

Full build: 1234 tests, 0 failures, 0 compile errors.
2026-09-03 13:06:11 +07:00
Dai Ha c4deef08be fleetd #249: withhold agentSessionId when the cwd is not a provisioned worktree
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m22s
A member spawned without a worktree inherits the lead's cwd, which holds
many old opencode session rows. sessionIdForDirectory picks the most
recently updated row for that directory, so a brand-new member - which
has not written its own row yet - resolves to somebody else's session.
Measured: a row three days old, from a different profile.

The damage was at the tool surface. fleet_list told the lead that
agentSessionId is the id to pass as resumeSessionId, so acting on it
would resume a stranger's conversation, with foreign context, and
nothing to distinguish that from a correct resume.

Fixed by refusing to answer rather than by making the heuristic smarter.
#234 already established the heuristic cannot be made reliable at that
layer, and its javadoc records why, so the SQL is untouched.

agentSessionId() now returns null for a non-provisioned cwd, and
spawn() refuses a resumeSessionId request for one outright, before
anything starts. isProvisionedWorktree moved to HerdrPeerLauncher so
both adapters share it. fleet_list and fleet_spawn descriptions no
longer describe the id as always safe to resume.

Verified rather than taken on trust:
- the refusal reaches the lead as a readable message, not a stack trace
  - FleetMcp.spawn already catches IllegalArgumentException and returns
  error(e.getMessage()).
- the message tells the lead to pass fleet_spawn{worktree:<slug>}, which
  is valid: worktree is typed string, 'true' or a ticket slug.
- 13 existing tests moved off a placeholder "/work/dir" onto a real
  provisioned-worktree fixture. They cover #175/#234 model-mismatch
  machinery and would otherwise have tripped the new gate incidentally.

Worker's mutation evidence, both reverted and diff-confirmed:
- gate at OpenCodeLauncher:812 -> if(false): RED at
  OpenCodeLauncherTest:448, expected <null> but was <ses_someone_elses>.
- resume refusal at OpenCodeLauncher:683 -> 'false &&': RED at
  OpenCodeLauncherTest:379, expected IllegalArgumentException.
Both 0 compile errors.

Closes #249. PR #253.
2026-09-03 12:55:29 +07:00
Dai Ha c796eac09c fleetd #176: subtract the lead's own subscription seat from free
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m16s
maxLoad counted panes, never subscription seats: a subscription:true
profile's lead is itself a live claude session on that same account,
so free overstated capacity by the lead's own seat (measured free:1
with a real ceiling of 0, and free:3 on an idle fleet with a real
ceiling of 2).

Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/
OutageSource) and Fleetd.leadSeatLookup, which derives the seat count
from fleet.leaders.<name>.profile matched against the target profile
by effectiveCredentialId() - no hardcoded "-1", and no new config key:
profile: already exists for this exact "which account does this lead
share" question. maxLoad itself is left untouched; only free (and a
new, additive-only leadSeats field) changes.

Exhaustion quarantine (cause 2 in the ticket) already forced free to 0
via the same BackendQuarantine capacityView already reads - confirmed
by reading the exhaustionSink wiring, no code change needed there.
2026-09-03 12:55:08 +07:00
Dai Ha e897e5257b fleetd #114: delete the drifted tool catalogue, keep the flows, guard the names
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
docs/MCP-Contract.md was written 2026-07-14, before any MCP code existed,
and never caught up. CLAUDE.md points every session in the fleet at it.

Audited against the code today. The drift was not confined to the tool
table the ticket reported:

  section 3  still described the OLD identity rule - "any connection that
             does not map to a known worker is treated as a primary".
             That was a real privilege bug, fixed since by the ancestry
             walk in #161. The page still taught it.
  section 4  names port 8080 (the mount is 8765) and says the pom does
             not yet carry an MCP dependency.
  section 5  named fleet_read and fleet_cancel, which do not exist, and
             omitted fleet_poll, fleet_ack, fleet_profiles, fleet_whoami
             and fleet_list, which do.
  section 8  says turn_id where the code says turnId, and has no row for
             the exhausted outcome CB-578 added.
  sections
  9, 10, 11  pre-build planning: "new work" columns, open decisions long
             since decided, CB-1xx placeholders.

Every one of those is the same defect: a hand-maintained second copy of
something the code already states. So the copy is deleted rather than
corrected - correcting it just restarts the clock.

What survives is the flows and the status gating, because a flow is a
shape rather than a name, and shapes are what this page was ever good
for. They are rewritten with the names checked against the code, and
extended with what has been learned since: the ~60s cap on a blocking
send, the ~55s ask window, and the three ways the turn-done fallback
loses a report (clipped, echoed brief, slow member).

389 lines -> 188.

The names that remain are guarded. McpContractDocTest fails if the page
names a fleet_* tool FleetMcp does not register, and - because an empty
set is a subset of everything - a second test pins that both sides
actually found names, so the check cannot pass by checking nothing. A
third pins the "this is not the tool reference" sentence, which is the
fix itself: without it someone helpfully re-adds a tool table.

Mutation-tested both ways, 0 compile errors each: adding `fleet_read` to
the doc fails theDocNamesNoToolThatDoesNotExist ("names [fleet_read] ...
Checked 6 name(s)"); removing the disclaimer fails
theDocStillDisclaimsBeingTheToolReference.

All 5 mermaid diagrams render under mermaid-cli.

CLAUDE.md's pointer said "section 6 only" and now names the guard
instead. It is in the project addendum, so the canonical block is
untouched - verified still byte-identical with the wiki template.

REST is split out to #252: 14 routes, documented nowhere, and it IS a
supported operator surface - one of them drains on read.

1232 tests, 0 failures.
2026-09-03 12:52:12 +07:00
Dai Ha 80092ff359 fleetd #247: stop writing a trust key Claude Code strips on every save
CI / contract (push) Successful in 1m26s
CI / build (push) Successful in 1m28s
seedTrustDialog wrote two keys into the shared .claude.json:
hasTrustDialogAccepted and hasCompletedProjectOnboarding. Only the first
one survives.

Measured live on 2026-09-03, minutes after a spawn seeded the file:

  hasTrustDialogAccepted:       28 of 28 project entries
  hasCompletedProjectOnboarding: 0 of 28 project entries

Our entry was written by the running jar and the key was already gone,
so it was written and then removed. It is absent from the 27 entries
Claude Code wrote for itself too, which says Claude Code normalises the
whole file when it saves and drops that key every time.

That reframes #247. I filed it as a race - a save landing between our
read and our ATOMIC_MOVE. It is not a race. The other writer removes
this key as its steady-state behaviour, with no window involved. So the
compare-and-swap retry proposed there would not have helped: it would
re-add a key that gets stripped again on the next save.

The seed's whole job is to stop the workspace-trust dialog blocking a
member (#149). The live probe reached idle with hasTrustDialogAccepted
alone, so the second key was never doing that job. Writing it only added
a contested key to a file two processes share, and made the next reader
think it mattered.

The atomic write and the lock stay. Both are still correct, both are
cheap, and hasTrustDialogAccepted is genuinely shared state.

The new assertion is assertFalse, not a deletion. Removing the old
assertion would leave nothing to stop someone re-adding the key later as
a plausible-looking completeness fix. Mutation-tested: restoring the
production line fails
seedTrustDialogWritesOnlyTheTrustFlagAndNotTheOnboardingKey:2198 with 0
compile errors.

1229 tests, 0 failures.
2026-09-03 12:41:20 +07:00
24 changed files with 2241 additions and 558 deletions
+8 -5
View File
@@ -215,11 +215,14 @@ must obey belongs in the charter, not here.
reads the list with `git config --worktree --get-all fleet.neutralizedConfig`, and the
consequence with `git config --worktree --get fleet.neutralizedConfigNote`. Never brief a worker
to edit one of these files: the edit cannot be committed, and it will not tell you so.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` **§6 only**. The rest of that page is a pre-build design
doc whose tool names, parameter names and REST paths never caught up with the code, so do not use
it as the tool reference (CB-609). Section 6 is kept out of this file because this file loads into
every session's context.
- **Flows and the error model** — rendezvous, `fleet_ask`, detached delivery, the turn-done
fallback and status gating — are diagrammed in `docs/MCP-Contract.md`. That page is now flows
only: its pre-build tool catalogue, parameter tables and REST paths were deleted rather than
corrected, because a hand-maintained second copy of the tool surface is what drifted for a month
while this line pointed every session at it (CB-609 / #114). **The live MCP schema is the tool
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
### Redeploying the daemon — the lead may do this (primary only)
+20 -4
View File
@@ -31,11 +31,27 @@ Environment=HERDR_SOCKET_PATH=%h/.config/herdr/herdr.sock
# spawns, so this line decides whether the fleet can run a build at all. systemd does not source a
# login shell, so without it the daemon — and every worker — gets a bare default with no JDK/Maven.
Environment=PATH=/usr/lib/jvm/temurin-25-jdk/bin:/usr/share/maven/bin:/usr/local/bin:/usr/bin:/bin
# Secrets are NOT set here — this file is committed. Put the API/worker tokens in a private
# drop-in that systemd reads with restrictive permissions:
# systemctl --user edit fleetd → [Service] / Environment=FLEETD_API_TOKEN=...
# or point EnvironmentFile at a 0600 file:
# Secrets are NOT set here — this file is committed. Put ALL three tokens in a private drop-in
# that systemd reads with restrictive permissions. In `systemctl --user edit fleetd`, add:
# [Service]
# Environment=FLEETD_API_TOKEN=...
# Environment=WORKER_GITEA_TOKEN=...
# Environment=AI_GATEWAY_TOKEN=...
# FLEETD_API_TOKEN protects fleetd's API. WORKER_GITEA_TOKEN lets members open pull requests; if
# it is missing, fleetd still starts, but a member fails when it later tries to open a pull request.
# AI_GATEWAY_TOKEN authenticates gateway profiles; if it is missing, fleetd still starts, but a
# gateway profile later returns HTTP 401. Or, put the same three variables in a 0600 file and add:
# EnvironmentFile=%h/.config/fleetd/env
# After starting, check which of them actually resolved. The daemon reports every secret a
# configured profile references, by name, never by value:
# journalctl --user -u fleetd | grep 'startup secret'
# A resolved one logs "startup secret NAME: set (profile 'x' tokenEnv)". A missing one logs
# "startup secret NAME: MISSING" at WARN — and the daemon starts anyway, which is the whole
# problem: without this grep the first sign is a member that cannot open a pull request, hours
# later and in a different component.
# Note what the report can and cannot tell you. It lists only names some profile actually
# references (tokenEnv, gitTokenEnv, and the broker uriEnv). A secret nothing references is never
# reported, because nothing needs it.
Restart=on-failure
RestartSec=10s
+123 -324
View File
@@ -1,311 +1,162 @@
# MCP Contract — `fleetd`'s unified gateway
# MCP flows and error model — `fleetd`
> **Status: 🔴 HISTORICAL DESIGN — do NOT use as the tool reference.** Written 2026-07-14, before
> any MCP code existed. The system shipped and this page never caught up, so **its tool names,
> parameter names and REST paths are wrong today**. Audited 2026-08-17; the specific drift:
> **What this page is.** The **flows**: how a delegation, a clarification, a detached task and a
> silent member each travel through `fleetd`. These shapes are what shipped, and they are hard to
> read off the code because they span the MCP face, the rendezvous registry, the `Injector` and
> herdr.
>
> - **Tools it names that do not exist:** `fleet_read`, `fleet_cancel`.
> - **Shipped tools it omits:** `fleet_poll`, `fleet_ack`, `fleet_profiles`, `fleet_whoami`.
> - **Parameter names are wrong nearly everywhere** — it says `message`/`target`/`timeout_seconds`/
> `block` where the code takes `content`/`sessionId`/`timeoutMs`/`wait`; `text` where
> `fleet_reply` takes `content`; `target` where `fleet_stop` takes `paneId`.
> - **REST paths are wrong:** it says `POST /workers` and `DELETE /workers/{paneId}`; the daemon
> serves `POST /members` and `DELETE /members/{paneId}`.
> **What this page is NOT: a tool reference.** It deliberately holds no tool catalogue, no
> parameter tables and no REST paths. **The live MCP schema is the authority** — each tool's own
> description and parameters, as mounted — with the intent→tool table in `CLAUDE.md` as the short
> form.
>
> **The authoritative tool surface is the live MCP schema** (each tool's own description and
> parameters, as mounted), with the intent→tool table in `CLAUDE.md` as the short form. Both were
> checked against `mcp/FleetMcp.java` on 2026-08-17 and are accurate.
> That absence is the fix for fleetd #114 (CB-609), and it is worth stating why. This page used to
> carry a full tool catalogue written in July 2026, before any MCP code existed. The code shipped;
> the page did not follow. By August it named two tools that do not exist, omitted five that do,
> had the wrong name for nearly every parameter, pointed at REST paths the daemon does not serve,
> and — worst — still described an identity model (*"any connection that does not map to a known
> worker is treated as a primary"*) that was a real privilege bug, fixed since by the ancestry
> walk in fleetd #161. Every one of those errors is the same error: **a second, hand-maintained
> copy of something the code already states**. So the second copy is gone rather than corrected.
> Only the flows remain, because a flow is a shape rather than a name, and shapes are what this
> page was ever good for.
>
> What is still worth reading here is **§6 — the flows and the error model** (rendezvous,
> `fleet_ask`, detached delivery, the turn-done fallback). The shapes it describes are the ones
> that shipped; only the names around them drifted. Rewriting this page is tracked as **CB-609**.
`fleetd` is the **sole communication gateway** for every Claude session in the bridge. Both
the **primary** (Opus, on subscription) and every **worker** (off-subscription Claude Code)
mount the *same* MCP server with a single `claude mcp add` line, and talk only through its
tools. No Claude session ever addresses a broker, a peer, or the network directly.
This document defines every MCP tool that face must expose, who may call it, its blocking
semantics, and how it maps onto the code already in the tree.
> The names that do appear below are checked by `McpContractDocTest`, which fails if this page
> names a `fleet_*` tool the server does not register. That test is the whole reason it is safe to
> write a tool name here at all.
---
## 1. Design constraints (non-negotiable)
## 1. Rendezvous flows
These come from the project's core invariants and bound every decision below.
### 1.1 Delegation — happy path
1. **One server, both roles.** The primary and all workers mount an identical server. The
catalog must serve both, and `fleetd` must decide *who is calling* from the connection —
never from a caller-supplied argument that could be spoofed.
2. **Subscription-safe by construction.** No MCP tool ever reads, sets, or forwards
`ANTHROPIC_BASE_URL`. Mounting the bridge cannot move a session off subscription.
Enforced today by [`SubscriptionGuard`](1-Architecture).
3. **Blocking rendezvous, no busy-poll.** The primary consumes a worker's reply through a
*single* MCP call that `fleetd` holds open — never a cross-turn poll loop that would burn
subscription quota.
4. **Status-gated delivery.** Anything that puts text into a worker flows through the existing
[`Injector`](1-Architecture): delivered only when the worker is `idle`/`blocked`, at most
one message per turn.
5. **`fleetd` owns policy; herdr owns PTYs.** MCP tools express *intent*; `fleetd`
translates it into guard checks, rendezvous bookkeeping, and herdr `agent.*` calls.
---
## 2. Topology
Both faces live in the one daemon. The **north face** is MCP (this document); the **south
face** is the herdr Unix socket. REST/SSE remains only for non-Claude clients and dashboards.
```mermaid
flowchart LR
OPUS["Opus — primary<br/>(Claude Code, env CLEAN)<br/>MCP client"]
subgraph BD["fleetd — standalone daemon"]
MCP["MCP server (north face)<br/>fleet_send · fleet_reply<br/>fleet_ask · fleet_status · lifecycle"]
RDV["rendezvous registry<br/>(blocking-call waiters)"]
INJ["Injector + StatusPoller<br/>(status-gated writer)"]
SOCK["herdr socket client (south face)"]
MCP --> RDV
RDV --> INJ
INJ --> SOCK
MCP --> SOCK
end
HERDR["herdr<br/>panes · agent-status"]
W["worker claude pane<br/>ANTHROPIC_BASE_URL set<br/>MCP client"]
OPUS -->|"fleet_send (blocks)"| MCP
W -.->|"fleet_reply / fleet_ask"| MCP
SOCK -->|"agent.start · agent.send<br/>agent.get · pane.close"| HERDR
HERDR -->|"drives PTY"| W
classDef ext fill:#2b6cb0,stroke:#1a365d,color:#ffffff;
classDef core fill:#2f855a,stroke:#22543d,color:#ffffff;
class OPUS,W ext
class MCP,RDV,INJ,SOCK core
```
---
## 3. Identity & addressing
Because the same server is mounted by everyone, `fleetd` resolves the caller's role on every
request — this is the linchpin of the whole contract and has no code yet.
- **Workers are known.** `fleetd` spawns every worker
([`WorkerService`](1-Architecture)) and records its herdr session UUID / `terminal_id` on
the returned [`Agent`]. When a call arrives on a connection that maps to a known worker,
the caller is *that* worker — so **workers never pass a target**; routing is implicit.
- **The primary is "not a worker".** Any connection that does not map to a known worker is
treated as a primary. It addresses workers **explicitly** by `target` — a session UUID,
a `terminal_id`, or a friendly `profile` name.
- **Turn correlation.** A blocking `fleet_send` registers a *waiter* keyed by worker
identity. A worker's later `fleet_reply` / `fleet_ask` on the same identity resolves that
waiter. A `turn_id` is minted per exchange so a clarification round-trip
(§6.2) rejoins the right turn.
---
## 4. Transport
`fleetd` is a long-lived daemon serving **multiple** concurrent clients (one primary + N
workers), so a per-client stdio child is the wrong shape. The recommended transport is
**streamable-HTTP / SSE** on the same bind as the REST face:
```bash
# identical on primary and every worker
claude mcp add --transport http fleetd http://127.0.0.1:8080/mcp
```
This adds an MCP-server dependency the pom does not yet carry. See [Open decisions](#10-open-decisions).
---
## 5. Tool catalog
| Tool | Caller | Blocks? | Backing (exists today?) |
|---|---|---|---|
| [`fleet_send`](#fleet_send) | primary | yes (default) | `Injector.enqueue` ✅ · rendezvous registry ❌ (CB-104) |
| [`fleet_reply`](#fleet_reply) | worker | no | rendezvous ❌ · pane injection via `Injector` ✅ |
| [`fleet_ask`](#fleet_ask) | worker | yes | reverse rendezvous ❌ |
| [`fleet_status`](#fleet_status) | either | no | `AgentControl.status` ✅ · `Injector.activeTargets` ✅ |
| [`fleet_spawn`](#lifecycle) | primary | no | `WorkerService.spawn` ✅ (`POST /workers`) |
| [`fleet_list`](#lifecycle) | either | no | `WorkerService.list` ✅ (`/agents`) |
| [`fleet_stop`](#lifecycle) | primary | no | `WorkerService.stop` ✅ (`DELETE /workers/{paneId}`) |
| [`fleet_read`](#fleet_read) | primary | no | `AgentControl.read` ✅ |
| [`fleet_cancel`](#fleet_cancel) | primary | no | — ❌ (future) |
### Core: delegation & rendezvous
#### `fleet_send`
*(primary → worker — the headline tool, CB-104)*
- **Params:** `message` (required); `target` (optional — defaults to the sole worker / default
profile); `timeout_seconds` (default 600); `block` (default `true`); `auto_spawn`
(default `true`); `turn_id` (optional — supplied when answering a worker's `fleet_ask`).
- **Blocking (`block:true`):** enqueue `message` via the `Injector`, then hold the call open
until exactly one of:
- worker calls `fleet_reply` → `{ outcome:"reply", text }`
- worker calls `fleet_ask` → `{ outcome:"question", text, turn_id }`
- worker's `agent_status` reaches done/idle with no reply → `{ outcome:"turn_done", text:<terminal tail> }`
- deadline elapses → `{ outcome:"timeout" }`
- worker gone → error `worker_gone`
- **Detached (`block:false`):** enqueue and return `{ outcome:"dispatched", dispatch_id }`
immediately. The eventual reply is injected into the primary's idle pane (§6.3), or drained
via `fleet_status` on a split-host primary.
#### `fleet_reply`
*(worker → primary)*
- **Params:** `text` (required); `final` (default `true`).
- **Behavior:** resolve the primary waiter registered against this worker with `text`. If no
waiter exists (detached delegation), `fleetd` **injects the primary's idle pane** instead.
Returns `{ delivered:true, mode:"resolved"|"injected" }`. No `target` — identity is implicit.
#### `fleet_ask`
*(worker → primary — the reverse rendezvous)*
- **Params:** `question` (required); `timeout_seconds`.
- **Behavior:** blocks the *worker's* call. Surfaces the question to the primary (resolving its
open `fleet_send` with `outcome:"question"`, or injecting its pane). When the primary
answers — a `fleet_send` carrying the matching `turn_id` — that unblocks this call and
returns `{ answer }` to the worker, which continues **in the same turn**.
### Worker lifecycle
<a id="lifecycle"></a>
Thin adapters over [`WorkerService`](1-Architecture) — parity with the existing REST routes.
- **`fleet_spawn`** — `{ profile? }` → worker view (`sessionId`, `terminalId`, `paneId`,
`status`). Guard-checked; a boundary breach returns error `subscription_boundary` (the
REST `403`).
- **`fleet_list`** — no params → all workers + `agent_status`. Read-only, either role.
- **`fleet_stop`** — `{ target }` → tears down the pane and its dedicated tab. Idempotent.
### Observability
#### `fleet_status`
*(either role — the README's 4th named tool)*
- **Params:** `target?`.
- **Behavior:** per-worker `agent_status`, queue depth (`Injector.activeTargets`), whether a
rendezvous is open, and ids. For the *calling* session it also reports/drains **pending
messages addressed to me** — the path a split-host primary's `Stop`-hook uses to wake and
collect replies without being injectable. Read-only, non-blocking.
#### `fleet_read`
*(primary)*
- **Params:** `target`; `source` ∈ `visible | recent | recent_unwrapped | detection`.
- **Behavior:** returns the worker's terminal text so the primary can peek at a *detached*
worker's progress. Adapter over `AgentControl.read`.
### Control (future)
#### `fleet_cancel`
*(primary)*
- **Params:** `target`. Interrupt the worker's current turn / abandon the rendezvous. No
backing code yet.
---
## 6. Rendezvous flows
### 6.1 Delegation — happy path
One blocking call, zero polls.
One blocking call, zero polls. The lead's call is held open by `fleetd` until the member answers.
```mermaid
sequenceDiagram
participant P as Primary (Opus)
participant B as fleetd (MCP + Injector)
participant P as "Lead (primary)"
participant B as "fleetd (MCP + Injector)"
participant H as herdr
participant W as Worker (Claude)
participant W as "Member"
P->>B: fleet_send("do X", target=w) — blocks
B->>B: register waiter(w)
B->>H: agent.send(w, "do X") (idle window)
H-->>W: prompt injected
W->>W: works the turn
W->>B: fleet_reply("result")
B->>B: resolve waiter(w)
B-->>P: { outcome:"reply", text:"result" }
P->>B: "fleet_send{sessionId, content} — blocks"
B->>B: "register waiter(sessionId)"
B->>H: "agent.send — only in an injectable window"
H-->>W: "prompt injected"
W->>W: "works the turn"
W->>B: "fleet_reply{content}"
B->>B: "resolve waiter"
B-->>P: "{ outcome: reply }"
```
### 6.2 Clarification — reverse rendezvous (`fleet_ask`)
**The cap that matters:** a blocking `fleet_send` is bounded by the *caller's own* MCP client
timeout, about 60 seconds — not by the task. Anything slower than that must use the detached flow
in §1.3, or the lead's call returns while the member is still working.
The worker pauses mid-turn to ask; the primary answers; the worker resumes in the same turn.
### 1.2 Clarification — reverse rendezvous
The member pauses mid-turn to ask, the lead answers, and the member resumes **the same turn** with
its context intact.
```mermaid
sequenceDiagram
participant P as Primary
participant P as "Lead"
participant B as fleetd
participant W as Worker
participant W as "Member"
P->>B: fleet_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>B: fleet_ask("which config?") — worker blocks
B-->>P: { outcome:"question", text:"which config?", turn_id }
P->>B: fleet_send("config.yaml", target=w, turn_id) — blocks again
B-->>W: resolve fleet_ask → { answer:"config.yaml" }
W->>W: resumes same turn
W->>B: fleet_reply("done")
B-->>P: { outcome:"reply", text:"done" }
P->>B: "fleet_send{sessionId, content} — blocks"
B-->>W: "content injected"
W->>B: "fleet_ask{question} — member blocks"
B-->>P: "{ outcome: question, turnId }"
P->>B: "fleet_send{turnId, content} — answers THIS turn"
B-->>W: "fleet_ask returns the answer"
W->>W: "resumes the same turn"
W->>B: "fleet_reply{content}"
B-->>P: "{ outcome: reply }"
```
### 6.3 Detached delegation — pane injection
**Answer with `turnId`, never `sessionId`.** A `sessionId` send starts a new turn; it does not
resolve the waiting `fleet_ask`.
The primary does not block; the reply arrives later in its idle pane.
**The window is about 55 seconds and no nudge extends it.** So never brief a member to "ask me":
decide before delegating, or give the member an explicit default to fall back on.
### 1.3 Detached delegation — the lead does not block
The lead gets a ticket immediately and collects the answer later. This is the flow for any real
task, because of the ~60s cap in §1.1.
```mermaid
sequenceDiagram
participant P as Primary
participant P as "Lead"
participant B as fleetd
participant W as Worker
participant W as "Member"
P->>B: fleet_send("do X", target=w, block=false)
B-->>P: { outcome:"dispatched", dispatch_id }
P->>P: continues its own work
W->>B: fleet_reply("result")
Note over B: no waiter → detached path
B->>B: Injector.enqueue(primary_pane, "result")
B-->>P: injected into idle pane (status-gated)
P->>B: "fleet_send{sessionId, content, wait:false}"
B-->>P: "accepted — ticket"
P->>P: "continues its own work"
W->>B: "fleet_reply{content}"
Note over B: "no waiter is blocked — the reply is held"
B->>B: "nudge the lead's own pane (status-gated)"
P->>B: "fleet_poll{ticket}"
B-->>P: "the member's report"
P->>B: "fleet_ack{target, msgId}"
```
### 6.4 Uncooperative worker — turn-done fallback
A terminal ticket nudges the lead's pane by itself, so a detached task does not need watching. The
nudge needs an injectable lead pane and is capped, so it is a convenience rather than a guarantee.
A worker that never calls `fleet_reply` still returns a result: `fleetd` reads its terminal
tail when the turn completes.
### 1.4 The member never replies — turn-done fallback
A member that ends its turn without `fleet_reply` still produces something: `fleetd` reads its
pane tail. This is a **fallback, not a channel** — it is lossy in three separate ways, and every
one of them has produced a wrong answer in practice.
```mermaid
sequenceDiagram
participant P as Primary
participant P as "Lead"
participant B as fleetd
participant W as Worker
participant W as "Member"
P->>B: fleet_send("do X", target=w) — blocks
B-->>W: "do X" (injected)
W->>W: works, never calls fleet_reply
B->>B: StatusPoller sees agent_status → idle/done
B->>B: AgentControl.read(w, "recent")
B-->>P: { outcome:"turn_done", text:<terminal tail> }
P->>B: "fleet_send — blocks or detaches"
B-->>W: "content injected"
W->>W: "works, never calls fleet_reply"
B->>B: "StatusPoller sees the turn end"
B->>B: "read the pane tail"
B->>B: "classify: exhausted? echoed brief? real report?"
B-->>P: "{ outcome: turn_done } or a named failure"
```
The three ways it goes wrong, and what each looks like now:
| What happened | What the lead used to get | What it gets today |
|---|---|---|
| The report is longer than the scrape window | The **end** silently cut off | Still clipped, but marked partial |
| The member never started — spent credential | The lead's **own brief** echoed back as a report | A named failure: backend exhausted |
| The member is simply slow | A tail of work in progress | Unchanged — read it as a hint, not a result |
The echoed-brief case is the one to remember: it reads as a long, on-topic report with nothing in
it from the member. It is suppressed now, but the general rule stands — **check the member's
worktree with `git log` before believing a report you did not watch arrive.**
---
## 7. Status gating
## 2. Status gating
Delivery only happens in a safe window. This is the state machine the `Injector` already
enforces via `AgentStatus.injectable()`; MCP `fleet_send` is simply its producer.
Delivery only happens in a safe window. `fleet_send` is a producer for the `Injector`, which
already enforces this through `AgentStatus.injectable()`.
```mermaid
stateDiagram-v2
[*] --> IDLE
IDLE --> WORKING: message delivered / picks up
WORKING --> IDLE: turn done
WORKING --> BLOCKED: awaits input
BLOCKED --> WORKING: input delivered
IDLE --> UNKNOWN: detection glitch
BLOCKED --> UNKNOWN: detection glitch
UNKNOWN --> IDLE: re-detected
IDLE --> WORKING: "message delivered, picked up"
WORKING --> IDLE: "turn done"
WORKING --> BLOCKED: "awaits input"
BLOCKED --> WORKING: "input delivered"
IDLE --> UNKNOWN: "detection glitch"
BLOCKED --> UNKNOWN: "detection glitch"
UNKNOWN --> IDLE: "re-detected"
note right of IDLE
injectable — deliver head of FIFO
@@ -321,69 +172,17 @@ stateDiagram-v2
end note
```
At most one message is delivered per turn: after a send the `Injector` waits for a `WORKING`
pickup before delivering the next, with a `PICKUP_GRACE_POLLS` fallback for turns faster than
the poll interval. A herdr `events.subscribe` stream can later replace the sampling without
touching this state machine.
**At most one message per turn.** After a send, the `Injector` waits for a `WORKING` pickup before
delivering the next, with a grace-poll fallback for turns that finish faster than the poll
interval.
---
Two consequences a lead feels directly:
## 8. Error model
- **A second send to a busy member never lands.** It reports as queued and times out. The member
is fine; the message simply waits, and then restarts the member when it next goes idle.
- **A spawned member is not deliverable until it has mounted the MCP.** Until then a send waits on
that gate for about 60 seconds and then fails without ever reaching the pane.
| Condition | `fleet_send` result | Notes |
|---|---|---|
| Worker replies | `{ outcome:"reply" }` | normal |
| Worker asks | `{ outcome:"question", turn_id }` | answer with `fleet_send(turn_id)` |
| Turn ends, no reply | `{ outcome:"turn_done" }` | terminal tail as text |
| Deadline elapsed | `{ outcome:"timeout" }` | message may still be queued/delivered |
| Worker vanished | error `worker_gone` | `Injector.drop` fails the queued future |
| Guard breach on spawn | error `subscription_boundary` | REST `403` parity |
| Delivery failed at herdr | error, message dropped | poisoned message not left blocking the FIFO |
`fleet_reply` from a worker with no open waiter is **not** an error — it falls through to
detached pane injection (§6.3).
---
## 9. Mapping to existing code
The MCP face is a thin adapter layer; nearly every capability already exists behind the REST
seam. Only the **rendezvous registry** and the **caller-identity resolver** are new.
| MCP tool | Existing collaborator | New work |
|---|---|---|
| `fleet_send` | `Injector.enqueue`, `AgentControl.send` | waiter registry, timeout, outcome mux (CB-104) |
| `fleet_reply` / `fleet_ask` | `Injector` (pane injection) | reverse rendezvous, identity resolver |
| `fleet_status` | `AgentControl.status`, `Injector.activeTargets` | pending-drain projection |
| `fleet_spawn` / `list` / `stop` | `WorkerService.{spawn,list,stop}` | MCP adapter only |
| `fleet_read` | `AgentControl.read` | MCP adapter only |
Because the REST routes in `FleetApp` already exercise the collaborators, MCP tools are
validated by **parity** against those routes, not by re-testing behavior.
---
## 10. Open decisions
1. **`fleet_ask` direction.** This page defines it as *worker-asks-primary* (a genuine reverse
channel, matching the "inject the primary's pane" language). The alternative — a synonym for
a blocking primary→worker send — is weaker and produces different plumbing. **Recommend
worker-asks-primary.**
2. **Detached delivery shape.** A `block:false` param on `fleet_send` (keeps the catalog
small) vs. a separate `fleet_dispatch` tool. **Recommend the param.**
3. **Auto-spawn on send.** `fleet_send` provisions a worker per profile when none exists
(simplest primary UX) vs. requiring an explicit `fleet_spawn` first. **Recommend
auto-spawn, defaulting on.**
4. **Transport & SDK.** Streamable-HTTP/SSE co-located with the REST bind (recommended) vs.
stdio. Requires choosing a Java MCP server SDK and adding it to the pom.
---
## 11. Implementation staging
- **CB-104** — blocking `fleet_send` + rendezvous registry + caller-identity resolver
(the producer that finally drives the inert `StatusPoller`).
- **CB-1xx** — `fleet_reply` / `fleet_ask` reverse rendezvous + detached pane injection.
- **CB-1xx** — lifecycle + observability adapters (`fleet_spawn/list/stop/status/read`).
- **CB-1xx** — transport wiring + `claude mcp add` docs; parity tests vs. REST.
- **Later** — `fleet_cancel`; swap `StatusPoller` for herdr `events.subscribe`.
`UNKNOWN` is deliberately neither injectable nor a pickup. A pane whose status cannot be read is
not a pane that is safe to write to — see fleetd #176 for what happens when a gate treats an
unreadable pane as a ready one.
+61 -7
View File
@@ -306,6 +306,41 @@ profiles:
# GOTCHA 2 — `maxLoad` is the ONLY throttle you have here. There is no metering, no budget
# and no refusal on cost; the cap on live members is the single thing standing between a
# fan-out and your monthly limit. Set it deliberately and keep it small.
#
# GOTCHA 3 (fleetd #176, corrected by fleetd #257) — `maxLoad` counts members, never the lead
# itself. The lead is a live `claude` session on this SAME account (a lead is never moved
# off-subscription, whatever its own profile says), so it already holds one seat before any
# member spawns. If a lead's `fleet.leaders.<name>.profile` names THIS profile — or ANY OTHER
# `subscription: true` profile that shares this one's account (see THE SENTINEL, just below,
# next to `credentialId:`) — `fleet_list` reports that seat count under `leadSeats`; see
# `profile:` under THE FLEET below. `free` itself is NEVER reduced by `leadSeats`: `free` means
# "what the real placement gate (`CompositePeerLauncher#enforceMaxLoad`) will actually grant a
# fresh `fleet_spawn` right now", and that gate only ever compares live members against
# `maxLoad` — it has no notion of the lead's own seat. An earlier cut of this feature
# subtracted `leadSeats` from `free` on the theory it made `free` describe the true ceiling on
# the account, but no backend seat ceiling shared with the lead has ever actually been
# measured, and the subtraction just made `free` disagree with the one thing it is supposed to
# describe — the fleetd #257 fix. `maxLoad: 3` means 3 member slots, full stop; a lead sharing
# the account is a fact you can see in `leadSeats`, not a reason `free` undercounts spawns that
# will, in practice, succeed.
#
# THE SENTINEL (fleetd #176 stage 2, correcting an inert stage 1 fix): every `subscription:
# true` profile that leaves `credentialId` unset shares ONE implicit account-wide credential
# id with every other such profile on this host — because a subscription profile doesn't
# authenticate with a credential of its own, it authenticates as the operator's own Claude
# login, and there is exactly one of those. So on a typical host, `opus` (the lead's profile)
# and `sonnet` (the members' profile) are linked automatically, with NOTHING to set here — that
# is what makes GOTCHA 3 above work without also writing matching `credentialId:` values on
# both. This linkage is not just cosmetic: it is the same key `BackendQuarantine`/cool-off use,
# so a usage-limit hit on `opus` now quarantines `sonnet` too (and vice versa) — correct, since
# they are one Claude account, but worth knowing before you wonder why an unrelated-looking
# profile went quarantined.
#
# WHEN TO OVERRIDE — set explicit, DIFFERENT `credentialId:` values on two `subscription: true`
# profiles only when they are genuinely two separate Claude logins on the same host (a real,
# if unusual, setup). An explicit `credentialId` always wins over the sentinel, so this is the
# one way to keep two subscription profiles from being treated as one account for lead-seat
# counting AND for quarantine/cool-off grouping alike.
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
@@ -503,6 +538,23 @@ fleet:
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `profile:` has a SECOND job as of fleetd #176, even for a recognise-only lead you never want
# auto-launched: it is also how fleetd learns which account this lead's own session shares. A
# `subscription: true` profile bills the operator's Claude account, and the lead itself is always
# a live `claude` session on that same account — `maxLoad` never counted that seat. If a lead
# entry here names a profile that shares a worker profile's account, `fleet_list` reports the
# lead's live seat(s) on that worker profile under `leadSeats` — informational only, as of fleetd
# #257 it is NEVER subtracted from `free` (see GOTCHA 3, next to `maxLoad:`, in THE WORKERS above,
# for why). "Shares the account" is decided by matching `effectiveCredentialId()`, which (fleetd
# #176 stage 2 — see THE SENTINEL, next to `credentialId:`, in THE WORKERS above) means: an
# explicit, matching `credentialId:` on both, OR — the common case, needing NO extra config — both
# being `subscription: true` with `credentialId` left unset, since those all share one implicit
# account-wide id. A lead on `opus` and workers on `sonnet` link automatically this way; they do
# NOT need the same profile name. Setting `profile:` on an already-running, recognise-only lead is
# safe — the daemon only launches the SHORTFALL below `instances`, so naming a profile here does
# not, by itself, start anything. Omit it and fleetd has no way to derive the sharing — there is
# no other reliable signal on the daemon's side — so that lead's seat never appears in `leadSeats`.
#
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
@@ -624,13 +676,15 @@ guard:
# every name here NOT also in `allow` is overlaid with a non-secret sentinel value before
# the pane's login shell runs — real protection only for names that shell does not itself
# re-export (see the ROUND-2 CORRECTION note above). Under allow-list: reporting only.
# sshAuthSock → whether SSH_AUTH_SOCK may pass through under allow-list ("allow") or must be
# blanked like any other non-derived name ("block", the default). This is a decision you
# have to make explicitly: SSH_AUTH_SOCK is a handle to YOUR ssh-agent, and a member
# holding it can sign with your keys — it sits in no secret file and looks like no
# credential, which is why it slipped past three earlier tickets (gitea #110). Blocking
# it breaks git over SSH inside members (push/fetch authenticate as you); use HTTPS
# remotes or scoped deploy keys instead of allowing it lightly.
# sshAuthSock → whether SSH_AUTH_SOCK may pass through under allow-list ("allow") or is omitted
# from the member environment ("block", the default). Blocking it only omits the
# inherited ssh-agent path. It discourages automatic use of the operator's agent.
# It does not deny same-user access to that socket. It also does not block SSH keys that
# are readable on disk. Git over SSH may still work from inside a member. Keep the block:
# it is correct and costs nothing, but it is not a control. A member runs as the same OS
# user as the lead. Inside one uid, ordinary Unix permissions provide no meaningful
# confidentiality boundary. A real boundary needs a different OS user or OS-level
# confinement, such as a container or VM. That is the open question in fleetd #184.
#
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
@@ -51,6 +51,7 @@ import dev.ltms.fleet.session.SessionReaper;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.member.OpenCodeLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
@@ -640,7 +641,8 @@ public final class Fleetd {
new FleetMcp.OutageSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, outagePolicy));
}, outagePolicy),
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)));
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
@@ -708,8 +710,11 @@ public final class Fleetd {
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
// fleetd #111: live (re-read-per-request) memberCredentials view for GET /member-credentials —
// same hot-reload shape as the memberCredentials supplier passed to ClaudeCodeLauncher above.
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics, deliverable).build();
callers, metrics, deliverable,
() -> MemberCredentialPolicyView.of(config.get().memberCredentials())).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("fleetd listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
@@ -765,6 +770,68 @@ public final class Fleetd {
.orElse(null);
}
/**
* fleetd #176: per-profile factory for {@link FleetMcp.LeadSeatSource} — how many seats a
* profile's own live LEAD session(s) hold on the same Claude subscription.
*
* <p>{@code maxLoad} counts only members; the lead itself is a live {@code claude} session that
* is never moved off-subscription ({@code LeadLauncher} strips {@code ANTHROPIC_BASE_URL}/
* {@code AUTH_TOKEN} from a lead's env whatever its profile says), so a {@code subscription:
* true} profile's real ceiling is lower than its configured {@code maxLoad} by exactly the
* number of lead seats sharing that same account.
*
* <p><b>The derivation, and why this route was chosen over a new config key.</b> The link is
* {@code fleet.leaders.<name>.profile} — the field the operator already sets to name which
* {@code profiles:} entry a lead runs on (CB-557; see {@code fleetd.example.yaml}) — matched
* against the profile passed in here via {@link FleetConfig.Profile#effectiveCredentialId()},
* the same grouping key {@link dev.ltms.fleet.placement.BackendQuarantine} already uses to say
* two profiles share one account. Nothing new is added to the config schema: this reuses a field
* that already exists and already means "the profile this lead's own session runs on". A lead
* entry that names no {@code profile:} (recognise-only, CB-558) says nothing about which account
* it shares, and there is no other reliable signal on the daemon's side to derive that from — so
* such a lead contributes no seats, exactly as before this ticket. Making that lead's seat count
* requires the operator to add one line (`profile: sonnet` under its {@code fleet.leaders} entry)
* — a config statement, not a code change, and the smallest one available given the field
* already exists for a closely related purpose.
*
* <p>Only counts leads {@code liveLeadTerminals} currently reports — CB-531's live tab scan (or
* the legacy {@code primary.terminal} pin) — never every configured lead: an entry whose
* {@code instances} nobody has actually started is not really competing for a seat, and must not
* shrink capacity for one that was never live.
*
* @param profiles the live profile map, normally {@code () -> config.get().profiles()}
* in {@code main} — hot, like every other {@code maxLoad}/
* {@code credentialId} read {@link FleetMcp.CapacitySource} already does
* @param leaders {@code fleet.leaders}, read once at startup like the rest of that
* block ({@code fleetd.example.yaml} notes it is not hot) — passed as a
* plain map, never re-read from {@code config.get()}
* @param liveLeadTerminals terminal_id → lead name for every CURRENTLY recognised lead, normally
* the same supplier {@link dev.ltms.fleet.auth.CallerResolver#leads()}
* and {@code LeadCoordLoop} already consult
*/
static Function<String, Integer> leadSeatLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
Map<String, FleetConfig.Leader> leaders, Supplier<Map<String, String>> liveLeadTerminals) {
return profileName -> {
FleetConfig.Profile target = profiles.get().get(profileName);
if (target == null || !target.isSubscription()) {
return 0;
}
String targetCredential = target.effectiveCredentialId();
int seats = 0;
for (String leadName : liveLeadTerminals.get().values()) {
FleetConfig.Leader lead = leaders.get(leadName);
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
continue;
}
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
if (leadProfile != null && targetCredential.equals(leadProfile.effectiveCredentialId())) {
seats++;
}
}
return seats;
};
}
/**
* fleetd #248 / fleetd#201 Unit 5: package-private factory for the per-target backend-error
* pattern lookup {@link CompletionResolver} classifies a pane scrape against. Closes over the
@@ -1082,12 +1149,14 @@ public final class Fleetd {
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
*/
static void reportMemberCredentialsGap(FleetConfig cfg) {
FleetConfig.MemberCredentials creds = cfg.memberCredentials();
if (creds != null && !creds.known().isEmpty()) {
// fleetd #111: the counts below come from MemberCredentialPolicyView, the same class the
// live GET /member-credentials endpoint reads — one place computes them, not two.
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(cfg.memberCredentials());
if (view.present()) {
log.info("memberCredentials: policy={}, {} known name(s), {} allowed — blocking {} on "
+ "every spawn{}",
creds.policy(), creds.known().size(), creds.allow().size(), creds.blockedSet().size(),
creds.isAllowList()
view.policy(), view.knownCount(), view.allowedCount(), view.blockedCount(),
cfg.memberCredentials().isAllowList()
? " (allow-list: known/allow are reporting only — the control is the derived ZDOTDIR scrub)"
: "");
return;
@@ -92,9 +92,15 @@ import java.util.regex.PatternSyntaxException;
* fleetd's own process, so fleetd's own {@code $SHELL} says nothing about what
* that pane runs. There is no channel to ask herdr for another user's shell, so
* this must be told, never guessed. {@code null}/blank (or a value not ending
* in {@code zsh}) is treated the same as "not zsh": the {@code
* memberCredentials.policy: allow-list} ZDOTDIR scrub is skipped in favour of
* the CB-596 sentinel overlay — a degraded control, never a refusal to spawn.
* in {@code zsh}) is treated the same as "not zsh". fleetd #155: what that means
* now depends on {@code memberCredentials.policy}. Under {@code deny-by-default}
* it stays a degraded control, never a refusal to spawn — the pane-creation
* overlay is unaffected by shell type, so the spawn proceeds with a WARN naming
* the shell. Under {@code allow-list} the spawn is REFUSED instead: that policy's
* whole point is a control a sourced file cannot undo, so silently falling back
* to the weaker overlay would be the same "control silently does nothing" defect
* #155 exists to remove — configure this field (or move the account to zsh, or
* switch policy back to {@code deny-by-default}) to unblock the spawn.
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
@@ -644,14 +650,46 @@ public record FleetConfig(
}
/**
* The credential group this profile quarantines with (CB-578 stage B): the configured
* {@link #credentialId} when set, else this profile's own name — so an unconfigured profile
* quarantines alone, exactly as it did before this field existed. Two profiles that set the
* same non-blank {@code credentialId} share one quarantine: a {@code BACKEND_EXHAUSTED}
* classification on either one quarantines both.
* The shared credential id every {@code subscription: true} profile falls back to when it
* sets no explicit {@link #credentialId} (fleetd #176 stage 2, correcting an inert first cut
* of that ticket). A subscription profile has no credential of its own to fall back to its
* name for: it authenticates as the operator's own Claude login, and a host has exactly one
* of those, whatever names the operator gives the profiles running on it. Falling back to the
* profile's own name (the way an ordinary off-subscription profile does) would keep two
* subscription profiles on one login apart from each other, which is the opposite of what
* "one account" means.
*
* <p>Measured live and what it broke: a lead on profile {@code opus}, members on profile
* {@code sonnet}, same Claude login, neither setting {@code credentialId}. Before this
* sentinel, {@code opus.effectiveCredentialId()} was {@code "opus"} and {@code sonnet
* .effectiveCredentialId()} was {@code "sonnet"} — so fleetd #176's lead-seat matcher (and,
* this sentinel now also fixes, {@code CompositePeerLauncher}'s quarantine/cool-off spawn
* refusal and {@code BackendOutagePolicy}'s incident grouping) silently never linked them: the
* fix shipped, and stayed inert on the one host it was written for.
*/
public static final String SUBSCRIPTION_CREDENTIAL_ID = "<subscription>";
/**
* The credential group this profile quarantines with (CB-578 stage B; extended fleetd #176
* stage 2 — see {@link #SUBSCRIPTION_CREDENTIAL_ID}): the configured {@link #credentialId}
* when set — that always wins, so an operator with two separate Claude logins on one host can
* still keep them apart. Otherwise, a {@code subscription: true} profile falls back to
* {@link #SUBSCRIPTION_CREDENTIAL_ID} rather than its own name; an ordinary off-subscription
* profile falls back to its own name, exactly as it did before this field existed, so an
* unconfigured off-subscription profile still quarantines alone.
*
* <p>A {@code BACKEND_EXHAUSTED} (or repeated backend-error) classification on any profile
* sharing the result quarantines/cools off every profile that shares it — including, now,
* every {@code subscription: true} profile with no explicit {@code credentialId}. That is
* intended, not incidental: one Claude subscription hitting a usage limit really does take out
* every profile running on it, the same way {@code credentialId: openai-shared} already lets
* two OpenAI-backed profiles share one quarantine.
*/
public String effectiveCredentialId() {
return (credentialId == null || credentialId.isBlank()) ? profile : credentialId;
if (credentialId != null && !credentialId.isBlank()) {
return credentialId;
}
return isSubscription() ? SUBSCRIPTION_CREDENTIAL_ID : profile;
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
@@ -946,7 +984,15 @@ public record FleetConfig(
* survives restarts of the agent inside it — so identity is now the tab label alone.
*
* @param profile the {@code profiles:} entry to launch this lead on when one must
* be created; {@code null} ⇒ recognise-only, never create
* be created; {@code null} ⇒ recognise-only, never create.
* <p>fleetd #176: also the field {@code Fleetd.leadSeatLookup} reads
* to learn which account this lead's own live session shares — set it
* (safely, even on an already-running recognise-only lead: naming a
* profile here never starts anything beyond {@code instances}) so a
* {@code subscription: true} worker profile sharing its
* {@code effectiveCredentialId()} has this lead's seat subtracted from
* {@code fleet_list}'s {@code free}. {@code null} here also means this
* lead's seat cannot be derived and is not counted.
* @param tab the exact tab label hosting this lead, matched case-insensitively;
* the only field identity depends on. Required — a lead with no
* {@code tab} can never be discovered, launched or not
@@ -96,6 +96,8 @@ public final class FleetMcp {
private final QuarantineSource quarantine;
/** fleetd #201 Unit 5: SEPARATE from {@link #quarantine} — see {@link OutageSource}'s doc. */
private final OutageSource outage;
/** fleetd #176: SEPARATE from both of the above — see {@link LeadSeatSource}'s doc. */
private final LeadSeatSource leadSeats;
/** CB-637: this daemon's lead-to-lead channel; {@code null} when no coordinator is configured. */
private final LeadChannel leadChannel;
@@ -135,6 +137,32 @@ public final class FleetMcp {
}
}
/**
* fleetd #176: the seats a profile's own live LEAD session(s) hold on the same Claude
* subscription — a fact {@code fleet_list} reports alongside {@code free} via the
* {@code leadSeats} key.
*
* <p>{@code maxLoad} counts only <em>members</em>, never the lead itself. But a
* {@code subscription: true} profile bills the operator's own Claude account, and the lead is
* always a live {@code claude} session on that same account (it is never moved off-subscription
* — see {@code LeadLauncher}). See {@code Fleetd.leadSeatLookup} for how the count is derived —
* from {@code fleet.leaders.<name>.profile} and each profile's {@code effectiveCredentialId()},
* never a hardcoded constant.
*
* <p>fleetd #257: this count is reported, never subtracted from {@code free}. An earlier cut of
* this feature subtracted it, on the theory that it made {@code free} describe the real ceiling
* on the account — but the real spawn gate ({@code CompositePeerLauncher#enforceMaxLoad}) never
* read this count at all, so the subtraction made {@code free} disagree with the one thing it is
* supposed to describe: what a fresh {@code fleet_spawn} will actually get. No backend seat
* ceiling shared with the lead has been measured either — see {@code fleetd.example.yaml}'s
* {@code maxLoad} docs. {@code free} now always equals {@code max(0, maxLoad - live)}, and
* {@code leadSeats} is reported purely as a fact the caller may act on however it likes.
*/
public record LeadSeatSource(Function<String, Integer> seatsFor) {
/** Inert source — no profile is ever reported as sharing a seat with a lead. */
public static LeadSeatSource none() { return new LeadSeatSource(_ -> 0); }
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
@@ -149,7 +177,7 @@ public final class FleetMcp {
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, null, OutageSource.none());
healthCoverage, quarantine, null, OutageSource.none(), LeadSeatSource.none());
}
/**
@@ -163,12 +191,12 @@ public final class FleetMcp {
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, OutageSource.none());
healthCoverage, quarantine, leadChannel, OutageSource.none(), LeadSeatSource.none());
}
/**
* As above, with fleetd #201 Unit 5 cool-off facts for {@code fleet_list}/{@code fleet_profiles}
* (see {@link OutageSource}). This is what {@code Fleetd.main} actually wires up.
* (see {@link OutageSource}).
*
* @param outage required — pass {@link OutageSource#none()} for a caller that does not want the
* feature, never a defaulting overload (the same rule {@code quarantine} follows).
@@ -177,10 +205,28 @@ public final class FleetMcp {
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage) {
this(messages, workers, sessions, identity, presence, primaryRegistry, callers, metrics, capacity,
healthCoverage, quarantine, leadChannel, outage, LeadSeatSource.none());
}
/**
* As above, with fleetd #176 lead-seat facts (see {@link LeadSeatSource}). This is what
* {@code Fleetd.main} actually wires up.
*
* @param leadSeats required — pass {@link LeadSeatSource#none()} for a caller that does not want
* the feature, never a defaulting overload (the same rule {@code quarantine} and
* {@code outage} follow).
*/
public FleetMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, LeadChannel leadChannel, OutageSource outage,
LeadSeatSource leadSeats) {
this.leadChannel = leadChannel;
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outage = Objects.requireNonNull(outage, "outage");
this.leadSeats = Objects.requireNonNull(leadSeats, "leadSeats");
this.healthCoverage = healthCoverage;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
@@ -310,7 +356,7 @@ public final class FleetMcp {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
callers == null ? Map.of() : callers.leads(),
leadSeats, callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange),
leadChannel == null ? null : leadChannel.selfCoordId());
};
@@ -994,7 +1040,8 @@ public final class FleetMcp {
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage, leads, selfTerm, null);
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, null);
}
/**
@@ -1011,7 +1058,7 @@ public final class FleetMcp {
QuarantineSource quarantine, Map<String, String> leads, String selfTerm,
String selfCoordId) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, OutageSource.none(),
leads, selfTerm, selfCoordId);
LeadSeatSource.none(), leads, selfTerm, selfCoordId);
}
/** As above, plus fleetd #201 Unit 5 cool-off facts (see {@link OutageSource}). */
@@ -1019,6 +1066,16 @@ public final class FleetMcp {
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
Map<String, String> leads, String selfTerm, String selfCoordId) {
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
LeadSeatSource.none(), leads, selfTerm, selfCoordId);
}
/** As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}). */
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, OutageSource outage,
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
String selfCoordId) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -1045,7 +1102,7 @@ public final class FleetMcp {
}
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong(), quarantine, outage)).toList());
capacity.clock().getAsLong(), quarantine, outage, leadSeats)).toList());
return text(json(result));
} catch (HerdrException e) {
return error("herdr error listing the fleet: " + e.getMessage());
@@ -1084,20 +1141,45 @@ public final class FleetMcp {
* {@code credentialId}/{@code coolingOffForSeconds}, but never {@code quarantinedForSeconds} —
* that key is added only when exhaustion quarantine is ALSO active for this profile, since the
* two checks are independent and either, both, or neither can be true.
*
* <p>fleetd #176: {@code maxLoad} counts panes, not subscription seats — it never counted the
* lead's own seat on a {@code subscription: true} profile's account. {@link LeadSeatSource}
* reports that count (0 for a non-subscription profile, or when no live lead shares its
* credential) via the {@code leadSeats} key, added only when the count is positive, for the
* same byte-identical-when-unused reason as the quarantine/cool-off keys above.
*
* <p>fleetd #257: {@code leadSeatCount} is reported, never subtracted from {@code free}.
* {@code free} means "what a fresh {@code fleet_spawn} on this profile will actually get", and
* the real gate ({@code CompositePeerLauncher#enforceMaxLoad}) only ever compares {@code live}
* against {@code maxLoad} — it has no notion of a lead's own seat. Subtracting
* {@code leadSeatCount} here made {@code free} disagree with the gate it is supposed to
* describe: it could report {@code free: 0} while a spawn on that exact profile still
* succeeded. {@code leadSeats} stays in the row as a fact the caller can act on however it
* likes, but it no longer changes what {@code free} means.
*/
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
Function<String, Integer> maxLoad, List<MemberSession> roster,
MessageService messages, long nowNanos, QuarantineSource quarantine,
OutageSource outage) {
OutageSource outage, LeadSeatSource leadSeats) {
Integer cap = maxLoad.apply(profile);
int live = liveCount.apply(profile);
int leadSeatCount = leadSeats.seatsFor().apply(profile);
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
.filter(s -> (s.state() == MemberSession.State.READY || s.state() == MemberSession.State.DONE))
.filter(s -> messages == null || (!messages.hasAcceptedDelivery(s.terminalId()) && !messages.hasInboxMessage(s.terminalId())))
.count();
Map<String, Object> row = new LinkedHashMap<>();
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
// fleetd #257: free must report what the real spawn gate (CompositePeerLauncher#enforceMaxLoad)
// will actually grant, and that gate never reads leadSeatCount — only maxLoad and live. Not
// subtracting the lead's seat here used to make free UNDERSTATE what a fresh fleet_spawn would
// get, so a lead believing free:0 gave up on a profile the gate would still spawn onto.
// leadSeatCount is still reported via the leadSeats key below, just never subtracted from free.
row.put("free", cap == null ? null : Math.max(0, cap - live));
row.put("reclaimable", reclaimable);
if (leadSeatCount > 0) {
row.put("leadSeats", leadSeatCount);
}
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
@@ -1298,8 +1380,13 @@ public final class FleetMcp {
+ "member spawned without one shares its directory with others and never reports "
+ "an id, however long it runs (fleetd #249). An empty 'members' "
+ "means no members are spawned; it says nothing about peers. When capacity "
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
+ "a quarantined profile's credential (see fleet_profiles), whatever its "
+ "facts are configured, a 'capacity' row per profile reports 'free' — the "
+ "slots a fresh fleet_spawn on that profile will actually be granted right "
+ "now (max(0, maxLoad - live)), the same check the spawn gate itself runs. A "
+ "'leadSeats' key, when present, reports how many of those live slots are a "
+ "lead session sharing this profile's subscription — informational only, "
+ "already NOT subtracted from 'free' (fleetd #257). It also reports free: 0 "
+ "for a quarantined profile's credential (see fleet_profiles), whatever its "
+ "maxLoad/live — with credentialId and quarantinedForSeconds naming the "
+ "quarantine, so 'free: 0, busy' can be told apart from 'free: 0, refusing "
+ "for N seconds'.",
@@ -19,6 +19,7 @@ import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.nio.file.attribute.PosixFileAttributeView;
import java.util.Arrays;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
@@ -495,7 +496,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* <p><b>Additive, not a rewrite.</b> {@code .claude.json} is large (tens of KB, dozens of
* projects) and Claude Code itself rewrites it while running, so this reads the file as a JSON
* tree (missing or unreadable → treated as an empty object) and changes only
* {@code projects.<cwd>.hasTrustDialogAccepted} / {@code .hasCompletedProjectOnboarding} —
* {@code projects.<cwd>.hasTrustDialogAccepted} (that key alone — see fleetd #247) —
* every other top-level key and every other project entry is written back untouched. Only the
* one project entry for {@code cwd} is replaced/created; an existing entry for a DIFFERENT cwd
* (or the operator's own project history) is never touched.
@@ -519,6 +520,26 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* the first. Both exist because of a real incident: see {@link HerdrPeerLauncher#isProvisionedWorktree}'s javadoc
* and {@link #writeAtomically}'s javadoc.
*
* <p><b>Compare-and-swap against a writer the lock cannot reach (fleetd #247).</b>
* {@code TRUST_JSON_LOCK} only serialises calls this launcher itself makes inside this one JVM.
* It does nothing about a writer outside it — and on a host where the profile's
* {@code configDir} is the operator's own {@code CLAUDE_CONFIG_DIR}, the file this method writes
* IS the operator's own live Claude Code session's config file, being read and written by that
* session while it runs. Measured 2026-09-03: its mtime moved minutes after a spawn while that
* session was active. A plain read-modify-write there is a routine lost update, not a rare one:
* fleetd reads v1, the operator's session reads v1 and writes v2 with their own change, fleetd's
* {@code ATOMIC_MOVE} then lands v3 built from v1 — atomic, but v2's change is gone. So before
* the move this method re-reads {@code target}'s exact bytes and compares them with the bytes it
* built its update from; a mismatch means someone else wrote in between, and it discards its
* work and rebuilds from the fresh bytes, up to {@link #MAX_TRUST_JSON_CAS_ATTEMPTS} times.
* <b>Exhausting the retries writes nothing</b> — see the WARN at the end of the loop for why
* that, not a last write-anyway, is the safe failure: the member shows the trust dialog and
* fails to reach an injectable state, which is visible, logged and recoverable; overwriting the
* operator's live config with a stale copy is neither. This narrows the lost-update window, it
* does not close it — a write landing between the final re-read and the {@code ATOMIC_MOVE}
* itself is still lost, because there is no OS-level compare-and-swap on a plain file, only this
* cooperative narrowing of the gap.
*
* @param configDir the profile's {@code CLAUDE_CONFIG_DIR} ({@code cfg.configDir()}), or
* {@code null}/blank to target the default {@code ~/.claude.json}
* @param cwd the spawn's resolved working directory — the exact key Claude Code will look
@@ -528,45 +549,118 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
if (!isProvisionedWorktree(cwd)) {
return;
}
Path target = (configDir == null || configDir.isBlank())
boolean unsetConfigDir = configDir == null || configDir.isBlank();
Path target = unsetConfigDir
? Path.of(System.getProperty("user.home"), ".claude.json")
: Path.of(configDir, ".claude.json");
if (unsetConfigDir) {
// fleetd #247: configDir unset is the ONLY path that targets ~/.claude.json — the
// operator's own home file, not a per-profile one — and it is the default, so a
// profile that simply forgot to set configDir gets no signal at all short of the
// operator noticing their own file changing. Say so loudly, every time it is about to
// happen, rather than only once ever: each occurrence is a live write to a real
// person's home config and deserves its own log line.
log.warn("seedTrustDialog: profile has no configDir set, so the workspace-trust seed "
+ "for cwd '{}' is about to write the operator's own default '{}' — set "
+ "configDir on this profile to target a per-member config file instead",
cwd, target);
}
synchronized (TRUST_JSON_LOCK) {
try {
if (target.getParent() != null) {
Files.createDirectories(target.getParent());
}
ObjectNode root = null;
if (Files.isRegularFile(target)) {
JsonNode existing = TRUST_JSON.readTree(target.toFile());
if (existing instanceof ObjectNode existingObject) {
root = existingObject;
for (int attempt = 1; attempt <= MAX_TRUST_JSON_CAS_ATTEMPTS; attempt++) {
byte[] before = Files.isRegularFile(target) ? Files.readAllBytes(target) : null;
ObjectNode root = parseTrustJsonOrEmpty(before);
JsonNode projectsNode = root.get("projects");
ObjectNode projects = projectsNode instanceof ObjectNode projectsObject
? projectsObject : TRUST_JSON.createObjectNode();
if (!(projectsNode instanceof ObjectNode)) {
root.set("projects", projects);
}
JsonNode projectNode = projects.get(cwd);
ObjectNode project = projectNode instanceof ObjectNode projectObject
? projectObject : TRUST_JSON.createObjectNode();
if (!(projectNode instanceof ObjectNode)) {
projects.set(cwd, project);
}
// fleetd #247: ONLY hasTrustDialogAccepted. We used to write
// hasCompletedProjectOnboarding beside it; do not put it back. Measured on
// 2026-09-03, minutes after a live spawn seeded this file: 28 of 28 project
// entries carried hasTrustDialogAccepted and 0 of 28 carried the onboarding key
// — including the 27 entries Claude Code wrote for itself. Claude Code
// normalises the whole file when it saves and drops that key every time, so
// writing it achieved nothing except making the next reader think it mattered.
// The member reached idle with the trust flag alone, which is the only outcome
// this seed exists for. If a future Claude Code needs the second flag the
// symptom returns as the trust dialog fleetd #149 describes — re-measure then,
// do not restore it on a guess.
project.put("hasTrustDialogAccepted", true);
String newContent = TRUST_JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root);
trustJsonCasTestHook.run();
// fleetd #247 CAS: re-read immediately before the move and compare with what
// this attempt built its update from. A mismatch means another writer (most
// plausibly the operator's own live Claude Code — see this method's javadoc)
// landed a change in between; discard this attempt's work and rebuild from the
// fresh bytes rather than blindly overwriting it.
byte[] atMove = Files.isRegularFile(target) ? Files.readAllBytes(target) : null;
if (!Arrays.equals(before, atMove)) {
continue;
}
writeAtomically(target, newContent);
return;
}
if (root == null) {
root = TRUST_JSON.createObjectNode();
}
JsonNode projectsNode = root.get("projects");
ObjectNode projects = projectsNode instanceof ObjectNode projectsObject
? projectsObject : TRUST_JSON.createObjectNode();
if (!(projectsNode instanceof ObjectNode)) {
root.set("projects", projects);
}
JsonNode projectNode = projects.get(cwd);
ObjectNode project = projectNode instanceof ObjectNode projectObject
? projectObject : TRUST_JSON.createObjectNode();
if (!(projectNode instanceof ObjectNode)) {
projects.set(cwd, project);
}
project.put("hasTrustDialogAccepted", true);
project.put("hasCompletedProjectOnboarding", true);
writeAtomically(target, TRUST_JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
// fleetd #247: deliberately do NOT write here. A member that starts without the
// seed still starts — it may hit the trust dialog fleetd #149 describes and fail to
// reach an injectable state, but that failure is visible (herdr reports it, the
// spawn-readiness gate times out) and recoverable (retry the spawn). Writing our
// stale copy over whatever the other writer left would be silent and, if that other
// writer is the operator's own live session, could destroy real configuration —
// fail toward the recoverable outcome, not the silent one.
log.warn("seedTrustDialog: gave up seeding workspace-trust for cwd '{}' into '{}' "
+ "after {} attempts — another writer (most plausibly the operator's own "
+ "live Claude Code sharing this file) kept changing it faster than we could "
+ "re-read it, so nothing was written; the member may show the trust dialog "
+ "instead", cwd, target, MAX_TRUST_JSON_CAS_ATTEMPTS);
} catch (Exception e) {
log.debug("cannot seed workspace-trust entry for cwd '{}' into '{}'", cwd, target, e);
}
}
}
/** Bound on {@link #seedTrustDialog}'s fleetd #247 compare-and-swap retry loop. */
private static final int MAX_TRUST_JSON_CAS_ATTEMPTS = 5;
/**
* Test-only seam for fleetd #247: invoked once per CAS attempt inside {@link #seedTrustDialog}'s
* retry loop, after that attempt has read the target's bytes and built its replacement content,
* but immediately before the final re-read/compare that decides whether to write. A no-op in
* production. Package-visible (not {@code private}) so {@code ClaudeCodeLauncherTest} can install
* a hook here that writes to the target file, deterministically simulating a writer racing
* fleetd's own read-modify-write at the exact instant the CAS is meant to catch — the same race
* an external process (most plausibly the operator's own live Claude Code) creates, without
* depending on real thread scheduling to land the interleaving. A test that sets this MUST
* restore it to the no-op default in a {@code finally} block — it is shared, static state.
*/
static Runnable trustJsonCasTestHook = () -> {};
/**
* Parse {@code bytes} as a {@code .claude.json} tree, or hand back a fresh empty object when
* {@code bytes} is {@code null} (no file yet) or does not parse to a JSON object — the same
* missing-or-unreadable-is-empty fallback {@link #seedTrustDialog} always used, factored out so
* the fleetd #247 CAS loop can call it once per attempt.
*/
private static ObjectNode parseTrustJsonOrEmpty(byte[] bytes) throws IOException {
if (bytes == null) {
return TRUST_JSON.createObjectNode();
}
JsonNode existing = TRUST_JSON.readTree(bytes);
return existing instanceof ObjectNode existingObject ? existingObject : TRUST_JSON.createObjectNode();
}
/**
* Write {@code content} to {@code target} atomically: serialise to a sibling temp file in the
* <strong>same directory</strong> as {@code target} (an atomic move is only guaranteed within
@@ -184,7 +184,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private final ConcurrentMap<String, Path> zdotdirByPane = new ConcurrentHashMap<>();
/** Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. */
/**
* Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. fleetd #155:
* only the {@code policy: deny-by-default} path still warns on a non-zsh shell — the {@code
* policy: allow-list} path refuses the spawn instead (see {@link #applyEnvironmentAllowListPolicy}).
*/
private final AtomicBoolean nonZshShellWarned = new AtomicBoolean();
/** Live config provides URI environment names that must never enter member panes. */
private final Supplier<FleetConfig> config;
@@ -1255,6 +1259,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (!creds.isAllowList()) {
overlayBlockedCredentials(workerEnv, creds);
logCredentialGap(creds, null);
// fleetd #155: this overlay itself does not depend on the shell (it lands in the
// pane-creation env map before any shell runs), so the spawn is never refused here —
// only policy=allow-list's ZDOTDIR scrub needs a login shell to run at all. Still worth
// telling the operator: the stronger post-shell control is unavailable on this shell.
warnNonZsh(memberLoginShell());
}
}
@@ -1307,10 +1316,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* dev.ltms.fleet.session.Worktrees#shareWithGroup} already uses, reused rather than
* inventing a second group key. Either one missing means the scrub cannot be guaranteed
* reachable by the member, which is the same "cannot guarantee the scrub runs" case as a
* non-zsh shell, so it gets the identical fallback.</li>
* non-zsh shell — that one still gets the fallback (see fleetd #155's javadoc note below
* on why a missing worktreeRoot/worktreeGroup is a different defect, out of that ticket's
* scope).</li>
* </ul>
* In every branch: never refuse to spawn. A degraded credential control must not become an
* outage for an opt-in feature.
* fleetd #155: the one exception to "never refuse to spawn" is a non-zsh login shell, right
* below. The operator asked for {@code policy: allow-list} specifically — its whole point is a
* control a sourced file cannot undo — so silently degrading to the weaker overlay is exactly
* the "silently does nothing" failure this ticket exists to remove. Every other branch in this
* method (worktreeRoot/worktreeGroup missing) keeps the old "degrade, never refuse" behaviour;
* that gap is real but is fleetd #213's scope, not this one.
*/
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
@@ -1324,20 +1339,21 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// must not even be called on this path, only the explicit memberLoginShell: config can
// answer it. With memberHerdrSocket absent, nothing here changes: fleetd's own $SHELL is
// still the input, exactly as before this fix.
String loginShell = memberHerdrSocket ? configuredMemberLoginShell() : resolveEnv("SHELL");
String loginShell = memberLoginShell();
boolean zsh = isZshShell(loginShell);
if (!zsh) {
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
// sentinel overlay over the enumerated known: names — weaker (a sourced file can undo
// it), but strictly better than nothing. Deliberately no "allowed N of M" line here: the
// scrub this count describes does not run on this path, so printing it would tell an
// operator that a fraction of names were blocked when the real number blocked is zero.
// logCredentialGap's WARN (below) is the only signal for this path.
warnNonZsh(loginShell);
overlayBlockedCredentials(launch.env(), creds);
logCredentialGap(creds, null);
return null;
// fleetd #155: a non-zsh login shell ignores ZDOTDIR entirely — NO scrub would run. The
// operator explicitly asked for policy=allow-list's blocking control, so degrading to
// the weaker CB-596 overlay and spawning anyway would be the same "control silently does
// nothing" defect this ticket exists to close. Refuse instead — the caller (FleetMcp.spawn)
// catches IllegalArgumentException and surfaces the message to the operator.
throw new IllegalArgumentException(
"memberCredentials policy=allow-list requires the member's login shell to be "
+ "zsh, so the ZDOTDIR scrub can run after it — refusing to spawn under "
+ "login shell '" + (loginShell == null ? "<unset>" : loginShell)
+ "'. Configure memberLoginShell: as a zsh path (only read when "
+ "memberHerdrSocket is set), move the member's OS account onto zsh, or "
+ "set memberCredentials.policy: deny-by-default instead.");
}
Path parentDir;
String group = null;
@@ -1392,6 +1408,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return cfg == null ? null : cfg.memberLoginShell();
}
/**
* fleetd #155: the member pane's login shell, using the same {@code memberHerdrSocket} routing
* {@link #applyEnvironmentAllowListPolicy} already used — fleetd's own {@code $SHELL} decides it
* when {@code memberHerdrSocket} is absent (today's only mode); the configured {@code
* memberLoginShell:} decides it otherwise, since fleetd's own {@code $SHELL} names a different
* user's shell once member panes run under a different OS user. Shared by both {@link
* #applyEnvironmentAllowListPolicy} (allow-list: refuses on non-zsh) and {@link
* #applyMemberCredentialPolicy} (deny-by-default: warns on non-zsh) so the two policies agree on
* what "the member's shell" means.
*/
private String memberLoginShell() {
return memberHerdrSocketConfigured() ? configuredMemberLoginShell() : resolveEnv("SHELL");
}
/**
* fleetd #213: {@code worktreeRoot}, as the ZDOTDIR scrub's parent directory under {@code
* memberHerdrSocket}, or {@code null} when unconfigured — the same "cannot guarantee the scrub
@@ -1476,15 +1506,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final String SSH_AUTH_SOCK = MemberEnvAllowList.SSH_AUTH_SOCK;
/**
* CB-633: a non-zsh login shell means the allow-list control CANNOT run — say so once per
* launcher instance, naming the shell, instead of failing silently.
* fleetd #155: called ONLY from the {@code policy: deny-by-default} branch of {@link
* #applyMemberCredentialPolicy} — under {@code policy: allow-list}, a non-zsh shell now REFUSES
* the spawn instead (see {@link #applyEnvironmentAllowListPolicy}), since that policy's whole
* point is a control a sourced file cannot undo. deny-by-default's own overlay does not depend
* on the shell (it lands in the pane-creation env map before any shell runs), so this is a
* heads-up, not a defect report: say once per launcher instance, naming the shell, that the
* stronger allow-list control is unavailable here — never silently.
*/
private void warnNonZsh(String shell) {
if (isZshShell(shell)) {
return;
}
if (nonZshShellWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=allow-list: member login shell '{}' is NOT zsh — "
+ "ZDOTDIR scrubbing cannot run, so members' inherited environment is "
+ "UNPROTECTED beyond the enumerated known: fallback. Move herdr onto a "
+ "zsh account or switch policy back to deny-by-default.",
log.warn("memberCredentials policy=deny-by-default: member login shell '{}' is NOT zsh — "
+ "the pane-creation credential overlay still applies here (it does not "
+ "depend on the shell), but policy=allow-list's stronger post-shell "
+ "ZDOTDIR scrub is unavailable on this shell.",
shell == null ? "<unset>" : shell);
}
}
@@ -0,0 +1,75 @@
package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import java.util.List;
/**
* fleetd #111 (CB-608): a single, testable read of the {@code memberCredentials:} policy — names
* and counts only, never a value. The daemon never holds a credential's <em>value</em> in the
* first place (only the names configured under {@code known:}/{@code allow:}), so there is
* nothing here to redact by construction; the point of this class is that it is the ONE place
* that turns a policy into names-and-counts, so nothing else hand-counts a second time.
*
* <p>Before this class, {@link dev.ltms.fleet.Fleetd#reportMemberCredentialsGap} computed these
* same counts inline for the startup log line, and {@code scripts/probe-member-credentials.sh}
* carried its own hardcoded {@code NAMES} array that the live policy could grow past silently
* (#111) — the exact "hand-maintained second copy drifts" shape #114 fixed for the tool
* catalogue. Both now read this class: the startup log via {@link
* dev.ltms.fleet.Fleetd#reportMemberCredentialsGap}, and a live daemon via the {@code
* GET /member-credentials} REST endpoint ({@link dev.ltms.fleet.rest.FleetApp}), which the probe
* script fetches instead of carrying its own list.
*
* @param present policy configured with at least one {@code known} name. {@code false} for an
* absent or empty {@code memberCredentials:} block — represented honestly as "no
* policy", never as "nothing blocked" (an empty {@link #blocked} could otherwise be
* misread as a clean bill of health).
* @param policy the normalized policy mode ({@link FleetConfig.MemberCredentials#policy()}), or
* {@code null} when {@link #present} is {@code false}.
* @param known every name the policy declares, in configured order. Names only, never a value.
* @param allowed the subset of {@link #known} explicitly let through. Names only.
* @param blocked {@link #known} minus {@link #allowed} — the names an actual spawn shadows. Names
* only.
*/
public record MemberCredentialPolicyView(boolean present, String policy, List<String> known,
List<String> allowed, List<String> blocked) {
private static final MemberCredentialPolicyView ABSENT =
new MemberCredentialPolicyView(false, null, List.of(), List.of(), List.of());
public MemberCredentialPolicyView {
known = known == null ? List.of() : List.copyOf(known);
allowed = allowed == null ? List.of() : List.copyOf(allowed);
blocked = blocked == null ? List.of() : List.copyOf(blocked);
}
/** The honest "no policy configured" view. */
public static MemberCredentialPolicyView absent() {
return ABSENT;
}
/**
* Build the view straight from the live config. {@code creds} may be {@code null} (no {@code
* memberCredentials:} block at all) — treated the same as a present-but-empty block, exactly
* like {@link dev.ltms.fleet.Fleetd#reportMemberCredentialsGap} already did.
*/
public static MemberCredentialPolicyView of(FleetConfig.MemberCredentials creds) {
if (creds == null || creds.known().isEmpty()) {
return ABSENT;
}
return new MemberCredentialPolicyView(true, creds.policy(), creds.known(), creds.allow(),
List.copyOf(creds.blockedSet()));
}
public int knownCount() {
return known.size();
}
public int allowedCount() {
return allowed.size();
}
public int blockedCount() {
return blocked.size();
}
}
@@ -13,6 +13,7 @@ import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.msg.MessageService;
@@ -32,6 +33,7 @@ import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
@@ -64,6 +66,10 @@ public final class FleetApp {
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
// fleetd #111: re-read per request, same hot-reload shape as every other live config read —
// absent() (the honest "no policy configured" view) for every constructor that does not wire
// a real one, so existing legacy call sites keep building without knowing this field exists.
private final Supplier<MemberCredentialPolicyView> memberCredentials;
private final ObjectMapper mapper = new ObjectMapper();
/**
@@ -111,6 +117,19 @@ public final class FleetApp {
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable) {
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics,
deliverable, MemberCredentialPolicyView::absent);
}
/**
* @param memberCredentials live {@code memberCredentials:} policy view (fleetd #111), re-read
* per request for {@code GET /member-credentials}; production wiring
* passes the same hot-reload shape as every other live config read
*/
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials) {
this.herdr = herdr;
this.memberHerdr = memberHerdr != null ? memberHerdr : herdr;
this.workers = workers;
@@ -120,6 +139,7 @@ public final class FleetApp {
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
this.memberCredentials = memberCredentials != null ? memberCredentials : MemberCredentialPolicyView::absent;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
@@ -148,6 +168,7 @@ public final class FleetApp {
app.get("/agents", this::agents);
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured backend profiles
app.get("/member-credentials", this::memberCredentials); // fleetd #111: policy names + counts, never a value
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
app.delete("/members/{paneId}", this::stopMember);
app.post("/sessions/{id}/message", this::sendMessage); // fleet_send (primary; blocking, wait:false, or answer via turnId)
@@ -346,6 +367,29 @@ public final class FleetApp {
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
}
/**
* fleetd #111 (CB-608): the live {@code memberCredentials:} policy as names and counts —
* NEVER a value. The daemon does not hold a credential's value in the first place (only the
* name it is configured under), so there is nothing to redact here beyond what {@link
* MemberCredentialPolicyView} already omits by construction. This is the source
* {@code scripts/probe-member-credentials.sh} reads instead of carrying its own hardcoded
* name list, which is exactly what let the list drift silently behind the real policy.
*/
private void memberCredentials(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
MemberCredentialPolicyView view = memberCredentials.get();
ctx.status(200).json(Map.of(
"present", view.present(),
"policy", view.policy() == null ? "" : view.policy(),
"known", view.known(),
"allowed", view.allowed(),
"knownCount", view.knownCount(),
"allowedCount", view.allowedCount(),
"blockedCount", view.blockedCount()));
}
/**
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
@@ -0,0 +1,156 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #176: {@link Fleetd#leadSeatLookup} is the factory {@code Fleetd.main} wires into {@code
* FleetMcp.LeadSeatSource} so {@code fleet_list}'s {@code free} can subtract the seat(s) a
* {@code subscription: true} profile's own live LEAD session holds on that same account —
* {@code maxLoad} never counted the lead, only members. {@code FleetdLeadSeatWiringTest} proves
* {@code main} still passes this factory's result in; this class proves the factory's own matching
* logic: subscription-only, credential-matched, and counting only CURRENTLY LIVE leads.
*/
class FleetdLeadSeatLookupTest {
private static FleetConfig.Profile subscriptionProfile(String name, String credentialId) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, credentialId, null);
}
private static FleetConfig.Profile offSubscriptionProfile(String name, String credentialId) {
return new FleetConfig.Profile(name, "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
null, "tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 2, false, null, credentialId, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("a live lead sharing the target profile's credential counts as one seat")
void liveLeadSharingCredentialCountsAsOneSeat() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(1, lookup.apply("sonnet"));
}
@Test
@DisplayName("no live lead names this profile ⇒ zero seats, exactly as before this ticket")
void noLiveLeadOnTheProfileCountsAsZero() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders, Map::of);
assertEquals(0, lookup.apply("sonnet"));
}
@Test
@DisplayName("a lead entry with no `profile:` (recognise-only) contributes no seats — cannot be derived")
void recogniseOnlyLeadWithNoProfileContributesNothing() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
"claude", "claude-sonnet-5");
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(0, lookup.apply("sonnet"));
}
@Test
@DisplayName("a non-subscription profile never has a lead seat subtracted, whatever the credential match")
void nonSubscriptionProfileIsNeverAdjusted() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"terra", offSubscriptionProfile("terra", "shared-openai"),
"sonnet", subscriptionProfile("sonnet", "shared-openai"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(0, lookup.apply("terra"), "terra is not subscription:true, so it must never be adjusted");
}
@Test
@DisplayName("explicit, different credentialIds still separate two subscription profiles (post fleetd #176 "
+ "stage 2 sentinel)")
void differentCredentialIsNotCounted() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"sonnet", subscriptionProfile("sonnet", "claude-account-a"),
"opus", subscriptionProfile("opus", "claude-account-b"));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(0, lookup.apply("sonnet"), "different accounts must never be conflated into one seat count "
+ "— an explicit credentialId on both sides must still win over the subscription sentinel, so an "
+ "operator with two separate Claude logins on one host can keep them apart");
}
/**
* fleetd #176 stage 2 — the exact live shape that shipped inert: a lead on subscription profile
* {@code opus}, members on a DIFFERENTLY NAMED subscription profile {@code sonnet}, same Claude
* login, and NEITHER profile sets {@code credentialId}. Every other test in this class puts the
* lead on the SAME profile name as the target, which happened to keep working even with the old
* fall-back-to-profile-name {@code effectiveCredentialId()} — this is the one that did not, and
* its absence is what let the stage-1 fix ship without ever catching the bug it was filed for.
*/
@Test
@DisplayName("[LIVE SHAPE] lead on a DIFFERENT subscription profile, same account, neither sets "
+ "credentialId ⇒ still counts as a seat")
void leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"opus", subscriptionProfile("opus", null),
"sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_primary", "primary"));
assertEquals(1, lookup.apply("sonnet"), "opus and sonnet are both subscription:true with no explicit "
+ "credentialId, so they share one Claude login and the lead's live seat on opus must be charged "
+ "against sonnet too — this is the live host's actual shape (fleetd #176 stage 2)");
}
@Test
@DisplayName("two live instances of the same lead count as two seats")
void twoLiveInstancesOfTheSameLeadCountAsTwoSeats() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders,
() -> Map.of("term_a", "primary", "term_b", "primary"));
assertEquals(2, lookup.apply("sonnet"));
}
@Test
@DisplayName("an unconfigured target profile resolves to zero, not a thrown exception")
void unconfiguredTargetProfileIsZero() {
Function<String, Integer> lookup = Fleetd.leadSeatLookup(Map::of, Map.of(), Map::of);
assertEquals(0, lookup.apply("ghost"));
}
@Test
@DisplayName("live leads are read through the supplier on every call, not snapshotted")
void liveLeadsAreReadThroughOnEveryCall() {
Map<String, FleetConfig.Profile> profiles = Map.of("sonnet", subscriptionProfile("sonnet", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("sonnet"));
java.util.Map<String, String> live = new java.util.HashMap<>();
Function<String, Integer> lookup = Fleetd.leadSeatLookup(() -> profiles, leaders, () -> live);
assertEquals(0, lookup.apply("sonnet"));
live.put("term_primary", "primary");
assertEquals(1, lookup.apply("sonnet"));
}
}
@@ -0,0 +1,43 @@
package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
* this ticket — green, because none of them go through {@code main} at all.
*
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
* {@code FleetMcp} and never runs {@code main}.
*/
class FleetdLeadSeatWiringTest {
private static String fleetdSource() throws Exception {
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
}
@Test
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
String source = fleetdSource();
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
+ "leaders, leads))"),
"FleetMcp's construction call must still pass a LeadSeatSource built from "
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
+ "existing behavioural test green — this source check is what must go red instead.");
}
}
@@ -6,11 +6,18 @@ import dev.ltms.fleet.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.ParameterizedType;
import java.lang.reflect.RecordComponent;
import java.lang.reflect.Type;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.TreeMap;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.*;
@@ -1366,9 +1373,18 @@ class FleetConfigTest {
}
/**
* Every optional knob the example documents must bind under the exact spelling used there.
* Keep this list in step with {@code fleetd.example.yaml}: a rename that updates the record
* but not the example (or vice versa) fails here instead of silently no-op'ing in production.
* A spot check that the knobs listed below bind under the exact spelling the example uses —
* it asserts real VALUES arrive in the record, which no name-matching guard can do.
*
* <p><b>This is NOT a coverage guard, and must not be read as one</b> (fleetd #113). The list
* inside it is hand-written, so it only ever covers what someone remembered to add. Coverage
* — "is every key the code reads documented, and does every documented key bind?" — comes
* from {@link #everyNestedConfigKeyIsDocumentedInTheExample} and
* {@link #everyLiveKeyInTheExampleBindsToARecordComponent}, both of which derive their key
* set from the record tree and therefore cannot drift.
*
* <p>Adding a knob here is optional. Leaving one out is not a coverage gap, because the two
* derived guards above already fail on an undocumented or unbindable key.
*/
@Test
void everyOptionalKnobDocumentedInTheExampleBinds(@TempDir Path dir) throws Exception {
@@ -1525,6 +1541,230 @@ class FleetConfigTest {
return p.matcher(yaml).find();
}
/**
* The nested half of {@link #everyKnownTopLevelKeyIsDocumentedInTheExample} (fleetd #113).
*
* <p>That guard walks {@link FleetConfig#KNOWN_TOP_LEVEL_KEYS} and anchors its regex at
* column 0, so it sees ONLY top-level keys. Every nested key — {@code profiles.<name>.model},
* {@code health.paneProbeIntervalSeconds} and a hundred others — is outside its scope, and
* neither its name nor its output says so. A green run then reads as "the example documents
* the schema" when most of the schema was never looked at.
*
* <p>This walks the record tree rather than a name list, so a key added to any nested record
* is covered the moment it compiles, with no edit here. That is the point: a hand-maintained
* second copy of a list always drifts from the thing it mirrors.
*
* <p><b>Scope, stated on purpose</b> (fleetd #113 criterion 3 — every check reports what it
* did and did not look at):
* <ul>
* <li>It checks each key NAME appears somewhere in the example as a YAML key, live or
* commented out. It does NOT check the key sits at the right path.</li>
* <li>It does NOT check a documented key is read by anything. {@code paneProbeIntervalSeconds}
* is parsed into {@link FleetConfig.Health} and used nowhere, and this guard passes it.
* Proving a key is live code needs a call graph, which this is not.</li>
* </ul>
*/
@Test
void everyNestedConfigKeyIsDocumentedInTheExample() throws Exception {
Path example = Path.of("fleetd.example.yaml");
assertTrue(Files.exists(example), "fleetd.example.yaml must ship next to the pom");
String text = Files.readString(example);
Map<String, String> pathByName = configKeyPaths();
// The denominator. An under-counting walk passes every subset check vacuously, which is
// the exact shape fleetd #113 collects — so the walk has to prove it descended at all.
// The floor is DERIVED, not a literal: the nested walk must find substantially more keys
// than the top-level set the old guard used, or it has not gone below the first level.
int topLevel = FleetConfig.KNOWN_TOP_LEVEL_KEYS.size();
assertTrue(pathByName.size() > topLevel * 2,
"the record walk found " + pathByName.size() + " config key(s) against "
+ topLevel + " top-level key(s) — it has stopped descending into the "
+ "nested records, so this guard would pass vacuously. Fix the walk "
+ "before trusting a green run.");
List<String> undocumented = pathByName.entrySet().stream()
.filter(e -> !keyDocumentedAnywhere(text, e.getKey()))
.map(Map.Entry::getValue)
.sorted()
.toList();
assertTrue(undocumented.isEmpty(), () -> "checked " + pathByName.size()
+ " config key(s) that FleetConfig can bind; " + undocumented.size()
+ " appear nowhere in fleetd.example.yaml: " + undocumented
+ " — document each one there, commented out if optional. fleetd.yaml is "
+ "gitignored, so the example is the only committed description of the schema.");
}
/**
* The other direction: a LIVE key in the example that {@link FleetConfig} cannot bind. That is
* a key an operator would copy into {@code fleetd.yaml} expecting it to do something, where it
* would be silently ignored.
*
* <p><b>Scope, stated on purpose:</b> only live (uncommented) keys are checked. Most of the
* example is commented-out prose, and that prose contains lines like {@code # mode: token}
* that are indistinguishable from keys by text alone. Parsing them would produce false
* failures, so they are deliberately out of scope — and saying so here is the point, rather
* than letting a reader assume the whole file was validated.
*/
@Test
void everyLiveKeyInTheExampleBindsToARecordComponent() throws Exception {
Path example = Path.of("fleetd.example.yaml");
String text = Files.readString(example);
List<List<String>> paths = liveKeyPaths(text);
assertTrue(paths.size() >= 20,
"only " + paths.size() + " live key path(s) were parsed out of the example — the "
+ "parser is not seeing the file, so this guard would pass vacuously.");
List<String> unbindable = paths.stream()
.filter(path -> !pathBinds(path))
.map(path -> String.join(".", path))
.distinct()
.sorted()
.toList();
assertTrue(unbindable.isEmpty(), () -> "checked " + paths.size()
+ " live key path(s) in fleetd.example.yaml; " + unbindable.size()
+ " bind to nothing in FleetConfig: " + unbindable
+ " — an operator copying one of these into fleetd.yaml gets silence, not an error.");
}
/**
* Every configuration key {@link FleetConfig} can bind, at every depth, as
* {@code name -> a dotted path to one place it appears}. Derived from the record components,
* so it cannot drift from the code.
*/
private static Map<String, String> configKeyPaths() {
Map<String, String> out = new TreeMap<>();
collectConfigKeys(FleetConfig.class, "", new HashSet<>(), out);
return out;
}
private static void collectConfigKeys(Class<?> type, String prefix, Set<String> seen,
Map<String, String> out) {
if (!type.isRecord() || !seen.add(type.getName())) {
return;
}
for (RecordComponent rc : type.getRecordComponents()) {
String path = prefix.isEmpty() ? rc.getName() : prefix + "." + rc.getName();
out.putIfAbsent(rc.getName(), path);
Class<?> nested = rc.getType();
if (nested.isRecord()) {
collectConfigKeys(nested, path, seen, out);
} else if (Map.class.isAssignableFrom(nested) || List.class.isAssignableFrom(nested)) {
Class<?> element = elementRecord(rc);
if (element != null) {
String childPrefix = Map.class.isAssignableFrom(nested)
? path + ".<name>" : path + "[]";
collectConfigKeys(element, childPrefix, seen, out);
}
}
}
}
/** The record type inside a {@code Map<String, X>} or {@code List<X>} component, else null. */
private static Class<?> elementRecord(RecordComponent rc) {
if (rc.getGenericType() instanceof ParameterizedType pt) {
Type[] args = pt.getActualTypeArguments();
if (args.length > 0 && args[args.length - 1] instanceof Class<?> c && c.isRecord()) {
return c;
}
}
return null;
}
/**
* True when {@code key} is documented in the example, in either of the two conventions that
* file actually uses:
* <ol>
* <li>as a YAML key at any indentation, live or commented out ({@code key:}); or</li>
* <li>in a prose block that describes a section's sub-keys, one per line, as
* {@code # key → what it does}.</li>
* </ol>
*
* <p>The second form is not decoration. {@code broker.uri} is documented ONLY that way, on
* purpose: writing it out as a copy-pasteable {@code uri: amqp://user:pass@host} invites an
* operator to paste a password into a file, which is the very thing {@code uriEnv} exists to
* avoid. A guard that demanded the key form would push the file toward doing that. So this
* encodes the convention the example really uses rather than imposing a new one.
*/
private static boolean keyDocumentedAnywhere(String yaml, String key) {
String quoted = Pattern.quote(key);
Pattern asYamlKey = Pattern.compile("(?m)^\\s*(?:#\\s*)?" + quoted + ":");
Pattern asProseEntry = Pattern.compile("(?m)^\\s*#\\s*" + quoted + "\\s+\u2192");
return asYamlKey.matcher(yaml).find() || asProseEntry.matcher(yaml).find();
}
/** Every live (uncommented) key in {@code yaml}, as a path from the document root. */
private static List<List<String>> liveKeyPaths(String yaml) {
Pattern keyLine = Pattern.compile("^(\\s*)([A-Za-z][A-Za-z0-9_]*):(\\s.*)?$");
List<String> stack = new ArrayList<>();
List<Integer> indents = new ArrayList<>();
List<List<String>> paths = new ArrayList<>();
for (String line : yaml.split("\n", -1)) {
if (line.isBlank() || line.stripLeading().startsWith("#")) {
continue;
}
Matcher m = keyLine.matcher(line);
if (!m.matches()) {
continue;
}
int indent = m.group(1).length();
while (!indents.isEmpty() && indents.get(indents.size() - 1) >= indent) {
indents.remove(indents.size() - 1);
stack.remove(stack.size() - 1);
}
indents.add(indent);
stack.add(m.group(2));
paths.add(List.copyOf(stack));
}
return paths;
}
/** True when a dotted YAML path resolves to something {@link FleetConfig} can bind. */
private static boolean pathBinds(List<String> path) {
Class<?> type = FleetConfig.class;
boolean nextSegmentIsAFreeFormName = false;
for (int i = 0; i < path.size(); i++) {
if (nextSegmentIsAFreeFormName) {
nextSegmentIsAFreeFormName = false;
continue;
}
RecordComponent rc = componentNamed(type, path.get(i));
if (rc == null) {
return false;
}
Class<?> t = rc.getType();
if (t.isRecord()) {
type = t;
} else if (Map.class.isAssignableFrom(t)) {
Class<?> element = elementRecord(rc);
if (element == null) {
return true; // Map<String,String>: its entries are data, not schema
}
type = element;
nextSegmentIsAFreeFormName = true;
} else {
// A scalar or a list of scalars: nothing may legitimately nest under it.
return i == path.size() - 1;
}
}
return true;
}
private static RecordComponent componentNamed(Class<?> type, String name) {
if (type == null || !type.isRecord()) {
return null;
}
for (RecordComponent rc : type.getRecordComponents()) {
if (rc.getName().equals(name)) {
return rc;
}
}
return null;
}
@Test
void placementDefaultsToFixedForExistingConfigs(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-placement.yaml");
@@ -2,6 +2,7 @@ package dev.ltms.fleet.inject;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.Fleetd;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -146,40 +147,16 @@ class BackendOutageFlowTest {
pushLoop = new ReplyPushLoop(registry, new AgentControl(leadClient), inbox, scheduler, 3, 50);
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>(pushLoop);
// --- mirrors Fleetd.main's backendErrorSink lambda EXACTLY: (1) mark BACKEND_ERROR,
// (2) resolve profile/credential via the roster, fail-loud + notify unmapped-target,
// (3) record in BackendOutagePolicy, (4) on a NEW incident, notify the lead. ------------
BackendErrorSink backendErrorSink = (target, matchedLine, reason) -> {
sessions.onBackendError(target, reason);
// fleetd #248 follow-up: this used to be a 30-line hand-copy of Fleetd.main's
// backendErrorSink lambda, with a comment promising it mirrored production "EXACTLY".
// That promise is exactly the problem: a copy proves the copy. Editing or deleting the
// real sink left this whole flow test green, because it never touched the real sink.
// #248 made Fleetd.backendErrorSink public precisely so a cross-package test could
// drive the real object, so this now calls it. Every assertion below is about
// production code again.
BackendErrorSink backendErrorSink =
Fleetd.backendErrorSink(sessions, () -> profiles, outagePolicy, pushLoopRef::get);
String profileName = sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.orElse(null);
FleetConfig.Profile profile = profileName == null ? null : profiles.get(profileName);
if (profile == null) {
ReplyPushLoop loop = pushLoopRef.get();
if (loop != null) {
loop.onBackendTargetUnmapped(target, reason);
}
return;
}
String credentialId = profile.effectiveCredentialId();
Optional<BackendOutagePolicy.Incident> incident = outagePolicy.record(credentialId, target, reason);
incident.ifPresent(inc -> {
List<String> affectedProfiles = profiles.values().stream()
.filter(p -> credentialId.equals(p.effectiveCredentialId()))
.map(FleetConfig.Profile::profile)
.sorted()
.toList();
ReplyPushLoop loop = pushLoopRef.get();
if (loop != null) {
loop.onBackendIncident(inc.id(), inc.targets(), credentialId, affectedProfiles,
(int) inc.remainingCoolOffSeconds());
}
});
};
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(),
ExhaustionSink.none(), patterns, backendErrorSink);
@@ -34,6 +34,7 @@ import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
@@ -651,6 +652,48 @@ class FleetMcpTest {
assertEquals(2, out.split("\"free\":0", -1).length - 1, out);
}
/**
* fleetd #176 stage 2 (correcting the inert stage 1): two {@code subscription: true} profiles,
* {@code opus} and {@code sonnet}, neither setting an explicit {@code credentialId} — the exact
* shape measured on the live Mac fleet. This is INTENDED, not a regression: a real Claude usage
* limit on the one login behind both profiles really does take out every profile running on it,
* the same way {@code credentialId: openai-shared} already lets two OpenAI-backed profiles share
* one quarantine (see {@code everyProfileSharingTheQuarantinedCredentialReportsZeroFree} above).
* The {@code credentialIdFor} function here is built the same way {@code Fleetd.main} wires it —
* {@code profile -> profiles.get(profile).effectiveCredentialId()} — so this proves the actual
* config-driven behaviour, not just {@code capacityView}'s arithmetic with a hand-picked string.
*/
@Test
void quarantiningOneSubscriptionProfileZeroesFreeOnTheOtherSharingTheSameAccount() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
Map<String, FleetConfig.Profile> profiles = Map.of(
"opus", new FleetConfig.Profile("opus", null, "claude-opus-4", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null),
"sonnet", new FleetConfig.Profile("sonnet", null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, 3, true, null, null, null));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine(FleetConfig.Profile.SUBSCRIPTION_CREDENTIAL_ID);
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(profile -> {
FleetConfig.Profile configured = profiles.get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 3,
() -> Set.of("opus", "sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertEquals(2, out.split("\"free\":0", -1).length - 1,
"opus and sonnet share one Claude login with neither setting credentialId, so quarantining "
+ "opus's account must also zero sonnet's free — this is intended, not a side effect: "
+ out);
assertEquals(2, out.split("\"credentialId\":\"" + FleetConfig.Profile.SUBSCRIPTION_CREDENTIAL_ID + "\"", -1)
.length - 1, out);
}
/** fleetd #201 Unit 5: cool-off forces {@code free:0} but never adds {@code quarantinedForSeconds}. */
@Test
void coolingOffProfileReportsZeroFreeButNeverQuarantinedForSeconds() {
@@ -767,6 +810,137 @@ class FleetMcpTest {
assertFalse(out.contains("quarantinedForSeconds"), out);
}
/**
* fleetd #257: this is the exact shape measured on the Mac fleet — {@code maxLoad:3, live:2},
* one lead sharing the subscription. The OLD formula ({@code max(0, cap - live - leadSeats)})
* reported {@code free:0} here, disagreeing with the real spawn gate (which never read
* {@code leadSeats} and would still grant one more spawn — see
* {@code freeMatchesWhatTheRealPlacementGateActuallyGrants} below for that proof against the
* actual gate). {@code free} must report {@code 1}: {@code leadSeats} is carried as a fact, not
* subtracted.
*/
@Test
void leadSeatIsReportedButNeverSubtractedFromFree() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(
profile -> "sonnet".equals(profile) ? 1 : 0);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 2, profile -> 3,
() -> Set.of("sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
assertTrue(out.contains("\"maxLoad\":3"), "maxLoad itself must be left untouched: " + out);
assertTrue(out.contains("\"live\":2"), out);
assertTrue(out.contains("\"free\":1"), "free is maxLoad - live only, never minus leadSeats: " + out);
assertTrue(out.contains("\"leadSeats\":1"), "leadSeats is still reported, just not subtracted: " + out);
}
/**
* fleetd #257: the OTHER measurement in the issue — a completely idle fleet with a lead sharing
* the subscription. {@code maxLoad:3, live:0} must report {@code free:3}, matching what the
* spawn gate (which has no notion of a lead's seat) would actually grant.
*/
@Test
void leadSeatDoesNotLowerFreeOnAnOtherwiseIdleSubscriptionProfile() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(
profile -> "sonnet".equals(profile) ? 1 : 0);
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 3,
() -> Set.of("sonnet"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":3"), "the real gate never subtracts the lead's seat: " + out);
assertTrue(out.contains("\"leadSeats\":1"), out);
}
/**
* fleetd #257 — the defect, driven against the REAL placement gate, not a copy of its
* arithmetic. {@code free} must equal exactly how many more spawns
* {@link CompositePeerLauncher#spawn} (routed through the same {@link SessionManager} fleet_spawn
* itself uses) will grant on this profile right now: this test reads whatever number
* {@code fleet_list}'s real {@code listFleet} call reports, then drives that many spawns through
* the REAL composite launcher and asserts every one succeeds, and the next one — one past what
* fleet_list promised — is refused. A test that instead hand-computed {@code cap - live} and
* compared it to {@code free} would pass even if both sides shared the same wrong formula (this
* repo has been bitten by exactly that before); this one only passes when fleet_list's number and
* the gate's real behaviour actually agree.
*
* <p>{@code liveCount} here is wired the same way {@code Fleetd.main} wires it in production: one
* function, read by both the {@link CompositePeerLauncher}'s {@code maxLoad} gate and
* {@code fleet_list}'s {@code CapacitySource}, off the SAME {@link SessionManager#roster()} — so
* the two paths cannot silently drift on what "live" means.
*/
@Test
void freeMatchesWhatTheRealPlacementGateActuallyGrants() {
FakeHerdr h = new FakeHerdr();
FleetConfig.Profile wcfg = new FleetConfig.Profile(
"sonnet", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null,
null, null, null, null, null, null, null, 3);
Map<String, FleetConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
new AgentControl(h), new WorkspaceControl(h), new SubscriptionGuard(Set.of("gx00.gw")),
profiles, wcfg.profile(), k -> "FLEETD_WORKER_TOKEN".equals(k) ? "tok" : null);
java.util.concurrent.atomic.AtomicReference<SessionManager> smRef =
new java.util.concurrent.atomic.AtomicReference<>();
Function<String, Integer> liveCount = profile -> (int) smRef.get().roster().stream()
.filter(s -> profile.equals(s.profile())).count();
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), liveCount);
SessionManager sm = new SessionManager(composite);
smRef.set(sm);
// Two members already live — the exact shape measured in fleetd #257 (maxLoad:3, live:2).
assertFalse(FleetMcp.spawn(sm, "sonnet").isError(), "setup: first live member must spawn cleanly");
assertFalse(FleetMcp.spawn(sm, "sonnet").isError(), "setup: second live member must spawn cleanly");
// A lead session shares this profile's subscription — leadSeats:1, same as the ticket.
FleetMcp.LeadSeatSource leadSeats = new FleetMcp.LeadSeatSource(p -> "sonnet".equals(p) ? 1 : 0);
String out = textOf(FleetMcp.listFleet(composite, sm, null,
new FleetMcp.CapacitySource(liveCount, p -> profiles.get(p).maxLoad(), profiles::keySet, () -> 0),
new FleetMcp.HealthCoverageSource(() -> "off"), FleetMcp.QuarantineSource.none(),
FleetMcp.OutageSource.none(), leadSeats, Map.of(), "", null));
int reportedFree = extractInt(out, "free");
for (int i = 0; i < reportedFree; i++) {
McpSchema.CallToolResult res = FleetMcp.spawn(sm, "sonnet");
assertFalse(res.isError(), "fleet_list promised free:" + reportedFree + "; spawn #" + (i + 1)
+ " of that many was refused by the real gate: " + textOf(res));
}
McpSchema.CallToolResult overflow = FleetMcp.spawn(sm, "sonnet");
assertTrue(overflow.isError(), "fleet_list reported free:" + reportedFree
+ " but the real placement gate granted at least one more spawn than that: " + textOf(overflow));
}
private static int extractInt(String json, String key) {
java.util.regex.Matcher m = java.util.regex.Pattern.compile("\"" + key + "\":(-?\\d+)").matcher(json);
assertTrue(m.find(), "no \"" + key + "\" field in: " + json);
return Integer.parseInt(m.group(1));
}
/** A profile with no lead seats reported must be byte-identical to before this ticket. */
@Test
void zeroLeadSeatsOmitsTheKeyAndLeavesFreeUnchanged() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
String out = textOf(FleetMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new FleetMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
Map.of(), "", null));
assertTrue(out.contains("\"free\":2"), out);
assertFalse(out.contains("leadSeats"), "no lead shares this profile's credential: " + out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
@@ -0,0 +1,121 @@
package dev.ltms.fleet.mcp;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashSet;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #114 (CB-609): the guard that lets {@code docs/MCP-Contract.md} name a tool at all.
*
* <p>That page was written in July 2026, before any MCP code existed, and then did not follow the
* code. By August it named two tools that had never been built, omitted five that shipped, had the
* wrong name for nearly every parameter, and still described a caller-identity rule that was a
* privilege bug by then. Nothing failed, because nothing checked it — and {@code CLAUDE.md} sends
* every session in the fleet to that page.
*
* <p>The fix was to delete the tool catalogue rather than correct it: a hand-maintained second copy
* of the tool surface is the defect, not the particular errors it had accumulated. What survives is
* the flows, which are shapes rather than names. But the flows still have to say {@code fleet_send}
* somewhere to be readable, and that is exactly the sentence that rots. This test is what makes it
* safe to write.
*
* <p><b>It checks source text, not behaviour.</b> It reads the Markdown and reads {@link FleetMcp}'s
* source, and it only catches a name in the doc that the server does not register. It cannot catch a
* flow that describes the wrong order, or a parameter name in prose — those are not name-shaped. The
* doc's own header carries that caveat for its readers.
*/
class McpContractDocTest {
/** Tests run with the module directory as cwd, so the repo-root doc is one level up. */
private static final Path DOC = Path.of("../docs/MCP-Contract.md");
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
private static Set<String> matches(Path file, String regex) throws Exception {
Matcher m = Pattern.compile(regex).matcher(Files.readString(file));
Set<String> found = new LinkedHashSet<>();
while (m.find()) {
found.add(m.group(1));
}
return found;
}
/** Every {@code fleet_*} the doc mentions, in prose or in a diagram. */
private static Set<String> toolsNamedInTheDoc() throws Exception {
return matches(DOC, "(fleet_[a-z_]+)");
}
/** Every tool {@link FleetMcp} actually registers, read from its {@code tool("…")} calls. */
private static Set<String> toolsTheServerRegisters() throws Exception {
return matches(MCP_SOURCE, "tool\\(\"(fleet_[a-z_]+)\"");
}
@Test
@DisplayName("[SOURCE TEXT] every fleet_* tool named in MCP-Contract.md is one the server registers")
void theDocNamesNoToolThatDoesNotExist() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> named = toolsNamedInTheDoc();
Set<String> unknown = new LinkedHashSet<>(named);
unknown.removeAll(registered);
assertTrue(unknown.isEmpty(),
"docs/MCP-Contract.md names " + unknown + ", which FleetMcp does not register. "
+ "Checked " + named.size() + " name(s) in the doc against " + registered.size()
+ " registered tool(s): " + registered + ". This is the fleetd #114 defect "
+ "recurring — the doc named fleet_read and fleet_cancel for weeks after the "
+ "code shipped without them. Either fix the name or drop it from the page; do "
+ "NOT weaken this test.");
}
/**
* The denominator guard. The check above passes trivially if the doc stops naming any tool at
* all — an empty set is a subset of everything. A checker that can silently check nothing is the
* fleetd #113 shape, so this pins that the doc really is still describing the flows, and that
* the registration scrape really did find the server's tools.
*/
@Test
@DisplayName("[SOURCE TEXT] the doc/server name check is not vacuous — both sides found names")
void theCheckActuallyHasSomethingToCheck() throws Exception {
Set<String> registered = toolsTheServerRegisters();
Set<String> named = toolsNamedInTheDoc();
assertTrue(registered.size() >= 10,
"scraped only " + registered.size() + " tool registrations from FleetMcp (" + registered
+ "); the server registers eleven, so the tool(\"…\") scrape has stopped matching "
+ "and the check above is now vacuous");
assertTrue(named.size() >= 4,
"docs/MCP-Contract.md names only " + named.size() + " fleet_* tool(s) (" + named + "). "
+ "The flows describe delegation, clarification, detached delivery and the "
+ "turn-done fallback, so it should name several. Too few means the page has been "
+ "gutted and this test is guarding nothing.");
}
/**
* fleetd #114's actual lesson. The catalogue was deleted on purpose; a well-meaning "let me just
* document the tools here" restores the exact second copy that drifted for a month.
*/
@Test
@DisplayName("[SOURCE TEXT] MCP-Contract.md still says it is not the tool reference")
void theDocStillDisclaimsBeingTheToolReference() throws Exception {
String doc = Files.readString(DOC);
assertTrue(doc.contains("**What this page is NOT: a tool reference.**"),
"docs/MCP-Contract.md must keep saying it is not the tool reference. That sentence is "
+ "the fix for fleetd #114: the page carried a hand-maintained tool catalogue that "
+ "drifted from the code for a month while CLAUDE.md pointed every session at it.");
assertEquals(0, countTables(doc.substring(0, doc.indexOf("## 1. Rendezvous flows"))),
"the header of docs/MCP-Contract.md must not grow a tool/parameter table — that is the "
+ "second copy fleetd #114 deleted");
}
private static int countTables(String markdown) {
return (int) markdown.lines().filter(l -> l.strip().startsWith("|")).count();
}
}
@@ -19,11 +19,14 @@ import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import org.junit.jupiter.api.AfterAll;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFileAttributeView;
@@ -32,10 +35,12 @@ import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.TreeSet;
import java.util.UUID;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Supplier;
@@ -147,18 +152,25 @@ class ClaudeCodeLauncherTest {
// every ide_* call to the member's own worktree via the charter.
/** A profile carrying an ideMcpUrl (plus optional bridge mcpUrl and cwd). ideMcpUrl is the last record component. */
private FleetConfig.Profile ideProfile(String mcpUrl, String ideMcpUrl, String cwd) {
/**
* {@code configDir} is first and mandatory on purpose (fleetd #258). A profile that sets no
* {@code configDir} sends {@code seedTrustDialog}'s write to the operator's real
* {@code ~/.claude.json}, and the fleetd #149 gate does not stop that when the fixture's
* {@code cwd} is worktree-shaped — which every IDE-overlay fixture's is. Pass a {@code @TempDir}
* whenever {@code cwd} has a {@code .git} FILE; {@code null} is only safe when it does not.
*/
private FleetConfig.Profile ideProfile(String configDir, String mcpUrl, String ideMcpUrl, String cwd) {
return new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
"ltms-local", "http://gx00.gw:8000", "coder", configDir, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "w #{n}", mcpUrl, cwd, null,
null, null, null, null, null, null, null, null, null, ideMcpUrl);
}
/** As {@link #ideProfile} but carrying the CB-634 auto-open fields (module subdir + open command). */
private FleetConfig.Profile ideProfileModule(String ideMcpUrl, String cwd, String ideProjectDir,
String ideOpenCommand) {
private FleetConfig.Profile ideProfileModule(String configDir, String ideMcpUrl, String cwd,
String ideProjectDir, String ideOpenCommand) {
return new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
"ltms-local", "http://gx00.gw:8000", "coder", configDir, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "w #{n}", null, cwd, null,
null, null, null, null, null, null, null, null, null, ideMcpUrl, ideProjectDir, ideOpenCommand);
}
@@ -171,7 +183,7 @@ class ClaudeCodeLauncherTest {
@Test
void mountsIdeMcpAsSecondServerWhenIdeMcpUrlSet() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, ideProfile("http://127.0.0.1:8765/mcp",
launcher(herdr, ideProfile(null, "http://127.0.0.1:8765/mcp",
"http://127.0.0.1:29170/index-mcp/streamable-http", null)).spawn();
List<String> args = spawnedArgs(herdr);
@@ -186,7 +198,8 @@ class ClaudeCodeLauncherTest {
@Test
void ideMcpUrlAloneStillEmitsTheMount() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, ideProfile(null, "http://127.0.0.1:29170/index-mcp/streamable-http", null)).spawn();
launcher(herdr, ideProfile(null, null, "http://127.0.0.1:29170/index-mcp/streamable-http", null))
.spawn();
List<String> args = spawnedArgs(herdr);
assertTrue(args.contains("--mcp-config"),
@@ -201,7 +214,7 @@ class ClaudeCodeLauncherTest {
FakeHerdr herdr = new FakeHerdr();
String roleCharter = "You review changes.";
String worktree = "/tmp/.fleet-worktrees/rev-1";
FleetConfig.Profile cfg = ideProfile("http://127.0.0.1:8765/mcp",
FleetConfig.Profile cfg = ideProfile(null, "http://127.0.0.1:8765/mcp",
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
@@ -226,6 +239,62 @@ class ClaudeCodeLauncherTest {
}, "the --append-system-prompt-file path must be a readable file");
}
// ---- fleetd #258: no fixture in this class may seed the operator's real ~/.claude.json ----
//
// seedTrustDialog writes <configDir>/.claude.json, or ~/.claude.json when the profile sets no
// configDir. The fleetd #149 gate (isProvisionedWorktree) closes the case where a fixture leaves
// cwd unset and it falls back to user.dir. It does NOT close the case where a fixture builds a
// worktree-shaped @TempDir on purpose — the gate opens, and a null configDir still points the
// write at the real home. Two IDE-overlay fixtures did exactly that on EVERY run; by 2026-09-03
// the operator's ~/.claude.json carried 116 dead JUnit temp paths, none of which still existed.
//
// DIFFERENTIAL, not absolute: it snapshots the temp-dir project keys already present and fails
// only on keys this class ADDS. An absolute check would fail on any host still carrying the
// historical entries, and a check that fails for a reason nobody can fix gets deleted, not fixed.
private static Set<String> tempProjectKeysBefore;
@BeforeAll
static void snapshotTempProjectKeysInTheDefaultClaudeJson() {
tempProjectKeysBefore = tempProjectKeysInDefaultClaudeJson();
}
@AfterAll
static void noFixtureSeededTheDefaultClaudeJson() {
Set<String> added = new TreeSet<>(tempProjectKeysInDefaultClaudeJson());
added.removeAll(tempProjectKeysBefore);
assertTrue(added.isEmpty(),
"a fixture in this class seeded the DEFAULT .claude.json (the operator's real file "
+ "when user.home is not redirected) with " + added.size() + " temp-dir "
+ "project entry/entries: " + added + ". Give that fixture's profile a "
+ "@TempDir configDir — see ideProfile's javadoc.");
}
/**
* Project keys under the JVM temp dir in {@code <user.home>/.claude.json}, or an empty set when
* the file is absent or unreadable. Only key NAMES are read; nothing in the operator's file is
* copied, asserted on, or written back.
*/
private static Set<String> tempProjectKeysInDefaultClaudeJson() {
Set<String> keys = new TreeSet<>();
Path target = Path.of(System.getProperty("user.home"), ".claude.json");
if (!Files.isRegularFile(target)) {
return keys;
}
String tmp = System.getProperty("java.io.tmpdir");
try {
JsonNode projects = new ObjectMapper().readTree(target.toFile()).path("projects");
projects.fieldNames().forEachRemaining(name -> {
if (name.startsWith(tmp) || name.contains("/junit-")) {
keys.add(name);
}
});
} catch (IOException e) {
return keys; // unreadable file proves nothing either way
}
return keys;
}
// CB-634: the IDE guidance is delivered as a CLAUDE.local.md overlay (written only into a
// provisioned worktree — cwd with a `.git` FILE) and registered in the repository's COMMON
// info/exclude. git reads a worktree's excludes from the common dir, not the per-worktree
@@ -241,8 +310,8 @@ class ClaudeCodeLauncherTest {
Files.writeString(worktree.resolve(".git"), "gitdir: " + gitDir);
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, ideProfile(null, "http://127.0.0.1:29170/index-mcp/streamable-http",
worktree.toString())).spawn();
launcher(herdr, ideProfile(root.toString(), null,
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree.toString())).spawn();
Path overlay = worktree.resolve("CLAUDE.local.md");
assertTrue(Files.exists(overlay), "the overlay is written beside the project's CLAUDE.md");
@@ -268,8 +337,9 @@ class ClaudeCodeLauncherTest {
FakeHerdr herdr = new FakeHerdr();
// ideProjectDir "fleetd" ⇒ the pin is <worktree>/fleetd, not <worktree>.
launcher(herdr, ideProfileModule("http://127.0.0.1:29170/index-mcp/streamable-http",
worktree.toString(), "fleetd", null)).spawn();
launcher(herdr, ideProfileModule(root.toString(),
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree.toString(), "fleetd", null))
.spawn();
Path overlay = worktree.resolve("CLAUDE.local.md");
assertTrue(Files.exists(overlay), "the overlay file still lives at the worktree root");
@@ -304,8 +374,8 @@ class ClaudeCodeLauncherTest {
Files.createDirectories(worktree.resolve(".git"));
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, ideProfile(null, "http://127.0.0.1:29170/index-mcp/streamable-http",
worktree.toString())).spawn();
launcher(herdr, ideProfile(root.toString(), null,
"http://127.0.0.1:29170/index-mcp/streamable-http", worktree.toString())).spawn();
assertFalse(Files.exists(worktree.resolve("CLAUDE.local.md")),
"the safety gate refuses to write into a non-worktree cwd (.git directory)");
@@ -1409,14 +1479,16 @@ class ClaudeCodeLauncherTest {
}
/**
* CB-633 follow-up (#192): the mirror of the zsh test above. On a non-zsh login shell {@code
* ZDOTDIR} is ignored, so no scrub ever runs — the report must keep the WARN wording (a name here
* really is inherited unblocked) rather than claiming a scrub protects it. This is the trap PR
* #174 fell into the other direction: keying the wording on the shell, not on {@code
* creds.isAllowList()}, is what keeps this branch correct.
* fleetd #155 (was CB-633 follow-up (#192) "allowListPolicyOnNonZshKeepsTheWarnWording"): on a
* non-zsh login shell {@code ZDOTDIR} is ignored, so no scrub could ever run there. Before #155
* the launcher degraded to the weaker overlay and spawned anyway; that is exactly the "control
* silently does nothing" failure #155 exists to close, since {@code policy: allow-list} is the
* operator explicitly asking for a control a sourced file cannot undo. Now the spawn is REFUSED
* instead — this test asserts the refusal, naming the shell, on the real spawn path ({@link
* ClaudeCodeLauncher#spawn}), not merely on the launcher method that computes it.
*/
@Test
void allowListPolicyOnNonZshKeepsTheWarnWording() {
void allowListPolicyOnNonZshRefusesTheSpawn() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
@@ -1430,27 +1502,12 @@ class ClaudeCodeLauncherTest {
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
}
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, svc::spawn,
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade "
+ "to the weaker overlay");
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("UNBLOCKED")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"no scrub runs on a non-zsh shell, so the WARN wording must be kept — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().toLowerCase(java.util.Locale.ROOT).contains("scrub")),
"nothing is scrubbed on this path, so the report must not claim otherwise — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
}
/**
@@ -2159,8 +2216,7 @@ class ClaudeCodeLauncherTest {
}
JsonNode root = new ObjectMapper().readTree(claudeJson.toFile());
JsonNode project = root.path("projects").path(worktree.toString());
seededBeforeStart.set(project.path("hasTrustDialogAccepted").asBoolean(false)
&& project.path("hasCompletedProjectOnboarding").asBoolean(false));
seededBeforeStart.set(project.path("hasTrustDialogAccepted").asBoolean(false));
} catch (IOException e) {
seededBeforeStart.set(false);
}
@@ -2178,7 +2234,7 @@ class ClaudeCodeLauncherTest {
}
@Test
void seedTrustDialogWritesBothTrustFlagsForTheResolvedCwd(
void seedTrustDialogWritesOnlyTheTrustFlagAndNotTheOnboardingKey(
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
markAsProvisionedWorktree(worktree);
FakeHerdr herdr = new FakeHerdr();
@@ -2192,7 +2248,13 @@ class ClaudeCodeLauncherTest {
JsonNode project = new ObjectMapper().readTree(claudeJson.toFile())
.path("projects").path(worktree.toString());
assertTrue(project.path("hasTrustDialogAccepted").asBoolean(false));
assertTrue(project.path("hasCompletedProjectOnboarding").asBoolean(false));
// fleetd #247: the onboarding key must NOT be written. Claude Code strips it on every
// save (measured: 0 of 28 live entries had it, including its own), so writing it only
// adds a contested key to a file two processes share. This assertion is the guard that
// stops it coming back as a plausible-looking "completeness" fix.
assertFalse(project.has("hasCompletedProjectOnboarding"),
"hasCompletedProjectOnboarding must not be written — Claude Code drops it on "
+ "every save, and the member reaches idle on hasTrustDialogAccepted alone");
}
/**
@@ -2238,7 +2300,6 @@ class ClaudeCodeLauncherTest {
JsonNode mine = root.path("projects").path(worktree.toString());
assertTrue(mine.path("hasTrustDialogAccepted").asBoolean(false));
assertTrue(mine.path("hasCompletedProjectOnboarding").asBoolean(false));
}
/**
@@ -2535,4 +2596,150 @@ class ClaudeCodeLauncherTest {
+ tornRead.get());
assertEquals(newContent, Files.readString(target), "the final content must be the new content");
}
// --- fleetd #247: CAS against a writer TRUST_JSON_LOCK cannot reach --------------------------
//
// fleetd #149's lock only serialises seedTrustDialog calls THIS launcher makes inside this one
// JVM. It does nothing about the one writer that actually shares this file on a real host: the
// operator's own live Claude Code, whose CLAUDE_CONFIG_DIR is routinely the very configDir this
// profile is given. A plain read-modify-write there is a routine lost update — fleetd reads v1,
// the operator's session writes v2, fleetd's ATOMIC_MOVE lands v3 built from v1, and v2 is gone,
// atomically. The fix re-reads the file's exact bytes immediately before the move and compares
// them with what the update was built from, retrying from fresh bytes on a mismatch.
//
// Both tests below drive the race through the real seedTrustDialog/spawn() path (not a
// hand-rolled call to some extracted primitive) using trustJsonCasTestHook — a seam fired once
// per CAS attempt, at the exact point between the read and the final compare, so the race is
// deterministic instead of depending on real thread timing.
@Test
void seedTrustDialogRetriesAndPreservesAConcurrentExternalWritersChange(
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
markAsProvisionedWorktree(worktree);
Path claudeJson = configDir.resolve(".claude.json");
Files.writeString(claudeJson, "{\"projects\":{}}");
// Fires exactly once, on the first CAS attempt — simulating the operator's own live Claude
// Code landing its own write to this SAME file in the gap between fleetd's read and write.
AtomicBoolean fired = new AtomicBoolean(false);
ClaudeCodeLauncher.trustJsonCasTestHook = () -> {
if (fired.compareAndSet(false, true)) {
try {
Files.writeString(claudeJson,
"{\"projects\":{\"/operator/own/project\":"
+ "{\"hasTrustDialogAccepted\":true}}}");
} catch (IOException e) {
throw new UncheckedIOException(e);
}
}
};
try {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
.spawn();
} finally {
ClaudeCodeLauncher.trustJsonCasTestHook = () -> {};
}
assertTrue(fired.get(), "the race hook must actually have fired during the spawn");
JsonNode root = new ObjectMapper().readTree(claudeJson.toFile());
assertTrue(root.path("projects").path("/operator/own/project")
.path("hasTrustDialogAccepted").asBoolean(false),
"the external writer's change, landed between fleetd's read and write, must SURVIVE "
+ "— this is the whole point of the CAS: without it, fleetd's stale-built "
+ "ATOMIC_MOVE would have silently discarded it. Final file: "
+ Files.readString(claudeJson));
assertTrue(root.path("projects").path(worktree.toString())
.path("hasTrustDialogAccepted").asBoolean(false),
"fleetd's own retry must still land its own trust entry, built from the fresh bytes");
}
/**
* Retry-exhaustion path: an external writer that changes the file on EVERY attempt (not just
* once) exhausts all {@code MAX_TRUST_JSON_CAS_ATTEMPTS} retries. fleetd must then write nothing
* at all — not even a partial/best-effort write — and log a WARN naming the file it gave up on.
*/
@Test
void seedTrustDialogWritesNothingAndWarnsWhenCasRetriesAreExhausted(
@TempDir Path configDir, @TempDir Path worktree) throws Exception {
markAsProvisionedWorktree(worktree);
Path claudeJson = configDir.resolve(".claude.json");
Files.writeString(claudeJson, "{\"marker\":\"start\"}");
AtomicInteger hookCalls = new AtomicInteger();
ClaudeCodeLauncher.trustJsonCasTestHook = () -> {
try {
Files.writeString(claudeJson,
"{\"marker\":\"race-" + hookCalls.incrementAndGet() + "\"}");
} catch (IOException e) {
throw new UncheckedIOException(e);
}
};
Logger logger = (Logger) LoggerFactory.getLogger(ClaudeCodeLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = trustProfile(configDir.toString(), worktree.toString());
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
.spawn();
} finally {
logger.detachAppender(appender);
ClaudeCodeLauncher.trustJsonCasTestHook = () -> {};
}
assertEquals(5, hookCalls.get(), "the race hook must fire exactly once per CAS attempt");
String finalContent = Files.readString(claudeJson);
assertEquals("{\"marker\":\"race-" + hookCalls.get() + "\"}", finalContent,
"the file must be left exactly as the external writer last left it — fleetd must not "
+ "have written at all once retries are exhausted");
assertFalse(finalContent.contains("hasTrustDialogAccepted"),
"the trust entry must never appear — its presence would mean the CAS gave up and "
+ "wrote anyway instead of skipping the seed");
assertTrue(appender.list.stream().anyMatch(e -> e.getLevel() == Level.WARN
&& e.getFormattedMessage().contains("gave up seeding workspace-trust")
&& e.getFormattedMessage().contains(worktree.toString())
&& e.getFormattedMessage().contains(claudeJson.toString())),
"a WARN naming both the cwd and the file it gave up on must be logged: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* The loud-default WARN (fleetd #247): a profile with no {@code configDir} targets the
* operator's real {@code ~/.claude.json} (redirected here to a {@code @TempDir}), and every time
* that happens must be logged, not just detected.
*/
@Test
void seedTrustDialogWarnsEveryTimeItTargetsTheDefaultClaudeJson(
@TempDir Path fakeHome, @TempDir Path worktree) throws Exception {
markAsProvisionedWorktree(worktree);
String originalHome = System.getProperty("user.home");
System.setProperty("user.home", fakeHome.toString());
Logger logger = (Logger) LoggerFactory.getLogger(ClaudeCodeLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = trustProfile(null, worktree.toString());
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
.spawn();
} finally {
logger.detachAppender(appender);
System.setProperty("user.home", originalHome);
}
assertTrue(appender.list.stream().anyMatch(e -> e.getLevel() == Level.WARN
&& e.getFormattedMessage().contains("no configDir set")
&& e.getFormattedMessage().contains(worktree.toString())),
"a WARN naming the cwd must fire when the profile sets no configDir: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
}
@@ -29,6 +29,7 @@ import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
@@ -99,18 +100,26 @@ class HerdrPeerLauncherAllowListWiringTest {
}
/**
* A non-zsh shell cannot read {@code ZDOTDIR} at all. The launcher must fall back rather than
* generate a directory nothing will ever read — a directory that would look like protection.
* fleetd #155: a non-zsh shell cannot read {@code ZDOTDIR} at all, so under {@code
* policy: allow-list} — the policy the operator picked specifically for a control a sourced file
* cannot undo — the launcher must REFUSE the spawn rather than silently generate a directory
* nothing will ever read (protection theatre) or fall back to the weaker overlay (exactly the
* "control silently does nothing" defect this ticket exists to close). Real path: through {@link
* HerdrPeerLauncher#spawn}, the method the daemon actually calls.
*/
@Test
void aNonZshShellGeneratesNothingAndFallsBack() {
void aNonZshShellUnderAllowListPolicyRefusesTheSpawn() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash");
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade");
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"bash ignores ZDOTDIR; setting it would be protection theatre");
"a refused spawn must not have generated (or wired in) a scrub directory: " + launcher.env);
}
private static Supplier<FleetConfig.MemberCredentials> allowList() {
@@ -246,11 +255,14 @@ class HerdrPeerLauncherAllowListWiringTest {
* "allowed N of M" line — which describes what the scrub does — must not be printed there either.
* Before this fix the line was logged BEFORE the zsh gate, so a non-zsh host printed e.g.
* "allowed 1 of 3" while blocking nothing at all, telling an operator a control ran when it did
* not. Real path: goes through {@link HerdrPeerLauncher#spawn}, same as the sibling test above,
* with the shell fixed to bash so the fallback branch is the one exercised.
* not. fleetd #155: the spawn itself is now refused on this path (see {@code
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}) rather than falling back — this test's own
* concern still holds under the refusal: the "allowed N of M" line describes a scrub that never
* ran here, so it must still never appear. Real path: goes through {@link
* HerdrPeerLauncher#spawn}, same as the sibling test above, with the shell fixed to bash.
*/
@Test
void noAllowedCountLineIsEmittedOnTheNonZshFallbackPath() {
void noAllowedCountLineIsEmittedOnTheNonZshRefusalPath() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of(INJECTED, "SOME_UNRELATED_NAME", "ANOTHER_UNRELATED_NAME");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash", () -> hostEnvNames);
@@ -262,7 +274,9 @@ class HerdrPeerLauncherAllowListWiringTest {
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"policy=allow-list on a non-zsh shell must refuse the spawn");
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
@@ -310,13 +324,20 @@ class HerdrPeerLauncherAllowListWiringTest {
* different OS user — {@link HerdrPeerLauncher#hostEnvNames} describes fleetd's own process, not
* that user's. Neither "inherits them UNBLOCKED" nor "scrub blanks them" is evidence-backed
* there, so neither may print; the single unknown-environment WARN must, naming the config key.
*
* <p>fleetd #155: {@code memberLoginShell} is now given explicitly as zsh, so the spawn reaches
* this WARN through the (still-degrading, not refusing) missing-{@code worktreeRoot}/{@code
* worktreeGroup} fallback rather than through the zsh gate, which now refuses instead of falling
* back — see {@code aNonZshShellUnderAllowListPolicyRefusesTheSpawn}. This test's own concern
* (the unknown-environment WARN) is orthogonal to which fallback reached {@code
* logCredentialGap}, so it still holds.
*/
@Test
void gapDetectorReportsUnknownInsteadOfAConclusionWhenMemberHerdrSocketIsConfigured() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
List<String> messages = spawnAndCaptureLogs(launcher);
@@ -340,8 +361,9 @@ class HerdrPeerLauncherAllowListWiringTest {
void theUnknownEnvironmentWarnFiresOnceNotOncePerSpawn() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
// fleetd #155: memberLoginShell given explicitly as zsh — see the sibling test above for why.
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
@@ -389,42 +411,44 @@ class HerdrPeerLauncherAllowListWiringTest {
}
/**
* fleetd #213 defect 1, acceptance criterion 1: {@code memberHerdrSocket} configured and {@code
* memberLoginShell} configured as non-zsh must fall back to the sentinel overlay exactly like a
* non-zsh {@code $SHELL} does today — and the "generated ZDOTDIR" INFO must not appear, since no
* scrub actually runs. The WiringLauncher's own {@code env("SHELL")} is deliberately set to
* {@code /bin/zsh} — the OPPOSITE of what {@code memberLoginShell} says — so a launcher that
* (incorrectly) fell back to fleetd's own {@code $SHELL} here would wrongly pass the gate and
* fail this test.
* fleetd #213 defect 1, acceptance criterion 1 — updated by fleetd #155: {@code
* memberHerdrSocket} configured and {@code memberLoginShell} configured as non-zsh must now
* REFUSE the spawn (not fall back to the sentinel overlay — see {@code
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}'s javadoc for why). The WiringLauncher's own
* {@code env("SHELL")} is deliberately set to {@code /bin/zsh} — the OPPOSITE of what {@code
* memberLoginShell} says — so a launcher that (incorrectly) fell back to fleetd's own {@code
* $SHELL} here would wrongly pass the gate and fail this test.
*/
@Test
void memberHerdrSocketWithNonZshMemberLoginShellFallsBackToTheOverlay() {
void memberHerdrSocketWithNonZshMemberLoginShellRefusesTheSpawn() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")),
"/bin/zsh", null,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/bash"));
List<String> messages = spawnAndCaptureLogs(launcher);
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"a non-zsh configured memberLoginShell under policy=allow-list must refuse the spawn");
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"a non-zsh memberLoginShell must fall back to the CB-596 sentinel overlay, exactly "
+ "like a non-zsh $SHELL does when memberHerdrSocket is absent");
assertTrue(e.getMessage().contains("/bin/bash"),
"the refusal must name the configured memberLoginShell, got: " + e.getMessage());
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
"a refused spawn must not have touched the pane env at all: " + launcher.env);
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub directory may be generated when the configured member login shell is not zsh");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"the 'generated ZDOTDIR' INFO must not appear when the scrub never runs — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 2: {@code memberHerdrSocket} configured and NO
* {@code memberLoginShell} configured must fall back exactly like criterion 1 above — AND
* fleetd's own {@code $SHELL} must never even be consulted (not merely "not decisive"). The
* fixture's {@code env} function reports {@code /bin/zsh} for {@code SHELL} — a value that would
* WRONGLY pass the zsh gate if the fix regressed to reading it — while flagging whether it was
* ever asked for at all, so this test fails loudly on either kind of regression.
* fleetd #213 defect 1, acceptance criterion 2 — updated by fleetd #155: {@code
* memberHerdrSocket} configured and NO {@code memberLoginShell} configured must now REFUSE the
* spawn exactly like criterion 1 above — AND fleetd's own {@code $SHELL} must never even be
* consulted (not merely "not decisive"). The fixture's {@code env} function reports {@code
* /bin/zsh} for {@code SHELL} — a value that would WRONGLY pass the zsh gate if the fix
* regressed to reading it — while flagging whether it was ever asked for at all, so this test
* fails loudly on either kind of regression.
*/
@Test
void memberHerdrSocketWithNoMemberLoginShellFallsBackAndNeverConsultsFleetdsOwnShell() {
void memberHerdrSocketWithNoMemberLoginShellRefusesAndNeverConsultsFleetdsOwnShell() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
@@ -437,15 +461,18 @@ class HerdrPeerLauncherAllowListWiringTest {
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")), env,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", null));
List<String> messages = spawnAndCaptureLogs(launcher);
assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"no configured memberLoginShell under policy=allow-list must refuse the spawn, same "
+ "as an explicit non-zsh one");
assertFalse(shellQueried.get(), "fleetd's own $SHELL must never be consulted once "
+ "memberHerdrSocket is configured — only memberLoginShell: may decide the gate");
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"no memberLoginShell configured must fall back to the sentinel overlay, same as a "
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
"a refused spawn must not have touched the pane env at all, same as a "
+ "configured non-zsh shell");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"no scrub may run without a configured memberLoginShell — got: " + messages);
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub may run without a configured memberLoginShell");
}
/**
@@ -700,7 +727,8 @@ class HerdrPeerLauncherAllowListWiringTest {
/**
* fleetd #213: as above, plus {@code worktreeRoot:}/{@code worktreeGroup:} — both required for
* the ZDOTDIR scrub to run at all once {@code memberHerdrSocket} is configured; either missing
* falls back to the sentinel overlay, same as a non-zsh {@code memberLoginShell}.
* falls back to the sentinel overlay (fleetd #155 left this branch alone — it is not a shell
* problem, so it is not this ticket's refusal).
*/
private static FleetConfig configWithMemberHerdrSocketRootAndGroup(String memberHerdrSocket,
String memberLoginShell, String worktreeRoot, String worktreeGroup) {
@@ -0,0 +1,99 @@
package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #111 (CB-608): {@link MemberCredentialPolicyView} is the one place that turns a {@code
* memberCredentials:} policy into names-and-counts, so the startup log line and {@code
* GET /member-credentials} cannot drift apart. These tests pin: the counts always match the
* policy that produced them, an absent/empty policy is represented honestly (never as "nothing
* blocked"), and the view carries names only — no value ever flows through it, because it is
* built only from {@link FleetConfig.MemberCredentials}, which itself never holds a value.
*/
class MemberCredentialPolicyViewTest {
@Test
void nullPolicyIsAbsentNotClean() {
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(null);
assertFalse(view.present(), "a null policy must be reported as absent");
assertEquals(0, view.knownCount());
assertEquals(0, view.allowedCount());
assertEquals(0, view.blockedCount());
assertTrue(view.known().isEmpty());
assertTrue(view.allowed().isEmpty());
assertTrue(view.blocked().isEmpty());
}
@Test
void emptyKnownListIsAbsentEvenWithAPolicyModeSet() {
// A memberCredentials: block can be present in YAML with policy: set but known: empty —
// that must still read as "no policy configured", the same as a fully absent block,
// because zero known names means the daemon blocks nothing either way.
FleetConfig.MemberCredentials creds =
new FleetConfig.MemberCredentials("deny-by-default", List.of(), List.of());
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertFalse(view.present());
assertEquals(0, view.knownCount());
}
@Test
void countsMatchARealPolicyExactly() {
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
"deny-by-default",
List.of("AI_GATEWAY_TOKEN"),
List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN"));
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertTrue(view.present());
assertEquals("deny-by-default", view.policy());
assertEquals(List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN"), view.known());
assertEquals(List.of("AI_GATEWAY_TOKEN"), view.allowed());
assertEquals(3, view.knownCount());
assertEquals(1, view.allowedCount());
// known minus allowed — the two names actually shadowed on a spawn.
assertEquals(2, view.blockedCount());
assertTrue(view.blocked().containsAll(List.of("GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN")));
}
@Test
void presentPolicyThatBlocksNothingIsStillDistinctFromAbsent() {
// known == allow => blockedCount is 0, exactly like an absent policy's blockedCount — the
// two must still be told apart by `present`, or a reader cannot tell "policy configured,
// nothing currently blocked" from "no policy at all".
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
"deny-by-default", List.of("X"), List.of("X"));
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertTrue(view.present());
assertEquals(1, view.knownCount());
assertEquals(0, view.blockedCount());
assertFalse(MemberCredentialPolicyView.absent().present());
}
@Test
void namesPassThroughUnchangedNeverAValue() {
// The view is built only from FleetConfig.MemberCredentials, which itself carries names,
// never values (see its javadoc) — so there is no code path here that could substitute a
// secret's value for its name. This pins the identity: what goes into `known`/`allow` is
// exactly what comes out, character for character.
List<String> known = List.of("SOME_TOKEN_NAME", "ANOTHER_NAME");
FleetConfig.MemberCredentials creds =
new FleetConfig.MemberCredentials("deny-by-default", List.of(), known);
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertEquals(known, view.known());
}
}
@@ -0,0 +1,147 @@
package dev.ltms.fleet.rest;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashSet;
import java.util.Locale;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #252: the REST surface has no supported operator entry point of its own — it is the
* documented fallback for when the MCP mount drops (see the operator wiki's REST page), and
* nothing has ever checked a hand-written route list against it. Proof it drifts: the ticket
* itself was filed with 14 routes, and {@link FleetApp} registers 15 — {@code GET
* /member-credentials} (fleetd #111) shipped hours before the ticket and was already missing from
* its list.
*
* <p>This test is modelled on {@code McpContractDocTest} (fleetd #114 / CB-609), which solved the
* same shape of problem for the MCP tool catalogue: read the source text with a regex instead of
* trusting a maintained copy. Here the "doc" is an explicit inventory written directly in this
* test rather than a separate Markdown file, because the operator wiki page lives in a submodule
* a worker cannot read reliably (see {@code CLAUDE.md} → Project addendum). Keeping the expected
* list in the test means it still fails loudly the moment {@link FleetApp} changes, which is the
* property that actually matters; a human keeps the wiki page in sync using the failure message
* below as the diff.
*
* <p><b>It checks source text, not behaviour.</b> It reads {@link FleetApp}'s source for {@code
* app.<verb>("<path>")} registrations and does not boot a server. It cannot catch a route that is
* registered through some other mechanism entirely (a filter, a redirect) — only ones shaped like
* the {@code app.get/post/delete/put/patch(...)} calls every route here actually uses.
*/
class RestRouteInventoryTest {
/** Tests run with the module directory as cwd. */
private static final Path REST_SOURCE = Path.of("src/main/java/dev/ltms/fleet/rest/FleetApp.java");
/**
* The REST surface as verified against {@link FleetApp} on 2026-09-03 (fleetd #252). Update
* this list AND the operator wiki's REST-surface entry together whenever a route is added,
* removed, or renamed — never one without the other.
*
* <p>{@code /mcp} is deliberately excluded: it is a raw Jetty {@code ServletHolder} mount
* (see {@code FleetApp.build()}, around the {@code modifyServletContextHandler} call), not an
* {@code app.<verb>(...)} route, so it is a different registration mechanism and this test's
* regex does not — and should not — see it. If {@code /mcp} ever moves to a Javalin route,
* add it here explicitly rather than relying on the regex to pick it up by accident.
*/
private static final Set<String> EXPECTED_ROUTES = Set.of(
"GET /healthz",
"GET /metrics",
"GET /sessions",
"GET /agents",
"GET /members",
"GET /profiles",
"GET /member-credentials",
"POST /members",
"DELETE /members/{paneId}",
"POST /sessions/{id}/message",
"POST /sessions/{id}/reply",
"GET /sessions/{id}/replies",
"POST /sessions/{id}/ask",
"GET /sessions/{id}/status",
"GET /tasks/{ticket}"
);
private static final Pattern ROUTE_CALL =
Pattern.compile("app\\.(get|post|delete|put|patch)\\(\\s*\"([^\"]+)\"");
/**
* Drop whole-line comments before scraping. Without this the scrape reads commented-out code as
* live: a registration disabled with {@code //} still matched, so the route stayed in the
* inventory while the server no longer served it — a silent false PASS, measured on 2026-09-03
* by commenting out {@code app.get("/tasks/{ticket}", ...)} and watching this test stay green.
* Deleting the same line was caught correctly, so only the commented-out shape was blind.
*
* <p>Only lines whose first non-blank characters are {@code //}, {@code *} or {@code /*} are
* dropped — deliberately NOT every {@code //} anywhere on a line, because that would also cut a
* string literal containing {@code //} (a URL) and could silently delete a real registration
* sharing that line. The remaining gap is a trailing comment on the same line as real code; no
* registration in this file has that shape, and the vacuity test below would catch a scrape that
* lost registrations wholesale.
*/
private static String withoutCommentLines(String source) {
return source.lines()
.filter(line -> {
String t = line.stripLeading();
return !(t.startsWith("//") || t.startsWith("*") || t.startsWith("/*"));
})
.collect(Collectors.joining("\n"));
}
/**
* Every {@code app.<verb>("<path>")} call in {@link FleetApp}'s source, as {@code "VERB path"}.
* This matches inside an {@code if (...) { ... }} block just as well as a top-level statement —
* the regex only looks for the call shape, not its surrounding control flow — which is what
* catches {@code GET /metrics} (registered conditionally on {@code metrics != null}).
*/
private static Set<String> routesTheServerRegisters() throws Exception {
Matcher m = ROUTE_CALL.matcher(withoutCommentLines(Files.readString(REST_SOURCE)));
Set<String> found = new LinkedHashSet<>();
while (m.find()) {
found.add(m.group(1).toUpperCase(Locale.ROOT) + " " + m.group(2));
}
return found;
}
@Test
@DisplayName("[SOURCE TEXT] FleetApp registers exactly the documented REST route inventory")
void theRegisteredRoutesMatchTheExpectedInventory() throws Exception {
Set<String> actual = routesTheServerRegisters();
Set<String> added = new LinkedHashSet<>(actual);
added.removeAll(EXPECTED_ROUTES);
Set<String> removed = new LinkedHashSet<>(EXPECTED_ROUTES);
removed.removeAll(actual);
assertTrue(added.isEmpty() && removed.isEmpty(),
"FleetApp's registered REST routes no longer match this test's expected inventory. "
+ "Added (in FleetApp, not in this test): " + added + ". "
+ "Removed (in this test, not in FleetApp): " + removed + ". "
+ "Update EXPECTED_ROUTES in RestRouteInventoryTest AND the operator wiki's "
+ "REST-surface entry together — this is the fleetd #252 defect: the route "
+ "list drifted for a month with nothing checking it. Do NOT weaken this test.");
}
/**
* The denominator guard (same shape as {@code McpContractDocTest}'s vacuity check). Pins that
* the regex really is still finding registrations, so a scrape that silently stops matching
* can't make the check above pass by finding nothing on both sides.
*/
@Test
@DisplayName("[SOURCE TEXT] the route scrape is not vacuous — it found the expected count")
void theScrapeActuallyFoundRoutes() throws Exception {
Set<String> actual = routesTheServerRegisters();
assertTrue(actual.size() >= EXPECTED_ROUTES.size(),
"scraped only " + actual.size() + " route registration(s) from FleetApp (" + actual
+ "), but this test expects at least " + EXPECTED_ROUTES.size()
+ "; the app.<verb>(\"...\") scrape has stopped matching and the check above "
+ "is now vacuous");
}
}
@@ -45,6 +45,9 @@ public final class FakeWorktrees implements Worktrees {
private final List<String> overlayShareOrder = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
/** Worktree paths that currently exist, mirroring GitWorktrees' {@code Files.exists} check for
* the already-gone case (CB-576 review, fleetd #116). */
private final Set<String> worktreePaths = ConcurrentHashMap.newKeySet();
private final AtomicLong snapshotSeq = new AtomicLong();
private volatile RuntimeException addFailure;
private volatile RuntimeException snapshotFailure;
@@ -115,7 +118,15 @@ public final class FakeWorktrees implements Worktrees {
}
// The branch already carries a unique nonce, so the derived path is distinct per acquire
// without an extra counter — keep it a pure function of the branch the test can predict.
return prefix + "/" + branch.replace('/', '_');
String path = prefix + "/" + branch.replace('/', '_');
worktreePaths.add(path);
return path;
}
/** Model an operator / {@code git worktree prune} removing the worktree before release. */
public FakeWorktrees markGone(String worktreePath) {
worktreePaths.remove(worktreePath);
return this;
}
@Override
@@ -125,6 +136,11 @@ public final class FakeWorktrees implements Worktrees {
@Override
public boolean hasUncommitted(String worktreePath) {
// A path that was never added, or was marked gone, is reported clean, mirroring
// GitWorktrees' already-gone guard — never an error, so teardown still completes.
if (!worktreePaths.contains(worktreePath)) {
return false;
}
return dirty;
}
@@ -258,6 +258,38 @@ class WorktreeSessionManagerTest {
}
}
/**
* CB-576 review (fleetd #116). A worktree that is already gone (operator cleanup,
* {@code git worktree prune}, an earlier half-completed release) must not break teardown.
* {@code hasUncommitted} reports the missing path clean (mirroring {@code GitWorktrees}), so
* {@code release} still runs {@code notifyReleased} (the CB-516 fast-fail for a blocked
* {@code fleet_send} caller) and {@code launcher.stop} (so the pane is not orphaned), and falls
* through to the already-gone-tolerant {@code remove}. Since CB-581 this all happens because the
* notify-and-stop work sits in {@code release}'s {@code finally}/post-try block rather than a
* checked branch — this test pins that shape by construction.
*/
@Test
void releaseStillStopsPaneAndNotifiesWhenWorktreeIsGone() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-576g", null));
worktrees.markGone(s.worktree());
sessions.release(s.paneId());
assertEquals(1, released.size(),
"notifyReleased must still fire when the worktree is already gone (CB-516)");
assertEquals(s.terminalId(), released.getFirst().terminalId());
assertTrue(herdr.called("pane.close"),
"the pane must still be stopped when the worktree is already gone");
assertEquals(1, worktrees.removeCalls().size(),
"release still calls the already-gone-tolerant remove");
}
@Test
void drainAllPreservesWorktreeOfIdleSession() {
FakeHerdr herdr = new FakeHerdr();
+142 -24
View File
@@ -22,8 +22,34 @@
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
# hash answers every question the prefix was for — is it set, is it the same value as over there,
# is it the CB-592 sentinel — and answers none of the ones it should not.
# * It never writes anywhere, never contacts the network, and never touches secrets.sh, which is
# the operator's file.
# * It never writes anywhere, never contacts the network except the daemon's own REST port (see
# below), and never touches secrets.sh, which is the operator's file.
#
# WHERE THE NAME LIST COMES FROM (fleetd #111 / CB-608)
#
# Earlier versions of this script carried their own hardcoded NAMES array, recorded by hand on
# 2026-08-16. The live memberCredentials: policy in fleetd.yaml grew past that list, and this probe
# never noticed — it kept checking the same 31 names, printed a clean-looking table, and exited 0.
# A verification tool that silently under-reports the thing it verifies is worse than no tool at
# all, because its "clean" output gets taken as proof rather than treated with the suspicion an
# absent tool would get.
#
# The fix is the same one #114 used for the drifted tool catalogue: delete the hand-maintained copy
# rather than update it. This script now fetches the policy's name list from the daemon itself, at
# `GET /member-credentials` (dev.ltms.fleet.member.MemberCredentialPolicyView via FleetApp) — names
# and counts only, the same way the daemon's own startup log line is computed, from the SAME class.
# If fleetd adds a name to memberCredentials.known tomorrow, this probe checks it tomorrow too,
# with no edit here required. There is no local fallback list. See fetch_policy() below for what
# happens when the daemon cannot be reached — it is a hard failure, on purpose (see next section).
#
# WHY AN UNREACHABLE DAEMON IS A HARD FAILURE, NOT A DEGRADED RUN
#
# An empty (or short) name list passes every subset check trivially — a probe that checked zero
# names would print "0 of 0 names are set" and look identical to a clean bill of health. That trap
# has bitten this project twice in one week (see docs/memory — "silent defaults disable features"
# and "a test on the seam does not prove the caller"). So the denominator is guarded explicitly:
# this script refuses to proceed unless it got a policy with at least one known name, and it refuses
# just as hard if the count it fetched does not match the count it is about to check.
#
# HOW TO RUN IT
#
@@ -32,34 +58,22 @@
# 2. For the comparison row, in your OWN shell — a lead, not a member:
# bash scripts/probe-member-credentials.sh --allow-outside-member
#
# Both readings need the daemon's REST port reachable (default http://127.0.0.1:8765; override with
# FLEETD_HOST). That is normally true in every pane this script is meant to run in.
#
# The two outputs side by side are the finding: any name whose hash matches between them is a
# credential the member holds in full.
#
set -uo pipefail
# The names ${SHARED_ENV}/tools/secrets.sh exports, recorded on 2026-08-16 (issue #82). Names only —
# this list contains no values and never should. If secrets.sh gains a name, this list goes stale and
# the probe silently stops asking about it; that staleness is itself part of what #82's criterion 4
# has to solve, so it is called out in the summary rather than hidden.
NAMES=(
AI_GATEWAY_TOKEN BESZEL_ADMIN_EMAIL BESZEL_ADMIN_PASSWORD
BESZEL_HUB_URL BESZEL_KEY BESZEL_UNIVERSAL_TOKEN
BRAIN_MCP_TOKEN CF_ACCOUNT_ID CF_API_TOKEN
CF_USER_TOKEN CONFLUENCE_API_TOKEN CONFLUENCE_USERNAME
CONTEXT7_TOKEN GITEA_HOST GITLAB_OAUTH_CLIENT_SECRET
GITLAB_PERSONAL_ACCESS_TOKEN GRAFANA_ADMIN_PASSWORD GRAFANA_ADMIN_USER
HASS_TOKEN HW_PASSWORD HW_USER
LTMS_API_KEY MEMORY_MCP_TOKEN METRICS_PUSH_TOKEN
OPENCODE_AUTOMODE_MODEL TELEGRAM_BOT_TOKEN TELEGRAM_CHAT_ID
TS_API_KEY TS_AUTHKEY WORKER_GITEA_TOKEN
GITEA_ACCESS_TOKEN
)
FLEETD_HOST="${FLEETD_HOST:-http://127.0.0.1:8765}"
POLICY_URL="${FLEETD_HOST%/}/member-credentials"
allow_outside=0
for arg in "$@"; do
case "$arg" in
--allow-outside-member) allow_outside=1 ;;
-h|--help) sed -n '2,40p' "$0"; exit 0 ;;
-h|--help) sed -n '2,60p' "$0"; exit 0 ;;
*) echo "unknown argument: $arg" >&2; exit 2 ;;
esac
done
@@ -75,6 +89,107 @@ EOF
exit 1
fi
# --- fetch the policy from the daemon (fleetd #111) — no local fallback, ever ------------------
#
# Prefer jq (a real JSON parser); fall back to python3 (present on every host this has run on so
# far); if neither exists, fail loudly rather than guess at the JSON with grep/sed, which is exactly
# the kind of "looks like it worked" degradation this ticket exists to remove.
#
# NOTE: jq's `//` alternative operator treats `false` AND `0` as "missing" and substitutes the
# default — so `.present // empty` silently turns a real `"present": false` into an empty string
# ("unknown"), not the false it actually is. Every extraction below reads its field directly
# instead, so a genuine false/0 is reported as exactly that, not swallowed into "unknown".
if ! command -v jq >/dev/null 2>&1 && ! command -v python3 >/dev/null 2>&1; then
echo "refusing to run: neither jq nor python3 is on PATH, and this probe will not guess at JSON" \
"with grep/sed. Install one of them, or run from a shell that has one." >&2
exit 1
fi
POLICY_JSON="$(curl -fsS --max-time 5 "$POLICY_URL" 2>/dev/null)"
CURL_STATUS=$?
if [ "$CURL_STATUS" -ne 0 ] || [ -z "$POLICY_JSON" ]; then
cat >&2 <<EOF
refusing to run: could not fetch the memberCredentials policy from $POLICY_URL (curl exit $CURL_STATUS).
This probe has NO built-in name list any more (fleetd #111) — it only checks what the live daemon
reports, so an unreachable daemon means it cannot check anything at all. It will not fall back to a
guessed or empty list, because an empty list would pass every check trivially and look clean.
Fix: confirm fleetd is up (curl \${FLEETD_HOST:-http://127.0.0.1:8765}/healthz) and that
FLEETD_HOST (if set) points at it, then re-run.
EOF
exit 1
fi
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
if command -v jq >/dev/null 2>&1; then
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | jq -r '
(.present | tostring),
(.policy // ""),
(.knownCount // 0 | tostring),
(.allowedCount // 0 | tostring),
(.blockedCount // 0 | tostring),
(.known[]? // empty)')
else
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
import json, sys
data = json.load(sys.stdin)
print(str(data.get("present")))
print(data.get("policy") or "")
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
for n in (data.get("known") or []):
print(n)
PY
)
fi
PRESENT="${_FIELDS[0]:-null}"
POLICY_MODE="${_FIELDS[1]:-}"
KNOWN_COUNT_REPORTED="${_FIELDS[2]:-0}"
ALLOWED_COUNT_REPORTED="${_FIELDS[3]:-0}"
BLOCKED_COUNT_REPORTED="${_FIELDS[4]:-0}"
NAMES=("${_FIELDS[@]:5}")
# knownCount must be a plain non-negative integer for the arithmetic guard below — a malformed or
# unparseable response must fail loudly, not be coerced into a number that happens to compare true.
case "$KNOWN_COUNT_REPORTED" in
''|*[!0-9]*)
echo "refusing to run: knownCount in the response ('$KNOWN_COUNT_REPORTED') is not a plain" \
"non-negative integer — the response could not be parsed as expected." >&2
exit 1
;;
esac
# --- guard the denominator explicitly — never proceed on a zero/short count ---------------------
#
# This is the exact trap named in the ticket: an empty (or truncated) NAMES array passes every
# subsequent "is it set" check vacuously and prints a table that LOOKS complete. So this is checked
# before anything else runs, with a message that says why, not just that it failed.
if [ "${#NAMES[@]}" -eq 0 ] || [ "$KNOWN_COUNT_REPORTED" -eq 0 ]; then
cat >&2 <<EOF
refusing to run: the policy fetched from $POLICY_URL contains 0 known names (present=${PRESENT:-unknown}).
Either memberCredentials: is absent/empty on the running daemon (nothing is protected — see fleetd's
own startup warning), or the response could not be parsed. Either way, checking zero names would
print a clean-looking table for a policy that protects nothing, or for a probe that read nothing.
This is refused rather than reported as a pass.
EOF
exit 1
fi
if [ "${#NAMES[@]}" -ne "$KNOWN_COUNT_REPORTED" ]; then
cat >&2 <<EOF
refusing to run: the policy reports knownCount=$KNOWN_COUNT_REPORTED but the known[] array this probe
parsed has ${#NAMES[@]} entries. That mismatch means the JSON was not parsed correctly, and this
probe will not check a name list it cannot trust to be complete.
EOF
exit 1
fi
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
# length only — degraded, but never a value.
hasher=""
@@ -95,12 +210,13 @@ else
where="NOT a member — comparison reading only"
fi
echo "CB-596 credential probe"
echo "CB-596 credential probe (fleetd #111: names sourced live from $POLICY_URL)"
echo "reading from : $where"
echo "shell : ${SHELL:-unknown}"
echo "hash : ${hasher:-none available — lengths only}"
# Only printed so the two readings can be told apart when they are pasted side by side.
echo "host : $(hostname 2>/dev/null || echo unknown)"
echo "policy : mode=${POLICY_MODE:-unknown} known=$KNOWN_COUNT_REPORTED allowed=${ALLOWED_COUNT_REPORTED:-?} blocked=${BLOCKED_COUNT_REPORTED:-?}"
echo
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
@@ -118,6 +234,7 @@ done
echo
echo "$set_count of ${#NAMES[@]} names are set in this shell."
echo "policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match."
echo
cat <<'EOF'
How to read this:
@@ -129,7 +246,8 @@ How to read this:
most urgent thing on this page.
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: fleetd.yaml names it in
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
* A name that is set here but is NOT in the list above will not appear at all. The list was
recorded on 2026-08-16 and does not update itself. Anything added to secrets.sh since then is
invisible to this probe — which is the same gap issue #82 criterion 4 asks to close properly.
* The name list above is fetched live from the running daemon's memberCredentials: policy
(fleetd #111) — it is never hand-maintained here, so it cannot go stale the way the old
hardcoded list did. If the daemon's policy changes, the next run of this script reflects it
with no edit to this file.
EOF