Compare commits

...

45 Commits

Author SHA1 Message Date
Dai Ha 7840e9adf6 CB-201: make unmapped target notice truthful
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m15s
2026-09-03 10:44:18 +07:00
Dai Ha ee932fd85b CB-201: nudge leads about backend outages
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Successful in 1m55s
2026-09-03 10:28:51 +07:00
ltms 0f08b93659 fleetd #176: spawn gate fails fast on a dead backend, and refines UNKNOWN safely
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m47s
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.

Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.

Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".

Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.

Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
2026-09-03 04:20:23 +02:00
Dai Ha 432c1d92d1 fleetd #176 review: cover the namePrefix guard, the adapter-refinement corroboration that had no test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m14s
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.

Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.

mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:17:10 +07:00
Dai Ha 60b7e67b42 fleetd #176: fail fast when the spawn gate's backend exits mid-wait, and safely refine its UNKNOWN status
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m26s
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.

Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.

Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
 - spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
 - spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
 - refinedIdleIsNotAcceptedWhenAgentTypeIsNull
 - refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
 - refinementNeverFiresWhenRawStatusIsAlreadyInjectable

mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:10:49 +07:00
ltms 3fbd43fe3f fleetd #175: read back the model opencode actually resolved, and quarantine a silent substitution
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m29s
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
2026-09-02 13:01:33 +02:00
Dai Ha 32ebf065ac fleetd #175 review: a missing providerID in opencode's model JSON is UNKNOWN, not a mismatch
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m13s
parseModel already tolerates a model JSON with an id but no providerID (a real shape
opencode can write). checkModelMatch's providerMatches check did not: a provider-
prefixed profile whose id matched but whose evidence had no providerID was reported
as a mismatch and quarantined on incomplete data, which acceptance rule 4 forbids.

Compare the provider only when BOTH the profile requested one AND the evidence has
one. A genuine id mismatch is still caught either way — narrows the check, does not
disable it.
2026-09-02 17:59:04 +07:00
Dai Ha 1178b3f684 fleetd #175: check opencode's actual model against the profile, quarantine on a real mismatch
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m49s
opencode does not fail on an unknown -m <model> flag — it silently falls back to a
default model, which can be a paid credential. Extends the existing late-resolve
path (#209's SessionManager -> handle.agentSessionId() re-poll) so that once the
opencode session row exists, OpenCodeSessionDiscovery also reads its `model` JSON
column and OpenCodeLauncher's SessionAwareHandle compares it against the profile's
configured model.

Comparison rule: split the profile's model on the first '/' into provider+id. Compare
id always; compare provider only when the profile specified one. A bare model name
with no '/' matches on id alone. Absent/unparseable evidence is UNKNOWN, never a
mismatch, so a working profile is never quarantined on missing data. A real mismatch
logs an ERROR naming both models and the profile, then quarantines through the
existing ExhaustionSink path (wired via an AtomicReference forwarding sink in
Fleetd.java to break the sessions/workers/adapters construction cycle).
2026-09-02 17:53:59 +07:00
ltms 2a434ced2f fleetd #222: keep the member's charter file out of fleetd's own java.io.tmpdir
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m28s
Verified by the lead before merge: read the full production diff, confirmed the memberHerdrSocket-absent branch is the literal unmodified Files.createTempFile call in its own branch, and that the refusal names the missing key. Measured behaviour recorded: claude 2.1.258 exits 1 immediately on an unreadable --append-system-prompt-file, so the pre-fix bug was the loud readiness-gate failure, not a silent charter-less member. The /tmp full-suite failure the worker reported was the two other #225 copies, fixed by #230 which is already on main; main is built and checked after this merge.
2026-09-02 02:56:36 +02:00
ltms cc919aa2b6 fleetd #226: reserve an architect slot before launch, so no member runs on a charter it will not hold
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m31s
Verified by the lead before merge. Read the full production diff: locking is consistent (every MemberRegistry method uses synchronized(terminalToSlot)), and the refusal happens before the launcher starts a process. Proved the restored fallback test is real by removing the DEV fallback from SessionManager and re-running it: it failed with "expected: <DEV> but was: <ARCHITECT>", then reverted. Independent build: BUILD SUCCESS, 1092 tests, 0 failures, 0 skipped.
2026-09-02 02:54:52 +02:00
Dai Ha 6417b0edd9 fleetd#222: fix currentUserGroup() to resolve the real primary group (fleetd#225)
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m43s
The helper I copied from OpenCodeLauncherTest read the CWD's owning
group instead of the process's real primary group, so it silently
picked up whatever group owns the directory Maven was started from
(staff in a home checkout, wheel under /private/tmp on macOS) rather
than a group the operator is actually in. Replaced with the id -gn
based resolution that landed on #230 for the other two copies of this
helper, same shape and skip wording.
2026-09-02 07:54:11 +07:00
ltms 39c7ce76f3 fleetd #224 + #225: make worktreeRoot group-traversable, and stop two tests depending on the checkout location
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m41s
Verified by the lead before merge: read the full production diff, confirmed `group` is normalised to null at GitWorktrees:148 so the `group == null` guard is complete, and confirmed the refusal runs before `git worktree add` so a failure leaves no half-made worktree. Independent build in the worker's worktree: BUILD SUCCESS, 1093 tests, 0 failures, 0 skipped.
2026-09-02 02:52:12 +02:00
Dai Ha b40f477210 #226 retain architect fallback coverage
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m33s
2026-09-02 07:50:34 +07:00
Dai Ha 5c56cb347f fleetd #224 / #225: share worktreeRoot with the group, and fix the group-detection test helper
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m42s
#224: GitWorktrees#add created worktreeRoot with the daemon's umask and never shared it with
worktreeGroup, even though shareWithGroup shares every child underneath it (each worktree, and
the repo's common git dir). Under memberHerdrSocket: the member pane runs as a different OS
user, which needs execute on every ancestor directory to reach anything underneath, no matter
how carefully each child is shared — so a member could not read the opencode.json #219 places
under this root, could not reach its own worktree, and could not read #213's ZDOTDIR scrub when
placed here either.

Fix: add() now calls a new shareRootWithGroup(root) right after creating the root, chgrp+chmod
g+x on the root itself (non-recursive — each child is still shared individually by its own call
site). No-op when worktreeGroup is unset, so behaviour is byte-identical in today's only live
mode. On failure (group missing, or operator not a member of it) the spawn is refused with a
WorktreeException naming the root, its current mode, and the group — mirroring shareWithGroup's
existing refusal shape — before `git worktree add` ever runs, so no partial worktree is left
behind.

Also adds the assertion the #221 reviewer flagged as missing: a test driving
EnvAllowListScrub#shareWithGroup directly against a directory holding several flat files
(opencode.json, member-charter.md, ide-rules.md, plus an unrelated one) and asserting every one
of them gets group-readable/never-group-writable permissions, not just the two files someone
happened to think of.

#225: OpenCodeLauncherTest/HerdrPeerLauncherAllowListWiringTest's currentUserGroup() read the
group that owns the current working directory, not the process's own primary group, despite its
comment claiming the latter. Those coincide only by accident: a home checkout is typically owned
by a group the operator belongs to (staff), while a checkout under /private/tmp on macOS is
group wheel, which the operator is usually not a member of — so the same test fails for real
depending on where the repo happens to be checked out, and the existing assumeTrue only guarded
against "no POSIX groups at all", never "a resolvable but wrong group". Fixed by resolving the
process's REAL primary group via `id -gn` instead, with assumeTrue (skip, not fail) only when
that itself cannot be resolved on the host. The permission assertions these tests exist for are
unchanged.

Verified `mvn clean install` green from both a home checkout and a /private/tmp copy (mirroring
the exact repro in #225): 1093 tests, 0 failures, 0 errors in both locations.
2026-09-02 07:47:13 +07:00
Dai Ha e694deace3 #226 reserve architect slots before launch
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Failing after 1m33s
2026-09-02 07:44:41 +07:00
Dai Ha 748367b7d6 fleetd#222: put the claude-code charter file where the member can read it
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m16s
ClaudeCodeLauncher#writeCharterFile used Files.createTempFile with no
directory argument, which resolves against fleetd's own java.io.tmpdir
(macOS: the per-user $TMPDIR, mode 0700). Under memberHerdrSocket: the
member pane runs as a different OS user and cannot read that directory,
and since #220 the charter file is the ONLY delivery path for
--append-system-prompt-file. Following #219's refusal decision (a
charter is the member's turn contract, not a degradable control): with
memberHerdrSocket configured, the charter now goes into a fresh
per-spawn directory under worktreeRoot, shared read-only via
EnvAllowListScrub.shareWithGroup (reusing #213/#219's mechanism); a
missing worktreeRoot/worktreeGroup refuses the spawn by name instead of
writing an unreadable file. With memberHerdrSocket absent the path is
unchanged.
2026-09-02 07:44:34 +07:00
ltms dcf5fb3be3 Merge pull request 'CB-619 / fleetd #123: refuse an architect spawn with no matching slot' (#223) from worker/cb-123-role-demotion-c600f7-2 into main
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m55s
2026-09-01 10:41:47 +02:00
ltms f9d2ee2a2b Merge pull request 'fleetd#219: fix OpenCodeLauncher config/discovery roots under memberHerdrSocket' (#221) from worker/cb-219-opencode-roots-1f677e-1 into main
CI / contract (push) Successful in 45s
CI / build (push) Successful in 2m16s
2026-09-01 10:33:04 +02:00
Dai Ha 866c7f2e9a CB-619 / fleetd #123: refuse an architect spawn with no matching slot
CI / build (pull_request) Successful in 1m16s
CI / contract (pull_request) Successful in 1m18s
An explicit-profile spawn bypasses role-pool placement (CompositePeerLauncher
only constrains an UNQUALIFIED spawn to fleet.<role>), so it was the one path
that could ask for role=architect on a profile no architect slot carries.
MemberRegistry silently held the session as a plain worker while GET /members
still reported the requested "architect" and only fleet_whoami (which reads
live bindings, not the request) told the truth.

- MemberLifecycle.requireSlotFor(role, profile): refuses the acquire before
  anything spawns when no configured architect slot carries the profile,
  naming the role, the profile, and the pools that do carry it. No-op for
  dev/reviewer, which are placement candidates only, never a live identity
  binding — refusing a profile mismatch there would break the documented
  fleet_spawn{profile:"opus"} (role defaults to dev) flow.
- MemberLifecycle.acquired(...) now returns the role the session actually
  holds, so a residual race (a slot exists but every instance is already
  bound to a different terminal) still falls back to dev honestly instead of
  lying — this case logs at WARN (was INFO), naming profile and terminal.
- SessionManager now records the role acquired() returns on MemberSession,
  never the requested role, so GET /members and fleet_list can no longer
  report a role the member does not hold; no changes needed to memberView/
  rosterView, which just read session.role().

An architect's identity IS the slot it is bound to — binding a role with no
slot to bind means inventing an identity out of nothing, which is the quiet
failure the whole role system exists to prevent.

Tests: SessionManagerTest and FleetMcpTest each drive a real spawn through
FleetMcp.spawn -> SessionManager.acquire -> the real ClaudeCodeLauncher (via
FakeHerdr), then assert on GET /members and fleet_whoami for that same
session — not on MemberRegistry.bind directly (fleetd issue #113's mistake).
2026-09-01 15:29:39 +07:00
Dai Ha d9168de43e fleetd#219: OpenCodeLauncher config/discovery roots must not assume fleetd's own filesystem
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m48s
Site 1 (config root): under memberHerdrSocket, writeConfig() now places the
ephemeral opencode.json directory under worktreeRoot and shares it read-only
with worktreeGroup, reusing EnvAllowListScrub#shareWithGroup (widened to
package-private and generalized) — the same mechanism #213 built for the
ZDOTDIR scrub, rather than a second copy. Unlike the ZDOTDIR scrub's
degrade-to-overlay fallback, a missing worktreeRoot/worktreeGroup here
REFUSES the spawn (IllegalStateException from buildLaunch): this file is the
member's only way to learn where the bridge MCP is, so writing it somewhere
unreadable would just produce an undeliverable member with no signal
pointing at the cause. memberHerdrSocket absent stays byte-identical.

Site 2 (discovery root): under memberHerdrSocket, agentSessionId() now
declares session discovery unavailable and logs one WARN per launcher
instance instead of silently scanning fleetd's own $HOME (opencode.db lives
under the MEMBER's home under this config key). Decision + reasoning for why
this is a declare-unavailable rather than a new config key is in
defaultDiscoveryRoot()'s javadoc.

Widened HerdrPeerLauncher#memberHerdrSocketConfigured/memberScrubParentDir/
memberGroup to package-private so OpenCodeLauncher reuses the exact same
config resolution rather than re-deriving it.

Same-shape finding (not fixed, out of scope): ClaudeCodeLauncher#writeCharterFile
(line ~465) writes the role-charter temp file via Files.createTempFile with no
directory argument, i.e. under java.io.tmpdir — the same site-1 shape, unfixed
for the Claude Code adapter.
2026-09-01 14:38:48 +07:00
Dai Ha c3fa1136d4 #220: keep the launch command inside the pane's 1024-byte line
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m37s
herdr does not exec a member's launch command — it TYPES it into the pane,
and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Past that the
tail is dropped and NOTHING reports it: herdr answers "agent started", the
backend exits on the mangled argument it was handed, the pane closes, and the
only symptom is the readiness gate timing out 20 seconds later with no reason.

That is what broke every claude-code spawn after #214. The reply charter rode
inline on --append-system-prompt, so the command was already 978 bytes; adding
--session-id <uuid> made it 1028, and the 4 bytes cut off the end turned
--autocompact 250000 into --autocompact 25, which claude rejects. Measured on
the live pane, the cut is at byte 1024 exactly.

- ClaudeCodeLauncher: the charter ALWAYS travels as --append-system-prompt-file.
  The file path already existed for the two-charter case; the inline form only
  ever saved a temp file, and it cost ~800 bytes of the line budget. This takes
  the prose off the command line for good.
- HerdrPeerLauncher.checkPaneCommandFits: refuse a command that cannot fit,
  naming the byte count and the longest argument, instead of spawning something
  that cannot work. The estimate is deliberately conservative — fleetd cannot
  see herdr's quoting, and an under-estimate would let the silent truncation
  back in.
- HerdrPeerLauncher.waitUntilInjectableOrThrow: log the pane tail and the last
  herdr status BEFORE stop() closes the pane. Without it the gate reports only
  that it timed out, which is true of every cause. This is what found the bug,
  and it stays.

The guard also catches a case that was already over the limit: a profile with
ideMcpUrl set assembles 1084 bytes. It is now impossible to ship that silently.

3 tests, all watched failing first: with the inline charter restored the guard
fires in the new fit test, in the pre-existing autocompact test and in the IDE
mount test. Full suite 1081 tests green. Proven live: sonnet spawns again, the
member obeys the file-delivered charter and ends its turn with fleet_reply, and
fleet_list reports the #214 agentSessionId.
2026-09-01 14:11:28 +07:00
Dai Ha cabcd87b66 #211 follow-up: the raw-scrape fallback must carry the pane, not just the matched line
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m50s
The normal backend-error path appends the pane tail to the failure reason on
purpose (fleetd#164): the BACKEND_ERROR pattern is a heuristic, and a member
that reported *about* an error while forgetting fleet_reply matches it too, so
dropping the rest of the pane destroys the report.

The new raw-scrape fallback did not do that. It matters more there, not less:
the fallback only runs when the trimmed assistant block was empty, so the raw
scrape is the ONLY copy of whatever the member managed to say. A lead read the
matched line and nothing else.

Clipped to the same cap the normal path uses, since a raw screen has no
boundary trimming to bound its size.

The assertion was watched failing without the fix:
  AssertionFailedError: fleetd#164: the failure must carry the pane, not only
  the matched line ... expected: <true> but was: <false>

mvn clean install: Tests run: 1079, Failures: 0, Errors: 0, Skipped: 0
2026-09-01 13:26:05 +07:00
Dai Ha bf0e09b1a2 #214: mint a session id on every claude-code spawn, so every member is resumable 2026-09-01 13:23:32 +07:00
Dai Ha e1eb50ce65 #213: fix ZDOTDIR credential scrub gate + directory under memberHerdrSocket 2026-09-01 13:23:32 +07:00
Dai Ha 049e7d9d54 fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback 2026-09-01 13:23:32 +07:00
Dai Ha ff3b49cd1e #214: mint a session id on every claude-code spawn, so every member is resumable
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m52s
2026-09-01 11:27:10 +07:00
Dai Ha 0373b6c41b #213: fix ZDOTDIR credential scrub gate + directory under memberHerdrSocket
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m17s
The memberCredentials.policy: allow-list ZDOTDIR scrub decided zsh-vs-not
using fleetd's own process $SHELL and wrote the generated scrub dir into
fleetd's own java.io.tmpdir. Under memberHerdrSocket: (member panes run as
a different OS user than fleetd's own process) this silently protects
nothing: the wrong shell decides the gate, and the directory can be
unreachable to the member.

- New FleetConfig.memberLoginShell: the member OS user's login shell,
  only ever read when memberHerdrSocket: is configured; fleetd's own
  $SHELL keeps deciding everything when memberHerdrSocket: is absent
  (byte-identical to before).
- HerdrPeerLauncher.applyEnvironmentAllowListPolicy: memberHerdrSocket +
  memberLoginShell not configured/non-zsh falls back to the CB-596
  sentinel overlay (warn loudly, never refuse to spawn). memberHerdrSocket
  + zsh memberLoginShell generates the ZDOTDIR under the configured
  worktreeRoot instead of java.io.tmpdir, shared with the existing
  worktreeGroup (reused, not a new key).
- EnvAllowListScrub: new generate(parentDir, allowedNames, group) overload
  shares the generated directory via pure-Java POSIX group ownership
  (rwxr-x--- dir, rw-r----- files) — no external process spawn.
- fleetd.example.yaml documents memberLoginShell: and worktreeGroup:'s
  reuse for the scrub directory (the live fleetd.yaml is gitignored).

4 new tests in HerdrPeerLauncherAllowListWiringTest cover the acceptance
criteria; 3 of the 4 were watched failing against the pre-fix code.
2026-09-01 11:07:11 +07:00
Dai Ha 445a45f6e1 fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m17s
CompletionResolver.resolve() returned an empty-scrape failure before the
BACKEND_EXHAUSTED / BACKEND_ERROR classification ever ran, whenever
lastAssistantBlock() found no usable text — most commonly a pane with no ⏺
marker at all, whose boundary scan starts at the top of the raw screen and
breaks immediately on the first line of TUI chrome. Since BACKEND_EXHAUSTED
is the only caller of exhaustionSink, this meant an exhausted backend was
recorded as "produced nothing" instead of being quarantined.

Fix: run the same two classifications against the raw (untrimmed) scrape as
a fallback, only inside the empty-scrape failure branch. A pane that already
yields a usable assistant block never reaches this branch, so the existing
narrow match is unchanged. lastAssistantBlock stays the sole source of the
reply text; only classification ever consults the raw scrape.
2026-09-01 10:50:10 +07:00
Dai Ha 966c58a3b8 #137: don't hand one stranded reply to every open async ticket (#215)
CI / contract (push) Successful in 1m31s
CI / build (push) Successful in 1m41s
abandon() drained the target's stranded reply once and then reused that same
Reply for every open task it walked past. Two open tickets on one target
therefore both came back REPLIED with the same text — one of them a reply the
worker never gave for that delegation.

One worker answer can settle at most one delegation. It now goes to the oldest
open task (lowest createdNanos) and every other open task keeps the ordinary
WORKER_FAILED path. If the chosen task turns out to be already resolved by
another path, the drained reply is published back to the inbox instead of being
dropped.

reply()'s matching side gets the same rule: more than one candidate means the
reply goes to the inbox rather than to a guess.

Reachability, checked rather than assumed:
- matching.size() >= 2 alone is reachable and was a real bug before this change.
- More than one candidate in reply() is not reachable today — hasAsyncQuestion
  matches any task with a stamped turnId, and answer()'s
  clearAsyncQuestion(turnId, false) leaves that stamp until the resumed turn
  resolves. Kept as defensive code, documented, no test seam added.
- matching.size() >= 2 together with a live strand is not reachable either:
  send() and answer() are the only two lock holders and both open a Rendezvous
  waiter inside the lock, so "lock held" and "waiter open" are one fact, and an
  acceptance always clears the strand first.

I checked the last point by building it in a scratch worktree: it can be forced
by closing an accepted send's waiter directly through Rendezvous, and it does go
red against the pre-fix loop — but that breaks the lock-and-waiter invariant
from outside the class, so no such test is added. A note in abandon()'s javadoc
says so, to save the next reader the same round trip.

Worker's REPORT-cb137.md left out of main.

mvn clean install: Tests run: 1071, Failures: 0, Errors: 0, Skipped: 0
2026-08-31 23:23:07 +07:00
Dai Ha de026b8f8a #137 follow-up: document defect-1 and defect-2 reachability, drop unconstructible test
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m28s
Defect 1 (reply()'s askAnsweredAsyncTasks returning >1 candidate): confirmed
unreachable today. Documented why in three places — hasAsyncQuestion matches
any task with a stamped turnId (not just an open question), and answer()'s
clearAsyncQuestion(turnId, false) leaves that stamp in place until the
resumed turn's own future resolves — so a second task can never reach the
same eligible state while a first one holds it. Kept the defensive
inbox-fallback branch as defence in depth against that guarantee weakening,
per review instruction; no test seam added.

Defect 2 (abandon()'s broader `matching` filter applying a stranded reply to
more than one task): matching.size() >= 2 alone IS reachable (already
covered by abandonFailsEveryPendingAsyncTicketForTheReleasedTarget) and was
a real pre-fix bug (97f6c33's parent reused one drained reply for every
matching task). But hadStrandedReply == true together with matching.size()
>= 2, at the instant abandon() runs, is not constructible through the public
API: send() and answer() are the only two sites that ever hold a target's
session lock, and both open a Rendezvous waiter for that target as the first
thing they do while holding it — so "lock held" and "waiter open" are the
same fact throughout this class, and reply()'s fast path always resolves an
open waiter directly instead of stranding. A strand can only be created
while no task is accepted, and the moment the lock is next taken, that
acceptance clears the strand again before abandon() can observe both facts
together. Documented this in abandon()'s javadoc and removed the earlier
attempt at a deterministic test for the conjunction, whose apparent failure
was an invalid premise (the "accepted" task's own acceptance silently
cleared the strand it was meant to race against), not the fix being absent.
2026-08-31 23:17:50 +07:00
Dai Ha 97f6c33a45 #137: don't guess when a target has more than one open async task
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m17s
The CB-205-recovery fix in #205 assumed a target has at most one open
async task, with no guard. Fix three consequences:

- reply(): askAnsweredAsyncTask -> askAnsweredAsyncTasks (List). Exactly
  one candidate completes it (unchanged). Zero falls to the inbox
  (unchanged). More than one now ALSO falls to the inbox instead of
  picking an arbitrary ConcurrentHashMap iteration order, and logs a
  WARN naming the target and every candidate ticket.

- abandon(): a stranded reply now settles at most one matching task —
  the oldest by Task#createdNanos (a new field, the tiebreaker). Every
  other matching task keeps WORKER_FAILED, same as today.

- abandon(): if the chosen recovery task's complete() loses a race
  (another path resolved it first), the drained reply is republished
  to the inbox instead of being silently dropped.

Single-task behavior is unchanged; only the ambiguous case changes.
2026-08-31 22:46:23 +07:00
Dai Ha 735c837604 #209: resolve agentSessionId lazily via the retained PeerHandle
CI / contract (push) Successful in 48s
CI / build (push) Successful in 2m18s
SessionManager called handle.agentSessionId() once at spawn and froze it in the
immutable MemberSession. For opencode that value is always null - the session
row does not exist yet when the pane is created - so fleet_list never reported
an agentSessionId and fleet_spawn{resumeSessionId} was unusable for that
backend. The handle's javadoc said 'the caller re-calls later'; no caller did,
and SessionManager did not even retain the handle.

Retain the PeerHandle per pane and re-resolve while the stored id is still
null, sticky once found, CAS-swapped into the registry.

roster() stays non-resolving - it is the roster supplier for LeadHeartbeatLoop
and FleetHealthMonitor, and resolving there would open opencode's 841MB SQLite
database on every tick, for every unresolved member, forever. rosterResolved()
carries the resolve and is used only by fleet_list and the REST roster, the two
surfaces that report the id. get(paneId) and release() resolve too, both
caller-driven.

A test pins the split: the plain roster() must never call agentSessionId()
again.
2026-08-31 22:21:16 +07:00
Dai Ha 457dc0330d #185 stage 2: the credential gap detector must not report on the wrong environment
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m59s
Under memberHerdrSocket: member panes run as a different OS user, so
hostEnvNames (fleetd's own environment) no longer describes what a member
inherits. logCredentialGap now reports "unknown, not clean" in that mode
instead of its usual UNBLOCKED/blanked conclusions, once per launcher, naming
the config key and scoping its count to fleetd's own process.

Byte-identical when memberHerdrSocket: is absent, pinned by a test.
2026-08-31 22:17:40 +07:00
Dai Ha a1a9015217 #199: /members returns rows under "members" 2026-08-31 22:17:35 +07:00
Dai Ha 5a811a3695 #199: GET /members returns its rows under "members", with "workers" kept as an alias
The endpoint became /members in the CB-634 rename but the body key stayed
"workers", so a caller that read "members" saw an empty fleet and reported
no members at all.

Emit the canonical "members" key. Keep "workers" as a deprecated alias so an
existing REST consumer keeps working - the out-of-band path a lead falls back
to when its MCP mount drops reads this endpoint.

The test now pins both keys and asserts they carry the same rows, so the alias
cannot silently drift. Watched failing without the fix:
"GET /members must return its rows under \"members\" ==> expected: not <null>"
2026-08-31 22:17:22 +07:00
Dai Ha 1fdaa74eb3 fleetd #209 follow-up: keep roster() off the resolve path
CI / build (pull_request) Successful in 1m20s
CI / contract (pull_request) Successful in 1m21s
roster() is the supplier for LeadHeartbeatLoop and FleetHealthMonitor
(both timer-driven) and for placement/exhaustion checks and the
metrics scrape — none of which read agentSessionId. Resolving there
meant every tick could open a lazy-resolving adapter's (opencode's)
on-disk session database once per member whose id was still unknown,
with no bound: a member whose id never appears would pay that cost for
the life of the process.

roster() goes back to its pre-#209 behavior (no resolve, no I/O). A
new rosterResolved() carries the resolve logic, and is used only by
the two surfaces that actually report agentSessionId to a caller:
fleet_list (FleetMcp.listFleet) and the REST roster
(FleetApp.listMembers). fleet_whoami's roster().stream() at
FleetMcp.java:768 does not surface the field, so it stays on the plain
roster(). get(paneId) (fleet_status) and the release() resolve are
caller-driven, not timers, and are unchanged.

Retargeted the roster-facing tests from #209 at rosterResolved(), and
added plainRosterDoesNotResolveAgentSessionId, which pins the split by
asserting the handle's agentSessionId() is not called again by
roster().
2026-08-31 22:15:16 +07:00
Dai Ha 7d4a4339c2 Merge branch 'worker/cb-137-ask-ticket-e7760c-2' of https://git.ltms.dev/fleet/fleetd into worker/cb137-ambiguous-task-4df3d8-4 2026-08-31 22:13:37 +07:00
Dai Ha 847e8bd3fa #185 stage 2: stop the credential-gap detector reporting on the wrong environment
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 1m39s
When memberHerdrSocket: is configured, member panes run under a different OS
user than fleetd's own process, so hostEnvNames (fleetd's own environment)
no longer describes what a member pane inherits. logCredentialGap now checks
for that config key and, when set, logs a single WARN saying the gap is
UNKNOWN (not clean) and names the key, instead of printing the "inherits
them UNBLOCKED" / "the scrub blanks them" conclusions as fact. Behaviour is
byte-identical when memberHerdrSocket is absent (the default and only mode
this host runs).
2026-08-31 22:12:56 +07:00
Dai Ha f9fb387427 fleetd #209: re-poll agentSessionId against the retained PeerHandle
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m40s
SessionManager used to call PeerHandle.agentSessionId() exactly once at
spawn and freeze the answer into the immutable MemberSession. For
opencode that call always came back null, because opencode has not
written its on-disk session row yet when the pane is created, and no
caller ever re-asked the handle — it went out of scope at the end of
the spawn method. fleet_list therefore never reported agentSessionId
for an opencode member, and resumeSessionId was unusable for it.

Retain each spawn's PeerHandle in SessionManager, keyed by paneId, and
re-resolve a still-null agentSessionId against it from roster(), get(),
and release() (so a released member's detail also carries a
late-resolved id). Resolution is bounded: only sessions with a still-
null id do any work, a resolved id is never looked up again, and a
throwing handle degrades to "unresolved" rather than breaking the
caller. MemberSession gains a withAgentSessionId wither in the same
style as withState/withActivity.
2026-08-31 22:04:20 +07:00
Dai Ha a55079afbd #185: opt-in worktreeGroup, so a member running as another OS user can write its worktree
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m39s
Stage 3 of #185. A provisioned worktree and the repo's git store are made
group-writable when worktreeGroup names an OS group; absent, nothing changes.
The share pass runs after overlayParity, not inside add(), because
overlayParity copies more files in after add() returns.

This isolates credentials, not the repository: a member in the group can
still write the operator's git objects and refs.
2026-08-31 21:47:03 +07:00
Dai Ha e18ad4723b Merge remote-tracking branch 'refs/remotes/origin/cb206' 2026-08-31 21:47:03 +07:00
Dai Ha 6de8ac8972 #206: read opencode session ids from opencode.db, and pin the read-only open
opencode migrated its session store from a JSON file tree to SQLite in
January. OpenCodeSessionDiscovery still scanned the frozen tree, so it
returned null for every member: agentSessionId was never known and
resumeSessionId silently did nothing for every opencode profile, through
57 member spawns, with nothing logging that the search found nothing.
2026-08-31 21:46:32 +07:00
Dai Ha 6d82ca95a4 #185: skip absent git paths, and resolve the git dir instead of assuming .git
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m38s
Two defects in the stage-3 share pass, both of which would have failed EVERY
provisioning spawn once worktreeGroup was set, not only the two-user case.

.git/logs was handed to chgrp unguarded while packed-refs was guarded. It does
not exist with core.logAllRefUpdates=false, or before the first ref update, and
chgrp on a missing path exits non-zero -- surfacing as a WorktreeException that
blames a group which is in fact fine. Every path is now skipped when absent.

repoRoot + "/.git" was hardcoded. That is a FILE, not a directory, when the
checkout is itself a linked worktree -- the very thing this class creates for
every member. It now asks git: rev-parse --git-common-dir, resolved against
repoRoot because git answers relatively for an ordinary checkout.

Both new tests were watched failing with the fix removed before being kept.
2026-08-31 21:34:49 +07:00
Dai Ha 8067ee4ec4 fleetd #185: opt-in worktreeGroup config for group-shared worktrees
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Successful in 1m45s
Adds worktreeGroup (top-level FleetConfig key), Worktrees.shareWithGroup
(GitWorktrees impl: git config core.sharedRepository group + one-time
chgrp/chmod g+rwX/setgid fix-up over the worktree, .git/objects, refs,
logs, worktrees, and packed-refs when present), and wires SessionManager
to call it AFTER overlayParity so overlay files are covered too. Off by
default (byte-identical behaviour when unset). Documents the
credentials-not-repository caveat in the javadoc and example config.
2026-08-31 15:53:01 +07:00
Dai Ha ee5f8b932b #137: complete an async ticket's own reply after answer() times out
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 1m44s
fleet_send{turnId} (MessageService.answer) blocks the primary only for its
own bounded MCP-call window (25s default, 120s max) — far shorter than a
worker's resumed turn can genuinely take. When that window expires, answer()
closes its rendezvous waiter, so the worker's eventual fleet_reply has no
live waiter to resolve and falls back to the session inbox. The async
ticket's future was never completed by that path, so fleet_poll{ticket}
stayed PENDING until fleet_stop's abandon() forced it FAILED with a
misleading "the worker session was released before it replied" reason,
even though the reply had genuinely arrived.

- MessageService.reply(): before falling to the inbox, look for the async
  task this exact turn belongs to (already answered — question cleared,
  turnId still stamped — but not yet resolved) and complete it directly
  with the real reply, so fleet_poll{ticket} returns it.
- MessageService.abandon(): defense in depth, independent of the above —
  never write a false WORKER_FAILED once a reply reached the inbox for
  this target; recover and use its real content instead.
- Two new tests drive the full delegation path (async send -> ask ->
  answer with a short timeout -> reply -> poll/abandon), not a reply sink
  directly; both fail with the fix disabled and pass with it restored.
2026-08-31 14:12:23 +07:00
35 changed files with 4596 additions and 160 deletions
+27
View File
@@ -117,6 +117,15 @@ herdrSocket: ~/.config/herdr/herdr.sock
# Optional socket for member panes. Omit this to use herdrSocket for both leads and members.
# memberHerdrSocket: /Users/member/.config/herdr/herdr.sock
# fleetd #213: the login shell the member OS user (memberHerdrSocket above) actually runs. ONLY
# read when memberHerdrSocket is set — fleetd's own $SHELL says nothing about a pane running
# under a different OS user, and there is no channel to ask herdr for that user's shell, so this
# must be told rather than guessed. Absent, blank, or anything not ending in "zsh" is treated the
# same as "not zsh": the memberCredentials.policy: allow-list ZDOTDIR scrub (see worktreeGroup
# below) is skipped in favour of the weaker CB-596 sentinel overlay — a degraded control, never a
# refusal to spawn. When memberHerdrSocket is absent this key is never consulted at all.
# memberLoginShell: /bin/zsh
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
@@ -634,6 +643,24 @@ guard:
# to a sibling directory of the repo root.
# worktreeRoot: /Users/me/src/.bridged-worktrees
# Worktree group sharing (fleetd #185 stage 3). OPTIONAL, off by default. Names an OS group
# that a provisioned worktree's repo is made group-writable for (git config
# core.sharedRepository group, plus a one-time chgrp/chmod/setgid fix-up), so a member spawned
# under a DIFFERENT OS user (see memberHerdrSocket) can write its own worktree, its
# per-worktree git metadata, and its own commit objects — without it, every file GitWorktrees
# creates is owned by fleetd's own uid and unwritable by another user.
# CAUTION: this isolates credentials, not the repository — a member in the group can still
# write the operator's git objects and refs in the shared repo. The operator running fleetd
# must already be a member of the named group, or every provisioning spawn fails loudly.
#
# fleetd #213: this is also the ONE group the memberCredentials.policy: allow-list ZDOTDIR scrub
# reuses when memberHerdrSocket is set — deliberately not a second config key. Under
# memberHerdrSocket, the scrub directory is generated under worktreeRoot (never java.io.tmpdir,
# which the member OS user cannot reach) and shared read-only with this group. If worktreeGroup
# is unset while memberHerdrSocket is set, the scrub cannot be guaranteed reachable by the member,
# so fleetd falls back to the weaker CB-596 sentinel overlay instead (a WARN names the gap).
# worktreeGroup: fleet-workers
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
# the feature disabled. By default the daemon never reaps, caps, or drains sessions.
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
@@ -169,6 +169,12 @@ public final class Fleetd {
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
// pointed at the real one once it exists.
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
ExhaustionSink forwardingExhaustionSink = (target, reason) -> exhaustionSinkRef.get().onExhausted(target, reason);
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
@@ -184,7 +190,7 @@ public final class Fleetd {
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials(), config::get));
() -> config.get().memberCredentials(), config::get, forwardingExhaustionSink));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
@@ -224,7 +230,7 @@ public final class Fleetd {
contextCap = cfg.lifecycle().contextCap();
}
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
@@ -334,6 +340,11 @@ public final class Fleetd {
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
//
// fleetd #175: this sink is now also the quarantine target for OpenCodeLauncher's
// model-mismatch check — it needs nothing profile-specific from the caller beyond `target`
// (a herdr terminal id) and `reason`, so reusing it here is exactly "the existing
// ExhaustionSink path", not a new mechanism.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
@@ -342,10 +353,12 @@ public final class Fleetd {
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
// now that `sessions` exists to resolve target -> session -> profile.
exhaustionSinkRef.set(exhaustionSink);
AgentControl agents = router.memberAgents();
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
@@ -5,17 +5,70 @@ import dev.ltms.fleet.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
/** A slot held before a member process starts. */
record SlotReservation(String slot, String profile) { }
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
public MemberRole acquired(MemberRole role, String profile, String terminal) {
return role; // no registry configured — nothing to bind against, so the request stands
}
@Override
public void released(String terminal) {
}
@Override
public void requireSlotFor(MemberRole role, String profile) {
// no registry configured — nothing to validate against, so nothing is refused
}
@Override
public SlotReservation reserve(MemberRole role, String profile) {
return null;
}
@Override
public boolean bind(SlotReservation reservation, String terminal) {
return false;
}
@Override
public void release(SlotReservation reservation) {
}
};
void acquired(MemberRole role, String profile, String terminal);
/**
* Try to bind a newly spawned {@code terminal} into the role it was granted.
*
* @return the role this session actually holds: {@code role} unchanged for a role with no
* live slot-binding semantics (dev, reviewer), or when the bind succeeded; a fallback
* role — never {@code role} — when a slot-bound role (architect) could not be bound.
* Callers must record THIS value on the session, never the requested {@code role}, so
* a later roster read never reports a role the session does not hold (CB-619). In
* normal operation this fallback should not happen once a reservation has been bound.
* It remains the honest answer if a caller has no reservation, or if binding a
* reservation unexpectedly fails.
*/
MemberRole acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
/**
* Refuse an acquire before anything spawns when {@code role} requires a live slot binding and
* no configured slot carries {@code profile} (CB-619 / fleetd #123). A no-op for a role with
* no slot-binding semantics.
*
* @throws IllegalArgumentException naming the role, the profile, and the pools that do carry it
*/
void requireSlotFor(MemberRole role, String profile);
/** Reserve a matching slot before launch, or refuse before a charter can be delivered. */
SlotReservation reserve(MemberRole role, String profile);
/** Convert a reservation into a live terminal binding. */
boolean bind(SlotReservation reservation, String terminal);
/** Return an unbound reservation after a failed launch. */
void release(SlotReservation reservation);
}
@@ -8,6 +8,7 @@ import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
@@ -55,8 +56,10 @@ public final class MemberRegistry implements MemberLifecycle {
}
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(FleetConfig.Fleet fleet) {
@@ -166,7 +169,7 @@ public final class MemberRegistry implements MemberLifecycle {
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
if (terminalToSlot.containsValue(slot) || reservedSlots.contains(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
@@ -206,19 +209,116 @@ public final class MemberRegistry implements MemberLifecycle {
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*
* <p>CB-619 / fleetd #123: the return value is the role this session actually holds, and the
* caller is required to record THAT — never the requested {@code role} — on the session. Before
* this fix the caller kept the requested role regardless of whether the bind below succeeded, so
* a demoted session's {@code GET /members} row still said {@code "architect"} while
* {@code fleet_whoami} (which reads the live binding, not the request) correctly said
* {@code "worker"} — three sources of truth that disagreed about one live member, silently.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
public MemberRole acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
return role;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
return MemberRole.ARCHITECT;
}
}
// fleetd #123: at least WARN — a role downgrade that the roster must now also reflect is
// not routine bookkeeping. requireSlotFor already refuses the config-gap case (no slot at
// all carries this profile) before a process ever spawns; reaching here means the config DID
// carry a matching slot but every one of them was already bound to a different terminal — a
// race this pre-spawn check cannot close on its own (see requireSlotFor's javadoc).
log.warn("member slot: no free architect slot for profile={} terminal={}; holding the session "
+ "as {} instead of the architect it asked for — every configured slot for this "
+ "profile is already bound to a different terminal", profile, terminal,
MemberRole.DEV.wireName());
return MemberRole.DEV;
}
/**
* CB-619 / fleetd #123: refuse an architect acquire before anything spawns when no configured
* slot carries {@code profile} — the config-gap case from the original defect report (a spawn
* asked for {@code role=architect, profile=sonnet}, and {@code fleet.architects} carried only
* {@code opus} and {@code sol}). A dev/reviewer acquire is always a no-op: those pools are
* placement candidates only (see {@code CompositePeerLauncher}), never a live identity binding,
* so there is nothing here to refuse — an explicit profile outside the pool for those roles is a
* documented operator override, not a defect.
*
* <p>This closes the config-gap case, not the live-capacity case: a profile that DOES carry a
* slot can still lose the race to a concurrent spawn between this check and the actual
* {@link #bind}, which is why {@link #acquired} must still answer honestly even after this
* check has passed.
*/
@Override
public void requireSlotFor(MemberRole role, String profile) {
if (role != MemberRole.ARCHITECT) {
return;
}
boolean hasSlot = slotsFor(MemberRole.ARCHITECT).values().stream()
.anyMatch(e -> Objects.equals(profile, e.profile()));
if (hasSlot) {
return;
}
List<String> pools = slotsFor(MemberRole.ARCHITECT).values().stream()
.map(Entry::profile)
.distinct()
.toList();
throw new IllegalArgumentException(
"no " + role.wireName() + " slot for profile '" + profile + "' — an architect's "
+ "identity IS the slot it is bound to, so there is nothing to bind this "
+ "session's identity to. fleet." + role.configKey() + " carries profiles: "
+ (pools.isEmpty() ? "(none configured)" : String.join(", ", pools))
+ "; add profile '" + profile + "' there, or spawn " + role.wireName()
+ " on one of those profiles instead");
}
@Override
public SlotReservation reserve(MemberRole role, String profile) {
if (role != MemberRole.ARCHITECT) {
return null;
}
synchronized (terminalToSlot) {
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if ((profile == null || profile.isBlank() || Objects.equals(profile, entry.profile()))
&& !terminalToSlot.containsValue(entry.key()) && reservedSlots.add(entry.key())) {
return new SlotReservation(entry.key(), entry.profile());
}
}
}
throw new IllegalArgumentException("no free architect slot for profile '" + profile
+ "' — every matching slot is already bound or reserved");
}
@Override
public boolean bind(SlotReservation reservation, String terminal) {
if (reservation == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!reservedSlots.remove(reservation.slot())) {
return false;
}
if (!isSlot(reservation.slot()) || terminalToSlot.containsKey(terminal)
|| terminalToSlot.containsValue(reservation.slot())) {
return false;
}
terminalToSlot.put(terminal, reservation.slot());
return true;
}
}
@Override
public void release(SlotReservation reservation) {
if (reservation != null) {
synchronized (terminalToSlot) {
reservedSlots.remove(reservation.slot());
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@@ -76,6 +76,26 @@ import java.util.Set;
* stay on {@code broker}'s vhost). {@code null} → no lead mailbox is opened.
* Config parsing + accessors only — nothing here wires it into a live
* {@code LeadMailbox}; that is a separate ticket. See {@link Coordinator}.
* @param worktreeGroup optional OS group name (fleetd #185 stage 3) that makes a provisioned
* worktree's repo group-shared, so a member running as a different OS user
* (see {@code memberHerdrSocket}) can write its own worktree, its per-worktree
* git metadata, and its own commit objects. {@code null}/blank/empty ⇒ off,
* today's behaviour unchanged (every file stays owned by fleetd's own uid).
* <strong>This isolates credentials, not the repository</strong>: a member in
* the group can still write the operator's git objects and refs in the shared
* repo. See {@link dev.ltms.fleet.session.Worktrees#shareWithGroup}.
* @param memberLoginShell fleetd #213: the login shell the member's OS user actually runs, ONLY
* meaningful (and only ever read) when {@code memberHerdrSocket} is
* configured — that mode spawns member panes under a different OS user than
* fleetd's own process, so fleetd's own {@code $SHELL} says nothing about what
* that pane runs. There is no channel to ask herdr for another user's shell, so
* this must be told, never guessed. {@code null}/blank (or a value not ending
* in {@code zsh}) is treated the same as "not zsh": the {@code
* memberCredentials.policy: allow-list} ZDOTDIR scrub is skipped in favour of
* the CB-596 sentinel overlay — a degraded control, never a refusal to spawn.
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -98,7 +118,33 @@ public record FleetConfig(
ConfigReload configReload,
Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials,
Coordinator coordinator) {
Coordinator coordinator,
String worktreeGroup,
String memberLoginShell) {
/** Back-compat form before the {@code memberLoginShell} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup, null);
}
/** Back-compat form before the {@code worktreeGroup} key was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, null, null);
}
/** Back-compat form before the {@code coordinator:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
@@ -109,7 +155,7 @@ public record FleetConfig(
MemberCredentials memberCredentials) {
this(bind, herdrSocket, null, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, null);
configReload, quarantineCooldownSeconds, memberCredentials, null, null);
}
/** Back-compat form before the CB-596 {@code memberCredentials:} block was added. */
@@ -1325,7 +1371,7 @@ public record FleetConfig(
"bind", "herdrSocket", "memberHerdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator");
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -1943,9 +1989,14 @@ public record FleetConfig(
: new MemberCredentials(null, List.of(), List.of());
// coordinator is left as-is, like broker/primary above: null keeps no LeadMailbox opened,
// and this ticket's Coordinator is config-only anyway (nothing yet reads it at startup).
// worktreeGroup is left as-is (fleetd #185 stage 3): null/blank is "off", and there is no
// sane non-null default — an OS group name is operator-specific.
// memberLoginShell is left as-is (fleetd #213), like worktreeGroup: null/blank is "not
// configured", and there is no sane non-null default — a member's login shell is
// operator-specific and only meaningful when memberHerdrSocket is also set.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator);
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell);
}
/**
@@ -237,11 +237,13 @@ public final class CompletionResolver implements TurnListener {
}
String tail;
String assistantBlock = null;
String rawScrape = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
rawScrape = agents.read(target, SCRAPE_SOURCE);
assistantBlock = lastAssistantBlock(rawScrape);
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
@@ -256,6 +258,20 @@ public final class CompletionResolver implements TurnListener {
// member, so a caller (including a lead deciding whether to delegate again) can tell a lost
// turn from a real empty answer.
if (scrapeFailed || tail.isEmpty()) {
// fleetd#211: lastAssistantBlock() found nothing usable — most often a pane with no ⏺
// marker at all, whose boundary scan then starts at the top of the raw screen and breaks
// immediately on the first line of TUI chrome (╭, │, ❯, …). Before giving up as a lost
// turn, run the same exhaustion/backend-error classification against the RAW scrape as a
// fallback, ONLY here. A pane that already yielded a usable block never reaches this
// branch, so the narrow (trimmed) match on the normal path below is completely unchanged
// — zero new false positives there. Every pane this fallback examines was already headed
// for the empty-scrape failure, so a wrong label here is strictly less bad than silently
// losing an exhaustion signal: the alternative outcome is already a failure, just one that
// never quarantines the credential. lastAssistantBlock stays the source of the reply
// TEXT everywhere else; only classification ever consults the raw scrape, and only here.
if (rawScrape != null && classifyRawScrapeFallback(target, turn, waiter, rawScrape)) {
return;
}
fail(target, turn, emptyScrapeReason(target, scrapeFailed));
return;
}
@@ -313,6 +329,43 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* fleetd#211: the raw-scrape fallback classification, run only when {@link #lastAssistantBlock}
* found nothing usable (see the call site in {@link #resolve}). Mirrors the two classifications
* the normal path already applies to the trimmed assistant block — exhaustion first, then the
* narrow {@link #BACKEND_ERROR} pattern — against {@code raw} instead, and reports whether one of
* them handled the turn (resolved the waiter or failed it) so the caller skips the empty-scrape
* failure. Never runs on the normal (non-empty-block) path, and never touches the reply text.
*/
private boolean classifyRawScrapeFallback(String target, InFlight turn,
CompletableFuture<Rendezvous.Resolution> waiter, String raw) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(raw, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED from the raw scrape (no "
+ "usable assistant block; no fleet_reply): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return true;
}
String backendError = firstMatchingLine(raw, BACKEND_ERROR);
if (backendError != null) {
// Carry the pane, not just the matched line — the same fleetd#164 rule the normal path
// above applies. Here it matters more, not less: the trimmed block was empty, so the raw
// scrape is the ONLY copy of whatever the member managed to say. Clipped to the same cap
// the normal path uses, since a raw screen has no boundary trimming to bound it.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + clip(raw));
return true;
}
return false;
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
@@ -951,7 +951,10 @@ public final class FleetMcp {
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
List<MemberSession> roster = sessions.roster();
// fleetd #209: this is the caller-driven fleet_list read that actually reports
// agentSessionId (via memberCapacityView -> SessionManager.rosterView), so it uses the
// resolving roster; the heartbeat/health/metrics timers stay on the plain sessions.roster().
List<MemberSession> roster = sessions.rosterResolved();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
@@ -278,18 +278,22 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
* session id — the resume handle. A resume passes the prior id via {@code -r} and returns
* that id; every other spawn mints a new UUID, passes it via {@code --session-id}, and
* returns the mint. The bridge's logical name rides along as {@code -n} when present.
*
* <p>fleetd #214: the mint is unconditional. A plain {@code fleet_spawn} passes neither
* sessionName nor resumeSessionId, yet the member must still be resumable, and this id is
* the only resume handle a claude-code member has — unlike opencode, nothing resolves it
* after the launch. Checked against the real binary (claude 2.1.252): the flag is safe on
* every spawn. The binary takes only a valid UUID — it refuses any other value at argument
* parsing ("Invalid session ID. Must be a valid UUID.") — so the {@code UUID.randomUUID()}
* mint is required, not incidental. The only flag interaction the binary documents is with
* {@code -r} (both claim the session id), and the resume branch above never combines the two.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
@@ -320,11 +324,19 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
*
* <p>CB-618: Claude Code refuses to start when BOTH {@code --append-system-prompt} and
* {@code --append-system-prompt-file} are on the command line ("Cannot use both ... Please use
* only one"), so the two charters can never travel on separate flags. When both are present they
* are concatenated into the one file, role charter first and reply charter last — last is where
* the reply rule must sit, because it is the rule that must survive. When only the reply charter
* is present it keeps its proven inline {@code --append-system-prompt} delivery, which is also
* the only form that reaches a member with no repo checkout.
* only one"), so the two charters can never travel on separate flags. They are concatenated
* into the one file, role charter first and reply charter last — last is where the reply rule
* must sit, because it is the rule that must survive.
*
* <p>fleetd #220: a lone reply charter used to ride inline on {@code --append-system-prompt},
* which put ~800 bytes of prose on the command line herdr types into the pane. That line is
* capped at {@value HerdrPeerLauncher#PANE_COMMAND_BYTE_LIMIT} bytes by the pty itself, and
* everything past the cap is dropped with no error from any layer. The charter alone left about
* 50 bytes of headroom, so adding one flag ({@code --session-id}, fleetd #214) truncated the
* LAST argument instead — {@code --autocompact 250000} arrived as {@code --autocompact 25},
* claude rejected it, and every claude-code spawn died as an unexplained readiness timeout.
* The charter now always travels as a file, which takes the prose off the command line for
* good; {@link HerdrPeerLauncher#checkPaneCommandFits} is the backstop for whatever grows next.
*/
private List<String> argvWithFleet(FleetConfig.Profile cfg, LaunchSpec spec) {
String roleCharter = nonBlank(spec.roleCharter());
@@ -341,16 +353,13 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
argv.add("--mcp-config");
argv.add(mcpConfigJson(cfg));
}
// Combine the charters in order role -> reply, dropping any that are absent. When two
// or more survive they must ride one --append-system-prompt-file (CB-618 forbids the inline
// flag and the file flag together). A lone reply charter keeps its proven inline delivery.
// Combine the charters in order role -> reply, dropping any that are absent. They ride one
// --append-system-prompt-file (CB-618 forbids the inline flag and the file flag together),
// always — fleetd #220: charter prose on the command line overruns the pane's byte cap.
List<String> charters = new java.util.ArrayList<>(2);
if (roleCharter != null) charters.add(roleCharter);
if (replyCharter != null) charters.add(replyCharter);
if (charters.size() == 1 && replyCharter != null && roleCharter == null) {
argv.add("--append-system-prompt");
argv.add(replyCharter);
} else if (!charters.isEmpty()) {
if (!charters.isEmpty()) {
argv.add("--append-system-prompt-file");
argv.add(writeCharterFile(String.join("\n\n", charters)).toString());
}
@@ -450,12 +459,83 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* {@link OpenCodeLauncher#writeConfig} already uses for its charter file, since the process that
* reads this file (the spawned peer) outlives this JVM call and there is no spawn-scoped teardown
* hook to delete it synchronously.
*
* <p><b>fleetd #222.</b> {@code Files.createTempFile(prefix, suffix)} with no directory argument
* resolves against {@code java.io.tmpdir} — on macOS the per-user {@code $TMPDIR} under
* {@code /var/folders/...}, mode {@code 0700}, both resolved against FLEETD's own OS user. Under
* {@code memberHerdrSocket:} the member pane runs as a DIFFERENT OS user, so that user cannot even
* traverse the directory, let alone read the file — and since fleetd #220 the charter file is the
* ONLY delivery path for {@code --append-system-prompt-file}, always, not merely the fallback it
* used to be. A member handed a path it cannot read is not degraded, it is broken: see this
* method's refusal branch below.
*
* <p><b>Measured severity (fleetd #222 real-binary check, claude 2.1.258):</b> an unreadable
* {@code --append-system-prompt-file} is the LOUD failure, not the silent one. {@code claude}
* checks the file before touching auth or the network — invoked with a bogus API key against a
* {@code chmod 000} file, it printed {@code Error reading append system prompt file: EACCES:
* permission denied, open '<path>'} and exited 1 immediately (a nonexistent path gets {@code
* Error: Append system prompt file not found: <path>}, same exit code). So the pre-fix bug did
* NOT leave a charter-less member silently occupying a pane and never calling {@code
* fleet_reply} — it made the herdr pane exit immediately, which the CB-306 spawn-readiness gate
* (this launcher's {@code spawnReadyTimeoutMs} poll) would have surfaced as "did not reach
* injectable state", the same unexplained-timeout shape fleetd #220 already describes. Still a
* real defect (every claude-code member under {@code memberHerdrSocket} would have failed to
* spawn), but not the worse, undetectable failure mode.
*
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only live mode): byte-identical to before this
* fix — {@code Files.createTempFile("fleetd-role-charter-", ".md")} with no directory
* argument, i.e. still resolved against {@code java.io.tmpdir}.</li>
* <li>{@code memberHerdrSocket} PRESENT: a fresh per-spawn directory is created under {@code
* worktreeRoot} (never {@code java.io.tmpdir}) holding just the charter file, then shared
* read-only with {@code worktreeGroup} via {@link EnvAllowListScrub#shareWithGroup} — the
* SAME mechanism fleetd #213 built for the ZDOTDIR scrub and fleetd #219 reused for {@link
* OpenCodeLauncher#writeConfig}'s {@code opencode.json} directory, reused here rather than
* duplicated a third time. A per-spawn subdirectory (not {@code worktreeRoot} itself) is the
* unit {@code shareWithGroup} chmods, so this never touches permissions on anything else
* under {@code worktreeRoot}.</li>
* </ul>
*
* <p><b>Follows fleetd #219's REFUSAL decision, not #213's degrade decision.</b> The ZDOTDIR
* scrub is a credential CONTROL — a degraded control (the CB-596 sentinel overlay) still has
* value, so #213 falls back rather than refusing. A charter is NOT a control, it is the member's
* TURN CONTRACT (the rule that ends every turn with {@code fleet_reply}). A member spawned with no
* charter is not degraded, it is broken: either claude-code exits on the unreadable
* {@code --append-system-prompt-file} path and the spawn dies at the readiness gate (loud), or it
* starts anyway with no charter and never calls {@code fleet_reply} — the sender silently gets
* nothing (silent). Neither outcome is worth trading for "spawn something." So a missing {@code
* worktreeRoot}/{@code worktreeGroup} under {@code memberHerdrSocket} refuses the spawn here,
* naming the missing key, exactly like {@link OpenCodeLauncher#configParentDir()}.
*
* @throws IllegalStateException when {@code memberHerdrSocket} is configured but {@code
* worktreeRoot} and/or {@code worktreeGroup} is not
*/
private static Path writeCharterFile(String charterText) {
private Path writeCharterFile(String charterText) {
try {
Path file = Files.createTempFile("fleetd-role-charter-", ".md");
if (!memberHerdrSocketConfigured()) {
Path file = Files.createTempFile("fleetd-role-charter-", ".md");
Files.writeString(file, charterText);
file.toFile().deleteOnExit();
return file;
}
Path parentDir = memberScrubParentDir();
String group = memberGroup();
if (parentDir == null || group == null) {
throw new IllegalStateException("memberHerdrSocket is configured, so the role/reply "
+ "charter file (mounted via --append-system-prompt-file) must be placed where "
+ "the member's OS user can read it — worktreeRoot, shared via worktreeGroup — "
+ "but " + (parentDir == null ? "worktreeRoot" : "worktreeGroup") + " is not "
+ "configured. Refusing to spawn rather than hand the member a charter path it "
+ "cannot read: that member's turn contract (the fleet_reply rule) would never "
+ "reach it. Configure both worktreeRoot and worktreeGroup to enable claude-code "
+ "member spawns under memberHerdrSocket.");
}
Path dir = Files.createTempDirectory(parentDir, "fleetd-role-charter-");
dir.toFile().deleteOnExit();
Path file = dir.resolve("charter.md");
Files.writeString(file, charterText);
file.toFile().deleteOnExit();
EnvAllowListScrub.shareWithGroup(dir, group);
return file;
} catch (IOException e) {
throw new UncheckedIOException("cannot write role charter temp file", e);
@@ -7,6 +7,9 @@ import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.GroupPrincipal;
import java.nio.file.attribute.PosixFileAttributeView;
import java.nio.file.attribute.PosixFilePermissions;
import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
@@ -116,6 +119,77 @@ public final class EnvAllowListScrub {
}
}
/**
* fleetd #213: as {@link #generate(Path, Set)}, plus share the generated directory with
* {@code group} — the member OS user's group (the operator's existing {@code worktreeGroup:}
* name, reused rather than inventing a second one) — so a member running under a different OS
* user than fleetd's own process can still read what it needs from a directory placed outside
* {@code java.io.tmpdir}. {@code group} null/blank ⇒ identical to {@link #generate(Path, Set)};
* this is the single-daemon (no {@code memberHerdrSocket}) shape, where the pane is fleetd's own
* uid and no group sharing is needed.
*
* @throws UncheckedIOException also when {@code group} does not resolve on this host, or a
* group-ownership/permission call is refused — the same "fail
* loudly rather than start unprotected" contract as above: a scrub
* the configured member user cannot even read is not a working
* control.
*/
public static Path generate(Path parentDir, Set<String> allowedNames, String group) {
Path dir = generate(parentDir, allowedNames);
if (group != null && !group.isBlank()) {
shareWithGroup(dir, group);
}
return dir;
}
/**
* chgrp/chmod-equivalent over a freshly generated directory and the flat files already written
* into it: owner keeps full access, {@code group} gets traverse+read on the directory ({@code
* rwxr-x---}, so a member process — a login shell reading it via {@code ZDOTDIR}, or another
* process simply opening a file under it — running under that group can find and read the
* files) and read-only on each file ({@code rw-r-----}) — deliberately no group WRITE anywhere,
* since a member never needs to add or change what fleetd generated. (For the ZDOTDIR scrub
* specifically, this also means the scrub script's own report write inside the pane fails
* closed rather than open — see {@code scrub.zsh}'s trailing {@code 2>/dev/null} — which {@link
* dev.ltms.fleet.member.HerdrPeerLauncher#releaseZdotdir} already treats as "cannot be
* confirmed to have run" rather than success.)
*
* <p>Package-private and named generically on purpose: fleetd #213 built this for the ZDOTDIR
* scrub directory, and fleetd #219 reuses it verbatim for {@link
* dev.ltms.fleet.member.OpenCodeLauncher}'s ephemeral {@code opencode.json} directory — both are
* "a fleetd-generated directory of flat files that a different-uid member process must read but
* never write," so the sharing mechanism is shared rather than copied a second time.
*/
static void shareWithGroup(Path dir, String group) {
try {
GroupPrincipal principal = dir.getFileSystem().getUserPrincipalLookupService()
.lookupPrincipalByGroupName(group);
setGroupAndPermissions(dir, principal, "rwxr-x---");
try (Stream<Path> entries = Files.list(dir)) {
for (Path file : entries.toList()) {
setGroupAndPermissions(file, principal, "rw-r-----");
}
}
} catch (IOException e) {
throw new UncheckedIOException("cannot share generated directory " + dir + " with group '"
+ group + "' — the group must exist, and the fleetd operator ("
+ System.getProperty("user.name") + ") must be a member of it", e);
} catch (UnsupportedOperationException e) {
throw new UncheckedIOException("cannot share generated directory " + dir + " with group '"
+ group + "' — this filesystem does not support POSIX group ownership",
new IOException(e));
}
}
private static void setGroupAndPermissions(Path path, GroupPrincipal group, String perms) throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(path, PosixFileAttributeView.class);
if (view == null) {
throw new IOException("POSIX file attributes are not supported for " + path);
}
view.setGroup(group);
Files.setPosixFilePermissions(path, PosixFilePermissions.fromString(perms));
}
/** One operator-sourcing startup file: source the {@code $HOME} counterpart, change nothing else. */
private static String homeSourcingFile(String name) {
return """
@@ -3,11 +3,13 @@ package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.StatusRefiner;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
@@ -112,6 +114,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* daemon's own process is started the same way (a login shell sourcing the same secret store —
* see CB-592's investigation of {@code secrets.sh}), so on a single-host deployment its env
* mirrors what the pane's login shell is about to export.
*
* <p>fleetd #185 stage 2: that mirroring assumption holds only while the member pane runs under
* the SAME OS user as the daemon. When {@code memberHerdrSocket:} is configured, member panes
* run on a second herdr owned by a different user — different {@code $HOME}, different {@code
* secrets.sh}, different environment entirely — so this field's data no longer describes what a
* member pane inherits. See {@link #logCredentialGap} for how that mode is handled.
*/
private final Supplier<Set<String>> hostEnvNames;
@@ -145,6 +153,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
/**
* fleetd #176 fix 2: the same UNKNOWN-refinement {@code dev.ltms.fleet.inject.StatusPoller}
* uses, reused here for the spawn-readiness gate. Constructed once from {@link #agents} — see
* {@link #refinedInjectable(String, Agent)} for the corroboration that keeps it from firing on
* a dead pane's bare shell prompt.
*/
private final StatusRefiner statusRefiner;
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
@@ -266,6 +282,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
this.memberCredentials = memberCredentials;
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
this.config = config;
this.statusRefiner = new StatusRefiner(agents);
}
// --- adapter seams -------------------------------------------------------------------------
@@ -676,6 +693,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
checkPaneCommandFits(cfg, argv);
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
@@ -691,6 +709,59 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw last;
}
/**
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line,
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent
* started", the backend exits on the mangled argument it was handed, the pane closes, and the
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
* would have hit anyway, named at the point it is still
* explainable
*/
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new PeerUnreachableException(
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}.
*/
static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
@@ -849,25 +920,135 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* Poll {@link AgentControl#get} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*
* <p>fleetd #176 fix 1: {@code agents.get} was previously unguarded here, so a herdr
* {@code *_not_found} answer — which is what happens when the backend process EXITED rather
* than being slow — propagated as a raw {@link HerdrException} instead of the
* {@link PeerUnreachableException} every other failure path of this gate produces, and skipped
* teardown ({@link #stop}) entirely, leaking the pane/tab. {@link #failFastOnGoneBackend} closes
* that gap: it stops waiting immediately (never burns the rest of the timeout), runs the same
* teardown the timeout path below runs, and throws with a message that says the backend exited
* rather than that the pane was slow. Any other {@link HerdrException} still propagates
* unchanged — this gate does not know how to recover from it.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
long start = nowMillis.getAsLong();
long deadline = start + spawnReadyTimeoutMs;
AgentStatus lastStatus = null;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
Agent sample;
try {
sample = agents.get(paneId);
} catch (HerdrException e) {
if (isAlreadyGone(e)) {
failFastOnGoneBackend(paneId, e, nowMillis.getAsLong() - start);
}
throw e; // any other herdr failure is not ours to interpret — let it propagate
}
lastStatus = sample.status();
if (lastStatus.injectable() || refinedInjectable(paneId, sample)) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
// fleetd #220: read the pane BEFORE stop() closes it. Without this the gate says only that
// it timed out, which is true of every cause — a backend that never launched, a binary that
// rejected an argument and exited, a trust prompt, a login shell that hung. The pane holds
// the one copy of that answer and it is destroyed a line later.
log.warn("peer pane={} did not become injectable within {}ms (last status {}) — closing. "
+ "Pane tail:\n{}",
paneId, spawnReadyTimeoutMs, lastStatus, readPaneQuietly(paneId));
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/**
* fleetd #176 fix 1: the backend process exited while the gate was still waiting — herdr
* answered {@code *_not_found} instead of ever reporting an injectable status. Fails
* immediately (never burns the rest of {@link #spawnReadyTimeoutMs}), runs the exact same
* teardown {@link #waitUntilInjectableOrThrow}'s timeout path runs, and throws a
* {@link PeerUnreachableException} whose message says the process exited rather than that the
* pane was slow.
*
* @throws PeerUnreachableException always — this method never returns normally
*/
private void failFastOnGoneBackend(String paneId, HerdrException cause, long elapsedMs) {
String tail = readPaneQuietly(paneId); // read before stop() closes the pane
log.warn("peer pane={} backend process exited after {}ms while waiting for injectable "
+ "state (herdr: {}) — closing. Pane tail:\n{}",
paneId, elapsedMs, cause.getMessage(), tail);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " backend process exited after " + elapsedMs
+ "ms while waiting to become injectable (herdr reported: " + cause.getMessage()
+ "). Pane tail:\n" + tail);
}
/**
* fleetd #176 fix 2: resolve a raw {@link AgentStatus#UNKNOWN} sample into a trustworthy
* injectable state via pane content, the same refinement {@code StatusPoller} applies during a
* peer's working life — but corroborated, because the trap this gate is exposed to that the
* poller is not: a pane whose backend has already exited settles at a plain shell prompt, and
* that prompt commonly contains the same {@code ❯} glyph {@link StatusRefiner#classify} treats
* as "idle at the Claude Code TUI prompt". Naively wiring the refiner in would turn "the backend
* died" into "ready to inject" — strictly worse than today's timeout.
*
* <p>Two guards, both required:
* <ul>
* <li>{@link StatusRefiner#classify} is written for the Claude Code TUI only (its own javadoc
* says so), so refinement only ever runs for the {@code claude} adapter — never for
* {@code opencode} or any future non-Claude backend, whatever its pane looks like.</li>
* <li>The refined status is accepted only when the <em>same</em> {@code agents.get} sample
* ({@code sample}, taken once by the caller) still reports a non-null
* {@link Agent#agentType()} — herdr's own "a supported backend is still detected here"
* signal, read from the very sample the raw status came from so the two can never
* disagree. A bare shell prompt left by an exited backend reports no {@code agentType}.</li>
* </ul>
*
* <p>Called only when {@code sample.status() == UNKNOWN}, so a healthy spawn — which never sees
* {@code UNKNOWN} — triggers zero extra herdr calls; only a persistently-{@code unknown} pane
* pays the one extra {@code agent.read} {@link StatusRefiner#refine} performs.
*/
private boolean refinedInjectable(String paneId, Agent sample) {
if (sample.status() != AgentStatus.UNKNOWN) {
return false;
}
if (!"claude".equals(namePrefix)) {
return false; // StatusRefiner.classify reads a Claude Code TUI prompt specifically
}
if (sample.agentType() == null) {
return false; // no corroborating liveness signal — could be a dead pane's bare shell
}
return statusRefiner.refine(paneId, AgentStatus.UNKNOWN, agents).injectable();
}
/**
* fleetd #220: the pane's recent output, clipped, for the readiness-gate timeout log — or a
* short note when it cannot be read. Best-effort by construction: this runs on a path that is
* already failing, so it must never replace the real error with one of its own.
*/
private String readPaneQuietly(String paneId) {
try {
String pane = agents.read(paneId, "recent");
if (pane == null || pane.isBlank()) {
return "<pane read returned nothing>";
}
return pane.length() <= SPAWN_FAILURE_PANE_CHARS
? pane
: pane.substring(pane.length() - SPAWN_FAILURE_PANE_CHARS);
} catch (RuntimeException e) {
return "<pane could not be read: " + e.getMessage() + ">";
}
}
/** How much of a failed spawn's pane the timeout log carries. */
private static final int SPAWN_FAILURE_PANE_CHARS = 4000;
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* the session identity the launch resolved (CB-547a): the bridge's logical name and the peer's
@@ -1071,6 +1252,29 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* name. It stays a one-off decision because it is a live handle to the operator's ssh-agent, not
* a value — a member holding it can sign with every key the agent holds, so letting it ride in
* on the generic {@code allow:} list would hand that out for an unrelated reason.
*
* <p>fleetd #213: which shell decides the zsh gate, and where the generated directory lives,
* both depend on whether {@code memberHerdrSocket:} is configured — see {@link
* #memberHerdrSocketConfigured()}'s javadoc for why fleetd's own {@code $SHELL} and {@code
* java.io.tmpdir} describe the wrong process once member panes run under a different OS user.
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only mode): byte-identical to before this
* fix — fleetd's own {@code $SHELL} decides zsh, and the directory is generated under
* {@code java.io.tmpdir}.</li>
* <li>{@code memberHerdrSocket} PRESENT: the configured {@code memberLoginShell:} decides
* zsh instead — fleetd's own {@code $SHELL} is never consulted, since it names a
* different user's shell, not the member's. Absent/non-zsh falls back exactly like the
* non-zsh case below. When it IS zsh, the directory still cannot go under {@code
* java.io.tmpdir} (mode 0700, unreadable by another uid — the exact gap fleetd #213
* exists to close), so it is generated under {@code worktreeRoot} instead and shared
* read-only with {@code worktreeGroup} — the same group {@link
* dev.ltms.fleet.session.Worktrees#shareWithGroup} already uses, reused rather than
* inventing a second group key. Either one missing means the scrub cannot be guaranteed
* reachable by the member, which is the same "cannot guarantee the scrub runs" case as a
* non-zsh shell, so it gets the identical fallback.</li>
* </ul>
* In every branch: never refuse to spawn. A degraded credential control must not become an
* outage for an opt-in feature.
*/
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
@@ -1078,8 +1282,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return null;
}
Set<String> allowed = derivedAllowedNames(creds, launch);
String loginShell = resolveEnv("SHELL");
boolean zsh = loginShell != null && (loginShell.endsWith("/zsh") || loginShell.equals("zsh"));
boolean memberHerdrSocket = memberHerdrSocketConfigured();
// fleetd #213 defect 1: under memberHerdrSocket the member pane runs as a DIFFERENT OS
// user, so fleetd's own $SHELL says nothing about what that pane runs — resolveEnv("SHELL")
// must not even be called on this path, only the explicit memberLoginShell: config can
// answer it. With memberHerdrSocket absent, nothing here changes: fleetd's own $SHELL is
// still the input, exactly as before this fix.
String loginShell = memberHerdrSocket ? configuredMemberLoginShell() : resolveEnv("SHELL");
boolean zsh = isZshShell(loginShell);
if (!zsh) {
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
@@ -1093,13 +1303,35 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
logCredentialGap(creds, null);
return null;
}
Path parentDir;
String group = null;
if (memberHerdrSocket) {
// fleetd #213 defect 2: java.io.tmpdir is fleetd's own per-user temp dir (mode 0700 on
// macOS) — a member running as a different uid cannot even traverse it, let alone read
// the generated files. worktreeRoot is the only configured location a different-uid
// member can be given access to, and only WITH worktreeGroup to grant that access —
// absent either, the scrub cannot be guaranteed reachable, so this falls back exactly
// like the non-zsh case above rather than generating a directory nothing can read.
parentDir = memberScrubParentDir();
group = memberGroup();
if (parentDir == null || group == null) {
warnCannotShareScrubDirectory();
overlayBlockedCredentials(launch.env(), creds);
logCredentialGap(creds, null);
return null;
}
} else {
parentDir = Path.of(System.getProperty("java.io.tmpdir"));
}
// Only reached when the scrub is actually about to run — the count below describes that
// scrub, so it must not be logged before this gate (see the non-zsh branch above). Same
// reasoning gates logCredentialGap's wording: passing the derived `allowed` set (non-null)
// here, and ONLY here, is what tells it the scrub will really blank an unkept name — #192.
logAllowListCoverage(allowed);
logCredentialGap(creds, allowed);
Path dir = EnvAllowListScrub.generate(Path.of(System.getProperty("java.io.tmpdir")), allowed);
Path dir = memberHerdrSocket
? EnvAllowListScrub.generate(parentDir, allowed, group)
: EnvAllowListScrub.generate(parentDir, allowed);
launch.env().put("ZDOTDIR", dir.toAbsolutePath().toString());
log.info("memberCredentials policy=allow-list: profile={} generated ZDOTDIR {} — derived "
+ "allow-list holds {} name(s); the pane reports allowed N of M at release",
@@ -1107,6 +1339,63 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return dir;
}
/** True when {@code shell} is a zsh login shell path or bare name — the ZDOTDIR gate. */
private static boolean isZshShell(String shell) {
return shell != null && (shell.endsWith("/zsh") || shell.equals("zsh"));
}
/**
* fleetd #213: the configured {@code memberLoginShell:}, or {@code null} when unconfigured.
* Called ONLY from the {@code memberHerdrSocket}-configured branch of {@link
* #applyEnvironmentAllowListPolicy} — fleetd's own {@code $SHELL} is never read on that path.
* {@link #config} being {@code null} (an older test call site, or a launcher that never
* threaded the full config through) is treated the same as "not configured".
*/
private String configuredMemberLoginShell() {
FleetConfig cfg = config == null ? null : config.get();
return cfg == null ? null : cfg.memberLoginShell();
}
/**
* fleetd #213: {@code worktreeRoot}, as the ZDOTDIR scrub's parent directory under {@code
* memberHerdrSocket}, or {@code null} when unconfigured — the same "cannot guarantee the scrub
* runs" gap as {@link #memberGroup()} being unset (see {@link
* #applyEnvironmentAllowListPolicy}). Deliberately no sibling-of-repo-root default here, unlike
* {@code GitWorktrees}' own {@code worktreeRoot} resolution: that default is a convenience for
* provisioning a worktree that will exist regardless, whereas an unconfigured value here means
* fleetd has no operator-endorsed location to put a credential-bearing directory a different OS
* user must reach, so falling back to the overlay is the honest answer, not a guess.
*
* <p>Package-private (fleetd #219) so {@link OpenCodeLauncher} can reuse the exact same
* "different OS user, put it under worktreeRoot instead of java.io.tmpdir" resolution for its
* own ephemeral {@code opencode.json} directory, rather than re-reading {@code config} a second
* time with a second copy of this null/blank handling.
*/
Path memberScrubParentDir() {
FleetConfig cfg = config == null ? null : config.get();
if (cfg == null || cfg.worktreeRoot() == null || cfg.worktreeRoot().isBlank()) {
return null;
}
return Path.of(cfg.worktreeRoot());
}
/**
* fleetd #213: the configured {@code worktreeGroup:}, or {@code null} when unset/blank. Reuses
* the group {@link dev.ltms.fleet.session.Worktrees#shareWithGroup} already establishes for
* provisioned worktrees, rather than a second group key — see {@link
* #applyEnvironmentAllowListPolicy}.
*
* <p>Package-private (fleetd #219) — reused by {@link OpenCodeLauncher} alongside {@link
* #memberScrubParentDir()}; see that method's javadoc.
*/
String memberGroup() {
FleetConfig cfg = config == null ? null : config.get();
if (cfg == null || cfg.worktreeGroup() == null || cfg.worktreeGroup().isBlank()) {
return null;
}
return cfg.worktreeGroup();
}
/**
* The full kept-name set for this spawn: the profile-derived names, unioned with {@code
* memberCredentials.allow:} (CB-633 follow-up — previously ignored by this whole policy), the
@@ -1164,6 +1453,30 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Guards {@link #warnCannotShareScrubDirectory} to one WARN per launcher instance. */
private final AtomicBoolean cannotShareScrubDirWarned = new AtomicBoolean();
/**
* fleetd #213: {@code memberHerdrSocket} is configured and the member login shell IS zsh, but
* {@code worktreeRoot} and/or {@code worktreeGroup} is missing, so the generated ZDOTDIR cannot
* be placed anywhere the member's OS user can reach — {@code java.io.tmpdir} is fleetd's own
* 0700 temp dir, unreadable by another uid, which is the exact gap this ticket exists to close.
* Say so once per launcher instance, instead of either generating a directory nothing can read
* (protection theatre) or refusing to spawn (turning a degraded credential control into an
* outage for an opt-in feature).
*/
private void warnCannotShareScrubDirectory() {
if (cannotShareScrubDirWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=allow-list: memberHerdrSocket is configured and the "
+ "member login shell is zsh, but worktreeRoot and/or worktreeGroup is not "
+ "configured — the generated ZDOTDIR cannot be placed where the member's OS "
+ "user can read it (java.io.tmpdir is fleetd's own, unreadable by another uid), "
+ "so the scrub cannot be guaranteed to run. Falling back to the CB-596 sentinel "
+ "overlay. Configure both worktreeRoot and worktreeGroup to enable the "
+ "allow-list scrub under memberHerdrSocket.");
}
}
/**
* CB-633 teardown half: read the pane's scrub report (the denominator report the generated
* scrub wrote) and delete the directory. Called from {@link #stop}, which is the one funnel
@@ -1230,6 +1543,71 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private final AtomicBoolean allowListGapLogged = new AtomicBoolean();
/**
* fleetd #185 stage 2: guards {@link #warnUnknownMemberEnvironment} to one WARN per launcher
* instance, not one per spawn — the same one-per-instance shape as {@link #unprotectedGapLogged}
* and {@link #allowListGapLogged}, kept as its own flag for the same reason those two are split:
* this mode is orthogonal to which of the other two branches would otherwise have fired.
*/
private final AtomicBoolean unknownMemberEnvironmentWarned = new AtomicBoolean();
/**
* fleetd #185 stage 2: whether {@code memberHerdrSocket:} is configured, i.e. member panes run
* on a second herdr owned by a different OS user than the daemon's own process. Re-read from the
* live config on every call (same hot-reload shape as {@link #memberCredentials}), never cached,
* so a config reload takes effect on the next spawn without a restart.
*
* <p>{@link #config} is {@code null} on any call site that never threaded the full config
* through (every production {@code HerdrPeerLauncher} does; a handful of older tests do not) —
* treated the same as "not configured", which is the correct, permissive default: it is exactly
* today's single-daemon behaviour.
*
* <p>Package-private (fleetd #219) — {@link OpenCodeLauncher} reuses this same gate to decide
* where its own ephemeral {@code opencode.json} directory (site 1) and its opencode session
* discovery (site 2) may run, rather than re-deriving "is this a multi-uid fleet" a second way.
*/
boolean memberHerdrSocketConfigured() {
if (config == null) {
return false;
}
FleetConfig cfg = config.get();
return cfg != null && cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank();
}
/**
* fleetd #185 stage 2: the single replacement WARN for {@link #logCredentialGap}'s usual
* conclusions when {@code memberHerdrSocket:} is configured. {@link #hostEnvNames} (and
* everything derived from it — {@code known}/{@code allow} coverage, the allow-list scrub's
* derived set) describes the DAEMON's own environment; under this config key member panes run as
* a different OS user with a different environment entirely, so neither "every member pane
* inherits them UNBLOCKED" nor "the scrub blanks them" is evidence-backed here — both would be
* reporting on the wrong process. Logged once, names the config key, and states the honest
* conclusion: the gap for member panes is UNKNOWN, not clean, so {@code memberCredentials} cannot
* be verified from this daemon. The one count it does report is scoped explicitly to fleetd's own
* environment, never presented as if it said anything about the member's — see {@link
* #logCredentialGap}'s javadoc for why this branch exists.
*/
private void warnUnknownMemberEnvironment(FleetConfig.MemberCredentials creds) {
if (!unknownMemberEnvironmentWarned.compareAndSet(false, true)) {
return;
}
Set<String> covered = new HashSet<>(creds.known());
covered.addAll(creds.allow());
Set<String> hostNames = hostEnvNames.get();
long gapInFleetdsOwnEnv = hostNames.stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.filter(name -> !covered.contains(name))
.count();
log.warn("memberCredentials gap: memberHerdrSocket is configured, so member panes run under "
+ "a different OS user than fleetd's own process, with a different environment "
+ "entirely — fleetd has no channel to read that user's environment. {} of the "
+ "{} names in fleetd's OWN environment are credential-shaped and not on "
+ "known:/allow:, but that count describes fleetd's process, not the member "
+ "herdr's. The credential gap for member panes is UNKNOWN, not clean, and "
+ "memberCredentials cannot be verified from here.",
gapInFleetdsOwnEnv, hostNames.size());
}
/**
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
* {@code allow} is not silently allowed — it is reported. {@link #hostEnvNames} enumerates the
@@ -1258,8 +1636,22 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* (same severity, and same guard, as the deny-by-default case — a name genuinely reaching a
* member unprotected is equally serious whichever path put it there), and the names it says are
* blanked keep the INFO.
*
* <p>fleetd #185 stage 2: everything above assumes the member pane runs under the same OS user
* as the daemon, so {@link #hostEnvNames} mirrors what the pane inherits — see that field's
* javadoc. When {@code memberHerdrSocket:} is configured that assumption is false: the member
* pane runs on a second herdr owned by a <em>different</em> user, and neither conclusion below
* ("inherits them UNBLOCKED" / "the scrub blanks them") is backed by evidence about that user's
* environment. So this method checks that first and, when configured, reports the honest
* "unknown, not clean" conclusion instead — see {@link #warnUnknownMemberEnvironment}. When
* {@code memberHerdrSocket:} is absent (the default, and the only mode this host runs) this
* branch is never taken and every line below is unchanged.
*/
private void logCredentialGap(FleetConfig.MemberCredentials creds, Set<String> effectiveAllowed) {
if (memberHerdrSocketConfigured()) {
warnUnknownMemberEnvironment(creds);
return;
}
Set<String> covered = new HashSet<>(creds.known());
covered.addAll(creds.allow());
List<String> gap = hostEnvNames.get().stream()
@@ -6,6 +6,7 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.PeerHandle;
@@ -20,6 +21,8 @@ import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.BooleanSupplier;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
@@ -71,6 +74,19 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
*/
private final OpenCodeSessionDiscovery discovery;
/**
* fleetd #175: notified when {@link SessionAwareHandle} detects, on the same late-resolve
* read that discovers the session id, that the live opencode session is running a DIFFERENT
* model than the profile requested — opencode does not fail on an unknown {@code -m}, it
* silently falls back to a default (potentially paid) model. Reused exactly as
* {@code CompletionResolver}'s {@code BACKEND_EXHAUSTED} path uses it: this launcher supplies
* only the herdr terminal id and a reason string; mapping target → session → profile →
* credential stays entirely the sink's job (see {@code Fleetd.main}'s wiring). Defaults to
* {@link ExhaustionSink#none()} for a caller (an older constructor, or a test not exercising
* this) that opts out.
*/
private final ExhaustionSink exhaustionSink;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
@@ -130,9 +146,25 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, spawnReadyPollMs,
fleet, memberCredentials, config, ExhaustionSink.none());
}
/**
* Production constructor, plus the live config for URI environment exclusions and the fleetd
* #175 model-mismatch quarantine sink.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config,
ExhaustionSink exhaustionSink) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials, config);
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials, config,
exhaustionSink);
}
/**
@@ -180,6 +212,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
this.exhaustionSink = ExhaustionSink.none();
}
/**
@@ -205,17 +238,72 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, nowMillis, sleeper,
configRoot, discoveryRoot, fleet, memberCredentials, config, ExhaustionSink.none());
}
/**
* Full testability constructor, plus the live config for URI environment exclusions and the
* fleetd #175 model-mismatch quarantine sink — the constructor a test drives directly to
* observe {@link ExhaustionSink#onExhausted} without going through {@code Fleetd.main}'s
* wiring.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot, Path discoveryRoot,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config,
ExhaustionSink exhaustionSink) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, null, config);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
this.exhaustionSink = exhaustionSink == null ? ExhaustionSink.none() : exhaustionSink;
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
/**
* The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir) —
* always FLEETD's OWN {@code user.home}, whichever OS user runs the daemon.
*
* <p><b>fleetd #219 site 2 — a decision, not a patch.</b> Under {@code memberHerdrSocket:} the
* member pane runs as a <em>different</em> OS user, and opencode writes {@code opencode.db}
* under <em>that</em> user's {@code $HOME}, not fleetd's. Scanning fleetd's own {@code
* user.home} is therefore looking in the wrong place — a wrong-LOCATION failure, not a
* wrong-PERMISSION one like site 1, and it fails quietly: {@link
* SessionAwareHandle#agentSessionId()} would keep returning {@code null} forever, which reads
* as "opencode does not support resume" rather than "fleetd looked in the wrong home." fleetd
* #209 is the reason that silence is unacceptable.
*
* <p>Three ways to close the gap were weighed:
* <ol>
* <li><b>Make the member's home configurable.</b> Correct in principle, but this ticket's
* scope is the two existing call sites, not a new config key — {@code memberHerdrSocket}
* already carries the second herdr's socket path, not its user's home, and inventing a
* parallel key here without also wiring it through discovery's actual callers is a
* half-shipped feature (the exact shape CB-596/CB-611 warn against).</li>
* <li><b>Derive it</b> (e.g. from {@code worktreeRoot}'s owner, or {@code getent passwd}).
* Rejected: nothing in this codebase resolves a Unix username to a home directory today,
* and guessing wrong would silently point discovery at a THIRD wrong location — worse
* than the current gap, because it would look like it should work.</li>
* <li><b>Declare discovery unavailable</b> under {@code memberHerdrSocket}, and say so once,
* loudly, instead of scanning a directory that structurally cannot hold the answer.</li>
* </ol>
*
* <p>Option 3 is taken — the one this ticket says to default to when unsure. {@link
* OpenCodeLauncher#spawn} routes {@link SessionAwareHandle#agentSessionId()} through {@link
* HerdrPeerLauncher#memberHerdrSocketConfigured()} before ever calling {@link
* OpenCodeSessionDiscovery#sessionIdForDirectory}, so under {@code memberHerdrSocket} the
* database at this root is never even opened, and one WARN per launcher instance names the gap
* instead of the {@code null} return reading as "unsupported." Capability advertising is
* unaffected: {@link #capabilities()} always includes {@code SESSION_RESUME}, since {@code
* memberHerdrSocket} absent (today's only live mode) is unchanged by this decision.
*/
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
@@ -331,7 +419,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
try {
Path dir = Files.createTempDirectory(configRoot, "fleetd-opencode-");
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
@@ -402,6 +490,14 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
if (memberHerdrSocketConfigured()) {
// fleetd #219: the same "different OS user" gap fleetd #213 closed for the ZDOTDIR
// scrub — share read-only with worktreeGroup rather than leaving the directory under
// fleetd's own 0700 java.io.tmpdir, where the member's OS user could not even
// traverse it. memberGroup() cannot be null here: configParentDir() above already
// refused this spawn if either worktreeRoot or worktreeGroup was missing.
EnvAllowListScrub.shareWithGroup(dir, memberGroup());
}
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
@@ -409,6 +505,57 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
}
/**
* fleetd #219 site 1: where {@link #writeConfig} creates its per-spawn directory.
*
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only mode): byte-identical to before this
* fix — always {@link #configRoot} (defaults to {@code java.io.tmpdir}, fleetd's own
* process).</li>
* <li>{@code memberHerdrSocket} PRESENT: {@code java.io.tmpdir} is fleetd's own per-user temp
* dir (mode {@code 0700} on macOS) — the member pane runs as a DIFFERENT OS user under
* this config key and cannot even traverse it, so the directory holding {@code
* opencode.json} (which tells the member where the bridge MCP is) and the member charter
* would be unreadable to the very process it is written for. The directory instead goes
* under {@code worktreeRoot}, shared read-only with {@code worktreeGroup} via {@link
* EnvAllowListScrub#shareWithGroup} — the SAME mechanism fleetd #213 built for the ZDOTDIR
* scrub, reused here rather than duplicated (see {@link
* HerdrPeerLauncher#memberScrubParentDir()}).</li>
* </ul>
*
* <p><b>Unlike the ZDOTDIR scrub, a missing {@code worktreeRoot}/{@code worktreeGroup} here
* REFUSES the spawn instead of degrading.</b> The ZDOTDIR scrub is a credential CONTROL: a
* degraded control (CB-596's sentinel overlay) is still worth having. This config file is not a
* control — it is the ONLY way the member learns where the bridge MCP lives. Writing it
* somewhere the member cannot read would not degrade anything; it would spawn a member that
* occupies a pane and never becomes deliverable, since {@code fleet_send} waits ~60s on the
* readiness gate and then fails with nothing pointing at a temp directory as the cause. Refusing
* up front, with a message that names the missing config key, is the honest failure — an
* undeliverable member is not a working spawn either way, so nothing is lost by refusing loudly
* instead of failing silently later.
*
* @throws IllegalStateException when {@code memberHerdrSocket} is configured but {@code
* worktreeRoot} and/or {@code worktreeGroup} is not
*/
private Path configParentDir() {
if (!memberHerdrSocketConfigured()) {
return configRoot;
}
Path root = memberScrubParentDir();
String group = memberGroup();
if (root == null || group == null) {
throw new IllegalStateException("memberHerdrSocket is configured, so opencode's config "
+ "directory (opencode.json + member charter) must be placed where the member's "
+ "OS user can read it — worktreeRoot, shared via worktreeGroup — but "
+ (root == null ? "worktreeRoot" : "worktreeGroup") + " is not configured. "
+ "Refusing to spawn rather than write a config the member cannot read: that "
+ "member would occupy a pane and never become deliverable, with nothing "
+ "pointing at the real cause. Configure both worktreeRoot and worktreeGroup to "
+ "enable opencode member spawns under memberHerdrSocket.");
}
return root;
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
@@ -517,11 +664,21 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/** One WARN per launcher instance for the fleetd #219 site-2 discovery-unavailable gap. */
private final AtomicBoolean discoveryUnavailableWarned =
new AtomicBoolean();
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
// fleetd #175: the same profile config buildLaunch resolved for this spawn (requireProfile
// is deterministic on req.profileName(), so re-resolving here costs a map lookup, not a
// second decision) — SessionAwareHandle needs cfg.model() to know what THIS session should
// be running.
FleetConfig.Profile cfg = requireProfile(req.profileName());
return new SessionAwareHandle(inner, discovery, effectiveCwd(req), cfg,
this::memberHerdrSocketConfigured, discoveryUnavailableWarned, exhaustionSink);
}
/**
@@ -531,16 +688,37 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* are untouched — only the opencode-specific identity answer is added. {@code sessionName()}
* stays null: opencode has no display-name seam, so the logical name lives only in the bridge's
* roster (see the SESSION_NAME capability).
*
* <p>fleetd #175: {@link #agentSessionId()} is also where the model-mismatch check lives (see
* {@link #checkModelMatch()}) — it is the one method the real late-resolve path
* ({@code SessionManager.resolveAgentSessionId}, via {@code get}/{@code rosterResolved}/
* {@code release}) actually calls, and only while the session id is still unknown. Putting the
* check anywhere else risks repeating PR #203's mistake: a check that runs before opencode has
* written the row it needs, and so never fires.
*/
private static final class SessionAwareHandle implements PeerHandle {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
private final FleetConfig.Profile cfg;
private final BooleanSupplier discoveryUnavailable;
private final AtomicBoolean discoveryUnavailableWarned;
private final ExhaustionSink exhaustionSink;
/** CAS'd true the first (and only) time a model mismatch is reported for this handle. */
private final AtomicBoolean modelMismatchReported = new AtomicBoolean();
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd,
FleetConfig.Profile cfg,
BooleanSupplier discoveryUnavailable,
AtomicBoolean discoveryUnavailableWarned,
ExhaustionSink exhaustionSink) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
this.cfg = cfg;
this.discoveryUnavailable = discoveryUnavailable;
this.discoveryUnavailableWarned = discoveryUnavailableWarned;
this.exhaustionSink = exhaustionSink;
}
@Override
@@ -565,11 +743,84 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public String agentSessionId() {
// fleetd #219 site 2: under memberHerdrSocket the member pane runs as a different OS
// user, so opencode.db lives under THAT user's $HOME, not the one discoveryRoot was
// built from (see OpenCodeLauncher#defaultDiscoveryRoot's javadoc for the full
// reasoning). Scanning fleetd's own $HOME under that config would only ever find "no
// row" and read as "resume unsupported" — declare it unavailable instead, once, loudly.
if (discoveryUnavailable.getAsBoolean()) {
if (discoveryUnavailableWarned.compareAndSet(false, true)) {
log.warn("opencode session discovery unavailable: memberHerdrSocket is "
+ "configured, so opencode's on-disk session database lives under the "
+ "MEMBER's own $HOME, not fleetd's ({}) — agentSessionId will stay null "
+ "for every opencode member under this config, and SESSION_RESUME "
+ "cannot be honored (fleetd #209/#219).",
System.getProperty("user.home"));
}
return null;
}
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
// appeared).
return discovery.sessionIdForDirectory(cwd);
String id = discovery.sessionIdForDirectory(cwd);
// fleetd #175: check on the SAME tick — while the caller (SessionManager's late-resolve
// step) is still re-polling because the id is unknown, the row this id came from (once
// it exists) is exactly the row that also carries the actual model. Once id resolves,
// the caller stops calling agentSessionId() for this session, so this is naturally a
// once-only check that happens right when the row first appears.
checkModelMatch();
return id;
}
/**
* Verify the live opencode session is running the model {@link #cfg} requested (fleetd
* #175) and, on a real mismatch, log an ERROR and quarantine through {@link
* #exhaustionSink}. A no-op when there is nothing to compare against — no model configured,
* already reported once for this handle, or the actual model is still UNKNOWN (no row yet,
* unreadable database, or unparseable evidence). UNKNOWN must never be treated as a
* mismatch: that is the single most important safety rule here — a false positive would
* quarantine a perfectly working profile's credential.
*/
private void checkModelMatch() {
if (modelMismatchReported.get() || cfg.model() == null || cfg.model().isBlank()) {
return;
}
OpenCodeSessionDiscovery.ActualModel actual = discovery.actualModelForDirectory(cwd);
if (actual == null) {
return; // UNKNOWN evidence — never a mismatch
}
String[] requestedParts = splitProviderModel(cfg.model());
String requestedProvider = requestedParts == null ? null : requestedParts[0];
String requestedId = requestedParts == null ? cfg.model() : requestedParts[1];
boolean idMatches = requestedId.equals(actual.id());
// Compare the provider ONLY when BOTH sides have one. requestedProvider == null covers
// a bare profile model with no "/" — the profile never asked for a specific provider.
// actual.provider() == null covers opencode's model JSON having an id but no providerID
// (a real shape parseModel accepts) — that is missing evidence, not a contradiction, and
// acceptance rule 4 says missing evidence is UNKNOWN, never a mismatch. Narrowing the
// provider comparison this way keeps the id comparison (the part that actually caught the
// xf bug) fully intact — a genuine id mismatch is still caught either way (fleetd #175
// review round 2).
boolean providerMatches = requestedProvider == null || actual.provider() == null
|| requestedProvider.equals(actual.provider());
if (idMatches && providerMatches) {
return;
}
if (!modelMismatchReported.compareAndSet(false, true)) {
return; // another thread already reported this exact mismatch
}
String actualDisplay = actual.provider() == null
? actual.id() : actual.provider() + "/" + actual.id();
log.error("opencode profile '{}' requested model '{}' but the live session is actually "
+ "running '{}' — opencode does not fail on an unknown -m, it silently "
+ "falls back to a default model, which may be a PAID credential "
+ "(fleetd #175); quarantining this profile's credential",
cfg.profile(), cfg.model(), actualDisplay);
exhaustionSink.onExhausted(delegate.terminalId(),
"opencode model mismatch: profile '" + cfg.profile() + "' requested '"
+ cfg.model() + "' but the live session is running '" + actualDisplay
+ "' (fleetd #175)");
}
@Override
@@ -1,9 +1,12 @@
package dev.ltms.fleet.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.sqlite.SQLiteConfig;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.sql.Connection;
@@ -44,6 +47,18 @@ import java.util.concurrent.atomic.AtomicBoolean;
final class OpenCodeSessionDiscovery {
private static final Logger log = LoggerFactory.getLogger(OpenCodeSessionDiscovery.class);
private static final ObjectMapper MAPPER = new ObjectMapper();
/**
* The model opencode actually ran a session on, parsed from the {@code session.model} JSON
* column (fleetd #175). {@code id} is never null/blank on a non-null {@code ActualModel} —
* {@link #parseModel} returns {@code null} instead when {@code id} cannot be determined, so a
* caller only ever sees a fully-known record or {@code null} (UNKNOWN). {@code provider} may
* still be {@code null} on its own when the profile that requested the session named no
* provider prefix, or opencode's JSON omitted {@code providerID}.
*/
record ActualModel(String provider, String id) {
}
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final Path databasePath;
@@ -117,4 +132,73 @@ final class OpenCodeSessionDiscovery {
storageRoot, directory);
return null;
}
/**
* The model opencode actually ran the {@code directory}'s most-recent session on (fleetd
* #175), read from the same row {@link #sessionIdForDirectory} matches — but via its own
* query and its own connection, deliberately kept independent so a database whose schema
* predates the {@code model} column (or any other read failure on this column alone) can
* never take {@link #sessionIdForDirectory}'s id resolution down with it. That would be a
* regression of the id-resolution feature #209 shipped; this method degrades on its own.
*
* <p>Never throws, and every failure mode — no matching row, a missing/unreadable database, a
* missing {@code model} column, a null/blank {@code model} value, or JSON that does not parse
* into {@code {"id": "...", "providerID": "..."}} with a non-blank {@code id} — resolves to
* {@code null}. That is UNKNOWN evidence, not a mismatch signal: the caller must never
* quarantine a profile on the strength of a {@code null} here.
*
* @param directory the worker's cwd, as resolved for this spawn
* @return the actual model, or {@code null} when unknown
*/
ActualModel actualModelForDirectory(String directory) {
if (directory == null || directory.isBlank()) {
return null;
}
if (!Files.isRegularFile(databasePath)) {
// sessionIdForDirectory already WARNs once (shared warnedMissingDatabase) for this
// exact condition — do not double-log it here.
return null;
}
String sql = "SELECT model FROM session WHERE directory = ? ORDER BY time_updated DESC LIMIT 1";
try (Connection connection = openReadOnly();
PreparedStatement statement = connection.prepareStatement(sql)) {
statement.setString(1, directory);
try (ResultSet rows = statement.executeQuery()) {
if (rows.next()) {
return parseModel(rows.getString("model"));
}
}
} catch (SQLException e) {
// Locked/corrupt database, or a `model` column this schema version does not have —
// never fatal, and never a mismatch signal. See the class doc above.
log.debug("opencode session model unreadable at {}: {}", databasePath, e.toString());
return null;
}
return null;
}
/**
* Parse opencode's {@code model} column — {@code {"id":"...","providerID":"..."}} — into an
* {@link ActualModel}, or {@code null} when {@code json} is null/blank, is not valid JSON, or
* parses without a non-blank {@code id}. {@code providerID} may be absent; that alone does not
* make the record unknown, since a caller comparing against a profile with no provider prefix
* never looks at it.
*/
private static ActualModel parseModel(String json) {
if (json == null || json.isBlank()) {
return null;
}
try {
JsonNode node = MAPPER.readTree(json);
String id = node.path("id").asText(null);
if (id == null || id.isBlank()) {
return null;
}
String provider = node.path("providerID").asText(null);
return new ActualModel(provider, id);
} catch (IOException e) {
log.debug("opencode session model JSON unparseable: {}", e.toString());
return null;
}
}
}
@@ -9,6 +9,7 @@ import dev.ltms.fleet.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
@@ -167,6 +168,12 @@ public final class MessageService {
private static final class Task {
private final String ticket;
private final String target;
/**
* When this task was created (#137 fix): the tiebreaker for which of several open tasks on
* one target gets a recovered reply in {@link #abandon} — the oldest, since it is the one
* that has been waiting longest.
*/
private final long createdNanos;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
/**
* When {@link #future} resolved, or {@code null} while it is still pending — the clock
@@ -184,6 +191,7 @@ public final class MessageService {
private Task(String ticket, String target, LongSupplier nowNanos) {
this.ticket = ticket;
this.target = target;
this.createdNanos = nowNanos.getAsLong();
future.whenComplete((reply, ex) -> completedNanos = nowNanos.getAsLong());
}
}
@@ -375,6 +383,19 @@ public final class MessageService {
* {@link Rendezvous#resolveQuestion} must keep today's {@code NO_WAITER} behaviour — questions
* are interactive and must never be queued.
*
* <p><strong>Ambiguous match also falls to the inbox.</strong> {@link #askAnsweredAsyncTasks}
* cannot actually return more than one entry today (see its own javadoc for why — in short,
* {@link #hasAsyncQuestion} keeps a target BUSY, so no second task can reach this state, for as
* long as an earlier one's {@code turnId} is still stamped). That is an emergent guarantee from
* two other facts, not one this method enforces, so this branch stays in as defence in depth
* rather than being removed as dead code: if it ever weakens, returning whichever candidate a
* {@code ConcurrentHashMap} iteration reaches first would let a genuine reply complete the
* <em>wrong</em> ticket — silently handing the lead something that reads like a correct answer to
* a delegation the worker never touched, which is worse than a failure because the lead acts on
* it. When more than one candidate exists, guessing is not safe: fall back to the inbox exactly
* as the zero-candidate case does, and let {@link #abandon} apply the eventual recovery
* deterministically instead.
*
* @return always {@code true} — the reply resolved a live send, completed a parked ticket, or
* was queued
*/
@@ -392,13 +413,21 @@ public final class MessageService {
// FAILED with a misleading "session released before it replied" reason, even though the reply
// had, in fact, arrived. Completing the matching ticket directly here means fleet_poll{ticket}
// sees the real reply instead.
Task orphan = askAnsweredAsyncTask(session);
if (orphan != null && orphan.future.complete(new Reply(Outcome.REPLIED, content))) {
if (orphan.turnId != null) {
asyncTasksByTurn.remove(orphan.turnId, orphan);
List<Task> candidates = askAnsweredAsyncTasks(session);
if (candidates.size() == 1) {
Task orphan = candidates.get(0);
if (orphan.future.complete(new Reply(Outcome.REPLIED, content))) {
if (orphan.turnId != null) {
asyncTasksByTurn.remove(orphan.turnId, orphan);
}
count(FleetMetrics.REPLIES, "path", "async-recovered");
return true; // the ticket itself took it — no inbox stranding at all
}
count(FleetMetrics.REPLIES, "path", "async-recovered");
return true; // the ticket itself took it — no inbox stranding at all
} else if (candidates.size() > 1) {
List<String> tickets = candidates.stream().map(t -> t.ticket).toList();
log.warn("reply from {} matches {} open async tickets {} — cannot tell which one it "
+ "answers, queuing to the inbox instead of guessing", session, candidates.size(),
tickets);
}
inbox.publish(session, UUID.randomUUID().toString(), content);
// CB-640: record the stranding itself (not just the reply text) so fleet health can see a
@@ -414,21 +443,38 @@ public final class MessageService {
}
/**
* The still-open async task on {@code target} whose {@code fleet_ask} was already answered — its
* {@link Task#turnId} is stamped but its {@link Task#question} was cleared by {@link #answer} —
* yet whose future is not resolved yet (#137). {@code null} if no such task exists, including the
* Every still-open async task on {@code target} whose {@code fleet_ask} was already answered —
* its {@link Task#turnId} is stamped but its {@link Task#question} was cleared by {@link #answer}
* — yet whose future is not resolved yet (#137). Empty if no such task exists, including the
* common case where {@code target}'s worker never used {@code fleet_ask} at all (a task that was
* never asked has {@code turnId == null}, so it can never match here and only ever completes
* through the ordinary rendezvous fast path in {@link #reply}).
*
* <p><strong>Returns at most one entry today — verified, not assumed.</strong> {@link #send}
* refuses to open a waiter on {@code target} while {@link #hasAsyncQuestion} is true, and that
* check matches ANY task whose {@code turnId} is still stamped in {@code asyncTasksByTurn} —
* not only while its question is still open. {@link #answer} deliberately leaves that stamp in
* place ({@code clearAsyncQuestion(turnId, false)}) until the resumed turn's own future actually
* resolves, at which point {@link #finishAsyncTask} both removes the stamp AND completes that
* task's future in the same call. So a second task can never reach "{@code turnId} stamped, future
* still open" — the exact pair this method matches on — while a first one already holds it: by
* the time the stamp is gone, so is the eligibility. This is an emergent property of those two
* facts holding together, not something this method (or its callers) enforces on its own — flip
* {@code forgetTurn} to {@code true} in that one {@link #answer} call and it silently stops being
* true, with nothing left to fail loudly. The callers below still handle "more than one" as
* defence in depth against exactly that, not because they exercise it today: {@link #reply}
* treats it as unresolvable and falls back to the inbox; {@link #abandon} would pick the oldest
* deterministically (its own {@code matching} list has no such guarantee — see its javadoc).
*/
private Task askAnsweredAsyncTask(String target) {
private List<Task> askAnsweredAsyncTasks(String target) {
List<Task> candidates = new ArrayList<>();
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null && task.turnId != null
&& !task.future.isDone()) {
return task;
candidates.add(task);
}
}
return null;
return candidates;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
@@ -473,16 +519,63 @@ public final class MessageService {
* outcome is counted, so a torn-down delegation stops being invisible to {@code /metrics}.
*
* <p><strong>#137 defence in depth.</strong> {@link #reply} already hands a worker's real
* {@code fleet_reply} straight to the async ticket it belongs to whenever one is still parked
* waiting for it (see {@link #askAnsweredAsyncTask}), so by the time a session is released its
* tasks are normally already resolved — this loop's {@code complete} calls are then harmless
* no-ops (a {@link CompletableFuture} can only resolve once). But should some other path someday
* strand a reply in the inbox without completing its ticket, checking
* {@code fleet_reply} straight to the async ticket it belongs to whenever exactly one is still
* parked waiting for it (see {@link #askAnsweredAsyncTasks}), so by the time a session is
* released its tasks are normally already resolved — this loop's {@code complete} calls are then
* harmless no-ops (a {@link CompletableFuture} can only resolve once). But should some other path
* someday strand a reply in the inbox without completing its ticket, checking
* {@link #hasStrandedReply(String)} here — before ever writing a failure — means a torn-down
* session whose worker in fact replied is still reported {@code REPLIED} with that reply's own
* text, never the misleading "the worker session was released before it replied" (which also
* means the snapshot/worktree recovery hint that follows it never prints once a reply exists).
*
* <p><strong>At most one task gets the recovered reply — and here, unlike {@link #reply}'s
* {@link #askAnsweredAsyncTasks}, {@code matching.size() >= 2} alone is reachable today.</strong>
* This method's {@code matching} filter has no {@code turnId != null} requirement, so it matches
* any plain (never-asked) open task too — and {@link #sendAsync} does not limit a target to one
* of those: a second {@code fleet_send{wait:false}} at a target that is still busy returns its own
* ticket immediately and simply parks its {@link #send} behind the target's session lock for up
* to {@link #ASYNC_TIMEOUT_MS}, exactly as {@code abandonFailsEveryPendingAsyncTicketForTheReleasedTarget}
* already proves. Before this fix, the loop below drained the strand once and then reused that
* same {@code Reply} for <em>every</em> task it walked past — so two open tasks really did both
* complete {@code REPLIED} with the same text (see the pre-fix loop in commit 97f6c33's parent).
* A stranded reply is one worker answer, so it can settle at most one open task on this target —
* never every open task, and never a guess. When more than one task is still open here, the
* recovered reply goes to the <em>oldest</em> (lowest {@link Task#createdNanos}) — it has been
* waiting longest, so it is the one most likely to be what the reply actually answers. Every
* other open task keeps the ordinary {@code WORKER_FAILED} path it would take without a stranded
* reply at all.
*
* <p><strong>{@code matching.size() >= 2} together with {@code hadStrandedReply} is a different
* question, and today it is defence in depth rather than a path this codebase's public API can
* drive.</strong> This class has exactly two sites that ever acquire a target's entry in
* {@code sessionLocks} — {@link #send} and {@link #answer} — and both open a {@link Rendezvous}
* waiter for that same target as the very first thing they do after acquiring the lock, then hold
* lock and waiter together for the rest of their critical section ({@link #send} also clears
* {@link #strandedReplies} right there, the instant it opens its waiter — before it ever enqueues
* delivery). So "the session lock is held" and "a live waiter is open for it" are the same fact
* throughout this class, and {@link #reply}'s fast path always resolves a currently-open waiter
* directly rather than stranding. The two facts this method wants therefore cannot be produced
* side by side: while the lock is held, a real reply resolves the open waiter directly and never
* reaches {@link #strandedReplies}; the instant the lock is free, any parked matching task's own
* {@link #send} that is scheduled next wins it and, by opening its waiter, clears the strand again
* before this method ever runs. There is no way to hold that lock open-but-unaccepted from outside
* {@link #send}/{@link #answer} to freeze a window in between. Constructing both facts at once
* through {@code sendAsync}/{@code reply}/{@code ask}/{@code answer} would need a race against
* virtual-thread scheduling, not a deterministic sequence — so the oldest-wins code below stays as
* defence in depth against a regression to that mechanism (e.g. clearing {@link #strandedReplies}
* on a narrower condition than "any acceptance"), not because today's test suite exercises the
* conjunction. {@code matching.size() >= 2} alone, without a strand, is exactly what
* {@code abandonFailsEveryPendingAsyncTicketForTheReleasedTarget} already covers.
*
* <p><strong>Do not "fix" that gap with a test that reaches past this class.</strong> A test can
* build both facts by calling {@link Rendezvous#close} itself on the waiter an accepted
* {@link #send} is still blocked on: the send keeps the lock, no waiter is registered any more, a
* second parked task stays open, and the next {@link #reply} then strands. That was checked, and
* such a test does go red against the pre-fix loop. But it only goes red because it broke the
* lock-and-waiter invariant above from outside — no caller of this class ever does that — so it
* pins a state production cannot reach, and would read to the next person as if it could.
*
* @return true if a live waiter or an async task was failed (never true for one recovered as a
* reply — see the note above)
*/
@@ -494,21 +587,39 @@ public final class MessageService {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
Reply recovered = null; // lazily drained at most once, only if a task actually needs it
List<Task> matching = new ArrayList<>();
for (Task task : tasks.values()) {
if (!target.equals(task.target) || task.question != null || task.future.isDone()) {
continue;
if (target.equals(task.target) && task.question == null && !task.future.isDone()) {
matching.add(task);
}
if (hadStrandedReply && recovered == null) {
recovered = recoverStrandedReply(target);
}
Task recoveryTask = null;
if (hadStrandedReply && !matching.isEmpty()) {
recoveryTask = matching.get(0);
for (Task candidate : matching) {
if (candidate.createdNanos < recoveryTask.createdNanos) {
recoveryTask = candidate;
}
}
Reply outcome = recovered != null ? recovered : new Reply(Outcome.WORKER_FAILED, reason);
}
Reply recovered = recoveryTask != null ? recoverStrandedReply(target) : null;
for (Task task : matching) {
boolean isRecovery = task == recoveryTask && recovered != null;
Reply outcome = isRecovery ? recovered : new Reply(Outcome.WORKER_FAILED, reason);
if (task.future.complete(outcome)) {
if (outcome.outcome() == Outcome.WORKER_FAILED) {
asyncFailed = true;
} else if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
} else if (isRecovery) {
// The recovered reply was already drained out of the inbox, but this task resolved
// through another path (e.g. a concurrent reply() or a second abandon() racing this
// one) between us choosing it and completing it here. Put the reply back rather than
// lose it silently — it may still belong to some other still-open task, or the next
// caller that drains this target's inbox.
inbox.publish(target, UUID.randomUUID().toString(), recovered.text());
}
}
if (failed) {
@@ -523,6 +634,10 @@ public final class MessageService {
* stranding fact raced away, e.g. a lead's own {@code fleet_poll} on the raw session already
* drained it first). When more than one message is queued, only the newest is the worker's actual
* final answer ({@link #drainReplies} returns them oldest-first).
*
* <p>This does drain (removes the messages from the inbox) before the caller knows whether the
* task it is recovering for will actually accept them — {@link #abandon} is the one that puts a
* reply back if its {@code complete} call turns out to lose the race.
*/
private Reply recoverStrandedReply(String target) {
var messages = drainReplies(target);
@@ -9,8 +9,10 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collection;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
@@ -68,6 +70,12 @@ public final class ReplyPushLoop {
static final String QUESTIONS_NUDGE_FORMAT =
"%d workers are paused on a question — run fleet_poll(ticket=...) for each, then answer "
+ "with fleet_send(turnId=..., content=...): %s";
static final String BACKEND_INCIDENT_NUDGE_FORMAT =
"Backend credential %s is cooling for %d remaining seconds; affected profiles: %s; "
+ "affected workers: %s. Run fleet_list to see more.";
static final String BACKEND_TARGET_UNMAPPED_NUDGE_FORMAT =
"A backend error on worker %s could not be mapped to a credential (%s) — no cool-off "
+ "was applied. Run fleet_list to check that worker.";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
@@ -93,6 +101,13 @@ public final class ReplyPushLoop {
* how depleted an older, still-open question's count is.
*/
private final ConcurrentHashMap<String, PendingQuestion> pendingQuestions = new ConcurrentHashMap<>();
/** Backend incidents awaiting one successful delivery, keyed by incident id and owning lead. */
private final ConcurrentHashMap<IncidentLead, PendingIncident> pendingIncidents = new ConcurrentHashMap<>();
/** Incident/lead pairs already delivered. They make repeated incident reports one-shot. */
private final Set<IncidentLead> deliveredIncidents = ConcurrentHashMap.newKeySet();
private final ConcurrentHashMap<UnmappedTargetLead, PendingUnmappedTarget> pendingUnmappedTargets =
new ConcurrentHashMap<>();
private final Set<UnmappedTargetLead> deliveredUnmappedTargets = ConcurrentHashMap.newKeySet();
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets and/or questions). */
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
@@ -178,7 +193,20 @@ public final class ReplyPushLoop {
* tracked per question, not per lead per source).
*/
private record PendingQuestion(String turnId, String ticket, String target, String lead,
String question, int nudgeCount) {
String question, int nudgeCount) {
}
private record IncidentLead(String incidentId, String lead) {
}
private record PendingIncident(IncidentLead key, String credential, List<String> profiles,
List<String> targets, int remainingCoolOffSeconds, int nudgeCount) {
}
private record UnmappedTargetLead(String target, String reason, String lead) {
}
private record PendingUnmappedTarget(UnmappedTargetLead key, int nudgeCount) {
}
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
@@ -192,6 +220,24 @@ public final class ReplyPushLoop {
.collect(Collectors.toUnmodifiableSet());
}
private List<PendingIncident> pendingIncidentsFor(String lead) {
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
private Set<IncidentLead> pendingIncidentKeysFor(String lead) {
return pendingIncidentsFor(lead).stream().map(PendingIncident::key)
.collect(Collectors.toUnmodifiableSet());
}
private List<PendingUnmappedTarget> pendingUnmappedTargetsFor(String lead) {
return pendingUnmappedTargets.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
private Set<UnmappedTargetLead> pendingUnmappedTargetKeysFor(String lead) {
return pendingUnmappedTargetsFor(lead).stream().map(PendingUnmappedTarget::key)
.collect(Collectors.toUnmodifiableSet());
}
/**
* The reply-source reminder count {@link #decide} should see for {@code lead} on this tick:
* the <em>minimum</em> nudge count among the reply targets currently pending for it (CB-598).
@@ -236,6 +282,15 @@ public final class ReplyPushLoop {
return min == Integer.MAX_VALUE ? 0 : min;
}
private int minIncidentNudgeCountFor(String lead) {
return pendingIncidentsFor(lead).stream().mapToInt(PendingIncident::nudgeCount).min().orElse(0);
}
private int minUnmappedTargetNudgeCountFor(String lead) {
return pendingUnmappedTargetsFor(lead).stream().mapToInt(PendingUnmappedTarget::nudgeCount)
.min().orElse(0);
}
/**
* Pure decision function: examine everything pending for {@code lead} — reply targets and
* tickets alike — and return what the loop should do.
@@ -272,17 +327,35 @@ public final class ReplyPushLoop {
* @param questionReminderCount the lowest nudge count among questions open for this lead
*/
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount) {
return decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
minIncidentNudgeCountFor(lead));
}
/** As above, with backend incidents as a fourth, independently bounded source. */
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount,
int incidentReminderCount) {
return decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
incidentReminderCount, minUnmappedTargetNudgeCountFor(lead));
}
/** As above, with unmapped backend targets as a fifth, independently bounded source. */
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount,
int incidentReminderCount, int unmappedTargetReminderCount) {
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
boolean hasQuestionWork = !pendingQuestionTurnIdsFor(lead).isEmpty();
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork) {
boolean hasIncidentWork = !pendingIncidentKeysFor(lead).isEmpty();
boolean hasUnmappedTargetWork = !pendingUnmappedTargetKeysFor(lead).isEmpty();
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork && !hasIncidentWork && !hasUnmappedTargetWork) {
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
return Action.STOP;
}
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
boolean questionEligible = hasQuestionWork && questionReminderCount < maxReminders;
if (!replyEligible && !ticketEligible && !questionEligible) {
boolean incidentEligible = hasIncidentWork && incidentReminderCount < maxReminders;
boolean unmappedTargetEligible = hasUnmappedTargetWork && unmappedTargetReminderCount < maxReminders;
if (!replyEligible && !ticketEligible && !questionEligible && !incidentEligible && !unmappedTargetEligible) {
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
maxReminders, lead);
countNudge("exhausted");
@@ -394,6 +467,48 @@ public final class ReplyPushLoop {
pendingQuestions.remove(turnId);
}
/**
* Queue a one-shot backend credential outage notice for every distinct lead that owns an
* affected worker. A successful injection records its {@code (incidentId, lead)} key, so a
* repeat report never reminds that lead again.
*/
public void onBackendIncident(String incidentId, Collection<String> targets, String credential,
Collection<String> profiles, int remainingCoolOffSeconds) {
Map<String, List<String>> targetsByLead = new ConcurrentHashMap<>();
for (String target : targets) {
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.warn("push: backend incident {} has no known lead for target {}", incidentId, target);
continue;
}
targetsByLead.computeIfAbsent(lead.get(), _ -> new ArrayList<>()).add(target);
}
List<String> profileNames = profiles.stream().sorted().toList();
for (var entry : targetsByLead.entrySet()) {
IncidentLead key = new IncidentLead(incidentId, entry.getKey());
if (deliveredIncidents.contains(key)) continue;
pendingIncidents.putIfAbsent(key, new PendingIncident(key, credential, profileNames,
entry.getValue().stream().sorted().toList(), remainingCoolOffSeconds, 0));
startOrCoalesce(entry.getKey());
}
}
/**
* Tell the owning lead that a classified backend error could not be tied to a credential.
* Without an owning lead, emit a warning because no control can act on the target.
*/
public void onBackendTargetUnmapped(String target, String reason) {
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.warn("push: backend target {} could not map to a credential: {}", target, reason);
return;
}
UnmappedTargetLead key = new UnmappedTargetLead(target, reason, lead.get());
if (deliveredUnmappedTargets.contains(key)) return;
pendingUnmappedTargets.putIfAbsent(key, new PendingUnmappedTarget(key, 0));
startOrCoalesce(lead.get());
}
// --- the schedule ----------------------------------------------------------------------------
/** Start a reminder schedule for {@code lead}, or join the one already running. */
@@ -425,18 +540,25 @@ public final class ReplyPushLoop {
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
Set<String> questionsBefore = pendingQuestionTurnIdsFor(lead);
Set<IncidentLead> incidentsBefore = pendingIncidentKeysFor(lead);
Set<UnmappedTargetLead> unmappedTargetsBefore = pendingUnmappedTargetKeysFor(lead);
int replyReminderCount = minReplyNudgeCountFor(lead);
int ticketReminderCount = minTicketNudgeCountFor(lead);
int questionReminderCount = minQuestionNudgeCountFor(lead);
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
int incidentReminderCount = minIncidentNudgeCountFor(lead);
int unmappedTargetReminderCount = minUnmappedTargetNudgeCountFor(lead);
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
incidentReminderCount, unmappedTargetReminderCount);
switch (action) {
case INJECT -> {
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount,
incidentReminderCount, unmappedTargetReminderCount);
scheduleNext(lead);
}
// Re-check after the configured backoff; the lead may become injectable soon.
case WAIT_BUSY -> scheduleNext(lead);
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore);
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, incidentsBefore,
unmappedTargetsBefore);
}
}
@@ -480,11 +602,25 @@ public final class ReplyPushLoop {
* decision-to-release window reclaims the schedule slot exactly like a raced-in reply or ticket.
*/
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore) {
Set<String> questionsBefore) {
stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, pendingIncidentKeysFor(lead));
}
private void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore, Set<IncidentLead> incidentsBefore) {
stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore, incidentsBefore,
pendingUnmappedTargetKeysFor(lead));
}
private void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore, Set<IncidentLead> incidentsBefore,
Set<UnmappedTargetLead> unmappedTargetsBefore) {
activeLeads.remove(lead);
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t))
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t));
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t))
|| pendingIncidentKeysFor(lead).stream().anyMatch(i -> !incidentsBefore.contains(i))
|| pendingUnmappedTargetKeysFor(lead).stream().anyMatch(i -> !unmappedTargetsBefore.contains(i));
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
scheduleNext(lead);
@@ -495,18 +631,22 @@ public final class ReplyPushLoop {
/** Send one combined nudge covering everything currently pending for {@code lead}. */
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount,
int questionReminderCount) {
int questionReminderCount, int incidentReminderCount,
int unmappedTargetReminderCount) {
// Re-read rather than threading it down from decide(): a reply can drain, a ticket be
// collected, or a question be answered (or another arrive), between the decision and the
// injection.
Set<String> replyTargets = pendingReplyTargetsFor(lead);
List<PendingTicket> tickets = pendingTicketsFor(lead);
List<PendingQuestion> questions = pendingQuestionsFor(lead);
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty()) {
List<PendingIncident> incidents = pendingIncidentsFor(lead);
List<PendingUnmappedTarget> unmappedTargets = pendingUnmappedTargetsFor(lead);
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty() && incidents.isEmpty()
&& unmappedTargets.isEmpty()) {
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
return;
}
String nudge = formatNudge(replyTargets, tickets, questions);
String nudge = formatNudge(replyTargets, tickets, questions, incidents, unmappedTargets);
try {
agents.send(lead, nudge);
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
@@ -515,6 +655,16 @@ public final class ReplyPushLoop {
questionReminderCount + 1, maxReminders,
replyTargets.size(), tickets.size(), questions.size());
countNudge("delivered");
for (PendingIncident incident : incidents) {
if (pendingIncidents.remove(incident.key(), incident)) {
deliveredIncidents.add(incident.key());
}
}
for (PendingUnmappedTarget unmappedTarget : unmappedTargets) {
if (pendingUnmappedTargets.remove(unmappedTarget.key(), unmappedTarget)) {
deliveredUnmappedTargets.add(unmappedTarget.key());
}
}
} catch (RuntimeException e) {
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
@@ -526,12 +676,13 @@ public final class ReplyPushLoop {
// item already at or over the cap keeps riding along in the text (still pending, still
// named) but its extra bumps here are inert: decide() already treats it as ineligible once
// its count reaches maxReminders.
bumpNudgeCounts(replyTargets, tickets, questions);
bumpNudgeCounts(replyTargets, tickets, questions, incidents, unmappedTargets);
}
/** Record that every one of these items was just named in a sent (or attempted) nudge. */
private void bumpNudgeCounts(Set<String> replyTargets, List<PendingTicket> tickets,
List<PendingQuestion> questions) {
List<PendingQuestion> questions, List<PendingIncident> incidents,
List<PendingUnmappedTarget> unmappedTargets) {
for (String target : replyTargets) {
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
}
@@ -544,6 +695,14 @@ public final class ReplyPushLoop {
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
e.nudgeCount() + 1));
}
for (PendingIncident incident : incidents) {
pendingIncidents.computeIfPresent(incident.key(), (id, e) -> new PendingIncident(e.key(),
e.credential(), e.profiles(), e.targets(), e.remainingCoolOffSeconds(), e.nudgeCount() + 1));
}
for (PendingUnmappedTarget unmappedTarget : unmappedTargets) {
pendingUnmappedTargets.computeIfPresent(unmappedTarget.key(), (id, e) ->
new PendingUnmappedTarget(e.key(), e.nudgeCount() + 1));
}
}
/** Schedule the next tick on the scheduler thread pool. */
@@ -556,7 +715,8 @@ public final class ReplyPushLoop {
/** Render everything pending for one lead as a single nudge line. */
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets,
List<PendingQuestion> questions) {
List<PendingQuestion> questions, List<PendingIncident> incidents,
List<PendingUnmappedTarget> unmappedTargets) {
List<String> parts = new ArrayList<>();
if (!replyTargets.isEmpty()) {
parts.add(formatRepliesNudge(replyTargets));
@@ -567,6 +727,12 @@ public final class ReplyPushLoop {
if (!questions.isEmpty()) {
parts.add(formatQuestionsNudge(questions));
}
if (!incidents.isEmpty()) {
parts.addAll(incidents.stream().map(ReplyPushLoop::formatBackendIncidentNudge).toList());
}
if (!unmappedTargets.isEmpty()) {
parts.addAll(unmappedTargets.stream().map(ReplyPushLoop::formatUnmappedTargetNudge).toList());
}
return String.join(" | ", parts);
}
@@ -606,6 +772,16 @@ public final class ReplyPushLoop {
return QUESTIONS_NUDGE_FORMAT.formatted(pending.size(), ids);
}
private static String formatBackendIncidentNudge(PendingIncident incident) {
return BACKEND_INCIDENT_NUDGE_FORMAT.formatted(incident.credential(), incident.remainingCoolOffSeconds(),
String.join(", ", incident.profiles()), String.join(", ", incident.targets()));
}
private static String formatUnmappedTargetNudge(PendingUnmappedTarget unmappedTarget) {
return BACKEND_TARGET_UNMAPPED_NUDGE_FORMAT.formatted(unmappedTarget.key().target(),
unmappedTarget.key().reason());
}
// --- lifecycle -----------------------------------------------------------------------------
/**
@@ -627,6 +803,10 @@ public final class ReplyPushLoop {
pendingReplies.clear();
pendingTickets.clear();
pendingQuestions.clear();
pendingIncidents.clear();
deliveredIncidents.clear();
pendingUnmappedTargets.clear();
deliveredUnmappedTargets.clear();
}
/** @see #stop() */
@@ -315,10 +315,18 @@ public final class FleetApp {
.map(Agent.class::cast)
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
// fleetd #209: this REST roster reports agentSessionId via SessionManager.rosterView, so it
// uses the resolving roster read (caller-driven, not a timer) rather than the plain one.
List<Map<String, Object>> out = sessions.rosterResolved().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
Map<String, Object> body = new LinkedHashMap<>();
// fleetd #199: the endpoint became /members in the CB-634 rename but the body key stayed
// "workers", so a caller that read "members" saw an empty fleet and reported no members at
// all. "members" is the canonical key; "workers" stays as a deprecated alias so an existing
// REST consumer keeps working — the out-of-band path a lead falls back to when its MCP mount
// drops reads this endpoint. Drop the alias once nothing reads it.
body.put("members", out);
body.put("workers", out);
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
@@ -13,6 +13,7 @@ import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.nio.file.attribute.PosixFilePermissions;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.HashSet;
@@ -23,6 +24,7 @@ import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Consumer;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
@@ -88,24 +90,64 @@ public final class GitWorktrees implements Worktrees {
);
private final String configuredRoot;
/** OS group name for {@link #shareWithGroup} (fleetd #185 stage 3); {@code null} ⇒ feature off. */
private final String group;
private final Consumer<String> afterWorktreeAdded;
/** How {@link #shareWithGroup}'s processes (git config / chgrp / chmod / find) actually run.
* Defaults to the real {@link #exec(String...)}. Package-private test seam so a unit test can
* prove "no group configured ⇒ zero processes spawned" and inspect exactly what a configured
* group runs, without a real second OS user or OS group on this host. */
private final Function<String[], String> shareGroupRunner;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
/** Default constructor: worktree root is derived per-repo as {@code <repoRoot>/../.bridged-worktrees}. */
public GitWorktrees() {
this(null);
this(null, (String) null);
}
/** @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of the repo root. */
/**
* @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of
* the repo root. No {@code worktreeGroup} configured — {@link #shareWithGroup}
* is a no-op.
*/
public GitWorktrees(String configuredRoot) {
this(configuredRoot, _ -> {});
this(configuredRoot, (String) null);
}
/**
* @param configuredRoot nullable absolute or relative path; null/blank derives a sibling of
* the repo root.
* @param group optional OS group name (fleetd #185 stage 3, {@code worktreeGroup:} in
* config); null/blank ⇒ {@link #shareWithGroup} is a no-op.
*/
public GitWorktrees(String configuredRoot, String group) {
this(configuredRoot, group, _ -> {});
}
/** Test seam for changing a real worktree between its creation and its security check. */
GitWorktrees(String configuredRoot, Consumer<String> afterWorktreeAdded) {
this(configuredRoot, null, afterWorktreeAdded);
}
/** Test seam combining a configurable {@code group} with {@link #afterWorktreeAdded}. */
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded) {
this(configuredRoot, group, afterWorktreeAdded, null);
}
/**
* Full test seam: also overrides how {@link #shareWithGroup}'s processes run (fleetd #185
* stage 3), so a unit test can prove "no group configured ⇒ no process spawned" and inspect
* exactly what commands a configured group runs, without a real second OS user/group.
*
* @param shareGroupRunner {@code null} ⇒ the real {@link #exec(String...)}.
*/
GitWorktrees(String configuredRoot, String group, Consumer<String> afterWorktreeAdded,
Function<String[], String> shareGroupRunner) {
this.configuredRoot = configuredRoot;
this.group = (group == null || group.isBlank()) ? null : group;
this.afterWorktreeAdded = afterWorktreeAdded == null ? _ -> {} : afterWorktreeAdded;
this.shareGroupRunner = shareGroupRunner != null ? shareGroupRunner : this::exec;
}
@Override
@@ -120,6 +162,7 @@ public final class GitWorktrees implements Worktrees {
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
shareRootWithGroup(root);
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
removeUserInfoFromHttpsOrigin(repoRoot);
@@ -607,6 +650,174 @@ public final class GitWorktrees implements Worktrees {
return deleted;
}
/**
* {@inheritDoc}
*
* <p>fleetd #185 stage 3. No-op — no process spawned, nothing logged — when {@link #group} is
* null/blank. Otherwise:
* <ol>
* <li>{@code git -C repoRoot config core.sharedRepository group} so every future write by
* either uid stays group-writable;</li>
* <li>a one-time {@code chgrp}/{@code chmod g+rwX} fix-up over the worktree directory and,
* under the repo's <em>common</em> git directory, {@code objects}, {@code refs},
* {@code logs}, {@code worktrees} and {@code packed-refs} — with setgid
* ({@code chmod g+s}) applied only to the directories among them, so files created later
* inherit the group;</li>
* <li>one INFO line naming the group and the paths touched.</li>
* </ol>
*
* <p><b>Every path is skipped when it does not exist.</b> {@code .git/logs} is absent in a repo
* with {@code core.logAllRefUpdates=false} or one that has had no ref update yet, and
* {@code packed-refs} is absent until refs are packed. Passing a missing path to {@code chgrp}
* exits non-zero, which would fail <em>every</em> provisioning spawn with a message blaming a
* group that is in fact fine.
*
* <p><b>The git directory is resolved, not assumed.</b> {@code <repoRoot>/.git} is a
* <em>file</em>, not a directory, when the checkout is itself a linked worktree — the very
* thing this class creates for every member. {@code git rev-parse --git-common-dir} gives the
* real shared store, and it may answer relatively, so it is resolved against {@code repoRoot}.
*
* <p><b>The fix-up re-runs on every spawn, by design.</b> {@code core.sharedRepository=group}
* governs only what git writes <em>after</em> it is set; the walk is what covers everything
* already on disk. It is not redundant work to optimise away — dropping it silently leaves
* pre-existing objects unreadable to the member. It costs three walks of the object store per
* spawn (about 3000 files in this repo, well under a second, but it grows with the repo).
*
* <p>This only fixes up file ownership/permissions on the operator's shared repo so a
* different-uid member can write to it — it isolates credentials, not the repository. A member
* in the group can still write the operator's git objects and refs.
*
* <p>Fails loudly: a missing group, or a {@code chgrp}/{@code chmod} refused because the
* operator is not a member of it, becomes a {@link WorktreeException} naming the group — never
* a silent skip that leaves a member unable to work with nothing in the log to explain why.
*/
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
if (group == null) {
return;
}
List<String> touched = new ArrayList<>();
try {
shareGroupRunner.apply(new String[]{"git", "-C", repoRoot, "config", "core.sharedRepository", "group"});
String commonDir = gitCommonDir(repoRoot);
shareGroupPathIfPresent(worktreePath, true, touched);
for (String name : List.of("objects", "refs", "logs", "worktrees")) {
shareGroupPathIfPresent(commonDir + "/" + name, true, touched);
}
shareGroupPathIfPresent(commonDir + "/packed-refs", false, touched);
} catch (WorktreeException e) {
throw new WorktreeException("cannot share worktree with group '" + group + "': "
+ e.getMessage() + " — the group must exist, and the fleetd operator ("
+ System.getProperty("user.name") + ") must be a member of it", e);
}
log.info("worktreeGroup={} shared repoRoot={} worktreePath={} paths={}",
group, repoRoot, worktreePath, touched);
}
/**
* fleetd #224: make {@code worktreeRoot} ITSELF group-traversable — established once, here,
* where the root is created, never at a use site. {@link #shareWithGroup} shares each worktree
* (and the repo's common git dir) with {@link #group}, but never the PARENT directory that
* contains every worktree — and under {@code memberHerdrSocket:} the member pane runs as a
* different OS user, which needs the execute bit on every ancestor directory to reach anything
* underneath, no matter how carefully each child is shared. Without this, a member cannot read
* the ephemeral {@code opencode.json} #219 places under this root, cannot reach its own
* worktree, and cannot read #213's ZDOTDIR scrub when placed here either.
*
* <p>No-op — no process spawned — when {@link #group} is null/blank, so behaviour with
* {@code worktreeGroup:} unset (today's only live mode) is unchanged. Only {@code root} itself
* is touched (non-recursive): each child underneath is shared individually, either by
* {@link #shareWithGroup} for a worktree or by the launcher that generates it (fleetd #213/#219)
* for a scrub/config directory — sharing this level again would just duplicate that policy in
* the wrong layer.
*
* <p>Fails loudly, naming {@code root}, its mode at the time of the attempt, and {@link #group}:
* a member that starts and then cannot see its own checkout is worse than a refused spawn, since
* nothing about that failure mode points at a directory's permission bits.
*/
private void shareRootWithGroup(Path root) {
if (group == null) {
return;
}
String mode = currentPosixMode(root);
try {
shareGroupRunner.apply(new String[]{"chgrp", group, root.toString()});
shareGroupRunner.apply(new String[]{"chmod", "g+x", root.toString()});
} catch (WorktreeException e) {
throw new WorktreeException("cannot make worktree root " + root + " (mode " + mode
+ ") group-traversable for group '" + group + "': " + e.getMessage()
+ " — the group must exist, and the fleetd operator (" + System.getProperty("user.name")
+ ") must be a member of it", e);
}
log.info("worktreeGroup={} made worktree root {} group-traversable (was mode {})", group, root, mode);
}
/** {@code root}'s current POSIX permission string, or {@code "unknown"} on a filesystem that does
* not support POSIX permissions — used only to name the mode in a refusal message. */
private static String currentPosixMode(Path root) {
try {
return PosixFilePermissions.toString(Files.getPosixFilePermissions(root));
} catch (IOException | UnsupportedOperationException e) {
return "unknown";
}
}
/**
* The repo's <em>common</em> git directory as an absolute path — where {@code objects},
* {@code refs} and {@code worktrees} actually live. {@code git rev-parse --git-common-dir}
* answers relative to {@code repoRoot} in the ordinary case ({@code .git}) and absolutely for a
* linked worktree, so the answer is resolved against {@code repoRoot} either way. Never
* hardcode {@code repoRoot + "/.git"}: that is a FILE when the checkout is itself a linked
* worktree.
*/
private String gitCommonDir(String repoRoot) {
String answer = shareGroupRunner.apply(
new String[]{"git", "-C", repoRoot, "rev-parse", "--git-common-dir"});
String trimmed = answer == null ? "" : answer.trim();
if (trimmed.isEmpty()) {
trimmed = ".git";
}
return Path.of(repoRoot).resolve(trimmed).normalize().toString();
}
/**
* {@link #shareGroupPath} when {@code path} exists, recording it in {@code touched}; otherwise
* nothing at all. A missing path is normal, not an error — see {@link #shareWithGroup}'s
* javadoc for which ones are routinely absent and why passing them to {@code chgrp} would fail
* every spawn.
*/
private void shareGroupPathIfPresent(String path, boolean recursive, List<String> touched) {
if (!Files.exists(Path.of(path))) {
return;
}
shareGroupPath(path, recursive);
touched.add(path);
}
/**
* {@code chgrp}/{@code chmod g+rwX} {@code path} to {@link #group}. When {@code recursive},
* also walks the directories under {@code path} (including {@code path} itself, when it is a
* directory) and sets setgid on each — directories only, per the javadoc on
* {@link #shareWithGroup}.
*/
private void shareGroupPath(String path, boolean recursive) {
List<String> chgrp = new ArrayList<>(List.of("chgrp"));
if (recursive) chgrp.add("-R");
chgrp.add(group);
chgrp.add(path);
shareGroupRunner.apply(chgrp.toArray(new String[0]));
List<String> chmod = new ArrayList<>(List.of("chmod"));
if (recursive) chmod.add("-R");
chmod.add("g+rwX");
chmod.add(path);
shareGroupRunner.apply(chmod.toArray(new String[0]));
if (recursive) {
shareGroupRunner.apply(new String[]{"find", path, "-type", "d", "-exec", "chmod", "g+s", "{}", "+"});
}
}
/**
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
@@ -89,4 +89,15 @@ public record MemberSession(
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
}
/**
* Return a copy with {@code agentSessionId} resolved to a non-null value (fleetd #209). Some
* adapters (opencode) cannot answer {@link dev.ltms.fleet.peer.PeerHandle#agentSessionId()} at
* spawn time — the peer has not persisted its session record yet — so the id is discovered on
* a later poll and swapped into the otherwise-immutable session via this wither.
*/
public MemberSession withAgentSessionId(String agentSessionId) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
}
@@ -46,6 +46,15 @@ public final class SessionManager implements TurnListener {
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, MemberSession> registry = new ConcurrentHashMap<>();
/**
* fleetd #209: the live {@link PeerHandle} for every registered pane, retained solely so
* {@link #resolveAgentSessionId} can re-poll {@link PeerHandle#agentSessionId()} after spawn.
* The handle used to go out of scope at the end of the spawn method, so a launcher that answers
* the id lazily (opencode — the on-disk session row is written after the pane is created) could
* never be re-asked, and {@code fleet_list}/{@code fleet_spawn resumeSessionId} never saw it.
* Populated on every spawn path, removed on {@link #release}.
*/
private final ConcurrentHashMap<String /*paneId*/, PeerHandle> handles = new ConcurrentHashMap<>();
private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
@@ -174,45 +183,51 @@ public final class SessionManager implements TurnListener {
String sessionName, String resumeSessionId) {
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
requireResumeCapability(profile, resumeSessionId);
// CB-619 / fleetd #123: an explicit profile bypasses placement (CompositePeerLauncher only
// constrains an UNQUALIFIED spawn to the role's pool), so it is the one path that can ask
// for a role with no slot to bind it to. Refuse before anything spawns. A blank profile is
// left to placement, which already restricts an unqualified spawn to the role's pool.
if (profile != null && !profile.isBlank()) {
memberLifecycle.requireSlotFor(memberRole, profile);
}
MemberLifecycle.SlotReservation reservation = memberLifecycle.reserve(memberRole, profile);
String launchProfile = reservation == null ? profile : reservation.profile();
if (wt == null) {
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
SpawnRequest req = new SpawnRequest(launchProfile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
PeerHandle handle;
boolean bound = false;
try {
handle = launcher.spawn(req);
String resolvedProfile = resolveProfile(handle, launchProfile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
MemberRole actualRole = acquired(memberRole, resolvedProfile, handle.terminalId(), reservation);
bound = reservation == null || actualRole == MemberRole.ARCHITECT;
MemberSession session = new MemberSession(
handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0,
MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId());
registry.put(handle.id(), session);
handles.put(handle.id(), handle);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={}: {}", profile, memberRole, e.getMessage());
if (!bound) memberLifecycle.release(reservation);
log.warn("spawn failed for profile={} role={}: {}", launchProfile, memberRole, e.getMessage());
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
memberRole,
cwd,
ownerTerminal,
now,
now,
0,
MemberSession.State.SPAWNING,
null,
null,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId);
try {
return acquireWithWorktree(launchProfile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId, reservation);
} catch (RuntimeException e) {
memberLifecycle.release(reservation);
throw e;
}
}
/**
@@ -257,6 +272,10 @@ public final class SessionManager implements TurnListener {
*/
private void release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
// fleetd #209: remove right alongside the registry entry so a released session's handle is
// never leaked — but keep the local reference below, so the id can still be resolved for
// the ReleaseDetail this teardown notifies with.
PeerHandle removedHandle = handles.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
String snapshotRef = null;
if (removed != null) {
@@ -303,8 +322,11 @@ public final class SessionManager implements TurnListener {
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
// CB-584 (issue #65 criterion 5): carry agentSessionId alongside them, so a lead can
// also resume the member's conversation, not just re-dispatch onto its files.
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
removed.branch(), snapshotRef, removed.agentSessionId()));
// fleetd #209: a late-resolving adapter (opencode) may only now have an id — resolve
// one last time so a released member's detail carries the id it now has.
MemberSession resolved = resolveAgentSessionId(removed, removedHandle);
notifyReleased(new ReleaseDetail(resolved.terminalId(), resolved.worktree(),
resolved.branch(), snapshotRef, resolved.agentSessionId()));
}
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
@@ -451,7 +473,8 @@ public final class SessionManager implements TurnListener {
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
String sessionName, String resumeSessionId,
MemberLifecycle.SlotReservation reservation) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
@@ -472,6 +495,10 @@ public final class SessionManager implements TurnListener {
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
// fleetd #185 stage 3: MUST run after overlayParity, not folded into add() — overlayParity
// copies more files into the worktree after add() returns, so sharing the group any earlier
// leaves those overlay files operator-owned and read-only for a different-uid member.
worktrees.shareWithGroup(repoRoot, path);
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
@@ -488,11 +515,14 @@ public final class SessionManager implements TurnListener {
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
// CB-619: see the no-worktree path above — bind before recording, and store the returned
// actual role, so this session's role is never a lie about what it actually holds.
MemberRole actualRole = acquired(memberRole, resolvedProfile, handle.terminalId(), reservation);
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
memberRole,
actualRole,
cwd,
ownerTerminal,
now,
@@ -504,13 +534,25 @@ public final class SessionManager implements TurnListener {
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
handles.put(handle.id(), handle);
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
private MemberRole acquired(MemberRole role, String profile, String terminal,
MemberLifecycle.SlotReservation reservation) {
if (reservation == null) {
return memberLifecycle.acquired(role, profile, terminal);
}
if (memberLifecycle.bind(reservation, terminal)) {
return role;
}
memberLifecycle.release(reservation);
return memberLifecycle.acquired(role, profile, terminal);
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
@@ -535,16 +577,73 @@ public final class SessionManager implements TurnListener {
return launcher.defaultProfile();
}
/** The session for {@code paneId}, if it is still registered and not released. */
/**
* The session for {@code paneId}, if it is still registered and not released. fleetd #209:
* resolves a still-unknown {@code agentSessionId} against the retained handle before returning,
* so {@code fleet_status} sees an id a lazy-resolving adapter has since written.
*/
public Optional<MemberSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
return Optional.ofNullable(registry.get(paneId)).map(this::resolveAgentSessionId);
}
/** Fleet-owned roster: all registered sessions (acquired minus released). */
/**
* Fleet-owned roster: all registered sessions (acquired minus released). Deliberately does
* <strong>not</strong> resolve {@code agentSessionId} (fleetd #209 follow-up) — this is the
* roster supplier on the heartbeat and health-tick timers ({@code LeadHeartbeatLoop},
* {@code FleetHealthMonitor} in {@code Fleetd}), on the placement/exhaustion paths, and on the
* metrics scrape ({@code FleetMetrics}), all called far more often than any caller actually
* reads {@code agentSessionId}. Resolving here would mean every tick opens a lazy-resolving
* adapter's (opencode's) on-disk session store once per member whose id is still unknown — and
* for a member whose id never appears, that cost never stops, for the life of the process. Use
* {@link #rosterResolved()} instead wherever the id must be current.
*/
public List<MemberSession> roster() {
return List.copyOf(registry.values());
}
/**
* {@link #roster()}, with each session's still-unknown {@code agentSessionId} re-resolved
* against its retained handle (fleetd #209) — so a caller that actually reports the id (
* {@code fleet_list}, the REST roster) sees one a lazy-resolving adapter (opencode) has since
* written, rather than the null frozen in at spawn time. Reserved for caller-driven reads, not
* timers: see {@link #roster()}'s javadoc for why the plain roster must stay non-resolving.
*/
public List<MemberSession> rosterResolved() {
return registry.values().stream().map(this::resolveAgentSessionId).toList();
}
/**
* Resolve {@code session}'s {@code agentSessionId} if still unknown, re-polling the retained
* {@link PeerHandle} for this pane (fleetd #209). A no-op — returning {@code session} unchanged
* — once the id is already known, once no handle is retained for this pane (never spawned, or
* already released), or if the handle throws while answering. A resolved id is best-effort
* CAS-swapped into the registry via {@link #replace}; a lost race just means another caller
* already applied the same update, so the resolved value is returned either way.
*/
private MemberSession resolveAgentSessionId(MemberSession session) {
return resolveAgentSessionId(session, handles.get(session.paneId()));
}
private MemberSession resolveAgentSessionId(MemberSession session, PeerHandle handle) {
if (session.agentSessionId() != null || handle == null) {
return session;
}
String resolved;
try {
resolved = handle.agentSessionId();
} catch (RuntimeException e) {
log.debug("agentSessionId lookup failed for pane={} terminal={}: {}",
session.paneId(), session.terminalId(), e.toString());
return session;
}
if (resolved == null) {
return session;
}
MemberSession updated = session.withAgentSessionId(resolved);
replace(session, updated); // best-effort; a lost CAS just means the resolved value stands anyway
return updated;
}
/**
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
@@ -98,4 +98,21 @@ public interface Worktrees {
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
record WipRefStats(int count, long costBytes) {
}
/**
* Make {@code repoRoot}'s git store and {@code worktreePath} writable by the configured group
* (fleetd #185 stage 3), so a member spawned as a different OS user (see
* {@code memberHerdrSocket}) can write its own worktree, its per-worktree git metadata, and
* its own commit objects. No-op when no group is configured.
*
* <p><strong>This isolates credentials, not the repository.</strong> A member in the group can
* still write the operator's git objects and refs in the shared repo — this only fixes file
* ownership/permissions so a different-uid member can work at all, it grants no narrower access
* than that.
*
* @param repoRoot the repository whose git store ({@code .git/objects}, {@code refs},
* {@code logs}, {@code worktrees}, {@code packed-refs}) needs sharing
* @param worktreePath the linked worktree's own directory
*/
void shareWithGroup(String repoRoot, String worktreePath);
}
@@ -1036,6 +1036,35 @@ class FleetConfigTest {
assertEquals(LeadMailbox.DEFAULT_PREFETCH, noEnv.prefetchOrDefault());
}
@Test
void absentWorktreeGroupLeavesItNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-worktree-group.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
FleetConfig cfg = FleetConfig.load(f);
assertNull(cfg.worktreeGroup(), "no worktreeGroup: key → null → GitWorktrees.shareWithGroup is a no-op");
}
@Test
void worktreeGroupKeyParses(@TempDir Path dir) throws Exception {
Path f = dir.resolve("worktree-group.yaml");
Files.writeString(f, "bind:\n port: 8080\nworktreeGroup: fleet-workers\n");
FleetConfig cfg = FleetConfig.load(f);
assertEquals("fleet-workers", cfg.worktreeGroup());
}
@Test
void absentWorktreeGroupSurvivesTheBackCompatConstructorChain() {
// fleetd #185 stage 3: withDefaults() (and every pre-existing call site) must not silently
// drop a live worktreeGroup by routing through a back-compat constructor that defaults it
// to null.
FleetConfig cfg = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, "fleet-workers");
assertEquals("fleet-workers", cfg.withDefaults().worktreeGroup(),
"withDefaults() must carry a configured worktreeGroup through unchanged");
}
@Test
void absentPrimaryBlockLeavesPrimaryNull(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-primary.yaml");
@@ -44,10 +44,13 @@ public final class FakeHerdr implements HerdrClient {
private String agentSendErrorCode = null;
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -103,6 +106,27 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the detected agent kind ({@code "agent"} field) that {@code agent.get} reports —
* {@code null} models herdr not (or no longer) detecting a supported backend in the pane, e.g.
* a bare shell prompt (fleetd #176 fix 2 corroboration test).
*/
public FakeHerdr agentType(String type) {
this.agentType = type;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
* fixture. {@code okCalls == 0} fails from the very first call.
*/
public FakeHerdr agentGetFailsWithAfter(int okCalls, String code) {
this.agentGetOkCalls = okCalls;
this.agentGetFailCode = code;
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
@@ -208,10 +232,21 @@ public final class FakeHerdr implements HerdrClient {
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
case "agent.get" -> mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
case "agent.get" -> {
if (agentGetFailCode != null) {
long getCalls = calls.stream().filter(c -> c.method().equals("agent.get")).count();
if (getCalls > agentGetOkCalls) {
throw new HerdrException(
"herdr error [" + agentGetFailCode + "]: agent target not found",
agentGetFailCode, null);
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
.formatted(agentField, agentStatus));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
@@ -633,6 +633,115 @@ class CompletionResolverTest {
"the failure carries the rest of the pane, not only the matched line: " + reason);
}
// --- fleetd#211: raw-scrape fallback classification when there is no usable assistant block ---
@Test
void anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink() {
// No ⏺ anywhere, and the first visible line is TUI chrome (╭). lastAssistantBlock's boundary
// scan starts at the top of the raw screen and breaks immediately, so the trimmed block is "".
// The fix: fall back to matching the RAW scrape so this doesn't get lost as an empty scrape.
String block = """
╭──────────────────────────────────────╮
The usage limit has been reached. Try again later.
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape match still resolves the blocked send");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"classified from the raw scrape even though the trimmed block was empty");
assertEquals(1, notified.size(),
"the sink is the whole point of this ticket — it must be notified: " + notified);
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
assertTrue(notified.get(0).contains("The usage limit has been reached"),
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape() {
String block = """
╭──────────────────────────────────────╮
API Error: 400 invalid request body
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape backend-error match still resolves the blocked send");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"classified BACKEND_ERROR from the raw scrape even though the trimmed block was empty");
assertTrue(waiter.getNow(null).text().contains("API Error: 400 invalid request body"),
"the failure carries the matched line: " + waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("--- pane tail ---"),
"fleetd#164: the failure must carry the pane, not only the matched line — with an "
+ "empty trimmed block the raw scrape is the only copy of what the member said: "
+ waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("╭"),
"the carried pane is the raw scrape, chrome included: " + waiter.getNow(null).text());
}
@Test
void anOrdinaryPaneWithANormalAssistantBlockIsUnaffectedByTheRawScrapeFallback() {
// Pin: on a pane that already yields a usable block, the fallback branch is never reached —
// same outcome, same text, sink not called — even though the raw screen around the marker
// would itself match the configured exhausted pattern.
String block = "The usage limit has been reached, but this is a leading TUI line above the "
+ "marker.\n⏺ complete report\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"unchanged: a usable assistant block never reaches the raw-scrape fallback");
assertEquals("complete report", waiter.getNow(null).text(), "the reply text is unaffected");
assertTrue(notified.isEmpty(), "the fallback never runs, so the sink is never called");
}
@Test
void aGenuinelyEmptyScrapeStillFailsAsEmptyAndNeverNotifiesTheSink() {
// The false-positive pin: no exhaustion or backend-error text anywhere on the pane (just
// chrome, no marker) — the raw-scrape fallback must not manufacture a classification, and
// the sink must stay untouched.
String block = """
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "an empty scrape must still resolve the send, not hang");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"no exhaustion or backend-error text anywhere ⇒ this stays the ordinary empty-scrape failure");
assertTrue(waiter.getNow(null).text().toLowerCase().contains("empty"),
"the failure still says the scrape was empty: " + waiter.getNow(null).text());
assertTrue(notified.isEmpty(), "a genuinely empty pane must never quarantine a credential");
}
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
@@ -1,10 +1,13 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.msg.MessageService;
@@ -1024,6 +1027,87 @@ class FleetMcpTest {
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
}
// ── CB-619 / fleetd #123: a spawn asking for a role its profile has no slot for must be
// refused, never silently demoted with the roster still lying about it ───────────────────
/** Two profiles on one launcher, so a role's pool can name one and exclude the other. */
private static ClaudeCodeLauncher architectCapableLauncher(FakeHerdr h) {
FleetConfig.Profile opus = new FleetConfig.Profile(
"opus", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.Profile sonnet = new FleetConfig.Profile(
"sonnet", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("opus", opus, "sonnet", sonnet),
"sonnet", _ -> "tok");
}
/** {@code fleet.architects} carries only {@code opus} — {@code sonnet} has no matching slot. */
private static MemberRegistry architectRegistry() {
return new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("opus", new FleetConfig.Slot("opus")), Map.of(), Map.of(), null));
}
/**
* The literal defect (fleetd #123): {@code role=architect, profile=sonnet}, where
* {@code fleet.architects} carries only {@code opus}. Drives the real path —
* {@link FleetMcp#spawn} calls {@link SessionManager#acquire}, which must refuse before ever
* reaching the real {@link ClaudeCodeLauncher} — never {@link MemberRegistry#bind} called
* directly, which would walk around the gate under test.
*/
@Test
void spawnRefusesAnArchitectWithNoMatchingSlotAndNeverTouchesTheLauncher() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(architectCapableLauncher(h));
sessions.setMemberLifecycle(architectRegistry());
McpSchema.CallToolResult res = FleetMcp.spawn(sessions, "sonnet", "architect",
null, null, null, null, null, null);
assertEquals(Boolean.TRUE, res.isError(), textOf(res));
String msg = textOf(res);
assertTrue(msg.contains("architect"), "names the role asked for: " + msg);
assertTrue(msg.contains("sonnet"), "names the profile: " + msg);
assertTrue(msg.contains("opus"), "names the pool that does carry the role: " + msg);
assertTrue(sessions.roster().isEmpty(), "a refused spawn must register no session");
assertTrue(h.calls.stream().noneMatch(c -> c.method().equals("agent.start")),
"a refused spawn must never reach the launcher — no process should ever start");
}
/**
* Positive control / parity check: when the profile DOES carry a slot, the spawn succeeds, and
* {@code GET /members} ({@link FleetMcp#listFleet}) and {@code fleet_whoami}
* ({@link FleetMcp#whoami}) — resolved through the SAME live {@link MemberRegistry} binding via
* a real {@link CallerResolver}, exactly as the daemon resolves a real MCP caller — must never
* disagree about this one live member's role.
*/
@Test
void rosterAndWhoamiAgreeOnceTheArchitectSlotBinds() {
FakeHerdr h = new FakeHerdr();
// Pin the spawn onto the one pane FakeHerdr's canned pane.process_info maps to WORKER_PID,
// so a CallerResolver can resolve THIS session's own terminal, not a fixture double.
h.pinNextStarts(1, "term_a", "w2:p7");
SessionManager sessions = new SessionManager(architectCapableLauncher(h));
MemberRegistry members = architectRegistry();
sessions.setMemberLifecycle(members);
McpSchema.CallToolResult spawnRes = FleetMcp.spawn(sessions, "opus", "architect",
null, null, null, null, null, null);
assertNotEquals(Boolean.TRUE, spawnRes.isError(), textOf(spawnRes));
assertTrue(textOf(spawnRes).contains("\"role\":\"architect\""), textOf(spawnRes));
String roster = textOf(FleetMcp.listFleet(architectCapableLauncher(h), sessions, Map.of(), ""));
assertTrue(roster.contains("\"role\":\"architect\""), "GET /members: " + roster);
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(h), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null, Map::of, members);
Principal caller = resolver.resolve("127.0.0.1", 42, null);
assertTrue(caller.isArchitect(), "fleet_whoami's own resolver must agree the terminal is bound");
String whoami = textOf(FleetMcp.whoami(caller, sessions));
assertTrue(whoami.contains("\"role\":\"architect\""), "fleet_whoami: " + whoami);
}
// ── CB-584: fleet_spawn accepts sessionName/resumeSessionId; roster shows agentSessionId ──
@Test
@@ -20,8 +20,10 @@ import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFilePermissions;
import java.util.List;
import java.util.Map;
import java.util.Set;
@@ -31,6 +33,7 @@ import java.util.function.Function;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
class ClaudeCodeLauncherTest {
@@ -61,8 +64,75 @@ class ClaudeCodeLauncherTest {
assertTrue(args.stream().noneMatch(a -> a.contains("\"bridge\"")),
"the mount is named fleet since CB-632 — a member addresses its tools as "
+ "mcp__fleet__*, and CLAUDE.md's role-detection ladder names that prefix");
assertTrue(args.contains("--append-system-prompt"));
assertTrue(args.stream().anyMatch(a -> a.contains("fleet_reply")), "reply charter present");
// fleetd #220: the charter travels as a FILE, never inline — charter prose on the command
// line overruns the byte cap of the pane line herdr types it into.
assertTrue(args.contains("--append-system-prompt-file"));
assertFalse(args.contains("--append-system-prompt"),
"the inline flag would put ~800 bytes of prose on the pane command line");
String charterFile = args.get(args.indexOf("--append-system-prompt-file") + 1);
assertTrue(readFile(charterFile).contains("fleet_reply"), "reply charter present in the file");
}
/** Read a charter file the launcher wrote, failing the test rather than the build on an IO error. */
private static String readFile(String path) {
try {
return java.nio.file.Files.readString(java.nio.file.Path.of(path));
} catch (java.io.IOException e) {
throw new AssertionError("charter file " + path + " is not readable", e);
}
}
/**
* fleetd #220 regression. herdr TYPES the launch command into the pane, and the pty line buffer
* holds 1024 bytes — past that the tail is dropped with no error anywhere, so the backend exits
* on a mangled argument and the spawn dies as an unexplained readiness timeout. That is what
* happened when #214 added --session-id to a command already 978 bytes long: --autocompact
* 250000 arrived as --autocompact 25. This asserts the whole assembled command still fits, with
* the flags a real spawn carries (MCP mount, charter, model, autocompact, session id).
*/
@Test
void theAssembledLaunchCommandFitsThePaneLineLimit() {
FakeHerdr herdr = new FakeHerdr();
// Shaped like the live sonnet profile, because the bug is a SUM: the charter alone fits,
// and so does every flag alone. Only model + autocompact + session id on top of the charter
// crossed the cap, which is why nothing caught it until a member failed to spawn.
FleetConfig.Profile cfg = new FleetConfig.Profile(
"sonnet", null, "claude-sonnet-5", null, null,
List.of("claude"), "tab", "fleetd-workers", "w #{n}", "http://127.0.0.1:8765/mcp",
null, null, null, null, null, null, null, null, true, null, null, null, null, null,
250000);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of()), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
.spawn();
List<String> args = spawnedArgs(herdr);
int bytes = "claude".length();
for (String arg : args) {
bytes += arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length + 3;
}
assertTrue(bytes <= HerdrPeerLauncher.PANE_COMMAND_BYTE_LIMIT,
"the launch command must fit the pane line: " + bytes + " bytes vs limit "
+ HerdrPeerLauncher.PANE_COMMAND_BYTE_LIMIT + " — args " + args);
}
/**
* fleetd #220: the guard refuses a command that cannot fit, instead of letting the pty drop the
* tail. The refusal must name the size and the argument to blame — a spawn that fails with
* "did not reach injectable state" tells the operator nothing, which is the whole reason this
* bug took a live pane scrape to find.
*/
@Test
void anOverlongLaunchCommandIsRefusedWithTheSizeAndTheCulprit() {
FakeHerdr herdr = new FakeHerdr();
String huge = "x".repeat(1500);
ClaudeCodeLauncher launcher = service(herdr, List.of("claude", huge), null);
PeerUnreachableException refused = assertThrows(PeerUnreachableException.class, launcher::spawn);
assertTrue(refused.getMessage().contains("1024"), "names the limit: " + refused.getMessage());
assertTrue(refused.getMessage().contains("1500"), "names the culprit's size: " + refused.getMessage());
assertFalse(herdr.called("agent.start"),
"nothing may be started — a truncated command is worse than no spawn");
}
// CB-634: a profile with ideMcpUrl set mounts the IDE Index MCP as a second server and pins
@@ -262,10 +332,16 @@ class ClaudeCodeLauncherTest {
Map<?, ?> start = (Map<?, ?>) herdr.lastCall("agent.start").params();
assertEquals("claude", start.get("kind"), "herdr launches the canonical executable by kind");
// CB-533: the shared fixture pins model "coder", so the model flag is the whole args list.
// What this test guards is that argv[0] is NOT repeated — herdr supplies it from `kind`.
assertEquals(List.of("--model", "coder"), start.get("args"),
"the configured executable is not repeated in args");
// fleetd #214: a plain spawn now also carries fleetd's own minted session id.
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("claude"), "the configured executable is not repeated in args");
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0, "a plain spawn mints a session id: " + args);
assertDoesNotThrow(() -> UUID.fromString(args.get(flag + 1)), "the minted id is a valid UUID");
assertEquals(List.of("--model", "coder"),
List.of(args.get(args.size() - 2), args.get(args.size() - 1)),
"the CB-533 model flag still trails the launch flags: " + args);
}
@Test
@@ -278,7 +354,15 @@ class ClaudeCodeLauncherTest {
assertFalse(args.contains("--append-system-prompt"), "no reply charter without mcpUrl");
// CB-533: the model flag is independent of the MCP mount — pinning the model is not part of
// "mount the bridge", so an unmounted worker still runs the model its profile names.
assertEquals(List.of("--verbose", "--model", "coder"), args,
// fleetd #214: a plain spawn now carries fleetd's own minted session id; drop the two
// mint elements when checking the rest of the argv.
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0, "a plain spawn mints a session id even without a bridge mount: " + args);
assertDoesNotThrow(() -> UUID.fromString(args.get(flag + 1)), "the minted id is a valid UUID");
List<String> rest = new java.util.ArrayList<>(args);
rest.remove(flag + 1);
rest.remove(flag);
assertEquals(List.of("--verbose", "--model", "coder"), rest,
"the operator's own args are preserved, in order, ahead of the model flag");
}
@@ -649,7 +733,7 @@ class ClaudeCodeLauncherTest {
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
// --- CB-547a: durable session identity (mint / resume / no-identity legacy) -----------------
// --- CB-547a / fleetd #214: durable session identity (always mint / resume) -----------------
@Test
void freshSpawnMintsASessionIdAndPassesTheName() {
@@ -684,18 +768,28 @@ class ClaudeCodeLauncherTest {
}
@Test
void noIdentitySpawnKeepsTheLegacyArgvAndCarriesNoSessionHandle() {
void plainSpawnMintsASessionIdSoEveryMemberIsResumable() {
// fleetd #214: a plain spawn passes no sessionName and no resumeSessionId, yet the member
// must still be resumable — the id is minted unconditionally, and it is the ONLY resume
// handle a claude-code member has (unlike opencode, nothing resolves it after the launch).
// The binary requires a valid UUID (checked against claude 2.1.252: a non-UUID is refused
// at argument parsing with "Invalid session ID. Must be a valid UUID.").
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "no identity → no --session-id");
assertFalse(args.contains("-n"), "no identity → no -n");
assertFalse(args.contains("-r"), "no identity → no -r");
assertNull(handle.agentSessionId(), "no identity → no resume handle");
assertNull(handle.sessionName(), "no identity → no logical name");
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0 && flag + 1 < args.size(),
"--session-id is minted even when no identity is requested: " + args);
String minted = args.get(flag + 1);
assertDoesNotThrow(() -> UUID.fromString(minted), "--session-id is a valid UUID: " + minted);
assertEquals(minted, handle.agentSessionId(),
"the resume handle is the minted id, so every member is resumable from fleet_list");
assertFalse(args.contains("-n"), "no sessionName was requested → no -n");
assertFalse(args.contains("-r"), "no resume was requested → no -r");
assertNull(handle.sessionName(), "no sessionName was requested → no logical name");
}
@Test
@@ -869,6 +963,152 @@ class ClaudeCodeLauncherTest {
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- fleetd #176 fix 1: fail fast when the backend process exits mid-wait -------------------
@Test
void spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout() {
// First agent.get sees UNKNOWN (one normal tick); the second reports the pane gone, exactly
// what herdr answers when the backend process has already exited. The timeout is generous
// (60s) so a test that wrongly falls through to the old unguarded call — and therefore
// waits out the whole window — is unambiguously distinguishable from one that fails fast.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(1, "pane_not_found");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().toLowerCase().contains("exited"),
"exception message says the backend exited, not that the pane was slow: "
+ ex.getMessage());
assertTrue(clock[0] < 60_000,
"the gate must not burn the rest of the 60s timeout: clock only reached " + clock[0]);
long getCalls = herdr.calls.stream().filter(c -> c.method().equals("agent.get")).count();
assertEquals(2, getCalls,
"exactly one normal poll then the not_found answer — no further polling after that: "
+ getCalls);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down (no orphan) even on the fail-fast path");
}
@Test
void spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged() {
// Fix 1 must only special-case a "*_not_found" answer. Any other herdr failure keeps
// propagating as-is — this gate does not know how to recover from it.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(0, "internal_error");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
dev.ltms.fleet.herdr.HerdrException ex = assertThrows(
dev.ltms.fleet.herdr.HerdrException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertEquals("internal_error", ex.code());
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"an error this gate does not recognize is not this gate's teardown to run");
}
// --- fleetd #176 fix 2: corroborated UNKNOWN refinement --------------------------------------
@Test
void refinedIdleIsNotAcceptedWhenAgentTypeIsNull() {
// The trap fix 2 must close: a pane sitting at a bare shell prompt after its backend exited
// still contains the same "❯" glyph StatusRefiner.classify treats as "idle at the Claude
// Code TUI prompt". Without the agentType corroboration this would be misread as injectable
// and the gate would hand back a peer that never started. herdr reports no agentType for
// that bare shell.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType(null); // no supported backend detected — could be a bare shell
herdr.readText("some-host:~ user$ ❯ "); // looks exactly like an idle Claude Code prompt
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
500, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 500,
"a null agentType must not let the '❯' prompt refine to injectable — the gate has "
+ "to wait out the full timeout: clock only reached " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down on timeout, same as any other never-injectable spawn");
}
@Test
void refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness() {
// The positive case: a genuinely live claude pane that herdr misreports as UNKNOWN (CB-115)
// still resolves to injectable once agentType corroborates that a supported backend is
// detected in the same sample.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN at the raw level
herdr.agentType("claude"); // herdr still detects a live claude backend
herdr.readText("? for shortcuts"); // StatusRefiner.classify's idle footer marker
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "a corroborated refined-IDLE sample lets the spawn succeed");
assertTrue(clock[0] < 60_000,
"refinement must resolve well before the timeout: clock reached " + clock[0]);
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane close — the peer is genuinely injectable");
}
@Test
void refinementNeverFiresWhenRawStatusIsAlreadyInjectable() {
// Acceptance criterion 4: refinement must trigger only on a raw UNKNOWN sample. Seed the
// pane content with an active-generation marker that StatusRefiner.classify would read as
// WORKING (never injectable) if refine() were wrongly invoked here — proving that a raw
// IDLE status short-circuits before refine() (and its extra agent.read call) ever runs.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("idle"); // already injectable at the raw level
herdr.agentType("claude");
herdr.readText("esc to interrupt"); // would classify as WORKING if refine() ran anyway
long[] clock = {0};
ClaudeCodeLauncher gated = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], () -> clock[0] += 300);
PeerHandle handle = gated.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "an already-injectable raw status succeeds without ever refining");
assertFalse(herdr.called("agent.read"),
"refine() must never run (and so never call agent.read) when the raw status is "
+ "already injectable");
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
@@ -1660,4 +1900,184 @@ class ClaudeCodeLauncherTest {
assertEquals(List.of("dev: sonnet #1", "[sonnet] dev 2"), tabLabels(herdr));
}
// ── fleetd #222: the charter file must not land under fleetd's own java.io.tmpdir when the ──
// ── member pane runs as a different OS user ─────────────────────────────────────────────────
/** A profile that mounts the bridge MCP (so a reply charter is always generated). */
private static FleetConfig.Profile charterCfg() {
return new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}",
"http://127.0.0.1:8765/mcp", null, null);
}
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
private static FleetConfig configWithMemberHerdrSocket(String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
"/tmp/other-user.sock", // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
null // memberLoginShell
).withDefaults();
}
private static ClaudeCodeLauncher serviceWithConfig(FakeHerdr herdr, FleetConfig.Profile cfg,
Supplier<FleetConfig> config) {
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, null, null, null, config);
}
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225). Those two coincide only by accident: a directory's
* group is whichever group happened to own the path Maven was started from — {@code staff} in
* a home checkout, {@code wheel} under {@code /private/tmp} on macOS — and the fix-up this test
* exercises then fails for real when the operator is not a member of that borrowed group,
* exactly the case {@code assumeTrue(view != null, ...)} never covered (it only detects a
* filesystem with no POSIX groups at all, not a resolvable-but-wrong one). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
try (java.io.BufferedReader r = new java.io.BufferedReader(
new java.io.InputStreamReader(p.getInputStream(), java.nio.charset.StandardCharsets.UTF_8))) {
out = r.lines().collect(java.util.stream.Collectors.joining("\n")).trim();
}
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !out.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/**
* fleetd #222 acceptance criterion 1: with {@code memberHerdrSocket} configured and both
* {@code worktreeRoot}/{@code worktreeGroup} set, the charter file lives in a fresh per-spawn
* directory under {@code worktreeRoot} — NEVER under {@code java.io.tmpdir} (fleetd's own 0700
* temp dir, unreadable by the member's different OS user) — and that directory is shared
* read-only with the group via the SAME mechanism (fleetd #213/#219's {@link
* EnvAllowListScrub#shareWithGroup}) the ZDOTDIR scrub and the opencode config directory use.
*/
@Test
void memberHerdrSocketWithWorktreeRootAndGroupPutsCharterUnderWorktreeRootAndSharesIt(
@TempDir Path worktreeRoot) throws Exception {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
serviceWithConfig(herdr, charterCfg(),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), group)).spawn();
List<String> args = spawnedArgs(herdr);
int fileFlag = args.indexOf("--append-system-prompt-file");
assertTrue(fileFlag >= 0, "the charter is still mounted via file: " + args);
Path charterFile = Path.of(args.get(fileFlag + 1));
Path dir = charterFile.getParent();
// NOTE: JUnit's own @TempDir provider places worktreeRoot itself under java.io.tmpdir on this
// host, so "not under java.io.tmpdir" is not a meaningful assertion here (it would hold by
// accident of the fixture, not by anything this method does). What this fix actually promises
// is that the directory is created UNDER worktreeRoot specifically — never resolved from the
// no-argument Files.createTempFile default (java.io.tmpdir) the pre-fix code always used — so
// that is the assertion: the parent is exactly worktreeRoot, whatever directory JUnit gave it.
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated directory's parent must be worktreeRoot, not java.io.tmpdir — got "
+ "parent " + dir.getParent());
assertEquals("rwxr-x---", PosixFilePermissions.toString(Files.getPosixFilePermissions(dir)),
"the directory must be group-traversable+readable, owner-only writable");
assertEquals("rw-r-----", PosixFilePermissions.toString(Files.getPosixFilePermissions(charterFile)),
"the charter file must be group-readable, never group-writable");
assertTrue(Files.readString(charterFile).contains("fleet_reply"),
"the charter content itself is unaffected by where it is written");
}
/**
* fleetd #222 acceptance criterion 2: with {@code memberHerdrSocket} configured but NEITHER
* {@code worktreeRoot} nor {@code worktreeGroup} set, the launcher must refuse the spawn rather
* than hand the member a {@code --append-system-prompt-file} path under {@code java.io.tmpdir}
* it cannot read — the member's whole turn contract would never reach it.
*/
@Test
void memberHerdrSocketWithoutWorktreeRootOrGroupRefusesTheSpawn() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher launcher = serviceWithConfig(herdr, charterCfg(),
() -> configWithMemberHerdrSocket(null, null));
IllegalStateException ex = assertThrows(IllegalStateException.class, launcher::spawn,
"a missing worktreeRoot/worktreeGroup must refuse the spawn, not write an unreadable charter");
assertTrue(ex.getMessage().contains("worktreeRoot"),
"the refusal must name the missing config key — got: " + ex.getMessage());
assertFalse(herdr.called("agent.start"),
"the spawn must be refused BEFORE the member is ever started — got calls: " + herdr.calls);
}
/**
* fleetd #222 acceptance criterion 2 (the other missing half): {@code worktreeRoot} set but
* {@code worktreeGroup} missing must ALSO refuse — either one alone is not enough to guarantee
* the member's OS user can read the charter file.
*/
@Test
void memberHerdrSocketWithWorktreeRootButNoGroupRefusesTheSpawn(@TempDir Path worktreeRoot) {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher launcher = serviceWithConfig(herdr, charterCfg(),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), null));
IllegalStateException ex = assertThrows(IllegalStateException.class, launcher::spawn);
assertTrue(ex.getMessage().contains("worktreeGroup"),
"worktreeRoot alone is not enough — got: " + ex.getMessage());
}
/**
* fleetd #222 acceptance criterion 3: with {@code memberHerdrSocket} ABSENT — even when a live,
* non-null {@code config} supplier is threaded through (not merely {@code config == null}, which
* every other test in this file already exercises) — the charter file must still be created
* directly under {@code java.io.tmpdir} via the same no-directory-argument
* {@code Files.createTempFile} call as before this fix, byte-identical to today.
*/
@Test
void memberHerdrSocketAbsentStaysUnderJavaIoTmpdirEvenWithALiveConfigSupplier() throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig config = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, null, null).withDefaults();
serviceWithConfig(herdr, charterCfg(), () -> config).spawn();
List<String> args = spawnedArgs(herdr);
int fileFlag = args.indexOf("--append-system-prompt-file");
assertTrue(fileFlag >= 0);
Path charterFile = Path.of(args.get(fileFlag + 1));
assertTrue(charterFile.startsWith(Path.of(System.getProperty("java.io.tmpdir"))),
"with memberHerdrSocket absent, the charter file must still land directly under "
+ "java.io.tmpdir, unchanged from before this fix");
assertTrue(charterFile.getFileName().toString().startsWith("fleetd-role-charter-"),
"same file-naming scheme as before this fix (no wrapping directory): " + charterFile);
}
}
@@ -114,6 +114,68 @@ class EnvAllowListScrubTest {
assertNull(EnvAllowListScrub.readReport(dir));
}
/**
* fleetd #224 criterion 5 (a gap the #221 reviewer flagged): {@link EnvAllowListScrub#shareWithGroup}
* lists one FLAT level of {@code dir} and shares every file it finds there — which does cover the
* OPTIONAL files a caller may or may not have written before calling it ({@code member-charter.md},
* {@code ide-rules.md} — both written by {@code OpenCodeLauncher#writeConfig}), but nothing
* pinned that down.
* Without this test, a future change that writes a file AFTER the sharing call, or into a
* subdirectory, would pass every existing test while quietly leaving that file unreadable to a
* different-uid member.
*
* <p>This drives {@code shareWithGroup} directly against a directory holding several flat files —
* not only {@code opencode.json}, but also the two optional ones named above plus a third,
* unrelated file, so the assertion is "every flat file", not "the two files someone thought of".
*/
@Test
void shareWithGroupCoversEveryFlatFileIncludingTheOptionalOnes(@TempDir Path dir) throws Exception {
String group = currentUserGroup();
Files.writeString(dir.resolve("opencode.json"), "{}\n");
Files.writeString(dir.resolve("member-charter.md"), "# charter\n");
Files.writeString(dir.resolve("ide-rules.md"), "# ide rules\n");
Files.writeString(dir.resolve("another-flat-file.txt"), "unrelated\n");
EnvAllowListScrub.shareWithGroup(dir, group);
assertEquals("rwxr-x---", java.nio.file.attribute.PosixFilePermissions.toString(
Files.getPosixFilePermissions(dir)),
"the directory itself must be group-traversable+readable, owner-only writable");
for (String name : List.of("opencode.json", "member-charter.md", "ide-rules.md", "another-flat-file.txt")) {
Path file = dir.resolve(name);
assertEquals("rw-r-----", java.nio.file.attribute.PosixFilePermissions.toString(
Files.getPosixFilePermissions(file)),
name + " must be group-readable, never group-writable");
}
}
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225: that reads wherever Maven happened to be started from,
* not the process's own group, and the two diverge outside a home checkout). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
String raw = new String(p.getInputStream().readAllBytes(), StandardCharsets.UTF_8).trim();
out = raw;
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !raw.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/**
* The same equality, for a shell that is INTERACTIVE but NOT a login shell — the shape herdr
* opens on Linux.
@@ -12,20 +12,25 @@ import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.Function;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* CB-633: proves the allow-list scrub is actually WIRED INTO the spawn path — not merely that its
@@ -119,6 +124,18 @@ class HerdrPeerLauncherAllowListWiringTest {
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, allow, List.of(), null);
}
/**
* Same as {@link #allowList()} but with an operator-configured {@code known:} list, so the
* fallback overlay (the non-zsh / no-memberLoginShell path) has something visible to shadow —
* {@link FleetConfig.MemberCredentials#blockedSet()} is {@code known - allow}, so an empty
* {@code known} (what {@link #allowList()} uses) blocks nothing and a fallback test would have
* no sentinel entry to assert on.
*/
private static Supplier<FleetConfig.MemberCredentials> allowListWithKnown(List<String> known) {
return () -> new FleetConfig.MemberCredentials(
FleetConfig.MemberCredentials.POLICY_ALLOW_LIST, List.of(), known, null);
}
/**
* CB-633 follow-up: a name that lives ONLY in {@code memberCredentials.allow:} — no profile
* mentions it — must survive the scrub the real spawn path generates. Calling {@code
@@ -257,6 +274,294 @@ class HerdrPeerLauncherAllowListWiringTest {
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* fleetd #185 stage 2 pin: with {@code memberHerdrSocket:} absent (today's only mode, and the
* default — this host runs no other), the gap detector's WARN/INFO conclusions read exactly as
* they did before this fix. Real path: {@code FLEETD_WORKER_TOKEN} (the test profile's own
* {@code tokenEnv}) is a name the derived allow-list keeps, so it gets the "UNBLOCKED" WARN;
* {@code SOME_UNKNOWN_SECRET_TOKEN} is not derived from anywhere, so it gets the "scrub blanks
* them" INFO. This is the exact shape #185 stage 2 must not touch on this path.
*/
@Test
void gapConclusionsAreByteIdenticalWhenMemberHerdrSocketIsAbsent() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> config(null));
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(messages.contains("memberCredentials gap: 1 credential-shaped env var name(s) are on "
+ "neither known: nor allow: — the derived allow-list keeps them anyway (a "
+ "profile's gitTokenEnv/gitHostEnv/tokenEnv/env: names one, or this spawn "
+ "injects it), so every member pane inherits them UNBLOCKED — [FLEETD_WORKER_TOKEN]. "
+ "Add each to memberCredentials.known (or .allow if a member legitimately needs "
+ "it), or remove it from whatever profile setting derives it in."),
"expected the pre-existing UNBLOCKED WARN unchanged, got: " + messages);
assertTrue(messages.contains("memberCredentials gap: 1 credential-shaped env var name(s) are on "
+ "neither known: nor allow: — [SOME_UNKNOWN_SECRET_TOKEN]. The allow-list scrub "
+ "blanks them anyway (they are not on the derived allow-list), so no member pane "
+ "keeps them; add each to memberCredentials.known or .allow to make that explicit."),
"expected the pre-existing 'scrub blanks them' INFO unchanged, got: " + messages);
}
/**
* fleetd #185 stage 2: with {@code memberHerdrSocket:} configured, member panes run under a
* different OS user — {@link HerdrPeerLauncher#hostEnvNames} describes fleetd's own process, not
* that user's. Neither "inherits them UNBLOCKED" nor "scrub blanks them" is evidence-backed
* there, so neither may print; the single unknown-environment WARN must, naming the config key.
*/
@Test
void gapDetectorReportsUnknownInsteadOfAConclusionWhenMemberHerdrSocketIsConfigured() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(messages.stream().anyMatch(m -> m.contains("memberHerdrSocket")
&& m.contains("UNKNOWN") && m.contains("cannot be verified")),
"expected the unknown-member-environment WARN naming memberHerdrSocket, got: " + messages);
assertFalse(messages.stream().anyMatch(m -> m.contains("UNBLOCKED")),
"the 'inherits them UNBLOCKED' conclusion must not print once the evidence is about "
+ "the wrong (daemon's own) environment — got: " + messages);
assertFalse(messages.stream().anyMatch(m -> m.contains("scrub blanks them")),
"the 'scrub blanks them' conclusion must not print once the evidence is about the "
+ "wrong (daemon's own) environment — got: " + messages);
}
/**
* fleetd #185 stage 2: the unknown-environment WARN is a standing fact about this launcher's
* configuration, not per-spawn news — it must fire once per launcher instance, the same shape as
* every other one-time WARN in this class (e.g. {@code warnNonZsh}).
*/
@Test
void theUnknownEnvironmentWarnFiresOnceNotOncePerSpawn() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.WARN);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
long count = appender.list.stream()
.filter(e -> e.getFormattedMessage().contains("cannot be verified from here"))
.count();
assertEquals(1, count, "the unknown-member-environment WARN must fire once per launcher "
+ "instance, not once per spawn — got " + count + " occurrence(s) among: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
* Hard constraint: the gap detector must never log an env var VALUE, only its NAME. {@code
* SOME_UNKNOWN_SECRET_TOKEN} resolves to a distinctive canary value through the same {@code env}
* lookup the launcher uses elsewhere (SHELL, PATH, token resolution) — proving the value IS
* resolvable does not mean the detector reads it, since {@link HerdrPeerLauncher#hostEnvNames}
* (names only) is its data source, never {@code env.apply(name)} for those names.
*/
@Test
void theGapDetectorNeverLogsAnEnvVarValueOnlyItsName() {
FakeHerdr herdr = new FakeHerdr();
String canary = "sekrit-value-CANARY-9f3a1b7c";
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> config(null), Map.of("SOME_UNKNOWN_SECRET_TOKEN", canary));
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(messages.stream().anyMatch(m -> m.contains("SOME_UNKNOWN_SECRET_TOKEN")),
"expected the credential-shaped NAME to appear in the log, got: " + messages);
assertFalse(messages.stream().anyMatch(m -> m.contains(canary)),
"the log must never contain an env var VALUE, only its NAME — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 1: {@code memberHerdrSocket} configured and {@code
* memberLoginShell} configured as non-zsh must fall back to the sentinel overlay exactly like a
* non-zsh {@code $SHELL} does today — and the "generated ZDOTDIR" INFO must not appear, since no
* scrub actually runs. The WiringLauncher's own {@code env("SHELL")} is deliberately set to
* {@code /bin/zsh} — the OPPOSITE of what {@code memberLoginShell} says — so a launcher that
* (incorrectly) fell back to fleetd's own {@code $SHELL} here would wrongly pass the gate and
* fail this test.
*/
@Test
void memberHerdrSocketWithNonZshMemberLoginShellFallsBackToTheOverlay() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")),
"/bin/zsh", null,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/bash"));
List<String> messages = spawnAndCaptureLogs(launcher);
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"a non-zsh memberLoginShell must fall back to the CB-596 sentinel overlay, exactly "
+ "like a non-zsh $SHELL does when memberHerdrSocket is absent");
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub directory may be generated when the configured member login shell is not zsh");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"the 'generated ZDOTDIR' INFO must not appear when the scrub never runs — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 2: {@code memberHerdrSocket} configured and NO
* {@code memberLoginShell} configured must fall back exactly like criterion 1 above — AND
* fleetd's own {@code $SHELL} must never even be consulted (not merely "not decisive"). The
* fixture's {@code env} function reports {@code /bin/zsh} for {@code SHELL} — a value that would
* WRONGLY pass the zsh gate if the fix regressed to reading it — while flagging whether it was
* ever asked for at all, so this test fails loudly on either kind of regression.
*/
@Test
void memberHerdrSocketWithNoMemberLoginShellFallsBackAndNeverConsultsFleetdsOwnShell() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
if ("SHELL".equals(name)) {
shellQueried.set(true);
return "/bin/zsh"; // would wrongly pass the zsh gate if this ever leaked through
}
return null;
};
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")), env,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", null));
List<String> messages = spawnAndCaptureLogs(launcher);
assertFalse(shellQueried.get(), "fleetd's own $SHELL must never be consulted once "
+ "memberHerdrSocket is configured — only memberLoginShell: may decide the gate");
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"no memberLoginShell configured must fall back to the sentinel overlay, same as a "
+ "configured non-zsh shell");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"no scrub may run without a configured memberLoginShell — got: " + messages);
}
/**
* fleetd #213 defect 2, acceptance criterion 3: with {@code memberHerdrSocket} configured, a
* zsh {@code memberLoginShell}, and {@code worktreeRoot}/{@code worktreeGroup} both configured,
* the generated scrub directory must live under {@code worktreeRoot} — NEVER under {@code
* java.io.tmpdir}, which is fleetd's own 0700 temp dir and unreadable by the member's different
* OS user. {@code worktreeGroup} is set to the CURRENT process's own primary group so {@code
* EnvAllowListScrub}'s group-sharing step resolves on whatever host runs this test, rather than
* hardcoding a group name that may not exist here.
*/
@Test
void memberHerdrSocketWithZshMemberLoginShellPutsTheScrubOutsideJavaIoTmpdir(@TempDir Path worktreeRoot)
throws IOException {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash-should-be-ignored", null,
() -> configWithMemberHerdrSocketRootAndGroup("/tmp/other-user-herdr.sock", "/bin/zsh",
worktreeRoot.toString(), group));
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
String zdotdir = launcher.env.get("ZDOTDIR");
assertNotNull(zdotdir, "a zsh memberLoginShell with worktreeRoot+worktreeGroup configured "
+ "must still generate a ZDOTDIR");
Path dir = Path.of(zdotdir);
// The immediate PARENT is asserted (not merely startsWith(java.io.tmpdir)), because a
// JUnit @TempDir is itself carved out of the JVM's java.io.tmpdir — startsWith alone would
// pass by coincidence of the test fixture, not because the launcher used worktreeRoot.
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated scrub directory's parent must be the configured worktreeRoot, not "
+ "System.getProperty(\"java.io.tmpdir\") — got parent " + dir.getParent());
}
/**
* fleetd #213, acceptance criterion 4: with {@code memberHerdrSocket} absent (today's only
* mode), behaviour must be byte-identical to before this fix — fleetd's own {@code $SHELL}
* still decides the gate (proven here, not merely assumed, by flagging the lookup), and the
* scrub still lands under {@code java.io.tmpdir}.
*/
@Test
void memberHerdrSocketAbsentStillConsultsFleetdsOwnShellAndBehavesAsBefore() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
if ("SHELL".equals(name)) {
shellQueried.set(true);
return "/bin/zsh";
}
return null;
};
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), env, null); // memberHerdrSocket absent
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
assertTrue(shellQueried.get(), "with memberHerdrSocket absent, fleetd's own $SHELL must "
+ "still decide the zsh gate, unchanged from before this fix");
String zdotdir = launcher.env.get("ZDOTDIR");
assertNotNull(zdotdir, "SHELL=/bin/zsh with memberHerdrSocket absent must still generate a "
+ "ZDOTDIR, as before this fix");
Path dir = Path.of(zdotdir);
assertTrue(dir.startsWith(Path.of(System.getProperty("java.io.tmpdir"))),
"with memberHerdrSocket absent the scrub directory must still be generated under "
+ "java.io.tmpdir, unchanged from before this fix: " + dir);
}
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225). Those two coincide only by accident: a directory's
* group is whichever group happened to own the path Maven was started from — {@code staff} in
* a home checkout, {@code wheel} under {@code /private/tmp} on macOS — and the fix-up this test
* exercises then fails for real when the operator is not a member of that borrowed group,
* exactly the case {@code assumeTrue(view != null, ...)} never covered (it only detects a
* filesystem with no POSIX groups at all, not a resolvable-but-wrong one). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
try (java.io.BufferedReader r = new java.io.BufferedReader(
new java.io.InputStreamReader(p.getInputStream(), java.nio.charset.StandardCharsets.UTF_8))) {
out = r.lines().collect(java.util.stream.Collectors.joining("\n")).trim();
}
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !out.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/** Spawn once through the real launcher path, capturing every INFO+ line this class logs. */
private static List<String> spawnAndCaptureLogs(HerdrPeerLauncher launcher) {
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList();
}
private static String readAll(Path p) {
try {
return Files.readString(p);
@@ -294,12 +599,40 @@ class HerdrPeerLauncherAllowListWiringTest {
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds, String shell,
Supplier<Set<String>> hostEnvNames, Supplier<FleetConfig> config) {
this(herdr, creds, shell, hostEnvNames, config, Map.of());
}
/**
* Plus a host-env value map (name → value), resolved through the same {@code env} lookup
* every adapter uses for {@code SHELL}/{@code PATH}/token resolution — fleetd #185 stage 2's
* "the gap detector never logs a value" tests use this to prove a value that IS resolvable
* for a credential-shaped name never reaches the log, since the detector only ever reads
* {@code hostEnvNames} (names), never {@code env.apply(name)} (values), for those names.
*/
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds, String shell,
Supplier<Set<String>> hostEnvNames, Supplier<FleetConfig> config,
Map<String, String> extraEnvValues) {
super("test", new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("test", profile()), "test",
name -> "SHELL".equals(name) ? shell : null,
name -> "SHELL".equals(name) ? shell
: (extraEnvValues != null && extraEnvValues.containsKey(name))
? extraEnvValues.get(name) : null,
0, () -> 0L, () -> { }, null, creds, hostEnvNames, config);
}
/**
* Full control over the {@code env} lookup, bypassing the {@code shell}/{@code
* extraEnvValues} convenience above entirely — fleetd #213's "fleetd's own $SHELL must
* never be consulted" tests need to OBSERVE whether {@code SHELL} was ever looked up, which
* a plain value substitution cannot do.
*/
WiringLauncher(FakeHerdr herdr, Supplier<FleetConfig.MemberCredentials> creds,
Function<String, String> env, Supplier<FleetConfig> config) {
super("test", new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("test", profile()), "test",
env, 0, () -> 0L, () -> { }, null, creds, null, config);
}
@Override
protected Launch buildLaunch(FleetConfig.Profile cfg, LaunchSpec spec) {
Map<String, String> launchEnv = baseEnv(cfg);
@@ -321,6 +654,82 @@ class HerdrPeerLauncherAllowListWiringTest {
null, null, null, null, null).withDefaults();
}
/**
* fleetd #185 stage 2: a config with {@code memberHerdrSocket:} set — member panes run on a
* second herdr owned by a different OS user, so {@link HerdrPeerLauncher#hostEnvNames} no
* longer describes what a member pane inherits.
*/
private static FleetConfig configWithMemberHerdrSocket(String memberHerdrSocket) {
return new FleetConfig(null, null, memberHerdrSocket, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null).withDefaults();
}
/**
* fleetd #213: as {@link #configWithMemberHerdrSocket(String)}, plus the {@code
* memberLoginShell:} the member's OS user actually runs — the config key {@link
* HerdrPeerLauncher#applyEnvironmentAllowListPolicy} must consult instead of fleetd's own
* {@code $SHELL} once {@code memberHerdrSocket} is configured.
*/
private static FleetConfig configWithMemberHerdrSocket(String memberHerdrSocket, String memberLoginShell) {
return new FleetConfig(
null, // bind
null, // herdrSocket
memberHerdrSocket, // memberHerdrSocket
Map.of(), // profiles
null, // guard
null, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
null, // worktreeGroup
memberLoginShell // memberLoginShell
).withDefaults();
}
/**
* fleetd #213: as above, plus {@code worktreeRoot:}/{@code worktreeGroup:} — both required for
* the ZDOTDIR scrub to run at all once {@code memberHerdrSocket} is configured; either missing
* falls back to the sentinel overlay, same as a non-zsh {@code memberLoginShell}.
*/
private static FleetConfig configWithMemberHerdrSocketRootAndGroup(String memberHerdrSocket,
String memberLoginShell, String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
memberHerdrSocket, // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
memberLoginShell // memberLoginShell
).withDefaults();
}
/** The generated directory is a temp directory; make sure the test does not leave a pile. */
@Test
void theGeneratedDirectoryIsRemovedWhenThePaneIsStopped() {
@@ -1,30 +1,43 @@
package dev.ltms.fleet.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFilePermissions;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
@@ -383,6 +396,46 @@ class OpenCodeLauncherTest {
assertNotNull(ex.getMessage());
}
/**
* fleetd #176 fix 2's per-adapter guard: {@code StatusRefiner.classify} is written for the
* Claude Code TUI only, and the spawn-readiness gate must never run it against an opencode
* pane. {@code agentType("opencode")} deliberately satisfies the OTHER guard (the corroborating
* liveness check) so it cannot be what makes this test pass — only the {@code namePrefix}
* check can be. If that check were ever removed, this pane's {@code ❯} content would refine
* straight to IDLE and the gate would report ready before the backend actually was.
*/
@Test
void opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType("opencode"); // non-null — satisfies the liveness guard on its own
herdr.readText("some-host:~ user$ ❯ "); // content StatusRefiner.classify reads as IDLE
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root, root);
assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000,
"an opencode pane must never refine to injectable, however its content reads — "
+ "the gate has to wait out the full timeout: clock only reached " + clock[0]);
// A blanket "agent.read is never called" does not hold here: the timeout path itself reads
// the pane tail (source=recent) for its own log message, on every timeout, regardless of
// adapter — see HerdrPeerLauncher.readPaneQuietly. So assert on the refiner's OWN probe
// source (StatusRefiner.PROBE_SOURCE = "detection") instead — that call happens only inside
// StatusRefiner.refine, so its absence proves the refiner itself was never reached for this
// opencode pane, not merely that its answer was discarded.
long detectionReads = herdr.calls.stream()
.filter(c -> c.method().equals("agent.read"))
.filter(c -> "detection".equals(((Map<?, ?>) c.params()).get("source")))
.count();
assertEquals(0, detectionReads,
"the refiner's own pane probe (source=detection) must never run against an "
+ "opencode pane — the namePrefix guard has to stop it before that call");
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
@@ -610,4 +663,474 @@ class OpenCodeLauncherTest {
assertTrue(json.path("mcp").path("intellij").isMissingNode(),
"no IDE server when ideMcpUrl is unset");
}
// --- fleetd #219: config root + discovery root under memberHerdrSocket ------------------------
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
private static FleetConfig configWithMemberHerdrSocket(String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
"/tmp/other-user.sock", // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
null // memberLoginShell
).withDefaults();
}
private static OpenCodeLauncher serviceWithConfig(FakeHerdr herdr, Path configRoot, Path discoveryRoot,
FleetConfig.Profile cfg, Supplier<FleetConfig> config) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, configRoot, discoveryRoot, null, null, config);
}
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225). Those two coincide only by accident: a directory's
* group is whichever group happened to own the path Maven was started from — {@code staff} in
* a home checkout, {@code wheel} under {@code /private/tmp} on macOS — and the fix-up this test
* exercises then fails for real when the operator is not a member of that borrowed group,
* exactly the case {@code assumeTrue(view != null, ...)} never covered (it only detects a
* filesystem with no POSIX groups at all, not a resolvable-but-wrong one). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
try (java.io.BufferedReader r = new java.io.BufferedReader(
new java.io.InputStreamReader(p.getInputStream(), java.nio.charset.StandardCharsets.UTF_8))) {
out = r.lines().collect(java.util.stream.Collectors.joining("\n")).trim();
}
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !out.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/**
* fleetd #219 site 1, acceptance criterion 1: with {@code memberHerdrSocket} configured and both
* {@code worktreeRoot}/{@code worktreeGroup} set, the generated {@code opencode.json} directory
* lives under {@code worktreeRoot} — NEVER under the injected {@code configRoot} (standing in for
* {@code java.io.tmpdir}, fleetd's own 0700 temp dir, unreadable by the member's different OS
* user) — and is shared read-only with the group via the SAME mechanism (fleetd #213's {@link
* EnvAllowListScrub#shareWithGroup}) the ZDOTDIR scrub uses.
*/
@Test
void memberHerdrSocketWithWorktreeRootAndGroupPutsConfigDirUnderWorktreeRootAndSharesIt(
@TempDir Path configRoot, @TempDir Path worktreeRoot) throws Exception {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), group)).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "the profile still needs a config file");
Path cfgFile = Path.of(cfgPath);
Path dir = cfgFile.getParent();
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated directory's parent must be worktreeRoot, not the injected configRoot "
+ "standing in for java.io.tmpdir — got parent " + dir.getParent());
assertFalse(dir.startsWith(configRoot),
"the generated directory must NOT be created under configRoot when memberHerdrSocket "
+ "is configured: " + dir);
assertEquals("rwxr-x---", PosixFilePermissions.toString(Files.getPosixFilePermissions(dir)),
"the directory must be group-traversable+readable, owner-only writable");
assertEquals("rw-r-----", PosixFilePermissions.toString(Files.getPosixFilePermissions(cfgFile)),
"opencode.json must be group-readable, never group-writable");
}
/**
* fleetd #219 site 1, acceptance criterion 2: with {@code memberHerdrSocket} configured but
* NEITHER {@code worktreeRoot} nor {@code worktreeGroup} set, the launcher must refuse the spawn
* rather than write a config under {@code java.io.tmpdir} the member cannot read — that member
* would occupy a pane and never become deliverable, with nothing pointing at the real cause.
*/
@Test
void memberHerdrSocketWithoutWorktreeRootOrGroupRefusesTheSpawn(@TempDir Path configRoot) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(null, null));
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> launcher.spawn(new SpawnRequest(null, null, null)),
"a missing worktreeRoot/worktreeGroup must refuse the spawn, not write an unreadable config");
assertTrue(ex.getMessage().contains("worktreeRoot"),
"the refusal must name the missing config key — got: " + ex.getMessage());
assertFalse(herdr.called("tab.create"),
"the spawn must be refused BEFORE any pane is created — got calls: " + herdr.calls);
}
/**
* fleetd #219 site 1, acceptance criterion 2 (the other missing half): {@code worktreeRoot} set
* but {@code worktreeGroup} missing must ALSO refuse — either one alone is not enough to
* guarantee the member's OS user can read the directory.
*/
@Test
void memberHerdrSocketWithWorktreeRootButNoGroupRefusesTheSpawn(
@TempDir Path configRoot, @TempDir Path worktreeRoot) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), null));
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> launcher.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("worktreeGroup"),
"worktreeRoot alone is not enough — got: " + ex.getMessage());
}
/**
* fleetd #219 site 1, acceptance criterion 3: with {@code memberHerdrSocket} ABSENT — even when
* a live, non-null {@code config} supplier is threaded through (not merely {@code config == null},
* which every other test in this file already exercises) — the generated directory must still
* land directly under the injected {@code configRoot}, byte-identical to before this fix.
*/
@Test
void memberHerdrSocketAbsentStaysUnderConfigRootEvenWithALiveConfigSupplier(@TempDir Path configRoot)
throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig config = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, null, null).withDefaults();
serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null), () -> config)
.spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath);
assertTrue(Path.of(cfgPath).startsWith(configRoot),
"with memberHerdrSocket absent, the config directory must still be created directly "
+ "under configRoot, unchanged from before this fix");
}
/**
* fleetd #219 site 2: under {@code memberHerdrSocket}, opencode session discovery must be
* declared unavailable rather than silently scanning fleetd's own {@code discoveryRoot} — which,
* under this config key, is NOT where the member's opencode actually writes its session
* database. This test proves the gate is real, not merely "no record yet": a matching record IS
* written to {@code discoveryRoot} (the exact fixture {@link
* #theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears} proves discovery would
* otherwise find), and {@code agentSessionId()} must still return {@code null} — proving the
* gate, not a coincidental absence of data, is what produced the null. One WARN is also logged,
* exactly once even across repeated calls.
*/
@Test
void discoveryIsUnavailableUnderMemberHerdrSocketEvenWhenARecordExists(
@TempDir Path configRoot, @TempDir Path worktreeRoot, @TempDir Path discRoot) throws Exception {
String group = currentUserGroup();
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_should_be_hidden", "/work/dir", 1000L);
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, configRoot, discRoot,
null, null, () -> configWithMemberHerdrSocket(worktreeRoot.toString(), group));
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
PeerHandle handle;
try {
handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
assertNull(handle.agentSessionId(),
"memberHerdrSocket configured: discovery must stay unavailable even though a "
+ "matching record exists in discoveryRoot");
assertNull(handle.agentSessionId(), "the gate must hold on a second call too");
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
List<String> warnings = appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.toList();
assertEquals(1, warnings.size(),
"exactly one WARN across two agentSessionId() calls — got: " + warnings);
assertTrue(warnings.get(0).contains("memberHerdrSocket"),
"the WARN must name memberHerdrSocket as the reason — got: " + warnings.get(0));
}
// --- fleetd #175: opencode silently substitutes a model on an unknown -m flag ----------------
private static OpenCodeLauncher serviceWithSink(FakeHerdr herdr, Path configRoot, Path discoveryRoot,
FleetConfig.Profile cfg, ExhaustionSink sink) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, configRoot, discoveryRoot,
null, null, null, sink);
}
/**
* fleetd #175 acceptance criterion 1: the check must fire on the REAL late-resolve path —
* {@code SessionManager}'s {@code get}/{@code roster} re-polling a retained {@link PeerHandle}
* (fleetd #209) — not on a handle built and queried directly. A handle whose row does not exist
* yet, then does, proves the check runs exactly where production runs it: PR #203 shipped a
* check that ran inside {@code spawn()}, before opencode had written the row, and every test
* passed anyway because none of them drove it through this path. This test would have caught
* that: {@code sessions.get(...)} before the row exists must show no mismatch, and the SAME
* call, re-driven after the row appears, must be what fires the sink.
*/
@Test
void theRealSessionManagerLateResolvePathCatchesAModelMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
FakeHerdr herdr = new FakeHerdr();
// xf's real shape (fleetd #175): weight:80, model "opencode/nemotron-3-ultra-free", no
// credentialId — the profile that actually escaped the fleet's accounting.
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(target + "|" + reason);
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, sink);
SessionManager sessions = new SessionManager(launcher);
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
// Real late-resolve path, driven BEFORE opencode has written its session row — same shape
// as production the instant a pane goes ready.
Optional<MemberSession> beforeRow = sessions.get(acquired.paneId());
assertTrue(beforeRow.isPresent());
assertNull(beforeRow.get().agentSessionId(), "no opencode row yet");
assertTrue(exhausted.isEmpty(), "no row yet → nothing to compare, the sink must stay silent");
// opencode writes its row late, running gpt-5.6-sol (a PAID credential) instead of the
// withdrawn free model the profile actually asked for — the exact fleetd #175 scenario.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
// Drive the SAME real late-resolve path again: sessions.get() -> resolveAgentSessionId ->
// the retained PeerHandle's agentSessionId() -> discovery -> checkModelMatch, all in one
// call, unmodified SessionManager code from fleetd #209.
Optional<MemberSession> afterRow = sessions.get(acquired.paneId());
assertEquals("ses_x", afterRow.get().agentSessionId(),
"the id itself still resolves correctly alongside the model check");
assertEquals(1, exhausted.size(), "the mismatch must fire exactly once through the real path");
assertTrue(exhausted.get(0).contains(cfg.model()), "reports the requested model: " + exhausted.get(0));
assertTrue(exhausted.get(0).contains("gpt-5.6-sol"), "reports the actual model: " + exhausted.get(0));
}
/** THE TRAP, row 1: a provider-prefixed request matching the DB's id AND provider is a match. */
@Test
void aProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "id and provider both match → never a mismatch: " + exhausted);
}
/** THE TRAP, row 2: same shape as row 1 with a different provider/id pair (the gx profile). */
@Test
void aGxProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("gx/deepseek-v4-flash", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "id and provider both match → never a mismatch: " + exhausted);
}
/**
* Incomplete evidence, not THE TRAP's provider mismatch: opencode's {@code model} JSON had an
* {@code id} but no {@code providerID} at all (a real shape {@code parseModel} accepts — see
* {@code OpenCodeSessionDiscoveryTest}). The id matches; the provider dimension is simply
* unknown, not contradicted. A provider-prefixed profile must NOT be quarantined on this —
* that would quarantine on incomplete evidence, which acceptance rule 4 forbids.
*/
@Test
void aMissingProviderIdInTheEvidenceIsUnknownNotAMismatchWhenTheIdMatches(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(),
"id matches and provider is simply unknown (absent), never a mismatch: " + exhausted);
}
/**
* The other half of the same fix: a missing {@code providerID} must NOT blind the check to a
* genuine id mismatch. This is what proves the fix narrows the comparison rather than switching
* the whole check off whenever {@code providerID} happens to be absent.
*/
@Test
void aMissingProviderIdInTheEvidenceStillCatchesARealIdMismatch(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\"}");
assertEquals("ses_x", handle.agentSessionId());
assertEquals(1, exhausted.size(),
"the id genuinely differs, so this must still quarantine even with providerID absent: "
+ exhausted);
assertTrue(exhausted.get(0).contains("gpt-5.6-sol") && exhausted.get(0).contains(cfg.model()),
"reports both the requested and actual model: " + exhausted.get(0));
}
/**
* THE TRAP, row 3: a profile that names no provider prefix (bare {@code "deepseek-v4-flash"})
* must match on id alone — the profile never asked for a specific provider, so opencode
* resolving it to {@code gx} is not evidence of anything wrong.
*/
@Test
void aBareModelWithNoProviderPrefixMatchesOnIdAloneAndIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("deepseek-v4-flash", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(),
"no provider was requested, so opencode's own provider resolution is not a mismatch: "
+ exhausted);
}
/**
* THE TRAP, row 4 (not a row in the table, the reason the table exists): a mismatched id, with
* the profile's requested and the actual model both named in the ERROR and the sink's reason —
* the exact fleetd #175 scenario (a withdrawn model silently falls back to a paid one).
*/
@Test
void aRealIdMismatchLogsAnErrorNamingBothModelsAndQuarantinesThroughTheSink(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(target + "|" + reason);
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
PeerHandle handle;
try {
handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertEquals("ses_x", handle.agentSessionId());
} finally {
logger.detachAppender(appender);
}
assertEquals(1, exhausted.size(), "a real id mismatch must reach the sink exactly once");
assertTrue(exhausted.get(0).contains("opencode/nemotron-3-ultra-free"),
"the sink reason must name the requested model: " + exhausted.get(0));
assertTrue(exhausted.get(0).contains("gpt-5.6-sol"),
"the sink reason must name the actual model: " + exhausted.get(0));
List<ILoggingEvent> errors = appender.list.stream()
.filter(e -> e.getLevel() == Level.ERROR)
.toList();
assertEquals(1, errors.size(), "exactly one ERROR for the mismatch: " + appender.list);
String message = errors.get(0).getFormattedMessage();
assertTrue(message.contains("opencode/nemotron-3-ultra-free") && message.contains("gpt-5.6-sol")
&& message.contains(cfg.profile()),
"the ERROR must name the requested model, the actual model, AND the profile: " + message);
// The check runs at most once per handle even if agentSessionId() is polled again.
handle.agentSessionId();
assertEquals(1, exhausted.size(), "no duplicate quarantine on a repeated call");
}
/**
* fleetd #175's central safety rule: absent, empty, or unparseable model evidence is UNKNOWN,
* never a mismatch — it must never quarantine a working profile. Covers every "no real
* evidence yet" shape: no row at all, a row with a null model, and a row with unparseable JSON.
*/
@Test
void unknownOrUnparseableModelEvidenceNeverQuarantines(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
// No row yet at all.
assertNull(handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "no row yet is UNKNOWN, not a mismatch: " + exhausted);
// A row exists (for a DIFFERENT directory) so a poll on ours still finds nothing.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_other", "/work/other", 500L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertNull(handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "a row for a different directory is UNKNOWN, not a mismatch: "
+ exhausted);
}
/**
* A profile with no {@code model:} configured has nothing to compare against — the check must
* stay silent no matter what opencode actually ran, since there is no requested value to be
* wrong about.
*/
@Test
void aProfileWithNoConfiguredModelIsNeverCheckedForAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg(null, null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"anything-at-all\",\"providerID\":\"anyone\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "no model configured → nothing to compare: " + exhausted);
}
}
@@ -27,19 +27,30 @@ class OpenCodeSessionDiscoveryTest {
* and insert one row. Static so {@link OpenCodeLauncherTest} can reuse it.
*/
static void writeRecord(Path root, String id, String directory, long timeUpdated) throws Exception {
writeRecord(root, id, directory, timeUpdated, null);
}
/**
* Same as {@link #writeRecord(Path, String, String, long)}, plus the {@code model} column
* fleetd #175 reads — the raw JSON opencode writes, e.g.
* {@code {"id":"gpt-5.6-terra","providerID":"openai"}}. {@code modelJson} may be {@code null}.
*/
static void writeRecord(Path root, String id, String directory, long timeUpdated, String modelJson)
throws Exception {
Path db = root.resolve("opencode.db");
try (Connection connection = DriverManager.getConnection("jdbc:sqlite:" + db)) {
try (Statement statement = connection.createStatement()) {
statement.execute("CREATE TABLE IF NOT EXISTS session ("
+ "id TEXT PRIMARY KEY, directory TEXT, time_updated INTEGER)");
+ "id TEXT PRIMARY KEY, directory TEXT, time_updated INTEGER, model TEXT)");
}
// Bound parameters, not string interpolation: the class under test uses a
// PreparedStatement, and a hand-escaped INSERT here is a pattern someone copies out.
try (PreparedStatement insert = connection.prepareStatement(
"INSERT INTO session (id, directory, time_updated) VALUES (?, ?, ?)")) {
"INSERT INTO session (id, directory, time_updated, model) VALUES (?, ?, ?, ?)")) {
insert.setString(1, id);
insert.setString(2, directory);
insert.setLong(3, timeUpdated);
insert.setString(4, modelJson);
insert.executeUpdate();
}
}
@@ -134,4 +145,96 @@ class OpenCodeSessionDiscoveryTest {
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"an unreadable database resolves to null, not an exception");
}
// --- fleetd #175: actualModelForDirectory / the model JSON column ---------------------------
@Test
void parsesTheModelJsonIntoProviderAndId(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
OpenCodeSessionDiscovery.ActualModel actual =
new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a");
assertNotNull(actual, "a well-formed model JSON parses");
assertEquals("gpt-5.6-terra", actual.id());
assertEquals("openai", actual.provider());
}
@Test
void prefersTheModelOfTheMostRecentlyUpdatedRow(@TempDir Path root) throws Exception {
writeRecord(root, "ses_old", "/w/a", 1000L, "{\"id\":\"old-model\",\"providerID\":\"openai\"}");
writeRecord(root, "ses_new", "/w/a", 5000L, "{\"id\":\"new-model\",\"providerID\":\"openai\"}");
assertEquals("new-model",
new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a").id(),
"the model of the row with the highest time_updated wins, same as the id");
}
@Test
void aNonMatchingDirectoryYieldsUnknownModelRatherThanAMismatch(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"x\",\"providerID\":\"y\"}");
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/other"),
"no row for this cwd yet → unknown, not a wrong model");
}
@Test
void aNullModelColumnYieldsUnknownWithoutThrowing(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, null);
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"),
"a row with no model value yet is unknown, not a mismatch");
}
@Test
void unparseableModelJsonYieldsUnknownWithoutThrowing(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "this is not json");
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"),
"JSON that fails to parse resolves to unknown, never an exception");
}
@Test
void modelJsonMissingIdYieldsUnknown(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"providerID\":\"openai\"}");
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"),
"no id in the JSON → unknown, since id is what a caller actually compares");
}
@Test
void aMissingDatabaseYieldsUnknownModelWithoutThrowing(@TempDir Path root) {
assertNull(new OpenCodeSessionDiscovery(root).actualModelForDirectory("/w/a"));
}
/**
* fleetd #175 site 2: a schema that predates the {@code model} column entirely — the exact
* shape a real opencode upgrade/downgrade could produce. This must degrade to UNKNOWN for the
* model, and — the property that actually matters — must NOT take id resolution down with it.
* A combined single query for both columns would fail this test; that is why
* {@link OpenCodeSessionDiscovery#actualModelForDirectory} runs its own independent query.
*/
@Test
void aMissingModelColumnYieldsUnknownButIdResolutionStillWorks(@TempDir Path root) throws Exception {
Path db = root.resolve("opencode.db");
try (Connection connection = DriverManager.getConnection("jdbc:sqlite:" + db)) {
try (Statement statement = connection.createStatement()) {
statement.execute("CREATE TABLE session (id TEXT PRIMARY KEY, directory TEXT, "
+ "time_updated INTEGER)");
}
try (PreparedStatement insert = connection.prepareStatement(
"INSERT INTO session (id, directory, time_updated) VALUES (?, ?, ?)")) {
insert.setString(1, "ses_aaa");
insert.setString(2, "/w/a");
insert.setLong(3, 1000L);
insert.executeUpdate();
}
}
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
assertEquals("ses_aaa", discovery.sessionIdForDirectory("/w/a"),
"id resolution must survive a database with no model column at all");
assertNull(discovery.actualModelForDirectory("/w/a"),
"no model column → unknown, not a throw and not a mismatch");
}
}
@@ -688,6 +688,32 @@ class MessageServiceTest {
assertFailedTicket(third, "agent target term_a not found");
}
// --- #137 follow-up: abandon() must not guess when more than one task is open ---------------
//
// A test combining a genuine stranded reply (hasStrandedReply(T)==true) with two simultaneously
// open matching tasks was attempted here and removed after investigation showed the combination
// is not reachable through the public API today, not merely hard to time right:
//
// This class has exactly two call sites that ever hold a target's entry in the session-lock map
// (send() and answer()), and both open a Rendezvous waiter for that same target as the first thing
// they do after acquiring the lock, holding lock and waiter together for their whole critical
// section. So "the lock is held" and "a live waiter is open" are the same fact throughout this
// class. reply()'s fast path always resolves a currently-open waiter directly instead of
// stranding — so a strand can only be created while NO task is accepted (lock free), and the
// instant the lock is next taken (by any parked matching task's own send(), the moment it is
// scheduled), that acceptance clears strandedReplies again (see send()'s CB-640 comment) before
// abandon() can ever observe both facts together. Confirmed empirically too: an earlier version of
// this test stranded a reply, then created an "accepted" task (awaitWaiting()) followed by a
// "parked" one — and the accepted task's own acceptance silently cleared the strand it was
// supposed to be racing against, so the parked task came back WORKER_FAILED instead of DONE, not
// because the fix was missing but because the test's premise could not be constructed.
//
// The reachable half — matching.size() >= 2 alone, no strand — is exactly what
// abandonFailsEveryPendingAsyncTicketForTheReleasedTarget already covers (all fail, none guess).
// The oldest-wins code in abandon() stays as defence in depth (see its own javadoc) against a
// regression that would make the conjunction reachable, e.g. clearing strandedReplies on a
// narrower condition than "any acceptance" — not because this suite exercises it today.
@Test
void abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
@@ -43,6 +43,7 @@ class ReplyPushLoopTest {
private static final String PRIMARY = "term_primary";
private static final String WORKER = "term_worker";
private static final String WORKER2 = "term_worker2";
private static final String OTHER_PRIMARY = "term_other_primary";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
@@ -529,6 +530,106 @@ class ReplyPushLoopTest {
assertTrue(nudge.contains("task-1"), "the ticket must not be dropped: " + nudge);
}
// --- CB-201: backend credential outage nudges -----------------------------------------------
@Test
void oneBackendIncidentForTwoWorkersOfOneLeadSendsOnceAndIsOneShot() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
registry.recordDelegation(WORKER, PRIMARY);
registry.recordDelegation(WORKER2, PRIMARY);
var loop = loop(2, 100);
loop.onBackendIncident("incident-1", List.of(WORKER, WORKER2), "cred-a",
List.of("terra", "codex"), 47);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one lead gets one outage injection");
Thread.sleep(200);
assertEquals(1, rec.sendCount(), "two workers for one lead must send only one outage notice");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("cred-a"));
assertTrue(nudge.contains("codex, terra"));
assertTrue(nudge.contains(WORKER) && nudge.contains(WORKER2));
assertTrue(nudge.contains("47 remaining seconds"));
assertTrue(nudge.contains("fleet_list"));
loop.onBackendIncident("incident-1", List.of(WORKER, WORKER2), "cred-a",
List.of("terra", "codex"), 47);
Thread.sleep(250);
assertEquals(1, rec.sendCount(), "a delivered incident id must never be sent again");
}
@Test
void oneBackendIncidentSendsOnceToEachAffectedLead() throws Exception {
var rec = recordingClient();
rec.sendLatch = new CountDownLatch(2);
agents = new AgentControl(rec);
registry.recordDelegation(WORKER, PRIMARY);
registry.recordDelegation(WORKER2, OTHER_PRIMARY);
var loop = loop(1, 50);
loop.onBackendIncident("incident-2", List.of(WORKER, WORKER2), "cred-a", List.of("terra"), 60);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "each affected lead gets its one notice");
assertEquals(2, rec.sendCount());
}
@Test
void backendIncidentWaitsForBusyLeadThenCombinesWithFailedTicket() throws Exception {
var rec = new BusyThenIdleHerdrClient(2);
agents = new AgentControl(rec);
var loop = loop(1, 50);
loop.onTicketTerminal("task-failed", WORKER, true);
loop.onBackendIncident("incident-3", List.of(WORKER), "cred-b", List.of("terra"), 31);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "busy lead should receive deferred combined nudge");
assertEquals(1, rec.sendCount(), "ticket and outage must use the same injection");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-failed") && nudge.contains("cred-b"),
"the one nudge must name both the failed ticket and outage: " + nudge);
}
@Test
void failedBackendIncidentSendRetriesWithinTheExistingReminderCap() throws Exception {
var rec = new ThrowOnceHerdrClient();
agents = new AgentControl(rec);
var loop = loop(2, 50);
loop.onBackendIncident("incident-4", List.of(WORKER), "cred-c", List.of("terra"), 20);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "failed first send must retry under maxReminders");
assertEquals(2, rec.promptAttempts.get(), "one failed send and one retry use the existing cap");
}
@Test
void unmappedBackendTargetUsesTheKnownLeadSchedule() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 50);
loop.onBackendTargetUnmapped(WORKER, "no credential matched");
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
String nudge = ((Map<?, ?>) rec.sentParams().getFirst().getValue()).get("text").toString();
assertEquals("A backend error on worker term_worker could not be mapped to a credential "
+ "(no credential matched) — no cool-off was applied. "
+ "Run fleet_list to check that worker.",
nudge);
}
@Test
void stopClearsPendingBackendIncidents() {
agents = agentWithStatus("idle");
var loop = loop(1, 100_000);
loop.onBackendIncident("incident-stop", List.of(WORKER), "cred-stop", List.of("terra"), 60);
loop.stop();
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0, 0),
"stop must clear a pending backend incident as well as older sources");
}
// --- CB-590 follow-up: per-source reminder budgets — the regression this round exists for ---
@Test
@@ -928,6 +1029,31 @@ class ReplyPushLoopTest {
return new RecordingHerdrClient();
}
/** Fails its first prompt call, then records the retry. */
private static final class ThrowOnceHerdrClient implements HerdrClient {
private final AtomicInteger promptAttempts = new AtomicInteger();
private final CountDownLatch sendLatch = new CountDownLatch(1);
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
return MAPPER.createObjectNode().set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY).put("agent_status", "idle"));
}
if ("agent.prompt".equals(method) && promptAttempts.getAndIncrement() == 0) {
throw new IllegalStateException("first prompt fails");
}
if ("agent.prompt".equals(method)) {
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
@Override
public void close() {
}
}
/**
* Thread-safe fake that reports {@code working} (not injectable) for its first
* {@code busyChecks} status calls, then {@code idle} forever after — used to prove work queued
@@ -218,9 +218,17 @@ class FleetAppTest {
HttpResponse<String> res = req(port, "GET", "/members");
assertEquals(200, res.statusCode());
JsonNode workers = mapper.readTree(res.body()).get("workers");
assertEquals(1, workers.size());
JsonNode w = workers.get(0);
JsonNode body = mapper.readTree(res.body());
// fleetd #199: "members" is the canonical key. The endpoint is /members, so a caller that
// reads "members" must not see an empty fleet. "workers" is kept only as a deprecated alias
// and must carry the same rows — assert both, or the alias can silently drift.
JsonNode members = body.get("members");
assertNotNull(members, "GET /members must return its rows under \"members\"");
assertEquals(1, members.size());
JsonNode workers = body.get("workers");
assertNotNull(workers, "the deprecated \"workers\" alias is still emitted");
assertEquals(members, workers, "the alias must carry the same rows as \"members\"");
JsonNode w = members.get(0);
assertEquals(spawned.get("terminalId").asText(), w.get("sessionId").asText());
assertEquals(paneId, w.get("paneId").asText());
assertEquals("ltms-local", w.get("profile").asText());
@@ -30,12 +30,19 @@ public final class FakeWorktrees implements Worktrees {
public record PruneCall(String repoRoot, long minAgeMillis) {
}
public record ShareCall(String repoRoot, String worktreePath) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
private final List<ShareCall> shareCalls = new CopyOnWriteArrayList<>();
/** Tags every {@code overlayParity}/{@code shareWithGroup} call in call order, so a test can
* pin that sharing runs after the overlay copy (fleetd #185 stage 3). */
private final List<String> overlayShareOrder = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private final AtomicLong snapshotSeq = new AtomicLong();
@@ -136,6 +143,13 @@ public final class FakeWorktrees implements Worktrees {
}
overlayCalls.add(new OverlayCall(repoRoot, worktreePath, List.copyOf(overlay),
List.copyOf(copied), List.copyOf(skipped)));
overlayShareOrder.add("overlay:" + worktreePath);
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
shareCalls.add(new ShareCall(repoRoot, worktreePath));
overlayShareOrder.add("share:" + worktreePath);
}
@Override
@@ -203,4 +217,17 @@ public final class FakeWorktrees implements Worktrees {
public SnapshotCall lastSnapshot() {
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
}
public List<ShareCall> shareCalls() {
return List.copyOf(shareCalls);
}
public ShareCall lastShare() {
return shareCalls.isEmpty() ? null : shareCalls.getLast();
}
/** Call-order tags ({@code "overlay:<path>"}/{@code "share:<path>"}) — see field javadoc. */
public List<String> overlayShareOrder() {
return List.copyOf(overlayShareOrder);
}
}
@@ -924,4 +924,257 @@ class GitWorktreesTest {
assertEquals(2, stats.count(), "two snapshot refs are reported");
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
}
/**
* fleetd #185 stage 3: a recording {@link java.util.function.Function} test seam stands in for
* every process {@link GitWorktrees#shareWithGroup} would run — no real second OS user/group
* exists on this host, so these are unit tests against that seam, not a live-group integration
* test (out of scope per the ticket).
*/
private static List<String> joined(String[] command) {
return List.of(command);
}
/** Remove {@code path} and anything under it. Tolerates an already-absent path. */
private static void deleteRecursively(Path path) throws Exception {
if (!Files.exists(path)) {
return;
}
if (Files.isDirectory(path)) {
try (java.util.stream.Stream<Path> children = Files.list(path)) {
for (Path child : children.toList()) {
deleteRecursively(child);
}
}
}
Files.delete(path);
}
/** {@code worktreeGroup} absent ⇒ zero processes spawned and no git config written. */
@Test
void shareWithGroupIsNoopWhenNoGroupConfigured(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString(), null, _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repo.toString(), repo.resolve("some-worktree").toString());
assertTrue(recorded.isEmpty(), "no group configured must spawn no process at all: " + recorded);
}
/** A configured group runs {@code git config core.sharedRepository group} first, then
* chgrp/chmod/setgid over every path {@link GitWorktrees#shareWithGroup} documents. */
@Test
void shareWithGroupRunsConfigThenChgrpChmodSetgidPerPath(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String repoRoot = repo.toString();
Path worktree = Files.createDirectories(repo.resolve("some-worktree"));
String worktreePath = worktree.toString();
Files.createDirectories(repo.resolve(".git/worktrees"));
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
// What real git answers for an ordinary (non-linked) checkout: relative to repoRoot.
return List.of(cmd).contains("--git-common-dir") ? ".git\n" : "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repoRoot, worktreePath);
assertEquals(List.of("git", "-C", repoRoot, "config", "core.sharedRepository", "group"), recorded.get(0),
"core.sharedRepository must be set first, so it keeps working after the one-time fix-up");
assertTrue(recorded.contains(List.of("git", "-C", repoRoot, "rev-parse", "--git-common-dir")),
"the git dir must be asked for, never hardcoded as <repoRoot>/.git — that is a FILE "
+ "when the checkout is itself a linked worktree: " + recorded);
for (String dir : List.of(worktreePath, repoRoot + "/.git/objects", repoRoot + "/.git/refs",
repoRoot + "/.git/logs", repoRoot + "/.git/worktrees")) {
assertTrue(recorded.contains(List.of("chgrp", "-R", "devteam", dir)), "missing chgrp -R for " + dir);
assertTrue(recorded.contains(List.of("chmod", "-R", "g+rwX", dir)), "missing chmod -R for " + dir);
assertTrue(recorded.contains(List.of("find", dir, "-type", "d", "-exec", "chmod", "g+s", "{}", "+")),
"missing setgid find pass for " + dir);
}
// packed-refs does not exist in a freshly-init'd repo (only git gc / pack-refs creates it) —
// tolerated absence, so it must not appear at all: no recursive/-R treatment for a plain file.
String packedRefs = repoRoot + "/.git/packed-refs";
assertTrue(recorded.stream().noneMatch(c -> c.contains(packedRefs)),
"packed-refs is absent here and must be skipped, not chgrp'd: " + recorded);
}
/**
* A path that does not exist is skipped, never handed to {@code chgrp}. {@code .git/logs} is
* absent whenever {@code core.logAllRefUpdates} is false or no ref has been updated yet, and
* {@code chgrp} on a missing path exits non-zero — which would fail EVERY provisioning spawn
* with a message blaming a group that is in fact fine.
*/
@Test
void shareWithGroupSkipsPathsThatDoNotExist(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String repoRoot = repo.toString();
deleteRecursively(repo.resolve(".git/logs"));
assertFalse(Files.exists(repo.resolve(".git/logs")), "fixture: .git/logs must be gone");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return List.of(cmd).contains("--git-common-dir") ? ".git\n" : "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repoRoot, repo.resolve("no-such-worktree").toString());
String logs = repoRoot + "/.git/logs";
assertTrue(recorded.stream().noneMatch(c -> c.contains(logs)),
"a missing .git/logs must be skipped, not chgrp'd: " + recorded);
assertTrue(recorded.stream().noneMatch(c -> c.contains(repo.resolve("no-such-worktree").toString())),
"a missing worktree path must be skipped too: " + recorded);
assertTrue(recorded.contains(List.of("chgrp", "-R", "devteam", repoRoot + "/.git/objects")),
"paths that DO exist are still shared: " + recorded);
}
/**
* The git store is located by {@code rev-parse --git-common-dir}, not by appending
* {@code /.git}. When git answers with an absolute path — what it does for a linked worktree,
* where {@code <repoRoot>/.git} is a file — every shared path must follow that answer.
*/
@Test
void shareWithGroupFollowsAnAbsoluteGitCommonDir(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path realGitDir = repo.resolve(".git");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return List.of(cmd).contains("--git-common-dir") ? realGitDir + "\n" : "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(tmp.resolve("some/linked/worktree").toString(),
repo.resolve("wt").toString());
assertTrue(recorded.contains(List.of("chgrp", "-R", "devteam", realGitDir + "/objects")),
"objects must be taken from the reported common dir, not <repoRoot>/.git: " + recorded);
}
/** {@code packed-refs}, when present, is chgrp/chmod'd but never setgid'd (it is a file, not a dir). */
@Test
void shareWithGroupIncludesPackedRefsWhenPresent(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String repoRoot = repo.toString();
Path packedRefsPath = repo.resolve(".git/packed-refs");
Files.writeString(packedRefsPath, "");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees =
new GitWorktrees(tmp.resolve("wts").toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.shareWithGroup(repoRoot, repo.resolve("some-worktree").toString());
String packedRefs = packedRefsPath.toString();
assertTrue(recorded.contains(List.of("chgrp", "devteam", packedRefs)),
"packed-refs must be chgrp'd non-recursively when present: " + recorded);
assertTrue(recorded.contains(List.of("chmod", "g+rwX", packedRefs)),
"packed-refs must be chmod'd non-recursively when present: " + recorded);
assertTrue(recorded.stream().noneMatch(c -> c.contains("find") && c.contains(packedRefs)),
"packed-refs (a file) must never get the recursive setgid pass: " + recorded);
}
/** A group that does not exist (or that the operator is not a member of) fails loudly, naming it. */
@Test
void shareWithGroupThrowsNamingTheGroupWhenChgrpFails(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString(), "cb185-nonexistent-group-zz");
String wt = new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-185-share", "HEAD");
WorktreeException e = assertThrows(WorktreeException.class,
() -> gitWorktrees.shareWithGroup(repo.toString(), wt));
assertTrue(e.getMessage().contains("cb185-nonexistent-group-zz"),
"exception must name the missing/refused group: " + e.getMessage());
}
// --- fleetd #224: worktreeRoot itself must be group-traversable, established in add() ---------
/**
* {@code add} must make the worktree ROOT itself group-traversable when a group is configured —
* established once here, where the root is created, never at a use site (never inside a
* launcher). A recording runner stands in for chgrp/chmod, the same seam
* {@link #shareWithGroupRunsConfigThenChgrpChmodSetgidPerPath} uses for the per-worktree share.
*/
@Test
void addSharesWorktreeRootWithGroupWhenConfigured(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path root = tmp.resolve("wts");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees = new GitWorktrees(root.toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.add(repo.toString(), "cb-224-branch", "HEAD");
assertTrue(recorded.contains(List.of("chgrp", "devteam", root.toString())),
"the worktree root itself must be chgrp'd to the configured group: " + recorded);
assertTrue(recorded.contains(List.of("chmod", "g+x", root.toString())),
"the worktree root itself must gain group-execute so a different-uid member can "
+ "traverse into it: " + recorded);
}
/** {@code worktreeGroup} unset (today's only live mode) ⇒ {@code add} spawns no share process
* for the root at all — behaviour must be byte-identical to before fleetd #224. */
@Test
void addSharesNothingForTheRootWhenNoGroupConfigured(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path root = tmp.resolve("wts");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees = new GitWorktrees(root.toString(), null, _ -> {}, recordingRunner);
gitWorktrees.add(repo.toString(), "cb-224-nogroup", "HEAD");
assertTrue(recorded.isEmpty(), "no group configured must spawn no share process for the "
+ "root at all: " + recorded);
}
/**
* fleetd #224, acceptance criterion 2/3: when the root cannot be made group-traversable — here
* because the configured group does not exist, the same real-failure shape
* {@link #shareWithGroupThrowsNamingTheGroupWhenChgrpFails} drives for the per-worktree share —
* the spawn is refused with a message naming the root, its current mode, and the group. This
* drives the REAL {@code chgrp} (no recording runner), and the refusal happens before {@code git
* worktree add} ever runs, so no partial worktree is left behind either.
*/
@Test
void addRefusesWhenWorktreeRootCannotBeMadeGroupTraversable(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path root = tmp.resolve("wts");
GitWorktrees gitWorktrees = new GitWorktrees(root.toString(), "cb224-nonexistent-group-zz");
WorktreeException e = assertThrows(WorktreeException.class,
() -> gitWorktrees.add(repo.toString(), "cb-224-refuse", "HEAD"));
assertTrue(e.getMessage().contains(root.toString()),
"refusal must name the worktree root: " + e.getMessage());
assertTrue(e.getMessage().contains("cb224-nonexistent-group-zz"),
"refusal must name the missing/refused group: " + e.getMessage());
assertTrue(Files.isDirectory(root), "the root is created before the group check runs");
String mode = java.nio.file.attribute.PosixFilePermissions.toString(Files.getPosixFilePermissions(root));
assertTrue(e.getMessage().contains(mode),
"refusal must name the root's current mode (" + mode + "): " + e.getMessage());
try (java.util.stream.Stream<Path> children = Files.list(root)) {
assertTrue(children.findAny().isEmpty(), "no worktree must be left behind under the root: "
+ "the refusal must happen before `git worktree add` ever runs");
}
}
}
@@ -4,6 +4,8 @@ import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.MemberLifecycle;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -148,6 +150,10 @@ class SessionManagerTest {
return 0;
}
@Override
public void shareWithGroup(String repoRoot, String worktreePath) {
}
List<String> removeCalls() {
return List.copyOf(removeCalls);
}
@@ -329,6 +335,110 @@ class SessionManagerTest {
}
}
/**
* fleetd #226: a contended slot is refused through the real {@link SessionManager#acquire}
* path before the real launcher can hand an architect charter to a process.
*/
@Test
void aContendedArchitectSlotRefusesBeforeTheLauncherStartsAMember() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberRegistry members = architectRegistry();
assertTrue(members.bind("architect:opus", "term_already_bound"),
"precondition: occupy the sole architect slot before the real spawn under test");
sessions.setMemberLifecycle(members);
assertThrows(IllegalArgumentException.class, () -> sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller", "term_primary", null));
assertFalse(herdr.called("agent.start"), "a refused slot must never reach the launcher");
assertTrue(sessions.roster().isEmpty(), "no member exists to receive the wrong charter");
}
@Test
void aSecondArchitectOnAnAlreadyBoundProfileIsHeldAsDevNotArchitectAndWarnsLoudly() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberRegistry members = architectRegistry();
sessions.setMemberLifecycle(bindFailureAfterReservation(members));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger registryLog = (ch.qos.logback.classic.Logger)
LoggerFactory.getLogger(MemberRegistry.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
registryLog.addAppender(appender);
registryLog.setLevel(Level.WARN);
try {
MemberSession session = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.DEV, session.role(), "a failed reservation bind must use the fallback");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no slot-exhaustion WARN logged");
assertTrue(warn.contains("ltms-local"), "the WARN names the profile: " + warn);
assertTrue(warn.contains(session.terminalId()), "the WARN names the terminal: " + warn);
} finally {
registryLog.detachAppender(appender);
}
}
@Test
void reservedArchitectSlotBindsTheLaunchedMember() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
sessions.setMemberLifecycle(architectRegistry());
MemberSession session = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.ARCHITECT, session.role());
}
private static MemberRegistry architectRegistry() {
return new MemberRegistry(new FleetConfig.Fleet(Map.of(), Map.of("opus", new FleetConfig.Slot("ltms-local")),
Map.of(), Map.of(), null));
}
private static MemberLifecycle bindFailureAfterReservation(MemberRegistry members) {
return new MemberLifecycle() {
@Override
public MemberRole acquired(MemberRole role, String profile, String terminal) {
return members.acquired(role, profile, terminal);
}
@Override
public void released(String terminal) {
members.released(terminal);
}
@Override
public void requireSlotFor(MemberRole role, String profile) {
members.requireSlotFor(role, profile);
}
@Override
public SlotReservation reserve(MemberRole role, String profile) {
return members.reserve(role, profile);
}
@Override
public boolean bind(SlotReservation reservation, String terminal) {
members.release(reservation);
assertTrue(members.bind(reservation.slot(), "term_racer"), "the racer takes the released slot");
return false;
}
@Override
public void release(SlotReservation reservation) {
members.release(reservation);
}
};
}
@Test
void rosterReflectsAcquiredMinusReleased() {
FakeHerdr herdr = new FakeHerdr();
@@ -595,6 +705,32 @@ class SessionManagerTest {
"no session is registered when spawn times out (roster empty)");
}
@Test
void failedArchitectLaunchReleasesItsReservationForTheNextLaunch() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
long[] clock = {0};
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
1, () -> clock[0], () -> clock[0] += 10);
SessionManager sessions = new SessionManager(workers, new GitWorktrees(), () -> 0L, 0);
sessions.setMemberLifecycle(architectRegistry());
assertThrows(PeerUnreachableException.class, () -> sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller", "term_primary", null));
herdr.agentStatus("idle");
MemberSession retry = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.ARCHITECT, retry.role(),
"the failed launch returned its reservation instead of silently shrinking the slot pool");
}
// --- CB-516: release must notify, so a blocked send can be failed --------------------------
@Test
@@ -883,15 +1019,18 @@ class SessionManagerTest {
}
@Test
void acquireWithNeitherSessionFieldLeavesAgentSessionIdNull() {
void acquireWithNeitherSessionFieldStillMintsAnAgentSessionId() {
// fleetd #214: the claude-code launcher mints a session id for EVERY spawn, so a member is
// resumable even when the spawn asked for no session identity.
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, null, null);
assertNull(s.agentSessionId(), "no identity requested — unchanged from before CB-584");
assertFalse(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"a null id is omitted from the roster, like charterSha256 for a receipt-less session");
assertNotNull(s.agentSessionId(),
"fleetd #214: a plain spawn mints a session id, so every member is resumable");
assertTrue(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"the minted id is in the roster, so fleet_list advertises every member's resume handle");
}
@Test
@@ -927,4 +1066,235 @@ class SessionManagerTest {
assertTrue(e.getMessage().contains("SESSION_RESUME"), e.getMessage());
assertTrue(e.getMessage().contains("stub-profile"), e.getMessage());
}
// ── fleetd #209: agentSessionId resolved lazily against the retained handle ─────────────────
//
// The opencode adapter cannot answer PeerHandle.agentSessionId() at spawn time — the on-disk
// session row is written only after the pane is live — so the id must be re-polled on a LATER
// call, against the SAME handle instance the launcher returned at spawn. SessionManager used
// to let that handle go out of scope at the end of the spawn method, so no caller ever re-asked
// it and fleet_list/fleet_status never saw the id. LazyIdHandle below reproduces exactly that
// shape: null on the first N calls (the spawn-time call included), a real id after.
/**
* A {@link PeerHandle} whose {@link #agentSessionId()} answers {@code null} for its first
* {@code nullCalls} invocations, then either a fixed id or a configured throw on every call
* after that — the shape of the opencode bug (fleetd #209): the session row is not written
* until after the pane is live, so early polls come back empty and a later one finds it.
*/
private static final class LazyIdHandle implements PeerHandle {
private final String id;
private final String terminalId;
private final int nullCalls;
private final String resolvedId;
private final java.util.concurrent.atomic.AtomicInteger calls =
new java.util.concurrent.atomic.AtomicInteger();
private volatile RuntimeException throwAfter;
LazyIdHandle(String id, String terminalId, int nullCalls, String resolvedId) {
this.id = id;
this.terminalId = terminalId;
this.nullCalls = nullCalls;
this.resolvedId = resolvedId;
}
/** After the null calls are exhausted, throw instead of answering the resolved id. */
LazyIdHandle throwing(RuntimeException e) {
this.throwAfter = e;
return this;
}
@Override
public String id() {
return id;
}
@Override
public String terminalId() {
return terminalId;
}
@Override
public String agentSessionId() {
int n = calls.incrementAndGet();
if (n <= nullCalls) {
return null;
}
if (throwAfter != null) {
throw throwAfter;
}
return resolvedId;
}
@Override
public CharterReceipt charterReceipt() {
return null;
}
int callCount() {
return calls.get();
}
}
/** A minimal {@link PeerLauncher} that hands out pre-built {@link LazyIdHandle}s, one per spawn. */
private static final class LazyIdLauncher implements PeerLauncher {
private final java.util.Deque<LazyIdHandle> queued = new java.util.ArrayDeque<>();
LazyIdLauncher queue(LazyIdHandle handle) {
queued.add(handle);
return this;
}
@Override
public Set<Capability> capabilities() {
return Set.of(Capability.WORKTREE, Capability.SESSION_RESUME);
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return capabilities();
}
@Override
public PeerHandle spawn(SpawnRequest req) {
LazyIdHandle handle = queued.poll();
if (handle == null) {
throw new IllegalStateException("no queued LazyIdHandle for this spawn");
}
return handle;
}
@Override
public Set<String> profiles() {
return Set.of("lazy");
}
@Override
public String defaultProfile() {
return "lazy";
}
@Override
public String effectiveCwd(SpawnRequest req) {
return "/cwd";
}
@Override
public List<String> parityOverlay(String profileName) {
return List.of();
}
@Override
public List<?> list() {
return List.of();
}
@Override
public int reapOrphanWorkers() {
return 0;
}
@Override
public void stop(String id) {
}
@Override
public boolean clearContext(String id) {
return false;
}
}
@Test
void plainRosterDoesNotResolveAgentSessionId() {
// fleetd #209 follow-up: roster() sits on the heartbeat/health-tick timers (and the metrics
// scrape), so it must never trigger the resolve lookup — for opencode that lookup opens an
// on-disk session database, and a member whose id never appears would pay that cost forever.
// rosterResolved() is the one to use when a caller actually reports the id.
LazyIdHandle handle = new LazyIdHandle("p0", "t0", 1, "oc-session-0");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
assertEquals(1, handle.callCount(), "sanity: only the spawn-time call happened so far");
List<MemberSession> roster = sessions.roster();
assertEquals(1, roster.size());
assertEquals(acquired.paneId(), roster.getFirst().paneId());
assertNull(roster.getFirst().agentSessionId(), "the plain roster must not resolve the id");
assertEquals(1, handle.callCount(),
"roster() must never call agentSessionId() again — it sits on the heartbeat/health timers");
}
@Test
void rosterResolvedResolvesALateAgentSessionIdFromTheRetainedHandle() {
LazyIdHandle handle = new LazyIdHandle("p1", "t1", 1, "oc-session-1");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
assertNull(acquired.agentSessionId(),
"opencode has not written its session row yet at spawn time");
List<MemberSession> roster = sessions.rosterResolved();
assertEquals(1, roster.size());
assertEquals("oc-session-1", roster.getFirst().agentSessionId(),
"fleet_list must see the id once the adapter can answer it");
Map<String, Object> view = SessionManager.rosterView(roster.getFirst(), null);
assertEquals("oc-session-1", view.get("agentSessionId"),
"rosterView renders whatever rosterResolved() resolved");
}
@Test
void getResolvesALateAgentSessionIdFromTheRetainedHandle() {
LazyIdHandle handle = new LazyIdHandle("p2", "t2", 1, "oc-session-2");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
MemberSession resolved = sessions.get(acquired.paneId()).orElseThrow();
assertEquals("oc-session-2", resolved.agentSessionId(),
"fleet_status (single-session lookup) must also see the late-resolved id");
}
@Test
void releaseCarriesALateResolvedAgentSessionIdIntoTheReleaseDetail() {
LazyIdHandle handle = new LazyIdHandle("p3", "t3", 1, "oc-session-3");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
MemberSession acquired = sessions.acquire("lazy", "/cwd", "/caller", null);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(detail -> released.add(detail.agentSessionId()));
sessions.release(acquired.paneId());
assertEquals(java.util.List.of("oc-session-3"), released,
"a released member's detail carries the id it has since resolved, not the null "
+ "frozen in at spawn time");
}
@Test
void aThrowingHandleDoesNotBreakRosterResolved() {
LazyIdHandle handle = new LazyIdHandle("p4", "t4", 1, "oc-session-4")
.throwing(new RuntimeException("sqlite locked"));
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
sessions.acquire("lazy", "/cwd", "/caller", null);
List<MemberSession> roster = assertDoesNotThrow(sessions::rosterResolved,
"a handle that throws resolving its id must not break the roster read");
assertEquals(1, roster.size());
assertNull(roster.getFirst().agentSessionId(), "the id stays unresolved when the lookup throws");
}
@Test
void aResolvedAgentSessionIdIsNotLookedUpAgain() {
LazyIdHandle handle = new LazyIdHandle("p5", "t5", 1, "oc-session-5");
SessionManager sessions = new SessionManager(new LazyIdLauncher().queue(handle));
sessions.acquire("lazy", "/cwd", "/caller", null);
assertEquals(1, handle.callCount(), "sanity: only the spawn-time call happened so far");
sessions.rosterResolved();
assertEquals(2, handle.callCount(), "the first rosterResolved() read resolves the id");
sessions.rosterResolved();
assertEquals(2, handle.callCount(), "once resolved, the id must not be looked up again");
}
}
@@ -95,10 +95,9 @@ class WorktreeSessionManagerTest {
null, "/caller/proj", null, null);
assertEquals("architect:architect", members.slotForTerminal(replacement.terminalId()));
MemberSession overflow = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
assertNull(members.slotForTerminal(overflow.terminalId()),
"a full slot pool must not stop the architect spawn");
assertThrows(IllegalArgumentException.class, () -> sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null),
"a full slot pool must stop the spawn before it can receive an architect charter");
}
@Test
@@ -163,6 +162,37 @@ class WorktreeSessionManagerTest {
"tracked copied paths are --skip-worktree'd");
}
/**
* fleetd #185 stage 3, THE TRAP: {@code overlayParity} copies more files into the worktree
* AFTER {@code add} returns, so {@code shareWithGroup} must run after it, not folded into
* {@code add()} — otherwise every overlay file lands operator-owned and unwritable for a
* different-uid member, with a green test suite hiding it.
*/
@Test
void shareWithGroupRunsAfterOverlayParityNotBeforeIt() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".envrc");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-185", null));
assertEquals(1, worktrees.overlayCalls().size(), "overlayParity ran exactly once");
assertEquals(1, worktrees.shareCalls().size(), "shareWithGroup ran exactly once");
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
FakeWorktrees.ShareCall share = worktrees.lastShare();
assertEquals(s.worktree(), overlay.worktreePath());
assertEquals(s.worktree(), share.worktreePath());
List<String> order = worktrees.overlayShareOrder();
int overlayIndex = order.indexOf("overlay:" + s.worktree());
int shareIndex = order.indexOf("share:" + s.worktree());
assertTrue(overlayIndex >= 0 && shareIndex >= 0, "both calls must be recorded: " + order);
assertTrue(overlayIndex < shareIndex,
"shareWithGroup MUST run after overlayParity, not before/inside add(): " + order);
}
@Test
void releaseRemovesWorktreeButDoesNotDeleteBranch() {
FakeHerdr herdr = new FakeHerdr();