Compare commits

...

21 Commits

Author SHA1 Message Date
Dai Ha 31b028e860 fleetd #234 round 4: invert ExhaustionSink's abstract method so the bug class is unrepresentable
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m23s
Round 3's factory fixed the two known call sites but the underlying shape
was still there: a lambda written against ExhaustionSink binds to whichever
overload is abstract, and the 2-arg form held that position, so ANY lambda
-- a call-site forwarder, a hand-built test double, a future caller who has
never heard of fleetd #234 -- could still silently take the hint-dropping
default. Two rounds shipped exactly that mistake in two different places.

Fix: made the 3-arg onExhausted(target, reason, profile) the interface's
single abstract method; the 2-arg form is now a default that delegates with
a null profile. A lambda declared against ExhaustionSink today is forced by
the compiler to take three parameters -- there is no overload left for it to
bind to that can drop the hint. This is enforced by the type system, not by
a test that has to remember to check for it.

Knock-on changes:
- ExhaustionSink.none() -- a 3-arg lambda, still a genuine no-op, now safe
  by construction rather than by care.
- ExhaustionSink.forwardingTo(...) -- collapses to a one-line 3-arg lambda;
  kept as a named factory (round 3's lesson: a test must call the real
  object, not rebuild its shape).
- Fleetd.java's real sink and the two OpenCodeLauncherTest sinks that used
  to be anonymous classes overriding both overloads are now plain lambdas
  too -- the 2-arg override each carried was pure boilerplate once the
  interface provides it as a default.
- CompletionResolver.java itself: UNCHANGED, zero diff (confirmed via
  `git diff --stat` before staging) -- its two call sites still call the
  2-arg onExhausted(target, reason), which is now the default and behaves
  identically. CompletionResolverTest (41 tests, 0 failures) proves this;
  its five ExhaustionSink lambdas needed a mechanical third parameter added
  to keep compiling against the new abstract method, no assertion changed.

Mutation proof, re-run against the new shape: forwardingTo's body edited to
call the 2-arg default instead of passing the hint through (the equivalent
of round 3's "delete the 3-arg override" now that there is only one method
to break) -- both new tests go red with the same assertions as round 3:

  ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
  OpenCodeLauncherTest...ForwardingHop:  expected: <true> but was: <false>
  Tests run: 68, Failures: 2

Restored, re-ran: green (Tests run: 109, Failures: 0, including
CompletionResolverTest).

Compiler proof (not committed -- a scratch file outside the worktree,
compiled with the real ExhaustionSink.java on the classpath, then deleted):

  ExhaustionSink forwarder = (target, reason) -> System.out.println(target + reason);

  error: incompatible types: incompatible parameter types in lambda expression

A 2-arg lambda against this interface no longer compiles at all.

mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:46:27 +07:00
Dai Ha c935b181dd fleetd #234 round 3: make ExhaustionSink's forwarder a shared factory, not a rebuilt-per-caller shape
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m36s
Round 2's tests never reached Fleetd.java at all: both new tests declared
their OWN local copy of the forwarding shape instead of calling production's.
Mutating Fleetd.java's real forwarder back into the broken lambda left those
copies untouched, so the whole suite stayed green while production had
regressed to exactly the bug being fixed -- proven live by the reviewer.

Fix: extracted the forwarding shape into one named factory,
ExhaustionSink.forwardingTo(Supplier<ExhaustionSink> target), with the
"why a lambda here is wrong" explanation moved onto it (the one place the
shape is now written). Fleetd.java's forwarder collapses to one line:

    ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);

Both new tests now call this same factory instead of rebuilding an anonymous
class inline, so they exercise the identical object production builds:
- ExhaustionSinkForwardingHazardTest: calls ExhaustionSink.forwardingTo
  directly and asserts the hint reaches the real sink through it.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
  same factory call, inside the full Fleetd-shaped construction order
  (forwarder built first, real sink pointed at via the AtomicReference
  afterward), driven through the real SessionManager.acquire() path.

Mutation proof, this time on production code only: deleted the factory's
3-arg override (falls back to the interface default, dropping the hint) --
both new tests go red with no test file touched:

  ExhaustionSinkForwardingHazardTest...: expected: <gx> but was: <null>
  OpenCodeLauncherTest...ForwardingHop:  expected: <true> but was: <false>
  Tests run: 68, Failures: 2

Restored, re-ran: green (Tests run: 68, Failures: 0). Confirmed Fleetd.java
carries no lambda ExhaustionSink anywhere (grep). ExhaustionSink.none() stays
a lambda on purpose -- both its overloads are true no-ops regardless of
arity, so there is no hint to drop.

mvn clean install: Tests run: 1129, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:30:37 +07:00
Dai Ha c325054242 fleetd #234 round 2: fix the ExhaustionSink forwarding hop Fleetd.java actually uses
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m18s
The round-1 fix was dead on the real production path. Fleetd.java:177 builds
a forwarding sink (needed because the adapters are constructed before
`sessions` exists, breaking a genuine cycle) as a LAMBDA:

    ExhaustionSink forwardingExhaustionSink =
            (target, reason) -> exhaustionSinkRef.get().onExhausted(target, reason);

A lambda can only implement the interface's one abstract method (the 2-arg
overload), so it silently inherited the 3-arg overload's default body, which
drops the profile hint and calls back into the 2-arg method. OpenCodeLauncher
is constructed with this forwarder, so the hint it supplies (its own
already-known profile name) was thrown away before it ever reached the real
sink built later in Fleetd.main -- reproducing the exact silent no-op round 1
was sent to fix. The 1127 tests from round 1 all injected a sink directly
into OpenCodeLauncher and never went through this forwarding hop, so none of
them could see it.

Fix: forwardingExhaustionSink is now an anonymous class overriding both
overloads, each delegating to whatever exhaustionSinkRef currently holds.

Audited every other ExhaustionSink value in main/: the only other one is
ExhaustionSink.none() (a lambda), which is safe regardless of arity since
both its 2-arg body and the inherited 3-arg default are true no-ops.

New tests:
- ExhaustionSinkForwardingHazardTest: isolates the hazard at the interface
  level (a lambda forwarder drops the hint; an anonymous-class forwarder
  does not), independent of Fleetd.java's specific wiring.
- OpenCodeLauncherTest#theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop:
  replicates Fleetd.java's actual construction order (forwarder built and
  handed to the launcher first, real sink built and pointed at via the
  AtomicReference afterward) and drives the quarantine through it via the
  real SessionManager.acquire() path.

Both proven by mutation: temporarily rewriting each fixed forwarder back
into the pre-fix lambda makes its test fail with a real assertion message
(both matched exactly: "expected: <gx> but was: <null>" for the interface
proof, "expected: <true> but was: <false>" for the composed-wiring test);
restoring makes it pass again. No reverts were committed.

mvn clean install: Tests run: 1130, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:22:31 +07:00
Dai Ha 4877992a70 fleetd #234: key the opencode model check on the resolved session id, and make the spawn-time quarantine actually happen
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m29s
Defect 1: OpenCodeSessionDiscovery.actualModelForDirectory queried
WHERE directory = ?, the same heuristic sessionIdForDirectory uses. Since a
default fleet_spawn (no worktree:) shares the lead's cwd with every other
worker and every past session ever run there, the model read-back could
silently compare against a DIFFERENT session's row. Renamed to
actualModelForSessionId(sessionId), keyed on the primary key id instead, and
made OpenCodeLauncher's SessionAwareHandle cache the resolved id once
non-null (AtomicReference) so a later sibling row in the same directory can
never flip which session's evidence is read. sessionIdForDirectory (#209) is
left directory-based on purpose, with a comment explaining why the heuristic
is unavoidable at that layer.

Defect 2: the ERROR log claimed "quarantining this profile's credential" but
Fleetd's ExhaustionSink lambda resolved target -> roster -> profile ->
credential, while OpenCodeLauncher's model-mismatch check fires from
agentSessionId() during SessionManager.acquire(), before the session is
registered in the roster -- the lookup found nothing and silently no-opped.
Added a default 3-arg ExhaustionSink.onExhausted(target, reason, profile)
overload (defaults to the 2-arg method, so CompletionResolver's two call
sites are unchanged); OpenCodeLauncher now passes its own already-known
profile name; Fleetd's sink became an anonymous class that tries the roster
first, falls back to the hint, and logs loudly at ERROR naming target/reason
when neither resolves, instead of silently no-oping.

Both fixes proven by mutation: reverting each independently makes its new
test fail with a real assertion message, restoring makes it pass again.

mvn clean install: Tests run: 1127, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
2026-09-03 10:09:26 +07:00
ltms 0f08b93659 fleetd #176: spawn gate fails fast on a dead backend, and refines UNKNOWN safely
CI / contract (push) Successful in 52s
CI / build (push) Successful in 1m47s
Two fixes to HerdrPeerLauncher.waitUntilInjectableOrThrow.

Fix 1: the status call is now guarded. A herdr *_not_found answer — what happens
when the backend process exited rather than being slow — used to escape as a raw
HerdrException, skipping stop() and leaking the pane and tab, with a message that
blamed a slow pane. It now fails immediately, runs the same teardown, and says the
process exited.

Fix 2: the gate can now resolve UNKNOWN with the same StatusRefiner the status
poller uses, behind two guards. It only runs for the claude adapter, because
StatusRefiner.classify reads a Claude Code TUI. And a refined result is accepted
only when the same agents.get sample still reports a non-null agentType. That
second guard closes a trap: a dead pane sits at a shell prompt containing the same
❯ glyph the classifier reads as idle, so refining without corroboration would turn
"the backend died" into "ready to inject".

Verified by the lead before merge: NAME_PREFIX really is "claude"/"opencode" so the
adapter split holds; AgentControl.status(t) was already get(t).status(), so moving
the loop to get() adds no herdr call; and a *_not_found already failed the spawn
before this change, so fix 1 improves an existing failure rather than creating one.
The adapter guard was mutation-tested independently (set it to `if (false)`, the new
opencode test fails with "Expected PeerUnreachableException to be thrown, but nothing
was thrown"; restored, it passes). Independent build: 1123 tests, 0 failures,
0 errors, 0 skipped.

Not fixed, and not claimed: that herdr reports a null agentType for a bare shell
pane is unverified against a live daemon. To be proved by a live spawn on both
backends after deploy. The seat-accounting suggestion in #176 was deliberately not
built — its cause was tested in that issue and not reproduced.
2026-09-03 04:20:23 +02:00
Dai Ha 432c1d92d1 fleetd #176 review: cover the namePrefix guard, the adapter-refinement corroboration that had no test
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m14s
The agentType corroboration (guard b) had positive and negative tests, but the
per-adapter namePrefix guard (guard a) had none — nothing proved that an
opencode pane can never reach StatusRefiner.classify, only that a claude pane
with a null agentType is rejected. Since a live opencode pane always reports a
non-null agentType ("opencode"), guard (b) alone cannot catch a broken guard
(a).

Added opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt
to OpenCodeLauncherTest: agentType("opencode") (non-null, satisfies guard b on
its own) + pane content containing "❯" (would classify as IDLE) + raw status
UNKNOWN throughout. Asserts the gate still times out, and that no agent.read
call used source=detection (StatusRefiner.PROBE_SOURCE) — proving the refiner
was never even reached, not just that its answer was discarded. A blanket
"agent.read is never called" does not hold here: HerdrPeerLauncher.readPaneQuietly
reads the pane tail (source=recent) for the timeout log on every timeout,
regardless of adapter, so the assertion is scoped to the refiner's own probe
source instead.

Mutation check performed and reverted before this commit: temporarily changed
`if (!"claude".equals(namePrefix))` to `if (false)` in
HerdrPeerLauncher.refinedInjectable — the new test failed
("Expected PeerUnreachableException to be thrown, but nothing was thrown"),
confirming it actually exercises the guard. Restored the guard and reran —
test passes again (1/1). No production code changed in this commit.

mvn clean install: Tests run: 1123, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:17:10 +07:00
Dai Ha 60b7e67b42 fleetd #176: fail fast when the spawn gate's backend exits mid-wait, and safely refine its UNKNOWN status
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m26s
Fix 1: agents.get(paneId) inside waitUntilInjectableOrThrow was unguarded, so a
herdr *_not_found answer (the backend process exited) propagated as a raw
HerdrException instead of PeerUnreachableException, and skipped teardown
entirely, leaking the pane/tab. Now caught via isAlreadyGone(); fails
immediately (does not burn the rest of the timeout), runs the same teardown
the timeout path runs, and the exception message says the process exited
rather than that the pane was slow. Any other HerdrException still
propagates unchanged.

Fix 2: the gate now resolves a raw UNKNOWN into StatusRefiner's pane-content
classification, like StatusPoller already does mid-life. Guarded against the
trap noted in the ticket: a pane whose backend exited settles at a bare shell
prompt that can also contain the "❯" glyph classify() reads as idle. So a
refined result is accepted only when the corroborating agentType from the
SAME agent.get sample is non-null, and only for the "claude" adapter (namePrefix)
since StatusRefiner.classify is written for the Claude Code TUI only. Refine
only runs when the raw status is UNKNOWN, so a healthy spawn adds zero extra
herdr calls.

Tests added to ClaudeCodeLauncherTest (FakeHerdr gained agentType()/
agentGetFailsWithAfter() fixtures):
 - spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout
 - spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged
 - refinedIdleIsNotAcceptedWhenAgentTypeIsNull
 - refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness
 - refinementNeverFiresWhenRawStatusIsAlreadyInjectable

mvn clean install: Tests run: 1122, Failures: 0, Errors: 0, Skipped: 0 -- BUILD SUCCESS
2026-09-03 09:10:49 +07:00
ltms 3fbd43fe3f fleetd #175: read back the model opencode actually resolved, and quarantine a silent substitution
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m29s
Verified by the lead before merge. Read the full production diff, checkModelMatch, parseModel and actualModelForDirectory. Confirmed the check runs on the real #209 late-resolve path (not at spawn, which is why #203 was closed), that unknown/incomplete evidence never quarantines, and that claude-code is structurally excluded because SessionAwareHandle is only built by OpenCodeLauncher.spawn(). Round 2 closed the one gap I found: model JSON with an id but no providerID used to read as a mismatch for a provider-prefixed profile. Independent build: BUILD SUCCESS, 1117 tests, 0 failures, 0 skipped.
2026-09-02 13:01:33 +02:00
Dai Ha 32ebf065ac fleetd #175 review: a missing providerID in opencode's model JSON is UNKNOWN, not a mismatch
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 1m13s
parseModel already tolerates a model JSON with an id but no providerID (a real shape
opencode can write). checkModelMatch's providerMatches check did not: a provider-
prefixed profile whose id matched but whose evidence had no providerID was reported
as a mismatch and quarantined on incomplete data, which acceptance rule 4 forbids.

Compare the provider only when BOTH the profile requested one AND the evidence has
one. A genuine id mismatch is still caught either way — narrows the check, does not
disable it.
2026-09-02 17:59:04 +07:00
Dai Ha 1178b3f684 fleetd #175: check opencode's actual model against the profile, quarantine on a real mismatch
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m49s
opencode does not fail on an unknown -m <model> flag — it silently falls back to a
default model, which can be a paid credential. Extends the existing late-resolve
path (#209's SessionManager -> handle.agentSessionId() re-poll) so that once the
opencode session row exists, OpenCodeSessionDiscovery also reads its `model` JSON
column and OpenCodeLauncher's SessionAwareHandle compares it against the profile's
configured model.

Comparison rule: split the profile's model on the first '/' into provider+id. Compare
id always; compare provider only when the profile specified one. A bare model name
with no '/' matches on id alone. Absent/unparseable evidence is UNKNOWN, never a
mismatch, so a working profile is never quarantined on missing data. A real mismatch
logs an ERROR naming both models and the profile, then quarantines through the
existing ExhaustionSink path (wired via an AtomicReference forwarding sink in
Fleetd.java to break the sessions/workers/adapters construction cycle).
2026-09-02 17:53:59 +07:00
ltms 2a434ced2f fleetd #222: keep the member's charter file out of fleetd's own java.io.tmpdir
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m28s
Verified by the lead before merge: read the full production diff, confirmed the memberHerdrSocket-absent branch is the literal unmodified Files.createTempFile call in its own branch, and that the refusal names the missing key. Measured behaviour recorded: claude 2.1.258 exits 1 immediately on an unreadable --append-system-prompt-file, so the pre-fix bug was the loud readiness-gate failure, not a silent charter-less member. The /tmp full-suite failure the worker reported was the two other #225 copies, fixed by #230 which is already on main; main is built and checked after this merge.
2026-09-02 02:56:36 +02:00
ltms cc919aa2b6 fleetd #226: reserve an architect slot before launch, so no member runs on a charter it will not hold
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m31s
Verified by the lead before merge. Read the full production diff: locking is consistent (every MemberRegistry method uses synchronized(terminalToSlot)), and the refusal happens before the launcher starts a process. Proved the restored fallback test is real by removing the DEV fallback from SessionManager and re-running it: it failed with "expected: <DEV> but was: <ARCHITECT>", then reverted. Independent build: BUILD SUCCESS, 1092 tests, 0 failures, 0 skipped.
2026-09-02 02:54:52 +02:00
Dai Ha 6417b0edd9 fleetd#222: fix currentUserGroup() to resolve the real primary group (fleetd#225)
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m43s
The helper I copied from OpenCodeLauncherTest read the CWD's owning
group instead of the process's real primary group, so it silently
picked up whatever group owns the directory Maven was started from
(staff in a home checkout, wheel under /private/tmp on macOS) rather
than a group the operator is actually in. Replaced with the id -gn
based resolution that landed on #230 for the other two copies of this
helper, same shape and skip wording.
2026-09-02 07:54:11 +07:00
ltms 39c7ce76f3 fleetd #224 + #225: make worktreeRoot group-traversable, and stop two tests depending on the checkout location
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m41s
Verified by the lead before merge: read the full production diff, confirmed `group` is normalised to null at GitWorktrees:148 so the `group == null` guard is complete, and confirmed the refusal runs before `git worktree add` so a failure leaves no half-made worktree. Independent build in the worker's worktree: BUILD SUCCESS, 1093 tests, 0 failures, 0 skipped.
2026-09-02 02:52:12 +02:00
Dai Ha b40f477210 #226 retain architect fallback coverage
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m33s
2026-09-02 07:50:34 +07:00
Dai Ha 5c56cb347f fleetd #224 / #225: share worktreeRoot with the group, and fix the group-detection test helper
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m42s
#224: GitWorktrees#add created worktreeRoot with the daemon's umask and never shared it with
worktreeGroup, even though shareWithGroup shares every child underneath it (each worktree, and
the repo's common git dir). Under memberHerdrSocket: the member pane runs as a different OS
user, which needs execute on every ancestor directory to reach anything underneath, no matter
how carefully each child is shared — so a member could not read the opencode.json #219 places
under this root, could not reach its own worktree, and could not read #213's ZDOTDIR scrub when
placed here either.

Fix: add() now calls a new shareRootWithGroup(root) right after creating the root, chgrp+chmod
g+x on the root itself (non-recursive — each child is still shared individually by its own call
site). No-op when worktreeGroup is unset, so behaviour is byte-identical in today's only live
mode. On failure (group missing, or operator not a member of it) the spawn is refused with a
WorktreeException naming the root, its current mode, and the group — mirroring shareWithGroup's
existing refusal shape — before `git worktree add` ever runs, so no partial worktree is left
behind.

Also adds the assertion the #221 reviewer flagged as missing: a test driving
EnvAllowListScrub#shareWithGroup directly against a directory holding several flat files
(opencode.json, member-charter.md, ide-rules.md, plus an unrelated one) and asserting every one
of them gets group-readable/never-group-writable permissions, not just the two files someone
happened to think of.

#225: OpenCodeLauncherTest/HerdrPeerLauncherAllowListWiringTest's currentUserGroup() read the
group that owns the current working directory, not the process's own primary group, despite its
comment claiming the latter. Those coincide only by accident: a home checkout is typically owned
by a group the operator belongs to (staff), while a checkout under /private/tmp on macOS is
group wheel, which the operator is usually not a member of — so the same test fails for real
depending on where the repo happens to be checked out, and the existing assumeTrue only guarded
against "no POSIX groups at all", never "a resolvable but wrong group". Fixed by resolving the
process's REAL primary group via `id -gn` instead, with assumeTrue (skip, not fail) only when
that itself cannot be resolved on the host. The permission assertions these tests exist for are
unchanged.

Verified `mvn clean install` green from both a home checkout and a /private/tmp copy (mirroring
the exact repro in #225): 1093 tests, 0 failures, 0 errors in both locations.
2026-09-02 07:47:13 +07:00
Dai Ha e694deace3 #226 reserve architect slots before launch
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Failing after 1m33s
2026-09-02 07:44:41 +07:00
Dai Ha 748367b7d6 fleetd#222: put the claude-code charter file where the member can read it
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m16s
ClaudeCodeLauncher#writeCharterFile used Files.createTempFile with no
directory argument, which resolves against fleetd's own java.io.tmpdir
(macOS: the per-user $TMPDIR, mode 0700). Under memberHerdrSocket: the
member pane runs as a different OS user and cannot read that directory,
and since #220 the charter file is the ONLY delivery path for
--append-system-prompt-file. Following #219's refusal decision (a
charter is the member's turn contract, not a degradable control): with
memberHerdrSocket configured, the charter now goes into a fresh
per-spawn directory under worktreeRoot, shared read-only via
EnvAllowListScrub.shareWithGroup (reusing #213/#219's mechanism); a
missing worktreeRoot/worktreeGroup refuses the spawn by name instead of
writing an unreadable file. With memberHerdrSocket absent the path is
unchanged.
2026-09-02 07:44:34 +07:00
ltms dcf5fb3be3 Merge pull request 'CB-619 / fleetd #123: refuse an architect spawn with no matching slot' (#223) from worker/cb-123-role-demotion-c600f7-2 into main
CI / contract (push) Successful in 1m11s
CI / build (push) Successful in 1m55s
2026-09-01 10:41:47 +02:00
ltms f9d2ee2a2b Merge pull request 'fleetd#219: fix OpenCodeLauncher config/discovery roots under memberHerdrSocket' (#221) from worker/cb-219-opencode-roots-1f677e-1 into main
CI / contract (push) Successful in 45s
CI / build (push) Successful in 2m16s
2026-09-01 10:33:04 +02:00
Dai Ha 866c7f2e9a CB-619 / fleetd #123: refuse an architect spawn with no matching slot
CI / build (pull_request) Successful in 1m16s
CI / contract (pull_request) Successful in 1m18s
An explicit-profile spawn bypasses role-pool placement (CompositePeerLauncher
only constrains an UNQUALIFIED spawn to fleet.<role>), so it was the one path
that could ask for role=architect on a profile no architect slot carries.
MemberRegistry silently held the session as a plain worker while GET /members
still reported the requested "architect" and only fleet_whoami (which reads
live bindings, not the request) told the truth.

- MemberLifecycle.requireSlotFor(role, profile): refuses the acquire before
  anything spawns when no configured architect slot carries the profile,
  naming the role, the profile, and the pools that do carry it. No-op for
  dev/reviewer, which are placement candidates only, never a live identity
  binding — refusing a profile mismatch there would break the documented
  fleet_spawn{profile:"opus"} (role defaults to dev) flow.
- MemberLifecycle.acquired(...) now returns the role the session actually
  holds, so a residual race (a slot exists but every instance is already
  bound to a different terminal) still falls back to dev honestly instead of
  lying — this case logs at WARN (was INFO), naming profile and terminal.
- SessionManager now records the role acquired() returns on MemberSession,
  never the requested role, so GET /members and fleet_list can no longer
  report a role the member does not hold; no changes needed to memberView/
  rosterView, which just read session.role().

An architect's identity IS the slot it is bound to — binding a role with no
slot to bind means inventing an identity out of nothing, which is the quiet
failure the whole role system exists to prevent.

Tests: SessionManagerTest and FleetMcpTest each drive a real spawn through
FleetMcp.spawn -> SessionManager.acquire -> the real ClaudeCodeLauncher (via
FakeHerdr), then assert on GET /members and fleet_whoami for that same
session — not on MemberRegistry.bind directly (fleetd issue #113's mistake).
2026-09-01 15:29:39 +07:00
22 changed files with 2300 additions and 104 deletions
+56 -13
View File
@@ -169,6 +169,18 @@ public final class Fleetd {
}
});
List<HerdrPeerLauncher> adapters = new ArrayList<>();
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
// pointed at the real one once it exists.
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
// fleetd #234, round 4: routed through the shared ExhaustionSink.forwardingTo factory
// rather than written inline here — not because a lambda at this call site is unsafe
// anymore (it is not: the 3-arg overload is now the interface's single abstract method, so
// there is no 2-arg overload left for any lambda to silently bind to instead), but so a
// test can call the exact same object this line builds, instead of asserting a copy of its
// shape (round 3's lesson).
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
// unless opencode is the only kind configured.
@@ -184,7 +196,7 @@ public final class Fleetd {
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet(),
() -> config.get().memberCredentials(), config::get));
() -> config.get().memberCredentials(), config::get, forwardingExhaustionSink));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
@@ -334,18 +346,49 @@ public final class Fleetd {
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.map(profileName -> config.get().profiles().get(profileName))
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
//
// fleetd #175: this sink is now also the quarantine target for OpenCodeLauncher's
// model-mismatch check — it needs nothing profile-specific from the caller beyond `target`
// (a herdr terminal id) and `reason`, so reusing it here is exactly "the existing
// ExhaustionSink path", not a new mechanism.
//
// fleetd #234: that check fires from SessionAwareHandle.agentSessionId(), which runs during
// SessionManager.acquire() BEFORE this session is registered in sessions.roster() — so the
// roster-only lookup below used to find nothing, .ifPresent silently no-op'd, and the
// ERROR the check had just logged ("quarantining this profile's credential") was a lie:
// nothing was quarantined, and nothing said so. Two changes: (1) OpenCodeLauncher now
// passes its OWN profile name via ExhaustionSink's 3-arg overload — it already has the
// FleetConfig.Profile in hand and does not need the roster at all — used here as a
// fallback whenever the roster lookup misses; (2) if a profile still cannot be resolved
// (neither the roster nor the hint names a configured one), this logs loudly at ERROR
// instead of silently doing nothing — a control that cannot act must say so.
// fleetd #234, round 4: the 3-arg overload is now ExhaustionSink's single abstract method,
// so this is safely a lambda — there is no separate 2-arg overload left for it to bind to
// instead and silently drop profileHint (that was rounds 1-3's whole hazard).
ExhaustionSink exhaustionSink = (target, reason, profileHint) -> {
String profileName = sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.orElse(profileHint);
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
if (profile == null) {
log.error("quarantine requested for target '{}' ({}) but no profile could be "
+ "resolved — the target is not (yet) in the roster, and {} — "
+ "credential NOT quarantined (fleetd #234)",
target, reason,
profileHint == null ? "no profile hint was given"
: "the hinted profile '" + profileHint + "' is not configured");
return;
}
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
};
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
// now that `sessions` exists to resolve target -> session -> profile.
exhaustionSinkRef.set(exhaustionSink);
AgentControl agents = router.memberAgents();
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
@@ -5,17 +5,70 @@ import dev.ltms.fleet.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
/** A slot held before a member process starts. */
record SlotReservation(String slot, String profile) { }
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
public MemberRole acquired(MemberRole role, String profile, String terminal) {
return role; // no registry configured — nothing to bind against, so the request stands
}
@Override
public void released(String terminal) {
}
@Override
public void requireSlotFor(MemberRole role, String profile) {
// no registry configured — nothing to validate against, so nothing is refused
}
@Override
public SlotReservation reserve(MemberRole role, String profile) {
return null;
}
@Override
public boolean bind(SlotReservation reservation, String terminal) {
return false;
}
@Override
public void release(SlotReservation reservation) {
}
};
void acquired(MemberRole role, String profile, String terminal);
/**
* Try to bind a newly spawned {@code terminal} into the role it was granted.
*
* @return the role this session actually holds: {@code role} unchanged for a role with no
* live slot-binding semantics (dev, reviewer), or when the bind succeeded; a fallback
* role — never {@code role} — when a slot-bound role (architect) could not be bound.
* Callers must record THIS value on the session, never the requested {@code role}, so
* a later roster read never reports a role the session does not hold (CB-619). In
* normal operation this fallback should not happen once a reservation has been bound.
* It remains the honest answer if a caller has no reservation, or if binding a
* reservation unexpectedly fails.
*/
MemberRole acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
/**
* Refuse an acquire before anything spawns when {@code role} requires a live slot binding and
* no configured slot carries {@code profile} (CB-619 / fleetd #123). A no-op for a role with
* no slot-binding semantics.
*
* @throws IllegalArgumentException naming the role, the profile, and the pools that do carry it
*/
void requireSlotFor(MemberRole role, String profile);
/** Reserve a matching slot before launch, or refuse before a charter can be delivered. */
SlotReservation reserve(MemberRole role, String profile);
/** Convert a reservation into a live terminal binding. */
boolean bind(SlotReservation reservation, String terminal);
/** Return an unbound reservation after a failed launch. */
void release(SlotReservation reservation);
}
@@ -8,6 +8,7 @@ import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
@@ -55,8 +56,10 @@ public final class MemberRegistry implements MemberLifecycle {
}
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(FleetConfig.Fleet fleet) {
@@ -166,7 +169,7 @@ public final class MemberRegistry implements MemberLifecycle {
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
if (terminalToSlot.containsValue(slot) || reservedSlots.contains(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
@@ -206,19 +209,116 @@ public final class MemberRegistry implements MemberLifecycle {
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*
* <p>CB-619 / fleetd #123: the return value is the role this session actually holds, and the
* caller is required to record THAT — never the requested {@code role} — on the session. Before
* this fix the caller kept the requested role regardless of whether the bind below succeeded, so
* a demoted session's {@code GET /members} row still said {@code "architect"} while
* {@code fleet_whoami} (which reads the live binding, not the request) correctly said
* {@code "worker"} — three sources of truth that disagreed about one live member, silently.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
public MemberRole acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
return role;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
return MemberRole.ARCHITECT;
}
}
// fleetd #123: at least WARN — a role downgrade that the roster must now also reflect is
// not routine bookkeeping. requireSlotFor already refuses the config-gap case (no slot at
// all carries this profile) before a process ever spawns; reaching here means the config DID
// carry a matching slot but every one of them was already bound to a different terminal — a
// race this pre-spawn check cannot close on its own (see requireSlotFor's javadoc).
log.warn("member slot: no free architect slot for profile={} terminal={}; holding the session "
+ "as {} instead of the architect it asked for — every configured slot for this "
+ "profile is already bound to a different terminal", profile, terminal,
MemberRole.DEV.wireName());
return MemberRole.DEV;
}
/**
* CB-619 / fleetd #123: refuse an architect acquire before anything spawns when no configured
* slot carries {@code profile} — the config-gap case from the original defect report (a spawn
* asked for {@code role=architect, profile=sonnet}, and {@code fleet.architects} carried only
* {@code opus} and {@code sol}). A dev/reviewer acquire is always a no-op: those pools are
* placement candidates only (see {@code CompositePeerLauncher}), never a live identity binding,
* so there is nothing here to refuse — an explicit profile outside the pool for those roles is a
* documented operator override, not a defect.
*
* <p>This closes the config-gap case, not the live-capacity case: a profile that DOES carry a
* slot can still lose the race to a concurrent spawn between this check and the actual
* {@link #bind}, which is why {@link #acquired} must still answer honestly even after this
* check has passed.
*/
@Override
public void requireSlotFor(MemberRole role, String profile) {
if (role != MemberRole.ARCHITECT) {
return;
}
boolean hasSlot = slotsFor(MemberRole.ARCHITECT).values().stream()
.anyMatch(e -> Objects.equals(profile, e.profile()));
if (hasSlot) {
return;
}
List<String> pools = slotsFor(MemberRole.ARCHITECT).values().stream()
.map(Entry::profile)
.distinct()
.toList();
throw new IllegalArgumentException(
"no " + role.wireName() + " slot for profile '" + profile + "' — an architect's "
+ "identity IS the slot it is bound to, so there is nothing to bind this "
+ "session's identity to. fleet." + role.configKey() + " carries profiles: "
+ (pools.isEmpty() ? "(none configured)" : String.join(", ", pools))
+ "; add profile '" + profile + "' there, or spawn " + role.wireName()
+ " on one of those profiles instead");
}
@Override
public SlotReservation reserve(MemberRole role, String profile) {
if (role != MemberRole.ARCHITECT) {
return null;
}
synchronized (terminalToSlot) {
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if ((profile == null || profile.isBlank() || Objects.equals(profile, entry.profile()))
&& !terminalToSlot.containsValue(entry.key()) && reservedSlots.add(entry.key())) {
return new SlotReservation(entry.key(), entry.profile());
}
}
}
throw new IllegalArgumentException("no free architect slot for profile '" + profile
+ "' — every matching slot is already bound or reserved");
}
@Override
public boolean bind(SlotReservation reservation, String terminal) {
if (reservation == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!reservedSlots.remove(reservation.slot())) {
return false;
}
if (!isSlot(reservation.slot()) || terminalToSlot.containsKey(terminal)
|| terminalToSlot.containsValue(reservation.slot())) {
return false;
}
terminalToSlot.put(terminal, reservation.slot());
return true;
}
}
@Override
public void release(SlotReservation reservation) {
if (reservation != null) {
synchronized (terminalToSlot) {
reservedSlots.remove(reservation.slot());
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@@ -1,5 +1,7 @@
package dev.ltms.fleet.inject;
import java.util.function.Supplier;
/**
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
@@ -10,22 +12,81 @@ package dev.ltms.fleet.inject;
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
* the sink's job — see {@code Fleetd.main}'s wiring, which resolves target → session → profile →
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
*
* <p><strong>The 3-arg overload is the single abstract method — fleetd #234, round 4.</strong> Two
* earlier rounds each shipped a caller that silently dropped the profile hint (see {@link
* #onExhausted(String, String, String)}): a lambda written against this interface can only ever
* implement whichever overload is abstract, and while the 2-arg form held that position, EVERY
* lambda site — a call-site forwarder in {@code Fleetd.java}, a hand-built test double — bound to
* it and silently inherited the profile-dropping default, whether or not its author remembered the
* hazard. Making the 3-arg form abstract instead removes the shape entirely: a lambda declared
* against this interface today is <em>forced</em> by the compiler to take {@code (target, reason,
* profile)}, so there is no overload left for it to bind to that can drop the hint. This is a type
* change, not a test — it holds even for a caller that has never heard of fleetd #234.
*/
@FunctionalInterface
public interface ExhaustionSink {
/**
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
* @param reason the matched-line reason carried by the classification
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
* @param reason the matched-line reason carried by the classification
* @param profile the profile the caller already knows should be quarantined, or {@code null}
* when the caller has no better answer than {@code target} alone (fleetd #234):
* a caller whose {@code target} is not yet resolvable through whatever roster
* the sink's implementation consults — {@link
* dev.ltms.fleet.member.OpenCodeLauncher}'s model-mismatch check fires from
* {@code SessionAwareHandle.agentSessionId()}, which runs during {@code
* SessionManager.acquire()} <em>before</em> that session is registered, so a
* target -> session -> profile lookup finds nothing at that point. That launcher
* already has its own {@code FleetConfig.Profile} in hand and does not need the
* roster to know which profile to quarantine, so it supplies this directly
* instead of leaving the sink to guess.
*/
void onExhausted(String target, String reason);
void onExhausted(String target, String reason, String profile);
/**
* Convenience for a caller with no profile to offer — every existing call site that predates
* the hint (fleetd #234): {@link dev.ltms.fleet.inject.CompletionResolver}'s two call sites
* always call with a {@code target} that IS live in the roster at the time of the call, so they
* need no hint and keep working exactly as before, unchanged by this default.
*
* @param target as {@link #onExhausted(String, String, String)}
* @param reason as {@link #onExhausted(String, String, String)}
*/
default void onExhausted(String target, String reason) {
onExhausted(target, reason, null);
}
/**
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
* exercising this feature) passes instead of a defaulting overload, exactly like
* {@link ExhaustedPatternLookup#none()}.
* {@link ExhaustedPatternLookup#none()}. Safe as a lambda: the 3-arg form is now the interface's
* single abstract method, so a lambda here has no other overload to silently bind to instead —
* it simply does nothing with all three arguments.
*/
static ExhaustionSink none() {
return (target, reason) -> { };
return (target, reason, profile) -> { };
}
/**
* A sink that forwards to whatever {@code target} currently supplies (fleetd #234). Exists to
* break a genuine construction-order cycle: {@code Fleetd.main} builds its adapters (including
* {@link dev.ltms.fleet.member.OpenCodeLauncher}) before {@code sessions} exists, so it cannot
* hand them the real sink yet — it hands them a forwarder pointed at an {@code
* AtomicReference<ExhaustionSink>} that starts at {@link #none()} and gets {@code .set()} to the
* real sink once {@code sessions} is built. {@code target} is evaluated on every call, never
* cached, so the forwarder keeps working after the reference is repointed.
*
* <p>Now safe as a one-line lambda (round 4): forwarding the single 3-arg abstract method
* forwards everything a caller can supply — there is no separate 2-arg overload left for a
* forwarder to bind to instead and silently lose the hint. Kept as a named factory rather than
* written inline at each call site anyway, so a test can call the exact object {@code
* Fleetd.java} builds instead of asserting a rebuilt copy of its shape (round 3's lesson).
*
* @param target supplies the sink to forward to, evaluated fresh on every call
* @return a sink whose call delegates to {@code target.get()}
*/
static ExhaustionSink forwardingTo(Supplier<ExhaustionSink> target) {
return (t, r, p) -> target.get().onExhausted(t, r, p);
}
}
@@ -459,12 +459,83 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* {@link OpenCodeLauncher#writeConfig} already uses for its charter file, since the process that
* reads this file (the spawned peer) outlives this JVM call and there is no spawn-scoped teardown
* hook to delete it synchronously.
*
* <p><b>fleetd #222.</b> {@code Files.createTempFile(prefix, suffix)} with no directory argument
* resolves against {@code java.io.tmpdir} — on macOS the per-user {@code $TMPDIR} under
* {@code /var/folders/...}, mode {@code 0700}, both resolved against FLEETD's own OS user. Under
* {@code memberHerdrSocket:} the member pane runs as a DIFFERENT OS user, so that user cannot even
* traverse the directory, let alone read the file — and since fleetd #220 the charter file is the
* ONLY delivery path for {@code --append-system-prompt-file}, always, not merely the fallback it
* used to be. A member handed a path it cannot read is not degraded, it is broken: see this
* method's refusal branch below.
*
* <p><b>Measured severity (fleetd #222 real-binary check, claude 2.1.258):</b> an unreadable
* {@code --append-system-prompt-file} is the LOUD failure, not the silent one. {@code claude}
* checks the file before touching auth or the network — invoked with a bogus API key against a
* {@code chmod 000} file, it printed {@code Error reading append system prompt file: EACCES:
* permission denied, open '<path>'} and exited 1 immediately (a nonexistent path gets {@code
* Error: Append system prompt file not found: <path>}, same exit code). So the pre-fix bug did
* NOT leave a charter-less member silently occupying a pane and never calling {@code
* fleet_reply} — it made the herdr pane exit immediately, which the CB-306 spawn-readiness gate
* (this launcher's {@code spawnReadyTimeoutMs} poll) would have surfaced as "did not reach
* injectable state", the same unexplained-timeout shape fleetd #220 already describes. Still a
* real defect (every claude-code member under {@code memberHerdrSocket} would have failed to
* spawn), but not the worse, undetectable failure mode.
*
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only live mode): byte-identical to before this
* fix — {@code Files.createTempFile("fleetd-role-charter-", ".md")} with no directory
* argument, i.e. still resolved against {@code java.io.tmpdir}.</li>
* <li>{@code memberHerdrSocket} PRESENT: a fresh per-spawn directory is created under {@code
* worktreeRoot} (never {@code java.io.tmpdir}) holding just the charter file, then shared
* read-only with {@code worktreeGroup} via {@link EnvAllowListScrub#shareWithGroup} — the
* SAME mechanism fleetd #213 built for the ZDOTDIR scrub and fleetd #219 reused for {@link
* OpenCodeLauncher#writeConfig}'s {@code opencode.json} directory, reused here rather than
* duplicated a third time. A per-spawn subdirectory (not {@code worktreeRoot} itself) is the
* unit {@code shareWithGroup} chmods, so this never touches permissions on anything else
* under {@code worktreeRoot}.</li>
* </ul>
*
* <p><b>Follows fleetd #219's REFUSAL decision, not #213's degrade decision.</b> The ZDOTDIR
* scrub is a credential CONTROL — a degraded control (the CB-596 sentinel overlay) still has
* value, so #213 falls back rather than refusing. A charter is NOT a control, it is the member's
* TURN CONTRACT (the rule that ends every turn with {@code fleet_reply}). A member spawned with no
* charter is not degraded, it is broken: either claude-code exits on the unreadable
* {@code --append-system-prompt-file} path and the spawn dies at the readiness gate (loud), or it
* starts anyway with no charter and never calls {@code fleet_reply} — the sender silently gets
* nothing (silent). Neither outcome is worth trading for "spawn something." So a missing {@code
* worktreeRoot}/{@code worktreeGroup} under {@code memberHerdrSocket} refuses the spawn here,
* naming the missing key, exactly like {@link OpenCodeLauncher#configParentDir()}.
*
* @throws IllegalStateException when {@code memberHerdrSocket} is configured but {@code
* worktreeRoot} and/or {@code worktreeGroup} is not
*/
private static Path writeCharterFile(String charterText) {
private Path writeCharterFile(String charterText) {
try {
Path file = Files.createTempFile("fleetd-role-charter-", ".md");
if (!memberHerdrSocketConfigured()) {
Path file = Files.createTempFile("fleetd-role-charter-", ".md");
Files.writeString(file, charterText);
file.toFile().deleteOnExit();
return file;
}
Path parentDir = memberScrubParentDir();
String group = memberGroup();
if (parentDir == null || group == null) {
throw new IllegalStateException("memberHerdrSocket is configured, so the role/reply "
+ "charter file (mounted via --append-system-prompt-file) must be placed where "
+ "the member's OS user can read it — worktreeRoot, shared via worktreeGroup — "
+ "but " + (parentDir == null ? "worktreeRoot" : "worktreeGroup") + " is not "
+ "configured. Refusing to spawn rather than hand the member a charter path it "
+ "cannot read: that member's turn contract (the fleet_reply rule) would never "
+ "reach it. Configure both worktreeRoot and worktreeGroup to enable claude-code "
+ "member spawns under memberHerdrSocket.");
}
Path dir = Files.createTempDirectory(parentDir, "fleetd-role-charter-");
dir.toFile().deleteOnExit();
Path file = dir.resolve("charter.md");
Files.writeString(file, charterText);
file.toFile().deleteOnExit();
EnvAllowListScrub.shareWithGroup(dir, group);
return file;
} catch (IOException e) {
throw new UncheckedIOException("cannot write role charter temp file", e);
@@ -3,11 +3,13 @@ package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.Tab;
import dev.ltms.fleet.herdr.Workspace;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.StatusRefiner;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
@@ -151,6 +153,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
/**
* fleetd #176 fix 2: the same UNKNOWN-refinement {@code dev.ltms.fleet.inject.StatusPoller}
* uses, reused here for the spawn-readiness gate. Constructed once from {@link #agents} — see
* {@link #refinedInjectable(String, Agent)} for the corroboration that keeps it from firing on
* a dead pane's bare shell prompt.
*/
private final StatusRefiner statusRefiner;
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
@@ -272,6 +282,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
this.memberCredentials = memberCredentials;
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
this.config = config;
this.statusRefiner = new StatusRefiner(agents);
}
// --- adapter seams -------------------------------------------------------------------------
@@ -909,16 +920,35 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* Poll {@link AgentControl#get} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*
* <p>fleetd #176 fix 1: {@code agents.get} was previously unguarded here, so a herdr
* {@code *_not_found} answer — which is what happens when the backend process EXITED rather
* than being slow — propagated as a raw {@link HerdrException} instead of the
* {@link PeerUnreachableException} every other failure path of this gate produces, and skipped
* teardown ({@link #stop}) entirely, leaking the pane/tab. {@link #failFastOnGoneBackend} closes
* that gap: it stops waiting immediately (never burns the rest of the timeout), runs the same
* teardown the timeout path below runs, and throws with a message that says the backend exited
* rather than that the pane was slow. Any other {@link HerdrException} still propagates
* unchanged — this gate does not know how to recover from it.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
Object lastStatus = null;
long start = nowMillis.getAsLong();
long deadline = start + spawnReadyTimeoutMs;
AgentStatus lastStatus = null;
while (nowMillis.getAsLong() < deadline) {
var status = agents.status(paneId);
lastStatus = status;
if (status.injectable()) {
Agent sample;
try {
sample = agents.get(paneId);
} catch (HerdrException e) {
if (isAlreadyGone(e)) {
failFastOnGoneBackend(paneId, e, nowMillis.getAsLong() - start);
}
throw e; // any other herdr failure is not ours to interpret — let it propagate
}
lastStatus = sample.status();
if (lastStatus.injectable() || refinedInjectable(paneId, sample)) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
@@ -937,6 +967,66 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
+ spawnReadyTimeoutMs + "ms");
}
/**
* fleetd #176 fix 1: the backend process exited while the gate was still waiting — herdr
* answered {@code *_not_found} instead of ever reporting an injectable status. Fails
* immediately (never burns the rest of {@link #spawnReadyTimeoutMs}), runs the exact same
* teardown {@link #waitUntilInjectableOrThrow}'s timeout path runs, and throws a
* {@link PeerUnreachableException} whose message says the process exited rather than that the
* pane was slow.
*
* @throws PeerUnreachableException always — this method never returns normally
*/
private void failFastOnGoneBackend(String paneId, HerdrException cause, long elapsedMs) {
String tail = readPaneQuietly(paneId); // read before stop() closes the pane
log.warn("peer pane={} backend process exited after {}ms while waiting for injectable "
+ "state (herdr: {}) — closing. Pane tail:\n{}",
paneId, elapsedMs, cause.getMessage(), tail);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " backend process exited after " + elapsedMs
+ "ms while waiting to become injectable (herdr reported: " + cause.getMessage()
+ "). Pane tail:\n" + tail);
}
/**
* fleetd #176 fix 2: resolve a raw {@link AgentStatus#UNKNOWN} sample into a trustworthy
* injectable state via pane content, the same refinement {@code StatusPoller} applies during a
* peer's working life — but corroborated, because the trap this gate is exposed to that the
* poller is not: a pane whose backend has already exited settles at a plain shell prompt, and
* that prompt commonly contains the same {@code ❯} glyph {@link StatusRefiner#classify} treats
* as "idle at the Claude Code TUI prompt". Naively wiring the refiner in would turn "the backend
* died" into "ready to inject" — strictly worse than today's timeout.
*
* <p>Two guards, both required:
* <ul>
* <li>{@link StatusRefiner#classify} is written for the Claude Code TUI only (its own javadoc
* says so), so refinement only ever runs for the {@code claude} adapter — never for
* {@code opencode} or any future non-Claude backend, whatever its pane looks like.</li>
* <li>The refined status is accepted only when the <em>same</em> {@code agents.get} sample
* ({@code sample}, taken once by the caller) still reports a non-null
* {@link Agent#agentType()} — herdr's own "a supported backend is still detected here"
* signal, read from the very sample the raw status came from so the two can never
* disagree. A bare shell prompt left by an exited backend reports no {@code agentType}.</li>
* </ul>
*
* <p>Called only when {@code sample.status() == UNKNOWN}, so a healthy spawn — which never sees
* {@code UNKNOWN} — triggers zero extra herdr calls; only a persistently-{@code unknown} pane
* pays the one extra {@code agent.read} {@link StatusRefiner#refine} performs.
*/
private boolean refinedInjectable(String paneId, Agent sample) {
if (sample.status() != AgentStatus.UNKNOWN) {
return false;
}
if (!"claude".equals(namePrefix)) {
return false; // StatusRefiner.classify reads a Claude Code TUI prompt specifically
}
if (sample.agentType() == null) {
return false; // no corroborating liveness signal — could be a dead pane's bare shell
}
return statusRefiner.refine(paneId, AgentStatus.UNKNOWN, agents).injectable();
}
/**
* fleetd #220: the pane's recent output, clipped, for the readiness-gate timeout log — or a
* short note when it cannot be read. Best-effort by construction: this runs on a path that is
@@ -6,6 +6,7 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.PeerHandle;
@@ -21,6 +22,7 @@ import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.BooleanSupplier;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -73,6 +75,19 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
*/
private final OpenCodeSessionDiscovery discovery;
/**
* fleetd #175: notified when {@link SessionAwareHandle} detects, on the same late-resolve
* read that discovers the session id, that the live opencode session is running a DIFFERENT
* model than the profile requested — opencode does not fail on an unknown {@code -m}, it
* silently falls back to a default (potentially paid) model. Reused exactly as
* {@code CompletionResolver}'s {@code BACKEND_EXHAUSTED} path uses it: this launcher supplies
* only the herdr terminal id and a reason string; mapping target → session → profile →
* credential stays entirely the sink's job (see {@code Fleetd.main}'s wiring). Defaults to
* {@link ExhaustionSink#none()} for a caller (an older constructor, or a test not exercising
* this) that opts out.
*/
private final ExhaustionSink exhaustionSink;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
@@ -132,9 +147,25 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, spawnReadyPollMs,
fleet, memberCredentials, config, ExhaustionSink.none());
}
/**
* Production constructor, plus the live config for URI environment exclusions and the fleetd
* #175 model-mismatch quarantine sink.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config,
ExhaustionSink exhaustionSink) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials, config);
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials, config,
exhaustionSink);
}
/**
@@ -182,6 +213,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
this.exhaustionSink = ExhaustionSink.none();
}
/**
@@ -207,10 +239,29 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs, nowMillis, sleeper,
configRoot, discoveryRoot, fleet, memberCredentials, config, ExhaustionSink.none());
}
/**
* Full testability constructor, plus the live config for URI environment exclusions and the
* fleetd #175 model-mismatch quarantine sink — the constructor a test drives directly to
* observe {@link ExhaustionSink#onExhausted} without going through {@code Fleetd.main}'s
* wiring.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, FleetConfig.Profile> profiles, String defaultProfile,
Function<String, String> env, long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot, Path discoveryRoot,
Supplier<FleetConfig.Fleet> fleet,
Supplier<FleetConfig.MemberCredentials> memberCredentials,
Supplier<FleetConfig> config,
ExhaustionSink exhaustionSink) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, null, config);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
this.exhaustionSink = exhaustionSink == null ? ExhaustionSink.none() : exhaustionSink;
}
private static Path defaultConfigRoot() {
@@ -622,8 +673,13 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req),
this::memberHerdrSocketConfigured, discoveryUnavailableWarned);
// fleetd #175: the same profile config buildLaunch resolved for this spawn (requireProfile
// is deterministic on req.profileName(), so re-resolving here costs a map lookup, not a
// second decision) — SessionAwareHandle needs cfg.model() to know what THIS session should
// be running.
FleetConfig.Profile cfg = requireProfile(req.profileName());
return new SessionAwareHandle(inner, discovery, effectiveCwd(req), cfg,
this::memberHerdrSocketConfigured, discoveryUnavailableWarned, exhaustionSink);
}
/**
@@ -633,22 +689,48 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* are untouched — only the opencode-specific identity answer is added. {@code sessionName()}
* stays null: opencode has no display-name seam, so the logical name lives only in the bridge's
* roster (see the SESSION_NAME capability).
*
* <p>fleetd #175: {@link #agentSessionId()} is also where the model-mismatch check lives (see
* {@link #checkModelMatch()}) — it is the one method the real late-resolve path
* ({@code SessionManager.resolveAgentSessionId}, via {@code get}/{@code rosterResolved}/
* {@code release}) actually calls, and only while the session id is still unknown. Putting the
* check anywhere else risks repeating PR #203's mistake: a check that runs before opencode has
* written the row it needs, and so never fires.
*/
private static final class SessionAwareHandle implements PeerHandle {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
private final FleetConfig.Profile cfg;
private final BooleanSupplier discoveryUnavailable;
private final AtomicBoolean discoveryUnavailableWarned;
private final ExhaustionSink exhaustionSink;
/** CAS'd true the first (and only) time a model mismatch is reported for this handle. */
private final AtomicBoolean modelMismatchReported = new AtomicBoolean();
/**
* The session id, once {@link OpenCodeSessionDiscovery#sessionIdForDirectory} first
* resolves a non-null answer for this handle (fleetd #234). Sticky on purpose: {@code
* directory} is a shared-cwd heuristic (see {@link OpenCodeSessionDiscovery}'s class
* javadoc) that can start returning a DIFFERENT row once another session shares the same
* directory and writes a newer one — re-deriving it on every call would let this handle's
* identity silently drift to a sibling's session. Once resolved, this IS the answer, and
* {@link #checkModelMatch} reads only the row this id names, never "whatever is newest in
* the directory right now."
*/
private final AtomicReference<String> resolvedSessionId = new AtomicReference<>();
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd,
FleetConfig.Profile cfg,
BooleanSupplier discoveryUnavailable,
AtomicBoolean discoveryUnavailableWarned) {
AtomicBoolean discoveryUnavailableWarned,
ExhaustionSink exhaustionSink) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
this.cfg = cfg;
this.discoveryUnavailable = discoveryUnavailable;
this.discoveryUnavailableWarned = discoveryUnavailableWarned;
this.exhaustionSink = exhaustionSink;
}
@Override
@@ -689,11 +771,93 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
return null;
}
// fleetd #234: once resolved, stay resolved. Re-deriving from `directory` on every call
// would let this handle's identity drift to a sibling session that later shares the
// same cwd and writes a newer row — see resolvedSessionId's javadoc.
String cached = resolvedSessionId.get();
if (cached != null) {
return cached;
}
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
// appeared).
return discovery.sessionIdForDirectory(cwd);
String id = discovery.sessionIdForDirectory(cwd);
if (id != null) {
resolvedSessionId.compareAndSet(null, id);
}
// fleetd #175: check on the SAME tick — while the caller (SessionManager's late-resolve
// step) is still re-polling because the id is unknown, the row this id came from (once
// it exists) is exactly the row that also carries the actual model. Once id resolves,
// the caller stops calling agentSessionId() for this session, so this is naturally a
// once-only check that happens right when the row first appears.
checkModelMatch(id);
return id;
}
/**
* Verify the live opencode session is running the model {@link #cfg} requested (fleetd
* #175) and, on a real mismatch, log an ERROR and quarantine through {@link
* #exhaustionSink}. A no-op when there is nothing to compare against — no model configured,
* already reported once for this handle, {@code sessionId} itself is not resolved yet
* (fleetd #234: absent evidence, not a mismatch), or the actual model is still UNKNOWN (no
* row yet, unreadable database, or unparseable evidence). UNKNOWN must never be treated as
* a mismatch: that is the single most important safety rule here — a false positive would
* quarantine a perfectly working profile's credential.
*
* @param sessionId the id {@link #agentSessionId()} just resolved (or had cached) for THIS
* handle — the model is read back for this exact session (fleetd #234's
* {@link OpenCodeSessionDiscovery#actualModelForSessionId}), never
* re-derived from {@code directory}
*/
private void checkModelMatch(String sessionId) {
if (modelMismatchReported.get() || cfg.model() == null || cfg.model().isBlank()) {
return;
}
if (sessionId == null || sessionId.isBlank()) {
return; // id not resolved yet — UNKNOWN, never a mismatch (fleetd #175's rule)
}
OpenCodeSessionDiscovery.ActualModel actual = discovery.actualModelForSessionId(sessionId);
if (actual == null) {
return; // UNKNOWN evidence — never a mismatch
}
String[] requestedParts = splitProviderModel(cfg.model());
String requestedProvider = requestedParts == null ? null : requestedParts[0];
String requestedId = requestedParts == null ? cfg.model() : requestedParts[1];
boolean idMatches = requestedId.equals(actual.id());
// Compare the provider ONLY when BOTH sides have one. requestedProvider == null covers
// a bare profile model with no "/" — the profile never asked for a specific provider.
// actual.provider() == null covers opencode's model JSON having an id but no providerID
// (a real shape parseModel accepts) — that is missing evidence, not a contradiction, and
// acceptance rule 4 says missing evidence is UNKNOWN, never a mismatch. Narrowing the
// provider comparison this way keeps the id comparison (the part that actually caught the
// xf bug) fully intact — a genuine id mismatch is still caught either way (fleetd #175
// review round 2).
boolean providerMatches = requestedProvider == null || actual.provider() == null
|| requestedProvider.equals(actual.provider());
if (idMatches && providerMatches) {
return;
}
if (!modelMismatchReported.compareAndSet(false, true)) {
return; // another thread already reported this exact mismatch
}
String actualDisplay = actual.provider() == null
? actual.id() : actual.provider() + "/" + actual.id();
log.error("opencode profile '{}' requested model '{}' but the live session is actually "
+ "running '{}' — opencode does not fail on an unknown -m, it silently "
+ "falls back to a default model, which may be a PAID credential "
+ "(fleetd #175); quarantining this profile's credential",
cfg.profile(), cfg.model(), actualDisplay);
// fleetd #234: pass our OWN profile name too. This check fires from agentSessionId(),
// called during SessionManager.acquire() BEFORE this session is registered in
// sessions.roster() — a roster-only sink (Fleetd's target -> session -> profile lookup)
// finds nothing at this point and silently no-ops (defect 2). We already know exactly
// which profile to quarantine without the roster; the sink is passed it explicitly.
exhaustionSink.onExhausted(delegate.terminalId(),
"opencode model mismatch: profile '" + cfg.profile() + "' requested '"
+ cfg.model() + "' but the live session is running '" + actualDisplay
+ "' (fleetd #175)",
cfg.profile());
}
@Override
@@ -1,9 +1,12 @@
package dev.ltms.fleet.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.sqlite.SQLiteConfig;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.sql.Connection;
@@ -29,11 +32,19 @@ import java.util.concurrent.atomic.AtomicBoolean;
* here, so a layout change, or a switch to the HTTP server, changes exactly one class and nothing
* in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every fleetd worker runs
* in its own unique git worktree, so the row's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
* the CLI listing does not even show the directory.
* <p><strong>The {@code directory} match is a heuristic, not an identity — fleetd #234.</strong> A
* worktree is opt-in: {@code fleet_spawn} only provisions one when the caller passes {@code
* worktree:}; the default spawn inherits the lead's own cwd, which every other worker spawned the
* same way (and every past session ever run there) shares. {@code directory} therefore does
* <em>not</em> identify a session unambiguously in general — only in the special case of a fresh,
* unique worktree does "most recently updated row for this directory" reliably mean "this worker's
* own row." {@link #sessionIdForDirectory} still has to use this heuristic (the id has to come from
* somewhere, and nothing else is available at this layer — see that method's javadoc), but a caller
* that already holds a resolved id must never re-derive evidence about that same session via
* {@code directory} again; see {@link #actualModelForSessionId}, which looks up by {@code id}
* instead for exactly this reason. We match on {@code directory} rather than diffing
* {@code opencode session list} before/after — that races under concurrent spawns, and the CLI
* listing does not even show the directory.
*
* <p>All reads are best-effort and never throw: a missing or unreadable database, a query that
* fails, or a directory with no row yet all yield {@code null}, and the caller (the session
@@ -44,6 +55,18 @@ import java.util.concurrent.atomic.AtomicBoolean;
final class OpenCodeSessionDiscovery {
private static final Logger log = LoggerFactory.getLogger(OpenCodeSessionDiscovery.class);
private static final ObjectMapper MAPPER = new ObjectMapper();
/**
* The model opencode actually ran a session on, parsed from the {@code session.model} JSON
* column (fleetd #175). {@code id} is never null/blank on a non-null {@code ActualModel} —
* {@link #parseModel} returns {@code null} instead when {@code id} cannot be determined, so a
* caller only ever sees a fully-known record or {@code null} (UNKNOWN). {@code provider} may
* still be {@code null} on its own when the profile that requested the session named no
* provider prefix, or opencode's JSON omitted {@code providerID}.
*/
record ActualModel(String provider, String id) {
}
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final Path databasePath;
@@ -74,8 +97,16 @@ final class OpenCodeSessionDiscovery {
/**
* The opencode session id whose row references {@code directory} (the worker's cwd), or
* {@code null} when no row matches yet. When several rows share the directory — e.g. repeated
* spawns into the same worktree — the row with the highest {@code time_updated} wins: it is
* the session the pane most likely corresponds to.
* spawns into the same worktree, OR several workers sharing one cwd because none of them was
* given a worktree (fleetd #234) — the row with the highest {@code time_updated} wins: it is
* the session the pane most likely corresponds to. That "most likely" is a real caveat, not a
* formality: when the directory is shared, this can and does pick another session's row (see
* the class javadoc). This is the one place fleetd resolves an opencode session id at all —
* nothing else is available at this layer to disambiguate further (no {@code opencode session
* list} entry names the directory, and diffing before/after races under concurrent spawns) — so
* the heuristic stays here unchanged. What must never happen is a SECOND, independent piece of
* evidence about the same session being re-derived via {@code directory} once an id has already
* come out of this method; see {@link #actualModelForSessionId}.
*
* <p>Never throws: a missing {@code opencode.db}, a locked/unreadable database, a query
* failure, or a directory that has not been persisted yet all resolve to {@code null} rather
@@ -117,4 +148,85 @@ final class OpenCodeSessionDiscovery {
storageRoot, directory);
return null;
}
/**
* The model opencode actually ran the session {@code sessionId} on (fleetd #175/#234), read by
* primary-key lookup — the ONE row that id names, and no other. Deliberately keyed on
* {@code id} rather than {@code directory}: two independent {@code WHERE directory = ? ORDER BY
* time_updated DESC LIMIT 1} queries (one for the id, one for the model) can each pick a
* DIFFERENT row once more than one session shares a directory (fleetd #234 — the default
* no-worktree spawn shares the lead's cwd with every other worker and every past session ever
* run there), silently comparing a profile's requested model against a session that is not even
* the one whose id was returned. Keying on {@code id} instead makes that impossible: the model
* read back is always the SAME session {@link #sessionIdForDirectory} (or a cached copy of its
* answer) already resolved.
*
* <p>Kept as its own query and its own connection, independent from {@link
* #sessionIdForDirectory}: a database whose schema predates the {@code model} column (or any
* other read failure on this column alone) can never take id resolution down with it. That
* would be a regression of the id-resolution feature #209 shipped; this method degrades on its
* own.
*
* <p>Never throws, and every failure mode — no matching row, a missing/unreadable database, a
* missing {@code model} column, a null/blank {@code model} value, or JSON that does not parse
* into {@code {"id": "...", "providerID": "..."}} with a non-blank {@code id} — resolves to
* {@code null}. That is UNKNOWN evidence, not a mismatch signal: the caller must never
* quarantine a profile on the strength of a {@code null} here. A blank/null {@code sessionId}
* (the id is not resolved yet) is UNKNOWN too, for the same reason — never call this with one.
*
* @param sessionId the session id already resolved by {@link #sessionIdForDirectory} for this
* spawn — never re-derived from {@code directory} here
* @return the actual model, or {@code null} when unknown
*/
ActualModel actualModelForSessionId(String sessionId) {
if (sessionId == null || sessionId.isBlank()) {
return null;
}
if (!Files.isRegularFile(databasePath)) {
// sessionIdForDirectory already WARNs once (shared warnedMissingDatabase) for this
// exact condition — do not double-log it here.
return null;
}
String sql = "SELECT model FROM session WHERE id = ?";
try (Connection connection = openReadOnly();
PreparedStatement statement = connection.prepareStatement(sql)) {
statement.setString(1, sessionId);
try (ResultSet rows = statement.executeQuery()) {
if (rows.next()) {
return parseModel(rows.getString("model"));
}
}
} catch (SQLException e) {
// Locked/corrupt database, or a `model` column this schema version does not have —
// never fatal, and never a mismatch signal. See the class doc above.
log.debug("opencode session model unreadable at {}: {}", databasePath, e.toString());
return null;
}
return null;
}
/**
* Parse opencode's {@code model} column — {@code {"id":"...","providerID":"..."}} — into an
* {@link ActualModel}, or {@code null} when {@code json} is null/blank, is not valid JSON, or
* parses without a non-blank {@code id}. {@code providerID} may be absent; that alone does not
* make the record unknown, since a caller comparing against a profile with no provider prefix
* never looks at it.
*/
private static ActualModel parseModel(String json) {
if (json == null || json.isBlank()) {
return null;
}
try {
JsonNode node = MAPPER.readTree(json);
String id = node.path("id").asText(null);
if (id == null || id.isBlank()) {
return null;
}
String provider = node.path("providerID").asText(null);
return new ActualModel(provider, id);
} catch (IOException e) {
log.debug("opencode session model JSON unparseable: {}", e.toString());
return null;
}
}
}
@@ -13,6 +13,7 @@ import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.nio.file.attribute.PosixFilePermissions;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.HashSet;
@@ -161,6 +162,7 @@ public final class GitWorktrees implements Worktrees {
} catch (IOException e) {
throw new WorktreeException("cannot create worktree root " + root + ": " + e.getMessage(), e);
}
shareRootWithGroup(root);
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
removeUserInfoFromHttpsOrigin(repoRoot);
@@ -712,6 +714,54 @@ public final class GitWorktrees implements Worktrees {
group, repoRoot, worktreePath, touched);
}
/**
* fleetd #224: make {@code worktreeRoot} ITSELF group-traversable — established once, here,
* where the root is created, never at a use site. {@link #shareWithGroup} shares each worktree
* (and the repo's common git dir) with {@link #group}, but never the PARENT directory that
* contains every worktree — and under {@code memberHerdrSocket:} the member pane runs as a
* different OS user, which needs the execute bit on every ancestor directory to reach anything
* underneath, no matter how carefully each child is shared. Without this, a member cannot read
* the ephemeral {@code opencode.json} #219 places under this root, cannot reach its own
* worktree, and cannot read #213's ZDOTDIR scrub when placed here either.
*
* <p>No-op — no process spawned — when {@link #group} is null/blank, so behaviour with
* {@code worktreeGroup:} unset (today's only live mode) is unchanged. Only {@code root} itself
* is touched (non-recursive): each child underneath is shared individually, either by
* {@link #shareWithGroup} for a worktree or by the launcher that generates it (fleetd #213/#219)
* for a scrub/config directory — sharing this level again would just duplicate that policy in
* the wrong layer.
*
* <p>Fails loudly, naming {@code root}, its mode at the time of the attempt, and {@link #group}:
* a member that starts and then cannot see its own checkout is worse than a refused spawn, since
* nothing about that failure mode points at a directory's permission bits.
*/
private void shareRootWithGroup(Path root) {
if (group == null) {
return;
}
String mode = currentPosixMode(root);
try {
shareGroupRunner.apply(new String[]{"chgrp", group, root.toString()});
shareGroupRunner.apply(new String[]{"chmod", "g+x", root.toString()});
} catch (WorktreeException e) {
throw new WorktreeException("cannot make worktree root " + root + " (mode " + mode
+ ") group-traversable for group '" + group + "': " + e.getMessage()
+ " — the group must exist, and the fleetd operator (" + System.getProperty("user.name")
+ ") must be a member of it", e);
}
log.info("worktreeGroup={} made worktree root {} group-traversable (was mode {})", group, root, mode);
}
/** {@code root}'s current POSIX permission string, or {@code "unknown"} on a filesystem that does
* not support POSIX permissions — used only to name the mode in a refusal message. */
private static String currentPosixMode(Path root) {
try {
return PosixFilePermissions.toString(Files.getPosixFilePermissions(root));
} catch (IOException | UnsupportedOperationException e) {
return "unknown";
}
}
/**
* The repo's <em>common</em> git directory as an absolute path — where {@code objects},
* {@code refs} and {@code worktrees} actually live. {@code git rev-parse --git-common-dir}
@@ -183,46 +183,51 @@ public final class SessionManager implements TurnListener {
String sessionName, String resumeSessionId) {
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
requireResumeCapability(profile, resumeSessionId);
// CB-619 / fleetd #123: an explicit profile bypasses placement (CompositePeerLauncher only
// constrains an UNQUALIFIED spawn to the role's pool), so it is the one path that can ask
// for a role with no slot to bind it to. Refuse before anything spawns. A blank profile is
// left to placement, which already restricts an unqualified spawn to the role's pool.
if (profile != null && !profile.isBlank()) {
memberLifecycle.requireSlotFor(memberRole, profile);
}
MemberLifecycle.SlotReservation reservation = memberLifecycle.reserve(memberRole, profile);
String launchProfile = reservation == null ? profile : reservation.profile();
if (wt == null) {
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
SpawnRequest req = new SpawnRequest(launchProfile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
PeerHandle handle;
boolean bound = false;
try {
handle = launcher.spawn(req);
String resolvedProfile = resolveProfile(handle, launchProfile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
MemberRole actualRole = acquired(memberRole, resolvedProfile, handle.terminalId(), reservation);
bound = reservation == null || actualRole == MemberRole.ARCHITECT;
MemberSession session = new MemberSession(
handle.id(), handle.terminalId(), resolvedProfile, actualRole, cwd, ownerTerminal, now, now, 0,
MemberSession.State.SPAWNING, null, null, handle.charterReceipt(), handle.agentSessionId());
registry.put(handle.id(), session);
handles.put(handle.id(), handle);
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={}: {}", profile, memberRole, e.getMessage());
if (!bound) memberLifecycle.release(reservation);
log.warn("spawn failed for profile={} role={}: {}", launchProfile, memberRole, e.getMessage());
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
memberRole,
cwd,
ownerTerminal,
now,
now,
0,
MemberSession.State.SPAWNING,
null,
null,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
handles.put(handle.id(), handle);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId);
try {
return acquireWithWorktree(launchProfile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId, reservation);
} catch (RuntimeException e) {
memberLifecycle.release(reservation);
throw e;
}
}
/**
@@ -468,7 +473,8 @@ public final class SessionManager implements TurnListener {
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
String sessionName, String resumeSessionId,
MemberLifecycle.SlotReservation reservation) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
@@ -509,11 +515,14 @@ public final class SessionManager implements TurnListener {
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
// CB-619: see the no-worktree path above — bind before recording, and store the returned
// actual role, so this session's role is never a lie about what it actually holds.
MemberRole actualRole = acquired(memberRole, resolvedProfile, handle.terminalId(), reservation);
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
memberRole,
actualRole,
cwd,
ownerTerminal,
now,
@@ -526,13 +535,24 @@ public final class SessionManager implements TurnListener {
handle.agentSessionId());
registry.put(handle.id(), session);
handles.put(handle.id(), handle);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
private MemberRole acquired(MemberRole role, String profile, String terminal,
MemberLifecycle.SlotReservation reservation) {
if (reservation == null) {
return memberLifecycle.acquired(role, profile, terminal);
}
if (memberLifecycle.bind(reservation, terminal)) {
return role;
}
memberLifecycle.release(reservation);
return memberLifecycle.acquired(role, profile, terminal);
}
private String slug(String raw) {
return raw == null ? "ticket" : raw.toLowerCase().replaceAll("[^a-z0-9]+", "-").replaceAll("^-+|-+$", "");
}
@@ -44,10 +44,13 @@ public final class FakeHerdr implements HerdrClient {
private String agentSendErrorCode = null;
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -103,6 +106,27 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the detected agent kind ({@code "agent"} field) that {@code agent.get} reports —
* {@code null} models herdr not (or no longer) detecting a supported backend in the pane, e.g.
* a bare shell prompt (fleetd #176 fix 2 corroboration test).
*/
public FakeHerdr agentType(String type) {
this.agentType = type;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
* fixture. {@code okCalls == 0} fails from the very first call.
*/
public FakeHerdr agentGetFailsWithAfter(int okCalls, String code) {
this.agentGetOkCalls = okCalls;
this.agentGetFailCode = code;
return this;
}
/** The text {@code agent.read} returns (the CB-106 completion scrape). */
public FakeHerdr readText(String text) {
this.readText = text;
@@ -208,10 +232,21 @@ public final class FakeHerdr implements HerdrClient {
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
case "agent.get" -> mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":"claude",
case "agent.get" -> {
if (agentGetFailCode != null) {
long getCalls = calls.stream().filter(c -> c.method().equals("agent.get")).count();
if (getCalls > agentGetOkCalls) {
throw new HerdrException(
"herdr error [" + agentGetFailCode + "]: agent target not found",
agentGetFailCode, null);
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
.formatted(agentField, agentStatus));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
@@ -499,7 +499,7 @@ class CompletionResolverTest {
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
@@ -520,7 +520,7 @@ class CompletionResolverTest {
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
@@ -648,7 +648,7 @@ class CompletionResolverTest {
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
@@ -701,7 +701,7 @@ class CompletionResolverTest {
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
@@ -728,7 +728,7 @@ class CompletionResolverTest {
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
ExhaustionSink sink = (target, reason, profile) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
@@ -0,0 +1,41 @@
package dev.ltms.fleet.inject;
import org.junit.jupiter.api.Test;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #234. Round 3: calls the REAL production factory, {@link ExhaustionSink#forwardingTo},
* rather than rebuilding its shape locally — a round-2 version of this test built its own copy of
* the forwarder inline, so mutating {@code Fleetd.java}'s actual forwarder back into a lambda left
* that copy untouched and the test kept passing while production had regressed. Calling the shared
* factory here means the same object under test IS the object {@code Fleetd.java} builds via
* {@code ExhaustionSink.forwardingTo(exhaustionSinkRef::get)}.
*
* <p>Round 4: the 3-arg overload is now {@link ExhaustionSink}'s single abstract method, so {@code
* forwardingTo} itself is a one-line lambda and every sink below can safely be one too — there is
* no 2-arg overload left for any of them to silently bind to instead. This test still catches a
* regression inside {@code forwardingTo}'s body (e.g. one that calls the 2-arg default and drops
* the hint that way) because it still goes through the shared factory rather than a rebuilt copy.
*/
class ExhaustionSinkForwardingHazardTest {
@Test
void forwardingToDeliversTheProfileHintToWhateverSinkTheSupplierCurrentlyReturns() {
AtomicReference<String> hintSeenByRealSink = new AtomicReference<>("NEVER CALLED");
ExhaustionSink real = (target, reason, profileHint) -> hintSeenByRealSink.set(profileHint);
// Same shape Fleetd.java uses: a reference that starts at none() and is repointed later.
AtomicReference<ExhaustionSink> ref = new AtomicReference<>(ExhaustionSink.none());
ExhaustionSink forwarding = ExhaustionSink.forwardingTo(ref::get);
ref.set(real);
forwarding.onExhausted("term_x", "model mismatch", "gx");
assertEquals("gx", hintSeenByRealSink.get(),
"ExhaustionSink.forwardingTo must deliver the profile hint to the sink the supplier "
+ "currently returns — the exact object Fleetd.java's forwarder is built from");
}
}
@@ -1,10 +1,13 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.msg.MessageService;
@@ -1024,6 +1027,87 @@ class FleetMcpTest {
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
}
// ── CB-619 / fleetd #123: a spawn asking for a role its profile has no slot for must be
// refused, never silently demoted with the roster still lying about it ───────────────────
/** Two profiles on one launcher, so a role's pool can name one and exclude the other. */
private static ClaudeCodeLauncher architectCapableLauncher(FakeHerdr h) {
FleetConfig.Profile opus = new FleetConfig.Profile(
"opus", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.Profile sonnet = new FleetConfig.Profile(
"sonnet", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("opus", opus, "sonnet", sonnet),
"sonnet", _ -> "tok");
}
/** {@code fleet.architects} carries only {@code opus} — {@code sonnet} has no matching slot. */
private static MemberRegistry architectRegistry() {
return new MemberRegistry(new FleetConfig.Fleet(Map.of(),
Map.of("opus", new FleetConfig.Slot("opus")), Map.of(), Map.of(), null));
}
/**
* The literal defect (fleetd #123): {@code role=architect, profile=sonnet}, where
* {@code fleet.architects} carries only {@code opus}. Drives the real path —
* {@link FleetMcp#spawn} calls {@link SessionManager#acquire}, which must refuse before ever
* reaching the real {@link ClaudeCodeLauncher} — never {@link MemberRegistry#bind} called
* directly, which would walk around the gate under test.
*/
@Test
void spawnRefusesAnArchitectWithNoMatchingSlotAndNeverTouchesTheLauncher() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(architectCapableLauncher(h));
sessions.setMemberLifecycle(architectRegistry());
McpSchema.CallToolResult res = FleetMcp.spawn(sessions, "sonnet", "architect",
null, null, null, null, null, null);
assertEquals(Boolean.TRUE, res.isError(), textOf(res));
String msg = textOf(res);
assertTrue(msg.contains("architect"), "names the role asked for: " + msg);
assertTrue(msg.contains("sonnet"), "names the profile: " + msg);
assertTrue(msg.contains("opus"), "names the pool that does carry the role: " + msg);
assertTrue(sessions.roster().isEmpty(), "a refused spawn must register no session");
assertTrue(h.calls.stream().noneMatch(c -> c.method().equals("agent.start")),
"a refused spawn must never reach the launcher — no process should ever start");
}
/**
* Positive control / parity check: when the profile DOES carry a slot, the spawn succeeds, and
* {@code GET /members} ({@link FleetMcp#listFleet}) and {@code fleet_whoami}
* ({@link FleetMcp#whoami}) — resolved through the SAME live {@link MemberRegistry} binding via
* a real {@link CallerResolver}, exactly as the daemon resolves a real MCP caller — must never
* disagree about this one live member's role.
*/
@Test
void rosterAndWhoamiAgreeOnceTheArchitectSlotBinds() {
FakeHerdr h = new FakeHerdr();
// Pin the spawn onto the one pane FakeHerdr's canned pane.process_info maps to WORKER_PID,
// so a CallerResolver can resolve THIS session's own terminal, not a fixture double.
h.pinNextStarts(1, "term_a", "w2:p7");
SessionManager sessions = new SessionManager(architectCapableLauncher(h));
MemberRegistry members = architectRegistry();
sessions.setMemberLifecycle(members);
McpSchema.CallToolResult spawnRes = FleetMcp.spawn(sessions, "opus", "architect",
null, null, null, null, null, null);
assertNotEquals(Boolean.TRUE, spawnRes.isError(), textOf(spawnRes));
assertTrue(textOf(spawnRes).contains("\"role\":\"architect\""), textOf(spawnRes));
String roster = textOf(FleetMcp.listFleet(architectCapableLauncher(h), sessions, Map.of(), ""));
assertTrue(roster.contains("\"role\":\"architect\""), "GET /members: " + roster);
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(h), _ -> FakeHerdr.WORKER_PID);
CallerResolver resolver = CallerResolver.withLeadsAndMembers(identity, false, null, Map::of, members);
Principal caller = resolver.resolve("127.0.0.1", 42, null);
assertTrue(caller.isArchitect(), "fleet_whoami's own resolver must agree the terminal is bound");
String whoami = textOf(FleetMcp.whoami(caller, sessions));
assertTrue(whoami.contains("\"role\":\"architect\""), "fleet_whoami: " + whoami);
}
// ── CB-584: fleet_spawn accepts sessionName/resumeSessionId; roster shows agentSessionId ──
@Test
@@ -20,8 +20,10 @@ import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFilePermissions;
import java.util.List;
import java.util.Map;
import java.util.Set;
@@ -31,6 +33,7 @@ import java.util.function.Function;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
class ClaudeCodeLauncherTest {
@@ -960,6 +963,152 @@ class ClaudeCodeLauncherTest {
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- fleetd #176 fix 1: fail fast when the backend process exits mid-wait -------------------
@Test
void spawnFailsFastWhenBackendProcessExitsMidWaitInsteadOfBurningTheTimeout() {
// First agent.get sees UNKNOWN (one normal tick); the second reports the pane gone, exactly
// what herdr answers when the backend process has already exited. The timeout is generous
// (60s) so a test that wrongly falls through to the old unguarded call — and therefore
// waits out the whole window — is unambiguously distinguishable from one that fails fast.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(1, "pane_not_found");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().toLowerCase().contains("exited"),
"exception message says the backend exited, not that the pane was slow: "
+ ex.getMessage());
assertTrue(clock[0] < 60_000,
"the gate must not burn the rest of the 60s timeout: clock only reached " + clock[0]);
long getCalls = herdr.calls.stream().filter(c -> c.method().equals("agent.get")).count();
assertEquals(2, getCalls,
"exactly one normal poll then the not_found answer — no further polling after that: "
+ getCalls);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down (no orphan) even on the fail-fast path");
}
@Test
void spawnLetsAnUnrelatedHerdrErrorPropagateUnchanged() {
// Fix 1 must only special-case a "*_not_found" answer. Any other herdr failure keeps
// propagating as-is — this gate does not know how to recover from it.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
herdr.agentGetFailsWithAfter(0, "internal_error");
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
dev.ltms.fleet.herdr.HerdrException ex = assertThrows(
dev.ltms.fleet.herdr.HerdrException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertEquals("internal_error", ex.code());
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"an error this gate does not recognize is not this gate's teardown to run");
}
// --- fleetd #176 fix 2: corroborated UNKNOWN refinement --------------------------------------
@Test
void refinedIdleIsNotAcceptedWhenAgentTypeIsNull() {
// The trap fix 2 must close: a pane sitting at a bare shell prompt after its backend exited
// still contains the same "❯" glyph StatusRefiner.classify treats as "idle at the Claude
// Code TUI prompt". Without the agentType corroboration this would be misread as injectable
// and the gate would hand back a peer that never started. herdr reports no agentType for
// that bare shell.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType(null); // no supported backend detected — could be a bare shell
herdr.readText("some-host:~ user$ ❯ "); // looks exactly like an idle Claude Code prompt
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
500, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 500,
"a null agentType must not let the '❯' prompt refine to injectable — the gate has "
+ "to wait out the full timeout: clock only reached " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"the pane is torn down on timeout, same as any other never-injectable spawn");
}
@Test
void refinedIdleIsAcceptedWhenAgentTypeCorroboratesLiveness() {
// The positive case: a genuinely live claude pane that herdr misreports as UNKNOWN (CB-115)
// still resolves to injectable once agentType corroborates that a supported backend is
// detected in the same sample.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN at the raw level
herdr.agentType("claude"); // herdr still detects a live claude backend
herdr.readText("? for shortcuts"); // StatusRefiner.classify's idle footer marker
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
60_000, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "a corroborated refined-IDLE sample lets the spawn succeed");
assertTrue(clock[0] < 60_000,
"refinement must resolve well before the timeout: clock reached " + clock[0]);
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane close — the peer is genuinely injectable");
}
@Test
void refinementNeverFiresWhenRawStatusIsAlreadyInjectable() {
// Acceptance criterion 4: refinement must trigger only on a raw UNKNOWN sample. Seed the
// pane content with an active-generation marker that StatusRefiner.classify would read as
// WORKING (never injectable) if refine() were wrongly invoked here — proving that a raw
// IDLE status short-circuits before refine() (and its extra agent.read call) ever runs.
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("idle"); // already injectable at the raw level
herdr.agentType("claude");
herdr.readText("esc to interrupt"); // would classify as WORKING if refine() ran anyway
long[] clock = {0};
ClaudeCodeLauncher gated = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], () -> clock[0] += 300);
PeerHandle handle = gated.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "an already-injectable raw status succeeds without ever refining");
assertFalse(herdr.called("agent.read"),
"refine() must never run (and so never call agent.read) when the raw status is "
+ "already injectable");
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
@@ -1751,4 +1900,184 @@ class ClaudeCodeLauncherTest {
assertEquals(List.of("dev: sonnet #1", "[sonnet] dev 2"), tabLabels(herdr));
}
// ── fleetd #222: the charter file must not land under fleetd's own java.io.tmpdir when the ──
// ── member pane runs as a different OS user ─────────────────────────────────────────────────
/** A profile that mounts the bridge MCP (so a reply charter is always generated). */
private static FleetConfig.Profile charterCfg() {
return new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("claude"), "tab", "fleetd-workers", "worker: {profile} #{n}",
"http://127.0.0.1:8765/mcp", null, null);
}
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
private static FleetConfig configWithMemberHerdrSocket(String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
"/tmp/other-user.sock", // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
null // memberLoginShell
).withDefaults();
}
private static ClaudeCodeLauncher serviceWithConfig(FakeHerdr herdr, FleetConfig.Profile cfg,
Supplier<FleetConfig> config) {
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, null, null, null, config);
}
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225). Those two coincide only by accident: a directory's
* group is whichever group happened to own the path Maven was started from — {@code staff} in
* a home checkout, {@code wheel} under {@code /private/tmp} on macOS — and the fix-up this test
* exercises then fails for real when the operator is not a member of that borrowed group,
* exactly the case {@code assumeTrue(view != null, ...)} never covered (it only detects a
* filesystem with no POSIX groups at all, not a resolvable-but-wrong one). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
try (java.io.BufferedReader r = new java.io.BufferedReader(
new java.io.InputStreamReader(p.getInputStream(), java.nio.charset.StandardCharsets.UTF_8))) {
out = r.lines().collect(java.util.stream.Collectors.joining("\n")).trim();
}
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !out.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/**
* fleetd #222 acceptance criterion 1: with {@code memberHerdrSocket} configured and both
* {@code worktreeRoot}/{@code worktreeGroup} set, the charter file lives in a fresh per-spawn
* directory under {@code worktreeRoot} — NEVER under {@code java.io.tmpdir} (fleetd's own 0700
* temp dir, unreadable by the member's different OS user) — and that directory is shared
* read-only with the group via the SAME mechanism (fleetd #213/#219's {@link
* EnvAllowListScrub#shareWithGroup}) the ZDOTDIR scrub and the opencode config directory use.
*/
@Test
void memberHerdrSocketWithWorktreeRootAndGroupPutsCharterUnderWorktreeRootAndSharesIt(
@TempDir Path worktreeRoot) throws Exception {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
serviceWithConfig(herdr, charterCfg(),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), group)).spawn();
List<String> args = spawnedArgs(herdr);
int fileFlag = args.indexOf("--append-system-prompt-file");
assertTrue(fileFlag >= 0, "the charter is still mounted via file: " + args);
Path charterFile = Path.of(args.get(fileFlag + 1));
Path dir = charterFile.getParent();
// NOTE: JUnit's own @TempDir provider places worktreeRoot itself under java.io.tmpdir on this
// host, so "not under java.io.tmpdir" is not a meaningful assertion here (it would hold by
// accident of the fixture, not by anything this method does). What this fix actually promises
// is that the directory is created UNDER worktreeRoot specifically — never resolved from the
// no-argument Files.createTempFile default (java.io.tmpdir) the pre-fix code always used — so
// that is the assertion: the parent is exactly worktreeRoot, whatever directory JUnit gave it.
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated directory's parent must be worktreeRoot, not java.io.tmpdir — got "
+ "parent " + dir.getParent());
assertEquals("rwxr-x---", PosixFilePermissions.toString(Files.getPosixFilePermissions(dir)),
"the directory must be group-traversable+readable, owner-only writable");
assertEquals("rw-r-----", PosixFilePermissions.toString(Files.getPosixFilePermissions(charterFile)),
"the charter file must be group-readable, never group-writable");
assertTrue(Files.readString(charterFile).contains("fleet_reply"),
"the charter content itself is unaffected by where it is written");
}
/**
* fleetd #222 acceptance criterion 2: with {@code memberHerdrSocket} configured but NEITHER
* {@code worktreeRoot} nor {@code worktreeGroup} set, the launcher must refuse the spawn rather
* than hand the member a {@code --append-system-prompt-file} path under {@code java.io.tmpdir}
* it cannot read — the member's whole turn contract would never reach it.
*/
@Test
void memberHerdrSocketWithoutWorktreeRootOrGroupRefusesTheSpawn() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher launcher = serviceWithConfig(herdr, charterCfg(),
() -> configWithMemberHerdrSocket(null, null));
IllegalStateException ex = assertThrows(IllegalStateException.class, launcher::spawn,
"a missing worktreeRoot/worktreeGroup must refuse the spawn, not write an unreadable charter");
assertTrue(ex.getMessage().contains("worktreeRoot"),
"the refusal must name the missing config key — got: " + ex.getMessage());
assertFalse(herdr.called("agent.start"),
"the spawn must be refused BEFORE the member is ever started — got calls: " + herdr.calls);
}
/**
* fleetd #222 acceptance criterion 2 (the other missing half): {@code worktreeRoot} set but
* {@code worktreeGroup} missing must ALSO refuse — either one alone is not enough to guarantee
* the member's OS user can read the charter file.
*/
@Test
void memberHerdrSocketWithWorktreeRootButNoGroupRefusesTheSpawn(@TempDir Path worktreeRoot) {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher launcher = serviceWithConfig(herdr, charterCfg(),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), null));
IllegalStateException ex = assertThrows(IllegalStateException.class, launcher::spawn);
assertTrue(ex.getMessage().contains("worktreeGroup"),
"worktreeRoot alone is not enough — got: " + ex.getMessage());
}
/**
* fleetd #222 acceptance criterion 3: with {@code memberHerdrSocket} ABSENT — even when a live,
* non-null {@code config} supplier is threaded through (not merely {@code config == null}, which
* every other test in this file already exercises) — the charter file must still be created
* directly under {@code java.io.tmpdir} via the same no-directory-argument
* {@code Files.createTempFile} call as before this fix, byte-identical to today.
*/
@Test
void memberHerdrSocketAbsentStaysUnderJavaIoTmpdirEvenWithALiveConfigSupplier() throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig config = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, null, null).withDefaults();
serviceWithConfig(herdr, charterCfg(), () -> config).spawn();
List<String> args = spawnedArgs(herdr);
int fileFlag = args.indexOf("--append-system-prompt-file");
assertTrue(fileFlag >= 0);
Path charterFile = Path.of(args.get(fileFlag + 1));
assertTrue(charterFile.startsWith(Path.of(System.getProperty("java.io.tmpdir"))),
"with memberHerdrSocket absent, the charter file must still land directly under "
+ "java.io.tmpdir, unchanged from before this fix");
assertTrue(charterFile.getFileName().toString().startsWith("fleetd-role-charter-"),
"same file-naming scheme as before this fix (no wrapping directory): " + charterFile);
}
}
@@ -114,6 +114,68 @@ class EnvAllowListScrubTest {
assertNull(EnvAllowListScrub.readReport(dir));
}
/**
* fleetd #224 criterion 5 (a gap the #221 reviewer flagged): {@link EnvAllowListScrub#shareWithGroup}
* lists one FLAT level of {@code dir} and shares every file it finds there — which does cover the
* OPTIONAL files a caller may or may not have written before calling it ({@code member-charter.md},
* {@code ide-rules.md} — both written by {@code OpenCodeLauncher#writeConfig}), but nothing
* pinned that down.
* Without this test, a future change that writes a file AFTER the sharing call, or into a
* subdirectory, would pass every existing test while quietly leaving that file unreadable to a
* different-uid member.
*
* <p>This drives {@code shareWithGroup} directly against a directory holding several flat files —
* not only {@code opencode.json}, but also the two optional ones named above plus a third,
* unrelated file, so the assertion is "every flat file", not "the two files someone thought of".
*/
@Test
void shareWithGroupCoversEveryFlatFileIncludingTheOptionalOnes(@TempDir Path dir) throws Exception {
String group = currentUserGroup();
Files.writeString(dir.resolve("opencode.json"), "{}\n");
Files.writeString(dir.resolve("member-charter.md"), "# charter\n");
Files.writeString(dir.resolve("ide-rules.md"), "# ide rules\n");
Files.writeString(dir.resolve("another-flat-file.txt"), "unrelated\n");
EnvAllowListScrub.shareWithGroup(dir, group);
assertEquals("rwxr-x---", java.nio.file.attribute.PosixFilePermissions.toString(
Files.getPosixFilePermissions(dir)),
"the directory itself must be group-traversable+readable, owner-only writable");
for (String name : List.of("opencode.json", "member-charter.md", "ide-rules.md", "another-flat-file.txt")) {
Path file = dir.resolve(name);
assertEquals("rw-r-----", java.nio.file.attribute.PosixFilePermissions.toString(
Files.getPosixFilePermissions(file)),
name + " must be group-readable, never group-writable");
}
}
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225: that reads wherever Maven happened to be started from,
* not the process's own group, and the two diverge outside a home checkout). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
String raw = new String(p.getInputStream().readAllBytes(), StandardCharsets.UTF_8).trim();
out = raw;
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !raw.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/**
* The same equality, for a shell that is INTERACTIVE but NOT a login shell — the shape herdr
* opens on Linux.
@@ -18,7 +18,6 @@ import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFileAttributeView;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
@@ -513,11 +512,37 @@ class HerdrPeerLauncherAllowListWiringTest {
+ "java.io.tmpdir, unchanged from before this fix: " + dir);
}
/** The current process's own primary group — resolvable on whatever host runs this test. */
private static String currentUserGroup() throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(Path.of("."), PosixFileAttributeView.class);
assumeTrue(view != null, "this host's filesystem does not support POSIX group ownership");
return view.readAttributes().group().getName();
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225). Those two coincide only by accident: a directory's
* group is whichever group happened to own the path Maven was started from — {@code staff} in
* a home checkout, {@code wheel} under {@code /private/tmp} on macOS — and the fix-up this test
* exercises then fails for real when the operator is not a member of that borrowed group,
* exactly the case {@code assumeTrue(view != null, ...)} never covered (it only detects a
* filesystem with no POSIX groups at all, not a resolvable-but-wrong one). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
try (java.io.BufferedReader r = new java.io.BufferedReader(
new java.io.InputStreamReader(p.getInputStream(), java.nio.charset.StandardCharsets.UTF_8))) {
out = r.lines().collect(java.util.stream.Collectors.joining("\n")).trim();
}
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !out.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/** Spawn once through the real launcher path, capturing every INFO+ line this class logs. */
@@ -10,12 +10,16 @@ import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.ExhaustionSink;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.session.MemberSession;
import dev.ltms.fleet.session.SessionManager;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
@@ -23,13 +27,15 @@ import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFileAttributeView;
import java.nio.file.attribute.PosixFilePermissions;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import java.util.concurrent.TimeUnit;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
@@ -48,6 +54,19 @@ class OpenCodeLauncherTest {
null, null, gitTokenEnv, null, FleetConfig.Profile.KIND_OPENCODE);
}
/**
* A profile carrying an explicit {@code credentialId} (fleetd #234 defect 2) — distinct from
* the profile's own name, so a test can assert on the credential precisely rather than relying
* on {@code effectiveCredentialId()}'s profile-name fallback.
*/
private static FleetConfig.Profile opencodeCfgWithCredential(String profileName, String model,
String credentialId) {
return new FleetConfig.Profile(profileName, null, model, null, "FLEETD_WORKER_TOKEN",
List.of("opencode"), "tab", "fleetd-workers", "opencode: {model} #{n}", null,
null, List.of(), null, null, FleetConfig.Profile.KIND_OPENCODE, Map.of(), 1.0f,
null, false, null, credentialId, null);
}
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
private static OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, FleetConfig.Profile cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
@@ -392,6 +411,46 @@ class OpenCodeLauncherTest {
assertNotNull(ex.getMessage());
}
/**
* fleetd #176 fix 2's per-adapter guard: {@code StatusRefiner.classify} is written for the
* Claude Code TUI only, and the spawn-readiness gate must never run it against an opencode
* pane. {@code agentType("opencode")} deliberately satisfies the OTHER guard (the corroborating
* liveness check) so it cannot be what makes this test pass — only the {@code namePrefix}
* check can be. If that check were ever removed, this pane's {@code ❯} content would refine
* straight to IDLE and the gate would report ready before the backend actually was.
*/
@Test
void opencodePaneIsNeverRefinedEvenWhenItsContentLooksLikeAnIdleClaudePrompt(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
herdr.agentType("opencode"); // non-null — satisfies the liveness guard on its own
herdr.readText("some-host:~ user$ ❯ "); // content StatusRefiner.classify reads as IDLE
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root, root);
assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000,
"an opencode pane must never refine to injectable, however its content reads — "
+ "the gate has to wait out the full timeout: clock only reached " + clock[0]);
// A blanket "agent.read is never called" does not hold here: the timeout path itself reads
// the pane tail (source=recent) for its own log message, on every timeout, regardless of
// adapter — see HerdrPeerLauncher.readPaneQuietly. So assert on the refiner's OWN probe
// source (StatusRefiner.PROBE_SOURCE = "detection") instead — that call happens only inside
// StatusRefiner.refine, so its absence proves the refiner itself was never reached for this
// opencode pane, not merely that its answer was discarded.
long detectionReads = herdr.calls.stream()
.filter(c -> c.method().equals("agent.read"))
.filter(c -> "detection".equals(((Map<?, ?>) c.params()).get("source")))
.count();
assertEquals(0, detectionReads,
"the refiner's own pane probe (source=detection) must never run against an "
+ "opencode pane — the namePrefix guard has to stop it before that call");
}
@Test
void spawnReturnsHandleWhenGateDisabled(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
@@ -657,11 +716,37 @@ class OpenCodeLauncherTest {
0, System::currentTimeMillis, () -> { }, configRoot, discoveryRoot, null, null, config);
}
/** The current process's own primary group — resolvable on whatever host runs this test. */
private static String currentUserGroup() throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(Path.of("."), PosixFileAttributeView.class);
assumeTrue(view != null, "this host's filesystem does not support POSIX group ownership");
return view.readAttributes().group().getName();
/**
* The current process's REAL primary group — resolved via {@code id -gn}, never by reading a
* directory's owning group (fleetd #225). Those two coincide only by accident: a directory's
* group is whichever group happened to own the path Maven was started from — {@code staff} in
* a home checkout, {@code wheel} under {@code /private/tmp} on macOS — and the fix-up this test
* exercises then fails for real when the operator is not a member of that borrowed group,
* exactly the case {@code assumeTrue(view != null, ...)} never covered (it only detects a
* filesystem with no POSIX groups at all, not a resolvable-but-wrong one). Skips (never fails)
* when {@code id} is unavailable or its primary group cannot be resolved on this host.
*/
private static String currentUserGroup() {
String out;
boolean ok;
try {
Process p = new ProcessBuilder("id", "-gn").redirectErrorStream(true).start();
try (java.io.BufferedReader r = new java.io.BufferedReader(
new java.io.InputStreamReader(p.getInputStream(), java.nio.charset.StandardCharsets.UTF_8))) {
out = r.lines().collect(java.util.stream.Collectors.joining("\n")).trim();
}
ok = p.waitFor(5, java.util.concurrent.TimeUnit.SECONDS) && p.exitValue() == 0 && !out.isBlank();
} catch (IOException e) {
out = null;
ok = false;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
out = null;
ok = false;
}
assumeTrue(ok, "cannot resolve this process's real primary group via `id -gn` on this host "
+ "— skipping a POSIX-group-dependent test rather than failing it");
return out;
}
/**
@@ -811,4 +896,452 @@ class OpenCodeLauncherTest {
assertTrue(warnings.get(0).contains("memberHerdrSocket"),
"the WARN must name memberHerdrSocket as the reason — got: " + warnings.get(0));
}
// --- fleetd #175: opencode silently substitutes a model on an unknown -m flag ----------------
private static OpenCodeLauncher serviceWithSink(FakeHerdr herdr, Path configRoot, Path discoveryRoot,
FleetConfig.Profile cfg, ExhaustionSink sink) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, configRoot, discoveryRoot,
null, null, null, sink);
}
/**
* fleetd #175 acceptance criterion 1: the check must fire on the REAL late-resolve path —
* {@code SessionManager}'s {@code get}/{@code roster} re-polling a retained {@link PeerHandle}
* (fleetd #209) — not on a handle built and queried directly. A handle whose row does not exist
* yet, then does, proves the check runs exactly where production runs it: PR #203 shipped a
* check that ran inside {@code spawn()}, before opencode had written the row, and every test
* passed anyway because none of them drove it through this path. This test would have caught
* that: {@code sessions.get(...)} before the row exists must show no mismatch, and the SAME
* call, re-driven after the row appears, must be what fires the sink.
*/
@Test
void theRealSessionManagerLateResolvePathCatchesAModelMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
FakeHerdr herdr = new FakeHerdr();
// xf's real shape (fleetd #175): weight:80, model "opencode/nemotron-3-ultra-free", no
// credentialId — the profile that actually escaped the fleet's accounting.
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(target + "|" + reason);
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, sink);
SessionManager sessions = new SessionManager(launcher);
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
// Real late-resolve path, driven BEFORE opencode has written its session row — same shape
// as production the instant a pane goes ready.
Optional<MemberSession> beforeRow = sessions.get(acquired.paneId());
assertTrue(beforeRow.isPresent());
assertNull(beforeRow.get().agentSessionId(), "no opencode row yet");
assertTrue(exhausted.isEmpty(), "no row yet → nothing to compare, the sink must stay silent");
// opencode writes its row late, running gpt-5.6-sol (a PAID credential) instead of the
// withdrawn free model the profile actually asked for — the exact fleetd #175 scenario.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
// Drive the SAME real late-resolve path again: sessions.get() -> resolveAgentSessionId ->
// the retained PeerHandle's agentSessionId() -> discovery -> checkModelMatch, all in one
// call, unmodified SessionManager code from fleetd #209.
Optional<MemberSession> afterRow = sessions.get(acquired.paneId());
assertEquals("ses_x", afterRow.get().agentSessionId(),
"the id itself still resolves correctly alongside the model check");
assertEquals(1, exhausted.size(), "the mismatch must fire exactly once through the real path");
assertTrue(exhausted.get(0).contains(cfg.model()), "reports the requested model: " + exhausted.get(0));
assertTrue(exhausted.get(0).contains("gpt-5.6-sol"), "reports the actual model: " + exhausted.get(0));
}
/** THE TRAP, row 1: a provider-prefixed request matching the DB's id AND provider is a match. */
@Test
void aProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "id and provider both match → never a mismatch: " + exhausted);
}
/** THE TRAP, row 2: same shape as row 1 with a different provider/id pair (the gx profile). */
@Test
void aGxProviderPrefixedModelMatchingBothIdAndProviderIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("gx/deepseek-v4-flash", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "id and provider both match → never a mismatch: " + exhausted);
}
/**
* Incomplete evidence, not THE TRAP's provider mismatch: opencode's {@code model} JSON had an
* {@code id} but no {@code providerID} at all (a real shape {@code parseModel} accepts — see
* {@code OpenCodeSessionDiscoveryTest}). The id matches; the provider dimension is simply
* unknown, not contradicted. A provider-prefixed profile must NOT be quarantined on this —
* that would quarantine on incomplete evidence, which acceptance rule 4 forbids.
*/
@Test
void aMissingProviderIdInTheEvidenceIsUnknownNotAMismatchWhenTheIdMatches(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(),
"id matches and provider is simply unknown (absent), never a mismatch: " + exhausted);
}
/**
* The other half of the same fix: a missing {@code providerID} must NOT blind the check to a
* genuine id mismatch. This is what proves the fix narrows the comparison rather than switching
* the whole check off whenever {@code providerID} happens to be absent.
*/
@Test
void aMissingProviderIdInTheEvidenceStillCatchesARealIdMismatch(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\"}");
assertEquals("ses_x", handle.agentSessionId());
assertEquals(1, exhausted.size(),
"the id genuinely differs, so this must still quarantine even with providerID absent: "
+ exhausted);
assertTrue(exhausted.get(0).contains("gpt-5.6-sol") && exhausted.get(0).contains(cfg.model()),
"reports both the requested and actual model: " + exhausted.get(0));
}
/**
* THE TRAP, row 3: a profile that names no provider prefix (bare {@code "deepseek-v4-flash"})
* must match on id alone — the profile never asked for a specific provider, so opencode
* resolving it to {@code gx} is not evidence of anything wrong.
*/
@Test
void aBareModelWithNoProviderPrefixMatchesOnIdAloneAndIsNotAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("deepseek-v4-flash", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(),
"no provider was requested, so opencode's own provider resolution is not a mismatch: "
+ exhausted);
}
/**
* THE TRAP, row 4 (not a row in the table, the reason the table exists): a mismatched id, with
* the profile's requested and the actual model both named in the ERROR and the sink's reason —
* the exact fleetd #175 scenario (a withdrawn model silently falls back to a paid one).
*/
@Test
void aRealIdMismatchLogsAnErrorNamingBothModelsAndQuarantinesThroughTheSink(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(target + "|" + reason);
FleetConfig.Profile cfg = opencodeCfg("opencode/nemotron-3-ultra-free", null, null);
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
PeerHandle handle;
try {
handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertEquals("ses_x", handle.agentSessionId());
} finally {
logger.detachAppender(appender);
}
assertEquals(1, exhausted.size(), "a real id mismatch must reach the sink exactly once");
assertTrue(exhausted.get(0).contains("opencode/nemotron-3-ultra-free"),
"the sink reason must name the requested model: " + exhausted.get(0));
assertTrue(exhausted.get(0).contains("gpt-5.6-sol"),
"the sink reason must name the actual model: " + exhausted.get(0));
List<ILoggingEvent> errors = appender.list.stream()
.filter(e -> e.getLevel() == Level.ERROR)
.toList();
assertEquals(1, errors.size(), "exactly one ERROR for the mismatch: " + appender.list);
String message = errors.get(0).getFormattedMessage();
assertTrue(message.contains("opencode/nemotron-3-ultra-free") && message.contains("gpt-5.6-sol")
&& message.contains(cfg.profile()),
"the ERROR must name the requested model, the actual model, AND the profile: " + message);
// The check runs at most once per handle even if agentSessionId() is polled again.
handle.agentSessionId();
assertEquals(1, exhausted.size(), "no duplicate quarantine on a repeated call");
}
/**
* fleetd #175's central safety rule: absent, empty, or unparseable model evidence is UNKNOWN,
* never a mismatch — it must never quarantine a working profile. Covers every "no real
* evidence yet" shape: no row at all, a row with a null model, and a row with unparseable JSON.
*/
@Test
void unknownOrUnparseableModelEvidenceNeverQuarantines(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
// No row yet at all.
assertNull(handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "no row yet is UNKNOWN, not a mismatch: " + exhausted);
// A row exists (for a DIFFERENT directory) so a poll on ours still finds nothing.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_other", "/work/other", 500L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertNull(handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "a row for a different directory is UNKNOWN, not a mismatch: "
+ exhausted);
}
/**
* A profile with no {@code model:} configured has nothing to compare against — the check must
* stay silent no matter what opencode actually ran, since there is no requested value to be
* wrong about.
*/
@Test
void aProfileWithNoConfiguredModelIsNeverCheckedForAMismatch(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg(null, null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"anything-at-all\",\"providerID\":\"anyone\"}");
assertEquals("ses_x", handle.agentSessionId());
assertTrue(exhausted.isEmpty(), "no model configured → nothing to compare: " + exhausted);
}
// --- fleetd #234 defect 1: key the model check on the RESOLVED id, not the shared directory --
/**
* The exact shape fleetd #234 reported: {@code fleet_spawn} with no {@code worktree:} shares
* the lead's cwd across every worker, so more than one session row can exist for the SAME
* {@code directory}. Once THIS handle's own session id is resolved, a sibling member spawned
* later into the same shared directory — writing a NEWER, unrelated row — must never make the
* already-resolved session look mismatched. Today's code re-derives "the newest row in this
* directory" on every call (both for the id AND, independently, for the model), so it would
* pick up the sibling's row on the second call and flag a false mismatch AND flip the returned
* id. The fix (fleetd #234) makes the id sticky once resolved and reads the model back for
* exactly that id (see {@link OpenCodeSessionDiscovery#actualModelForSessionId}) — never
* "whatever is newest in the directory right now."
*/
@Test
void modelCheckReadsTheResolvedSessionsOwnRowNotWhateverIsNewestInTheSharedDirectory(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
List<String> exhausted = new ArrayList<>();
ExhaustionSink sink = (target, reason, profile) -> exhausted.add(reason);
FleetConfig.Profile cfg = opencodeCfg("openai/gpt-5.6-terra", null, null);
PeerHandle handle = serviceWithSink(new FakeHerdr(), configRoot, discRoot, cfg, sink)
.spawn(new SpawnRequest(null, "/work/dir", null));
// Our own session's row, correctly matching the profile's requested model.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_ours", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
assertEquals("ses_ours", handle.agentSessionId(), "resolves to our own session");
assertTrue(exhausted.isEmpty(), "matching model → no mismatch on first resolve: " + exhausted);
// A sibling member, spawned later into the SAME shared directory (no worktree, fleetd
// #234's default), writes a newer row running a totally different model.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_sibling", "/work/dir", 9000L,
"{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
assertEquals("ses_ours", handle.agentSessionId(),
"the session id, once resolved, must not flip to a sibling sharing the directory");
assertTrue(exhausted.isEmpty(),
"a sibling's later, unrelated row in the same shared directory must never be read "
+ "as OUR session's model: " + exhausted);
}
// --- fleetd #234 defect 2: the quarantine the ERROR announces must actually happen -----------
/**
* fleetd #234: {@link OpenCodeLauncher.SessionAwareHandle#checkModelMatch} fires from {@code
* agentSessionId()}, which {@code SessionManager.acquire()} calls to build the very first
* {@code MemberSession} record — BEFORE that session is put into the registry {@code
* sessions.roster()} reads. A sink that resolves {@code target -> profile} ONLY through the
* roster (today's {@code Fleetd.java} code, before this fix) therefore finds nothing at this
* exact moment and silently does not quarantine, even though it just logged an ERROR saying it
* would. This test drives the REAL path — {@code SessionManager.acquire()} — not the sink
* directly, because the bug is entirely about this ordering; a direct-sink test cannot see it
* (and is exactly why #175's own test suite, which only ever called the sink directly or after
* registration, never caught this).
*
* <p>The sink under test mirrors {@code Fleetd.main()}'s real wiring after the fix: resolve
* via the roster first (unchanged for {@code CompletionResolver}'s two call sites), falling
* back to the profile hint {@link OpenCodeLauncher} now supplies via {@link
* ExhaustionSink#onExhausted(String, String, String)} when the roster lookup misses.
*/
@Test
void aSpawnTimeModelMismatchActuallyQuarantinesTheCredentialThroughTheRealAcquirePath(
@TempDir Path configRoot, @TempDir Path discRoot) throws Exception {
FleetConfig.Profile cfg = opencodeCfgWithCredential(
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
Map<String, FleetConfig.Profile> profiles = Map.of(cfg.profile(), cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
// fleetd #234, round 4: safely a lambda now — the 3-arg overload is the interface's single
// abstract method, so there is no 2-arg overload left to bind to instead.
ExhaustionSink sink = (target, reason, profileHint) -> {
FleetConfig.Profile profile = profileHint == null ? null : profiles.get(profileHint);
if (profile != null) {
quarantine.quarantine(profile.effectiveCredentialId());
}
};
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, sink);
SessionManager sessions = new SessionManager(launcher);
// The mismatching row exists BEFORE the spawn — reproducing fleetd #234's exact timing:
// opencode's session table already carries evidence by the moment acquire() first asks.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertFalse(quarantine.isQuarantined("openai-shared"), "nothing quarantined before the spawn");
// The real production entrypoint: acquire() builds the MemberSession by calling
// handle.agentSessionId() BEFORE registry.put() runs.
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
assertTrue(quarantine.isQuarantined("openai-shared"),
"the mismatch fires DURING acquire(), before roster registration, and must still "
+ "reach the quarantine via the profile hint — not silently no-op");
}
/**
* The other half of the same proof: a sink that resolves {@code target -> profile} ONLY
* through the roster (i.e. ignores the profile hint entirely, ~today's pre-fix {@code
* Fleetd.java}) drops the SAME spawn-time mismatch silently — the credential is never
* quarantined even though {@link OpenCodeLauncher} logged the mismatch ERROR. This is the
* failure fleetd #234 reported, reproduced through the real {@code SessionManager.acquire()}
* path rather than asserted by inspecting the fix.
*/
@Test
void aRosterOnlySinkSilentlyDropsTheSpawnTimeQuarantine(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
FleetConfig.Profile cfg = opencodeCfgWithCredential(
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
// Deliberately ignores the profile hint — the pre-fix shape: only a roster lookup (modelled
// here as always empty, since acquire() has not registered the session yet either way).
ExhaustionSink rosterOnlySink = (target, reason, profile) -> { };
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, rosterOnlySink);
SessionManager sessions = new SessionManager(launcher);
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
assertFalse(quarantine.isQuarantined("openai-shared"),
"a roster-only sink cannot see this target yet — the quarantine silently never "
+ "happens, which is exactly fleetd #234 defect 2");
}
/**
* fleetd #234, round 2: the two tests above inject a sink DIRECTLY into {@link OpenCodeLauncher},
* which is not what {@code Fleetd.main} actually does. Production has one more hop: {@code
* Fleetd.java} builds the adapters (including {@code OpenCodeLauncher}) before {@code sessions}
* exists — a genuine construction-order cycle — so it hands the launcher a <em>forwarding</em>
* sink pointed at an {@code AtomicReference<ExhaustionSink>}, and only later {@code .set(...)}s
* that reference to the real sink once {@code sessions} is built.
*
* <p><strong>Round 3:</strong> a round-2 version of this test built its own copy of that
* forwarder's shape as an anonymous class. Mutating {@code Fleetd.java}'s ACTUAL forwarder back
* into a broken lambda left this test's own copy untouched, so it kept passing while production
* had regressed to the exact bug being fixed. This version instead calls {@link
* ExhaustionSink#forwardingTo}, the same factory {@code Fleetd.java} calls — the identical
* object, not a rebuilt copy of its shape — so a regression at either the {@code Fleetd.java}
* call site or inside {@code forwardingTo} itself has nowhere left to hide. See {@code
* ExhaustionSinkForwardingHazardTest} for the same factory exercised in isolation.
*/
@Test
void theSpawnTimeQuarantineSurvivesTheFleetdStyleForwardingHop(@TempDir Path configRoot,
@TempDir Path discRoot) throws Exception {
FleetConfig.Profile cfg = opencodeCfgWithCredential(
"terra", "opencode/nemotron-3-ultra-free", "openai-shared");
Map<String, FleetConfig.Profile> profiles = Map.of(cfg.profile(), cfg);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.SECONDS.toNanos(1800));
// Fleetd.java:176 — the forwarding sink is built BEFORE the real one can exist, and the
// launcher below is constructed against this forwarder, exactly like Fleetd.main. Calling
// the SAME factory Fleetd.java calls — ExhaustionSink.forwardingTo — rather than rebuilding
// the forwarder's shape here is the whole point (fleetd #234, round 3): a round-2 version of
// this test built its own copy, so mutating Fleetd.java's real forwarder back into a lambda
// left this test untouched. Routed through the shared factory, a regression at either the
// Fleetd.java call site (reverting to a hand-written lambda) or inside the factory body
// itself has nowhere left to hide from this test.
java.util.concurrent.atomic.AtomicReference<ExhaustionSink> exhaustionSinkRef =
new java.util.concurrent.atomic.AtomicReference<>(ExhaustionSink.none());
ExhaustionSink forwardingExhaustionSink = ExhaustionSink.forwardingTo(exhaustionSinkRef::get);
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithSink(herdr, configRoot, discRoot, cfg, forwardingExhaustionSink);
SessionManager sessions = new SessionManager(launcher);
// Fleetd.java:359-390 — the real sink is only built and pointed to AFTER `sessions` exists,
// same order as production. Safely a lambda (round 4): see the note on the forwarder above.
ExhaustionSink realSink = (target, reason, profileHint) -> {
FleetConfig.Profile profile = profileHint == null ? null : profiles.get(profileHint);
if (profile != null) {
quarantine.quarantine(profile.effectiveCredentialId());
}
};
exhaustionSinkRef.set(realSink);
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_x", "/work/dir", 1000L,
"{\"id\":\"gpt-5.6-sol\",\"providerID\":\"openai\"}");
assertFalse(quarantine.isQuarantined("openai-shared"), "nothing quarantined before the spawn");
MemberSession acquired = sessions.acquire(cfg.profile(), "/work/dir", null, null);
assertEquals("ses_x", acquired.agentSessionId(), "the id itself still resolves correctly");
assertTrue(quarantine.isQuarantined("openai-shared"),
"the profile hint must survive the Fleetd-style forwarding hop and reach the real "
+ "sink — a lambda forwarder drops it and this must go red");
}
}
@@ -27,19 +27,30 @@ class OpenCodeSessionDiscoveryTest {
* and insert one row. Static so {@link OpenCodeLauncherTest} can reuse it.
*/
static void writeRecord(Path root, String id, String directory, long timeUpdated) throws Exception {
writeRecord(root, id, directory, timeUpdated, null);
}
/**
* Same as {@link #writeRecord(Path, String, String, long)}, plus the {@code model} column
* fleetd #175 reads — the raw JSON opencode writes, e.g.
* {@code {"id":"gpt-5.6-terra","providerID":"openai"}}. {@code modelJson} may be {@code null}.
*/
static void writeRecord(Path root, String id, String directory, long timeUpdated, String modelJson)
throws Exception {
Path db = root.resolve("opencode.db");
try (Connection connection = DriverManager.getConnection("jdbc:sqlite:" + db)) {
try (Statement statement = connection.createStatement()) {
statement.execute("CREATE TABLE IF NOT EXISTS session ("
+ "id TEXT PRIMARY KEY, directory TEXT, time_updated INTEGER)");
+ "id TEXT PRIMARY KEY, directory TEXT, time_updated INTEGER, model TEXT)");
}
// Bound parameters, not string interpolation: the class under test uses a
// PreparedStatement, and a hand-escaped INSERT here is a pattern someone copies out.
try (PreparedStatement insert = connection.prepareStatement(
"INSERT INTO session (id, directory, time_updated) VALUES (?, ?, ?)")) {
"INSERT INTO session (id, directory, time_updated, model) VALUES (?, ?, ?, ?)")) {
insert.setString(1, id);
insert.setString(2, directory);
insert.setLong(3, timeUpdated);
insert.setString(4, modelJson);
insert.executeUpdate();
}
}
@@ -134,4 +145,107 @@ class OpenCodeSessionDiscoveryTest {
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"an unreadable database resolves to null, not an exception");
}
// --- fleetd #175/#234: actualModelForSessionId / the model JSON column -----------------------
@Test
void parsesTheModelJsonIntoProviderAndId(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
OpenCodeSessionDiscovery.ActualModel actual =
new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa");
assertNotNull(actual, "a well-formed model JSON parses");
assertEquals("gpt-5.6-terra", actual.id());
assertEquals("openai", actual.provider());
}
@Test
void looksUpByIdEvenWhenAnotherRowInTheSameDirectoryIsNewer(@TempDir Path root) throws Exception {
writeRecord(root, "ses_ours", "/work/dir", 1000L, "{\"id\":\"gpt-5.6-terra\",\"providerID\":\"openai\"}");
writeRecord(root, "ses_sibling", "/work/dir", 9000L, "{\"id\":\"deepseek-v4-flash\",\"providerID\":\"gx\"}");
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
assertEquals("gpt-5.6-terra", discovery.actualModelForSessionId("ses_ours").id(),
"querying by id reads OUR row, not the directory's newest row");
assertEquals("deepseek-v4-flash", discovery.actualModelForSessionId("ses_sibling").id(),
"each id resolves to its own row independently of time_updated ordering");
}
@Test
void anUnknownSessionIdYieldsUnknownModelRatherThanAMismatch(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"x\",\"providerID\":\"y\"}");
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_no_such_row"),
"no row for this id yet → unknown, not a wrong model");
}
@Test
void aBlankOrNullSessionIdYieldsUnknownModel(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"id\":\"x\",\"providerID\":\"y\"}");
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
assertNull(discovery.actualModelForSessionId(null));
assertNull(discovery.actualModelForSessionId(" "));
}
@Test
void aNullModelColumnYieldsUnknownWithoutThrowing(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, null);
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"),
"a row with no model value yet is unknown, not a mismatch");
}
@Test
void unparseableModelJsonYieldsUnknownWithoutThrowing(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "this is not json");
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"),
"JSON that fails to parse resolves to unknown, never an exception");
}
@Test
void modelJsonMissingIdYieldsUnknown(@TempDir Path root) throws Exception {
writeRecord(root, "ses_aaa", "/w/a", 1000L, "{\"providerID\":\"openai\"}");
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"),
"no id in the JSON → unknown, since id is what a caller actually compares");
}
@Test
void aMissingDatabaseYieldsUnknownModelWithoutThrowing(@TempDir Path root) {
assertNull(new OpenCodeSessionDiscovery(root).actualModelForSessionId("ses_aaa"));
}
/**
* fleetd #175 site 2: a schema that predates the {@code model} column entirely — the exact
* shape a real opencode upgrade/downgrade could produce. This must degrade to UNKNOWN for the
* model, and — the property that actually matters — must NOT take id resolution down with it.
* A combined single query for both columns would fail this test; that is why
* {@link OpenCodeSessionDiscovery#actualModelForSessionId} runs its own independent query.
*/
@Test
void aMissingModelColumnYieldsUnknownButIdResolutionStillWorks(@TempDir Path root) throws Exception {
Path db = root.resolve("opencode.db");
try (Connection connection = DriverManager.getConnection("jdbc:sqlite:" + db)) {
try (Statement statement = connection.createStatement()) {
statement.execute("CREATE TABLE session (id TEXT PRIMARY KEY, directory TEXT, "
+ "time_updated INTEGER)");
}
try (PreparedStatement insert = connection.prepareStatement(
"INSERT INTO session (id, directory, time_updated) VALUES (?, ?, ?)")) {
insert.setString(1, "ses_aaa");
insert.setString(2, "/w/a");
insert.setLong(3, 1000L);
insert.executeUpdate();
}
}
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
assertEquals("ses_aaa", discovery.sessionIdForDirectory("/w/a"),
"id resolution must survive a database with no model column at all");
assertNull(discovery.actualModelForSessionId("ses_aaa"),
"no model column → unknown, not a throw and not a mismatch");
}
}
@@ -1099,4 +1099,82 @@ class GitWorktreesTest {
assertTrue(e.getMessage().contains("cb185-nonexistent-group-zz"),
"exception must name the missing/refused group: " + e.getMessage());
}
// --- fleetd #224: worktreeRoot itself must be group-traversable, established in add() ---------
/**
* {@code add} must make the worktree ROOT itself group-traversable when a group is configured —
* established once here, where the root is created, never at a use site (never inside a
* launcher). A recording runner stands in for chgrp/chmod, the same seam
* {@link #shareWithGroupRunsConfigThenChgrpChmodSetgidPerPath} uses for the per-worktree share.
*/
@Test
void addSharesWorktreeRootWithGroupWhenConfigured(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path root = tmp.resolve("wts");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees = new GitWorktrees(root.toString(), "devteam", _ -> {}, recordingRunner);
gitWorktrees.add(repo.toString(), "cb-224-branch", "HEAD");
assertTrue(recorded.contains(List.of("chgrp", "devteam", root.toString())),
"the worktree root itself must be chgrp'd to the configured group: " + recorded);
assertTrue(recorded.contains(List.of("chmod", "g+x", root.toString())),
"the worktree root itself must gain group-execute so a different-uid member can "
+ "traverse into it: " + recorded);
}
/** {@code worktreeGroup} unset (today's only live mode) ⇒ {@code add} spawns no share process
* for the root at all — behaviour must be byte-identical to before fleetd #224. */
@Test
void addSharesNothingForTheRootWhenNoGroupConfigured(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path root = tmp.resolve("wts");
List<List<String>> recorded = new java.util.ArrayList<>();
java.util.function.Function<String[], String> recordingRunner = cmd -> {
recorded.add(joined(cmd));
return "";
};
GitWorktrees gitWorktrees = new GitWorktrees(root.toString(), null, _ -> {}, recordingRunner);
gitWorktrees.add(repo.toString(), "cb-224-nogroup", "HEAD");
assertTrue(recorded.isEmpty(), "no group configured must spawn no share process for the "
+ "root at all: " + recorded);
}
/**
* fleetd #224, acceptance criterion 2/3: when the root cannot be made group-traversable — here
* because the configured group does not exist, the same real-failure shape
* {@link #shareWithGroupThrowsNamingTheGroupWhenChgrpFails} drives for the per-worktree share —
* the spawn is refused with a message naming the root, its current mode, and the group. This
* drives the REAL {@code chgrp} (no recording runner), and the refusal happens before {@code git
* worktree add} ever runs, so no partial worktree is left behind either.
*/
@Test
void addRefusesWhenWorktreeRootCannotBeMadeGroupTraversable(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
Path root = tmp.resolve("wts");
GitWorktrees gitWorktrees = new GitWorktrees(root.toString(), "cb224-nonexistent-group-zz");
WorktreeException e = assertThrows(WorktreeException.class,
() -> gitWorktrees.add(repo.toString(), "cb-224-refuse", "HEAD"));
assertTrue(e.getMessage().contains(root.toString()),
"refusal must name the worktree root: " + e.getMessage());
assertTrue(e.getMessage().contains("cb224-nonexistent-group-zz"),
"refusal must name the missing/refused group: " + e.getMessage());
assertTrue(Files.isDirectory(root), "the root is created before the group check runs");
String mode = java.nio.file.attribute.PosixFilePermissions.toString(Files.getPosixFilePermissions(root));
assertTrue(e.getMessage().contains(mode),
"refusal must name the root's current mode (" + mode + "): " + e.getMessage());
try (java.util.stream.Stream<Path> children = Files.list(root)) {
assertTrue(children.findAny().isEmpty(), "no worktree must be left behind under the root: "
+ "the refusal must happen before `git worktree add` ever runs");
}
}
}
@@ -4,6 +4,8 @@ import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.MemberLifecycle;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -333,6 +335,110 @@ class SessionManagerTest {
}
}
/**
* fleetd #226: a contended slot is refused through the real {@link SessionManager#acquire}
* path before the real launcher can hand an architect charter to a process.
*/
@Test
void aContendedArchitectSlotRefusesBeforeTheLauncherStartsAMember() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberRegistry members = architectRegistry();
assertTrue(members.bind("architect:opus", "term_already_bound"),
"precondition: occupy the sole architect slot before the real spawn under test");
sessions.setMemberLifecycle(members);
assertThrows(IllegalArgumentException.class, () -> sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller", "term_primary", null));
assertFalse(herdr.called("agent.start"), "a refused slot must never reach the launcher");
assertTrue(sessions.roster().isEmpty(), "no member exists to receive the wrong charter");
}
@Test
void aSecondArchitectOnAnAlreadyBoundProfileIsHeldAsDevNotArchitectAndWarnsLoudly() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberRegistry members = architectRegistry();
sessions.setMemberLifecycle(bindFailureAfterReservation(members));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger registryLog = (ch.qos.logback.classic.Logger)
LoggerFactory.getLogger(MemberRegistry.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
registryLog.addAppender(appender);
registryLog.setLevel(Level.WARN);
try {
MemberSession session = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.DEV, session.role(), "a failed reservation bind must use the fallback");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no slot-exhaustion WARN logged");
assertTrue(warn.contains("ltms-local"), "the WARN names the profile: " + warn);
assertTrue(warn.contains(session.terminalId()), "the WARN names the terminal: " + warn);
} finally {
registryLog.detachAppender(appender);
}
}
@Test
void reservedArchitectSlotBindsTheLaunchedMember() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
sessions.setMemberLifecycle(architectRegistry());
MemberSession session = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.ARCHITECT, session.role());
}
private static MemberRegistry architectRegistry() {
return new MemberRegistry(new FleetConfig.Fleet(Map.of(), Map.of("opus", new FleetConfig.Slot("ltms-local")),
Map.of(), Map.of(), null));
}
private static MemberLifecycle bindFailureAfterReservation(MemberRegistry members) {
return new MemberLifecycle() {
@Override
public MemberRole acquired(MemberRole role, String profile, String terminal) {
return members.acquired(role, profile, terminal);
}
@Override
public void released(String terminal) {
members.released(terminal);
}
@Override
public void requireSlotFor(MemberRole role, String profile) {
members.requireSlotFor(role, profile);
}
@Override
public SlotReservation reserve(MemberRole role, String profile) {
return members.reserve(role, profile);
}
@Override
public boolean bind(SlotReservation reservation, String terminal) {
members.release(reservation);
assertTrue(members.bind(reservation.slot(), "term_racer"), "the racer takes the released slot");
return false;
}
@Override
public void release(SlotReservation reservation) {
members.release(reservation);
}
};
}
@Test
void rosterReflectsAcquiredMinusReleased() {
FakeHerdr herdr = new FakeHerdr();
@@ -599,6 +705,32 @@ class SessionManagerTest {
"no session is registered when spawn times out (roster empty)");
}
@Test
void failedArchitectLaunchReleasesItsReservationForTheNextLaunch() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown");
long[] clock = {0};
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
1, () -> clock[0], () -> clock[0] += 10);
SessionManager sessions = new SessionManager(workers, new GitWorktrees(), () -> 0L, 0);
sessions.setMemberLifecycle(architectRegistry());
assertThrows(PeerUnreachableException.class, () -> sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller", "term_primary", null));
herdr.agentStatus("idle");
MemberSession retry = sessions.acquire("ltms-local", MemberRole.ARCHITECT, null,
"/caller", "term_primary", null);
assertEquals(MemberRole.ARCHITECT, retry.role(),
"the failed launch returned its reservation instead of silently shrinking the slot pool");
}
// --- CB-516: release must notify, so a blocked send can be failed --------------------------
@Test
@@ -95,10 +95,9 @@ class WorktreeSessionManagerTest {
null, "/caller/proj", null, null);
assertEquals("architect:architect", members.slotForTerminal(replacement.terminalId()));
MemberSession overflow = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
assertNull(members.slotForTerminal(overflow.terminalId()),
"a full slot pool must not stop the architect spawn");
assertThrows(IllegalArgumentException.class, () -> sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null),
"a full slot pool must stop the spawn before it can receive an architect charter");
}
@Test