Compare commits

..

8 Commits

Author SHA1 Message Date
Dai Ha d9168de43e fleetd#219: OpenCodeLauncher config/discovery roots must not assume fleetd's own filesystem
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m48s
Site 1 (config root): under memberHerdrSocket, writeConfig() now places the
ephemeral opencode.json directory under worktreeRoot and shares it read-only
with worktreeGroup, reusing EnvAllowListScrub#shareWithGroup (widened to
package-private and generalized) — the same mechanism #213 built for the
ZDOTDIR scrub, rather than a second copy. Unlike the ZDOTDIR scrub's
degrade-to-overlay fallback, a missing worktreeRoot/worktreeGroup here
REFUSES the spawn (IllegalStateException from buildLaunch): this file is the
member's only way to learn where the bridge MCP is, so writing it somewhere
unreadable would just produce an undeliverable member with no signal
pointing at the cause. memberHerdrSocket absent stays byte-identical.

Site 2 (discovery root): under memberHerdrSocket, agentSessionId() now
declares session discovery unavailable and logs one WARN per launcher
instance instead of silently scanning fleetd's own $HOME (opencode.db lives
under the MEMBER's home under this config key). Decision + reasoning for why
this is a declare-unavailable rather than a new config key is in
defaultDiscoveryRoot()'s javadoc.

Widened HerdrPeerLauncher#memberHerdrSocketConfigured/memberScrubParentDir/
memberGroup to package-private so OpenCodeLauncher reuses the exact same
config resolution rather than re-deriving it.

Same-shape finding (not fixed, out of scope): ClaudeCodeLauncher#writeCharterFile
(line ~465) writes the role-charter temp file via Files.createTempFile with no
directory argument, i.e. under java.io.tmpdir — the same site-1 shape, unfixed
for the Claude Code adapter.
2026-09-01 14:38:48 +07:00
Dai Ha c3fa1136d4 #220: keep the launch command inside the pane's 1024-byte line
CI / build (push) Successful in 1m37s
CI / contract (push) Successful in 1m37s
herdr does not exec a member's launch command — it TYPES it into the pane,
and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Past that the
tail is dropped and NOTHING reports it: herdr answers "agent started", the
backend exits on the mangled argument it was handed, the pane closes, and the
only symptom is the readiness gate timing out 20 seconds later with no reason.

That is what broke every claude-code spawn after #214. The reply charter rode
inline on --append-system-prompt, so the command was already 978 bytes; adding
--session-id <uuid> made it 1028, and the 4 bytes cut off the end turned
--autocompact 250000 into --autocompact 25, which claude rejects. Measured on
the live pane, the cut is at byte 1024 exactly.

- ClaudeCodeLauncher: the charter ALWAYS travels as --append-system-prompt-file.
  The file path already existed for the two-charter case; the inline form only
  ever saved a temp file, and it cost ~800 bytes of the line budget. This takes
  the prose off the command line for good.
- HerdrPeerLauncher.checkPaneCommandFits: refuse a command that cannot fit,
  naming the byte count and the longest argument, instead of spawning something
  that cannot work. The estimate is deliberately conservative — fleetd cannot
  see herdr's quoting, and an under-estimate would let the silent truncation
  back in.
- HerdrPeerLauncher.waitUntilInjectableOrThrow: log the pane tail and the last
  herdr status BEFORE stop() closes the pane. Without it the gate reports only
  that it timed out, which is true of every cause. This is what found the bug,
  and it stays.

The guard also catches a case that was already over the limit: a profile with
ideMcpUrl set assembles 1084 bytes. It is now impossible to ship that silently.

3 tests, all watched failing first: with the inline charter restored the guard
fires in the new fit test, in the pre-existing autocompact test and in the IDE
mount test. Full suite 1081 tests green. Proven live: sonnet spawns again, the
member obeys the file-delivered charter and ends its turn with fleet_reply, and
fleet_list reports the #214 agentSessionId.
2026-09-01 14:11:28 +07:00
Dai Ha cabcd87b66 #211 follow-up: the raw-scrape fallback must carry the pane, not just the matched line
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m50s
The normal backend-error path appends the pane tail to the failure reason on
purpose (fleetd#164): the BACKEND_ERROR pattern is a heuristic, and a member
that reported *about* an error while forgetting fleet_reply matches it too, so
dropping the rest of the pane destroys the report.

The new raw-scrape fallback did not do that. It matters more there, not less:
the fallback only runs when the trimmed assistant block was empty, so the raw
scrape is the ONLY copy of whatever the member managed to say. A lead read the
matched line and nothing else.

Clipped to the same cap the normal path uses, since a raw screen has no
boundary trimming to bound its size.

The assertion was watched failing without the fix:
  AssertionFailedError: fleetd#164: the failure must carry the pane, not only
  the matched line ... expected: <true> but was: <false>

mvn clean install: Tests run: 1079, Failures: 0, Errors: 0, Skipped: 0
2026-09-01 13:26:05 +07:00
Dai Ha bf0e09b1a2 #214: mint a session id on every claude-code spawn, so every member is resumable 2026-09-01 13:23:32 +07:00
Dai Ha e1eb50ce65 #213: fix ZDOTDIR credential scrub gate + directory under memberHerdrSocket 2026-09-01 13:23:32 +07:00
Dai Ha 049e7d9d54 fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback 2026-09-01 13:23:32 +07:00
Dai Ha ff3b49cd1e #214: mint a session id on every claude-code spawn, so every member is resumable
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m52s
2026-09-01 11:27:10 +07:00
Dai Ha 445a45f6e1 fleetd#211: classify BACKEND_EXHAUSTED/BACKEND_ERROR from the raw scrape as a fallback
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m17s
CompletionResolver.resolve() returned an empty-scrape failure before the
BACKEND_EXHAUSTED / BACKEND_ERROR classification ever ran, whenever
lastAssistantBlock() found no usable text — most commonly a pane with no ⏺
marker at all, whose boundary scan starts at the top of the raw screen and
breaks immediately on the first line of TUI chrome. Since BACKEND_EXHAUSTED
is the only caller of exhaustionSink, this meant an exhausted backend was
recorded as "produced nothing" instead of being quarantined.

Fix: run the same two classifications against the raw (untrimmed) scrape as
a fallback, only inside the empty-scrape failure branch. A pane that already
yields a usable assistant block never reaches this branch, so the existing
narrow match is unchanged. lastAssistantBlock stays the sole source of the
reply text; only classification ever consults the raw scrape.
2026-09-01 10:50:10 +07:00
9 changed files with 751 additions and 57 deletions
@@ -237,11 +237,13 @@ public final class CompletionResolver implements TurnListener {
}
String tail;
String assistantBlock = null;
String rawScrape = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
rawScrape = agents.read(target, SCRAPE_SOURCE);
assistantBlock = lastAssistantBlock(rawScrape);
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
@@ -256,6 +258,20 @@ public final class CompletionResolver implements TurnListener {
// member, so a caller (including a lead deciding whether to delegate again) can tell a lost
// turn from a real empty answer.
if (scrapeFailed || tail.isEmpty()) {
// fleetd#211: lastAssistantBlock() found nothing usable — most often a pane with no ⏺
// marker at all, whose boundary scan then starts at the top of the raw screen and breaks
// immediately on the first line of TUI chrome (╭, │, ❯, …). Before giving up as a lost
// turn, run the same exhaustion/backend-error classification against the RAW scrape as a
// fallback, ONLY here. A pane that already yielded a usable block never reaches this
// branch, so the narrow (trimmed) match on the normal path below is completely unchanged
// — zero new false positives there. Every pane this fallback examines was already headed
// for the empty-scrape failure, so a wrong label here is strictly less bad than silently
// losing an exhaustion signal: the alternative outcome is already a failure, just one that
// never quarantines the credential. lastAssistantBlock stays the source of the reply
// TEXT everywhere else; only classification ever consults the raw scrape, and only here.
if (rawScrape != null && classifyRawScrapeFallback(target, turn, waiter, rawScrape)) {
return;
}
fail(target, turn, emptyScrapeReason(target, scrapeFailed));
return;
}
@@ -313,6 +329,43 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* fleetd#211: the raw-scrape fallback classification, run only when {@link #lastAssistantBlock}
* found nothing usable (see the call site in {@link #resolve}). Mirrors the two classifications
* the normal path already applies to the trimmed assistant block — exhaustion first, then the
* narrow {@link #BACKEND_ERROR} pattern — against {@code raw} instead, and reports whether one of
* them handled the turn (resolved the waiter or failed it) so the caller skips the empty-scrape
* failure. Never runs on the normal (non-empty-block) path, and never touches the reply text.
*/
private boolean classifyRawScrapeFallback(String target, InFlight turn,
CompletableFuture<Rendezvous.Resolution> waiter, String raw) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(raw, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED from the raw scrape (no "
+ "usable assistant block; no fleet_reply): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return true;
}
String backendError = firstMatchingLine(raw, BACKEND_ERROR);
if (backendError != null) {
// Carry the pane, not just the matched line — the same fleetd#164 rule the normal path
// above applies. Here it matters more, not less: the trimmed block was empty, so the raw
// scrape is the ONLY copy of whatever the member managed to say. Clipped to the same cap
// the normal path uses, since a raw screen has no boundary trimming to bound it.
fail(target, turn, "member " + target + " ended on a backend error: " + backendError
+ "\n--- pane tail ---\n" + clip(raw));
return true;
}
return false;
}
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
@@ -278,18 +278,22 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
* session id — the resume handle. A resume passes the prior id via {@code -r} and returns
* that id; every other spawn mints a new UUID, passes it via {@code --session-id}, and
* returns the mint. The bridge's logical name rides along as {@code -n} when present.
*
* <p>fleetd #214: the mint is unconditional. A plain {@code fleet_spawn} passes neither
* sessionName nor resumeSessionId, yet the member must still be resumable, and this id is
* the only resume handle a claude-code member has — unlike opencode, nothing resolves it
* after the launch. Checked against the real binary (claude 2.1.252): the flag is safe on
* every spawn. The binary takes only a valid UUID — it refuses any other value at argument
* parsing ("Invalid session ID. Must be a valid UUID.") — so the {@code UUID.randomUUID()}
* mint is required, not incidental. The only flag interaction the binary documents is with
* {@code -r} (both claim the session id), and the resume branch above never combines the two.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
@@ -320,11 +324,19 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
*
* <p>CB-618: Claude Code refuses to start when BOTH {@code --append-system-prompt} and
* {@code --append-system-prompt-file} are on the command line ("Cannot use both ... Please use
* only one"), so the two charters can never travel on separate flags. When both are present they
* are concatenated into the one file, role charter first and reply charter last — last is where
* the reply rule must sit, because it is the rule that must survive. When only the reply charter
* is present it keeps its proven inline {@code --append-system-prompt} delivery, which is also
* the only form that reaches a member with no repo checkout.
* only one"), so the two charters can never travel on separate flags. They are concatenated
* into the one file, role charter first and reply charter last — last is where the reply rule
* must sit, because it is the rule that must survive.
*
* <p>fleetd #220: a lone reply charter used to ride inline on {@code --append-system-prompt},
* which put ~800 bytes of prose on the command line herdr types into the pane. That line is
* capped at {@value HerdrPeerLauncher#PANE_COMMAND_BYTE_LIMIT} bytes by the pty itself, and
* everything past the cap is dropped with no error from any layer. The charter alone left about
* 50 bytes of headroom, so adding one flag ({@code --session-id}, fleetd #214) truncated the
* LAST argument instead — {@code --autocompact 250000} arrived as {@code --autocompact 25},
* claude rejected it, and every claude-code spawn died as an unexplained readiness timeout.
* The charter now always travels as a file, which takes the prose off the command line for
* good; {@link HerdrPeerLauncher#checkPaneCommandFits} is the backstop for whatever grows next.
*/
private List<String> argvWithFleet(FleetConfig.Profile cfg, LaunchSpec spec) {
String roleCharter = nonBlank(spec.roleCharter());
@@ -341,16 +353,13 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
argv.add("--mcp-config");
argv.add(mcpConfigJson(cfg));
}
// Combine the charters in order role -> reply, dropping any that are absent. When two
// or more survive they must ride one --append-system-prompt-file (CB-618 forbids the inline
// flag and the file flag together). A lone reply charter keeps its proven inline delivery.
// Combine the charters in order role -> reply, dropping any that are absent. They ride one
// --append-system-prompt-file (CB-618 forbids the inline flag and the file flag together),
// always — fleetd #220: charter prose on the command line overruns the pane's byte cap.
List<String> charters = new java.util.ArrayList<>(2);
if (roleCharter != null) charters.add(roleCharter);
if (replyCharter != null) charters.add(replyCharter);
if (charters.size() == 1 && replyCharter != null && roleCharter == null) {
argv.add("--append-system-prompt");
argv.add(replyCharter);
} else if (!charters.isEmpty()) {
if (!charters.isEmpty()) {
argv.add("--append-system-prompt-file");
argv.add(writeCharterFile(String.join("\n\n", charters)).toString());
}
@@ -143,17 +143,24 @@ public final class EnvAllowListScrub {
}
/**
* chgrp/chmod-equivalent over the freshly generated directory and the startup files already
* written into it: owner keeps full access, {@code group} gets traverse+read on the directory
* ({@code rwxr-x---}, so a login shell under that group can find and source the files) and
* read-only on each file ({@code rw-r-----}) — deliberately no group WRITE anywhere, since a
* member never needs to add or change fleetd's own generated scrub. (The scrub script's own
* report write inside the pane consequently fails closed rather than open — see {@code
* scrub.zsh}'s trailing {@code 2>/dev/null} — which {@link
* chgrp/chmod-equivalent over a freshly generated directory and the flat files already written
* into it: owner keeps full access, {@code group} gets traverse+read on the directory ({@code
* rwxr-x---}, so a member process — a login shell reading it via {@code ZDOTDIR}, or another
* process simply opening a file under it — running under that group can find and read the
* files) and read-only on each file ({@code rw-r-----}) — deliberately no group WRITE anywhere,
* since a member never needs to add or change what fleetd generated. (For the ZDOTDIR scrub
* specifically, this also means the scrub script's own report write inside the pane fails
* closed rather than open — see {@code scrub.zsh}'s trailing {@code 2>/dev/null} — which {@link
* dev.ltms.fleet.member.HerdrPeerLauncher#releaseZdotdir} already treats as "cannot be
* confirmed to have run" rather than success.)
*
* <p>Package-private and named generically on purpose: fleetd #213 built this for the ZDOTDIR
* scrub directory, and fleetd #219 reuses it verbatim for {@link
* dev.ltms.fleet.member.OpenCodeLauncher}'s ephemeral {@code opencode.json} directory — both are
* "a fleetd-generated directory of flat files that a different-uid member process must read but
* never write," so the sharing mechanism is shared rather than copied a second time.
*/
private static void shareWithGroup(Path dir, String group) {
static void shareWithGroup(Path dir, String group) {
try {
GroupPrincipal principal = dir.getFileSystem().getUserPrincipalLookupService()
.lookupPrincipalByGroupName(group);
@@ -164,11 +171,11 @@ public final class EnvAllowListScrub {
}
}
} catch (IOException e) {
throw new UncheckedIOException("cannot share generated ZDOTDIR " + dir + " with group '"
throw new UncheckedIOException("cannot share generated directory " + dir + " with group '"
+ group + "' — the group must exist, and the fleetd operator ("
+ System.getProperty("user.name") + ") must be a member of it", e);
} catch (UnsupportedOperationException e) {
throw new UncheckedIOException("cannot share generated ZDOTDIR " + dir + " with group '"
throw new UncheckedIOException("cannot share generated directory " + dir + " with group '"
+ group + "' — this filesystem does not support POSIX group ownership",
new IOException(e));
}
@@ -682,6 +682,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
checkPaneCommandFits(cfg, argv);
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
@@ -697,6 +698,59 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw last;
}
/**
* fleetd #220: herdr does not exec the launch command — it TYPES it into the pane as one line,
* and a pty line buffer holds only {@value #PANE_COMMAND_BYTE_LIMIT} bytes (BSD/macOS {@code
* MAX_CANON}). Everything past that byte is dropped. Nothing reports it: herdr answers "agent
* started", the backend exits on the mangled argument it was handed, the pane closes, and the
* only symptom is {@link #waitUntilInjectableOrThrow} timing out 20 seconds later with no
* reason. That is exactly how #214 broke every claude-code spawn — one 50-byte flag pushed a
* 978-byte command to 1028, and the tail that got cut was {@code --autocompact 250000}.
*
* <p>So measure it here and refuse, loudly and immediately, rather than spawn something that
* cannot work. The estimate is deliberately conservative: fleetd cannot see herdr's quoting, so
* every argument is charged its own bytes plus a separator and a quote pair. An over-estimate
* costs a clear error at a length that was already unsafe; an under-estimate would let the
* silent truncation back in.
*
* @throws PeerUnreachableException when the command cannot fit — the same failure the spawn
* would have hit anyway, named at the point it is still
* explainable
*/
private void checkPaneCommandFits(FleetConfig.Profile cfg, List<String> argv) {
int bytes = 0;
String longest = null;
int longestBytes = 0;
for (String arg : argv) {
int argBytes = arg == null ? 0 : arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length;
bytes += argBytes + QUOTING_OVERHEAD_PER_ARG;
if (argBytes > longestBytes) {
longestBytes = argBytes;
longest = arg;
}
}
if (bytes <= PANE_COMMAND_BYTE_LIMIT) {
return;
}
String culprit = longest == null ? "<none>"
: longest.substring(0, Math.min(longest.length(), 60)) + (longest.length() > 60 ? "…" : "");
throw new PeerUnreachableException(
"launch command for profile " + cfg.profile() + " is about " + bytes + " bytes, over the "
+ PANE_COMMAND_BYTE_LIMIT + "-byte limit of the pane line herdr types it into. "
+ "The pty would drop the tail silently and the backend would exit on a mangled "
+ "argument. Longest argument is " + longestBytes + " bytes: " + culprit
+ " — move it off the command line (a file flag) or shorten it.");
}
/**
* The pty line buffer herdr types a launch command into: BSD/macOS {@code MAX_CANON}. Not a
* fleetd choice and not configurable — see {@link #checkPaneCommandFits}.
*/
static final int PANE_COMMAND_BYTE_LIMIT = 1024;
/** Per-argument allowance for the separating space and a shell quote pair fleetd cannot see. */
private static final int QUOTING_OVERHEAD_PER_ARG = 3;
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
@@ -860,20 +914,51 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
Object lastStatus = null;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
var status = agents.status(paneId);
lastStatus = status;
if (status.injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
// fleetd #220: read the pane BEFORE stop() closes it. Without this the gate says only that
// it timed out, which is true of every cause — a backend that never launched, a binary that
// rejected an argument and exited, a trust prompt, a login shell that hung. The pane holds
// the one copy of that answer and it is destroyed a line later.
log.warn("peer pane={} did not become injectable within {}ms (last status {}) — closing. "
+ "Pane tail:\n{}",
paneId, spawnReadyTimeoutMs, lastStatus, readPaneQuietly(paneId));
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/**
* fleetd #220: the pane's recent output, clipped, for the readiness-gate timeout log — or a
* short note when it cannot be read. Best-effort by construction: this runs on a path that is
* already failing, so it must never replace the real error with one of its own.
*/
private String readPaneQuietly(String paneId) {
try {
String pane = agents.read(paneId, "recent");
if (pane == null || pane.isBlank()) {
return "<pane read returned nothing>";
}
return pane.length() <= SPAWN_FAILURE_PANE_CHARS
? pane
: pane.substring(pane.length() - SPAWN_FAILURE_PANE_CHARS);
} catch (RuntimeException e) {
return "<pane could not be read: " + e.getMessage() + ">";
}
}
/** How much of a failed spawn's pane the timeout log carries. */
private static final int SPAWN_FAILURE_PANE_CHARS = 4000;
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* the session identity the launch resolved (CB-547a): the bridge's logical name and the peer's
@@ -1190,8 +1275,13 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* provisioning a worktree that will exist regardless, whereas an unconfigured value here means
* fleetd has no operator-endorsed location to put a credential-bearing directory a different OS
* user must reach, so falling back to the overlay is the honest answer, not a guess.
*
* <p>Package-private (fleetd #219) so {@link OpenCodeLauncher} can reuse the exact same
* "different OS user, put it under worktreeRoot instead of java.io.tmpdir" resolution for its
* own ephemeral {@code opencode.json} directory, rather than re-reading {@code config} a second
* time with a second copy of this null/blank handling.
*/
private Path memberScrubParentDir() {
Path memberScrubParentDir() {
FleetConfig cfg = config == null ? null : config.get();
if (cfg == null || cfg.worktreeRoot() == null || cfg.worktreeRoot().isBlank()) {
return null;
@@ -1204,8 +1294,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* the group {@link dev.ltms.fleet.session.Worktrees#shareWithGroup} already establishes for
* provisioned worktrees, rather than a second group key — see {@link
* #applyEnvironmentAllowListPolicy}.
*
* <p>Package-private (fleetd #219) — reused by {@link OpenCodeLauncher} alongside {@link
* #memberScrubParentDir()}; see that method's javadoc.
*/
private String memberGroup() {
String memberGroup() {
FleetConfig cfg = config == null ? null : config.get();
if (cfg == null || cfg.worktreeGroup() == null || cfg.worktreeGroup().isBlank()) {
return null;
@@ -1378,8 +1471,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* through (every production {@code HerdrPeerLauncher} does; a handful of older tests do not) —
* treated the same as "not configured", which is the correct, permissive default: it is exactly
* today's single-daemon behaviour.
*
* <p>Package-private (fleetd #219) — {@link OpenCodeLauncher} reuses this same gate to decide
* where its own ephemeral {@code opencode.json} directory (site 1) and its opencode session
* discovery (site 2) may run, rather than re-deriving "is this a multi-uid fleet" a second way.
*/
private boolean memberHerdrSocketConfigured() {
boolean memberHerdrSocketConfigured() {
if (config == null) {
return false;
}
@@ -20,6 +20,8 @@ import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.BooleanSupplier;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
@@ -215,7 +217,43 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
/**
* The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir) —
* always FLEETD's OWN {@code user.home}, whichever OS user runs the daemon.
*
* <p><b>fleetd #219 site 2 — a decision, not a patch.</b> Under {@code memberHerdrSocket:} the
* member pane runs as a <em>different</em> OS user, and opencode writes {@code opencode.db}
* under <em>that</em> user's {@code $HOME}, not fleetd's. Scanning fleetd's own {@code
* user.home} is therefore looking in the wrong place — a wrong-LOCATION failure, not a
* wrong-PERMISSION one like site 1, and it fails quietly: {@link
* SessionAwareHandle#agentSessionId()} would keep returning {@code null} forever, which reads
* as "opencode does not support resume" rather than "fleetd looked in the wrong home." fleetd
* #209 is the reason that silence is unacceptable.
*
* <p>Three ways to close the gap were weighed:
* <ol>
* <li><b>Make the member's home configurable.</b> Correct in principle, but this ticket's
* scope is the two existing call sites, not a new config key — {@code memberHerdrSocket}
* already carries the second herdr's socket path, not its user's home, and inventing a
* parallel key here without also wiring it through discovery's actual callers is a
* half-shipped feature (the exact shape CB-596/CB-611 warn against).</li>
* <li><b>Derive it</b> (e.g. from {@code worktreeRoot}'s owner, or {@code getent passwd}).
* Rejected: nothing in this codebase resolves a Unix username to a home directory today,
* and guessing wrong would silently point discovery at a THIRD wrong location — worse
* than the current gap, because it would look like it should work.</li>
* <li><b>Declare discovery unavailable</b> under {@code memberHerdrSocket}, and say so once,
* loudly, instead of scanning a directory that structurally cannot hold the answer.</li>
* </ol>
*
* <p>Option 3 is taken — the one this ticket says to default to when unsure. {@link
* OpenCodeLauncher#spawn} routes {@link SessionAwareHandle#agentSessionId()} through {@link
* HerdrPeerLauncher#memberHerdrSocketConfigured()} before ever calling {@link
* OpenCodeSessionDiscovery#sessionIdForDirectory}, so under {@code memberHerdrSocket} the
* database at this root is never even opened, and one WARN per launcher instance names the gap
* instead of the {@code null} return reading as "unsupported." Capability advertising is
* unaffected: {@link #capabilities()} always includes {@code SESSION_RESUME}, since {@code
* memberHerdrSocket} absent (today's only live mode) is unchanged by this decision.
*/
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
@@ -331,7 +369,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
*/
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
try {
Path dir = Files.createTempDirectory(configRoot, "fleetd-opencode-");
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
@@ -402,6 +440,14 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// carries operator-supplied values (URL, model id, api key), so escaping must be real.
Files.writeString(cfgFile, JSON.writerWithDefaultPrettyPrinter().writeValueAsString(root));
cfgFile.toFile().deleteOnExit();
if (memberHerdrSocketConfigured()) {
// fleetd #219: the same "different OS user" gap fleetd #213 closed for the ZDOTDIR
// scrub — share read-only with worktreeGroup rather than leaving the directory under
// fleetd's own 0700 java.io.tmpdir, where the member's OS user could not even
// traverse it. memberGroup() cannot be null here: configParentDir() above already
// refused this spawn if either worktreeRoot or worktreeGroup was missing.
EnvAllowListScrub.shareWithGroup(dir, memberGroup());
}
return cfgFile;
} catch (IOException e) {
throw new UncheckedIOException(
@@ -409,6 +455,57 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
}
/**
* fleetd #219 site 1: where {@link #writeConfig} creates its per-spawn directory.
*
* <ul>
* <li>{@code memberHerdrSocket} ABSENT (today's only mode): byte-identical to before this
* fix — always {@link #configRoot} (defaults to {@code java.io.tmpdir}, fleetd's own
* process).</li>
* <li>{@code memberHerdrSocket} PRESENT: {@code java.io.tmpdir} is fleetd's own per-user temp
* dir (mode {@code 0700} on macOS) — the member pane runs as a DIFFERENT OS user under
* this config key and cannot even traverse it, so the directory holding {@code
* opencode.json} (which tells the member where the bridge MCP is) and the member charter
* would be unreadable to the very process it is written for. The directory instead goes
* under {@code worktreeRoot}, shared read-only with {@code worktreeGroup} via {@link
* EnvAllowListScrub#shareWithGroup} — the SAME mechanism fleetd #213 built for the ZDOTDIR
* scrub, reused here rather than duplicated (see {@link
* HerdrPeerLauncher#memberScrubParentDir()}).</li>
* </ul>
*
* <p><b>Unlike the ZDOTDIR scrub, a missing {@code worktreeRoot}/{@code worktreeGroup} here
* REFUSES the spawn instead of degrading.</b> The ZDOTDIR scrub is a credential CONTROL: a
* degraded control (CB-596's sentinel overlay) is still worth having. This config file is not a
* control — it is the ONLY way the member learns where the bridge MCP lives. Writing it
* somewhere the member cannot read would not degrade anything; it would spawn a member that
* occupies a pane and never becomes deliverable, since {@code fleet_send} waits ~60s on the
* readiness gate and then fails with nothing pointing at a temp directory as the cause. Refusing
* up front, with a message that names the missing config key, is the honest failure — an
* undeliverable member is not a working spawn either way, so nothing is lost by refusing loudly
* instead of failing silently later.
*
* @throws IllegalStateException when {@code memberHerdrSocket} is configured but {@code
* worktreeRoot} and/or {@code worktreeGroup} is not
*/
private Path configParentDir() {
if (!memberHerdrSocketConfigured()) {
return configRoot;
}
Path root = memberScrubParentDir();
String group = memberGroup();
if (root == null || group == null) {
throw new IllegalStateException("memberHerdrSocket is configured, so opencode's config "
+ "directory (opencode.json + member charter) must be placed where the member's "
+ "OS user can read it — worktreeRoot, shared via worktreeGroup — but "
+ (root == null ? "worktreeRoot" : "worktreeGroup") + " is not configured. "
+ "Refusing to spawn rather than write a config the member cannot read: that "
+ "member would occupy a pane and never become deliverable, with nothing "
+ "pointing at the real cause. Configure both worktreeRoot and worktreeGroup to "
+ "enable opencode member spawns under memberHerdrSocket.");
}
return root;
}
/**
* Declare a custom OpenAI-compatible provider so the worker talks to a pinned endpoint (a local
* vLLM, say) instead of opencode's default gateway (CB-508).
@@ -517,11 +614,16 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/** One WARN per launcher instance for the fleetd #219 site-2 discovery-unavailable gap. */
private final AtomicBoolean discoveryUnavailableWarned =
new AtomicBoolean();
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
return new SessionAwareHandle(inner, discovery, effectiveCwd(req),
this::memberHerdrSocketConfigured, discoveryUnavailableWarned);
}
/**
@@ -536,11 +638,17 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
private final BooleanSupplier discoveryUnavailable;
private final AtomicBoolean discoveryUnavailableWarned;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd,
BooleanSupplier discoveryUnavailable,
AtomicBoolean discoveryUnavailableWarned) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
this.discoveryUnavailable = discoveryUnavailable;
this.discoveryUnavailableWarned = discoveryUnavailableWarned;
}
@Override
@@ -565,6 +673,22 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public String agentSessionId() {
// fleetd #219 site 2: under memberHerdrSocket the member pane runs as a different OS
// user, so opencode.db lives under THAT user's $HOME, not the one discoveryRoot was
// built from (see OpenCodeLauncher#defaultDiscoveryRoot's javadoc for the full
// reasoning). Scanning fleetd's own $HOME under that config would only ever find "no
// row" and read as "resume unsupported" — declare it unavailable instead, once, loudly.
if (discoveryUnavailable.getAsBoolean()) {
if (discoveryUnavailableWarned.compareAndSet(false, true)) {
log.warn("opencode session discovery unavailable: memberHerdrSocket is "
+ "configured, so opencode's on-disk session database lives under the "
+ "MEMBER's own $HOME, not fleetd's ({}) — agentSessionId will stay null "
+ "for every opencode member under this config, and SESSION_RESUME "
+ "cannot be honored (fleetd #209/#219).",
System.getProperty("user.home"));
}
return null;
}
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
@@ -633,6 +633,115 @@ class CompletionResolverTest {
"the failure carries the rest of the pane, not only the matched line: " + reason);
}
// --- fleetd#211: raw-scrape fallback classification when there is no usable assistant block ---
@Test
void anExhaustionLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrapeAndNotifiesTheSink() {
// No ⏺ anywhere, and the first visible line is TUI chrome (╭). lastAssistantBlock's boundary
// scan starts at the top of the raw screen and breaks immediately, so the trimmed block is "".
// The fix: fall back to matching the RAW scrape so this doesn't get lost as an empty scrape.
String block = """
╭──────────────────────────────────────╮
The usage limit has been reached. Try again later.
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape match still resolves the blocked send");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"classified from the raw scrape even though the trimmed block was empty");
assertEquals(1, notified.size(),
"the sink is the whole point of this ticket — it must be notified: " + notified);
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
assertTrue(notified.get(0).contains("The usage limit has been reached"),
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aBackendErrorLineWithNoMarkerAndLeadingChromeIsClassifiedFromTheRawScrape() {
String block = """
╭──────────────────────────────────────╮
API Error: 400 invalid request body
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a raw-scrape backend-error match still resolves the blocked send");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"classified BACKEND_ERROR from the raw scrape even though the trimmed block was empty");
assertTrue(waiter.getNow(null).text().contains("API Error: 400 invalid request body"),
"the failure carries the matched line: " + waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("--- pane tail ---"),
"fleetd#164: the failure must carry the pane, not only the matched line — with an "
+ "empty trimmed block the raw scrape is the only copy of what the member said: "
+ waiter.getNow(null).text());
assertTrue(waiter.getNow(null).text().contains("╭"),
"the carried pane is the raw scrape, chrome included: " + waiter.getNow(null).text());
}
@Test
void anOrdinaryPaneWithANormalAssistantBlockIsUnaffectedByTheRawScrapeFallback() {
// Pin: on a pane that already yields a usable block, the fallback branch is never reached —
// same outcome, same text, sink not called — even though the raw screen around the marker
// would itself match the configured exhausted pattern.
String block = "The usage limit has been reached, but this is a leading TUI line above the "
+ "marker.\n⏺ complete report\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone());
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"unchanged: a usable assistant block never reaches the raw-scrape fallback");
assertEquals("complete report", waiter.getNow(null).text(), "the reply text is unaffected");
assertTrue(notified.isEmpty(), "the fallback never runs, so the sink is never called");
}
@Test
void aGenuinelyEmptyScrapeStillFailsAsEmptyAndNeverNotifiesTheSink() {
// The false-positive pin: no exhaustion or backend-error text anywhere on the pane (just
// chrome, no marker) — the raw-scrape fallback must not manufacture a classification, and
// the sink must stay untouched.
String block = """
╭──────────────────────────────────────╮
│ > │
╰──────────────────────────────────────╯
""";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "an empty scrape must still resolve the send, not hang");
assertEquals(Rendezvous.Kind.FAILED, waiter.getNow(null).kind(),
"no exhaustion or backend-error text anywhere ⇒ this stays the ordinary empty-scrape failure");
assertTrue(waiter.getNow(null).text().toLowerCase().contains("empty"),
"the failure still says the scrape was empty: " + waiter.getNow(null).text());
assertTrue(notified.isEmpty(), "a genuinely empty pane must never quarantine a credential");
}
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
@@ -61,8 +61,75 @@ class ClaudeCodeLauncherTest {
assertTrue(args.stream().noneMatch(a -> a.contains("\"bridge\"")),
"the mount is named fleet since CB-632 — a member addresses its tools as "
+ "mcp__fleet__*, and CLAUDE.md's role-detection ladder names that prefix");
assertTrue(args.contains("--append-system-prompt"));
assertTrue(args.stream().anyMatch(a -> a.contains("fleet_reply")), "reply charter present");
// fleetd #220: the charter travels as a FILE, never inline — charter prose on the command
// line overruns the byte cap of the pane line herdr types it into.
assertTrue(args.contains("--append-system-prompt-file"));
assertFalse(args.contains("--append-system-prompt"),
"the inline flag would put ~800 bytes of prose on the pane command line");
String charterFile = args.get(args.indexOf("--append-system-prompt-file") + 1);
assertTrue(readFile(charterFile).contains("fleet_reply"), "reply charter present in the file");
}
/** Read a charter file the launcher wrote, failing the test rather than the build on an IO error. */
private static String readFile(String path) {
try {
return java.nio.file.Files.readString(java.nio.file.Path.of(path));
} catch (java.io.IOException e) {
throw new AssertionError("charter file " + path + " is not readable", e);
}
}
/**
* fleetd #220 regression. herdr TYPES the launch command into the pane, and the pty line buffer
* holds 1024 bytes — past that the tail is dropped with no error anywhere, so the backend exits
* on a mangled argument and the spawn dies as an unexplained readiness timeout. That is what
* happened when #214 added --session-id to a command already 978 bytes long: --autocompact
* 250000 arrived as --autocompact 25. This asserts the whole assembled command still fits, with
* the flags a real spawn carries (MCP mount, charter, model, autocompact, session id).
*/
@Test
void theAssembledLaunchCommandFitsThePaneLineLimit() {
FakeHerdr herdr = new FakeHerdr();
// Shaped like the live sonnet profile, because the bug is a SUM: the charter alone fits,
// and so does every flag alone. Only model + autocompact + session id on top of the charter
// crossed the cap, which is why nothing caught it until a member failed to spawn.
FleetConfig.Profile cfg = new FleetConfig.Profile(
"sonnet", null, "claude-sonnet-5", null, null,
List.of("claude"), "tab", "fleetd-workers", "w #{n}", "http://127.0.0.1:8765/mcp",
null, null, null, null, null, null, null, null, true, null, null, null, null, null,
250000);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of()), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null)
.spawn();
List<String> args = spawnedArgs(herdr);
int bytes = "claude".length();
for (String arg : args) {
bytes += arg.getBytes(java.nio.charset.StandardCharsets.UTF_8).length + 3;
}
assertTrue(bytes <= HerdrPeerLauncher.PANE_COMMAND_BYTE_LIMIT,
"the launch command must fit the pane line: " + bytes + " bytes vs limit "
+ HerdrPeerLauncher.PANE_COMMAND_BYTE_LIMIT + " — args " + args);
}
/**
* fleetd #220: the guard refuses a command that cannot fit, instead of letting the pty drop the
* tail. The refusal must name the size and the argument to blame — a spawn that fails with
* "did not reach injectable state" tells the operator nothing, which is the whole reason this
* bug took a live pane scrape to find.
*/
@Test
void anOverlongLaunchCommandIsRefusedWithTheSizeAndTheCulprit() {
FakeHerdr herdr = new FakeHerdr();
String huge = "x".repeat(1500);
ClaudeCodeLauncher launcher = service(herdr, List.of("claude", huge), null);
PeerUnreachableException refused = assertThrows(PeerUnreachableException.class, launcher::spawn);
assertTrue(refused.getMessage().contains("1024"), "names the limit: " + refused.getMessage());
assertTrue(refused.getMessage().contains("1500"), "names the culprit's size: " + refused.getMessage());
assertFalse(herdr.called("agent.start"),
"nothing may be started — a truncated command is worse than no spawn");
}
// CB-634: a profile with ideMcpUrl set mounts the IDE Index MCP as a second server and pins
@@ -262,10 +329,16 @@ class ClaudeCodeLauncherTest {
Map<?, ?> start = (Map<?, ?>) herdr.lastCall("agent.start").params();
assertEquals("claude", start.get("kind"), "herdr launches the canonical executable by kind");
// CB-533: the shared fixture pins model "coder", so the model flag is the whole args list.
// What this test guards is that argv[0] is NOT repeated — herdr supplies it from `kind`.
assertEquals(List.of("--model", "coder"), start.get("args"),
"the configured executable is not repeated in args");
// fleetd #214: a plain spawn now also carries fleetd's own minted session id.
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("claude"), "the configured executable is not repeated in args");
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0, "a plain spawn mints a session id: " + args);
assertDoesNotThrow(() -> UUID.fromString(args.get(flag + 1)), "the minted id is a valid UUID");
assertEquals(List.of("--model", "coder"),
List.of(args.get(args.size() - 2), args.get(args.size() - 1)),
"the CB-533 model flag still trails the launch flags: " + args);
}
@Test
@@ -278,7 +351,15 @@ class ClaudeCodeLauncherTest {
assertFalse(args.contains("--append-system-prompt"), "no reply charter without mcpUrl");
// CB-533: the model flag is independent of the MCP mount — pinning the model is not part of
// "mount the bridge", so an unmounted worker still runs the model its profile names.
assertEquals(List.of("--verbose", "--model", "coder"), args,
// fleetd #214: a plain spawn now carries fleetd's own minted session id; drop the two
// mint elements when checking the rest of the argv.
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0, "a plain spawn mints a session id even without a bridge mount: " + args);
assertDoesNotThrow(() -> UUID.fromString(args.get(flag + 1)), "the minted id is a valid UUID");
List<String> rest = new java.util.ArrayList<>(args);
rest.remove(flag + 1);
rest.remove(flag);
assertEquals(List.of("--verbose", "--model", "coder"), rest,
"the operator's own args are preserved, in order, ahead of the model flag");
}
@@ -649,7 +730,7 @@ class ClaudeCodeLauncherTest {
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
// --- CB-547a: durable session identity (mint / resume / no-identity legacy) -----------------
// --- CB-547a / fleetd #214: durable session identity (always mint / resume) -----------------
@Test
void freshSpawnMintsASessionIdAndPassesTheName() {
@@ -684,18 +765,28 @@ class ClaudeCodeLauncherTest {
}
@Test
void noIdentitySpawnKeepsTheLegacyArgvAndCarriesNoSessionHandle() {
void plainSpawnMintsASessionIdSoEveryMemberIsResumable() {
// fleetd #214: a plain spawn passes no sessionName and no resumeSessionId, yet the member
// must still be resumable — the id is minted unconditionally, and it is the ONLY resume
// handle a claude-code member has (unlike opencode, nothing resolves it after the launch).
// The binary requires a valid UUID (checked against claude 2.1.252: a non-UUID is refused
// at argument parsing with "Invalid session ID. Must be a valid UUID.").
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "no identity → no --session-id");
assertFalse(args.contains("-n"), "no identity → no -n");
assertFalse(args.contains("-r"), "no identity → no -r");
assertNull(handle.agentSessionId(), "no identity → no resume handle");
assertNull(handle.sessionName(), "no identity → no logical name");
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0 && flag + 1 < args.size(),
"--session-id is minted even when no identity is requested: " + args);
String minted = args.get(flag + 1);
assertDoesNotThrow(() -> UUID.fromString(minted), "--session-id is a valid UUID: " + minted);
assertEquals(minted, handle.agentSessionId(),
"the resume handle is the minted id, so every member is resumable from fleet_list");
assertFalse(args.contains("-n"), "no sessionName was requested → no -n");
assertFalse(args.contains("-r"), "no resume was requested → no -r");
assertNull(handle.sessionName(), "no sessionName was requested → no logical name");
}
@Test
@@ -1,5 +1,9 @@
package dev.ltms.fleet.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.config.FleetConfig;
@@ -14,9 +18,13 @@ import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.PosixFileAttributeView;
import java.nio.file.attribute.PosixFilePermissions;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ExecutorService;
@@ -25,6 +33,7 @@ import java.util.concurrent.Future;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* The opencode adapter's launch build: a file-based MCP mount + reply-charter instructions (no
@@ -610,4 +619,196 @@ class OpenCodeLauncherTest {
assertTrue(json.path("mcp").path("intellij").isMissingNode(),
"no IDE server when ideMcpUrl is unset");
}
// --- fleetd #219: config root + discovery root under memberHerdrSocket ------------------------
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
private static FleetConfig configWithMemberHerdrSocket(String worktreeRoot, String worktreeGroup) {
return new FleetConfig(
null, // bind
null, // herdrSocket
"/tmp/other-user.sock", // memberHerdrSocket
Map.of(), // profiles
null, // guard
worktreeRoot, // worktreeRoot
null, // lifecycle
null, // spawnReadyTimeoutMs
null, // spawnReadyPollMs
null, // broker
null, // primary
null, // fleet
null, // leadHeartbeat
null, // health
null, // placement
null, // auth
null, // configReload
null, // quarantineCooldownSeconds
null, // memberCredentials
null, // coordinator
worktreeGroup, // worktreeGroup
null // memberLoginShell
).withDefaults();
}
private static OpenCodeLauncher serviceWithConfig(FakeHerdr herdr, Path configRoot, Path discoveryRoot,
FleetConfig.Profile cfg, Supplier<FleetConfig> config) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, configRoot, discoveryRoot, null, null, config);
}
/** The current process's own primary group — resolvable on whatever host runs this test. */
private static String currentUserGroup() throws IOException {
PosixFileAttributeView view = Files.getFileAttributeView(Path.of("."), PosixFileAttributeView.class);
assumeTrue(view != null, "this host's filesystem does not support POSIX group ownership");
return view.readAttributes().group().getName();
}
/**
* fleetd #219 site 1, acceptance criterion 1: with {@code memberHerdrSocket} configured and both
* {@code worktreeRoot}/{@code worktreeGroup} set, the generated {@code opencode.json} directory
* lives under {@code worktreeRoot} — NEVER under the injected {@code configRoot} (standing in for
* {@code java.io.tmpdir}, fleetd's own 0700 temp dir, unreadable by the member's different OS
* user) — and is shared read-only with the group via the SAME mechanism (fleetd #213's {@link
* EnvAllowListScrub#shareWithGroup}) the ZDOTDIR scrub uses.
*/
@Test
void memberHerdrSocketWithWorktreeRootAndGroupPutsConfigDirUnderWorktreeRootAndSharesIt(
@TempDir Path configRoot, @TempDir Path worktreeRoot) throws Exception {
String group = currentUserGroup();
FakeHerdr herdr = new FakeHerdr();
serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), group)).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "the profile still needs a config file");
Path cfgFile = Path.of(cfgPath);
Path dir = cfgFile.getParent();
assertEquals(worktreeRoot.toAbsolutePath().normalize(), dir.getParent(),
"the generated directory's parent must be worktreeRoot, not the injected configRoot "
+ "standing in for java.io.tmpdir — got parent " + dir.getParent());
assertFalse(dir.startsWith(configRoot),
"the generated directory must NOT be created under configRoot when memberHerdrSocket "
+ "is configured: " + dir);
assertEquals("rwxr-x---", PosixFilePermissions.toString(Files.getPosixFilePermissions(dir)),
"the directory must be group-traversable+readable, owner-only writable");
assertEquals("rw-r-----", PosixFilePermissions.toString(Files.getPosixFilePermissions(cfgFile)),
"opencode.json must be group-readable, never group-writable");
}
/**
* fleetd #219 site 1, acceptance criterion 2: with {@code memberHerdrSocket} configured but
* NEITHER {@code worktreeRoot} nor {@code worktreeGroup} set, the launcher must refuse the spawn
* rather than write a config under {@code java.io.tmpdir} the member cannot read — that member
* would occupy a pane and never become deliverable, with nothing pointing at the real cause.
*/
@Test
void memberHerdrSocketWithoutWorktreeRootOrGroupRefusesTheSpawn(@TempDir Path configRoot) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(null, null));
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> launcher.spawn(new SpawnRequest(null, null, null)),
"a missing worktreeRoot/worktreeGroup must refuse the spawn, not write an unreadable config");
assertTrue(ex.getMessage().contains("worktreeRoot"),
"the refusal must name the missing config key — got: " + ex.getMessage());
assertFalse(herdr.called("tab.create"),
"the spawn must be refused BEFORE any pane is created — got calls: " + herdr.calls);
}
/**
* fleetd #219 site 1, acceptance criterion 2 (the other missing half): {@code worktreeRoot} set
* but {@code worktreeGroup} missing must ALSO refuse — either one alone is not enough to
* guarantee the member's OS user can read the directory.
*/
@Test
void memberHerdrSocketWithWorktreeRootButNoGroupRefusesTheSpawn(
@TempDir Path configRoot, @TempDir Path worktreeRoot) {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> configWithMemberHerdrSocket(worktreeRoot.toString(), null));
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> launcher.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("worktreeGroup"),
"worktreeRoot alone is not enough — got: " + ex.getMessage());
}
/**
* fleetd #219 site 1, acceptance criterion 3: with {@code memberHerdrSocket} ABSENT — even when
* a live, non-null {@code config} supplier is threaded through (not merely {@code config == null},
* which every other test in this file already exercises) — the generated directory must still
* land directly under the injected {@code configRoot}, byte-identical to before this fix.
*/
@Test
void memberHerdrSocketAbsentStaysUnderConfigRootEvenWithALiveConfigSupplier(@TempDir Path configRoot)
throws Exception {
FakeHerdr herdr = new FakeHerdr();
FleetConfig config = new FleetConfig(null, null, null, Map.of(), null, null, null, null, null,
null, null, null, null, null, null, null, null, null, null, null, null, null).withDefaults();
serviceWithConfig(herdr, configRoot, configRoot,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null), () -> config)
.spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath);
assertTrue(Path.of(cfgPath).startsWith(configRoot),
"with memberHerdrSocket absent, the config directory must still be created directly "
+ "under configRoot, unchanged from before this fix");
}
/**
* fleetd #219 site 2: under {@code memberHerdrSocket}, opencode session discovery must be
* declared unavailable rather than silently scanning fleetd's own {@code discoveryRoot} — which,
* under this config key, is NOT where the member's opencode actually writes its session
* database. This test proves the gate is real, not merely "no record yet": a matching record IS
* written to {@code discoveryRoot} (the exact fixture {@link
* #theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears} proves discovery would
* otherwise find), and {@code agentSessionId()} must still return {@code null} — proving the
* gate, not a coincidental absence of data, is what produced the null. One WARN is also logged,
* exactly once even across repeated calls.
*/
@Test
void discoveryIsUnavailableUnderMemberHerdrSocketEvenWhenARecordExists(
@TempDir Path configRoot, @TempDir Path worktreeRoot, @TempDir Path discRoot) throws Exception {
String group = currentUserGroup();
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "ses_should_be_hidden", "/work/dir", 1000L);
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, configRoot, discRoot,
null, null, () -> configWithMemberHerdrSocket(worktreeRoot.toString(), group));
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
Level original = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
PeerHandle handle;
try {
handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
assertNull(handle.agentSessionId(),
"memberHerdrSocket configured: discovery must stay unavailable even though a "
+ "matching record exists in discoveryRoot");
assertNull(handle.agentSessionId(), "the gate must hold on a second call too");
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
}
List<String> warnings = appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.toList();
assertEquals(1, warnings.size(),
"exactly one WARN across two agentSessionId() calls — got: " + warnings);
assertTrue(warnings.get(0).contains("memberHerdrSocket"),
"the WARN must name memberHerdrSocket as the reason — got: " + warnings.get(0));
}
}
@@ -887,15 +887,18 @@ class SessionManagerTest {
}
@Test
void acquireWithNeitherSessionFieldLeavesAgentSessionIdNull() {
void acquireWithNeitherSessionFieldStillMintsAnAgentSessionId() {
// fleetd #214: the claude-code launcher mints a session id for EVERY spawn, so a member is
// resumable even when the spawn asked for no session identity.
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, null, null);
assertNull(s.agentSessionId(), "no identity requested — unchanged from before CB-584");
assertFalse(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"a null id is omitted from the roster, like charterSha256 for a receipt-less session");
assertNotNull(s.agentSessionId(),
"fleetd #214: a plain spawn mints a session id, so every member is resumable");
assertTrue(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"the minted id is in the roster, so fleet_list advertises every member's resume handle");
}
@Test